PubMed Health⌕ Search

PubMed · 11514040

Uniform power method for sample size calculation in historical control studies with binary response.

Abstract

Makuch and Simon gave a sample size calculation formula for historical control (HC) studies that assumed that the observed response rate in the control group is the true response rate. We dropped this assumption and computed the expected power and expected sample size to evaluate the performance of the procedure under the omniscient model. When there is uncertainty in the HC response rate but this uncertainty is not considered, Makuch and Simon's method produces a sample size that gives a considerably lower power than that specified. Even the larger sample size obtained from the randomized design formula and applied to the HC setting does not guarantee the advertised power in the HC setting. We developed a new uniform power method to search for the sample size required for the experimental group to yield an exact power without relying on the estimated HC response rate being perfectly correct. The new method produces the correct uniform predictive power for all permissible response rates. The resulting sample size is closer to the sample size needed for the randomized design than Makuch and Simon's method, especially when there is a small difference in response rates or a limited sample size in the HC group. HC design may be a viable option in clinical trials when the patient selection bias and the outcome evaluation bias can be minimized. However, the common perception of the extra sample size savings is largely unjustified without the strong assumption that the observed HC response rate is equal to the true control response rate. Generally speaking, results from HC studies need to be confirmed by studies with concurrent controls and cannot be used for making definitive decisions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J J Lee, C Tseng. 2001. Uniform power method for sample size calculation in historical control studies with binary response.. https://doi.org/10.1016/s0197-2456(01)00143-x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Assessment of blinding in pharmacotherapy and noninvasive neuromodulation randomized controlled trials for neuropathic pain in adults.

In randomized controlled trials (RCTs), study participants and research personnel are often blinded to minimize biases related to knowing treatment allocation. To determine if blinding was effective, participants may be asked which treatment they believe they received ("treatment guess"). This descriptive review characterized blinding assessment (BA) reporting in pharmacotherapy and neuromodulation neuropathic pain RCTs. Of 288 papers, 36 (12.5%) reported a BA. One paper reported the results of 2 studies, so in total 37 studies with a BA were assessed. Of these, 19 were crossover, 17 parallel, and 1 partial crossover in design. All 37 studies assessed participant blinding, and 10 also assessed investigator blinding. Approximately 27% included an "unsure" answer option for treatment guess, and 38% asked the reason for the guess. There were no clear patterns in BA reporting across time nor based on treatment type. Seventeen trials provided sufficient data to calculate Bang Blinding Index (BI) to determine blinding success. Participants remained blinded (BI = 0 &#xb1; 0.2) in 10/17 placebo and 10/17 treatment arms, 6 placebo and 5 treatment arms had a BI > 0.2 suggesting possible unblinding, whereas 1 placebo and 2 treatment arms had a BI < -0.2 suggesting misinformed guessing. Overall, we found that BAs are done in a minority of published neuropathic pain trials and with variable methodology. Given the importance of minimizing risk of bias because of treatment unblinding, future studies should consider including BAs, and further consensus building is necessary to determine if and how BAs should be conducted and interpreted in analgesic clinical trials.

Bias↗

Sources of variation and bias in studies of diagnostic accuracy: a systematic review.

BACKGROUND: Studies of diagnostic accuracy are subject to different sources of bias and variation than studies that evaluate the effectiveness of an intervention. Little is known about the effects of these sources of bias and variation. PURPOSE: To summarize the evidence on factors that can lead to bias or variation in the results of diagnostic accuracy studies. DATA SOURCES: MEDLINE, EMBASE, and BIOSIS, and the methodologic databases of the Centre for Reviews and Dissemination and the Cochrane Collaboration. Methodologic experts in diagnostic tests were contacted. STUDY SELECTION: Studies that investigated the effects of bias and variation on measures of test performance were eligible for inclusion, which was assessed by one reviewer and checked by a second reviewer. Discrepancies were resolved through discussion. DATA EXTRACTION: Data extraction was conducted by one reviewer and checked by a second reviewer. DATA SYNTHESIS: The best-documented effects of bias and variation were found for demographic features, disease prevalence and severity, partial verification bias, clinical review bias, and observer and instrument variation. For other sources, such as distorted selection of participants, absent or inappropriate reference standard, differential verification bias, and review bias, the amount of evidence was limited. Evidence was lacking for other features, including incorporation bias, treatment paradox, arbitrary choice of threshold value, and dropouts. CONCLUSIONS: Many issues in the design and conduct of diagnostic accuracy studies can lead to bias or variation; however, the empirical evidence about the size and effect of these issues is limited.

Bias↗

Test bias in a cognitive test: differential item functioning in the CASI.

Assessment of test bias is important to establish the construct validity of tests. Assessment of differential item functioning (DIF) is an important first step in this process. DIF is present when examinees from different groups have differing probabilities of success on an item, after controlling for overall ability level. Here, we present analysis of DIF in the Cognitive Assessment Screening Instrument (CASI) using data from a large cohort study of elderly adults. We developed an ordinal logistic regression modelling technique to assess test items for DIF. Estimates of cognitive ability were obtained in two ways based on responses to CASI items: using traditional CASI scoring according to the original test instructions as well as using item response theory (IRT) scoring. Several demographic characteristics were examined for potential DIF, including ethnicity and gender (entered into the model as dichotomous variables), and years of education and age (entered as continuous variables). We found that a disappointingly large number of items had DIF with respect to at least one of these demographic variables. More items were found to have DIF with traditional CASI scoring than with IRT scoring. This study demonstrates a powerful technique for the evaluation of DIF in psychometric tests. The finding that so many CASI items had DIF suggests that previous findings of differences between groups in cognitive functioning as measured by the CASI may be due to biased test items rather than true differences between groups. The finding that IRT scoring diminished the impact of DIF is discussed. Some preliminary suggestions for how to deal with items found to have DIF in cognitive tests are made. The advantages of the DIF detection techniques we developed are discussed in relation to other techniques for the evaluation of DIF.

Bias↗