Peer review of statistics in medical research. Reporting power calculations is important.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Kenneth F Schulz.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Biases in systematic reviews and meta-analyses may be examined in 'meta-epidemiological' studies, in which the influence of trial characteristics such as measures of study quality on treatment effect estimates is explored. Published studies to date have analysed data from collections of meta-analyses with binary outcomes, using logistic regression models that assume that there is no between- or within-meta-analysis heterogeneity. Using data from a study of publication bias (39 meta-analyses, 394 published and 88 unpublished trials) and language bias (29 meta-analyses, 297 English language trials and 52 non-English language trials), we compare results from logistic regression models, with and without robust standard errors to allow for clustering on meta-analysis, with results using a 'meta-meta-analytic' approach that can allow for between- and within-meta-analysis heterogeneity. We also consider how to allow for the confounding effects of different trial characteristics. We show that both within- and between meta-analysis heterogeneity may be of importance in the analysis of meta-epidemiological studies, and that confounding exists between the effects of publication status and trial quality.
Explore the source record for details and available documents.
We cringe at the pervasive notion that a randomised trial needs to yield equal sample sizes in the comparison groups. Unfortunately, that conceptual misunderstanding can lead to bias by investigators who force equality, especially if by non-scientific means. In simple, unrestricted, randomised trials (analogous to repeated coin-tossing), the sizes of groups should indicate random variation. In other words, some discrepancy between the numbers in the comparison groups would be expected. The appeal of equal group sizes in a simple randomised controlled trial is cosmetic, not scientific. Moreover, other randomisation schemes, termed restricted randomisation, force equality by departing from simple randomisation. Forcing equal group sizes, however, potentially harms the unpredictability of treatment assignments, especially when using permuted-block randomisation in non-double-blinded trials. Diminished unpredictability can allow bias to creep into a trial. Overall, investigators underuse simple randomisation and overuse fixed-block randomisation. For non-double-blinded trials larger than 200 participants, investigators should use simple randomisation more often and accept moderate disparities in group sizes. Such unpredictability reflects the essence of randomness. We endorse the generation of mildly unequal group sizes and encourage an appreciation of such inequalities. For non-double-blinded randomised controlled trials with a sample size of less than 200 overall or within any principal stratum or subgroup, urn randomisation enhances unpredictability compared with blocking. A simpler alternative, our mixed randomisation approach, attains unpredictability within the context of the currently understood simple randomisation and permuted-block methods. Simple randomisation contributes the unpredictability whereas permuted-block randomisation contributes the balance, but avoids the perfect balance that can result in selection bias.
Screening tests are ubiquitous in contemporary practice, yet the principles of screening are widely misunderstood. Screening is the testing of apparently well people to find those at increased risk of having a disease or disorder. Although an earlier diagnosis generally has intuitive appeal, earlier might not always be better, or worth the cost. Four terms describe the validity of a screening test: sensitivity, specificity, and predictive value of positive and negative results. For tests with continuous variables--eg, blood glucose--sensitivity and specificity are inversely related; where the cutoff for abnormal is placed should indicate the clinical effect of wrong results. The prevalence of disease in a population affects screening test performance: in low-prevalence settings, even very good tests have poor predictive value positives. Hence, knowledge of the approximate prevalence of disease is a prerequisite to interpreting screening test results. Tests are often done in sequence, as is true for syphilis and HIV-1 infection. Lead-time and length biases distort the apparent value of screening programmes; randomised controlled trials are the only way to avoid these biases. Screening can improve health; strong indirect evidence links cervical cytology programmes to declines in cervical cancer mortality. However, inappropriate application or interpretation of screening tests can rob people of their perceived health, initiate harmful diagnostic testing, and squander health-care resources.
Proper randomisation means little if investigators cannot include all randomised participants in the primary analysis. Participants might ignore follow-up, leave town, or take aspartame when instructed to take aspirin. Exclusions before randomisation do not bias the treatment comparison, but they can hurt generalisability. Eligibility criteria for a trial should be clear, specific, and applied before randomisation. Readers should assess whether any of the criteria make the trial sample atypical or unrepresentative of the people in which they are interested. In principle, assessment of exclusions after randomisation is simple: none are allowed. For the primary analysis, all participants enrolled should be included and analysed as part of the original group assigned (an intent-to-treat analysis). In reality, however, losses frequently occur. Investigators should, therefore, commit adequate resources to develop and implement procedures to maximise retention of participants. Moreover, researchers should provide clear, explicit information on the progress of all randomised participants through the trial by use of, for instance, a trial profile. Investigators can also do secondary analyses on, for instance, per-protocol or as-treated participants. Such analyses should be described as secondary and non-randomised comparisons. Mishandling of exclusions causes serious methodological difficulties. Unfortunately, some explanations for mishandling exclusions intuitively appeal to readers, disguising the seriousness of the issues. Creative mismanagement of exclusions can undermine trial validity.
Blinding embodies a rich history spanning over two centuries. Most researchers worldwide understand blinding terminology, but confusion lurks beyond a general comprehension. Terms such as single blind, double blind, and triple blind mean different things to different people. Moreover, many medical researchers confuse blinding with allocation concealment. Such confusion indicates misunderstandings of both. The term blinding refers to keeping trial participants, investigators (usually health-care providers), or assessors (those collecting outcome data) unaware of the assigned intervention, so that they will not be influenced by that knowledge. Blinding usually reduces differential assessment of outcomes (information bias), but can also improve compliance and retention of trial participants while reducing biased supplemental care or treatment (sometimes called co-intervention). Many investigators and readers naïvely consider a randomised trial as high quality simply because it is double blind, as if double-blinding is the sine qua non of a randomised controlled trial. Although double blinding (blinding investigators, participants, and outcome assessors) indicates a strong design, trials that are not double blinded should not automatically be deemed inferior. Rather than solely relying on terminology like double blinding, researchers should explicitly state who was blinded, and how. We recommend placing greater credence in results when investigators at least blind outcome assessments, except with objective outcomes, such as death, which leave little room for bias. If investigators properly report their blinding efforts, readers can judge them. Unfortunately, many articles do not contain proper reporting. If an article claims blinding without any accompanying clarification, readers should remain sceptical about its effect on bias reduction.
Proper randomisation rests on adequate allocation concealment. An allocation concealment process keeps clinicians and participants unaware of upcoming assignments. Without it, even properly developed random allocation sequences can be subverted. Within this concealment process, the crucial unbiased nature of randomised controlled trials collides with their most vexing implementation problems. Proper allocation concealment frequently frustrates clinical inclinations, which annoys those who do the trials. Randomised controlled trials are anathema to clinicians. Many involved with trials will be tempted to decipher assignments, which subverts randomisation. For some implementing a trial, deciphering the allocation scheme might frequently become too great an intellectual challenge to resist. Whether their motives indicate innocent or pernicious intents, such tampering undermines the validity of a trial. Indeed, inadequate allocation concealment leads to exaggerated estimates of treatment effect, on average, but with scope for bias in either direction. Trial investigators will be crafty in any potential efforts to decipher the allocation sequence, so trial designers must be just as clever in their design efforts to prevent deciphering. Investigators must effectively immunise trials against selection and confounding biases with proper allocation concealment. Furthermore, investigators should report baseline comparisons on important prognostic variables. Hypothesis tests of baseline characteristics, however, are superfluous and could be harmful if they lead investigators to suppress reporting any baseline imbalances.
The randomised controlled trial sets the gold standard of clinical research. However, randomisation persists as perhaps the least-understood aspect of a trial. Moreover, anything short of proper randomisation courts selection and confounding biases. Researchers should spurn all systematic, non-random methods of allocation. Trial participants should be assigned to comparison groups based on a random process. Simple (unrestricted) randomisation, analogous to repeated fair coin-tossing, is the most basic of sequence generation approaches. Furthermore, no other approach, irrespective of its complexity and sophistication, surpasses simple randomisation for prevention of bias. Investigators should, therefore, use this method more often than they do, and readers should expect and accept disparities in group sizes. Several other complicated restricted randomisation procedures limit the likelihood of undesirable sample size imbalances in the intervention groups. The most frequently used restricted sequence generation procedure is blocked randomisation. If this method is used, investigators should randomly vary the block sizes and use larger block sizes, particularly in an unblinded trial. Other restricted procedures, such as urn randomisation, combine beneficial attributes of simple and restricted randomisation by preserving most of the unpredictability while achieving some balance. The effectiveness of stratified randomisation depends on use of a restricted randomisation approach to balance the allocation sequences for each stratum. Generation of a proper randomisation sequence takes little time and effort but affords big rewards in scientific accuracy and credibility. Investigators should devote appropriate resources to the generation of properly randomised trials and reporting their methods clearly.
Explore the source record for details and available documents.
Epidemiologists benefit greatly from having case-control study designs in their research armamentarium. Case-control studies can yield important scientific findings with relatively little time, money, and effort compared with other study designs. This seemingly quick road to research results entices many newly trained epidemiologists. Indeed, investigators implement case-control studies more frequently than any other analytical epidemiological study. Unfortunately, case-control designs also tend to be more susceptible to biases than other comparative studies. Although easier to do, they are also easier to do wrong. Five main notions guide investigators who do, or readers who assess, case-control studies. First, investigators must explicitly define the criteria for diagnosis of a case and any eligibility criteria used for selection. Second, controls should come from the same population as the cases, and their selection should be independent of the exposures of interest. Third, investigators should blind the data gatherers to the case or control status of participants or, if impossible, at least blind them to the main hypothesis of the study. Fourth, data gatherers need to be thoroughly trained to elicit exposure in a similar manner from cases and controls; they should use memory aids to facilitate and balance recall between cases and controls. Finally, investigators should address confounding in case-control studies, either in the design stage or with analytical techniques. Devotion of meticulous attention to these points enhances the validity of the results and bolsters the reader's confidence in the findings.
A cohort study tracks two or more groups forward from exposure to outcome. This type of study can be done by going ahead in time from the present (prospective cohort study) or, alternatively, by going back in time to comprise the cohorts and following them up to the present (retrospective cohort study). A cohort study is the best way to identify incidence and natural history of a disease, and can be used to examine multiple outcomes after a single exposure. However, this type of study is less useful for examination of rare events or those that take a long time to develop. A cohort study should provide specific definitions of exposures and outcomes: determination of both should be as objective as possible. The control group (unexposed) should be similar in all important respects to the exposed, with the exception of not having the exposure. Observational studies, however, rarely achieve such a degree of similarity, so investigators need to measure and control for confounding factors. Reduction of loss to follow-up over time is a challenge, since differential losses to follow-up introduce bias. Variations on the cohort theme include the before-after study and nested case-control study (within a cohort study). Strengths of a cohort study include the ability to calculate incidence rates, relative risks, and 95% CIs. This format is the preferred way of presenting study results, rather that with p values.
Readers of medical literature need to consider two types of validity, internal and external. Internal validity means that the study measured what it set out to; external validity is the ability to generalise from the study to the reader's patients. With respect to internal validity, selection bias, information bias, and confounding are present to some degree in all observational research. Selection bias stems from an absence of comparability between groups being studied. Information bias results from incorrect determination of exposure, outcome, or both. The effect of information bias depends on its type. If information is gathered differently for one group than for another, bias results. By contrast, non-differential misclassification tends to obscure real differences. Confounding is a mixing or blurring of effects: a researcher attempts to relate an exposure to an outcome but actually measures the effect of a third factor (the confounding variable). Confounding can be controlled in several ways: restriction, matching, stratification, and more sophisticated multivariate techniques. If a reader cannot explain away study results on the basis of selection, information, or confounding bias, then chance might be another explanation. Chance should be examined last, however, since these biases can account for highly significant, though bogus results. Differentiation between spurious, indirect, and causal associations can be difficult. Criteria such as temporal sequence, strength and consistency of an association, and evidence of a dose-response effect lend support to a causal link.
Descriptive studies often represent the first scientific toe in the water in new areas of inquiry. A fundamental element of descriptive reporting is a clear, specific, and measurable definition of the disease or condition in question. Like newspapers, good descriptive reporting answers the five basic W questions: who, what, why, when, where. and a sixth: so what? Case reports, case-series reports, cross-sectional studies, and surveillance studies deal with individuals, whereas ecological correlational studies examine populations. The case report is the least-publishable unit in medical literature. Case-series reports aggregate individual cases in one publication. Clustering of unusual cases in a short period often heralds a new epidemic, as happened with AIDS. Cross-sectional (prevalence) studies describe the health of populations. Surveillance can be thought of as watchfulness over a community; feedback to those who need to know is an integral component of surveillance. Ecological correlational studies look for associations between exposures and outcomes in populations-eg, per capita cigarette sales and rates of coronary artery disease-rather than in individuals. Three important uses of descriptive studies include trend analysis, health-care planning, and hypothesis generation. A frequent error in reports of descriptive studies is overstepping the data: studies without a comparison group allow no inferences to be drawn about associations, causal or otherwise. Hypotheses about causation from descriptive studies are often tested in rigorous analytical studies.
Many clinicians report that they cannot read the medical literature critically. To address this difficulty, we provide a primer of clinical research for clinicians and researchers alike. Clinical research falls into two general categories: experimental and observational, based on whether the investigator assigns the exposures or not. Experimental trials can also be subdivided into two: randomised and non-randomised. Observational studies can be either analytical or descriptive. Analytical studies feature a comparison (control) group, whereas descriptive studies do not. Within analytical studies, cohort studies track people forward in time from exposure to outcome. By contrast, case-control studies work in reverse, tracing back from outcome to exposure. Cross-sectional studies are like a snapshot, which measures both exposure and outcome at one time point. Descriptive studies, such as case-series reports, do not have a comparison group. Thus, in this type of study, investigators cannot examine associations, a fact often forgotten or ignored. Measures of association, such as relative risk or odds ratio, are the preferred way of expressing results of dichotomous outcomes-eg, sick versus healthy. Confidence intervals around these measures indicate the precision of these results. Measures of association with confidence intervals reveal the strength, direction, and a plausible range of an effect as well as the likelihood of chance occurrence. By contrast, p values address only chance. Testing null hypotheses at a p value of 0.05 has no basis in medicine and should be discouraged.
Side effects caused by oral contraceptives discourage compliance with and continuation of oral contraceptives. A suggested disadvantage of biphasic oral contraceptive pills compared to triphasic oral contraceptive pills is an increase in breakthrough bleeding. We examined this potential disadvantage by conducting a systematic review comparing biphasic oral contraceptives with triphasic oral contraceptives in terms of efficacy, cycle control, and discontinuation because of side effects. We included randomized, controlled trials comparing any biphasic oral contraceptive with any triphasic oral contraceptive when used to prevent pregnancy. Only two trials of limited quality met our inclusion criteria. Larranaga compared two biphasic and one triphasic pills, each containing levonorgestrel and ethinyl estradiol. No important differences emerged, and the frequency of discontinuation because of medical problems was similar with all three pills. Percival-Smith compared a biphasic pill containing norethindrone (Ortho 10/11) with a triphasic pill containing levonorgestrel (Triphasil) and another triphasic pill containing norethindrone (Ortho 7/7/7). The biphasic pill had inferior cycle control compared with the levonorgestrel triphasic pill. The available evidence is limited and of poor quality; the internal validity of these trials is questionable. Given that caveat, the biphasic pill containing norethindrone was associated with inferior cycle control compared with the triphasic pill containing levonorgestrel. This suggests that the choice of progestin may be more important that the phasic regimen in determining bleeding patterns.
BACKGROUND: Side-effects caused by oral contraceptives discourage compliance with and continuation of oral contraceptives. Three approaches have been used to decrease these adverse effects: reduction of steroid dose, development of new steroids, and new formulas and schedules of administration. The third strategy led to the biphasic oral contraceptive pill. We compared biphasic oral contraceptives with monophasic oral contraceptives in terms of efficacy, cycle control and discontinuation due to side-effects. Our a priori hypotheses were: (i) biphasic oral contraceptives are less effective in preventing pregnancy than monophasic oral contraceptives, and (ii) biphasic oral contraceptives cause more side-effects, give poorer cycle control and have lower continuation rates. METHODS: We searched computerized databases Medline, Embase, Popline and the Cochrane Controlled Trial Register. Additionally, we searched the reference lists of all potentially relevant articles and book chapters. We also contacted the authors of relevant studies and pharmaceutical companies in Europe and the USA. We included randomized controlled trials comparing any biphasic oral contraceptive with any monophasic oral contraceptive when used to prevent pregnancy. We examined the studies found during the various literature searches for possible inclusion and assessed their methodological quality using the Cochrane guidelines. We contacted the authors of all included studies and of possibly randomized studies for supplementary information about the study methods and outcomes. We entered the data in RevMan 3.1, imported the data into RevMan 4.1, and calculated Peto odds ratios for the incidence of intermenstrual bleeding, absence of withdrawal bleeding and study discontinuation due to intermenstrual bleeding. RESULTS: Only one trial of limited quality compared a biphasic and monophasic preparation. This trial examined 533 user cycles of a biphasic pill (norethindrone 500 microg/ethinyl estradiol 35 microg for 10 days, followed by norethindrone 1000 microg/ethinyl estradiol 35 microg for 11 days) and 481 user cycles of a monophasic contraceptive pill (norethindrone acetate 1500 microg/ethinyl estradiol 30 microg daily). The study found no significant differences in intermenstrual bleeding, amenorrhoea and study discontinuation due to intermenstrual bleeding between the biphasic and monophasic oral contraceptive pills. CONCLUSIONS: Conclusions are limited by the identification of only one trial, the methodological shortcomings of that trial and the absence of data on accidental pregnancies. However, the trial found no important differences in bleeding patterns between the biphasic and monophasic preparations studied. Since no clear rationale exists for biphasic pills and since extensive evidence is available for monophasic pills, the latter are preferred.