PubMed HealthSearch

PubMed · 8796936

Basic statistical testing, including interim analysis.

Abstract

Clinical trials, due to the randomization process, require statistical methods for testing key hypotheses that are straightforward and involve simple comparisons of group proportions or group means. Differences between group proportions are tested using a chi 2 statistic or a Fisher's exact test. Differences between group means are tested using a two-sample t statistic. When there are more than two groups, the F statistic is calculated. When ordinal data or data not normally distributed are analyzed, nonparametric testing is performed. A critical decision prior to analysis is the choice of the endpoint, which must be clearly defined and consistently applied. Monitoring a clinical trial is important, and may lead to interim analysis. This should be done by independent study monitors. Interim results play an important role in analyzing clinical trials; however, they create problems due to multiple testing. An early stoppage rule is frequently applied when independent interim results are obtained. The statistical methods for interim analysis are discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D S Guzick. 1996. Basic statistical testing, including interim analysis.. https://doi.org/10.1055/s-2007-1016321

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Variation of sensitivity, specificity, likelihood ratios and predictive values with disease prevalence.

The sensitivity, specificity and likelihood ratios of binary diagnostic tests are often thought of as being independent of disease prevalence. Empirical studies, however, have frequently revealed substantial variation of these measures for the same diagnostic test in different populations. One reason for this discrepancy is related to the fact that only few diagnostic tests are inherently dichotomous. The majority of tests are based on categorization of individuals according to one or several underlying continuous traits. For these tests, the magnitude of diagnostic misclassification depends not only on the magnitude of the measurement or perception error of the underlying trait(s), but also on the distribution of the underlying trait(s) in the population relative to the diagnostic cutpoint. Since this distribution also determines prevalence of the disease in the population, diagnostic misclassification and disease prevalence are related for this type of test. We assess the variation of various measures of validity of diagnostic tests with disease prevalence for simple models of the distribution of the underlying trait(s) and the measurement or perception error. We illustrate that variation with disease prevalence is typically strong for sensitivity and specificity, and even more so for the likelihood ratios. Although positive and negative predictive values also strongly vary with disease prevalence, this variation is usually less pronounced than one would expect if sensitivity and specificity were independent of disease prevalence.

Data Interpretation, Statistical

Using numerical results from systematic reviews in clinical practice.

Systematic reviews summarize large amounts of information and are more likely than individual trials to describe the true clinical effect of an intervention. Traditional statistical outputs from systematic reviews cannot immediately be applied to clinical practice. The number needed to treat (NNT) has that clinical immediacy. This number can be calculated easily from raw data or from statistical outputs, and the principle involved in its calculation can be applied to different outcomes: treatment efficacy, adverse events (harm), or other end points. The NNT defines the treatment-specific effect of an intervention, and we suggest it as a currency for making decisions about individual patients. Knowing the NNT for different interventions that have the same outcome for the same disorder can help shape individual and institutional practice. Knowing or estimating the number needed to harm is also an important part of the equation. Knowing or estimating an individual patient's risk can, with the NNT, be a guide to the overall or net value of a prophylactic intervention. We advocate an approach to systematic reviews that distills information into, in effect, one number: the NNT. This is simple to remember and directly supports efforts to work with patients to make the best possible clinical decisions for their care.

Data Interpretation, Statistical

Blinded subjective rankings as a method of assessing treatment effect: a large sample example from the Systolic Hypertension in the Elderly Program (SHEP).

Because many randomized clinical trials study more than one important outcome variable, evaluation of efficacy is often difficult and not completely satisfactory. This paper considers the use of a procedure for endpoint determination described by Follmann et al., that allows raters to integrate subjectively all relevant information about an individual's clinical course into a single univariate assessment. To explore the method's feasibility, we tested the procedure with data from a completed clinical trial, the Systolic Hypertension in the Elderly Program (SHEP). We provided raters blinded to treatment assignment with cards that schematically represent the clinical trajectories of SHEP study participants. The raters independently ranked these trajectories. The method combined ranks across raters to determine a single rank for each study participant; we used a rank procedure to test treatment effect. The major findings were: (i) the raters showed a high level of concordance of rankings; (ii) tests of treatment effect were highly statistically significant; (iii) three statistical methods were effective for implementing the ranking in the large study size case. These methods were use of: (a) scoring rules; (b) incomplete block designs, and (c) categorical ranking.

Data Interpretation, Statistical