PubMed Health⌕ Search

PubMed · 11523072

Applying sample survey methods to clinical trials data.

Abstract

This paper outlines the utility of statistical methods for sample surveys in analysing clinical trials data. Sample survey statisticians face a variety of complex data analysis issues deriving from the use of multi-stage probability sampling from finite populations. One such issue is that of clustering of observations at the various stages of sampling. Survey data analysis approaches developed to accommodate clustering in the sample design have more general application to clinical studies in which repeated measures structures are encountered. Situations where these methods are of interest include multi-visit studies where responses are observed at two or more time points for each patient, multi-period cross-over studies, and epidemiological studies for repeated occurrences of adverse events or illnesses. We describe statistical procedures for fitting multiple regression models to sample survey data that are more effective for repeated measures studies with complicated data structures than the more traditional approaches of multivariate repeated measures analysis. In this setting, one can specify a primary sampling unit within which repeated measures have intraclass correlation. This intraclass correlation is taken into account by sample survey regression methods through robust estimates of the standard errors of the regression coefficients. Regression estimates are obtained from model fitting estimation equations which ignore the correlation structure of the data (that is, computing procedures which assume that all observational units are independent or are from simple random samples). The analytic approach is straightforward to apply with logistic models for dichotomous data, proportional odds models for ordinal data, and linear models for continuously scaled data, and results are interpretable in terms of population average parameters. Through the features summarized here, the sample survey regression methods have many similarities to the broader family of methods based on generalized estimating equations (GEE). Sample survey methods for the analysis of time-to-event data have more recently been developed and implemented in the context of finite probability sampling. Given the importance of survival endpoints in late phase studies for drug development, these methods have clear utility in the area of clinical trials data analysis. A brief overview of methods for sample survey data analysis is first provided, followed by motivation for applying these methods to clinical trials data. Examples drawn from three clinical studies are provided to illustrate survey methods for logistic regression, proportional odds regression and proportional hazards regression. Potential problems with the proposed methods and ways of addressing them are discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

L M LaVange, G G Koch, T A Schwartz. Applying sample survey methods to clinical trials data.. https://doi.org/10.1002/sim.732

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

On the exact interval estimation for the difference in paired areas under the ROC curves.

An important measure for comparison of accuracy between two diagnostic procedures is the difference in paired areas under the receiver operating characteristic (ROC) curves. Non-parametric and maximum likelihood methods have been proposed for interval estimation for the difference in paired areas under ROC curves. However, these two methods are asymptotic procedures and their performance in finite sample sizes has not been thoroughly investigated. We propose to use the concept of generalized pivotal quantities (GPQs) to construct an exact confidence interval for the difference in paired areas under ROC curves. A simulation study is conducted to empirically investigate the probability coverage and expected length of the three methods for various combinations of sample sizes, values of the area under the ROC curve and correlations. Simulation results demonstrate that the exact confidence interval based on the concept of GPQs provides not only sufficient probability coverage but also reasonable expected length. Numerical examples using published data sets illustrate the proposed method.

Clinical Trials as Topic↗

An efficient test for the analysis of dichotomized variables when the reliability is known.

A difference in an outcome variable between the treatment groups in a trial does not necessarily mean that there is a difference in the number of patients who experience relevant improvement on that variable. When the relevant improvement corresponds with an outcome or change in outcome that exceeds a certain threshold, the outcome variable can be dichotomized. A responder is a patient whose outcome exceeds the threshold. Comparisons can be made between the number of responders in the two treatment groups using logistic regression, or some other method to evaluate binary outcomes. An important disadvantage of this approach is the loss of power. In general, it is more efficient to test the difference between the mean values. We developed a statistical test that compares response rates for a dichotomized variable. It requires that an estimate of the reliability of the outcome variable is available. Simulations showed that the test was valid and robust over a wide range of distributions and sample sizes. The power was greater than the power of a chi(2) test, which would enable substantial reduction in the sample size.

Clinical Trials as Topic↗

Sample size determination for logistic regression revisited.

There is no consensus on the approach to compute the power and sample size with logistic regression. Some authors use the likelihood ratio test; some use the test on proportions; some suggest various approximations to handle the multivariate case. We advocate the use of the Wald test since the Z-score is routinely used for statistical significance testing of regression coefficients. The null-variance formula became popular from early studies, which contradicts modern software, which utilizes the method of maximum likelihood estimation (MLE), when the variance of the MLE is estimated at the MLE, not at the null. We derive general Wald-based power and sample size formulas for logistic regression and then apply them to binary exposure and confounder to obtain a closed-form expression. These formulas are applied to minimize the total sample size in a case-control study to achieve a given power by optimizing the ratio of controls to cases. Approximately, the optimal number of controls to cases is equal to the square root of the alternative odds ratio. Our sample size and power calculations can be carried out online at www.dartmouth.edu/ approximately eugened.

Clinical Trials as Topic↗