PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “imputation of missing not at random values”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Intent-to-treat analysis for longitudinal studies with drop-outs.

We consider intent-to-treat (IT) analysis of clinical trials involving longitudinal data subject to drop-out. Common methods, such as Last Observation Carried Forward imputation or incomplete-data methods based on models that assume random dropout, have serious drawbacks in the IT setting. We propose a method that involves multiple imputation of the missing values following drop-out based on an "as treated" model, using actual dose after drop-out if this is known, or imputed doses that incorporate a variety of plausible alternative assumptions if unknown. The multiply-imputed data sets are then analyzed using IT methods, were subjects are classified by randomization group rather than by the dose actually received. Results from the multiply-imputed data sets are combined using the methods of Rubin (1987, Multiple Imputation for Nonresponse in Surveys). A novel feature of the proposed method is that the models for imputation differ from the model used for the analysis of the filled-in data. The method is applied to data on a clinical trial for Tacrine in the treatment of Alzheimer's disease.

Alzheimer Disease↗

A SAS macro for a simulation study of imputation methods for missing values--an application of Bebbington's algorithm.

This paper presents a SAS macro for a simulation study of comparing a new variant of hot-deck imputation with mean imputation for missing values, in which a simple algorithm proposed by Bebbington (Applied Statistics, 1975) for carrying out simple random sampling without replacement was employed to draw repeated random samples efficiently. A simulated example of drawing repeated random samples from a regional survey of obesity in school children was used to demonstrate the SAS macro.

Algorithms↗

The long-term course of low-density lipoprotein cholesterol after initiation of statin treatment: retrospective database analysis over 3 years in health maintenance organization enrollees.

OBJECTIVES: Our primary objective was to obtain robust estimates of the low-density lipoprotein cholesterol (LDL-C) decrease from baseline over a long period (ie, 3 years) after initiation of statin treatment in a usual-care setting. Our secondary objective was to investigate the predictors of the LDL-C time course. METHODS: We retrospectively analyzed the data for a sample of enrollees in a health maintenance organization (HMO) who started statin treatment between October 1, 1995, and December 31, 1998. Using the HMO's claims database, we examined the LDL-C change from baseline (as measured at the prescribing physicians' discretion) and computed mean estimates every 6 months up to 3 years. We investigated potential predictors of the LDL-C time course (ie, age, sex, baseline LDL-C, previous treatment, prescribing physician's specialty, and most recent treatment) with a mixed model applied to longitudinal data. This model enabled us to impute missing data for all enrollees still being followed, including those who had stopped treatment, and to discuss the robustness of our findings. We applied 2 methods of imputation, assuming either of the following: (1) data were missing at random but could be estimated from the parameters in the mixed model, or (2) LDL-C returned to the baseline value > or =15 days after treatment cessation. RESULTS: We examined data from 3193 individuals. In most cases, the statin used was fluvastatin or pravastatin. The observed mean (95% CI) LDL-C decrease from baseline widened progressively from 23.6% (23.0%-24.3%) at 6 months to 28.0% (27.1%-28.9%) at 18 months and 30.2% (28.7%-31.7%) at 36 months after treatment initiation. These results remained robust after the imputation of missing data, with mean LDL-C reduction consistently >20% at each 6-month time point during the 3 years after treatment initiation. Variations as a function of baseline characteristics were limited (demographics) or explicable by extraneous factors (baseline LDL-C). There were predictable variations as a function of the most recent treatment. CONCLUSIONS: This analysis indicates a long-term reduction in LDL-C among a sample of HMO enrollees who initiated statin treatment in a usual-care setting. The results were robust after imputation of missing data, with mean decrease from baseline consistently >20% over 3 years. However, given the retrospective design of our study and the absence of a control group, we cannot determine how much of the decrease was attributable to treatment.

Aged↗

A comparison of the random-effects pattern mixture model with last-observation-carried-forward (LOCF) analysis in longitudinal clinical trials with dropouts.

The last-observation-carried-forward imputation method is commonly used for imputting data missing due to dropouts in longitudinal clinical trials. The method assumes that outcome remains constant at the last observed value after dropout, which is unlikely in many clinical trials. Recently, random-effects regression models have become popular for analysis of longitudinal clinical trial data with dropouts. However, inference obtained from random-effects regression models is valid when the missing-at-random dropout process is present. The random-effects pattern-mixture model, on the other hand, provides an approach that is valid under more general missingness mechanisms. In this article we describe the use of random-effects pattern-mixture models under different patterns for dropouts. First, subjects are divided into groups depending on their missing-data patterns, and then model parameters are estimated for each pattern. Finally, overall estimates are obtained by averaging over the missing-data patterns and corresponding standard errors are obtained using the delta method. A typical longitudinal clinical trial data set is used to illustrate and compare the above methods of data analyses in the presence of missing data due to dropouts.

Clinical Trials as Topic↗

The effect of correlation structure on treatment contrasts estimated from incomplete clinical trial data with likelihood-based repeated measures compared with last observation carried forward ANOVA.

Valid analyses of longitudinal data can be problematic, particularly when subjects dropout prior to completing the trial for reasons related to the outcome. Regulatory agencies often favor the last observation carried forward (LOCF) approach for imputing missing values in the primary analysis of clinical trials. However, recent evidence suggests that likelihood-based analyses developed under the missing at random framework provide viable alternatives. The within-subject error correlation structure is often the means by which such methods account for the bias from missing data. The objective of this study was to extend previous work that used only one correlation structure by including several common correlation structures in order to assess the effect of the correlation structure in the data, and how it is modeled, on Type I error rates and power from a likelihood-based repeated measures analysis (MMRM), using LOCF for comparison. Data from four realistic clinical trial scenarios were simulated using autoregressive, compound symmetric and unstructured correlation structures. When the correct correlation structure was fit, MMRM provided better control of Type I error and power than LOCF. Although misfitting the correlation structure in MMRM inflated Type I error and altered power, misfitting the structure was typically less deleterious than using LOCF. In fact, simply specifying an unstructured matrix for use in MMRM, regardless of the true correlation structure, yielded superior control of Type I error than LOCF in every scenario. The present and previous investigations have shown that the bias in LOCF is influenced by several factors and interactions between them. Hence, it is difficult to precisely anticipate the direction and magnitude of bias from LOCF in practical situations. However, in scenarios where the overall tendency is for patient improvement, LOCF tends to: 1) overestimate a drug's advantage when dropout is higher in the comparator and underestimate the advantage when dropout is lower in the comparator; 2) overestimate a drug's advantage when the advantage is maximum at intermediate time points and underestimate the advantage when the advantage increases over time; and 3) have a greater likelihood of overestimating a drug's advantage when the advantage is small. In scenarios in which the overall tendency is for patient worsening, the above biases are reversed. In the simulation scenarios considered in this study, which were patterned after acute phase neuropsychiatric clinical trials, the likelihood-based repeated measures approach, implemented with standard software, was more robust to the bias from missing data than LOCF, and choice of correlation structure was not an impediment to its implementation.

Analysis of Variance↗

Are missing outcome data adequately handled? A review of published randomized controlled trials in major medical journals.

BACKGROUND: Randomized controlled trials almost always have some individuals with missing outcomes. Inadequate handling of these missing data in the analysis can cause substantial bias in the treatment effect estimates. We examine how missing outcome data are handled in randomized controlled trials in order to assess whether adequate steps have been taken to reduce nonresponse bias and to identify ways to improve procedures for missing data. METHODS: We reviewed all randomized trials published between July and December 2001 in BMJ, JAMA, Lancet and New England Journal of Medicine, excluding trials in which the primary outcome was described as a time-to-event. We focused on trial designs, how missing outcome data were described and the statistical methods used to deal with the missing outcome data, including sensitivity analyses. RESULTS: We identified 71 trials of which 63 (89%) reported having partly missing outcome data: 13 trials had more than 20% of patients with missing outcomes. In 26 trials that measured the outcome at a single time point, 92% performed a complete case analysis and 8% imputed the missing outcomes using baseline values or the worst case value. In 37 trials with repeated measures of the outcome, 46% performed complete case analyses, potentially excluding individuals with some follow-up data, while 14% performed a repeated measures analysis, 19% used the last observation carried forward, 11% imputed with the worst case value and 2% imputed using regression predictions. Thirteen (21%) of trials with missing data reported a sensitivity analysis. CONCLUSIONS: Our review shows that missing outcome data are a common problem in randomized controlled trials, and are often inadequately handled in the statistical analysis in the top tier medical journals. Authors should explicitly state the assumptions underlying the handling of the missing outcomes and justify them through data descriptions and sensitivity analyses.

Bias↗

Imputation of missing values is superior to complete case analysis and the missing-indicator method in multivariable diagnostic research: a clinical example.

BACKGROUND AND OBJECTIVES: To illustrate the effects of different methods for handling missing data--complete case analysis, missing-indicator method, single imputation of unconditional and conditional mean, and multiple imputation (MI)--in the context of multivariable diagnostic research aiming to identify potential predictors (test results) that independently contribute to the prediction of disease presence or absence. METHODS: We used data from 398 subjects from a prospective study on the diagnosis of pulmonary embolism. Various diagnostic predictors or tests had (varying percentages of) missing values. Per method of handling these missing values, we fitted a diagnostic prediction model using multivariable logistic regression analysis. RESULTS: The receiver operating characteristic curve area for all diagnostic models was above 0.75. The predictors in the final models based on the complete case analysis, and after using the missing-indicator method, were very different compared to the other models. The models based on MI did not differ much from the models derived after using single conditional and unconditional mean imputation. CONCLUSION: In multivariable diagnostic research complete case analysis and the use of the missing-indicator method should be avoided, even when data are missing completely at random. MI methods are known to be superior to single imputation methods. For our example study, the single imputation methods performed equally well, but this was most likely because of the low overall number of missing values.

Adult↗

Accounting for dropout bias using mixed-effects models.

Treatment effects are often evaluated by comparing change over time in outcome measures. However, valid analyses of longitudinal data can be problematic when subjects discontinue (dropout) prior to completing the study. This study assessed the merits of likelihood-based repeated measures analyses (MMRM) compared with fixed-effects analysis of variance where missing values were imputed using the last observation carried forward approach (LOCF) in accounting for dropout bias. Comparisons were made in simulated data and in data from a randomized clinical trial. Subject dropout was introduced in the simulated data to generate ignorable and nonignorable missingness. Estimates of treatment group differences in mean change from baseline to endpoint from MMRM were, on average, markedly closer to the true value than estimates from LOCF in every scenario simulated. Standard errors and confidence intervals from MMRM accurately reflected the uncertainty of the estimates, whereas standard errors and confidence intervals from LOCF underestimated uncertainty.

Analysis of Variance↗

A comparison of imputation techniques for handling missing data.

Researchers are commonly faced with the problem of missing data. This article presents theoretical and empirical information for the selection and application of approaches for handling missing data on a single variable. An actual data set of 492 cases with no missing values was used to create a simulated yet realistic data set with missing at random (MAR) data. The authors compare and contrast five approaches (listwise deletion, mean substitution, simple regression, regression with an error term, and the expectation maximization [EM] algorithm) for dealing with missing data, and compare the effects of each method on descriptive statistics and correlation coefficients for the imputed data (n = 96) and the entire sample (n = 492) when imputed data are inculded. All methods had limitations, although our findings suggest that mean substitution was the least effective and that regression with an error term and the EM algorithm produced estimates closest to those of the original variables.

Algorithms↗

Imputing missing repeated measures data: how should we proceed?

OBJECTIVE: This paper compares six missing data methods that can be used for carrying out statistical tests on repeated measures data: listwise deletion, last value carried forward (LVCF), standardized score imputation, regression and two versions of a closest match method. METHOD: The efficacy of each was investigated under a variety of sample sizes and with differing levels of missingness. Randomly selected samples from a dataset (n = 804) were used to compare the methods using t-tests. Efficacy was defined as the closeness of the estimated t-values to the true t-values from the complete dataset. RESULTS: The results suggest a reliable and efficacious basis for imputation method for repeated measures data is to substitute a missing datum with a value from another individual who has the closest scores on the same variable measured at other timepoints, or the average value of four individuals who have the closest scores on the same variable at other timepoints. The LVCF and standardized score methods performed relatively poorly, which is of concern since these are often recommended. Listwise deletion was also an inefficient missing data method. CONCLUSIONS: Researchers should consider using closest match missing data imputation. Since listwise deletion performed poorly, is widely reported and is the default method in many statistical software packages, the findings have broad implications.

Alcoholism↗

A practical guide to applying the intention-to-treat principle to clinical trials in HIV infection.

It is recommended that randomized controlled trials be analyzed on an intention-to-treat (ITT) basis whereby patients are analyzed in the group to which they were originally assigned, irrespective of the treatment actually received. However, in trials of antiretroviral therapy, it is quite common for patients to withdraw from the trial, and information on virological or immunological endpoints may not be available. The way in which this missing information is dealt with in the analysis can have a large effect on the results of a trial and, therefore, the principle of ITT may be adhered to more closely in some studies than in others. This article describes some simple approaches commonly taken to impute missing data values and discusses the possible effects of these approaches on the results of a trial.

Anti-HIV Agents↗

Use of a single global assessment to reduce missing data in a clinical trial with follow-up at one year.

We conducted a randomized controlled trial (ISRCTN96537534) to assess the effects of acupuncture on migraine and chronic tension headache. Patients (n=401) completed a diary of headache severity four times a day for 4 weeks at baseline, immediately following a 3-month treatment period and 1 year after randomization. During the trial, it appeared that dropout might be higher than expected. We therefore obtained a rapid global assessment of headache from participants to aid imputation of missing data. Patients were contacted by telephone and asked to rate current and baseline headache on a 0-10 scale. Use of global assessment reduced the number of patients from whom we obtained no follow-up headache data from 69 (17%) to 24 (6%). Analysis of patients who provided both a diary and a global assessment demonstrated excellent properties of global assessment, with very similar results to the full diary. We therefore used the global assessment to help impute missing 1-year diary scores. Rapid global assessment can be easily implemented in any trial and aids imputation of missing data, though it should not be used instead of more intensive methods of assessment. Further research might usefully examine the value of global assessment for imputation of missing data in a variety of different settings.

Acupuncture↗

Cost-effectiveness of a low-carbohydrate diet and a standard diet in severe obesity.

OBJECTIVE: Low-carbohydrate diets have become a popular alternative to standard diets for weight loss. Our aim was to compare the cost-effectiveness of these two diets. RESEARCH METHODS AND PROCEDURES: The patient population included 129 severely obese subjects (BMI = 42.9) from a randomized trial; participants had a high prevalence of diabetes or metabolic syndrome. We compared within-trial costs, quality-adjusted life years (QALYs), and the incremental cost-effectiveness ratio (CER) for the two study groups. We imputed missing values for QALYs. The CER was bootstrapped to derive 95% confidence intervals and to define acceptability cut-offs. We took a societal perspective for our analysis. RESULTS: Total costs during the one year of the trial were 6742 dollars +/- 6675 and 6249 dollars +/- 5100 for the low-carbohydrate and standard groups, respectively (p = 0.78). Participants experienced 0.64 +/- 0.02 and 0.61 +/- 0.02 QALYs during the one year of the study, respectively (p = 0.17 for difference). The point estimate of the incremental CER was -1225 dollars/QALY (i.e., the low-carbohydrate diet dominated the standard diet). However, in the bootstrap analysis, the wide spread of CERs caused the 95% confidence interval to be undefined. The probabilities that the low-carbohydrate diet was acceptable, using cut-offs of 50,000 dollars/QALY, 100,000 dollars/QALY, and 150,000 dollars/QALY, were 72.4% 78.6%, and 79.8%, respectively. DISCUSSION: The low-carbohydrate diet was not more cost-effective for weight loss than the standard diet in the patient population studied. Larger studies are needed to better assess the cost-effectiveness of dietary therapies for weight loss.

Adult↗

Applications of multiple imputation in medical studies: from AIDS to NHANES.

Rubin's multiple imputation is a three-step method for handling complex missing data, or more generally, incomplete-data problems, which arise frequently in medical studies. At the first step, m (> 1) completed-data sets are created by imputing the unobserved data m times using m independent draws from an imputation model, which is constructed to reasonably approximate the true distributional relationship between the unobserved data and the available information, and thus reduce potentially very serious nonresponse bias due to systematic difference between the observed data and the unobserved ones. At the second step, m complete-data analyses are performed by treating each completed-data set as a real complete-data set, and thus standard complete-data procedures and software can be utilized directly. At the third step, the results from the m complete-data analyses are combined in a simple, appropriate way to obtain the so-called repeated-imputation inference, which properly takes into account the uncertainty in the imputed values. This paper reviews three applications of Rubin's method that are directly relevant for medical studies. The first is about estimating the reporting delay in acquired immune deficiency syndrome (AIDS) surveillance systems for the purpose of estimating survival time after AIDS diagnosis. The second focuses on the issue of missing data and noncompliance in randomized experiments, where a school choice experiment is used as an illustration. The third looks at handling nonresponse in United States National Health and Nutrition Examination Surveys (NHANES). The emphasis of our review is on the building of imputation models (i.e. the first step), which is the most fundamental aspect of the method.

Acquired Immunodeficiency Syndrome↗

The relationship between hot-deck multiple imputation and weighted likelihood.

Hot-deck imputation is an intuitively simple and popular method of accommodating incomplete data. Users of the method will often use the usual multiple imputation variance estimator which is not appropriate in this case. However, no variance expression has yet been derived for this easily implemented method applied to missing covariates in regression models. The simple hot-deck method is in fact asymptotically equivalent to the mean-score method for the estimation of a regression model parameter, so that hot-deck can be understood in the context of likelihood methods. Both of these methods accommodate data where missingness may depend on the observed variables but not on the unobserved value of the incomplete covariate, that is, missing at random (MAR). The asymptotic properties of hot-deck are derived here for the case where the fully observed variables are categorical, though the incomplete covariate(s) may be continuous. Simulation studies indicate that the two methods compare well in small samples and for small numbers of imputations. Current users of hot-deck may now conduct their analysis using mean-score, which is a weighted likelihood method and can thus be implemented by a single pass through the data using any standard package which accommodates weighted regression models. Valid inference is now straightforward using the variance expression provided here. The equivalence of mean-score and hot-deck is illustrated using three clinical data sets where an important covariate is missing for a large number of study subjects.

Angioplasty, Balloon, Coronary↗

A computer program for multivariate ratio analysis (MISCAT).

Analysts must deal frequently with missing data in multivariate analysis. In such cases, estimating the covariance maxtrix V of the dependent variables usually involves initial estimation and iterative adjustment of imputed missing data values, and/or smoothing of an estimate V which is not necessarily positive semi-definite. This paper presents an alternative procedure for computing estimates of relevant multivariate parameters in situations where missing data occur at random and with small probability. MISCAT is a computer program which computes multivariate ratio estimates of the means and a corresponding positive semi-definite estimate of the covariance matrix. It is an extension of GENCAT, which is a program for the generalizaed least squares analysis of categorical data. Thus, one advantage of dealing with missing data in this manner is that variation among the ratio estimates may be conveniently analyzed within MISCAT using asymptotic regression methodology, provided that sample sizes are sufficiently large. An example is given to illustrate such analysis for longitudinal data from a multicenter clinical trial.

Computers↗

Intention-to-treat approach to data from randomized controlled trials: a sensitivity analysis.

The intention-to-treat (ITT) approach to randomized controlled trials analyzes data on the basis of treatment assignment, not treatment receipt. Alternative approaches make comparisons according to the treatment received at the end of the trial (as-treated analysis) or using only subjects who did not deviate from the assigned treatment (adherers-only analysis). Using a sensitivity analysis on data for a hypothetical trial, we compare these different analytical approaches in the context of two common protocol deviations: loss to follow-up and switching across treatments. In each case, two rates of deviation are considered: 10% and 30%. The analysis shows that biased estimates of effect may occur when deviation is nonrandom, when a large percentage of participants switch treatments or are lost to follow-up, and when the method of estimating missing values accounts inadequately for the process causing loss to follow-up. In general, ITT analysis attenuates between-group effects. Trialists should use sensitivity analyses on their data and should compare the characteristics of participants who do and those who do not deviate from the trial protocol. The ITT approach is not a remedy for unsound design, and imputation of missing values is not a substitute for complete, good quality data.

Data Interpretation, Statistical↗

Handling missing values in support vector machine classifiers.

This paper discusses the task of learning a classifier from observed data containing missing values amongst the inputs which are missing completely at random. A non-parametric perspective is adopted by defining a modified risk taking into account the uncertainty of the predicted outputs when missing values are involved. It is shown that this approach generalizes the approach of mean imputation in the linear case and the resulting kernel machine reduces to the standard Support Vector Machine (SVM) when no input values are missing. Furthermore, the method is extended to the multivariate case of fitting additive models using componentwise kernel machines, and an efficient implementation is based on the Least Squares Support Vector Machine (LS-SVM) classifier formulation.

Algorithms↗