PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “imputation of missing not at random values”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Compliance in clinical trials: impact on design, analysis and interpretation.

A positive association between compliance and clinical outcome has been observed in several randomized, controlled, clinical trials. This association, seen in the placebo-treated group as well as the active-treatment group, clarifies the possibility that data analyses incorporating estimates of protocol adherence are potentially biased. In the presence of non-compliance, or missing data from any cause, several statistical analyses may seem plausible, with none clearly superior to the others. These may include an analysis of all patients randomized, with imputed values for missing data, and an analysis restricted to protocol-adherent patients. The recommended approach is a conservative one that examines consistency among the plausible analyses. Using compliance data in trial conduct can also introduce bias into trial results by inducing differential treatment of compliers and non-compliers. This possibility arises, for instance, when adherence is affected by the randomized treatment. Non-compliance can have a substantial impact on statistical power and sample size requirements in a clinical trial. Under certain assumptions, required sample sizes are doubled with 30% non-compliance and tripled with 40% non-compliance.

Clinical Trials as Topic↗

Dealing with missing data in a multi-question depression scale: a comparison of imputation methods.

BACKGROUND: Missing data present a challenge to many research projects. The problem is often pronounced in studies utilizing self-report scales, and literature addressing different strategies for dealing with missing data in such circumstances is scarce. The objective of this study was to compare six different imputation techniques for dealing with missing data in the Zung Self-reported Depression scale (SDS). METHODS: 1580 participants from a surgical outcomes study completed the SDS. The SDS is a 20 question scale that respondents complete by circling a value of 1 to 4 for each question. The sum of the responses is calculated and respondents are classified as exhibiting depressive symptoms when their total score is over 40. Missing values were simulated by randomly selecting questions whose values were then deleted (a missing completely at random simulation). Additionally, a missing at random and missing not at random simulation were completed. Six imputation methods were then considered; 1) multiple imputation, 2) single regression, 3) individual mean, 4) overall mean, 5) participant's preceding response, and 6) random selection of a value from 1 to 4. For each method, the imputed mean SDS score and standard deviation were compared to the population statistics. The Spearman correlation coefficient, percent misclassified and the Kappa statistic were also calculated. RESULTS: When 10% of values are missing, all the imputation methods except random selection produce Kappa statistics greater than 0.80 indicating 'near perfect' agreement. MI produces the most valid imputed values with a high Kappa statistic (0.89), although both single regression and individual mean imputation also produced favorable results. As the percent of missing information increased to 30%, or when unbalanced missing data were introduced, MI maintained a high Kappa statistic. The individual mean and single regression method produced Kappas in the 'substantial agreement' range (0.76 and 0.74 respectively). CONCLUSION: Multiple imputation is the most accurate method for dealing with missing data in most of the missind data scenarios we assessed for the SDS. Imputing the individual's mean is also an appropriate and simple method for dealing with missing data that may be more interpretable to the majority of medical readers. Researchers should consider conducting methodological assessments such as this one when confronted with missing data. The optimal method should balance validity, ease of interpretability for readers, and analysis expertise of the research team.

Alberta↗

Usefulness of imputation for the analysis of incomplete otoneurologic data.

The usefulness of imputation in the treatment of missing values of an otoneurologic database for the discriminant analysis was evaluated on the basis of the agreement of imputed values and the analysis results. The data consisted of six patient groups with vertigo (N=564). There were 38 variables and 11% of the data was missing. Missing values were filled in with the means, regression and Expectation-Maximisation (EM) imputation methods and a random imputation method provided the baseline results. Means, regression and EM methods agreed on 41-42% of the imputed missing values. The level of agreement between these and the random method was 20-22%. Despite the moderate agreement between the means, regression and EM methods, the discriminant functions were similar and accurate (prediction accuracy 83-99%). The discriminant functions obtained from the randomly imputed data were also accurate having prediction accuracy 88-97%. Imputation seems to be a useful method for treating the missing data in this database. However, a lot of data was missing in otoneurologic tests, which are likely to be of less importance in the diagnosis of vertiginous patients. Consequently, the disagreement of the methods did not affect clearly the discriminant analysis, and, therefore, future research requires more complete data and advanced imputation methods.

Data Collection↗

Tools for statistical analysis with missing data: application to a large medical database.

Missing data is a common feature of large data sets in general and medical data sets in particular. Depending on the goal of statistical analysis, various techniques can be used to tackle this problem. Imputation methods consist in substituting the missing values with plausible or predicted values so that the completed data can then be analysed with any chosen data mining procedure. In this work, we study imputation in the context of multivariate data and we evaluate a number of methods which can be used by today's standard statistical software packages. Imputation using multivariate classification, multiple imputation and imputation by factorial analysis are compared using simulated data and a large medical database (from the diabetes field) with numerous missing values. Our main result is to provide a control chart for assessing data quality after the imputation process. To this end, we developed an algorithm for which the input is a set of parameters describing the underlying data (e.g., covariance matrix, distribution) and the output is a chart which plots the change in the prediction error with respect to the proportion of missing values. The chart is built by means of an iterative algorithm involving four steps: (1) a sample of simulated data is drawn by using the input parameters; (2) missing values are randomly generated; (3) an imputation method is used to fill in the missing data and (4) the prediction error is computed. Steps 1 to 4 are repeated in order to estimate the distribution of the prediction error. The control chart was established for the 3 imputation methods studied here, assuming a multivariate normal distribution of data. The use of this tool on a large medical database was then investigated. We show how the control chart can be used to assess the quality of the imputation process in the pre-processing step upstream of data mining procedures.

Algorithms↗

Using the outcome for imputation of missing predictor values was preferred.

BACKGROUND AND OBJECTIVE: Epidemiologic studies commonly estimate associations between predictors (risk factors) and outcome. Most software automatically exclude subjects with missing values. This commonly causes bias because missing values seldom occur completely at random (MCAR) but rather selectively based on other (observed) variables, missing at random (MAR). Multiple imputation (MI) of missing predictor values using all observed information including outcome is advocated to deal with selective missing values. This seems a self-fulfilling prophecy. METHODS: We tested this hypothesis using data from a study on diagnosis of pulmonary embolism. We selected five predictors of pulmonary embolism without missing values. Their regression coefficients and standard errors (SEs) estimated from the original sample were considered as "true" values. We assigned missing values to these predictors--both MCAR and MAR--and repeated this 1,000 times using simulations. Per simulation we multiple imputed the missing values without and with the outcome, and compared the regression coefficients and SEs to the truth. RESULTS: Regression coefficients based on MI including outcome were close to the truth. MI without outcome yielded very biased--underestimated--coefficients. SEs and coverage of the 90% confidence intervals were not different between MI with and without outcome. Results were the same for MCAR and MAR. CONCLUSION: For all types of missing values, imputation of missing predictor values using the outcome is preferred over imputation without outcome and is no self-fulfilling prophecy.

Adult↗

Multiple imputation for body mass index: lessons from the Australian Longitudinal Study on Women's Health.

In large epidemiological studies missing data can be a problem, especially if information is sought on a sensitive topic or when a composite measure is calculated from several variables each affected by missing values. Multiple imputation is the method of choice for 'filling in' missing data based on associations among variables. Using an example about body mass index from the Australian Longitudinal Study on Women's Health, we identify a subset of variables that are particularly useful for imputing values for the target variables. Then we illustrate two uses of multiple imputation. The first is to examine and correct for bias when data are not missing completely at random. The second is to impute missing values for an important covariate; in this case omission from the imputation process of variables to be used in the analysis may introduce bias. We conclude with several recommendations for handling issues of missing data.

Australia↗

Missing data imputation in quality-of-life assessment: imputation for WHOQOL-BREF.

INTRODUCTION: This study investigated the effects of imputing missing data in the WHO Quality of Life Abbreviated Questionnaire (WHOQOL-BREF). The imputation results from both the item and domain levels were compared and the impact of the missing data rate and the number of items included for imputation were examined. METHODS: An empirical analysis and a simulation study were used to examine the effects of missing data rates and the number of items used for imputation on the accuracy for imputation. In the empirical analysis, both item-level and domain-level imputations were performed, and the missing values were imputed using different amounts of data. In the simulation study, sets of 2%, 5% and 10% of the data were drawn randomly and replaced with missing values. Twenty datasets were generated for each situation. The data were imputed and the accuracy of the imputation was reported. RESULTS: In the empirical study, the number of items used for imputation had only a small impact on the accuracy of imputation. Furthermore, in the simulation study, the accuracy rates of imputation did not significantly change as the proportions of missing data increased. However, the number of items used in the computation did contribute to some extent to the missing values imputed. Extreme responses had the worst computations and the lowest accuracy rates. CONCLUSION: It is recommended that as many items as possible be included for imputation within the same domain. However, it is not particularly helpful to use items from different domains for imputation. Researchers should exercise extra caution in interpreting the imputed values of extreme responses.

Data Interpretation, Statistical↗

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics↗

Estimating treatment effects from longitudinal clinical trial data with missing values: comparative analyses using different methods.

The selection of a method for estimating treatment effects in an intent-to-treat analysis from clinical trial data with missing values often depends on the field of practice. The last observation carried forward (LOCF) analysis assumes that the responses do not change after dropout. Such an assumption is often unrealistic. Analysis with completers only requires that missing values occur completely at random (MCAR). Ignorable maximum likelihood (IML) and multiple imputation (MI) methods require that data are missing at random (MAR). We applied these four methods to a randomized clinical trial comparing anti-depressant effects in an elderly depressed group of patients using a mixed model to describe the course of the treatment effects. Results from an explanatory approach showed a significant difference between the treatments using LOCF and IML methods. Statistical tests indicate violation of the MCAR assumption favoring the flexible IML and MI methods. IML and MI methods were repeated under the pragmatic approach, using data collected after termination of protocol treatment and compared with previously reported results using piecewise splines and rescue (treatment adjustment) pragmatic analysis. No significant treatment differences were found. We conclude that attention to the missing-data mechanism should be an integral part in analysis of clinical trial data.

Aged↗

Treatment of missing values with imputation for the analysis of otologic data.

Usefulness of imputation in the treatment of missing values in an otologic database was studied. Missing values were filled in with means (ME), regression (LR) and Expectation-Maximization (EM) imputation methods. A random imputation method (RA) provided baseline results. ME, LR and EM methods agreed on 41-42% of the imputed missing values. The level of agreement between these and RA method was 20-22%. Despite the moderate agreement, discriminant functions were similar and accurate (prediction accuracy 83-99%) for each diagnosis. A lot of data were missing in otoneurotologic tests which have less weight in the diagnosis of vertiginous patients. Consequently, the disagreement of the methods did not affect discriminant analysis. Inputation seems to be a useful method to treat missing data in this database, but future research requires more complete data and advanced imputation methods.

Data Collection↗

Evaluating the effects of tubal sterilization on menstrual function: selected issues in data analysis.

We examined selected issues in data analysis in the Collaborative Review of Sterilization (CREST). CREST is a multicentre, prospective, observational study of women undergoing tubal sterilization. We analysed menstrual function after sterilization in over 5000 women who were enrolled in the period 1978-1983 and followed for 5 years with yearly follow-up interviews. To take into account the dependency among repeated responses from the same individuals, we used the generalized estimating equations (GEE) approach to longitudinal data analysis. Marginal modelling resulted in a statistically significant increase in the odds of menstrual dysfunction at 5 years after tubal sterilization. Transitional modelling produced rates of menstrual dysfunction given a woman's menstrual function at baseline, after adjusting for other baseline characteristics such as method of contraception before sterilization. To examine the direction of the bias that could result from non-random missing data, we refitted our models using imputed values. The models with imputed values showed the same trends as the original models.

Bias↗

Methods for handling missing data in palliative care research.

Missing data is a common problem in palliative care research due to the special characteristics (deteriorating condition, fatigue and cachexia) of the population. Using data from a palliative study, we illustrate the problems that missing data can cause and show some approaches for dealing with it. Reasons for missing data and ways to deal with missing data (including complete case analysis, imputation and modelling procedures) are explored. Possible mechanisms behind the missing data are: missing completely at random, missing at random or missing not at random. In the example study, data are shown to be missing at random. Imputation of missing data is commonly used (including last value carried forward, regression procedures and simple mean). Imputation affects subsequent summary statistics and analyses, and can have a substantial impact on estimated group means and standard deviations. The choice of imputation method should be carried out with caution and the effects reported.

Adult↗

A preprocessing method for improving data mining techniques. Application to a large medical diabetes database.

The Knowledge Discovery in Databases (KDD) methodology seems to be attractive on the analyze of large clinical databases. In the KDD process, the preprocessing step (data cleaning and handling of missing values) is paramount since it conditions the quality of the results obtained by data mining procedures and represents about 80% of the whole project time. The aims of the present study were to analyze this step and provide tools to handle inconsistent data and missing values. We have broken down the process into 3 main stages: data cleaning--explanatory study of missing values--choice of the procedure used for handling missing values. The data cleaning stage was based on a system of logical rules to correct mistakes and on cluster analysis to discard the poorly filled files. The missing-data mechanism was analyzed by means of multivariate statistical procedures. Two methods to deal with missing values were compared: imputation by the most common value (mode) and imputation using decision trees. This study was performed on a large medical diabetes database (23,601 patients) including numerous missing values. A system of logical rules allowed to correct mistakes on essential parameters (for example, the type of diabetes). Cluster analysis allowed to identify 10% of poorly filled files. After multivariate analysis, the missing-data mechanism could be considered as random. For variables with low number of missing values (< 10%) and categories (< 4), imputation using decision trees provided better results than imputation by mode.

Data Interpretation, Statistical↗

The advantages of community-randomized trials for evaluating lifestyle modification.

Observational studies may provide suggestive evidence for the results of behavior change and lifestyle modification, but they do not replace randomized trials for comparing interventions. To obtain a valid comparison of competing intervention strategies, randomized trials of adequate size are the recommended approach. Randomization avoids bias, achieves balance (on average) of both known and unknown predictive factors between intervention and comparison groups, and provides the basis of statistical tests. The value of randomization is as relevant when investigating community interventions as it is for studies that are directed at individuals. Randomization by group is less efficient statistically than randomization by individual, but there are reasons why randomization by group (such as community) may be chosen, including feasibility of delivery of the intervention, political and administrative considerations, avoiding contamination between individuals allocated to competing interventions, and the very nature of the intervention. One example is the Community Intervention Trial for Smoking Cessation (COMMIT), which involved 11 matched pairs of communities and randomized within these pairs to active community-level intervention versus comparison. For analysis of results, community-level permutation tests (and corresponding test-based confidence intervals) can be designed based on the randomization distribution. The advantages of this approach are that it is robust, and the unit of randomization is the unit of analysis, yet it can incorporate individual-level covariates. Such covariates can play a role in imputation for missing values, adjustment for imbalances, and separate analyses in demographic subsets (with appropriate tests for interaction). A community-randomized trial can investigate a multichannel community-based approach to lifestyle modification, thus providing generalizability coupled with a rigorous evaluation of the intervention.

Adult↗

Review: a gentle introduction to imputation of missing values.

In most situations, simple techniques for handling missing data (such as complete case analysis, overall mean imputation, and the missing-indicator method) produce biased results, whereas imputation techniques yield valid results without complicating the analysis once the imputations are carried out. Imputation techniques are based on the idea that any subject in a study sample can be replaced by a new randomly chosen subject from the same source population. Imputation of missing data on a variable is replacing that missing by a value that is drawn from an estimate of the distribution of this variable. In single imputation, only one estimate is used. In multiple imputation, various estimates are used, reflecting the uncertainty in the estimation of this distribution. Under the general conditions of so-called missing at random and missing completely at random, both single and multiple imputations result in unbiased estimates of study associations. But single imputation results in too small estimated standard errors, whereas multiple imputation results in correctly estimated standard errors and confidence intervals. In this article we explain why all this is the case, and use a simple simulation study to demonstrate our explanations. We also explain and illustrate why two frequently used methods to handle missing data, i.e., overall mean imputation and the missing-indicator method, almost always result in biased estimates.

Bias↗

Interplay between design and analysis for behavioral intervention trials with community as the unit of randomization.

This paper outlines an approach for the design and analysis of randomized controlled trials investigating community-based interventions for behavioral change aimed at health promotion. The approach is illustrated using the Community Intervention Trial for Smoking Cessation (COMMIT), conducted from 1988 to 1993, involving 11 pairs of communities in North America, matched on geographic location, size, and sociodemographic factors. The situation discussed is when assignment to intervention is done at the community level; for COMMIT, the very nature of the intervention required this. The number of communities as a key determinant of the statistical power of the trial. The use of matched pairs of communities can achieve a gain in statistical efficiency. Randomization is used to obtain an unbiased assessment of the intervention effect; randomization also provides the basis for statistical analysis. Permutation tests (and corresponding test-based confidence intervals), using community as the unit of analysis, follow directly from the randomization distribution. Within this framework, individual-level covariates can be used for imputation of missing values and for adjusting analyses of intervention effect.

Health Behavior↗

[Roaming through methodology. XVI. What to do about missing data].

Medical scientific research involving multiple measurements in patients is usually complicated by missing values. In case of missing values the choice is to limit the analysis to the complete cases or to analyse all available data. Both methods may suffer from substantial bias and may only be applied in a valid way if the rather strong assumption of 'missing completely at random' holds for the missing values, i.e. the missing value is not related to the other measured data nor to unmeasured data. Two other statistical methods may be applied to deal with missing values: the likelihood approach and the multiple imputation method. These methods make efficient use of all available data and take into account information implied by the available data. These methods are valid under the less stringent assumption of 'missing at random', i.e. the missing value is related to the other measured data, but not to unmeasured data. The best approach is to ensure that no data are missing.

Clinical Trials as Topic↗

The influence of missing value imputation on detection of differentially expressed genes from microarray data.

MOTIVATION: Missing values are problematic for the analysis of microarray data. Imputation methods have been compared in terms of the similarity between imputed and true values in simulation experiments and not of their influence on the final analysis. The focus has been on missing at random, while entries are missing also not at random. RESULTS: We investigate the influence of imputation on the detection of differentially expressed genes from cDNA microarray data. We apply ANOVA for microarrays and SAM and look to the differentially expressed genes that are lost because of imputation. We show that this new measure provides useful information that the traditional root mean squared error cannot capture. We also show that the type of missingness matters: imputing 5% missing not at random has the same effect as imputing 10-30% missing at random. We propose a new method for imputation (LinImp), fitting a simple linear model for each channel separately, and compare it with the widely used KNNimpute method. For 10% missing at random, KNNimpute leads to twice as many lost differentially expressed genes as LinImp. AVAILABILITY: The R package for LinImp is available at http://folk.uio.no/idasch/imp.

Algorithms↗