PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “imputation of missing not at random values”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A comparison of anlaytic methods for non-random missingness of outcome data.

Missing outcome values occur frequently in survey data and are rarely missing randomly. Depending on the pattern of missingness, the choice of analytic method has implications for accuracy of the estimated outcome distribution as well as multivariate models. Data from a study of patterns of care in disabled elders were used to evaluate several common methods when missingness of the outcome was nonrandom. Results from single and multiple model-based imputation were compared with results from complete-case analysis and mean imputation. By ignoring nonrespondents' covariate information, the latter two methods yielded biased estimates of population means. Mean imputation and single model-based imputation underestimated standard errors by treating imputed values as if they were observed. Mean imputation also distorted the relationship between the outcome and predictors. Multiple model-based imputation provided an easily implemented method of adjustment for non-random non-response in both univariate and multivariate analyses.

Activities of Daily Living↗

Imputation of missing data when measuring physical activity by accelerometry.

PURPOSE: We consider the issue of summarizing accelerometer activity count data accumulated over multiple days when the time interval in which the monitor is worn is not uniform for every subject on every day. The fact that counts are not being recorded during periods in which the monitor is not worn means that many common estimators of daily physical activity are biased downward. METHODS: Data from the Trial for Activity in Adolescent Girls (TAAG), a multicenter group-randomized trial to reduce the decline in physical activity among middle-school girls, were used to illustrate the problem of bias in estimation of physical activity due to missing accelerometer data. The effectiveness of two imputation procedures to reduce bias was investigated in a simulation experiment. Count data for an entire day, or a segment of the day were deleted at random or in an informative way with higher probability of missingness at upper levels of body mass index (BMI) and lower levels of physical activity. RESULTS: When data were deleted at random, estimates of activity computed from the observed data and those based on a data set in which the missing data have been imputed were equally unbiased; however, imputation estimates were more precise. When the data were deleted in a systematic fashion, the bias in estimated activity was lower using imputation procedures. Both imputation techniques, single imputation using the EM algorithm and multiple imputation (MI), performed similarly, with no significant differences in bias or precision. CONCLUSIONS: Researchers are encouraged to take advantage of software to implement missing value imputation, as estimates of activity are more precise and less biased in the presence of intermittent missing accelerometer data than those derived from an observed data analysis approach.

Acceleration↗

CD4+ lymphocytes and tuberculin skin test as survival predictors in pulmonary tuberculosis HIV-infected patients.

BACKGROUND: We analyse whether the tuberculin skin test is a good survival marker in a cohort of pulmonary tuberculosis patients with HIV infection (PTB/HIV). In all, 494 PTB/HIV patients were enrolled in Barcelona (Spain) between January 1992 and December 1994 in the Tuberculosis Program of Barcelona. The main data problem was the large proportion of missing values in the covariates percentage of T CD4+ lymphocytes and the tuberculin test results: only 157 patients (31.8%) had both covariates recorded. METHODS: Patients were dichotomized into two groups according to their level of immunosuppression (< or = 14 and >14% T CD4+ cells). First, we carried out the semiparametric and parametric complete case analysis. After this, we analysed the data assuming a missing at random non-response pattern. We developed a bootstrap approach where missing data in the markers are imputed via a two-way linear model. Using Weibull regression estimation, we used a multiple imputation scheme to estimate the parameters of interest. RESULTS: We found significative differences for the most immunosuppressed group when comparing positive tuberculin patients with those who were tuberculin negative. From a complete case approach and through a multivariate Cox analysis, we obtained a significant relative hazard of 0.3657 (95% CI: 0.13-1.02; P = 0.054). When a Weibull model was fitted, we estimated a constant relative percentile value of pR = 4.1329 (95% CI: 0.97-17.59). From a missing data approach, we obtain a higher constant relative percentile 5.48 (P = 0.079). CONCLUSIONS: The imputation method allows us to assess the protective character of positivity for the tuberculin test for the lowest CD4+ level. These findings strongly suggest the value of the tuberculin skin test as a qualitative measure of the immunological response and its interest for developing countries where specific laboratory tests are not affordable.

Adult↗

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem↗

Missing... presumed at random: cost-analysis of incomplete data.

When collecting patient-level resource use data for statistical analysis, for some patients and in some categories of resource use, the required count will not be observed. Although this problem must arise in most reported economic evaluations containing patient-level data, it is rare for authors to detail how the problem was overcome. Statistical packages may default to handling missing data through a so-called 'complete case analysis', while some recent cost-analyses have appeared to favour an 'available case' approach. Both of these methods are problematic: complete case analysis is inefficient and is likely to be biased; available case analysis, by employing different numbers of observations for each resource use item, generates severe problems for standard statistical inference. Instead we explore imputation methods for generating 'replacement' values for missing data that will permit complete case analysis using the whole data set and we illustrate these methods using two data sets that had incomplete resource use information.

Algorithms↗

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross&#x2010;taxon transferability↗

Collateral missing value imputation: a new robust missing value estimation algorithm for microarray data.

MOTIVATION: Microarray data are used in a range of application areas in biology, although often it contains considerable numbers of missing values. These missing values can significantly affect subsequent statistical analysis and machine learning algorithms so there is a strong motivation to estimate these values as accurately as possible before using these algorithms. While many imputation algorithms have been proposed, more robust techniques need to be developed so that further analysis of biological data can be accurately undertaken. In this paper, an innovative missing value imputation algorithm called collateral missing value estimation (CMVE) is presented which uses multiple covariance-based imputation matrices for the final prediction of missing values. The matrices are computed and optimized using least square regression and linear programming methods. RESULTS: The new CMVE algorithm has been compared with existing estimation techniques including Bayesian principal component analysis imputation (BPCA), least square impute (LSImpute) and K-nearest neighbour (KNN). All these methods were rigorously tested to estimate missing values in three separate non-time series (ovarian cancer based) and one time series (yeast sporulation) dataset. Each method was quantitatively analyzed using the normalized root mean square (NRMS) error measure, covering a wide range of randomly introduced missing value probabilities from 0.01 to 0.2. Experiments were also undertaken on the yeast dataset, which comprised 1.7% actual missing values, to test the hypothesis that CMVE performed better not only for randomly occurring but also for a real distribution of missing values. The results confirmed that CMVE consistently demonstrated superior and robust estimation capability of missing values compared with other methods for both series types of data, for the same order of computational complexity. A concise theoretical framework has also been formulated to validate the improved performance of the CMVE algorithm. AVAILABILITY: The CMVE software is available upon request from the authors.

Algorithms↗

Drotrecogin alfa (activated) (recombinant human activated protein C) reduces host coagulopathy response in patients with severe sepsis.

Drotrecogin alfa (activated) improved survival in patients with severe sepsis in PROWESS, a double-blind, study of 1690 adult patients randomized to drotrecogin alfa (activated) at 24 microg/kg/h (N=850) or placebo (N=840) infused for 96 hours. Pharmacodynamic effects of drotrecogin alfa (activated) were assessed with 15 prospectively defined systemic biomarkers of hemostasis, inflammation and endothelial injury. The last-observation-carried-forward (LOCF) method of imputation for missing observations was the prospectively defined statistical method. The results were also analyzed with only the observed values without imputation for missing data (repeated measures analysis). With both statistical methods, drotrecogin alfa (activated)-treated patients demonstrated antithrombotic (reduced markers of thrombin generation and accelerated normalization of anticoagulant factor, protein C and fibrinolytic factors) and anticoagulant (prolonged PT and APTT) effects compared with placebo. A profibrinolytic (reduction in plasminogen activator inhibitor-1) effect was significant only with the LOCF imputation method in observed case and percent change from baseline analyses. An anti-inflammatory (reduction in interleukin-6) effect was significant only with the LOCF imputation method in change from baseline and percent change from baseline analyses. Drotrecogin alfa (activated) is a new and promising agent for treatment of patients with severe sepsis. The extensive analysis of systemic biomarkers confirms the previously published antithrombotic effects. However, the present results using different statistical methods do not provide a strong basis for systemic anti-inflammatory or pro-fibrinolytic effects. These latter two effects may occur at the local or cellular level. The systemic biomarkers reported here might not be the most appropriate approach to demonstrate these potential effects of drotrecogin alfa (activated).

Anti-Inflammatory Agents, Non-Steroidal↗

Random regression with imputed values for dropouts.

The random regression model (RRM) has been advocated as a potential solution to problems of statistical analysis posed by dropouts in clinical trials. However, the power of the RRM tests for differences in rates of change can be seriously attenuated by presence of dropouts. The use of imputed scores and other modifications are examined in an attempt to render a simple growth-curve form of the RRM analysis more robust against dropouts. Methods that extrapolate from an individual's own performance were found effective, although inclusion of time-in-treatment as a covariate was documented to be important under identifiable conditions. Of the methods evaluated, those that used group data to impute missing values for dropouts produced nonconservative bias. The results suggest the importance of careful evaluation of potential bias when integrating any group-based imputation procedure into the RRM analyses.

Clinical Trials as Topic↗

Determinants of power-frequency magnetic fields in residences located away from overhead power lines.

The Wertheimer-Leeper wire code, originally developed as a surrogate for magnetic-field exposure, has been associated with childhood leukemia in several epidemiologic investigations. However, these and other studies indicate that most between-residence variability in measured magnetic fields remains unexplained by wire codes. To better understand this remaining variability, engineering and demographic data were examined for 333 underground (UG) and very-low current configuration (VLCC) single-family or duplex residences, selected from a database of nearly 1000 residences specifically because their magnetic fields are most likely affected negligibly by overhead power lines. Using linear regression techniques, four factors predictive of the log-transformed residential field were identified: the square-root of the 24-h average net service drop current (this current is equivalent to the current in the grounding system), the log of the number of service drops on the same secondary serving the residence, residence age (four categories), and area type (rural, suburban, or urban). Complete data on ground current and service drops, the two factors with the strongest individual relationships to measured fields, were available for only half of the residences in the sample. However, these data were determined to be "missing at random" according to established statistical criteria. The full-sample or "composite" models thus relied on a method similar to regression imputation, accounting for missing data with binary dummy variables. When applied to the samples from which they were derived, these models accounted for 25% of the variance of the log-spot-measured magnetic field values in the full sample, while models that considered only those residences with complete data (n = 167) explained about 35%. The model validated well against a sample of 201 ordinary low current configuration (OLCC) homes selected from the same database.

Child↗

Cost-effectiveness inferences from bootstrap quadrant confidence levels: three degrees of dominance.

When with at least 95% confidence a new treatment is shown to be not only less costly (LC), but also more effective (ME), than a current treatment, that new treatment can be said to "strictly dominate" the current treatment statistically. But what can be said when head-to-head treatment comparisons turn out to be less clear-cut than this? Here, we propose two additional sets of specific LC and/or ME confidence thresholds to define the concepts of "some dominance" and "much dominance." Confidence levels associated with entire quadrants of the incremental cost-effectiveness (ICE) plane are easily computed using the same bootstrapping techniques used to estimate an "acceptability curve." Our two proposed additional "degrees" of dominance, although less stringent than strict dominance, are nevertheless more stringent than commonly accepted approaches using ICE ratio or net benefit calculations. To illustrate analysis concepts, we use data from a randomized, double-blind, placebo- and active comparator-controlled clinical registration trial for treatment of major depressive disorder (MDD). As is typical, our case study is rather small and short term, providing outcome information for a total of only 264 patients during their initial 8 weeks of acute-phase MDD treatment. Thus, we focus attention on sensitivity analyses, showing that the bootstrap distribution of cost-effectiveness uncertainty is robust across two alternative ways of measuring overall effectiveness and three alternative ways of imputing missing values. Evaluation of the balance between cost and benefit is particularly difficult when a new pharmacological treatment is first introduced, yet information of this sort is highly desired by decision makers. We show that, even with only a relatively modest amount of clinical trial information, sensitivity analyses can still confirm that cost-effectiveness comparisons are being made in a consistent fashion. In contrast, extensive follow-up comparisons using data from actual clinical practice will almost always ultimately be needed to better inform health policy makers.

Adult↗

An imputation method for non-ignorable missing data in studies of blood pressure.

In studies with repeated measures of blood pressure (BP), particularly in trials of hypertension prevention, BP measurements often become censored once a participant commences antihypertensive medication. When prescribed by non-study physicians under uncontrolled conditions, the missing data mechanism is non-ignorable and may bias the BP effects of interest. I propose a method that models the distribution of BPs measured by non-study physicians and their relation to study BPs using random effects models. If treated for hypertension, I assume that BP measured outside the study is greater than a clinical cutpoint, such as diastolic BP > or = 90 mmHg. I then compute estimates for the missing study BPs conditional on previously observed study BPs and treatment for hypertension. Multiple imputation is used to model the variability of the BP values and adjust the standard error estimates of the parameters. Examples are given using simulated data and data from the weight loss intervention of phase I of the Trials of Hypertension Prevention.

Antihypertensive Agents↗

Impact of missing data due to dropouts on estimates of the treatment effect in a randomized trial of antiretroviral therapy for HIV-infected individuals. Canadian HIV Trials Network A002 Study Group.

PURPOSE: To evaluate the impact of missing data due to nonrandom dropout on estimates of the effect of treatment on the CD4 count in a clinical trial of antiretroviral therapy for HIV infected individuals. METHODS: The effect of treatment on CD4 counts in a recent study of continued ZDV versus ddI in HIV-infected individuals was estimated from the observed data and after imputing missing CD4 counts for patients who dropped out of the study. Imputation methods studied were (a) carrying forward the last observed CD4 count, (b) predicting missing CD4 counts from regression models, and (c) assuming that CD4 counts of patients who dropped out declined at a rate of 100 cells per year. RESULTS: Of the 245 patients enrolled in the study, 52% completed the planned 48 weeks of follow-up. Patients with lower CD4 counts were more likely to drop out of the study (RR = 1.77; p = 0.0001). Patients receiving ZDV had a greater tendency to drop out than patients receiving ddI (p = 0.07). Mean CD4 counts calculated after imputing missing data were lower than those obtained from the observed data at all follow-up times for both treatment groups. Imputing CD4 counts with regression models yielded higher estimates of the effect of treatment than were obtained using the observed data. CONCLUSION: Missing outcome data due to dropouts can result in an underestimation of the treatment effect and overly optimistic statements about the outcome of participants on both treatment arms due to the selective dropout of participants with lower or decreasing CD4 counts. When there are significant dropout rates in randomized trials, imputation is a useful technique to assess the range of plausible values of the treatment effect.

Antiviral Agents↗

Intent-to-treat analysis for longitudinal clinical trials: coping with the challenge of missing values.

Drop-out is a common phenomenon in clinical trials of drug treatments involving longitudinal assessments for a fixed duration of follow-up. For these trials intent-to-treat (IT) analysis is usually preferred because time effects are seen in practice. The IT analysis mandates that all subjects randomized to a treatment arm should be included in the analysis. The purpose of the present paper is to acquaint both clinicians and statisticians with recent statistical methodological advances in handling drop-outs and their usage for IT analysis. We discuss a sensitivity analysis of 12-month outcome data to investigate the efficacy of drug therapy from a longitudinal double-blind placebo-controlled clinical trial in the maintenance therapy of geriatric major depressive illness. Outcome measures consist of monthly Hamilton depression scores. The sensitivity analysis includes endpoint analysis, last observation carried forward analysis, repeated measures models and imputation models. Imputation models are based on multiple imputations of missing responses deriving from an 'as-treated' model. The model used imputed doses from a plausible treatment scenario after drop-out and a 'propensity-adjusted' model where the imputations for the drop-outs were obtained from the adhering subjects with the same probability to remain on study (propensity) given the observed trajectory prior to withdrawal. Issues related to bias and efficiency of the estimates obtained by different analyses are discussed. We recommend a more widespread use of imputation models for the IT analysis.

Aged↗

Indoor radon and lung cancer risk in connecticut and utah.

Radon is a well-established cause of lung cancer in miners. Residents of homes with high levels of radon are potentially also at risk. Although most individual studies of indoor radon have failed to demonstrate significant risks, results have generally been consistent with estimates from studies of miners. We studied 1474 incident lung cancer cases aged 40-79 yr in Connecticut, Utah, and southern Idaho. Population controls (n = 1811) were identified by random telephone screening and from lists of Medicare recipients, and were selected to be similar to cases on age, gender, and smoking 10 yr before diagnosis/interview using randomized recruitment. Complete residential histories and information on known lung cancer risk factors were obtained by in-person and telephone interviews. Radon was measured on multiple levels of past and current homes using 12-mo alpha-track etch detectors. Missing data were imputed using mean radon concentrations for informative subgroups of controls. Average radon exposures were lower than anticipated, with median values of 23 Bq/m3 in Connecticut and 45 Bq/m3 in Utah/southern Idaho. Overall, there was little association between time-weighted average radon exposures 5 to 25 yr prior to diagnosis/interview and lung cancer risk. The excess relative risk (ERR) associated with a 100-Bq/m3 increase in radon level was 0.002 (95% CI -0.21, 0.21) in the overall population, 0.134 (95% CI -0.23, 0.50) in Connecticut, and -0.112 (95% CI -0.34, 0.11) in Utah/Idaho. ERRs were higher for some subgroups less prone to misclassification, but there was no group with a statistically significant linear increase in risk. While results were consistent with the estimates from studies of miners, this study provides no evidence of an increased risk for lung cancer at the exposure levels observed.

Adult↗

The effects of non-response on statistical inference.

Surveys have been, and will most likely continue to be, the source of data for many empirical articles. Likewise, the difficulty of making valid statistical inferences in the face of missing data will continue to plague researchers. In an ideal situation, all potential survey participants would respond; in reality, the goal of an 80 to 90% response rate is very difficult to achieve. When nonresponse is systematic, the combination of low response rate and systematic differences can severely bias inferences that are made by the researcher to the population. It is important for the researcher to assess the potential causes of nonresponse and the differences between the observed values in the sample compared to what may have been gained if the sample was complete, particularly when the response rate is low. There are methods available that substitute imputed values for missing data, but these methods are useless if the researcher lacks knowledge of how the responders and nonresponders may differ. With regard to statistical inference, the researcher also should be aware of the difference between a convenient sample and a probability sample. Valid statistical inference assumes that the probability of characteristics observed in the sample bear some relationship to their occurrence in the population. For example, in a simple random sample each member of the accessible population has an equal chance of inclusion in the sample. A convenient sample lacks the statistical properties of a probability sample that allow the validity of its inferences to be assessed strictly from a mathematical framework. The context of the research and the type of data being gathered greatly affect the validity of any generalizations the researcher makes with regard to the population the convenient sample attempts to represent.

Bias↗

Backfilling missing microbial concentrations in a riverine database using artificial neural networks.

Predicting peak pathogen loadings can provide a basis for watershed and water treatment plant management decisions that can minimize microbial risk to the public from contact or ingestion. Artificial neural network models (ANN) have been successfully applied to the complex problem of predicting peak pathogen loadings in surface waters. However, these data-driven models require substantial, multiparameter databases upon which to train, and missing input values for pathogen indicators must often be estimated. In this study, ANN models were evaluated for backfilling values for individual observations of indicator bacterial concentrations in a river from 44 other related physical, chemical, and bacteriological data contained in a multi-year database. The ANN modeling approach provided slightly superior predictions of actual microbial concentrations when compared to conventional imputation and multiple linear regression models. The ANN model provided excellent classification of 300 randomly selected, individual data observations into two defined ranges for fecal coliform concentrations with 97% overall accuracy. The application of the relative strength effect (RSE) concept for selection of input variables for ANN modeling and an approach for identifying anomalous data observations utilizing cross validation with ANN model are also presented.

Bacteria↗

Methods for addressing missing data in psychiatric and developmental research.

OBJECTIVE: First, to provide information about best practices in handling missing data so that readers can judge the quality of research studies. Second, to provide more detailed information about missing data analysis techniques and software on the Journal's Web site at www.jaacap.com. METHOD: We focus our review of techniques on those that are based on the "Missing at Random" assumption and are either extremely popular because of their convenience or that are harder to employ but yield more precise inferences. RESULTS: The literature regarding missing data indicates that deletion of observations with missing data can yield biased findings. Other popular methods for handling missing data, notably replacing missing values with means, can lead to confidence intervals that are too narrow as well as false identifications of significant differences (type I statistical errors). Methods such as multiple imputation and direct maximum likelihood estimation are often superior to deleting observations and other popular methods for handling missing data problems. CONCLUSIONS: Psychiatric and developmental researchers should consider using multiple imputation and direct maximum likelihood estimation rather than deleting observations with missing values.

Adolescent↗