PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

How to evaluate and improve the quality and credibility of an outcomes database: validation and feedback study on the UK Cardiac Surgery Experience.

OBJECTIVES: To assess the quality and completeness of a database of clinical outcomes after cardiac surgery and to determine whether a process of validation, monitoring, and feedback could improve the quality of the database. DESIGN: Stratified sampling of retrospective data followed by prospective re-sampling of database after intervention of monitoring, validation, and feedback. SETTING: Ten tertiary care cardiac surgery centres in the United Kingdom. INTERVENTION: Validation of data derived from a stratified sample of case notes (recording of deaths cross checked with mortuary records), monitoring of completeness and accuracy of data entry, feedback to local data managers and lead surgeons. MAIN OUTCOME MEASURES: Average percentage missing data, average kappa coefficient, and reliability score by centre for 17 variables required for assignment of risk scores. Actual minus risk adjusted mortality in each centre. RESULTS: The database was incomplete, with a mean (SE) of 24.96% (0.09%) of essential data elements missing, whereas only 1.18% (0.06%) were missing in the patient records (P<0.0001). Intervention was associated with (a) significantly less missing data (9.33% (0.08%) P<0.0001); (b) marginal improvement in reliability of data and mean (SE) overall centre reliability score (0.53 (0.15) v 0.44 (0.17)); and (c) improved accuracy of assigned Parsonnet risk scores (kappa 0.84 v 0.70). Mortality scores (actual minus risk adjusted mortality) for all participating centres fell within two standard deviations of the mean score. CONCLUSION: A short period of independent validation, monitoring, and feedback improved the quality of an outcomes database and improved the process of risk adjustment, but with substantial room for further improvement. Wider application of this approach should increase the credibility of similar databases before their public release.

Cardiac Surgical Procedures↗

Ignorability and parameter estimation in longitudinal pharmacokinetic studies.

In the analysis of longitudinal pharmacokinetic data, both balanced (equal number of samples per subject) and unbalanced data are used. It is implicitly assumed that the process that caused the missing data can be ignored. A simulation study was performed to determine the effect of ignoring the missing data (i.e., "ignorability") on the accuracy and precision of parameter estimation in longitudinal pharmacokinetic studies. A two-compartment model with multiple intravenous bolus inputs was assumed. Subjects with balanced data sets had six samples, and those with unbalanced data had 1 to 5 samples missing (i.e., supplied in a decreasing order from 5 to 1 samples). The proportion of subjects with 1 to 5 samples missing varied from 25% to 75% in a fixed sample size of 100. The effect of ignorability was studied at intersubject variability ranging from 15% to 60% for a drug assumed to be dosed at its elimination half-life. One hundred replicate data sets of 100 subjects each were simulated for each missing data scenario. The accuracy of parameter estimation was not significantly affected by the amount of ignorable missing data at any given level of variability. However, the precision of parameter estimation was affected by the degree of "missingness."

Humans↗

Reparameterizing the pattern mixture model for sensitivity analyses under informative dropout.

Pattern mixture models are frequently used to analyze longitudinal data where missingness is induced by dropout. For measured responses, it is typical to model the complete data as a mixture of multivariate normal distributions, where mixing is done over the dropout distribution. Fully parameterized pattern mixture models are not identified by incomplete data; Little (1993, Journal of the American Statistical Association 88, 125-134) has characterized several identifying restrictions that can be used for model fitting. We propose a reparameterization of the pattern mixture model that allows investigation of sensitivity to assumptions about nonidentified parameters in both the mean and variance, allows consideration of a wide range of nonignorable missing-data mechanisms, and has intuitive appeal for eliciting plausible missing-data mechanisms. The parameterization makes clear an advantage of pattern mixture models over parametric selection models, namely that the missing-data mechanism can be varied without affecting the marginal distribution of the observed data. To illustrate the utility of the new parameterization, we analyze data from a recent clinical trial of growth hormone for maintaining muscle strength in the elderly. Dropout occurs at a high rate and is potentially informative. We undertake a detailed sensitivity analysis to understand the impact of missing-data assumptions on the inference about the effects of growth hormone on muscle strength.

Aged↗

Imputing missing repeated measures data: how should we proceed?

OBJECTIVE: This paper compares six missing data methods that can be used for carrying out statistical tests on repeated measures data: listwise deletion, last value carried forward (LVCF), standardized score imputation, regression and two versions of a closest match method. METHOD: The efficacy of each was investigated under a variety of sample sizes and with differing levels of missingness. Randomly selected samples from a dataset (n = 804) were used to compare the methods using t-tests. Efficacy was defined as the closeness of the estimated t-values to the true t-values from the complete dataset. RESULTS: The results suggest a reliable and efficacious basis for imputation method for repeated measures data is to substitute a missing datum with a value from another individual who has the closest scores on the same variable measured at other timepoints, or the average value of four individuals who have the closest scores on the same variable at other timepoints. The LVCF and standardized score methods performed relatively poorly, which is of concern since these are often recommended. Listwise deletion was also an inefficient missing data method. CONCLUSIONS: Researchers should consider using closest match missing data imputation. Since listwise deletion performed poorly, is widely reported and is the default method in many statistical software packages, the findings have broad implications.

Alcoholism↗

Pattern-mixture models for multivariate incomplete data with covariates.

Pattern-mixture models stratify incomplete data by the pattern of missing values and formulate distinct models within each stratum. Pattern-mixture models are developed for analyzing a random sample on continuous variables y(1), y(2) when values of y(2) are nonrandomly missing. Methods for scalar y(1) and y(2) are here generalized to vector y(1) and y(2) with additional fixed covariates x. Parameters in these models are identified by alternative assumptions about the missing-data mechanism. Models may be underidentified (in which case additional assumptions are needed), just-identified, or overidentified. Maximum likelihood and Bayesian methods are developed for the latter two situations, using the EM and SEM algorithms, direct and interactive simulation methods. The methods are illustrated on a data set involving alternative dosage regimens for the treatment of schizophrenia using haloperidol and on a regression example. Sensitivity to alternative assumptions about the missing-data mechanism is assessed, and the new methods are compared with complete-case analysis and maximum likelihood for a probit selection model.

Algorithms↗

Robust outcome prediction for intensive-care patients.

Missing data are a major plague of medical databases in general, and of Intensive Care Unit databases in particular. The time pressure of work in an Intensive Care Unit pushes the physicians to omit randomly or selectively record data. These different omission strategies give rise to different patterns of missing data and the recommended approach of completing the database using median imputation and fitting a logistic regression model can lead to significant biases. This paper applies a new classification method, called robust Bayes classifier, which does not rely on any particular assumption about the pattern of missing data and compares it to the median imputation approach using a database of 324 Intensive Care Unit patients.

Bayes Theorem↗

A critical look at methods for handling missing covariates in epidemiologic regression analyses.

Epidemiologic studies often encounter missing covariate values. While simple methods such as stratification on missing-data status, conditional-mean imputation, and complete-subject analysis are commonly employed for handling this problem, several studies have shown that these methods can be biased under reasonable circumstances. The authors review these results in the context of logistic regression and present simulation experiments showing the limitations of the methods. The method based on missing-data indicators can exhibit severe bias even when the data are missing completely at random, and regression (conditional-mean) imputation can be inordinately sensitive to model misspecification. Even complete-subject analysis can outperform these methods. More sophisticated methods, such as maximum likelihood, multiple imputation, and weighted estimating equations, have been given extensive attention in the statistics literature. While these methods are superior to simple methods, they are not commonly used in epidemiology, no doubt due to their complexity and the lack of packaged software to apply these methods. The authors contrast the results of multiple imputation to simple methods in the analysis of a case-control study of endometrial cancer, and they find a meaningful difference in results for age at menarche. In general, the authors recommend that epidemiologists avoid using the missing-indicator method and use more sophisticated methods whenever a large proportion of data are missing.

Case-Control Studies↗

New control chart for multivariate data with missing values.

A new multivariate statistical quality control method has been developed. It is an extension of the method developed by Kume, which is able to find abnormal values in multivariate biochemical data of a clinical laboratory. The present method makes use of the difference between two sets of data measured from the samples of the same patient obtained on different days. The Mahalanobis' distance between two samples can be calculated from the difference of their observations. If the Mahalanobis' distance of the two data is larger than the critical value decided in advance, the reliability of the measurement is doubtful. The characteristic of the present method is that it can apply to data with missing values by estimating them from measured data. Some numerical examples are shown to demonstrate the availability of the method.

Algorithms↗

Methodology for quality-of-life assessment: a critical appraisal.

The methodology for QOL assessment covers a wide range of topics. It involves a proper choice of instruments with appropriate psychometric properties, the administration of these instruments, frequency of measurements, missing data problems, and the method of analysis. There are currently debates on the meaning and interpretation of the HRQOL domains taking the form of arguing how to define minimal clinically meaningful difference and whether this can be used in regulatory approval for drug development. From a practical point of view, the authors proposed that a disease-specific checklist or symptom domains incorporated within a HRQOL questionnaire may be a middle ground to gain general agreements among academic institutions, manufacturers, and regulatory agencies to use a specific symptom checklist or domain as the primary end point for clinical trials together with other HRQOL domains as ancillary data for the study. Antiemetic trial with HRQOL assessments is an example. Most would agree, however, that no matter what HRQOL domains or symptoms are being studied, it should be based on a patient self-administered questionnaire as shown by the lack of sensitivity in the example in this article. Missing data are a problem in the data collection and handling. The authors have examined a few commonly used approaches and performed simulation to study their properties. The subscale-mean method when one has more than 50% of the information on a subscale generally reflects the true values. In practice, one still would have missing data that cannot be handled completely by imputation. The method of analysis must be flexible enough to incorporate the nature of these data. Two approaches have been discussed, and they are both flexible in terms of using all available information being obtained in a longitudinal fashion with variable visiting schedules and potential missing data. The HRQOL response variable approach is simple and easy to understand. The growth curve models approach provides more detailed information on average trends between treatment arms. In general, these two methods agree on the results of the example. They can be used to report clinical trial results using HRQOL data as end points.

Health Status↗

The long-term course of low-density lipoprotein cholesterol after initiation of statin treatment: retrospective database analysis over 3 years in health maintenance organization enrollees.

OBJECTIVES: Our primary objective was to obtain robust estimates of the low-density lipoprotein cholesterol (LDL-C) decrease from baseline over a long period (ie, 3 years) after initiation of statin treatment in a usual-care setting. Our secondary objective was to investigate the predictors of the LDL-C time course. METHODS: We retrospectively analyzed the data for a sample of enrollees in a health maintenance organization (HMO) who started statin treatment between October 1, 1995, and December 31, 1998. Using the HMO's claims database, we examined the LDL-C change from baseline (as measured at the prescribing physicians' discretion) and computed mean estimates every 6 months up to 3 years. We investigated potential predictors of the LDL-C time course (ie, age, sex, baseline LDL-C, previous treatment, prescribing physician's specialty, and most recent treatment) with a mixed model applied to longitudinal data. This model enabled us to impute missing data for all enrollees still being followed, including those who had stopped treatment, and to discuss the robustness of our findings. We applied 2 methods of imputation, assuming either of the following: (1) data were missing at random but could be estimated from the parameters in the mixed model, or (2) LDL-C returned to the baseline value > or =15 days after treatment cessation. RESULTS: We examined data from 3193 individuals. In most cases, the statin used was fluvastatin or pravastatin. The observed mean (95% CI) LDL-C decrease from baseline widened progressively from 23.6% (23.0%-24.3%) at 6 months to 28.0% (27.1%-28.9%) at 18 months and 30.2% (28.7%-31.7%) at 36 months after treatment initiation. These results remained robust after the imputation of missing data, with mean LDL-C reduction consistently >20% at each 6-month time point during the 3 years after treatment initiation. Variations as a function of baseline characteristics were limited (demographics) or explicable by extraneous factors (baseline LDL-C). There were predictable variations as a function of the most recent treatment. CONCLUSIONS: This analysis indicates a long-term reduction in LDL-C among a sample of HMO enrollees who initiated statin treatment in a usual-care setting. The results were robust after imputation of missing data, with mean decrease from baseline consistently >20% over 3 years. However, given the retrospective design of our study and the absence of a control group, we cannot determine how much of the decrease was attributable to treatment.

Aged↗

Evaluating Medicaid HMOs when encounter data are missing: case of developmentally delayed children.

In evaluating Medicaid Health Maintenance Organizations (HMOs), crucial information regarding severity of illness of patients is often missing--in part because encounter data are not available. If we assume that patients are either in the HMO or in fee-for-service (FFS) plans (i.e., no in or out migration); then severity of HMO patients can be deduced from encounters of FFS patients. We applied this approach to effectiveness of HMO services for developmentally delayed children. Data supported the assumption of a closed system. Data also showed that over 12 months, severity of FFS patients declined. Therefore, we inferred that the HMO was attracting sicker patients. The HMO was paid less than FFS plan, despite the fact that it attracted sicker patients.

Child↗

Culling before testing in swine: identification of culling strategy and estimation of culling precision.

The aim of this simulation study was to identify culling strategy and to estimate culling precision based on various characteristics available in field data in order to evaluate the ability to detect situations in which adjustment for missing data should be applied in genetic evaluation. Data were simulated for age at 100 kg of live weight (AGE) measured on the farm. Culling was done within (C-W/IN) or over (C-OVER) litters by deleting records from the simulated datasets with culling intensities of .33 and .67. The culling variate (CVAR) used indicated the culling precision and had genetic and phenotypic correlations of 1.00, .75, .50, .25, or .00 with AGE (r(CVAR,AGE)). We were able to distinguish between culling strategies C-OVER and C-W/IN by means of decision rules based on proportion of tested animals per litter. Estimates of r(CVAR,AGE) were obtained from calibration curves for linear regression coefficients of litter average or within-litter variance for AGE on proportion of tested animals, and within- and between-litter variance (V(W) and V(B)) for AGE. Moderate to high r(CVAR,AGE) could be identified with little error by using V(W) or V(B) in C-W/IN and V(W) in C-OVER. Within-litter variance and the weighted average of the estimates from all four characteristics were well able to detect r(CVAR,AGE) values of .50 and higher in both C-W/IN and C-OVER. In conclusion, characteristics of swine field data with missing observations contain information that makes it possible to determine culling strategy, intensity, and precision. This information can be used to decide whether missing data should be replaced by their expected values in genetic evaluation.

Animal Husbandry↗

Evaluation of the emergency department logbook for population-based surveillance of firearm-related injury.

STUDY OBJECTIVE: To evaluate existing emergency department logbooks as a source of population-based data on firearm-related injuries. METHODS: We examined the logbooks of the 24 acute care and specialty-hospital EDs in Allegheny County, Pennsylvania, to determine the number and type of data variables each contained and the completeness of reporting of each variable for selected firearm-related cases. The amount of missing data for certain variables was determined and the cause for the missing data described. RESULTS: Logbooks from 18 of the 24 eligible hospitals were reviewed. We identified 785 cases of firearm-related injury recorded between January 1, 1992, and December 31, 1993. Of the variables we selected for analysis, only date (100%), chief complaint or diagnosis (100%), name (98%), and time of admission (97%) were consistently documented. In 37% of cases the patient's county of residence could not be determined. Similarly incomplete data were found for body part injured (31%), race (28%), age (26%), sex (22%), and mode of arrival (21%). The factor most responsible for the high percentage of incomplete data was the considerable variation in the data elements contained in the different hospitals' logbooks. CONCLUSION: Missing data resulting from inconsistencies in the variables contained in different EDs' logbooks and errors of omission prevent ED logbooks, in their current state, from providing population-based data for surveillance of firearm-related injury. Standardization of such variables in ED logbooks would yield a more useful source of information for injury and disease surveillance. In lieu of standardized logbooks, multiple sources of data are necessary to establish a more comprehensive and useful system of surveillance of firearm-related injury.

Emergency Service, Hospital↗

Summary measures and statistics for comparison of quality of life in a clinical trial of cancer therapy.

Assessment of health related quality of life (QOL) has become an important endpoint in many clinical trials of cancer therapy. Most of these studies entail multiple QOL scales that are assessed repeatedly over time. As a result, the problem of multiple comparisons is a primary analytic challenge with these trials. The use of summary measures and statistics both reduces the number of hypotheses tested and facilitates the interpretation of trial results where the primary question is 'Does the overall QOL differ between treatment arms?' I present two classes of summary measures that are sensitive to consistent trends in the same direction across multiple assessment times or multiple QOL scales. Missing data strongly influences the choice between the two classes, where one class handles missing data on an individual basis, while the other class uses model-based strategies. I present the results from a clinical trial of adjuvant therapy for breast cancer that use summary measures with a focus on the practical issues that affect these analysis strategies, such as missing data and integration of QOL with efficacy endpoints such as survival.

Analysis of Variance↗

Comparison of data quality for reports and ratings of ambulatory care by African American and White Medicare managed care enrollees.

OBJECTIVE: Compare missing data and reliability of health care evaluations between African Americans and Whites in Medicare managed care health plans. METHOD: Consumer Assessment of Healthcare Providers and Systems (CAHPS) 3.0 health plan survey data collected from 109,980 Medicare managed care enrollees (101,189 Whites, 8,791 African Americans) in 321 plans. Participants self-administered the survey and four single-item global ratings of care. RESULTS: Missing data rates were significantly higher for African Americans than Whites on all CAHPS items (p < .0001). Internal consistency reliability estimates for the CAHPS scales did not differ significantly between African Americans and Whites, but plan-level reliability estimates for the scales and global rating items were significantly lower for African Americans than Whites. DISCUSSION: Higher missing data rates and lower plan-level reliability estimates for African American Medicare managed care enrollees suggest caution in making race/ethnicity comparisons. Future efforts are needed to enhance the quality of data collected from older African Americans.

Adult↗

Incorporation of potential for multimedia exposure into chemical hazard scores for pollution prevention.

We are reporting a chemical hazard score for pollution prevention, called the Purdue score. The Purdue score provides a relative quantitative measure combining a variety of chemical hazards into a single quantitative hazard weighting factor for the non-expert to use. The main expected uses are to design safer products, assist in implementing and measuring achievement in pollution prevention, and as an adjunct for reporting Toxic Release Inventory data to the U.S. Government. Scoring results are presented for 200 Superfund chemicals, rank ordered by the worker hazard part of the score, by the environmental hazard part, and by combined worker and environmental hazard scores. We have reviewed the extent to which the Purdue score presently incorporates potential for multimedia pathway and multiroute absorption exposure. Until other possible uses have been carefully tested, peer-reviewed and published, users are advised to limit use of this system to planning, implementing and measuring pollution prevention and to enhancing the interpretation of Toxic Release Inventory data. The objective of this report is to look at how the structure of this score handles exposure to chemicals, both via multi-compartment pathways and multi-routes for contact or absorption health damage, as well as how it handles habitat degradation by chemicals. For all of these, the approach is built on inherent properties of each chemical, which are true for all sites and scenarios. The biggest obstacle to scoring is lack of measured chemical property data needed for scoring. We handle missing data by regression, quantitative structure activity relationship estimations, and a missing data default rule. The limitations of chemical hazard scoring are reviewed. At present, there is no widely accepted single measure of relative chemical hazard, against which to calibrate this hazard score for accuracy, except experience from industrial use. However, despite limitations, we suggest there is a strong value added for industry and society in availability of a concise, simple-to-use measure of relative chemical hazard. The Purdue score enables separate or combined consideration of chemical hazard to workers and to the natural environment. The Purdue score has potential for major cost savings in relative hazard ranking and business decision making regarding little-studied organic chemicals, because of the extensive use of advanced property estimation software. We conclude that there is societal need to warrant advanced development of this risk management tool, which is now ready for pilot use by industry. The Purdue score is mainly intended to assist and encourage businesses to implement and measure pollution prevention-especially small businesses--in a cost-effective way. The Purdue score relies strongly on sublethal toxicity, and there is practical potential for it to be used with thousands of chemicals.

Animals↗

Riluzole for amyotrophic lateral sclerosis (ALS)/motor neuron disease (MND).

BACKGROUND: Riluzole has been approved for treatment of patients with amyotrophic lateral sclerosis (ALS) in some countries but not others. Questions persist about its clinical utility because of high cost, modest efficacy and concern over adverse effects. OBJECTIVES: To examine the efficacy of riluzole in prolonging survival, and in delaying the use of surrogates (tracheostomy and mechanical ventilation) to sustain survival. SEARCH STRATEGY: Search of the Cochrane Neuromuscular Disease Group Register for randomized trials and enquiry from authors of trials and other experts in the field. The most recent search was conducted in June 1999. SELECTION CRITERIA: Types of studies: randomized trials TYPES OF PARTICIPANTS: adults with a diagnosis of ALS Types of interventions: treatment with riluzole or placebo Types of outcome measures: Primary: per cent mortality at 12 months with riluzole 100 mg Secondary: per cent mortality as a function of time with 100 mg and with all doses of riluzole, scales of neurologic function, quality of life, muscle strength and adverse events. DATA COLLECTION AND ANALYSIS: We identified two randomized trials. Each reviewer graded them for methodological quality. Data extraction was performed by a single reviewer and checked by the other two. We obtained some missing data from investigators. We performed meta-analyses with RevMan software using a fixed effects model. MAIN RESULTS: The two eligible trials included a total of 794 riluzole treated patients and 320 placebo treated patients. The methodological quality was acceptable and the trials were easily comparable. There were significant differences between the riluzole and placebo groups of both trials, in terms of the primary outcome measure, which was per cent mortality at 12 months with the 100 mg dose of riluzole. The odds ratio for the combined studies was 0.57 (95%CI 0.41 to 0.80) at 12 months. In the secondary outcome measures, there was a survival advantage with riluzole 100 mg at six, nine, 12 and 15 months, but not at three or 18 months. Pooled data from the 50, 100 and 200mg dose groups in the larger trial showed a lower per cent mortality with riluzole compared to placebo only at 12 months (odds ratio (OR) 0.64, 95% CI 0.47 to 0.88). There was no beneficial effect on bulbar function, or muscle strength. There were scant data on quality of life, but patients treated with riluzole remained in a more moderately affected health state significantly longer than placebo-treated patients (weighted mean difference (WMD) 35.5 days, 95% CI 5.9 to 65. 0). A threefold increase in serum alanine transferase was more frequent in riluzole treated patients than controls (WMD 2.65, 95% CI 1.51 to 4.65). REVIEWER'S CONCLUSIONS: Riluzole 100 mg per day appears to be modestly effective in prolonging survival for patients with ALS.

Amyotrophic Lateral Sclerosis↗

Analysis of pregnancy and other factors on detection of human papilloma virus (HPV) infection using weighted estimating equations for follow-up data.

Generalized estimating equations have been well established to draw inference for the marginal mean from follow-up data. Many studies suffer from missing data that may result in biased parameter estimates if the data are not missing completely at random. Robins and co-workers proposed using weighted estimating equations (WEE) in estimating the mean structure if drop-out occurs missing at random. We illustrate the differences between the WEE and the commonly applied available case analysis in a simulation study. We apply the WEE and reanalyse data of a longitudinal study of pregnancy and human papilloma virus (HPV) infection. We estimate the response probabilities and demonstrate that the data are not missing completely at random. Upon use of the WEE, we are able to show that pregnant women have an increased odds for an HPV infection compared with non-pregnant women after delivery (p=0.027). We conclude that the WEE are useful for dealing with monotone missing data due to drop-outs in follow-up data.

Cohort Studies↗