PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Mediational analysis in HIV/AIDS research: estimating multivariate path analytic models in a structural equation modeling framework.

Mediational analyses have been recognized as useful in answering two broad questions that arise in HIV/AIDS research, those of theoretical model testing and of the effectiveness of multicomponent interventions. This article serves as a primer for those wishing to use mediation techniques in their own research, with a specific focus on mediation applied in the context of path analysis within a structural equation modeling (SEM) framework. Mediational analyses and the SEM framework are reviewed at a general level, followed by a discussion of the techniques as applied to complex research designs, such as models with multiple mediators, multilevel or longitudinal data, categorical outcomes, and problematic data (e.g., missing data, nonnormally distributed variables). Issues of statistical power and of testing the significance of the mediated effect are also discussed. Concrete examples that include computer syntax and output are provided to demonstrate the application of these techniques to testing a theoretical model and to the evaluation of a multicomponent intervention.

Data Interpretation, Statistical↗

Autoregressive spectral models of heart rate variability. Practical issues.

Autoregressive time series model-based spectral estimates of heart period sequences can provide a parsimonious and visually attractive representation of the dynamics of interbeat intervals. While a corollary to Wold's decomposition theorem implies that the discrete Fourier periodogram spectral estimate and the autoregressive spectral estimate converge asymptotically, there are practical differences between the two approaches when applied to short blocks of data. Autoregressive spectra can achieve good frequency resolution and excellent statistical stability on short segments of heart period data of sinus origin. However, the order of the autoregressive model (number of free parameters to be estimated) must be explicitly chosen, a decision that influences the trade-off of frequency resolution with statistical stability. Akaike's Information Criterion (AIC), an information-theoretic rule for picking the optimum order, is sensitive to the aggregate amount of data in the analysis. Thus, the best model order for estimating the spectrum of a 4-minute segment of data will generally be lower than the best order for estimating an hourly spectrum based on averaging 15 4-minute spectra. A major advantage of the autoregressive model approach to spectral analysis is the ease with which it can be extended to handle messy data frequently seen in heart rate variability studies. A number of autoregressive-based robust-resistant techniques are available for the analysis of heart period sequences that contain a high volume of nonsinus and other unusual beats intervals. A theoretically satisfying framework is also available for spectral analysis of unevenly sampled data and missing data.

Fourier Analysis↗

Consent in resuscitation trials: benefit or harm for patients and society?

CONTEXT: Research in an emergency setting is challenging because there may not be sufficient opportunity or time to obtain informed consent from the patient or their legally authorized representative. Such research can be conducted without prior consent if specific criteria are met. However consent is sometimes required for continued participation and may bias the results of the study. OBJECTIVE: To review regulations related to waiver of consent in emergency research, and evidence of whether such regulations introduce bias. RESULTS: Emergency research can be conducted without consent, either through community disclosure and consultation followed by patient or family notification and consent for continued participation after the intervention was applied, or under a minimal risk waiver. Review of the clinical record is necessary to determine important outcomes such as survival to discharge. If consent is required for this review but not granted, then these data are missing during analysis. If seriously ill or disadvantaged patients are less likely to assent, then investigators cannot determine reliably whether these vulnerable patients were harmed by the intervention. If missing data are different from complete data, then the analysis is susceptible to bias, and the conclusions could be misleading. Extrapolation from non-consent rates in resuscitation studies to results from the DAVID trial demonstrates that the rate of absence of data and information due to lack of assent can influence whether there is a significant difference between treatment groups (survival of control versus intervention: p=0.04 for complete data; p=0.08 for 10.8% lack of assent; p=0.40 for 19.7% lack of assent). CONCLUSIONS: Exception from consent for emergency research should extend to review of the hospital record as the standard in emergency research. The only potential risk to patients associated with review of the clinical record after the intervention is loss of privacy and confidentiality. Appropriate safeguards can be taken to minimize this risk.

Clinical Protocols↗

Evaluation of audit-based performance measures for dental care plans.

OBJECTIVES: Although a set of clinical performance measures, i.e., a report card for dental plans, has been designed for use with administrative data, most plans do not have administrative data systems containing the data needed to calculate the measures. Therefore, we evaluated the use of a set of proxy clinical performance measures calculated from data obtained through chart audits. METHODS: Chart audits were conducted in seven dental programs--three public health clinics, two dental health maintenance organizations (DHMO), and two preferred provider organizations (PPO). In all instances audits were completed by clinical staff who had been trained using telephone consultation and a self-instructional audit manual. The performance measures were calculated for the seven programs, audit reliability was assessed in four programs, and for one program the audit-based proxy measures were compared to the measures calculated using administrative data. RESULTS: The audit-based measures were sensitive to known differences in program performance. The chart audit procedures yielded reasonably reliable data. However, missing data in patient charts rendered the calculation of some measures problematic--namely, caries and periodontal disease assessment and experience. Agreement between administrative and audit-based measures was good for most, but not all, measures in one program. CONCLUSIONS: The audit-based proxy measures represent a complex but feasible approach to the calculation of performance measures for those programs lacking robust administrative data systems. However, until charts contain more complete diagnostic information (i.e., periodontal charting and diagnostic codes or reason-for-treatment codes), accurate determination of these aspects of clinical performance will be difficult.

Adolescent↗

A comparison of health utility measures for the evaluation of multiple sclerosis treatments.

OBJECTIVES: To evaluate the practical application and psychometric properties of three health utility measures in a sample of MS patients with a broad range of neurological disability as measured by the Extended Disability Status Scale (EDSS). METHODS: Patients randomly selected from two MS clinic registries were assessed using standard clinical methods and completed three generic measures of health utility (EQ-5D, HUI Mark III, SF-6D). The proportion of missing data, test/retest reliability, and construct validity of each health utility measure were examined. RESULTS: The assessments were completed by 187 patients. Less than 10% of data were missing for the subscales of the SF-6D (< 3.2%), HUI Mark III (<1.6%), and EQ-5D (< or =7.5%). Severely disabled patients were more likely to omit physical function questions for the SF-6D (20%), and EQ-5D (43%). Retest reliability for the SF-6D (ICC = 0.83), EQ-5D (ICC = 0.81), and HUI Mark III (ICC = 0.87) were adequate for population surveys. Correlations between assessment of clinical function and each health utility measure were strongest for the HUI Mark III (HUI Mark III EDSS rho = -0.77, HUI Mark III ambulation index rho = -0.76, HUI Mark III timed 25 foot walk rho = -0.73, HUI Mark III nine hole peg test rho = -0.65). CONCLUSIONS: The health utility measures were generally feasible and reliable but the HUI Mark III demonstrated highest concordance with the EDSS across the full range of neurological disability. Of the three measures studied, the HUI Mark III may be the most appropriate for cost effectiveness evaluations of MS therapies.

Adjuvants, Immunologic↗

Genome-wide linkage analysis of the tracking of systolic blood pressure using a mixed model.

BACKGROUND: Elevated blood pressure in middle age is a major risk factor for subsequent cardiovascular complications. An important longitudinal characteristic of blood pressure is the "tracking phenomenon". Tracking is defined as the persistence of the rank of a person's blood pressure level in a group over a long period of time. In this analysis, we used the Framingham data to investigate whether there are some genes responsible for this phenomenon. RESULTS: Both two-point and multipoint linkage analyses were applied to family members with complete data only and to all family data with missing values imputed by a Gaussian model. The results of two-point linkage analysis indicated that two loci for linkage with the intercept were on chromosomes 10 and 13, and two loci for linkage with both slope and intercept were on chromosomes 1 and 3. Multipoint linkage analysis indicated only one region, 200-240 cM on chromosome 1, to be linked with both intercept and slope. For the intercept of SBP, the highest LOD (4.43) was found at 214 cM when missing data were imputed, and the highest LOD (2.81) was at 231 cM for the complete case data. For the slope of SBP, the highest multipoint LODs were 3.63 at 227 cM and 2.02 at 234 cM for the complete case data and imputation data, respectively. CONCLUSION: One or more genes in the range of 200-240 cM on chromosome 1 may be related to the tracking phenomenon of SBP.

Adult↗

Vaginal misoprostol for cervical ripening and labour induction in late pregnancy.

BACKGROUND: Although not yet registered for such use, misoprostol has been widely used for obstetric and gynaecological indications, such as induction of abortion and of labour. OBJECTIVES: The objective of this review was to assess the effects of vaginal misoprostol for third trimester cervical ripening or induction of labour. SEARCH STRATEGY: The Cochrane Pregnancy and Childbirth Group trials register, the Cochrane Controlled Trials Register and reference lists of articles were searched. SELECTION CRITERIA: Randomised trials comparing vaginal misoprostol with other methods of cervical ripening or labour induction, placebo or no treatment in women due for third trimester induction of labour. DATA COLLECTION AND ANALYSIS: Trial quality assessment and data extraction were done by both reviewers. MAIN RESULTS: Twenty-six studies were included. Compared to placebo, misoprostol was associated with increased cervical ripening (relative risk of unfavourable or unchanged cervix after 12 to 24 hours with misoprostol 0.20, 95% confidence interval 0.07 to 0.61). It was also associated with a reduced need for oxytocin (relative risk 0.47, 95% confidence interval 0.23 to 0.96). Misoprostol was more effective than prostaglandin E2 for labour induction (relative risk of failure to achieve vaginal delivery in 24 hours 0.70, 95% confidence interval 0.62 to 0.79). Oxytocin augmentation was used less often with misoprostol than with prostaglandin E2 (relative risk 0.64, 95% confidence interval 0.58 to 0.71). Uterine hyperstimulation and meconium stained liquor were more common with misoprostol than with prostaglandin E2. Lower doses of misoprostol compared to higher doses did not show significant differences except for more need for oxytocin augmentation and less uterine hyperstimulation without fetal heart rate changes. REVIEWER'S CONCLUSIONS: Vaginal misoprostol appears to be more effective in inducing labour than conventional methods of cervical ripening and labour induction. The apparent increase in uterine hyperstimulation is of concern. The studies were not large enough to exclude the possibility of rare but serious adverse effects.

Cervical Ripening↗

Review: a gentle introduction to imputation of missing values.

In most situations, simple techniques for handling missing data (such as complete case analysis, overall mean imputation, and the missing-indicator method) produce biased results, whereas imputation techniques yield valid results without complicating the analysis once the imputations are carried out. Imputation techniques are based on the idea that any subject in a study sample can be replaced by a new randomly chosen subject from the same source population. Imputation of missing data on a variable is replacing that missing by a value that is drawn from an estimate of the distribution of this variable. In single imputation, only one estimate is used. In multiple imputation, various estimates are used, reflecting the uncertainty in the estimation of this distribution. Under the general conditions of so-called missing at random and missing completely at random, both single and multiple imputations result in unbiased estimates of study associations. But single imputation results in too small estimated standard errors, whereas multiple imputation results in correctly estimated standard errors and confidence intervals. In this article we explain why all this is the case, and use a simple simulation study to demonstrate our explanations. We also explain and illustrate why two frequently used methods to handle missing data, i.e., overall mean imputation and the missing-indicator method, almost always result in biased estimates.

Bias↗

Strategies for the analysis of imputed data from a sample survey. The National Medical Care Utilization and Expenditure Survey.

Missing data in sample surveys is virtually unavoidable, whether it is an entire unit that is missing or only an item for a responding unit. Compensation for unit nonresponse is usually made through the assignments of weights to responding units; for item nonresponse, the compensation often is by an imputation procedure. This paper reviews the extent of missing data in a large federal survey, the National Medical Care Utilization and Expenditure Survey, and the imputation procedures used to compensate for item missing data. The effects of imputation on several types of estimates from the survey are examined. In addition, several methods for analyzing survey data with imputed values are reviewed, and recommendations about preferred strategies are made for selected circumstances.

Data Collection↗

Pragmatic randomised controlled trials in parenting research: the issue of intention to treat.

STUDY OBJECTIVE: To evaluate trials of parenting programmes, regarding their use of intention to treat (ITT). DESIGN: Individual trials included in two relevant Cochrane systematic reviews were scrutinised by two independent reviewers. Data on country of origin, target audience, trial type, treatment violations, use of ITT, and the management of missing data were extracted. MAIN RESULTS: Thirty trial reports were reviewed. Three reported the use of an ITT approach to data analysis. Nineteen reported losing subjects to follow up although the implications of this were rarely considered. Insufficient detail in reports meant it was difficult to identify study drop outs, the nature of treatment violations, and those failing to provide outcome assessments. In two trials, study drop outs were considered as additional control groups, violating the basic principle of ITT. CONCLUSIONS: It is recommended that future trial reports adhere to CONSORT guidelines. In particular ITT should be used for the main analyses, with strategies for managing treatment violations and handling missing data being reported a priori. Those conducting trials need to acknowledge the social nature of these programmes can sometimes result in erratic parent attendance and participation, which would only increase the chances of missing data. The use of approaches that can limit the proportion of missing data is therefore recommended.

Humans↗

Inductive learning of thyroid functional states using the ID3 algorithm. The effect of poor examples on the learning result.

The ID3 algorithm for inductive learning was tested using preclassified material for patients suspected to have a thyroid illness. Classification followed a rule-based expert system for the diagnosis of thyroid function. Thus, the knowledge to be learned was limited to the rules existing in the knowledge base of that expert system. The learning capability of the ID3 algorithm was tested with an unselected learning material (with some inherent missing data) and with a selected learning material (no missing data). The selected learning material was a subgroup which formed a part of the unselected learning material. When the number of learning cases was increased, the accuracy of the program improved. When the learning material was large enough, an increase in the learning material did not improve the results further. A better learning result was achieved with the selected learning material not including missing data as compared to unselected learning material. With this material we demonstrate a weakness in the ID3 algorithm: it can not find available information from good example cases if we add poor examples to the data.

Algorithms↗

A computer program for multivariate ratio analysis (MISCAT).

Analysts must deal frequently with missing data in multivariate analysis. In such cases, estimating the covariance maxtrix V of the dependent variables usually involves initial estimation and iterative adjustment of imputed missing data values, and/or smoothing of an estimate V which is not necessarily positive semi-definite. This paper presents an alternative procedure for computing estimates of relevant multivariate parameters in situations where missing data occur at random and with small probability. MISCAT is a computer program which computes multivariate ratio estimates of the means and a corresponding positive semi-definite estimate of the covariance matrix. It is an extension of GENCAT, which is a program for the generalizaed least squares analysis of categorical data. Thus, one advantage of dealing with missing data in this manner is that variation among the ratio estimates may be conveniently analyzed within MISCAT using asymptotic regression methodology, provided that sample sizes are sufficiently large. An example is given to illustrate such analysis for longitudinal data from a multicenter clinical trial.

Computers↗

Multiple imputation for multivariate data with missing and below-threshold measurements: time-series concentrations of pollutants in the Arctic.

Many chemical and environmental data sets are complicated by the existence of fully missing values or censored values known to lie below detection thresholds. For example, week-long samples of airborne particulate matter were obtained at Alert, NWT, Canada, between 1980 and 1991, where some of the concentrations of 24 particulate constituents were coarsened in the sense of being either fully missing or below detection limits. To facilitate scientific analysis, it is appealing to create complete data by filling in missing values so that standard complete-data methods can be applied. We briefly review commonly used strategies for handling missing values and focus on the multiple-imputation approach, which generally leads to valid inferences when faced with missing data. Three statistical models are developed for multiply imputing the missing values of airborne particulate matter. We expect that these models are useful for creating multiple imputations in a variety of incomplete multivariate time series data sets.

Air Pollutants↗

A systematic review identifies a lack of standardization in methods for handling missing variance data.

BACKGROUND AND OBJECTIVES: To describe and critically appraise available methods for handling missing variance data in meta-analysis (MA). METHODS: Systematic review. MEDLINE, EMBASE, Web of Science, MathSciNet, Current Index to Statistics, BMJ SearchAll, The Cochrane Library and Cochrance Colloquium proceedings, MA texts and references were searched. Any form of text was included: MA, method chapter, or otherwise. Descriptions of how to implement each method, the theoretic basis and/or ad hoc motivation(s), and the input and output variable(s) were extracted and assessed. Methods may be: true imputations, methods that obviate the need for a standard deviation (SD), or methods that recalculate the SD. RESULTS: Eight classes of methods were identified: algebraic recalculations, approximate algebraic recalculations, imputed study-level SDs, imputed study-level SDs from nonparametric summaries, imputed study-level correlations (e.g., for change-from-baseline SD), imputed MA-level effect sizes, MA-level tests, and no-impute methods. CONCLUSION: This work aggregates the ideas of many investigators. The abundance of methods suggests a lack of consistency within the systematic review community. Appropriate use of methods is sometimes suspect; consulting a statistician, early in the review process, is recommended. Further work is required to optimize method choice to alleviate any potential for bias and improve accuracy. Improved reporting is also encouraged.

Data Interpretation, Statistical↗

Multiple imputation for body mass index: lessons from the Australian Longitudinal Study on Women's Health.

In large epidemiological studies missing data can be a problem, especially if information is sought on a sensitive topic or when a composite measure is calculated from several variables each affected by missing values. Multiple imputation is the method of choice for 'filling in' missing data based on associations among variables. Using an example about body mass index from the Australian Longitudinal Study on Women's Health, we identify a subset of variables that are particularly useful for imputing values for the target variables. Then we illustrate two uses of multiple imputation. The first is to examine and correct for bias when data are not missing completely at random. The second is to impute missing values for an important covariate; in this case omission from the imputation process of variables to be used in the analysis may introduce bias. We conclude with several recommendations for handling issues of missing data.

Australia↗

Bayesian modeling of air pollution health effects with missing exposure data.

The authors propose a new statistical procedure that utilizes measurement error models to estimate missing exposure data in health effects assessment. The method detailed in this paper follows a Bayesian framework that allows estimation of various parameters of the model in the presence of missing covariates in an informative way. The authors apply this methodology to study the effect of household-level long-term air pollution exposures on lung function for subjects from the Southern California Children's Health Study pilot project, conducted in the year 2000. Specifically, they propose techniques to examine the long-term effects of nitrogen dioxide (NO2) exposure on children's lung function for persons living in 11 southern California communities. The effect of nitrogen dioxide exposure on various measures of lung function was examined, but, similar to many air pollution studies, no completely accurate measure of household-level long-term nitrogen dioxide exposure was available. Rather, community-level nitrogen dioxide was measured continuously over many years, but household-level nitrogen dioxide exposure was measured only during two 2-week periods, one period in the summer and one period in the winter. From these incomplete measures, long-term nitrogen dioxide exposure and its effect on health must be inferred. Results show that the method improves estimates when compared with standard frequentist approaches.

Air Pollutants↗

A comparison of the random-effects pattern mixture model with last-observation-carried-forward (LOCF) analysis in longitudinal clinical trials with dropouts.

The last-observation-carried-forward imputation method is commonly used for imputting data missing due to dropouts in longitudinal clinical trials. The method assumes that outcome remains constant at the last observed value after dropout, which is unlikely in many clinical trials. Recently, random-effects regression models have become popular for analysis of longitudinal clinical trial data with dropouts. However, inference obtained from random-effects regression models is valid when the missing-at-random dropout process is present. The random-effects pattern-mixture model, on the other hand, provides an approach that is valid under more general missingness mechanisms. In this article we describe the use of random-effects pattern-mixture models under different patterns for dropouts. First, subjects are divided into groups depending on their missing-data patterns, and then model parameters are estimated for each pattern. Finally, overall estimates are obtained by averaging over the missing-data patterns and corresponding standard errors are obtained using the delta method. A typical longitudinal clinical trial data set is used to illustrate and compare the above methods of data analyses in the presence of missing data due to dropouts.

Clinical Trials as Topic↗

Natural history of morbid obesity without surgical intervention.

BACKGROUND: To study the mortality among morbidly obese patients qualifying for bariatric surgery. Mortality from bariatric surgery for morbid obesity has been widely reported; however, little is known about the mortality in morbidly obese patients who defer surgery. METHODS: Consecutive patients evaluated for bariatric surgery with an initial encounter between 1997 and 2004 were identified. The Social Security Death Index and office records were used to identify mortality through 2006. We conducted telephone interviews to determine whether the 305 patients who did not undergo bariatric surgery at our institution had undergone the surgery elsewhere. Using Cox proportional hazards models, we compared the mortality in patients undergoing surgery with that of those who did not. To evaluate bias resulting from missing data, we conducted analyses assuming that all patients with missing data had (1) undergone surgery and (2) not undergone surgery. RESULTS: A total of 908 patients underwent bariatric surgery (880 patients at our institution and 28 patients elsewhere). A total of 112 patients did not undergo surgery. Data regarding surgery on 165 patients could not be obtained. The mortality in those patients who did not undergo surgery was 14.3% compared with 2.9% for those who did undergo surgery. Adjusting for age, gender, and body mass index, patients who had undergone surgery had an 82% reduction in mortality (hazard ratio 0.18, 95% confidence interval 0.09-0.35, P <.0001). Sensitivity analysis, assuming that all patients with missing data received surgery resulted in an 85% mortality reduction (P <.001) and assuming that patients did not receive surgery resulted in a 50% mortality reduction (P = .04). CONCLUSIONS: Mortality among morbidly obese patients without surgery was 14.3% during the study period. Surgical intervention offered a 50%-85% mortality reduction benefit.

Adult↗