PubMed HealthSearch

PubMed · 3168432

Hierarchical time-oriented approaches to missing data inference.

Abstract

In practice clinical data are nearly always incomplete. When confronted with such data, a physician or investigator must make inferences about missing information. Possible strategies for inference include (1) interpolation, (2) extrapolation, (3) repeating the nearest value, (4) repeating the previous value, (5) patient-specific mean values, (6) patient-specific linear regression over time, (7) disease-specific mean values, (8) normal values, and (9) linear regression of correlated co-recorded variables. This study analyzes these strategies in a time-oriented data bank of patients with systemic lupus erythematosus, demonstrating that more accurate inferences of missing data are obtained when (1) strategies are tailored to the characteristics of the individual variable, (2) time-oriented strategies (e.g., interpolation) rather than non-time-oriented strategies (e.g., disease mean) are incorporated, (3) a ranked set of strategies is incorporated in a hierarchical stepwise fashion, and (4) the degree to which missing data are "nonrandomly" missing is assessed to allow estimation of bias. Interpolation is the best single technique with these data while linear regression of correlated co-recorded variables is a relatively weak technique. Inferences made by these hierarchical time-oriented approaches show significantly smaller mean differences from the actual values than do results from typical statistical package strategies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

K M Albridge, J Standish, J F Fries. 1988. Hierarchical time-oriented approaches to missing data inference.. https://doi.org/10.1016/0010-4809(88)90050-x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Internal validity analysis: a method for adjusting capture-recapture estimates of prevalence.

The authors propose a method for adjusting results of log-linear multi-source capture-recapture estimates of total population. The method compares the totals in some subpopulations of known size with estimates derived from various capture-recapture approaches to these subpopulations. The authors term such an approach an "internal validity analysis". Trends in the ratios of the estimates to the known true values of these subpopulations provide a plausible indicator of the bias of some types of estimates of the total population especially when underlying assumptions of the methods used have not been met in analysis of the total population. The authors apply this method to published data on an open population of injection drug users that had been previously analyzed with a standard capture-recapture analysis as if it were a closed population. Internal validity analysis suggests that the size of this population is about 15% greater than that previously estimated.

Data Interpretation, Statistical

Estimating sample sizes for binary, ordered categorical, and continuous outcomes in two group comparisons.

Sample size calculations are now mandatory for many research protocols, but the ones useful in common situations are not all easily accessible. This paper outlines the ways of calculating sample sizes in two group studies for binary, ordered categorical, and continuous outcomes. Formulas and worked examples are given. Maximum power is usually achieved by having equal numbers in the two groups. However, this is not always possible and calculations for unequal group sizes are given.

Data Interpretation, Statistical