PubMed Health⌕ Search

PubMed · 10673586

[Analysis of longitudinal Gaussian data with missing data on the response variable].

Abstract

BACKGROUND: Using an application and a simulation study we show the bias induced by missing data in the outcome in longitudinal studies and discuss suitable statistical methods according to the type of missing responses when the variable under study is gaussian. METHOD: The model used for the analysis of gaussian longitudinal data is the mixed effects linear model. When the probability of response does not depend on the missing values of the outcome and on the parameters of the linear model, missing data are ignorable, and parameters of the mixed effects linear model may be estimated by the maximum likelihood method with classical softwares. When the missing data are non ignorable, several methods have been proposed. We describe the method proposed by Diggle and Kenward (1994) (DK method) for which a software is available. This model consists in the combination of a linear mixed effects model for the outcome variable and a logistic model for the probability of response which depends on the outcome variable. RESULTS: A simulation study shows the efficacy of this method and its limits when the data are not normal. In this case, estimators obtained by the DK approach may be more biased than estimators obtained under the hypothesis of ignorable missing data even if the data are non ignorable. Data of the Paquid cohort about the evolution of the scores to a neuropsychological test among elderly subjects show the bias of a naive analysis using all available data. Although missing responses are not ignorable in this study, estimates of the linear mixed effects model are not very different using the DK approach and the hypothesis of ignorable missing data. CONCLUSION: Statistical methods for longitudinal data including non ignorable missing responses are sensitive to hypotheses difficult to verify. Thus, it will be better in practical applications to perform an analysis under the hypothesis of ignorable missing responses and compare the results obtained with several approaches for non ignorable missing data. However, such a strategy requires development of new softwares.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

H Jacqmin-Gadda, D Commenges, J Dartigues. 1999. [Analysis of longitudinal Gaussian data with missing data on the response variable].. https://pubmed.ncbi.nlm.nih.gov/10673586/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Analysis of ulcer data using hierarchical generalized linear models.

In multi-centre clinical trials, heterogeneities in individual hospital treatment effects can be modelled as random effects. Estimates of the individual hospital treatment effects and estimate of the mean treatment effect, allowing for the presence of overall hospital differences, are required, together with some measure of their uncertainty. Systematic inferences from the hierarchical-likelihood are now possible, using hierarchical generalized linear models. We show how to construct profile likelihoods for the treatment effects of individual hospitals.

Data Interpretation, Statistical↗

Statistical power and measurement allocation in ergonomic intervention studies assessing upper trapezius EMG amplitude. A case study of assembly work.

The present study aimed at exploring the statistical power of ergonomic intervention studies using electromyography (EMG) from the upper trapezius muscle. Data from a previous study of cyclic assembly work were reanalyzed with respect to exposure variability between subjects, between days, and within days. On basis of this information, the precision and power of different data collection strategies were explored. A sampling strategy comprising four registrations of about two min each (i.e. two work cycles) for one day per subject resulted in coefficients of variation between subjects on the 10-, 50-, and 90-APDF-percentiles of 0.44, 0.31, and 0.29, respectively. The corresponding necessary numbers of subjects in a study aiming at detecting a 20% exposure difference between two independent groups of equal size were 154, 78, and 68, respectively (p< or = 0.05, power 0.80). Multiple measurement days per subject would improve power, but only to a marginal extent beyond 4 days of recording. Increasing the number of recordings per day would have minor effects. Bootstrap resampling of the data set revealed that estimates of variability and power were associated with considerable uncertainty. The present results in combination with an overview of other occupational studies showed that common-size investigations using trapezius EMG percentiles are at great risk of suffering from insufficient statistical power, even if the expected intervention effect is substantial. The paper suggests a procedure of how to retrieve and use exposure variability information as an aid when studies are planned, and how to allocate measurements efficiently.

Data Interpretation, Statistical↗