PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “longitudinal data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A transitional model for longitudinal binary data subject to nonignorable missing data.

Binary longitudinal data are often collected in clinical trials when interest is on assessing the effect of a treatment over time. Our application is a recent study of opiate addiction that examined the effect of a new treatment on repeated urine tests to assess opiate use over an extended follow-up. Drug addiction is episodic, and a new treatment may affect various features of the opiate-use process such as the proportion of positive urine tests over follow-up and the time to the first occurrence of a positive test. Complications in this trial were the large amounts of dropout and intermittent missing data and the large number of observations on each subject. We develop a transitional model for longitudinal binary data subject to nonignorable missing data and propose an EM algorithm for parameter estimation. We use the transitional model to derive summary measures of the opiate-use process that can be compared across treatment groups to assess treatment effect. Through analyses and simulations, we show the importance of properly accounting for the missing data mechanism when assessing the treatment effect in our example.

Algorithms↗

Comparison of methods for the analysis of longitudinal interval count data.

Longitudinal studies are often concerned with estimating the recurrence rate of a non-fatal event. In many cases, only the total number of events occurring during successive time intervals is known. We compared a mixed Poisson-gamma regression method proposed by Thall and a quasi-likelihood method proposed by Zeger and Liang for the analysis of such data, in the case where the mean was correctly specified, using simulation techniques with large samples. Both methods produced similar standard errors in most situations, except in the case of time-dependent covariates with non-Poisson-gamma data where they were seriously underestimated by the Thall method. A simple method for discriminating between the variance forms of the two methods is described. The findings are applied to the analyses of clinical trials of non-melanoma skin cancer and familial polyposis. This study extends the findings of Breslow concerning variance misspecification in overdispersed Poisson and quasi-likelihood models to the longitudinal setting.

Adenomatous Polyposis Coli↗

On the analysis of mixed longitudinal growth data.

Mixed longitudinal growth data consists of several observations on a characteristic over a limited age range for each individual in a study. This data is then combined to model growth over the total age range of all individuals in the study. The limited data collected on each individual precludes a subject-specific approach to modelling so that a population-based approach must be used. Here we propose a method for the analysis of mixed longitudinal data using linear models constructed from cubic splines to model both the mean and variance curves of an observed characteristic. The method is illustrated on growth in height and head circumference for children in a recently collected mixed longitudinal data set concerning the growth of Victorian schoolchildren, where particular concern was the timing of growth spurts and increases in variability.

Adolescent↗

Mass media messages and reproductive behaviour in Nigeria.

This paper examines the effects of exposure to mass media messages promoting family planning on the reproductive behaviour of married women in Nigeria using cross-sectional data. Longitudinal data are also used to ensure that exposure to media messages pre-dates the indicators of reproductive behaviour. Cross-sectional analysis suggests that: (1) contraceptive use and intention are positively associated with exposure to mass media messages, and (2) women who are exposed to media messages are more likely to desire fewer children than those who are not exposed to such messages. Similarly, analysis of the longitudinal data shows that exposure to mass media messages is a significant predictor of contraceptive use. Thus, exposure to mass media messages about family planning may be a powerful tool for influencing reproductive behaviour in Nigeria.

Adolescent↗

Marginal analysis of incomplete longitudinal binary data: a cautionary note on LOCF imputation.

In recent years there has been considerable research devoted to the development of methods for the analysis of incomplete data in longitudinal studies. Despite these advances, the methods used in practice have changed relatively little, particularly in the reporting of pharmaceutical trials. In this setting, perhaps the most widely adopted strategy for dealing with incomplete longitudinal data is imputation by the "last observation carried forward" (LOCF) approach, in which values for missing responses are imputed using observations from the most recently completed assessment. We examine the asymptotic and empirical bias, the empirical type I error rate, and the empirical coverage probability associated with estimators and tests of treatment effect based on the LOCF imputation strategy. We consider a setting involving longitudinal binary data with longitudinal analyses based on generalized estimating equations, and an analysis based simply on the response at the end of the scheduled follow-up. We find that for both of these approaches, imputation by LOCF can lead to substantial biases in estimators of treatment effects, the type I error rates of associated tests can be greatly inflated, and the coverage probability can be far from the nominal level. Alternative analyses based on all available data lead to estimators with comparatively small bias, and inverse probability weighted analyses yield consistent estimators subject to correct specification of the missing data process. We illustrate the differences between various methods of dealing with drop-outs using data from a study of smoking behavior.

Adolescent↗

Covariance components models for longitudinal family data.

A longitudinal family study is an epidemiological design that involves repeated measurements over time in a sample that includes families. Such studies, that may also include relative pairs and unrelated individuals, allow closer investigation of not only the factors that cause a disease to arise, but also the genetic and environmental determinants that modulate the subsequent progression of that disease. Knowledge of such determinants may pay high dividends in terms of prognostic assessment and in the development of new treatments that may be tailored to the prognostic profile of individual patients. Unfortunately longitudinal family studies are difficult to analyse. They conflate the complex within-family correlation structure of a cross-sectional family study with the correlation over time that is intrinsic to longitudinal repeated measures. Here we describe an approach to analysis that is relatively straightforward to implement, yet is flexible in its application. It represents a natural extension of a Gibbs-sampling-based approach to the analysis of cross-sectional family studies that we have described previously. The approach can be applied to pedigrees of arbitrary complexity. It is applicable to continuous traits, repeated binary disease states, and repeated counts or rates with a Poisson distribution. It not only supports the analysis of observed determinants, including measured genotypes, but also allows decomposition of the correlation structure, thereby permitting conclusions to be drawn about the effect of unobserved genes and environment on key features of disease progression, and hence to estimate the heritability of these features. We demonstrate the efficacy of our methods using a range of simulated data analyses, and illustrate its practical application to longitudinal blood pressure data measured in families from the Framingham Heart Study.

Blood Pressure↗

Prescription duration after drug copay changes in older people: methodological aspects.

OBJECTIVES: Impact assessment of drug benefits policies is a growing field of research that is increasingly relevant to health care planning for older people. Some cost-containment policies are thought to increase noncompliance. This paper examines mechanisms that can produce spurious reductions in drug utilization measures after drug policy changes when relying on pharmacy dispensing data. Reference pricing, a copayment for expensive medications above a fixed limit, for angiotensin-converting enzyme(ACE) inhibitors in older British Columbia residents, is used as a case example. DESIGN: Time series of 36 months of individual claims data. Longitudinal data analysis, adjusting for autoregressive data. SETTING: Pharmacare, the drug benefits program covering all patients aged 65 and older in the province of British Columbia, Canada. PARTICIPANTS: All noninstitutionalized Pharmacare beneficiaries aged 65 and older who used ACE inhibitors between 1995 and 1997 (N = 119,074). INTERVENTION: The introduction of reference drug pricing for ACE inhibitors for patients aged 65 and older. MEASUREMENTS: Timing and quantity of drug use from a claims database. RESULTS: We observed a transitional sharp decline of 110% t a standard error of 30% (P = .02) in the overall utilization rate of all ACE inhibitors after the policy implementation; five months later, utilization rates had increased, but remained under the predicted prepolicy trend. Coinciding with the sharp decrease, we observed a reduction in prescription duration by 31% in patients switching to no-cost drugs. This reduction may be attributed to increased monitoring for intolerance or treatment failure in switchers, which in turn led to a spurious reduction in total drug utilization. We ruled out the extension of medication use over the prescribed duration through reduced daily doses (prescription stretching) by a quantity-adjusted analysis of prescription duration. CONCLUSION: The analysis of prescription duration after drug policy interventions may provide alternative explanations to apparent short-term reductions in drug utilization and adds important insights to time trend analyses of drug utilization data in the evaluation of drug benefit policy changes.

Age Factors↗

Latent variable models for clustered ordinal data.

Existing methods for the analysis of clustered, ordinal data are inappropriate for certain applications. We propose latent variable models for clustered ordinal data which are derived as natural extensions of latent variable models for clustered binary data (Qu, Williams, Beck, and Medendorp, 1992. Biometrics 48, 1095-1102). These models can be applied to repeated measures data, familial data, longitudinal data, and data with both cluster specific and occasion specific covariates with a wide range of correlation structures.

Analysis of Variance↗

Assessing missing data assumptions in longitudinal studies: an example using a smoking cessation trial.

Due to the chaotic nature of the clinical disorder, longitudinal data analysis in substance abuse research is plagued by missing values. To obtain an unbiased estimation on intervention effects, different longitudinal modeling strategies require various assumptions on the patterns and mechanisms of missing data. By defining missingness as intermittent missingness (occasional omission) and dropout (premature withdrawal), this article demonstrates statistical ways for assessing missing data assumptions using evidence from a clinical trial. Within the framework of multiple imputation, intermittent missing data are imputed first so that dropouts can be isolated and treated specifically. A computational tool called "pattern reduction resampling" is proposed to simplify missing data methods when the number of intra-subject repeated measures is large. To test whether missingness patterns are nondifferential across treatment conditions, a formal testing approach treats indicators of missingness as a special type of repeated measures (e.g., 0: intermittent missing, 1: observed, and 2: dropout missing). After reviewing the idea of ignorability for missing data and of classifying missingness mechanisms into subcategories, the article provides an example for assessing common assumptions on missingness mechanisms and how these assumptions affect model selection for significance testing. A carbon monoxide longitudinal data set in a smoking cessation study is used for illustration.

Bias↗

Random regression models for male and female fertility evaluation using longitudinal binary data.

A longitudinal Bayesian threshold analysis of insemination outcomes was carried out using 2 random regression models with 3 (Model 1) and 5 (Model 2) parameters to model the additive genetic values at the liability scale. All insemination events of first-parity Holstein cows were used. The outcome of an insemination event was treated as a binary response of either a success (1) or a failure (0). Thus, all breeding information for a cow, including all service sires, was included, thereby allowing for a joint evaluation of male and female fertility. An edited data set of 369,353 insemination records from 210,373 first-lactation cows was used. On the liability scale, both models included the systematic effects of herd-year, month of insemination, technician, and regressions on age of service sire and milk yield during the first 100 d of lactation. The random effects in the model were the 3 or 5 random regression coefficients specific to each cow, the permanent effect of the cow, and the service sire effect. Using Model 1, the estimated heritability of an insemination outcome decreased from 0.035 at d 50 to 0.032 at d 140 and then increased continuously with DIM. The genetic correlations for insemination success at different time points ranged from 0.83 to 0.99, and their magnitude decreased with an increase in the interval between inseminations. A similar trend was observed for heritability and genetic correlations using Model 2. However, the average estimate of heritability was much higher (0.058) than those obtained using Model 1 or a repeatability model. In addition, the estimated genetic correlations followed the same trend as Model 1, but were lower and with a higher rate of decrease when the interval between inseminations increased. The posterior mean of service sire variance was 0.01 for both models, and permanent environmental variance was 0.05 and 0.02 for Models 1 and 2, respectively. Model comparison based on the Bayes factor indicated that Model 1 was more plausible, given the data.

Animals↗

Genetic analysis of male and female fertility using longitudinal binary data.

A longitudinal Bayesian threshold analysis of insemination events during the first 250 d after calving of first-parity Holsteins was carried out. The outcome of an insemination event was treated as a binary response of either a success (1) or a failure (0). Thus, all breeding information for a cow, including all service sires, was included, thereby allowing for a joint evaluation of male and female fertility. An edited data set of 297,823 insemination records from 151,758 first lactation cows was used. On the liability scale, the model included the systematic effects of herd-year of insemination, technician, month of insemination, and regressions on age of service sire, 3 test days in the first 100 d of lactation (early milk yield), and days in milk at insemination. The random effects in the model were the additive breeding value, the permanent effect of the cow, and the service sire effect. Posterior mean (standard deviation) of the dispersion parameters in the model were 0.034 (0.006), 0.009 (0.001), and 0.171 (0.013) for the additive, service sire, and permanent environmental variances, respectively. The residual variance was fixed at 1, as a result of the nonidentifiability of the threshold model. The posterior mean (standard deviation) of heritability was 0.028 (0.005). This point estimate of heritability is well within the range of available estimates for the trait. Thus, these estimates suggest that some genetic variation exists that can potentially be used to improve reproductive performance or at least avoid its further deterioration. The estimate of the regression coefficient on age of service sire was 0.001, indicating better fertility among older bulls. However, this result has to be interpreted with caution given the preferential use of proven bulls on well-managed cows (as opposed to problem breeders). The estimate of the regression coefficient was negative (-0.005) for early milk yield, as expected, and positive (0.003) for days in milk at insemination. This suggests that high-producing cows are less likely to conceive at the beginning of lactation.

Age Factors↗

A method of presenting longitudinal growth data.

1. Longitudinal growth profiles contain much information but are difficult to incorporate into mathematical and statistical analyses. 2. A growth function, which is a weighted average of growth achievement at different ages, is proposed. 3. This function is a non-dimensional number with defined statistical properties, and emphasizes growth achievement in early life. It can be used to compare the growth of individuals and populations.

Aging↗

Using administrative data for longitudinal research: comparisons with primary data collection.

This paper discusses the advantages and disadvantages of using administrative data for longitudinal research, focusing on loss to follow-up. Comparisons between research relying on primary data collection and that using data bases are made. After development of a suitable framework, follow-up in several well-known projects based on primary data collection (the Seven Countries project on coronary heart disease, the Massachusetts research on long-term care and the Pittsburgh clinical trial of tonsillectomy) is compared with follow-up using the Health Services Commission data base in Manitoba, Canada. Overall follow-up in the Manitoba research compares favorably with participation and follow-up rates in other studies based on primary data collection. Initial nonresponse and nonlocation are major problems with studies using primary data; failure to locate earlier respondents in subsequent waves results in a wide range of overall response rates. Data bases do not require researchers to contact individuals and hence follow-up is simplified. Eight year follow-up rates in the Manitoba data base are almost always over 80% and often over 90%. Because records can be flexibly summarized for each individual over time, data bases facilitate certain types of longitudinal studies which would be difficult, if not impossible, to perform using other methodologies. If the desired data are available and recorded with acceptable accuracy, administrative data banks hold considerable promise for the health care researcher.

Adolescent↗

The use of GEE for analyzing longitudinal binomial data: a primer using data from a tobacco intervention.

Longitudinal study designs in addictive behaviors research are common as researchers have focused increasingly on how various explanatory variables affect responses over time. In particular, such designs are used in intervention studies that have multiple follow-up points. These designs typically involve repeated measurement of participants' responses, and thus correlation within each participant is expected. Correct inferences can only be obtained by taking into account this within-participant correlation between repeated measurements, which can complicate the analysis of longitudinal data. In recent years, generalized estimating equations (GEE) has become a standard method for analyzing non-normal longitudinal data, yet it often is not utilized by addiction researchers. The goal of this article is to provide an overview of the GEE approach for analyzing correlated binary data for behavioral researchers, using data from an intervention study on the prevention of relapse to tobacco smoking.

Adult↗

Statistical analysis of longitudinal psychiatric data with dropouts.

Longitudinal studies are used in psychiatric research to address outcome changes over time within and between individuals. However, because participants may drop out of a study prematurely, ignoring the nature of dropout often leads to biased inference, and in turn, wrongful conclusions. The purpose of the present paper is: (1) to review several dropout processes, corresponding inferential issues and recent methodological advances; (2) to evaluate the impact of assumptions regarding the dropout processes on inference by simulation studies and an illustrative example using psychiatric data; and (3) to provide a general strategy for practitioners to perform analyses of longitudinal data with dropouts, using software available commercially or in the public domain. The statistical methods used in this paper are maximum likelihood, multiple imputation and semi-parametric regression methods for inference, as well as Little's test and index of sensitivity to nonignorability (ISNI) for assessing statistical dropout mechanisms. We show that accounting for the nature of the dropout process influences results and that sensitivity analysis is useful in assessing the robustness of parameter estimates and related uncertainties. We conclude that recording the causes of dropouts should be an integral part of any statistical analysis with longitudinal psychiatric data, and we recommend performing a sensitivity analysis when the exact nature of the dropout process cannot be discerned.

Antidepressive Agents↗

[Analysis of longitudinal Gaussian data with missing data on the response variable].

BACKGROUND: Using an application and a simulation study we show the bias induced by missing data in the outcome in longitudinal studies and discuss suitable statistical methods according to the type of missing responses when the variable under study is gaussian. METHOD: The model used for the analysis of gaussian longitudinal data is the mixed effects linear model. When the probability of response does not depend on the missing values of the outcome and on the parameters of the linear model, missing data are ignorable, and parameters of the mixed effects linear model may be estimated by the maximum likelihood method with classical softwares. When the missing data are non ignorable, several methods have been proposed. We describe the method proposed by Diggle and Kenward (1994) (DK method) for which a software is available. This model consists in the combination of a linear mixed effects model for the outcome variable and a logistic model for the probability of response which depends on the outcome variable. RESULTS: A simulation study shows the efficacy of this method and its limits when the data are not normal. In this case, estimators obtained by the DK approach may be more biased than estimators obtained under the hypothesis of ignorable missing data even if the data are non ignorable. Data of the Paquid cohort about the evolution of the scores to a neuropsychological test among elderly subjects show the bias of a naive analysis using all available data. Although missing responses are not ignorable in this study, estimates of the linear mixed effects model are not very different using the DK approach and the hypothesis of ignorable missing data. CONCLUSION: Statistical methods for longitudinal data including non ignorable missing responses are sensitive to hypotheses difficult to verify. Thus, it will be better in practical applications to perform an analysis under the hypothesis of ignorable missing responses and compare the results obtained with several approaches for non ignorable missing data. However, such a strategy requires development of new softwares.

Data Interpretation, Statistical↗

A shared random effect parameter approach for longitudinal dementia data with non-ignorable missing data.

A significant source of missing data in longitudinal epidemiologic studies on elderly individuals is death. It is generally believed that these missing data by death are non-ignorable to likelihood based inference. Inference based on data only from surviving participants in the study may lead to biased results. In this paper we model both the probability of disease and the probability of death using shared random effect parameters. We also propose to use the Laplace approximation for obtaining an approximate likelihood function so that high dimensional integration over the distributions of the random effect parameters is not necessary. Parameter estimates can be obtained by maximizing the approximate log-likelihood function. Data from a longitudinal dementia study will be used to illustrate the approach. A small simulation is conducted to compare parameter estimates from the proposed method to the 'naive' method where missing data is considered at random.

Aged↗

Analysis of longitudinal binary data with missing data due to dropouts.

Longitudinal binary data from clinical trials with missing observations are frequently analyzed by using the Last Observation Carry Forward (LOCF) method for imputing missing values at a visit (e.g., the prospectively defined primary visit time point for analysis at the end of treatment period). Usually, to understand time trend in treatment response, analyses are also performed separately on data at intermediate time points. The objective of such analyses is to estimate the proportion of "response" at a time point and then to compare two treatment groups (e.g., drug vs. placebo) by testing for the difference in the two proportions of response. The commonly used methods are Fisher's exact test, chi-squared test, Cochran-Mantel-Haenszel test, and logistic regression. Analyses based on the Observed Cases (OC) data are usually also performed and compared with those obtained by LOCF. Another approach that is gaining popularity (after the introduction of PROC GENMOD by the SAS Institute) is to use the method of Generalized Estimating Equations (GEE) with a view to include all repeated observations in the analysis in a more comprehensive manner. It is now well recognized, however, that results obtained by these methods are susceptible to bias, depending on the "missing data mechanism." Of particular concern is the bias introduced by NMAR dropouts. Because there is no one method to satisfactorily handle dropouts in data analysis, consensus is gathering toward doing analyses by several methods (including methods to handle NMAR dropouts) to evaluate sensitivity of results to model assumptions. In this article, we demonstrate application of the following methods for handling dropouts in longitudinal binary data: Generalized Linear Mixture Models (GLMM) (for handling NMAR dropouts), Weighted GEE (for handling MAR dropouts), and GEE (MCAR dropouts). The results are also compared with those obtained by logistic regression (univariate) on both LOCF and OC data.

Clinical Trials as Topic↗