PubMed Health⌕ Search

PubMed · 9169176

A simple program to create exact person-time data in cohort analyses.

Abstract

BACKGROUND: Before disease rates can be calculated a tabulation of the length of follow-up for each person in the cohort has to be made. In complicated analyses such tabulations are often stratified by many characteristics, some which show no change with time, such as gender or year of birth, and some which do change with time, such as age or cumulative exposure. Available computer programs often restrict the way these tables can be made, particularly when handling time-dependent variables. METHODS: This paper presents a simple computer program which calculates the length of follow-up for each person in a study. RESULTS: Person-time data can be tabulated by a large number of variables using this method. This program is extremely flexible in the way that time-dependent variables can be created, can categorize observations by any unit of person-time, and will run on a range of platforms including a personal computer. CONCLUSIONS: This method should simplify the task of creating person-time data for analyses of disease rates in epidemiological studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J Wood, D Richardson, S Wing. 1997. A simple program to create exact person-time data in cohort analyses.. https://doi.org/10.1093/ije%2F26.2.395

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A boosting approach to flexible semiparametric mixed models.

In linear mixed models the influence of covariates is restricted to a strictly parametric form. With the rise of semi- and non-parametric regression also the mixed model has been expanded to allow for additive predictors. The common approach uses the representation of additive models as mixed models. An alternative approach that is proposed in the present paper is likelihood based boosting. Boosting originates in the machine learning community where it has been proposed as a technique to improve classification procedures by combining estimates with reweighted observations. Likelihood based boosting is a general method which may be seen as an extension of L2 boost. In additive mixed models the advantage of boosting techniques in the form of componentwise boosting is that it is suitable for high dimensional settings where many explanatory variables are present. It allows to fit additive models for many covariates with implicit selection of relevant variables and automatic selection of smoothing parameters. Moreover, boosting techniques may be used to incorporate the subject-specific variation of smooth influence functions by specifying 'random slopes' on smooth effects. This results in flexible semiparametric mixed models which are appropriate in cases where a simple random intercept is unable to capture the variation of effects across subjects.

Cohort Studies↗

Estimation of attributable number of deaths and standard errors from simple and complex sampled cohorts.

Estimates of the attributable number of deaths (AD) from all causes can be obtained by first estimating population attributable risk (AR) adjusted for confounding covariates, and then multiplying the AR by the number of deaths determined from vital mortality statistics that occurred in the population for a specific time period. Proportional hazard regression estimates of adjusted relative hazards obtained from mortality follow-up data from a cohort is combined with a joint distribution of risk factor and confounders to compute an adjusted AR. Two estimators of adjusted AR are examined. These estimators differ according to which reference population is used to obtain the joint distribution of risk factor and confounders. Two types of reference populations were considered: (i) the population represented by the baseline cohort and (ii) a population that is external to the cohort. Methods used in survey sampling are applied to obtain estimates of the variance of the AD estimator. These variances can be applied to data that range from simple random samples to multistage stratified cluster samples, which are used in national household surveys. The variance estimation of AD is illustrated in an analysis of excess deaths due to having a non-ideal body mass index using the second National Health and Examination Survey (NHANES) Mortality Study and the 1999-2002 NHANES. These methods can also be used to estimate the attributable number of cause-specific deaths and their standard errors when the time period for the accrual of deaths is short.

Cohort Studies↗

Longitudinal variable selection by cross-validation in the case of many covariates.

Longitudinal models are commonly used for studying data collected on individuals repeatedly through time. While there are now a variety of such models available (marginal models, mixed effects models, etc.), far fewer options exist for the closely related issue of variable selection. In addition, longitudinal data typically derive from medical or other large-scale studies where often large numbers of potential explanatory variables and hence even larger numbers of candidate models must be considered. Cross-validation is a popular method for variable selection based on the predictive ability of the model. Here, we propose a cross-validation Markov chain Monte Carlo procedure as a general variable selection tool which avoids the need to visit all candidate models. Inclusion of a 'one-standard error' rule provides users with a collection of good models as is often desired. We demonstrate the effectiveness of our procedure both in a simulation setting and in a real application.

Cohort Studies↗