PubMed Health⌕ Search

PubMed · 11796089

DINER (Data Into Nutrients for Epidemiological Research) - a new data-entry program for nutritional analysis in the EPIC-Norfolk cohort and the 7-day diary method.

Abstract

BACKGROUND AND OBJECTIVE: A new data-entry system (DINER - Data Into Nutrients for Epidemiological Research) for food record methods has been devised for the European Prospective Investigation into Cancer (EPIC) cohort study of 25,000 men and women in Norfolk. DINER has been developed to address the problems of efficiency and consistency of data entry, comparability of data, maximising information and future flexibility in large long-term population studies of diet and disease that use record methods to assess dietary intakes. DINER captures more detail than traditional systems and enables provision of new variables for specific food types or groups. The system has been designed to be fully flexible and easy to update. Analysis of consistency of data entry was tested in a group of 3525 participants entered by 25 coders. RESULTS: A food list of 9000 food items and values for 24,000 portion sizes have been incorporated into the database, using information from the 5979 diaries coded since 1995. Analysis of consistency of entry indicated that this has largely been achieved. The effect of coders in a multivariate regression model was significant only if the three coders involved in early use of the program were included (P < 0.013). CONCLUSIONS: The development of DINER has facilitated the use of more accurate record methods in large-scale epidemiological studies of diet and disease. Furthermore, the retention of original information as an extensive food list allows greater flexibility in later analyses of data of multiple dietary hypotheses.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A A Welch, A McTaggart, A A Mulligan, R Luben, N Walker, K T Khaw, N E Day, S A Bingham. 2001. DINER (Data Into Nutrients for Epidemiological Research) - a new data-entry program for nutritional analysis in the EPIC-Norfolk cohort and the 7-day diary method.. https://doi.org/10.1079/phn2001196

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A boosting approach to flexible semiparametric mixed models.

In linear mixed models the influence of covariates is restricted to a strictly parametric form. With the rise of semi- and non-parametric regression also the mixed model has been expanded to allow for additive predictors. The common approach uses the representation of additive models as mixed models. An alternative approach that is proposed in the present paper is likelihood based boosting. Boosting originates in the machine learning community where it has been proposed as a technique to improve classification procedures by combining estimates with reweighted observations. Likelihood based boosting is a general method which may be seen as an extension of L2 boost. In additive mixed models the advantage of boosting techniques in the form of componentwise boosting is that it is suitable for high dimensional settings where many explanatory variables are present. It allows to fit additive models for many covariates with implicit selection of relevant variables and automatic selection of smoothing parameters. Moreover, boosting techniques may be used to incorporate the subject-specific variation of smooth influence functions by specifying 'random slopes' on smooth effects. This results in flexible semiparametric mixed models which are appropriate in cases where a simple random intercept is unable to capture the variation of effects across subjects.

Cohort Studies↗

Estimation of attributable number of deaths and standard errors from simple and complex sampled cohorts.

Estimates of the attributable number of deaths (AD) from all causes can be obtained by first estimating population attributable risk (AR) adjusted for confounding covariates, and then multiplying the AR by the number of deaths determined from vital mortality statistics that occurred in the population for a specific time period. Proportional hazard regression estimates of adjusted relative hazards obtained from mortality follow-up data from a cohort is combined with a joint distribution of risk factor and confounders to compute an adjusted AR. Two estimators of adjusted AR are examined. These estimators differ according to which reference population is used to obtain the joint distribution of risk factor and confounders. Two types of reference populations were considered: (i) the population represented by the baseline cohort and (ii) a population that is external to the cohort. Methods used in survey sampling are applied to obtain estimates of the variance of the AD estimator. These variances can be applied to data that range from simple random samples to multistage stratified cluster samples, which are used in national household surveys. The variance estimation of AD is illustrated in an analysis of excess deaths due to having a non-ideal body mass index using the second National Health and Examination Survey (NHANES) Mortality Study and the 1999-2002 NHANES. These methods can also be used to estimate the attributable number of cause-specific deaths and their standard errors when the time period for the accrual of deaths is short.

Cohort Studies↗

Longitudinal variable selection by cross-validation in the case of many covariates.

Longitudinal models are commonly used for studying data collected on individuals repeatedly through time. While there are now a variety of such models available (marginal models, mixed effects models, etc.), far fewer options exist for the closely related issue of variable selection. In addition, longitudinal data typically derive from medical or other large-scale studies where often large numbers of potential explanatory variables and hence even larger numbers of candidate models must be considered. Cross-validation is a popular method for variable selection based on the predictive ability of the model. Here, we propose a cross-validation Markov chain Monte Carlo procedure as a general variable selection tool which avoids the need to visit all candidate models. Inclusion of a 'one-standard error' rule provides users with a collection of good models as is often desired. We demonstrate the effectiveness of our procedure both in a simulation setting and in a real application.

Cohort Studies↗