PubMed HealthSearch

Biomedical subjects

R J Carroll

Publications and source records attributed to R J Carroll.

At least 19 recordsLinked to original sources

Measurement error and dietary intake.

This chapter reviews work of Carroll, Freedman, Kipnis, and Li (1998) on the statistical analysis of the relationship between dietary intake and health outcomes. In the area of nutritional epidemiology, there is some evidence from biomarker studies that the usual statistical model for dietary measurements may break down due to two causes: (a) systematic biases depending on a person's body mass index; and (b) an additional random component of bias, so that the error structure is the same as a one-way random effects model. We investigate this problem, in the context of (1) the estimation of the distribution of usual nutrient intake; (2) estimating the correlation between a nutrient instrument and usual nutrient intake; and (3) estimating the true relative risk from an estimated relative risk using the error-prone covariate. While systematic bias due to body mass index appears to have little effect, the additional random effect in the variance structure is shown to have a potentially important impact on overall results, both on corrections for relative risk estimates and in estimating the distribution usual of nutrient intake. Our results point to a need for new experiments aimed at estimation of a crucial parameter.

Bias

Measurement error, biases, and the validation of complex models for blood lead levels in children.

Measurement error causes biases in regression fits. If one could accurately measure exposure to environmental lead media, the line obtained would differ in important ways from the line obtained when one measures exposure with error. The effects of measurement error vary from study to study. It is dangerous to take measurement error corrections derived from one study and apply them to data from entirely different studies or populations. Measurement error can falsely invalidate a correct (complex mechanistic) model. If one builds a model such as the integrated exposure uptake biokinetic model carefully, using essentially error-free lead exposure data, and applies this model in a different data set with error-prone exposures, the complex mechanistic model will almost certainly do a poor job of prediction, especially of extremes. Although mean blood lead levels from such a process may be accurately predicted, in most cases one would expect serious underestimates or overestimates of the proportion of the population whose blood lead level exceeds certain standards.

Algorithms

Statistical design of calibration studies.

We investigated some design aspects of calibration studies. The specific situation addressed was one in which a large group is evaluated with a food-frequency questionnaire and a smaller calibration study is conducted through use of repeated food records or recalls, with the subjects in the calibration study constituting a random sample of those in the large group. In designing a calibration study, one may use large sample sizes and few food records per individual or smaller samples and more records per subject. Neither strategy is always preferable. Instead, the optimal method for a given study depends on the survey instrument used (24-h recalls or multiple-day food records) and the variables of interest.

Computer Simulation

Categorical regression analysis of acute exposure to tetrachloroethylene.

Exposure-response analysis of acute noncancer risks should consider both concentration (C) and duration (T) of exposure, as well as severity of response. Stratified categorical regression is a form of meta-analysis that addresses these needs by combining studies and analyzing response data expressed as ordinal severity categories. A generalized linear model for ordinal data was used to estimate the probability of response associated with exposure and severity category. Stratification of the regression model addresses systematic differences among studies by allowing one or more model parameters to vary across strata defined, for example, by species and sex. The ability to treat partial information addresses the difficulties in assigning consistent severity scores. Studies containing information on acute effects of tetrachloroethylene in rats, mice, and humans were analyzed. The mouse data were highly uncertain due to lack of data on effects of low concentrations and were excluded from the analysis. A model with species-specific concentration intercept terms for rat and human central nervous system data improved fit to the data compared with the base model (combined species). More complex models with strata defined by sex and species did not improve the fit. The stratified regression model allows human effect levels to be identified more confidently by basing the intercept on human data and the slope parameters on the combined data (on a C x T plot). This analysis provides an exposure-response function for acute exposures to tetrachloroethylene using categorical regression analysis.

Administration, Inhalation

Transformations to additivity in measurement error models.

In many problems, one wants to model the relationship between a response Y and a covariate X. Sometimes it is difficult, expensive, or even impossible to observe X directly, but one can instead observe a substitute variable W that is easier to obtain. By far, the most common model for the relationship between the actual covariate of interest X and the substitute W is W = X + U, where the variable U represents measurement error. This assumption of additive measurement error may be unreasonable for certain data sets. We propose a new model, namely h(W) = h(X) + U, where h(.) is a monotone transformation function selected from some family H of monotone functions. The idea of the new model is that, in the correct scale, measurement error is additive. We propose two possible transformation families H. One is based on selecting a transformation that makes the within-sample mean and standard deviation of replicated W's uncorrelated. The second is based on selecting the transformation so that the errors (U's) fit a prespecified distribution. Transformation families used are the parametric power transformations and a cubic spline family. Several data examples are presented to illustrate the methods.

Adult

Analyzing bivariate continuous data grouped into categories defined by empirical quantiles of marginal distributions.

Epidemiologists sometimes study the association between two measurements of exposure on the same subjects by grouping the original bivariate continuous data into categories that are defined by the empirical quantiles of the two marginal distributions. Although such grouped data are presented in a two-way contingency table, the cell counts in this table do not have a multinomial distribution. We describe the joint distribution of counts in such a table by the term empirical bivariate quantile-partitioned (EBQP) distribution. Blomqvist (1950, Annals of Mathematical Statistics 21, 539-600) gave an asymptotic EBQP theory for bivariate data partitioned by the sample medians. We demonstrate that his asymptotic theory is not correct, however, except in special cases. We present a general asymptotic theory for tables of arbitrary dimensions and apply this theory to construct confidence intervals for the kappa statistic. We show by simulations that the confidence interval procedures we propose have near nominal coverage for sample sizes exceeding 60 for both 2 x 2 and 3 x 3 tables. These simulations also illustrate that the asymptotic theory of Blomqvist (1950) and the methods that Fleiss, Cohen, and Everitt (1969, Psychological Bulletin 72, 323-327) give for multinomial tables can yield subnominal coverage for kappa calculated from EBQP tables, although in some cases the coverage for these procedures is near nominal levels.

Analysis of Variance

Design aspects of calibration studies in nutrition, with analysis of missing data in linear measurement error models.

Motivated by an example in nutritional epidemiology, we investigate some design and analysis aspects of linear measurement error models with missing surrogate data. The specific problem investigated consists of an initial large sample in which the response (a food frequency questionnaire, FFQ) is observed and then a smaller calibration study in which replicates of the error prone predictor are observed (food records or recalls, FR). The difference between our analysis and most of the measurement error model literature is that, in our study, the selection into the calibration study can depend on the value of the response. Rationale for this type of design is given. Two major problems are investigated. In the design of a calibration study, one has the option of larger sample sizes and fewer replicates or smaller sample sizes and more replicates. Somewhat surprisingly, neither strategy is uniformly preferable in cases of practical interest. The answers depend on the instrument used (recalls or records) and the parameters of interest. The second problem investigated is one of analysis. In the usual linear model with no missing data, method of moments estimates and normal-theory maximum likelihood estimates are approximately equivalent, with the former method in most use because it can be calculated easily and explicitly. Both estimates are valid without any distributional assumptions. In contrast, in the missing data problem under consideration, only the moments estimate is distribution-free, but the maximum likelihood estimate has at least 50% greater precision in practical situations when normality obtains. Implications for the design of nutritional calibration studies are discussed.

Aged

On design considerations and randomization-based inference for community intervention trials.

This paper discusses design considerations and the role of randomization-based inference in randomized community intervention trials. We stress that longitudinal follow-up of cohorts within communities often yields useful information on the effects of intervention on individuals, whereas cross-sectional surveys can usefully assess the impact of intervention on group indices of health. We also discuss briefly special design considerations, such as sampling cohorts from targeted subpopulations (for example, heavy smokers), matching the communities, calculating sample size, and other practical issues. We present randomization tests for matched and unmatched cohort designs. As is well known, these tests necessarily have proper size under the strong null hypothesis that treatment has no effect on any community response. It is less well known, however, that the size of randomization tests can exceed nominal levels under the 'weak' null hypothesis that intervention does not affect the average community response. Because this weak null hypothesis is of interest in community intervention trials, we study the size of randomization tests by simulation under conditions in which the weak null hypothesis holds but the strong null hypothesis does not. In unmatched studies, size may exceed nominal levels under the weak null hypothesis if there are more intervention than control communities and if the variance among community responses is larger among control communities than among intervention communities; size may also exceed nominal levels if there are more control than intervention communities and if the variance among community responses is larger among intervention communities. Otherwise, size is likely near nominal levels. To avoid such problems, we recommend use of the same numbers of control and intervention communities in unmatched designs. Pair-matched designs usually have size near nominal levels, even under the weak null hypothesis. We have identified some extreme cases, unlikely to arise in practice, in which even the size of pair-matched studies can exceed nominal levels. These simulations, however, tend to confirm the robustness of randomization tests for matched and unmatched community intervention trials, particularly if the latter designs have equal numbers of intervention and control communities. We also describe adaptations of randomization tests to allow for covariate adjustment, missing data, and application to cross-sectional surveys. We show that covariate adjustment can increase power, but such power gains diminish as the random component of variation among communities increases, which corresponds to increasing intraclass correlation of responses within communities. We briefly relate our results to model-based methods of inference for community intervention trials that include hierarchical models such as an analysis of variance model with random community effects and fixed intervention effects. Although we have tailored this paper to the design of community intervention trials, many of the ideas apply to other experiments in which one allocates groups or clusters of subjects at random to intervention or control treatments.

Clinical Trials as Topic

Use of semiquantitative food frequency questionnaires to estimate the distribution of usual intake.

The authors consider whether semiquantitative food frequency questionnaires can be used to survey a population to estimate the distribution of usual intake. They take as an assumption that, if they were possible to obtain, the mean of many food records or recalls would be an accurate representation of an individual's usual diet. They then assume that nutrient intake as measured by a questionnaire follows a linear regression model when regressed against the usual intake of that nutrient. If the coefficients in this regression relation were known, then the distribution of usual intake could be constructed from the responses to the questionnaire. Since one generally does not know the values of the coefficients, they need to be estimated from a calibration study in which respondents complete the questionnaire together with multiple food records or recalls. This can be done either through an internal subset of the data or through an independent external study. With an internal substudy, the authors find that food frequency questionnaires typically provide little information about the distribution of usual intake in addition to that obtained from the multiple records or recalls in the substudy. When the substudy is external, if it is small then having very large numbers of subjects completing food frequency questionnaires in the survey is no more efficient than having a few subjects completing food records or recalls. However, if the external substudy is large and accurately characterizes the relation between the questionnaire response and usual intake, food frequency questionnaires can provide a cost-efficient way of estimating the distribution of usual intake. These results do not apply to the different problem of correcting relative risks for the effects of measurement error.

Bias

Quasilikelihood estimation in measurement error models with correlated replicates.

We consider quasilikelihood models when some of the predictors are measured with error. In many cases, the true but fallible predictor is impossible to measure, and the best one can do is to obtain replicates of the fallible predictor. We consider the case that the replicates are not independent. If one assumes that replicates are independent and they are not, one typically underestimates the extent of the measurement error, leading to an inconsistent errors in variables correction. We devise techniques for estimating the measurement error covariance matrix. In addition, we discuss how one might perform a quasilikelihood analysis by computing the mean and variance functions of the observed data, both using approximations and also exactly through a Monte Carlo method. The methods are illustrated on a data set involving systolic blood pressure and urinary sodium chloride, where the measurement errors appear to be approximately normally distributed but highly correlated, and the distribution of the true predictor is reasonably modeled as a mixture of normals.

Analysis of Variance

Estimating the reliability of an exposure variable in the presence of confounders.

In this paper we discuss estimation of the reliability of an exposure variable in the presence of confounders measured without error. We give an explicit formula that shows how the exposure becomes less reliable as the degree of correlation between the exposure and confounders increases. We also discuss biases in the corresponding logistic regression estimates and methods for correction. Data from a matched case-control study of hormone levels and risk of breast cancer are used to illustrate the methods.

Breast Neoplasms

A transgene coding for a human insulin analog has a mitogenic effect on murine embryonic beta cells.

We have investigated the mitogenic effect of three mutant forms of human insulin on insulin-producing beta cells of the developing pancreas. We examined transgenic embryonic and adult mice expressing (i) human [AspB10]-proinsulin/insulin ([AspB10]ProIN/IN), produced by replacement of histidine by aspartic acid at position 10 of the B chain and characterized by an increased affinity for the insulin receptor; (ii) human [LeuA3]insulin, produced by the substitution of leucine for valine in position 3 of the A chain, which exhibits decreased receptor binding affinity; and (iii) human [LeuA3, AspB10]insulin "double" mutation. During development, beta cells of AspB10 embryos were twice as abundant and had a 3 times higher rate of proliferation compared with beta cells of littermate controls. The mitogenic effect of [AspB10]ProIN/IN was specific for embryonic beta cells because the rate of proliferation of beta cells of adults and of glucagon (alpha) cells and adrenal chromaffin cells of embryos was similar in AspB10 mice and controls. In contrast to AspB10 embryos, the number of beta cells in the LeuA3 and "double" mutant lines was similar to the number in controls. These findings indicate that the [AspB10]ProIN/IN analog increased the rate of fetal beta-cell proliferation. The mechanism or mechanisms that mediate this mitogenic effect remain to be determined.

Amino Acid Sequence

International comparison of waiting times for selected cardiovascular procedures.

OBJECTIVES: This study was designed to compare waiting times for cardiovascular procedures in five different health care delivery/financing systems. BACKGROUND: A recurrent criticism of national health care systems is long waiting times, or "queues," for high technology procedures. However, no objective data exist comparing waiting times in the United States with those in other systems. METHODS: Directors of cardiac catheterization laboratories, directors of cardiac surgery in the United States, U.S. Department of Veterans Affairs (VA) system, Canada and the United Kingdom and directors of cardiology clinics in Sweden were asked to respond to a mailed questionnaire as to how long it would take to obtain coronary angiography or coronary artery bypass surgery, or both, for specified case scenarios at their institutions. RESULTS: Significant differences in waiting times (p < 0.00001) were found among the systems for all four scenarios (elective and urgent angiography, elective and urgent bypass surgery). Compared with non-VA hospitals in the United States, waiting times were significantly longer in all systems, with the exception of waiting times for urgent surgery in the U.S. VA hospitals (p = 0.9). The longest waiting times for all four procedures were reported in the United Kingdom, Sweden and Canada, with some waiting times for elective procedures > 9 months. CONCLUSIONS: Physicians report that patients treated in health care systems structured differently from the non-VA hospital system in the United States wait significantly longer for cardiac catheterization and coronary artery bypass surgery.

Canada

Hyperproinsulinaemia in obese fat/fat mice associated with a carboxypeptidase E mutation which reduces enzyme activity.

Mice homozygous for the fat mutation develop obesity and hyperglycaemia that can be suppressed by treatment with exogenous insulin. The fat mutation maps to mouse chromosome 8, very close to the gene for carboxypeptidase E (Cpe), which encodes an enzyme (CPE) that processes prohormone intermediates such as proinsulin. We now demonstrate a defect in proinsulin processing associated with the virtual absence of CPE activity in extracts of fat/fat pancreatic islets and pituitaries. A single Ser202Pro mutation distinguishes the mutant Cpe allele, and abolishes enzymatic activity in vitro. Thus, the fat mutation represents the first demonstration of an obesity-diabetes syndrome elicited by a genetic defect in a prohormone processing pathway.

Amino Acid Sequence

Adjusting for time trends when estimating the relationship between dietary intake obtained from a food frequency questionnaire and true average intake.

In measuring food intake, three common methods are used: 24-hour recalls, food frequency questionnaires and food records. Food records or 24-hour recalls are often thought to be the most reliable, but they are difficult and expensive to obtain. The question of interest to us is to use the food records or 24-hour recalls to examine possible systematic biases in questionnaires as a measure of usual food intake. In Freedman, et al. (1991), this problem is addressed through a linear errors in variables analysis. Their model assumes that all measurements on a given individual have the same mean and variance. However, such assumptions may be violated in at least two circumstances, as in for example the Women's Health Trial Vanguard Study and in the Finnish Smokers' Study. First, some studies occur over a period of years, and diets may change over the course of the study. Second, measurements might be taken at different times of the year, and it is known that diets differ on the basis of seasonal factors. In this paper, we will suggest new models incorporating mean and variance offsets, i.e., changes in the population mean and variance for observations taken at different time points. The parameters in the model are estimated by simple methods, and the theory of unbiased estimating equations (M-estimates) is used to derive asymptotic covariance matrix estimates. The methods are illustrated with data from the Women's Health Trial Vanguard Study.

Analysis of Variance

Analysis of tomato root initiation using a normal mixture distribution.

We attempt to identify the number of underlying physical phenomena behind tomato lateral root initiation by using a normal mixture distribution coupled with the Box-Cox power transformation. An initial analysis of the data suggested the possibility of two (possibly more) subpopulations, but upon taking reciprocals, the data appear to be very nearly Gaussian. A simulation study explores the possibility of erroneously detecting a second subpopulation by fitting data which are improperly scaled. A power calculation suggests that only unrealistically large sample sizes can detect the unbalanced mixtures one might expect with data of this type.

Biometry

Measurement error, instrumental variables and corrections for attenuation with applications to meta-analyses.

MacMahon et al. present a meta-analysis of the effect of blood pressure on coronary heart disease, as well as new methods for estimation in measurement error models for the case when a replicate or second measurement is made of the fallible predictor. The correction for attenuation used by these authors is compared to others already existing in the literature, as well as to a new instrumental variable method. The assumptions justifying the various methods are examined and their efficiencies are studied via simulation. Compared to the methods we discuss, the method of MacMahon et al. may have bias in some circumstances because it does not take into account: (i) possible correlations among the predictors within a study; (ii) possible bias in the second measurement; or (iii) possibly differing marginal distributions of the predictors or measurement errors across studies. A unifying asymptotic theory using estimating equations is also presented.

Bias