PubMed Health⌕ Search

Biomedical subjects

A E Raftery

Publications and source records attributed to A E Raftery.

6 recordsLinked to original sources

Variable selection and Bayesian model averaging in case-control studies.

Covariate and confounder selection in case-control studies is often carried out using a statistical variable selection method, such as a two-step method or a stepwise method in logistic regression. Inference is then carried out conditionally on the selected model, but this ignores the model uncertainty implicit in the variable selection process, and so may underestimate uncertainty about relative risks. We report on a simulation study designed to be similar to actual case-control studies. This shows that p-values computed after variable selection can greatly overstate the strength of conclusions. For example, for our simulated case-control studies with 1000 subjects, of variables declared to be 'significant' with p-values between 0.01 and 0.05, only 49 per cent actually were risk factors when stepwise variable selection was used. We propose Bayesian model averaging as a formal way of taking account of model uncertainty in case-control studies. This yields an easily interpreted summary, the posterior probability that a variable is a risk factor, and our simulation study indicates this to be reasonably well calibrated in the situations simulated. The methods are applied and compared in the context of a case-control study of cervical cancer.

Analysis of Variance↗

Model-based clustering and data transformations for gene expression data.

MOTIVATION: Clustering is a useful exploratory technique for the analysis of gene expression data. Many different heuristic clustering algorithms have been proposed in this context. Clustering algorithms based on probability models offer a principled alternative to heuristic algorithms. In particular, model-based clustering assumes that the data is generated by a finite mixture of underlying probability distributions such as multivariate normal distributions. The issues of selecting a 'good' clustering method and determining the 'correct' number of clusters are reduced to model selection problems in the probability framework. Gaussian mixture models have been shown to be a powerful tool for clustering in many applications. RESULTS: We benchmarked the performance of model-based clustering on several synthetic and real gene expression data sets for which external evaluation criteria were available. The model-based approach has superior performance on our synthetic data sets, consistently selecting the correct model and the number of clusters. On real expression data, the model-based approach produced clusters of quality comparable to a leading heuristic clustering algorithm, but with the key advantage of suggesting the number of clusters and an appropriate model. We also explored the validity of the Gaussian mixture assumption on different transformations of real data. We also assessed the degree to which these real gene expression data sets fit multivariate Gaussian distributions both before and after subjecting them to commonly used data transformations. Suitably chosen transformations seem to result in reasonable fits. AVAILABILITY: MCLUST is available at http://www.stat.washington.edu/fraley/mclust. The software for the diagonal model is under development. CONTACT: kayee@cs.washington.edu. SUPPLEMENTARY INFORMATION: http://www.cs.washington.edu/homes/kayee/model.

Algorithms↗

Bayesian information criterion for censored survival models.

We investigate the Bayesian Information Criterion (BIC) for variable selection in models for censored survival data. Kass and Wasserman (1995, Journal of the American Statistical Association 90, 928-934) showed that BIC provides a close approximation to the Bayes factor when a unit-information prior on the parameter space is used. We propose a revision of the penalty term in BIC so that it is defined in terms of the number of uncensored events instead of the number of observations. For a simple censored data model, this revision results in a better approximation to the exact Bayes factor based on a conjugate unit-information prior. In the Cox proportional hazards regression model, we propose defining BIC in terms of the maximized partial likelihood. Using the number of deaths rather than the number of individuals in the BIC penalty term corresponds to a more realistic prior on the parameter space and is shown to improve predictive performance for assessing stroke risk in the Cardiovascular Health Study.

Aged↗

Are births underreported in rural China? Manipulation of statistical records in response to China's population policies.

Under the current family planning policy in China, the criterion for evaluating all parties involved in the birth planning system provides an incentive for everyone to see that the policy is met, either in reality through strict enforcement of family planning regulations, or statistically through manipulation of statistical records. We investigate underreporting of births in four rural counties of northern China, using data from a 1992 sample survey featuring a reproductive history. To clarify the mechanisms of underreporting, we focus on the ways in which reporting errors may affect the distribution of first births by time since marriage. The results of our investigation suggest that in three of the four counties, first-birth intervals are lengthened by underreporting of girl babies and by replacing them with second births reported as first births.

China↗

Demand or ideation? Evidence from the Iranian marital fertility decline.

Is the onset of fertility decline caused by structural socioeconomic changes or by the transmission of new ideas? The decline of marital fertility in Iran provides a quasi-experimental setting for addressing this question. Massive economic growth started in 1955; measurable ideational changes took place in 1967. We argue that the decline is described more precisely by demand theory than by ideation theory. It began around 1959, just after the onset of massive economic growth but well before the ideational changes. It paralleled the rapid growth of participation in primary education, and we found no evidence that the 1967 events had any effect on the decline. More than one-quarter of the decline can be attributed to the reduction in child mortality, a key mechanism of demand theory. Several other findings support this main conclusion.

Adolescent↗

Forecasting spore concentrations: a time series approach.

Fungal basidiospores and Cladosporium spores are the two most numerous spore types in the air of Dublin and its surroundings. They are known to have allergenic components, and the aim of the study described here is to develop a predictive model for these spores. A very simple model, which combines an estimated diurnal rhythm with a simple, one-parameter time series model, provided good short-term forecasts. The one-step prediction error variance was reduced by 88% for Cladosporium spores and by 98% for basidiospores.

Air Microbiology↗