PubMed Health⌕ Search

PubMed · 17156280

Feature-specific penalized latent class analysis for genomic data.

Abstract

Genomic data are often characterized by a moderate to large number of categorical variables observed for relatively few subjects. Some of the variables may be missing or noninformative. An example of such data is loss of heterozygosity (LOH), a dichotomous variable, observed on a moderate number of genetic markers. We first consider a latent class model where, conditional on unobserved membership in one of k classes, the variables are independent with probabilities determined by a regression model of low dimension q. Using a family of penalties including the ridge and LASSO, we extend this model to address higher-dimensional problems. Finally, we present an orthogonal map that transforms marker space to a space of "features" for which the constrained model has better predictive power. We demonstrate these methods on LOH data collected at 19 markers from 93 brain tumor patients. For this data set, the existing unpenalized latent class methodology does not produce estimates. Additionally, we show that posterior classes obtained from this method are associated with survival for these patients.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

E Andrés Houseman, Brent A Coull, Rebecca A Betensky. 2006. Feature-specific penalized latent class analysis for genomic data.. https://doi.org/10.1111/j.1541-0420.2006.00566.x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Tutorial in biostatistics: competing risks and multi-state models.

Standard survival data measure the time span from some time origin until the occurrence of one type of event. If several types of events occur, a model describing progression to each of these competing risks is needed. Multi-state models generalize competing risks models by also describing transitions to intermediate events. Methods to analyze such models have been developed over the last two decades. Fortunately, most of the analyzes can be performed within the standard statistical packages, but may require some extra effort with respect to data preparation and programming. This tutorial aims to review statistical methods for the analysis of competing risks and multi-state models. Although some conceptual issues are covered, the emphasis is on practical issues like data preparation, estimation of the effect of covariates, and estimation of cumulative incidence functions and state and transition probabilities. Examples of analysis with standard software are shown.

Biometry↗

The role of education in biostatistical consulting.

Medical students, residents, postdoctoral fellows, and faculty commonly consult with biostatistical experts about study design and data analysis when conducting clinical research. The role of biostatistical training during these consultations is examined, and characterizations of the connections between biostatistical consultation and education are reviewed. The presence and kinds of teaching efforts during biostatistical consults at four academic research institutions over various periods of time between 1999 and 2005 (237 consultations in total) were recorded and are described. By site, 67, 70, 78, and 100 per cent of the consulting sessions included biostatistical training, with an overall 78 per cent (95 per cent CI: 73-83 per cent) of consultations including an educational component when all consultations were combined. Training covered a wide range of biostatistical topics. Seventy-five per cent of the consultations with faculty (120/161), 79 per cent with fellows and residents (31/39), and 100 per cent with medical students (10/10) included some degree of instruction in study design or statistical analysis topics. Results show that both the need and the opportunity exist for specialized biostatistical instruction during one-on-one sessions between a consulting biostatistician and physicians, medical students, and research staff. Academic researchers are ideally positioned to absorb this kind of training when they initiate a request for assistance with their own research project.

Biometry↗

Improving the quality of patient care using reliability measures: a classification tree approach.

This paper considers the application and interpretation of new reliability measures for a classification tree-based medical risk assessment tool. Following the construction of a classification tree reliability measures may then be used to provide an estimate of the precision of the classification and the probability in each terminal node of the classification tree. Identification of unreliable nodes (those that have low precision) in this application may indicate patient groups requiring closer monitoring or scenarios in which further information about the patient is required, thereby providing medical practitioners with an avenue for more informed decision making.

Biometry↗