PubMed Health⌕ Search

PubMed · 16284340

Microarray gene expression data with linked survival phenotypes: diffuse large-B-cell lymphoma revisited.

Abstract

Diffuse large-B-cell lymphoma (DLBCL) is an aggressive malignancy of mature B lymphocytes and is the most common type of lymphoma in adults. While treatment advances have been substantial in what was formerly a fatal disease, less than 50% of patients achieve lasting remission. In an effort to predict treatment success and explain disease heterogeneity clinical features have been employed for prognostic purposes, but have yielded only modest predictive performance. This has spawned a series of high-profile microarray-based gene expression studies of DLBCL, in the hope that molecular-level information could be used to refine prognosis. The intent of this paper is to reevaluate these microarray-based prognostic assessments, and extend the statistical methodology that has been used in this context. Methodological challenges arise in using patients' gene expression profiles to predict survival endpoints on account of the large number of genes and their complex interdependence. We initially focus on the Lymphochip data and analysis of Rosenwald et al. (2002). After describing relationships between the analyses performed and gene harvesting (Hastie et al., 2001a), we argue for the utility of penalized approaches, in particular least angle regression-least absolute shrinkage and selection operator (Efron et al., 2004). While these techniques have been extended to the proportional hazards/partial likelihood framework, the resultant algorithms are computationally burdensome. We develop residual-based approximations that eliminate this burden yet perform similarly. Comparisons of predictive accuracy across both methods and studies are effected using time-dependent receiver operating characteristic curves. These indicate that gene expression data, in turn, only delivers modest predictions of posttherapy DLBCL survival. We conclude by outlining possibilities for further work.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mark R Segal. 2005-11-11. Microarray gene expression data with linked survival phenotypes: diffuse large-B-cell lymphoma revisited.. https://doi.org/10.1093/biostatistics%2Fkxj006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Tutorial in biostatistics: competing risks and multi-state models.

Standard survival data measure the time span from some time origin until the occurrence of one type of event. If several types of events occur, a model describing progression to each of these competing risks is needed. Multi-state models generalize competing risks models by also describing transitions to intermediate events. Methods to analyze such models have been developed over the last two decades. Fortunately, most of the analyzes can be performed within the standard statistical packages, but may require some extra effort with respect to data preparation and programming. This tutorial aims to review statistical methods for the analysis of competing risks and multi-state models. Although some conceptual issues are covered, the emphasis is on practical issues like data preparation, estimation of the effect of covariates, and estimation of cumulative incidence functions and state and transition probabilities. Examples of analysis with standard software are shown.

Biometry↗

The role of education in biostatistical consulting.

Medical students, residents, postdoctoral fellows, and faculty commonly consult with biostatistical experts about study design and data analysis when conducting clinical research. The role of biostatistical training during these consultations is examined, and characterizations of the connections between biostatistical consultation and education are reviewed. The presence and kinds of teaching efforts during biostatistical consults at four academic research institutions over various periods of time between 1999 and 2005 (237 consultations in total) were recorded and are described. By site, 67, 70, 78, and 100 per cent of the consulting sessions included biostatistical training, with an overall 78 per cent (95 per cent CI: 73-83 per cent) of consultations including an educational component when all consultations were combined. Training covered a wide range of biostatistical topics. Seventy-five per cent of the consultations with faculty (120/161), 79 per cent with fellows and residents (31/39), and 100 per cent with medical students (10/10) included some degree of instruction in study design or statistical analysis topics. Results show that both the need and the opportunity exist for specialized biostatistical instruction during one-on-one sessions between a consulting biostatistician and physicians, medical students, and research staff. Academic researchers are ideally positioned to absorb this kind of training when they initiate a request for assistance with their own research project.

Biometry↗

Improving the quality of patient care using reliability measures: a classification tree approach.

This paper considers the application and interpretation of new reliability measures for a classification tree-based medical risk assessment tool. Following the construction of a classification tree reliability measures may then be used to provide an estimate of the precision of the classification and the probability in each terminal node of the classification tree. Identification of unreliable nodes (those that have low precision) in this application may indicate patient groups requiring closer monitoring or scenarios in which further information about the patient is required, thereby providing medical practitioners with an avenue for more informed decision making.

Biometry↗