PubMed Health⌕ Search

Biomedical subjects

Yihui Luan

Publications and source records attributed to Yihui Luan.

4 recordsLinked to original sources

Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray data.

MOTIVATION: An important area of research in the postgenomics era is to relate high-dimensional genetic or genomic data to various clinical phenotypes of patients. Due to large variability in time to certain clinical events among patients, studying possibly censored survival phenotypes can be more informative than treating the phenotypes as categorical variables. Due to high dimensionality and censoring, building a predictive model for time to event is more difficult than the classification/linear regression problem. We propose to develop a boosting procedure using smoothing splines for estimating the general proportional hazards models. Such a procedure can potentially be used for identifying non-linear effects of genes on the risk of developing an event. RESULTS: Our empirical simulation studies showed that the procedure can indeed recover the true functional forms of the covariates and can identify important variables that are related to the risk of an event. Results from predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed method can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. In addition, there is clear evidence of non-linear effects of some genes on survival time.

Algorithms↗

Clustering of time-course gene expression data using a mixed-effects model with B-splines.

MOTIVATION: Time-course gene expression data are often measured to study dynamic biological systems and gene regulatory networks. To account for time dependency of the gene expression measurements over time and the noisy nature of the microarray data, the mixed-effects model using B-splines was introduced. This paper further explores such mixed-effects model in analyzing the time-course gene expression data and in performing clustering of genes in a mixture model framework. RESULTS: After fitting the mixture model in the framework of the mixed-effects model using an EM algorithm, we obtained the smooth mean gene expression curve for each cluster. For each gene, we obtained the best linear unbiased smooth estimate of its gene expression trajectory over time, combining data from that gene and other genes in the same cluster. Simulated data indicate that the methods can effectively cluster noisy curves into clusters differing in either the shapes of the curves or the times to the peaks of the curves. We further demonstrate the proposed method by clustering the yeast genes based on their cell cycle gene expression data and the human genes based on the temporal transcriptional response of fibroblasts to serum. Clear periodic patterns and varying times to peaks are observed for different clusters of the cell-cycle regulated genes. Results of the analysis of the human fibroblasts data show seven distinct transcriptional response profiles with biological relevance. AVAILABILITY: Matlab programs are available on request from the authors.

Algorithms↗

Kernel Cox regression models for linking gene expression profiles to censored survival data.

In functional genomics, one important problem is to relate the microarray gene expression profiles to various clinical phenotypes from patients. The success has been demonstrated in molecular classification of cancer in which gene expression data serve as predictors and different types of cancer are the binary or multi-categorical outcome variable. However, there has been less research in linking gene expression profiles to other types of phenotypes, in particular, the censored survival data such as patients' overall survival or cancer relapse times. In the paper, we develop a kernel Cox regression model for relating gene expression profiles to censored phenotypes in the framework the penalization method in terms of function estimation in reproducing kernel Hilbert spaces. To circumvent the problem of censoring, we use the negative partial likelihood as a loss function in the estimation procedure. The functional combinations of the original gene expression data identified by the method are highly correlated with the patients' survival times and at the same time account for the variability in the gene expression levels. We apply our method to data sets from diffuse large B-cell lymphoma, lung adenocarcinoma and breast carcinoma studies to verify its effectiveness. The results from these analyses indicate that the proposed method works very well in identifying subgroups of patients with different risks of death or relapse and in predicting the risk of relapse or death based on the gene expression profiles measured from the tumor samples taken from the patients.

Artificial Intelligence↗

Statistical methods for analysis of time course gene expression data.

Since many biological systems or regulatory networks are dynamic systems, gene expression levels measured over different time points during a given biological process can often provide more insights about the underlying system. These gene expression data measured over time are often called the time-course gene expression data. One unique feature of such data is the time dependency of the gene expression levels for a given gene at different times or between two different genes. Statistical analysis needs to account for such dependency in order to make valid inferences. This paper presents several statistical methods for analyzing such time-course gene expression data, including the time-lagged correlation coefficient for analyzing the relationship between genes, a mixed-effects model with splines for clustering genes and for estimating missing gene expression data, and a new method for aligning gene expression profiles obtained under two experimental conditions and for identifying gene clusters that show significant changes between two experimental conditions. We used the yeast cell cycle gene expression data sets to illustrate these methods and obtained the biologically meaningful conclusions from these analyses.

Cell Cycle Proteins↗