PubMed Health⌕ Search

Biomedical subjects

Birong Liao

Publications and source records attributed to Birong Liao.

4 recordsLinked to original sources

Too much data, but little inter-changeability: a lesson learned from mining public data on tissue specificity of gene expression.

BACKGROUND: The tissue expression pattern of a gene often provides an important clue to its potential role in a biological process. A vast amount of gene expression data have been and are being accumulated in public repository through different technology platforms. However, exploitations of these rich data sources remain limited in part due to issues of technology standardization. Our objective is to test the data comparability between SAGE and microarray technologies, through examining the expression pattern of genes under normal physiological states across variety of tissues. RESULTS: There are 42-54% of genes showing significant correlations in tissue expression patterns between SAGE and GeneChip, with 30-40% of genes whose expression patterns are positively correlated and 10-15% of genes whose expression patterns are negatively correlated at a statistically significant level (p = 0.05). Our analysis suggests that the discrepancy on the expression patterns derived from technology platforms is not likely from the heterogeneity of tissues used in these technologies, or other spurious correlations resulting from microarray probe design, abundance of genes, or gene function. The discrepancy can be partially explained by errors in the original assignment of SAGE tags to genes due to the evolution of sequence databases. In addition, sequence analysis has indicated that many SAGE tags and Affymetrix array probe sets are mapped to different splice variants or different sequence regions although they represent the same gene, which also contributes to the observed discrepancies between SAGE and array expression data. CONCLUSION: To our knowledge, this is the first report attempting to mine gene expression patterns across tissues using public data from different technology platforms. Unlike previous similar studies that only demonstrated the discrepancies between the two gene expression platforms, we carried out in-depth analysis to further investigate the cause for such discrepancies. Our study shows that the exploitation of rich public expression resource requires extensive knowledge about the technologies, and experiment. Informatic methodologies for better interoperability among platforms still remain a gap. One of the areas that can be improved practically is the accurate sequence mapping of SAGE tags and array probes to full-length genes.

Journal Article↗

Comparative analysis and integrative classification of NCI60 cell lines and primary tumors using gene expression profiling data.

BACKGROUND: NCI60 cell lines are derived from cancers of 9 tissue origins and have been invaluable in vitro models for cancer research and anti-cancer drug screen. Although extensive studies have been carried out to assess the molecular features of NCI60 cell lines related to cancer and their sensitivities to more than 100,000 chemical compounds, it remains unclear if and how well these cell lines represent or model their tumor tissues of origin. Identification and confirmation of correct origins of NCI60 cell lines are critical to their usage as model systems and to translate in vitro studies into clinical potentials. Here we report a direct comparison between NCI60 cell lines and primary tumors by analyzing global gene expression profiles. RESULTS: Comparative analysis suggested that 51 of 59 cell lines we analyzed represent their presumed tumors of origin. Taking advantage of available clinical information of primary tumor samples used to generate gene expression profiling data, we further classified those cell lines with the correct origins into different subtypes of cancer or different stages in cancer development. For example, 6 of 7 non-small cell lung cancer cell lines were classified as lung adenocarcinomas and all of them were classified into late stages in tumor progression. CONCLUSION: Taken together, we developed and applied a novel approach for systematic comparative analysis and integrative classification of NCI60 cell lines and primary tumors. Our results could provide guidance to the selection of appropriate cell lines for cancer research and pharmaceutical compound screenings. Moreover, this gene expression profile based approach can be generally applied to evaluate experimental model systems such as cell lines and animal models for human diseases.

Carcinoma, Non-Small-Cell Lung↗

A flexible integration and visualisation system for biomarker discovery.

Biological data have accumulated at an unprecedented pace as a result of improvements in molecular technologies. However, the translation of data into information, and subsequently into knowledge, requires the intricate interplay of data access, visualisation and interpretation. Biological data are complex and are organised either hierarchically or non-hierarchically. For non-hierarchically organised data, it is difficult to view relationships among biological facts. In addition, it is difficult to make changes in underlying data storage without affecting the visualisation interface. Here, we demonstrate a platform where non-hierarchically organised data can be visualised through the application of a customised hierarchy incorporating medical subject headings (MeSH) classifications. This platform gives users flexibility in updating and manipulation. It can also facilitate fresh scientific insight by highlighting biological impacts across different hierarchical branches. An example of the integration of biomarker information from the curated Proteome database using MeSH and the StarTree visualisation tool is presented.

Algorithms↗

Evaluating methods for classifying expression data.

An attractive application of expression technologies is to predict drug efficacy or safety using expression data of biomarkers. To evaluate the performance of various classification methods for building predictive models, we applied these methods on six expression datasets. These datasets were from studies using microarray technologies and had either two or more classes. From each of the original datasets, two subsets were generated to simulate two scenarios in biomarker applications. First, a 50-gene subset was used to simulate a candidate gene approach when it might not be practical to measure a large number of genes/biomarkers. Next, a 2000-gene subset was used to simulate a whole genome approach. We evaluated the relative performance of several classification methods by using leave-one-out cross-validation and bootstrap cross-validation. Although all methods perform well in both subsets for a relative easy dataset with two classes, differences in performance do exist among methods for other datasets. Overall, partial least squares discriminant analysis (PLS-DA) and support vector machines (SVM) outperform all other methods. We suggest a practical approach to take advantage of multiple methods in biomarker applications.

Algorithms↗