PubMed Health⌕ Search

Biomedical subjects

Li M Fu

Publications and source records attributed to Li M Fu.

5 recordsLinked to original sources

Evaluation of gene importance in microarray data based upon probability of selection.

BACKGROUND: Microarray devices permit a genome-scale evaluation of gene function. This technology has catalyzed biomedical research and development in recent years. As many important diseases can be traced down to the gene level, a long-standing research problem is to identify specific gene expression patterns linking to metabolic characteristics that contribute to disease development and progression. The microarray approach offers an expedited solution to this problem. However, it has posed a challenging issue to recognize disease-related genes expression patterns embedded in the microarray data. In selecting a small set of biologically significant genes for classifier design, the nature of high data dimensionality inherent in this problem creates substantial amount of uncertainty. RESULTS: Here we present a model for probability analysis of selected genes in order to determine their importance. Our contribution is that we show how to derive the P value of each selected gene in multiple gene selection trials based on different combinations of data samples and how to conduct a reliability analysis accordingly. The importance of a gene is indicated by its associated P value in that a smaller value implies higher information content from information theory. On the microarray data concerning the subtype classification of small round blue cell tumors, we demonstrate that the method is capable of finding the smallest set of genes (19 genes) with optimal classification performance, compared with results reported in the literature. CONCLUSION: In classifier design based on microarray data, the probability value derived from gene selection based on multiple combinations of data samples enables an effective mechanism for reducing the tendency of fitting local data particularities.

Algorithms↗

Multi-class cancer subtype classification based on gene expression signatures with reliability analysis.

Differential diagnosis among a group of histologically similar cancers poses a challenging problem in clinical medicine. Constructing a classifier based on gene expression signatures comprising multiple discriminatory molecular markers derived from microarray data analysis is an emerging trend for cancer diagnosis. To identify the best genes for classification using a small number of samples relative to the genome size remains the bottleneck of this approach, despite its promise. We have devised a new method of gene selection with reliability analysis, and demonstrated that this method can identify a more compact set of genes than other methods for constructing a classifier with optimum predictive performance for both small round blue cell tumors and leukemia. High consensus between our result and the results produced by methods based on artificial neural networks and statistical techniques confers additional evidence of the validity of our method. This study suggests a way for implementing a reliable molecular cancer classifier based on gene expression signatures.

Artificial Intelligence↗

TSGDB: a database system for tumor suppressor genes.

UNLABELLED: A Web-based database system was constructed and implemented that contains 174 tumor suppressor genes. The database homepage was created to accommodate these genes in a pull-down window so that each gene can be viewed individually in a separate Web page. Information displayed on each page includes gene name, aliases, source organism, chromosome location, expression cells/tissues, gene structure, protein size, gene functions and major reference sources. Queries to the database can be conducted through a user-friendly interface, and query results are returned in the HTML format on dynamically generated web pages. AVAILABILITY: The database is available at http://www.cise.ufl.edu/~yy1/HTML-TSGDB/Homepage.html (data files also at http://www.patcar.org/Databases/Tumor_Suppressor_Genes)

Abstracting and Indexing↗

Improving reliability of gene selection from microarray functional genomics data.

Constructing a classifier based on microarray gene expression data has recently emerged as an important problem for cancer classification. Recent results have suggested the feasibility of constructing such a classifier with reasonable predictive accuracy under the circumstance where only a small number of cancer tissue samples of known type are available. Difficulty arises from the fact that each sample contains the expression data of a vast number of genes and these genes may interact with one another. Selection of a small number of critical genes is fundamental to correctly analyze the otherwise overwhelming data. It is essential to use a multivariate approach for capturing the correlated structure in the data. However, the curse of dimensionality leads to the concern about the reliability of selected genes. Here, we present a new gene selection method in which error and repeatability of selected genes are assessed within the context of M-fold cross-validation. In particular, we show that the method is able to identify source variables underlying data generation.

Algorithms↗

Genome comparison of Mycobacterium tuberculosis and other bacteria.

The availability of the complete genome sequence of Mycobacterium tuberculosis allows its phylogenetic analysis based on the whole genome rather than single genes. As a genome-based tree is more representative of whole organisms and less inconsistent than single-gene trees, it could provide a better index for interpretation and inference about the origin and nature of species. The standard bacterial phylogeny based on 16S ribosomal RNA sequence comparison shows that M. tuberculosis is more related to Gram-positive than to Gram-negative bacteria. Our results based on genome comparison in terms of shared orthologous genes challenge this implication. We demonstrate that M. tuberculosis is more related to Gram-negative than to Gram-positive bacteria by a quantitative analysis on the genome tree. The numerical distance data derived from genome comparison and those from 16S rRNA comparison show high significant correlation, implying that conserved gene content carries a strong phylogenetic signature in evolution.

Animals↗