PubMed Health⌕ Search

Biomedical subjects

Ming-Chih J Kao

Publications and source records attributed to Ming-Chih J Kao.

7 recordsLinked to original sources

Functional annotation and network reconstruction through cross-platform integration of microarray data.

The rapid accumulation of microarray data translates into a need for methods to effectively integrate data generated with different platforms. Here we introduce an approach, 2(nd)-order expression analysis, that addresses this challenge by first extracting expression patterns as meta-information from each data set (1(st)-order expression analysis) and then analyzing them across multiple data sets. Using yeast as a model system, we demonstrate two distinct advantages of our approach: we can identify genes of the same function yet without coexpression patterns and we can elucidate the cooperativities between transcription factors for regulatory network reconstruction by overcoming a key obstacle, namely the quantification of activities of transcription factors. Experiments reported in the literature and performed in our lab support a significant number of our predictions.

Algorithms↗

HumanUpstream and MouseUpstream: databases of promoter sequences in the human and mouse genomes.

Large-scale genome annotations, based largely on gene prediction programs, may be inaccurate in their predictions of transcription start sites, so that the identification of promoter regions remains unreliable. Here we focus on the identification of reliable gene promoter regions, critical to the understanding of transcriptional regulation. We report the construction of databases of upstream sequences Human Upstream and Mouse Upstream based on information from both the human and mouse genomes and the database of expressed sequence tags (dbEST). Using the ENSEMBL generic genome annotation system, our approach allows more reliable identification of transcript start sites, and therefore extraction of more reliable promoters regions. The Human Upstream and Human Upstream databases are available free of charge.

Animals↗

Determination of local statistical significance of patterns in Markov sequences with application to promoter element identification.

High-level eukaryotic genomes present a particular challenge to the computational identification of transcription factor binding sites (TFBSs) because of their long noncoding regions and large numbers of repeat elements. This is evidenced by the noisy results generated by most current methods. In this paper, we present a p-value-based scoring scheme using probability generating functions to evaluate the statistical significance of potential TFBSs. Furthermore, we introduce the local genomic context into the model so that candidate sites are evaluated based both on their similarities to known binding sites and on their contrasts against their respective local genomic contexts. We demonstrate that our approach is advantageous in the prediction of myogenin and MEF2 binding sites in the human genome. We also apply LMM to large-scale human binding site sequences in situ and found that, compared to current popular methods, LMM analysis can reduce false positive errors by more than 50% without compromising sensitivity. This improvement will be of importance to any subsequent algorithm that aims to detect regulatory modules based on known PSSMs.

Algorithms↗

GoSurfer: a graphical interactive tool for comparative analysis of large gene sets in Gene Ontology space.

UNLABELLED: The analysis of complex patterns of gene regulation is central to understanding the biology of cells, tissues and organisms. Patterns of gene regulation pertaining to specific biological processes can be revealed by a variety of experimental strategies, particularly microarrays and other highly parallel methods, which generate large datasets linking many genes. Although methods for detecting gene expression have improved substantially in recent years, understanding the physiological implications of complex patterns in gene expression data is a major challenge. This article presents GoSurfer, an easy-to-use graphical exploration tool with built-in statistical features that allow a rapid assessment of the biological functions represented in large gene sets. GoSurfer takes one or two list(s) of gene identifiers (Affymetrix probe set ID) as input and retrieves all the Gene Ontology (GO) terms associated with the input genes. GoSurfer visualises these GO terms in a hierarchical tree format. With GoSurfer, users can perform statistical tests to search for the GO terms that are enriched in the annotations of the input genes. These GO terms can be highlighted on the GO tree. Users can manipulate the GO tree in various ways and interactively query the genes associated with any GO term. The user-generated graphics can be saved as graphics files, and all the GO information related to the input genes can be exported as text files. AVAILABILITY: GoSurfer is a Windows-based program freely available for noncommercial use and can be downloaded at http://www.gosurfer.org. Datasets used to construct the trees shown in the figures in this article are available at http://www.gosurfer.org/download/GoSurfer.zip.

Computer Graphics↗

Novel mechanisms of T-cell and dendritic cell activation revealed by profiling of psoriasis on the 63,100-element oligonucleotide array.

A global picture of gene expression in the common immune-mediated skin disease, psoriasis, was obtained by interrogating the full set of Affymetrix GeneChips with psoriatic and control skin samples. We identified 1,338 genes with potential roles in psoriasis pathogenesis/maintenance and revealed many perturbed biological processes. A novel method for identifying transcription factor binding sites was also developed and applied to this dataset. Many of the identified sites are known to be involved in immune response and proliferation. An in-depth study of immune system genes revealed the presence of many regulating cytokines and chemokines within involved skin, and markers of dendritic cell (DC) activation in uninvolved skin. The combination of many CCR7+ T cells, DCs, and regulating chemokines in psoriatic lesions, together with the detection of DC activation markers in nonlesional skin, strongly suggests that the spatial organization of T cells and DCs could sustain chronic T-cell activation and persistence within focal skin regions.

Cell Separation↗

Chemical genetic modifier screens: small molecule trichostatin suppressors as probes of intracellular histone and tubulin acetylation.

Histone deacetylase (HDAC) inhibitors are being developed as new clinical agents in cancer therapy, in part because they interrupt cell cycle progression in transformed cell lines. To examine cell cycle arrest induced by HDAC inhibitor trichostatin A (TSA), a cytoblot cell-based screen was used to identify small molecule suppressors of this process. TSA suppressors (ITSAs) counteract TSA-induced cell cycle arrest, histone acetylation, and transcriptional activation. Hydroxamic acid-based HDAC inhibitors like TSA and suberoylanilide hydroxamic acid (SAHA) promote acetylation of cytoplasmic alpha-tubulin as well as histones, a modification also suppressed by ITSAs. Although tubulin acetylation appears irrelevant to cell cycle progression and transcription, it may play a role in other cellular processes. Small molecule suppressors such as the ITSAs, available from chemical genetic suppressor screens, may prove to be valuable probes of many biological processes.

Acetylation↗

Transitive functional annotation by shortest-path analysis of gene expression data.

Current methods for the functional analysis of microarray gene expression data make the implicit assumption that genes with similar expression profiles have similar functions in cells. However, among genes involved in the same biological pathway, not all gene pairs show high expression similarity. Here, we propose that transitive expression similarity among genes can be used as an important attribute to link genes of the same biological pathway. Based on large-scale yeast microarray expression data, we use the shortest-path analysis to identify transitive genes between two given genes from the same biological process. We find that not only functionally related genes with correlated expression profiles are identified but also those without. In the latter case, we compare our method to hierarchical clustering, and show that our method can reveal functional relationships among genes in a more precise manner. Finally, we show that our method can be used to reliably predict the function of unknown genes from known genes lying on the same shortest path. We assigned functions for 146 yeast genes that are considered as unknown by the Saccharomyces Genome Database and by the Yeast Proteome Database. These genes constitute around 5% of the unknown yeast ORFome.

Cell Nucleus↗