PubMed Health⌕ Search

Biomedical subjects

X Shirley Liu

Publications and source records attributed to X Shirley Liu.

15 recordsLinked to original sources

High-throughput mapping of the chromatin structure of human promoters.

Our understanding of how chromatin structure influences cellular processes such as transcription and replication has been limited by a lack of nucleosome-positioning data in human cells. We describe a high-resolution microarray approach combined with an analysis algorithm to examine nucleosome positioning in 3,692 promoters within seven human cell lines. Unlike unexpressed genes without transcription-preinitiation complexes at their promoters, expressed genes or genes containing preinitiation complexes exhibit characteristic nucleosome-free regions at their transcription start sites. The combination of these nucleosome data with chromatin immunoprecipitation-chip analyses reveals that the melanocyte master regulator microphthalmia-associated transcription factor (MITF) predominantly binds nucleosome-free regions, supporting the model that nucleosomes limit sequence accessibility. This study presents a global view of human nucleosome positioning and provides a high-throughput tool for analyzing chromatin structure in development and disease.

Algorithms↗

Genome-wide analysis of estrogen receptor binding sites.

The estrogen receptor is the master transcriptional regulator of breast cancer phenotype and the archetype of a molecular therapeutic target. We mapped all estrogen receptor and RNA polymerase II binding sites on a genome-wide scale, identifying the authentic cis binding sites and target genes, in breast cancer cells. Combining this unique resource with gene expression data demonstrates distinct temporal mechanisms of estrogen-mediated gene regulation, particularly in the case of estrogen-suppressed genes. Furthermore, this resource has allowed the identification of cis-regulatory sites in previously unexplored regions of the genome and the cooperating transcription factors underlying estrogen signaling in breast cancer.

Adaptor Proteins, Signal Transducing↗

Model-based analysis of tiling-arrays for ChIP-chip.

We propose a fast and powerful analysis algorithm, titled Model-based Analysis of Tiling-arrays (MAT), to reliably detect regions enriched by transcription factor chromatin immunoprecipitation (ChIP) on Affymetrix tiling arrays (ChIP-chip). MAT models the baseline probe behavior by considering probe sequence and copy number on each array. It standardizes the probe value through the probe model, eliminating the need for sample normalization. MAT uses an innovative function to score regions for ChIP enrichment, which allows robust P value and false discovery rate calculations. MAT can detect ChIP regions from a single ChIP sample, multiple ChIP samples, or multiple ChIP samples with controls with increasing accuracy. The single-array ChIP region detection feature minimizes the time and monetary costs for laboratories newly adopting ChIP-chip to test their protocols and antibodies and allows established ChIP-chip laboratories to identify samples with questionable quality that might contaminate their data. MAT is developed in open-source Python and is available at http://chip.dfci.harvard.edu/ approximately wli/MAT. The general framework presented here can be extended to other oligonucleotide microarrays and tiling array platforms.

Algorithms↗

Genome-wide in silico identification and analysis of cis natural antisense transcripts (cis-NATs) in ten species.

We developed a fast, integrative pipeline to identify cis natural antisense transcripts (cis-NATs) at genome scale. The pipeline mapped mRNAs and ESTs in UniGene to genome sequences in GoldenPath to find overlapping transcripts and combining information from coding sequence, poly(A) signal, poly(A) tail and splicing sites to deduce transcription orientation. We identified cis-NATs in 10 eukaryotic species, including 7830 candidate sense-antisense (SA) genes in 3915 SA pairs in human. The abundance of SA genes is remarkably low in worm and does not seem to be caused by the prevalence of operons. Hundreds of SA pairs are conserved across different species, even maintaining the same overlapping patterns. The convergent SA class is prevalent in fly, worm and sea squirt, but not in human or mouse as reported previously. The percentage of SA genes among imprinted genes in human and mouse is 24-47%, a range between the two previous reports. There is significant shortage of SA genes on Chromosome X in human and mouse but not in fly or worm, supporting X-inactivation in mammals as a possible cause. SA genes are over-represented in the catalytic activities and basic metabolism functions. All candidate cis-NATs can be downloaded from http://nats.cbi.pku.edu.cn/download/.

Animals↗

Identification of human STAT5-dependent gene regulatory elements based on interspecies homology.

STAT5 is a transcription factor essential for hematopoietic physiology. STAT5 functions to transduce signals from cytokines to the nucleus where it regulates gene expression. Although several important transcriptional targets of STAT5 are known, most remain unidentified. To identify novel STAT5 targets, we searched chromosomes 21 and 22 for clusters of STAT5 binding sites contained within regions of interspecies homology. We identified four such regions, including one with tandem STAT5 binding sites in the first intron of the NCAM2 gene. Unlike known STAT5 binding sites, this site is found within a very large intron and resides approximately 200 kb from the first coding exon of NCAM2. We demonstrate that this region confers STAT5-dependent transcriptional activity. We show that STAT5 binds in vivo to the NCAM2 intron in the NKL natural killer cell line and that this binding is induced by cytokines that activate STAT5. Neither STAT1 nor STAT3 bind to this region, despite sharing a consensus binding sequence with STAT5. Activation of STAT4 and STAT5 causes the accumulation of both of these STATs to the NCAM2 regulatory region. Therefore, using an informatics based approach to identify STAT5 targets, we have identified NCAM2 as both a STAT4- and STAT5-regulated gene, and we show that its expression is regulated by cytokines essential for natural killer cell survival and differentiation. This strategy may be an effective way to identify functional binding regions for transcription factors with known cognate binding sites anywhere in the genome.

Base Sequence↗

CEAS: cis-regulatory element annotation system.

The recent availability of high-density human genome tiling arrays enables biologists to conduct ChIP-chip experiments to locate the in vivo-binding sites of transcription factors in the human genome and explore the regulatory mechanisms. Once genomic regions enriched by transcription factor ChIP-chip are located, genome-scale downstream analyses are crucial but difficult for biologists without strong bioinformatics support. We designed and implemented the first web server to streamline the ChIP-chip downstream analyses. Given genome-scale ChIP regions, the cis-regulatory element annotation system (CEAS) retrieves repeat-masked genomic sequences, calculates GC content, plots evolutionary conservation, maps nearby genes and identifies enriched transcription factor-binding motifs. Biologists can utilize CEAS to retrieve useful information for ChIP-chip validation, assemble important knowledge to include in their publication and generate novel hypotheses (e.g. transcription factor cooperative partner) for further study. CEAS helps the adoption of ChIP-chip in mammalian systems and provides insights towards a more comprehensive understanding of transcriptional regulatory mechanisms. The URL of the server is http://ceas.cbi.pku.edu.cn.

Binding Sites↗

Genomic localization of RNA binding proteins reveals links between pre-mRNA processing and transcription.

Pre-mRNA processing often occurs in coordination with transcription thereby coupling these two key regulatory events. As such, many proteins involved in mRNA processing associate with the transcriptional machinery and are in proximity to DNA. This proximity allows for the mapping of the genomic associations of RNA binding proteins by chromatin immunoprecipitation (ChIP) as a way of determining their sites of action on the encoded mRNA. Here, we used ChIP combined with high-density microarrays to localize on the human genome three functionally distinct RNA binding proteins: the splicing factor polypyrimidine tract binding protein (PTBP1/hnRNP I), the mRNA export factor THO complex subunit 4 (ALY/THOC4), and the 3' end cleavage stimulation factor 64 kDa (CSTF2). We observed interactions at promoters, internal exons, and 3' ends of active genes. PTBP1 had biases toward promoters and often coincided with RNA polymerase II (RNA Pol II). The 3' processing factor, CSTF2, had biases toward 3' ends but was also observed at promoters. The mRNA processing and export factor, ALY, mapped to some exons but predominantly localized to introns and did not coincide with RNA Pol II. Because the RNA binding proteins did not consistently coincide with RNA Pol II, the data support a processing mechanism driven by reorganization of transcription complexes as opposed to a scanning mechanism. In sum, we present the mapping in mammalian cells of RNA binding proteins across a portion of the genome that provides insight into the transcriptional assembly of RNA-protein complexes.

Chromatin Immunoprecipitation↗

Transcriptional regulatory networks downstream of TAL1/SCL in T-cell acute lymphoblastic leukemia.

Aberrant expression of 1 or more transcription factor oncogenes is a critical component of the molecular pathogenesis of human T-cell acute lymphoblastic leukemia (T-ALL); however, oncogenic transcriptional programs downstream of T-ALL oncogenes are mostly unknown. TAL1/SCL is a basic helix-loop-helix (bHLH) transcription factor oncogene aberrantly expressed in 60% of human T-ALLs. We used chromatin immunoprecipitation (ChIP) on chip to identify 71 direct transcriptional targets of TAL1/SCL. Promoters occupied by TAL1 were also frequently bound by the class I bHLH proteins E2A and HEB, suggesting that TAL1/E2A as well as TAL1/HEB heterodimers play a role in transformation of T-cell precursors. Using RNA interference, we demonstrated that TAL1 is required for the maintenance of the leukemic phenotype in Jurkat cells and showed that TAL1 binding can be associated with either repression or activation of genes whose promoters occupied by TAL1, E2A, and HEB. In addition, oligonucleotide microarray analysis of RNA from 47 primary T-ALL samples showed specific expression signatures involving TAL1 targets in TAL1-expressing compared with -nonexpressing human T-ALLs. Our results indicate that TAL1 may act as a bifunctional transcriptional regulator (activator and repressor) at the top of a complex regulatory network that disrupts normal T-cell homeostasis and contributes to leukemogenesis.

Basic Helix-Loop-Helix Proteins↗

Chromosome-wide mapping of estrogen receptor binding reveals long-range regulation requiring the forkhead protein FoxA1.

Estrogen plays an essential physiologic role in reproduction and a pathologic one in breast cancer. The completion of the human genome has allowed the identification of the expressed regions of protein-coding genes; however, little is known concerning the organization of their cis-regulatory elements. We have mapped the association of the estrogen receptor (ER) with the complete nonrepetitive sequence of human chromosomes 21 and 22 by combining chromatin immunoprecipitation (ChIP) with tiled microarrays. ER binds selectively to a limited number of sites, the majority of which are distant from the transcription start sites of regulated genes. The unbiased sequence interrogation of the genuine chromatin binding sites suggests that direct ER binding requires the presence of Forkhead factor binding in close proximity. Furthermore, knockdown of FoxA1 expression blocks the association of ER with chromatin and estrogen-induced gene expression demonstrating the necessity of FoxA1 in mediating an estrogen response in breast cancer cells.

Animals↗

A boosting approach for motif modeling using ChIP-chip data.

MOTIVATION: Building an accurate binding model for a transcription factor (TF) is essential to differentiate its true binding targets from those spurious ones. This is an important step toward understanding gene regulation. RESULTS: This paper describes a boosting approach to modeling TF-DNA binding. Different from the widely used weight matrix model, which predicts TF-DNA binding based on a linear combination of position-specific contributions, our approach builds a TF binding classifier by combining a set of weight matrix based classifiers, thus yielding a non-linear binding decision rule. The proposed approach was applied to the ChIP-chip data of Saccharomyces cerevisiae. When compared with the weight matrix method, our new approach showed significant improvements on the specificity in a majority of cases.

Algorithms↗

A hidden Markov model for analyzing ChIP-chip experiments on genome tiling arrays and its application to p53 binding sequences.

MOTIVATION: Transcription factors (TFs) regulate gene expression by recognizing and binding to specific regulatory regions on the genome, which in higher eukaryotes can occur far away from the regulated genes. Recently, Affymetrix developed the high-density oligonucleotide arrays that tile all the non-repetitive sequences of the human genome at 35 bp resolution. This new array platform allows for the unbiased mapping of in vivo TF binding sequences (TFBSs) using Chromatin ImmunoPrecipitation followed by microarray experiments (ChIP-chip). The massive dataset generated from these experiments pose great challenges for data analysis. RESULTS: We developed a fast, scalable and sensitive method to extract TFBSs from ChIP-chip experiments on genome tiling arrays. Our method takes advantage of tiling array data from many experiments to normalize and model the behavior of each individual probe, and identifies TFBSs using a hidden Markov model (HMM). When applied to the data of p53 ChIP-chip experiments from an earlier study, our method discovered many new high confidence p53 targets including all the regions verified by quantitative PCR. Using a de novo motif finding algorithm MDscan, we also recovered the p53 motif from our HMM identified p53 target regions. Furthermore, we found substantial p53 motif enrichment in these regions comparing with both genomic background and the TFBSs identified earlier. Several of the newly identified p53 TFBSs are in the promoter region of known genes or associated with previously characterized p53-responsive genes. SUPPLEMENTARY INFORMATION: Available at the following URL http://genome.dfci.harvard.edu/~xsliu/HMMTiling/index.html.

Amino Acid Motifs↗

A suite of web-based programs to search for transcriptional regulatory motifs.

The identification of regulatory motifs is important for the study of gene expression. Here we present a suite of programs that we have developed to search for regulatory sequence motifs: (i) BioProspector, a Gibbs-sampling-based program for predicting regulatory motifs from co-regulated genes in prokaryotes or lower eukaryotes; (ii) CompareProspector, an extension to BioProspector which incorporates comparative genomics features to be used for higher eukaryotes; (iii) MDscan, a program for finding protein-DNA interaction sites from ChIP-on-chip targets. All three programs examine a group of sequences that may share common regulatory motifs and output a list of putative motifs as position-specific probability matrices, the individual sites used to construct the motifs and the location of each site on the input sequences. The web servers and executables can be accessed at http://seqmotifs.stanford.edu.

Algorithms↗

Eukaryotic regulatory element conservation analysis and identification using comparative genomics.

Comparative genomics is a promising approach to the challenging problem of eukaryotic regulatory element identification, because functional noncoding sequences may be conserved across species from evolutionary constraints. We systematically analyzed known human and Saccharomyces cerevisiae regulatory elements and discovered that human regulatory elements are more conserved between human and mouse than are background sequences. Although S. cerevisiae regulatory elements do not appear to be more conserved by comparison of S. cerevisiae to Schizosaccharomyces pombe, they are more conserved when compared with multiple other yeast genomes (Saccharomyces paradoxus, Saccharomyces mikatae, and Saccharomyces bayanus). Based on these analyses, we developed a sequence-motif-finding algorithm called CompareProspector, which extends Gibbs sampling by biasing the search in regions conserved across species. Using human-mouse comparison, CompareProspector identified known motifs for transcription factors Mef2, Myf, Srf, and Sp1 from a set of human-muscle-specific genes. It also discovered the NFAT motif from genes up-regulated by CD28 stimulation in T-cells, which implies the direct involvement of NFAT in mediating the CD28 stimulatory signal. Using Caenorhabditis elegans-Caenorhabditis briggsae comparison, CompareProspector found the PHA-4 motif and the UNC-86 motif. CompareProspector outperformed many other computational motif-finding programs, demonstrating the power of comparative genomics-based biased sampling in eukaryotic regulatory element identification.

Algorithms↗

Integrating regulatory motif discovery and genome-wide expression analysis.

We propose motif regressor for discovering sequence motifs upstream of genes that undergo expression changes in a given condition. The method combines the advantages of matrix-based motif finding and oligomer motif-expression regression analysis, resulting in high sensitivity and specificity. motif regressor is particularly effective in discovering expression-mediating motifs of medium to long width with multiple degenerate positions. When applied to Saccharomyces cerevisiae, motif regressor identified the ROX1 and YAP1 motifs from Rox1p and Yap1p overexpression experiments, respectively; predicted that Gcn4p may have increased activity in YAP1 deletion mutants; reported a group of motifs (including GCN4, PHO4, MET4, STRE, USR1, RAP1, M3A, and M3B) that may mediate the transcriptional response to amino acid starvation; and found all of the known cell-cycle regulation motifs from 18 expression microarrays over two cell cycles.

Algorithms↗

An algorithm for finding protein-DNA binding sites with applications to chromatin-immunoprecipitation microarray experiments.

Chromatin immunoprecipitation followed by cDNA microarray hybridization (ChIP-array) has become a popular procedure for studying genome-wide protein-DNA interactions and transcription regulation. However, it can only map the probable protein-DNA interaction loci within 1-2 kilobases resolution. To pinpoint interaction sites down to the base-pair level, we introduce a computational method, Motif Discovery scan (MDscan), that examines the ChIP-array-selected sequences and searches for DNA sequence motifs representing the protein-DNA interaction sites. MDscan combines the advantages of two widely adopted motif search strategies, word enumeration and position-specific weight matrix updating, and incorporates the ChIP-array ranking information to accelerate searches and enhance their success rates. MDscan correctly identified all the experimentally verified motifs from published ChIP-array experiments in yeast (STE12, GAL4, RAP1, SCB, MCB, MCM1, SFF, and SWI5), and predicted two motif patterns for the differential binding of Rap1 protein in telomere regions. In our studies, the method was faster and more accurate than several established motif-finding algorithms. MDscan can be used to find DNA motifs not only in ChIP-array experiments but also in other experiments in which a subgroup of the sequences can be inferred to contain relatively abundant motif sites. The MDscan web server can be accessed at http://BioProspector.stanford.edu/MDscan/.

Algorithms↗