PubMed Health⌕ Search

Biomedical subjects

Albin Sandelin

Publications and source records attributed to Albin Sandelin.

16 recordsLinked to original sources

Transcriptional and structural impact of TATA-initiation site spacing in mammalian core promoters.

BACKGROUND: The TATA box, one of the most well studied core promoter elements, is associated with induced, context-specific expression. The lack of precise transcription start site (TSS) locations linked with expression information has impeded genome-wide characterization of the interaction between TATA and the pre-initiation complex. RESULTS: Using a comprehensive set of 5.66 x 10(6) sequenced 5' cDNA ends from diverse tissues mapped to the mouse genome, we found that the TATA-TSS distance is correlated with the tissue specificity of the downstream transcript. To achieve tissue-specific regulation, the TATA box position relative to the TSS is constrained to a narrow window (-32 to -29), where positions -31 and -30 are the optimal positions for achieving high tissue specificity. Slightly larger spacings can be accommodated only when there is no optimally spaced initiation signal; in contrast, the TATA box like motifs found downstream of position -28 are generally nonfunctional. The strength of the TATA binding protein-DNA interaction plays a subordinate role to spacing in terms of tissue specificity. Furthermore, promoters with different TATA-TSS spacings have distinct features in terms of consensus sequence around the initiation site and distribution of alternative TSSs. Unexpectedly, promoters that have two dominant, consecutive TSSs are TATA depleted and have a novel GGG initiation site consensus. CONCLUSION: In this report we present the most comprehensive characterization of TATA-TSS spacing and functionality to date. The coupling of spacing to tissue specificity at the transcriptome level provides important clues as to the function of core promoters and the choice of TSS by the pre-initiation complex.

Animals↗

The complexity of the mammalian transcriptome.

A comprehensive understanding of protein and regulatory networks is strictly dependent on the complete description of the transcriptome of cells. After the determination of the genome sequence of several mammalian species, gene identification is based on in silico predictions followed by evidence of transcription. Conservative estimates suggest that there are about 20,000 protein-encoding genes in the mammalian genome. In the last few years the combination of full-length cDNA cloning, cap-analysis gene expression (CAGE) tag sequencing and tiling arrays experiments have unveiled unexpected additional complexities in the transcriptome. Here we describe the current view of the mammalian transcriptome focusing on transcripts diversity, the growing non-coding RNA world, the organization of transcriptional units in the genome and promoter structures. In-depth analysis of the brain transcriptome has been challenging due to the cellular complexity of this organ. Here we present a computational analysis of CAGE data from different regions of the central nervous system, suggesting distinctive mechanisms of brain-specific transcription.

Alternative Splicing↗

A global genomic transcriptional code associated with CNS-expressed genes.

Highly conserved non-coding DNA regions (HCNR) occur frequently in vertebrate genomes, but their functional roles remain unclear. Here, we provide evidence that a large portion of HCNRs are enriched for binding sites for Sox, POU and Homeodomain transcription factors, and such HCNRs can act as cis-regulatory regions active in neural stem cells. Strikingly, these HCNRs are linked to several hundreds of genes expressed in the developing CNS and they may exert locus-wide regulatory effects on multiple genes flanking their genomic location. Moreover, these data imply a unifying transcriptional logic for a large set of CNS-expressed genes in which Sox and POU proteins act as generic promoters of transcription while Homeodomain proteins control the spatial expression of genes through active repression.

Animals↗

Evolutionary turnover of mammalian transcription start sites.

Alignments of homologous genomic sequences are widely used to identify functional genetic elements and study their evolution. Most studies tacitly equate homology of functional elements with sequence homology. This assumption is violated by the phenomenon of turnover, in which functionally equivalent elements reside at locations that are nonorthologous at the sequence level. Turnover has been demonstrated previously for transcription-factor-binding sites. Here, we show that transcription start sites of equivalent genes do not always reside at equivalent locations in the human and mouse genomes. We also identify two types of partial turnover, illustrating evolutionary pathways that could lead to complete turnover. These findings suggest that the signals encoding transcription start sites are highly flexible and evolvable, and have cautionary implications for the use of sequence-level conservation to detect gene regulatory elements.

Animals↗

Genome-wide analysis of mammalian promoter architecture and evolution.

Mammalian promoters can be separated into two classes, conserved TATA box-enriched promoters, which initiate at a well-defined site, and more plastic, broad and evolvable CpG-rich promoters. We have sequenced tags corresponding to several hundred thousand transcription start sites (TSSs) in the mouse and human genomes, allowing precise analysis of the sequence architecture and evolution of distinct promoter classes. Different tissues and families of genes differentially use distinct types of promoters. Our tagging methods allow quantitative analysis of promoter usage in different tissues and show that differentially regulated alternative TSSs are a common feature in protein-coding genes and commonly generate alternative N termini. Among the TSSs, we identified new start sites associated with the majority of exons and with 3' UTRs. These data permit genome-scale identification of tissue-specific promoters and analysis of the cis-acting elements associated with them.

3' Untranslated Regions↗

A new generation of JASPAR, the open-access repository for transcription factor binding site profiles.

JASPAR is the most complete open-access collection of transcription factor binding site (TFBS) matrices. In this new release, JASPAR grows into a meta-database of collections of TFBS models derived by diverse approaches. We present JASPAR CORE--an expanded version of the original, non-redundant collection of annotated, high-quality matrix-based transcription factor binding profiles, JASPAR FAM--a collection of familial TFBS models and JASPAR phyloFACTS--a set of matrices computationally derived from statistically overrepresented, evolutionarily conserved regulatory region motifs from mammalian genomes. JASPAR phyloFACTS serves as a non-redundant extension to JASPAR CORE, enhancing the overall breadth of JASPAR for promoter sequence analysis. The new release of JASPAR is available at http://jaspar.genereg.net.

Animals↗

Dynamic usage of transcription start sites within core promoters.

BACKGROUND: Mammalian promoters do not initiate transcription at single, well defined base pairs, but rather at multiple, alternative start sites spread across a region. We previously characterized the static structures of transcription start site usage within promoters at the base pair level, based on large-scale sequencing of transcript 5' ends. RESULTS: In the present study we begin to explore the internal dynamics of mammalian promoters, and demonstrate that start site selection within many mouse core promoters varies among tissues. We also show that this dynamic usage of start sites is associated with CpG islands, broad and multimodal promoter structures, and imprinting. CONCLUSION: Our results reveal a new level of biologic complexity within promoters--fine-scale regulation of transcription starting events at the base pair level. These events are likely to be related to epigenetic transcriptional regulation.

Animals↗

Exploring hepatic hormone actions using a compilation of gene expression profiles.

BACKGROUND: Microarray analysis is attractive within the field of endocrine research because regulation of gene expression is a key mechanism whereby hormones exert their actions. Knowledge discovery and testing of hypothesis based on information-rich expression profiles promise to accelerate discovery of physiologically relevant hormonal mechanisms of action. However, most studies so-far concentrate on the analysis of actions of single hormones and few examples exist that attempt to use compilation of different hormone-regulated expression profiles to gain insight into how hormone act to regulate tissue physiology. This report illustrates how a meta-analysis of multiple transcript profiles obtained from a single tissue, the liver, can be used to evaluate relevant hypothesis and discover novel mechanisms of hormonal action. We have evaluated the differential effects of Growth Hormone (GH) and estrogen in the regulation of hepatic gender differentiated gene expression as well as the involvement of sterol regulatory element-binding proteins (SREBPs) in the hepatic actions of GH and thyroid hormone. RESULTS: Little similarity exists between liver transcript profiles regulated by 17-alpha-ethinylestradiol and those induced by the continuos infusion of bGH. On the other hand, strong correlations were found between both profiles and the female enriched transcript profile. Therefore, estrogens have feminizing effects in male rat liver which are different from those induced by GH. The similarity between bGH and T3 were limited to a small group of genes, most of which are involved in lipogenesis. An in silico promoter analysis of genes rapidly regulated by thyroid hormone predicted the activation of SREBPs by short-term treatment in vivo. It was further demonstrated that proteolytic processing of SREBP1 in the endoplasmic reticulum might contribute to the rapid actions of T3 on these genes. CONCLUSION: This report illustrates how a meta-analysis of multiple transcript profiles can be used to link knowledge concerning endocrine physiology to hormonally induced changes in gene expression. We conclude that both GH and estrogen are important determinants of gender-related differences in hepatic gene expression. Rapid hepatic thyroid hormone effects affect genes involved in lipogenesis possibly through the induction of SREBP1 proteolytic processing.

Animals↗

Arrays of ultraconserved non-coding regions span the loci of key developmental genes in vertebrate genomes.

BACKGROUND: Evolutionarily conserved sequences within or adjoining orthologous genes often serve as critical cis-regulatory regions. Recent studies have identified long, non-coding genomic regions that are perfectly conserved between human and mouse, termed ultra-conserved regions (UCRs). Here, we focus on UCRs that cluster around genes involved in early vertebrate development; genes conserved over 450 million years of vertebrate evolution. RESULTS: Based on a high resolution detection procedure, our UCR set enables novel insights into vertebrate genome organization and regulation of developmentally important genes. We find that the genomic positions of deeply conserved UCRs are strongly associated with the locations of genes encoding key regulators of development, with particularly strong positional correlation to transcription factor-encoding genes. Of particular importance is the observation that most UCRs are clustered into arrays that span hundreds of kilobases around their presumptive target genes. Such a hallmark signature is present around several uncharacterized human genes predicted to encode developmentally important DNA-binding proteins. CONCLUSION: The genomic organization of UCRs, combined with previous findings, suggests that UCRs act as essential long-range modulators of gene expression. The exceptional sequence conservation and clustered structure suggests that UCR-mediated molecular events involve greater complexity than traditional DNA binding by transcription factors. The high-resolution UCR collection presented here provides a wealth of target sequences for future experimental studies to determine the nature of the biochemical mechanisms involved in the preservation of arrays of nearly identical non-coding sequences over the course of vertebrate evolution.

Animals↗

Prediction of nuclear hormone receptor response elements.

The nuclear receptor (NR) class of transcription factors controls critical regulatory events in key developmental processes, homeostasis maintenance, and medically important diseases and conditions. Identification of the members of a regulon controlled by a NR could provide an accelerated understanding of development and disease. New bioinformatics methods for the analysis of regulatory sequences are required to address the complex properties associated with known regulatory elements targeted by the receptors because the standard methods for binding site prediction fail to reflect the diverse target site configurations. We have constructed a flexible Hidden Markov Model framework capable of predicting NHR binding sites. The model allows for variable spacing and orientation of half-sites. In a genome-scale analysis enabled by the model, we show that NRs in Fugu rubripes have a significant cross-regulatory potential. The model is implemented in a web interface, freely available for academic researchers, available at http://mordor.cgb.ki.se/NHR-scan.

Algorithms↗

ConSite: web-based prediction of regulatory elements using cross-species comparison.

ConSite is a user-friendly, web-based tool for finding cis-regulatory elements in genomic sequences. Predictions are based on the integration of binding site prediction generated with high-quality transcription factor models and cross-species comparison filtering (phylogenetic footprinting). By incorporating evolutionary constraints, selectivity is increased by an order of magnitude as compared to single-sequence analysis. ConSite offers several unique features, including an interactive expert system for retrieving orthologous regulatory sequences. Programming modules and biological databases that form the foundation of the ConSite service are freely available to the research community. ConSite is available at http:/www.phylofoot.org/consite.

Animals↗

Constrained binding site diversity within families of transcription factors enhances pattern discovery bioinformatics.

Diverse computational and experimental efforts are required to elucidate the control circuitry regulating the transcription of human genes. The fusion of gene-specific promoter analyses with large microarray studies and bioinformatics advances has produced optimism that significant progress can be made in unravelling this complex network. Within bioinformatics, past emphasis for improved pattern discovery has been placed upon "phylogenetic footprinting", the identification of sequences conserved over moderate periods of evolution (e.g. human and mouse comparisons). We introduce a new direction in bioinformatics based on the constraints imposed by the structures of DNA-binding proteins. For most structurally related families of transcription factors, there are clear similarities in the sequences of the sites to which they bind. On the basis of this observation, we construct familial binding profiles for well-characterized transcription factor families. The profiles are shown to classify correctly the structural class of mediating transcription factors for novel motifs in 88% of cases. By incorporating the familial profiles into pattern discovery procedures, we demonstrate that functional binding sites can be found in genomic sequences of dramatically greater length than is possible otherwise. Thus, incorporating familial models can overcome the signal-to-noise challenge that has hindered the transition from microarray data to regulatory control sequences for human genes. Biochemically motivated constraints upon sequence diversity of binding sites will complement the genetically motivated constraints imposed in "phylogenetic footprinting" algorithms.

Algorithms↗

JASPAR: an open-access database for eukaryotic transcription factor binding profiles.

The analysis of regulatory regions in genome sequences is strongly based on the detection of potential transcription factor binding sites. The preferred models for representation of transcription factor binding specificity have been termed position-specific scoring matrices. JASPAR is an open-access database of annotated, high-quality, matrix-based transcription factor binding site profiles for multicellular eukaryotes. The profiles were derived exclusively from sets of nucleotide sequences experimentally demonstrated to bind transcription factors. The database is complemented by a web interface for browsing, searching and subset selection, an online sequence analysis utility and a suite of programming tools for genome-wide and comparative genomic analysis of regulatory regions. JASPAR is available at http://jaspar. cgb.ki.se.

Animals↗

Integrated analysis of yeast regulatory sequences for biologically linked clusters of genes.

Dramatic progress in deciphering the regulatory controls in Saccharomyces cerevisiae has been enabled by the fusion of high-throughput genomics technologies with advanced sequence analysis algorithms. Sets of genes likely to function together and with similar expression profiles have been identified in diverse studies. By fusing an advanced pattern recognition algorithm for identification of transcription factor binding sites with a new method for the quantitative comparison of binding properties of transcription factors, we provide an integrated means to move from expression data to biological insights. The Yeast Regulatory Sequence Analysis system, YRSA, combines standard functions with a novel pattern characterization procedure in an intuitive interface designed for use by a broad range of scientists. The features of the system include automated retrieval of user-defined promoter sequences, binding site discovery by pattern recognition, graphical displays of the observed pattern and positions of similar sequences in the specified genes, and comparison of the new pattern against a collection of binding patterns for characterized transcription factors. The comprehensive YRSA system was used to study the regulatory mechanisms of yeast regulons. Analysis of the regulatory controls of a battery of genes induced by DNA damaging agents supports a putative mediating role for the cell-cycle checkpoint regulatory element MCB. YRSA is available at http://yrsa.cgb.ki.se. [YRSA: ancient Scandinavian name meaning old she-bear (Latin Ursus arctos = brown bear/grizzly).]

Algorithms↗

Identification of conserved regulatory elements by comparative genome analysis.

BACKGROUND: For genes that have been successfully delineated within the human genome sequence, most regulatory sequences remain to be elucidated. The annotation and interpretation process requires additional data resources and significant improvements in computational methods for the detection of regulatory regions. One approach of growing popularity is based on the preferential conservation of functional sequences over the course of evolution by selective pressure, termed 'phylogenetic footprinting'. Mutations are more likely to be disruptive if they appear in functional sites, resulting in a measurable difference in evolution rates between functional and non-functional genomic segments. RESULTS: We have devised a flexible suite of methods for the identification and visualization of conserved transcription-factor-binding sites. The system reports those putative transcription-factor-binding sites that are both situated in conserved regions and located as pairs of sites in equivalent positions in alignments between two orthologous sequences. An underlying collection of metazoan transcription-factor-binding profiles was assembled to facilitate the study. This approach results in a significant improvement in the detection of transcription-factor-binding sites because of an increased signal-to-noise ratio, as demonstrated with two sets of promoter sequences. The method is implemented as a graphical web application, ConSite, which is at the disposal of the scientific community at http://www.phylofoot.org/. CONCLUSIONS: Phylogenetic footprinting dramatically improves the predictive selectivity of bioinformatic approaches to the analysis of promoter sequences. ConSite delivers unparalleled performance using a novel database of high-quality binding models for metazoan transcription factors. With a dynamic interface, this bioinformatics tool provides broad access to promoter analysis with phylogenetic footprinting.

Algorithms↗