PubMed Health⌕ Search

Biomedical subjects

Martin C Frith

Publications and source records attributed to Martin C Frith.

At least 19 recordsLinked to original sources

Evolutionary turnover of mammalian transcription start sites.

Alignments of homologous genomic sequences are widely used to identify functional genetic elements and study their evolution. Most studies tacitly equate homology of functional elements with sequence homology. This assumption is violated by the phenomenon of turnover, in which functionally equivalent elements reside at locations that are nonorthologous at the sequence level. Turnover has been demonstrated previously for transcription-factor-binding sites. Here, we show that transcription start sites of equivalent genes do not always reside at equivalent locations in the human and mouse genomes. We also identify two types of partial turnover, illustrating evolutionary pathways that could lead to complete turnover. These findings suggest that the signals encoding transcription start sites are highly flexible and evolvable, and have cautionary implications for the use of sequence-level conservation to detect gene regulatory elements.

Animals↗

Genome-wide analysis of mammalian promoter architecture and evolution.

Mammalian promoters can be separated into two classes, conserved TATA box-enriched promoters, which initiate at a well-defined site, and more plastic, broad and evolvable CpG-rich promoters. We have sequenced tags corresponding to several hundred thousand transcription start sites (TSSs) in the mouse and human genomes, allowing precise analysis of the sequence architecture and evolution of distinct promoter classes. Different tissues and families of genes differentially use distinct types of promoters. Our tagging methods allow quantitative analysis of promoter usage in different tissues and show that differentially regulated alternative TSSs are a common feature in protein-coding genes and commonly generate alternative N termini. Among the TSSs, we identified new start sites associated with the majority of exons and with 3' UTRs. These data permit genome-scale identification of tissue-specific promoters and analysis of the cis-acting elements associated with them.

3' Untranslated Regions↗

Pseudo-messenger RNA: phantoms of the transcriptome.

The mammalian transcriptome harbours shadowy entities that resist classification and analysis. In analogy with pseudogenes, we define pseudo-messenger RNA to be RNA molecules that resemble protein-coding mRNA, but cannot encode full-length proteins owing to disruptions of the reading frame. Using a rigorous computational pipeline, which rules out sequencing errors, we identify 10,679 pseudo-messenger RNAs (approximately half of which are transposon-associated) among the 102,801 FANTOM3 mouse cDNAs: just over 10% of the FANTOM3 transcriptome. These comprise not only transcribed pseudogenes, but also disrupted splice variants of otherwise protein-coding genes. Some may encode truncated proteins, only a minority of which appear subject to nonsense-mediated decay. The presence of an excess of transcripts whose only disruptions are opal stop codons suggests that there are more selenoproteins than currently estimated. We also describe compensatory frameshifts, where a segment of the gene has changed frame but remains translatable. In summary, we survey a large class of non-standard but potentially functional transcripts that are likely to encode genetic information and effect biological processes in novel ways. Many of these transcripts do not correspond cleanly to any identifiable object in the genome, implying fundamental limits to the goal of annotating all functional elements at the genome sequence level.

Animals↗

Clusters of internally primed transcripts reveal novel long noncoding RNAs.

Non-protein-coding RNAs (ncRNAs) are increasingly being recognized as having important regulatory roles. Although much recent attention has focused on tiny 22- to 25-nucleotide microRNAs, several functional ncRNAs are orders of magnitude larger in size. Examples of such macro ncRNAs include Xist and Air, which in mouse are 18 and 108 kilobases (Kb), respectively. We surveyed the 102,801 FANTOM3 mouse cDNA clones and found that Air and Xist were present not as single, full-length transcripts but as a cluster of multiple, shorter cDNAs, which were unspliced, had little coding potential, and were most likely primed from internal adenine-rich regions within longer parental transcripts. We therefore conducted a genome-wide search for regional clusters of such cDNAs to find novel macro ncRNA candidates. Sixty-six regions were identified, each of which mapped outside known protein-coding loci and which had a mean length of 92 Kb. We detected several known long ncRNAs within these regions, supporting the basic rationale of our approach. In silico analysis showed that many regions had evidence of imprinting and/or antisense transcription. These regions were significantly associated with microRNAs and transcripts from the central nervous system. We selected eight novel regions for experimental validation by northern blot and RT-PCR and found that the majority represent previously unrecognized noncoding transcripts that are at least 10 Kb in size and predominantly localized in the nucleus. Taken together, the data not only identify multiple new ncRNAs but also suggest the existence of many more macro ncRNAs like Xist and Air.

Animals↗

The abundance of short proteins in the mammalian proteome.

Short proteins play key roles in cell signalling and other processes, but their abundance in the mammalian proteome is unknown. Current catalogues of mammalian proteins exhibit an artefactual discontinuity at a length of 100 aa, so that protein abundance peaks just above this length and falls off sharply below it. To clarify the abundance of short proteins, we identify proteins in the FANTOM collection of mouse cDNAs by analysing synonymous and non-synonymous substitutions with the computer program CRITICA. This analysis confirms that there is no real discontinuity at length 100. Roughly 10% of mouse proteins are shorter than 100 aa, although the majority of these are variants of proteins longer than 100 aa. We identify many novel short proteins, including a "dark matter" subset containing ones that lack detectable homology to other known proteins. Translation assays confirm that some of these novel proteins can be translated and localised to the secretory pathway.

Amino Acid Sequence↗

Discrimination of non-protein-coding transcripts from protein-coding mRNA.

Several recent studies indicate that mammals and other organisms produce large numbers of RNA transcripts that do not correspond to known genes. It has been suggested that these transcripts do not encode proteins, but may instead function as RNAs. However, discrimination of coding and non-coding transcripts is not straightforward, and different laboratories have used different methods, whose ability to perform this discrimination is unclear. In this study, we examine ten bioinformatic methods that assess protein-coding potential and compare their ability and congruency in the discrimination of non-coding from coding sequences, based on four underlying principles: open reading frame size, sequence similarity to known proteins or protein domains, statistical models of protein-coding sequence, and synonymous versus non-synonymous substitution rates. Despite these different approaches, the methods show broad concordance, suggesting that coding and non-coding transcripts can, in general, be reliably discriminated, and that many of the recently discovered extra-genic transcripts are indeed non-coding. Comparison of the methods indicates reasons for unreliable predictions, and approaches to increase confidence further. Conversely and surprisingly, our analyses also provide evidence that as much as approximately 10% of entries in the manually curated protein database Swiss-Prot are erroneous translations of actually non-coding transcripts.

Algorithms↗

Dynamic usage of transcription start sites within core promoters.

BACKGROUND: Mammalian promoters do not initiate transcription at single, well defined base pairs, but rather at multiple, alternative start sites spread across a region. We previously characterized the static structures of transcription start site usage within promoters at the base pair level, based on large-scale sequencing of transcript 5' ends. RESULTS: In the present study we begin to explore the internal dynamics of mammalian promoters, and demonstrate that start site selection within many mouse core promoters varies among tissues. We also show that this dynamic usage of start sites is associated with CpG islands, broad and multimodal promoter structures, and imprinting. CONCLUSION: Our results reveal a new level of biologic complexity within promoters--fine-scale regulation of transcription starting events at the base pair level. These events are likely to be related to epigenetic transcriptional regulation.

Animals↗

Experimental validation of the regulated expression of large numbers of non-coding RNAs from the mouse genome.

Recent large-scale analyses of mainly full-length cDNA libraries generated from a variety of mouse tissues indicated that almost half of all representative cloned sequences did not contain an apparent protein-coding sequence, and were putatively derived from non-protein-coding RNA (ncRNA) genes. However, many of these clones were singletons and the majority were unspliced, raising the possibility that they may be derived from genomic DNA or unprocessed pre-mRNA contamination during library construction, or alternatively represent nonspecific "transcriptional noise." Here we show, using reverse transcriptase-dependent PCR, microarray, and Northern blot analyses, that many of these clones were derived from genuine transcripts of unknown function whose expression appears to be regulated. The ncRNA transcripts have larger exons and fewer introns than protein-coding transcripts. Analysis of the genomic landscape around these sequences indicates that some cDNA clones were produced not from terminal poly(A) tracts but internal priming sites within longer transcripts, only a minority of which is encompassed by known genes. A significant proportion of these transcripts exhibit tissue-specific expression patterns, as well as dynamic changes in their expression in macrophages following lipopolysaccharide stimulation. Taken together, the data provide strong support for the conclusion that ncRNAs are an important, regulated component of the mammalian transcriptome.

Animals↗

Rapid evolution of noncoding RNAs: lack of conservation does not mean lack of function.

The mammalian transcriptome contains many non-protein-coding RNAs (ncRNAs), but most of these are of unclear significance and lack strong sequence conservation, prompting suggestions that they might be non-functional. However, certain long functional ncRNAs such as Air and Xist are also poorly conserved. In this article, we systematically analyzed the conservation of several groups of functional ncRNAs, including miRNAs, snoRNAs and longer ncRNAs whose function has been either documented or confidently predicted. As expected, miRNAs and snoRNAs were highly conserved. By contrast, the longer functional non-micro, non-sno ncRNAs were much less conserved with many displaying rapid sequence evolution. Our findings suggest that longer ncRNAs are under the influence of different evolutionary constraints and that the lack of conservation displayed by the thousands of candidate ncRNAs does not necessarily signify an absence of function.

Animals↗

Assessing computational tools for the discovery of transcription factor binding sites.

The prediction of regulatory elements is a problem where computational methods offer great hope. Over the past few years, numerous tools have become available for this task. The purpose of the current assessment is twofold: to provide some guidance to users regarding the accuracy of currently available tools in various settings, and to provide a benchmark of data sets for assessing future tools.

Amino Acid Motifs↗

CARRIE web service: automated transcriptional regulatory network inference and interactive analysis.

We present an intuitive and interactive web service for CARRIE (Computational Ascertainment of Regulatory Relationships Inferred from Expression). CARRIE is a computational method that analyzes microarray and promoter sequence data to infer a transcriptional regulatory network from the response to a specific stimulus. This service displays an interactive graph of the inferred network and provides easy access to the evidence for the involvement of each gene in the network. We provide functionality to include network data in KEGG XML (KGML) format in this graph. Our service also provides Gene Ontology annotation to aid the user in forming hypotheses about the role of each gene in the cellular response. The CARRIE web service is freely available at http://zlab.bu.edu/CARRIE-web.

Binding Sites↗

MotifViz: an analysis and visualization tool for motif discovery.

Detecting overrepresented known transcription factor binding motifs in a set of promoter sequences of co-regulated genes has become an important approach to deciphering transcriptional regulatory mechanisms. In this paper, we present an interactive web server, MotifViz, for three motif discovery programs, Clover, Rover and Motifish, covering most available flavors of algorithms for achieving this goal. For comparison, we have also implemented the simple motif-matching program Possum. MotifViz provides uniform and intuitive input and output formats for all four programs. It can be accessed at http://biowulf.bu.edu/MotifViz.

Algorithms↗

Genomic targets of nuclear estrogen receptors.

Estrogen influences the physiology of many target tissues in both women and men. The long-term effects of estrogen are mediated predominantly by nuclear estrogen receptors (ERs) functioning as DNA-binding transcription factors. Tissue-specific responses to estrogen therefore result from regulation of different sets of genes. However, it remains perplexing as to what regulatory sequence contexts specify distinct genomic responses. First, this review classifies estrogen response sequences in mammalian target genes. Of note, around one third of known human target genes associate only indirectly with ER, through intermediary transcription factor(s). Then, computational approaches are presented both for refining direct ER-binding sites and for formulating hypotheses regarding the overall genomic expression pattern. Surprisingly, limited evolutionary conservation of specific estrogen-responsive sites is observed between human and mouse. Finally, consideration of the cellular functions of regulated human genes suggests links between particular biological roles and specific types of estrogen response elements, although with the important caveat that only a restricted set of target genes is available. These analyses support the view that specific, hormone-driven gene expression programs can result from the interplay of environmental and cellular cues with the distinct types of estrogen-response sequences.

Animals↗

Characterization of genomic organization of the adenosine A2A receptor gene by molecular and bioinformatics analyses.

The adenosine A(2A) receptor (A(2A)R) is abundantly expressed in brain and emerging as an important therapeutic target for Parkinson's disease and potentially other neuropsychiatric disorders. To understand the molecular mechanisms of A(2A)R gene expression, we have characterized the genomic organization of the mouse and human A(2A)R genes by molecular and bioinformatic analyses. Three new exons (m1A, m1B and m1C) encoding the 5' untranslated regions (5'-UTRs) of mouse A(2A)R mRNA were identified by rapid amplification of 5' cDNA end (5' RACE), RT-PCR analysis and genome sequence analyses. Similar bioinformatics analysis also suggested six variants of the non-coding "exon 1" (h1A, h1B, h1C, h1D, h1E and h1F) in the human A(2A)R gene, which were confirmed by RT-PCR analysis, while three of the human exon 1 variants (h1D, h1E and h1F) were likewise verified by 5' oligonucleotide capping analysis suggesting multiple transcription start sites. Importantly, RT-PCR and quantitative PCR analysis demonstrated that the A(2A)R transcripts with different exon 1 variants displayed tissue-specific expression patterns. For instance, the mouse exon m1A mRNA was detected only in brain (specifically striatum) and the human exon h1D mRNA in lymphoreticular system. Furthermore, the determination of the three new transcription start sites of human A(2A)R gene by 5' oligonucleotide capping and bioinformatics analyses led to the identification of three corresponding promoter regions which contain several important cis elements, providing additional target for further molecular dissection of A(2A)R gene expression. Finally, our analysis indicates that A(2A)R mRNA and a novel transcript partially overlapping with the 3' exon h3, but in opposite orientation to the A(2A)R gene, could conceivably form duplexes to mutually regulate transcript expression. Thus, combined molecular and bioinformatics analyses revealed a new A(2A)R genomic structure, with conserved coding exons 2 and 3 and divergent, tissue-specific exon 1 variants encoding for 5'-UTR. This raises the possibility of generating multiple tissue-specific A(2A)R mRNA species by alternative promoters with varying regulatory susceptibility.

Animals↗

Detection of functional DNA motifs via statistical over-representation.

The interaction of proteins with DNA recognition motifs regulates a number of fundamental biological processes, including transcription. To understand these processes, we need to know which motifs are present in a sequence and which factors bind to them. We describe a method to screen a set of DNA sequences against a precompiled library of motifs, and assess which, if any, of the motifs are statistically over- or under-represented in the sequences. Over-represented motifs are good candidates for playing a functional role in the sequences, while under-representation hints that if the motif were present, it would have a harmful dysregulatory effect. We apply our method (implemented as a computer program called Clover) to dopamine-responsive promoters, sequences flanking binding sites for the transcription factor LSF, sequences that direct transcription in muscle and liver, and Drosophila segmentation enhancers. In each case Clover successfully detects motifs known to function in the sequences, and intriguing and testable hypotheses are made concerning additional motifs. Clover compares favorably with an ab initio motif discovery algorithm based on sequence alignment, when the motif library includes only a homolog of the factor that actually regulates the sequences. It also demonstrates superior performance over two contingency table based over-representation methods. In conclusion, Clover has the potential to greatly accelerate characterization of signals that regulate transcription.

Animals↗

Site2genome: locating short DNA sequences in whole genomes.

SUMMARY: Many biological papers describe short, functional DNA sites without specifying their exact positions in the genome. We have developed a Web server that automates the tedious task of locating such sites in eukaryotic genomes, thus giving access to the context of rich annotations that are increasingly available for genome sequences. AVAILABILITY: http://zlab.bu.edu/site2genome/

Algorithms↗

Finding functional sequence elements by multiple local alignment.

Algorithms that detect and align locally similar regions of biological sequences have the potential to discover a wide variety of functional motifs. Two theoretical contributions to this classic but unsolved problem are presented here: a method to determine the width of the aligned motif automatically; and a technique for calculating the statistical significance of alignments, i.e. an assessment of whether the alignments are stronger than those that would be expected to occur by chance among random, unrelated sequences. Upon exploring variants of the standard Gibbs sampling technique to optimize the alignment, we discovered that simulated annealing approaches perform more efficiently. Finally, we conduct failure tests by applying the algorithm to increasingly difficult test cases, and analyze the manner of and reasons for eventual failure. Detection of transcription factor-binding motifs is limited by the motifs' intrinsic subtlety rather than by inadequacy of the alignment optimization procedure.

Algorithms↗