PubMed Health⌕ Search

Biomedical subjects

John S Mattick

Publications and source records attributed to John S Mattick.

At least 19 recordsLinked to original sources

RNAdb 2.0--an expanded database of mammalian non-coding RNAs.

RNAdb is a comprehensive database of mammalian non-protein-coding RNAs (ncRNAs). There is increasing recognition that ncRNAs play important regulatory roles in multicellular organisms, and there is an expanding rate of discovery of novel ncRNAs as well as an increasing allocation of function. In this update to RNAdb, we provide nucleotide sequences and annotations for tens of thousands of non-housekeeping ncRNAs, including a wide range of mammalian microRNAs, small nucleolar RNAs and larger mRNA-like ncRNAs. Some of these have documented functions and/or expression patterns, but the majority remain of unclear significance, and include PIWI-interacting RNAs, ncRNAs identified from the latest rounds of large-scale cDNA sequencing projects, putative antisense transcripts, as well as ncRNAs predicted on the basis of structural features and alignments. Improvements to the database comprise not only new and updated ncRNA datasets, but also provision of microarray-based expression data and closer interface with more specialized ncRNA resources such as miRBase and snoRNA-LBME-db. To access RNAdb, visit http://research.imb.uq.edu.au/RNAdb.

Animals↗

Effect of site-specific mutations in different phosphotransfer domains of the chemosensory protein ChpA on Pseudomonas aeruginosa motility.

The virulence of Pseudomonas aeruginosa and other surface pathogens involves the coordinate expression of a wide range of virulence determinants, including type IV pili. These surface filaments are important for the colonization of host epithelial tissues and mediate bacterial attachment to, and translocation across, surfaces by a process known as twitching motility. This process is controlled in part by a complex signal transduction system whose central component, ChpA, possesses nine potential sites of phosphorylation, including six histidine-containing phosphotransfer (HPt) domains, one serine-containing phosphotransfer domain, one threonine-containing phosphotransfer domain, and one CheY-like receiver domain. Here, using site-directed mutagenesis, we show that normal twitching motility is entirely dependent on the CheY-like receiver domain and partially dependent on two of the HPt domains. Moreover, under different assay conditions, point mutations in several of the phosphotransfer domains of ChpA give rise to unusual "swarming" phenotypes, possibly reflecting more subtle perturbations in the control of P. aeruginosa motility that are not evident from the conventional twitching stab assay. Together, these results suggest that ChpA plays a central role in the complex regulation of type IV pilus-mediated motility in P. aeruginosa.

Bacterial Proteins↗

RNA: Networks & Imaging.

The past few years have brought about a fundamental change in our understanding and definition of the RNA world and its role in the functional and regulatory architecture of the cell. The discovery of small RNAs that regulate many aspects of differentiation and development have joined the already known non-coding RNAs that are involved in chromosome dosage compensation, imprinting, and other functions to become key players in regulating the flow of genetic information. It is also evident that there are tens or even hundreds of thousands of other non-coding RNAs that are transcribed from the mammalian genome, as well as many other yet-to-be-discovered small regulatory RNAs. In the recent symposium RNA: Networks & Imaging held in Heidelberg, the dual roles of RNA as a messenger and a regulator in the flow of genetic information were discussed and new molecular genetic and imaging methods to study RNA presented.

Animals↗

Non-coding RNAs in the nervous system.

Increasing evidence suggests that the development and function of the nervous system is heavily dependent on RNA editing and the intricate spatiotemporal expression of a wide repertoire of non-coding RNAs, including micro RNAs, small nucleolar RNAs and longer non-coding RNAs. Non-coding RNAs may provide the key to understanding the multi-tiered links between neural development, nervous system function, and neurological diseases.

Animals↗

Clusters of internally primed transcripts reveal novel long noncoding RNAs.

Non-protein-coding RNAs (ncRNAs) are increasingly being recognized as having important regulatory roles. Although much recent attention has focused on tiny 22- to 25-nucleotide microRNAs, several functional ncRNAs are orders of magnitude larger in size. Examples of such macro ncRNAs include Xist and Air, which in mouse are 18 and 108 kilobases (Kb), respectively. We surveyed the 102,801 FANTOM3 mouse cDNA clones and found that Air and Xist were present not as single, full-length transcripts but as a cluster of multiple, shorter cDNAs, which were unspliced, had little coding potential, and were most likely primed from internal adenine-rich regions within longer parental transcripts. We therefore conducted a genome-wide search for regional clusters of such cDNAs to find novel macro ncRNA candidates. Sixty-six regions were identified, each of which mapped outside known protein-coding loci and which had a mean length of 92 Kb. We detected several known long ncRNAs within these regions, supporting the basic rationale of our approach. In silico analysis showed that many regions had evidence of imprinting and/or antisense transcription. These regions were significantly associated with microRNAs and transcripts from the central nervous system. We selected eight novel regions for experimental validation by northern blot and RT-PCR and found that the majority represent previously unrecognized noncoding transcripts that are at least 10 Kb in size and predominantly localized in the nucleus. Taken together, the data not only identify multiple new ncRNAs but also suggest the existence of many more macro ncRNAs like Xist and Air.

Animals↗

Non-coding RNA.

The term non-coding RNA (ncRNA) is commonly employed for RNA that does not encode a protein, but this does not mean that such RNAs do not contain information nor have function. Although it has been generally assumed that most genetic information is transacted by proteins, recent evidence suggests that the majority of the genomes of mammals and other complex organisms is in fact transcribed into ncRNAs, many of which are alternatively spliced and/or processed into smaller products. These ncRNAs include microRNAs and snoRNAs (many if not most of which remain to be identified), as well as likely other classes of yet-to-be-discovered small regulatory RNAs, and tens of thousands of longer transcripts (including complex patterns of interlacing and overlapping sense and antisense transcripts), most of whose functions are unknown. These RNAs (including those derived from introns) appear to comprise a hidden layer of internal signals that control various levels of gene expression in physiology and development, including chromatin architecture/epigenetic memory, transcription, RNA splicing, editing, translation and turnover. RNA regulatory networks may determine most of our complex characteristics, play a significant role in disease and constitute an unexplored world of genetic variation both within and between species.

Animals↗

Discrimination of non-protein-coding transcripts from protein-coding mRNA.

Several recent studies indicate that mammals and other organisms produce large numbers of RNA transcripts that do not correspond to known genes. It has been suggested that these transcripts do not encode proteins, but may instead function as RNAs. However, discrimination of coding and non-coding transcripts is not straightforward, and different laboratories have used different methods, whose ability to perform this discrimination is unclear. In this study, we examine ten bioinformatic methods that assess protein-coding potential and compare their ability and congruency in the discrimination of non-coding from coding sequences, based on four underlying principles: open reading frame size, sequence similarity to known proteins or protein domains, statistical models of protein-coding sequence, and synonymous versus non-synonymous substitution rates. Despite these different approaches, the methods show broad concordance, suggesting that coding and non-coding transcripts can, in general, be reliably discriminated, and that many of the recently discovered extra-genic transcripts are indeed non-coding. Comparison of the methods indicates reasons for unreliable predictions, and approaches to increase confidence further. Conversely and surprisingly, our analyses also provide evidence that as much as approximately 10% of entries in the manually curated protein database Swiss-Prot are erroneous translations of actually non-coding transcripts.

Algorithms↗

Evidence for control of splicing by alternative RNA secondary structures in Dipteran homothorax pre-mRNA.

In a recent study that identified highly evolutionary conserved sequences in three genomes of Diptera species we described an ultraconserved element found at an internal exon-intron junction of the Drosophila melanogaster homothorax (hth) gene that appeared to be involved in the control of hth pre-mRNA splicing. We also discussed a possible role of RNA secondary structure at this site in the regulation of hth pre-mRNA splicing. In this report we identify a shorter evolutionary conserved intronic element within the hth gene that is located downstream of the first element and has sequence complementarity to it. We demonstrate that intramolecular interactions between these two elements would give rise to alternative RNA secondary structures, which in turn may result in differential control of homothorax pre-mRNA splicing. We also provide additional comparative genomic data from several newly available insect genomes supporting our original conclusion that these conserved elements are important in the post-transcriptional regulation of homothorax gene expression in Diptera.

Alternative Splicing↗

GONOME: measuring correlations between GO terms and genomic positions.

BACKGROUND: Current methods to find significantly under- and over-represented gene ontology (GO) terms in a set of genes consider the genes as equally probable "balls in a bag", as may be appropriate for transcripts in micro-array data. However, due to the varying length of genes and intergenic regions, that approach is inappropriate for deciding if any GO terms are correlated with a set of genomic positions. RESULTS: We present an algorithm--GONOME--that can determine which GO terms are significantly associated with a set of genomic positions given a genome annotated with (at least) the starts and ends of genes. We show that certain GO terms may appear to be significantly associated with a set of randomly chosen positions in the human genome if gene lengths are not considered, and that these same terms have been reported as significantly over-represented in a number of recent papers. This apparent over-representation disappears when gene lengths are considered, as GONOME does. For example, we show that, when gene length is taken into account, the term "development" is not significantly enriched in genes associated with human CpG islands, in contradiction to a previous report. We further demonstrate the efficacy of GONOME by showing that occurrences of the proteosome-associated control element (PACE) upstream activating sequence in the S. cerevisiae genome associate significantly to appropriate GO terms. An extension of this approach yields a whole-genome motif discovery algorithm that allows identification of many other promoter sequences linked to different types of genes, including a large group of previously unknown motifs significantly associated with the terms 'translation' and 'translational elongation'. CONCLUSION: GONOME is an algorithm that correctly extracts over-represented GO terms from a set of genomic positions. By explicitly considering gene size, GONOME avoids a systematic bias toward GO terms linked to large genes. Inappropriate use of existing algorithms that do not take gene size into account has led to erroneous or suspect conclusions. Reciprocally GONOME may be used to identify new features in genomes that are significantly associated with particular categories of genes.

Algorithms↗

Transposon-free regions in mammalian genomes.

Despite the presence of over 3 million transposons separated on average by approximately 500 bp, the human and mouse genomes each contain almost 1000 transposon-free regions (TFRs) over 10 kb in length. The majority of human TFRs correlate with orthologous TFRs in the mouse, despite the fact that most transposons are lineage specific. Many human TFRs also overlap with orthologous TFRs in the marsupial opossum, indicating that these regions have remained refractory to transposon insertion for long evolutionary periods. Over 90% of the bases covered by TFRs are noncoding, much of which is not highly conserved. Most TFRs are not associated with unusual nucleotide composition, but are significantly associated with genes encoding developmental regulators, suggesting that they represent extended regions of regulatory information that are largely unable to tolerate insertions, a conclusion difficult to reconcile with current conceptions of gene regulation.

Animals↗

Experimental validation of the regulated expression of large numbers of non-coding RNAs from the mouse genome.

Recent large-scale analyses of mainly full-length cDNA libraries generated from a variety of mouse tissues indicated that almost half of all representative cloned sequences did not contain an apparent protein-coding sequence, and were putatively derived from non-protein-coding RNA (ncRNA) genes. However, many of these clones were singletons and the majority were unspliced, raising the possibility that they may be derived from genomic DNA or unprocessed pre-mRNA contamination during library construction, or alternatively represent nonspecific "transcriptional noise." Here we show, using reverse transcriptase-dependent PCR, microarray, and Northern blot analyses, that many of these clones were derived from genuine transcripts of unknown function whose expression appears to be regulated. The ncRNA transcripts have larger exons and fewer introns than protein-coding transcripts. Analysis of the genomic landscape around these sequences indicates that some cDNA clones were produced not from terminal poly(A) tracts but internal priming sites within longer transcripts, only a minority of which is encompassed by known genes. A significant proportion of these transcripts exhibit tissue-specific expression patterns, as well as dynamic changes in their expression in macrophages following lipopolysaccharide stimulation. Taken together, the data provide strong support for the conclusion that ncRNAs are an important, regulated component of the mammalian transcriptome.

Animals↗

Rapid evolution of noncoding RNAs: lack of conservation does not mean lack of function.

The mammalian transcriptome contains many non-protein-coding RNAs (ncRNAs), but most of these are of unclear significance and lack strong sequence conservation, prompting suggestions that they might be non-functional. However, certain long functional ncRNAs such as Air and Xist are also poorly conserved. In this article, we systematically analyzed the conservation of several groups of functional ncRNAs, including miRNAs, snoRNAs and longer ncRNAs whose function has been either documented or confidently predicted. As expected, miRNAs and snoRNAs were highly conserved. By contrast, the longer functional non-micro, non-sno ncRNAs were much less conserved with many displaying rapid sequence evolution. Our findings suggest that longer ncRNAs are under the influence of different evolutionary constraints and that the lack of conservation displayed by the thousands of candidate ncRNAs does not necessarily signify an absence of function.

Animals↗

The functional genomics of noncoding RNA.

Large numbers of noncoding RNA transcripts (ncRNAs) are being revealed by complementary DNA cloning and genome tiling array studies in animals. The big and as yet largely unanswered question is whether these transcripts are relevant. A paper by Willingham et al. shows the way forward by developing a strategy for large-scale functional screening of ncRNAs, involving small interfering RNA knockdowns in cell-based screens, which identified a previously unidentified ncRNA repressor of the transcription factor NFAT. It appears likely that ncRNAs constitute a critical hidden layer of gene regulation in complex organisms, the understanding of which requires new approaches in functional genomics.

Animals↗

Ultraconserved elements in insect genomes: a highly conserved intronic sequence implicated in the control of homothorax mRNA splicing.

Recently, we identified a large number of ultraconserved (uc) sequences in noncoding regions of human, mouse, and rat genomes that appear to be essential for vertebrate and amniote ontogeny. Here, we used similar methods to identify ultraconserved genomic regions between the insect species Drosophila melanogaster and Drosophila pseudoobscura, as well as the more distantly related Anopheles gambiae. As with vertebrates, ultraconserved sequences in insects appear to occur primarily in intergenic and intronic sequences, and at intron-exon junctions. The sequences are significantly associated with genes encoding developmental regulators and transcription factors, but are less frequent and are smaller in size than in vertebrates. The longest identical, nongapped orthologous match between the three genomes was found within the homothorax (hth) gene. This sequence spans an internal exon-intron junction, with the majority located within the intron, and is predicted to form a highly stable stem-loop RNA structure. Real-time quantitative PCR analysis of different hth splice isoforms and Northern blotting showed that the conserved element is associated with a high incidence of intron retention in hth pre-mRNA, suggesting that the conserved intronic element is critically important in the post-transcriptional regulation of hth expression in Diptera.

Animals↗

Small regulatory RNAs in mammals.

Mammalian cells harbor numerous small non-protein-coding RNAs, including small nucleolar RNAs (snoRNAs), microRNAs (miRNAs), short interfering RNAs (siRNAs) and small double-stranded RNAs, which regulate gene expression at many levels including chromatin architecture, RNA editing, RNA stability, translation, and quite possibly transcription and splicing. These RNAs are processed by multistep pathways from the introns and exons of longer primary transcripts, including protein-coding transcripts. Most show distinctive temporal- and tissue-specific expression patterns in different tissues, including embryonal stem cells and the brain, and some are imprinted. Small RNAs control a wide range of developmental and physiological pathways in animals, including hematopoietic differentiation, adipocyte differentiation and insulin secretion in mammals, and have been shown to be perturbed in cancer and other diseases. The extent of transcription of non-coding sequences and the abundance of small RNAs suggests the existence of an extensive regulatory network on the basis of RNA signaling which may underpin the development and much of the phenotypic variation in mammals and other complex organisms and which may have different genetic signatures from sequences encoding proteins.

Animals↗

RNAdb--a comprehensive mammalian noncoding RNA database.

In recent years, there have been increasing numbers of transcripts identified that do not encode proteins, many of which are developmentally regulated and appear to have regulatory functions. Here, we describe the construction of a comprehensive mammalian noncoding RNA database (RNAdb) which contains over 800 unique experimentally studied non-coding RNAs (ncRNAs), including many associated with diseases and/or developmental processes. The database is available at http://research.imb.uq.edu.au/RNAdb and is searchable by many criteria. It includes microRNAs and snoRNAs, but not infrastructural RNAs, such as rRNAs and tRNAs, which are catalogued elsewhere. The database also includes over 1100 putative antisense ncRNAs and almost 20,000 putative ncRNAs identified in high-quality murine and human cDNA libraries, with more to be added in the near future. Many of these RNAs are large, and many are spliced, some alternatively. The database will be useful as a foundation for the emerging field of RNomics and the characterization of the roles of ncRNAs in mammalian gene expression and regulation.

Animals↗