PubMed Health⌕ Search

Biomedical subjects

Ivo L Hofacker

Publications and source records attributed to Ivo L Hofacker.

At least 19 recordsLinked to original sources

RNAs everywhere: genome-wide annotation of structured RNAs.

Starting with the discovery of microRNAs and the advent of genome-wide transcriptomics, non-protein-coding transcripts have moved from a fringe topic to a central field research in molecular biology. In this contribution we review the state of the art of "computational RNomics", i.e., the bioinformatics approaches to genome-wide RNA annotation. Instead of rehashing results from recently published surveys in detail, we focus here on the open problem in the field, namely (functional) annotation of the plethora of putative RNAs. A series of exploratory studies are used to provide non-trivial examples for the discussion of some of the difficulties.

Biological Evolution↗

Partition function and base pairing probabilities of RNA heterodimers.

BACKGROUND: RNA has been recognized as a key player in cellular regulation in recent years. In many cases, non-coding RNAs exert their function by binding to other nucleic acids, as in the case of microRNAs and snoRNAs. The specificity of these interactions derives from the stability of inter-molecular base pairing. The accurate computational treatment of RNA-RNA binding therefore lies at the heart of target prediction algorithms. METHODS: The standard dynamic programming algorithms for computing secondary structures of linear single-stranded RNA molecules are extended to the co-folding of two interacting RNAs. RESULTS: We present a program, RNAcofold, that computes the hybridization energy and base pairing pattern of a pair of interacting RNA molecules. In contrast to earlier approaches, complex internal structures in both RNAs are fully taken into account. RNAcofold supports the calculation of the minimum energy structure and of a complete set of suboptimal structures in an energy band above the ground state. Furthermore, it provides an extension of McCaskill's partition function algorithm to compute base pairing probabilities, realistic interaction energies, and equilibrium concentrations of duplex structures.

Journal Article↗

The expansion of the metazoan microRNA repertoire.

BACKGROUND: MicroRNAs have been identified as crucial regulators in both animals and plants. Here we report on a comprehensive comparative study of all known miRNA families in animals. We expand the MicroRNA Registry 6.0 by more than 1000 new homologs of miRNA precursors whose expression has been verified in at least one species. Using this uniform data basis we analyze their evolutionary history in terms of individual gene phylogenies and in terms of preservation of genomic nearness across species. This allows us to reliably identify microRNA clusters that are derived from a common transcript. RESULTS: We identify three episodes of microRNA innovation that correspond to major developmental innovations: A class of about 20 miRNAs is common to protostomes and deuterostomes and might be related to the advent of bilaterians. A second large wave of innovations maps to the branch leading to the vertebrates. The third significant outburst of miRNA innovation coincides with placental (eutherian) mammals. In addition, we observe the expected expansion of the microRNA inventory due to genome duplications in early vertebrates and in an ancestral teleost. The non-local duplications in the vertebrate ancestor are predated by local (tandem) duplications leading to the formation of about a dozen ancient microRNA clusters. CONCLUSION: Our results suggest that microRNA innovation is an ongoing process. Major expansions of the metazoan miRNA repertoire coincide with the advent of bilaterians, vertebrates, and (placental) mammals.

Animals↗

Algebraic comparison of metabolic networks, phylogenetic inference, and metabolic innovation.

BACKGROUND: Comparison of metabolic networks is typically performed based on the organisms' enzyme contents. This approach disregards functional replacements as well as orthologies that are misannotated. Direct comparison of the structure of metabolic networks can circumvent these problems. RESULTS: Metabolic networks are naturally represented as directed hypergraphs in such a way that metabolites are nodes and enzyme-catalyzed reactions form (hyper)edges. The familiar operations from set algebra (union, intersection, and difference) form a natural basis for both the pairwise comparison of networks and identification of distinct metabolic features of a set of algorithms. We report here on an implementation of this approach and its application to the procaryotes. CONCLUSION: We demonstrate that metabolic networks contain valuable phylogenetic information by comparing phylogenies obtained from network comparisons with 16S RNA phylogenies. The algebraic approach to metabolic networks is suitable to study metabolic innovations in two sets of organisms, free living microbes and Pyrococci, as well as obligate intracellular pathogens.

Algorithms↗

Memory efficient folding algorithms for circular RNA secondary structures.

BACKGROUND: A small class of RNA molecules, in particular the tiny genomes of viroids, are circular. Yet most structure prediction algorithms handle only linear RNAs. The most straightforward approach is to compute circular structures from 'internal' and 'external' substructures separated by a base pair. This is incompatible, however, with the memory-saving approach of the Vienna RNA Package which builds a linear RNA structure from shorter (internal) structures only. RESULT: Here we describe how circular secondary structures can be obtained without additional memory requirements as a kind of 'post-processing' of the linear structures. AVAILABILITY: The circular folding algorithm is implemented in the current version of the of RNAfold program of the Vienna RNA Package, which can be downloaded from http://www.tbi.univie.ac.at/RNA/

Algorithms↗

Thermodynamics of RNA-RNA binding.

BACKGROUND: Reliable prediction of RNA-RNA binding energies is crucial, e.g. for the understanding on RNAi, microRNA-mRNA binding and antisense interactions. The thermodynamics of such RNA-RNA interactions can be understood as the sum of two energy contributions: (1) the energy necessary to 'open' the binding site and (2) the energy gained from hybridization. METHODS: We present an extension of the standard partition function approach to RNA secondary structures that computes the probabilities Pu[i, j] that a sequence interval [i, j] is unpaired. RESULTS: Comparison with experimental data shows that Pu[i, j] can be applied as a significant determinant of local target site accessibility for RNA interference (RNAi). Furthermore, these quantities can be used to rigorously determine binding free energies of short oligomers to large mRNA targets. The resource consumption is comparable with a single partition function computation for the large target molecule. We can show that RNAi efficiency correlates well with the binding energies of siRNAs to their respective mRNA target. AVAILABILITY: RNAup will be distributed as part of the Vienna RNA Package, www.tbi.univie.ac.at/~ivo/RNA/

Algorithms↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗

Local RNA base pairing probabilities in large sequences.

SUMMARY: The genome-wide search for non-coding RNAs requires efficient methods to compute and compare local secondary structures. Since the exact boundaries of such putative transcripts are typically unknown, arbitrary sequence windows have to be used in practice. Here we present a method for robustly computing the probabilities of local base pairs from long RNA sequences independent of the exact positions of the sequence window. AVAILABILITY: The program RNAplfold is part of the Vienna RNA Package and can be downloaded from http://www.tbi.univie.ac.at/RNA/.

Algorithms↗

Multiple sequence alignments of partially coding nucleic acid sequences.

BACKGROUND: High quality sequence alignments of RNA and DNA sequences are an important prerequisite for the comparative analysis of genomic sequence data. Nucleic acid sequences, however, exhibit a much larger sequence heterogeneity compared to their encoded protein sequences due to the redundancy of the genetic code. It is desirable, therefore, to make use of the amino acid sequence when aligning coding nucleic acid sequences. In many cases, however, only a part of the sequence of interest is translated. On the other hand, overlapping reading frames may encode multiple alternative proteins, possibly with intermittent non-coding parts. Examples are, in particular, RNA virus genomes. RESULTS: The standard scoring scheme for nucleic acid alignments can be extended to incorporate simultaneously information on translation products in one or more reading frames. Here we present a multiple alignment tool, codaln, that implements a combined nucleic acid plus amino acid scoring model for pairwise and progressive multiple alignments that allows arbitrary weighting for almost all scoring parameters. Resource requirements of codaln are comparable with those of standard tools such as ClustalW. CONCLUSION: We demonstrate the applicability of codaln to various biologically relevant types of sequences (bacteriophage Levivirus and Vertebrate Hox clusters) and show that the combination of nucleic acid and amino acid sequence information leads to improved alignments. These, in turn, increase the performance of analysis tools that depend strictly on good input alignments such as methods for detecting conserved RNA secondary structure elements.

Algorithms↗

Fast and reliable prediction of noncoding RNAs.

We report an efficient method for detecting functional RNAs. The approach, which combines comparative sequence analysis and structure prediction, already has yielded excellent results for a small number of aligned sequences and is suitable for large-scale genomic screens. It consists of two basic components: (i) a measure for RNA secondary structure conservation based on computing a consensus secondary structure, and (ii) a measure for thermodynamic stability, which, in the spirit of a z score, is normalized with respect to both sequence length and base composition but can be calculated without sampling from shuffled sequences. Functional RNA secondary structures can be identified in multiple sequence alignments with high sensitivity and high specificity. We demonstrate that this approach is not only much more accurate than previous methods but also significantly faster. The method is implemented in the program rnaz, which can be downloaded from www.tbi.univie.ac.at/~wash/RNAz. We screened all alignments of length n > or = 50 in the Comparative Regulatory Genomics database, which compiles conserved noncoding elements in upstream regions of orthologous genes from human, mouse, rat, Fugu, and zebrafish. We recovered all of the known noncoding RNAs and cis-acting elements with high significance and found compelling evidence for many other conserved RNA secondary structures not described so far to our knowledge.

Algorithms↗

Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.

In contrast to the fairly reliable and complete annotation of the protein coding genes in the human genome, comparable information is lacking for noncoding RNAs (ncRNAs). We present a comparative screen of vertebrate genomes for structural noncoding RNAs, which evaluates conserved genomic DNA sequences for signatures of structural conservation of base-pairing patterns and exceptional thermodynamic stability. We predict more than 30,000 structured RNA elements in the human genome, almost 1,000 of which are conserved across all vertebrates. Roughly a third are found in introns of known genes, a sixth are potential regulatory elements in untranslated regions of protein-coding mRNAs and about half are located far away from any known gene. Only a small fraction of these sequences has been described previously. A comparison with recent tiling array data shows that more than 40% of the predicted structured RNAs overlap with experimentally detected sites of transcription. The widespread conservation of secondary structure points to a large number of functional ncRNAs and cis-acting mRNA structures in the human genome.

Animals↗

Unorthodox mRNA start site to extend the highly structured leader of retrotransposon Tto1 mRNA increases transposition rate.

Retroelement RNAs serve as templates for both translation and reverse transcription into extrachromosomal DNA. DNA copies may be inserted into the host genome to multiply element sequences. This transpositional activity of retroelements is usually restricted to specific conditions, particularly to conditions that impose stress on the host organism. In this work, we examined how the mRNA initiation point, and features of primary and secondary structure, of tobacco retrotransposon Tto1 RNA influence its transpositional activity. We found that the most abundant Tto1 RNA is not a substrate for reverse transcription. It is poorly translated, and its 5'-end does not contain a region of redundancy with the most prominent 3'-end. In contrast, expression of an mRNA with the 5'-end extended by 28 nucleotides allows translation and gives rise to transposition events in the heterologous host, Arabidopsis thaliana. In addition, the presence of extended hairpins and of two short open reading frames in the 5'-leader sequence of Tto1 mRNA suggests that translation does not involve ribosome scanning from the mRNA 5'-end to the translation initiation site.

3' Untranslated Regions↗

Consensus folding of aligned sequences as a new measure for the detection of functional RNAs by comparative genomics.

Facing the ever-growing list of newly discovered classes of functional RNAs, it can be expected that further types of functional RNAs are still hidden in recently completed genomes. The computational identification of such RNA genes is, therefore, of major importance. While most known functional RNAs have characteristic secondary structures, their free energies are generally not statistically significant enough to distinguish RNA genes from the genomic background. Additional information is required. Considering the wide availability of new genomic data of closely related species, comparative studies seem to be the most promising approach. Here, we show that prediction of consensus structures of aligned sequences can be a significant measure to detect functional RNAs. We report a new method to test multiple sequence alignments for the existence of an unusually structured and conserved fold. We show for alignments of six types of well-known functional RNA that an energy score consisting of free energy and a covariation term significantly improves sensitivity compared to single sequence predictions. We further test our method on a number of non-coding RNAs from Caenorhabditis elegans/Caenorhabditis briggsae and seven Saccharomyces species. Most RNAs can be detected with high significance. We provide a Perl implementation that can be used readily to score single alignments and discuss how the methods described here can be extended to allow for efficient genome-wide screens.

Algorithms↗

Alignment of RNA base pairing probability matrices.

MOTIVATION: Many classes of functional RNA molecules are characterized by highly conserved secondary structures but little detectable sequence similarity. Reliable multiple alignments can therefore be constructed only when the shared structural features are taken into account. Since multiple alignments are used as input for many subsequent methods of data analysis, structure-based alignments are an indispensable necessity in RNA bioinformatics. RESULTS: We present here a method to compute pairwise and progressive multiple alignments from the direct comparison of base pairing probability matrices. Instead of attempting to solve the folding and the alignment problem simultaneously as in the classical Sankoff's algorithm, we use McCaskill's approach to compute base pairing probability matrices which effectively incorporate the information on the energetics of each sequences. A novel, simplified variant of Sankoff's algorithms can then be employed to extract the maximum-weight common secondary structure and an associated alignment. AVAILABILITY: The programs pmcomp and pmmulti described in this contribution are implemented in Perl and can be downloaded together with the example datasets from http://www.tbi.univie.ac.at/RNA/PMcomp/. A web server is available at http://rna.tbi.univie.ac.at/cgi-bin/pmcgi.pl

Algorithms↗

Structural and evolutionary analysis of the transcribed sequence of Boudicca, a Schistosoma mansoni retrotransposon.

Boudicca is a gypsy-like, long terminal repeat (LTR) retrotransposon that has colonized the genome of the human blood fluke, Schistosoma mansoni. Previous studies have indicated that more than 1000 copies of Boudicca reside within the S. mansoni genome, although many of them may be degenerate and inactive. Messenger RNAs transcribed from genomic copies of Boudicca were investigated by reverse transcription PCR. Overlapping RT-PCR products corresponding to the gag and pol polyproteins of Boudicca, along with relevant sequences of genomic fragments of Boudicca, were assembled into contigs. Consensus sequences from these contigs were used to predict the sequence and structure of transpositionally active copies of the Boudicca retrotransposon. They verified that Boudicca has a kabuki-like Cys-His box motif at the active site of its gag protein, a classic DTG motif as the active site of the protease domain of the pol ORF2, and indicated a contiguous integrase domain at the C-terminus of pol with strong identity to integrase from the LTR retrotransposons CsRn1 and kabuki, as well as to the conserved integrase core domain, GenBank rve (). Models of the secondary structure of the Boudicca transcript suggested that the first AUG was occluded by a stem loop structure, which in turn suggested a method of regulation of expression, at the level of translation, of Boudicca proteins. In addition, phylogenetic analysis targeting discrete domains of Boudicca revealed a generalized radiation in sequences among the multiple copies of Boudicca resident in the schistosome genome.

Amino Acid Sequence↗

Conserved RNA secondary structures in Flaviviridae genomes.

Presented here is a comprehensive computational survey of evolutionarily conserved secondary structure motifs in the genomic RNAs of the family Flaviviridae: This virus family consists of the three genera Flavivirus, Pestivirus and Hepacivirus and the group of GB virus C/hepatitis G virus with a currently uncertain taxonomic classification. Based on the control of replication and translation, two subgroups were considered separately: the genus Flavivirus, with its type I cap structure at the 5' untranslated region (UTR) and a highly structured 3' UTR, and the remaining three groups, which exhibit translation control by means of an internal ribosomal entry site (IRES) in the 5' UTR and a much shorter less-structured 3' UTR. The main findings of this survey are strong hints for the possibility of genome cyclization in hepatitis C virus and GB virus C/hepatitis G virus in addition to the flaviviruses; a surprisingly large number of conserved RNA motifs in the coding regions; and a lower level of detailed structural conservation in the IRES and 3' UTR motifs than reported in the literature. An electronic atlas organizes the information on the more than 150 conserved, and therefore putatively functional, RNA secondary structure elements.

3' Untranslated Regions↗

Vienna RNA secondary structure server.

The Vienna RNA secondary structure server provides a web interface to the most frequently used functions of the Vienna RNA software package for the analysis of RNA secondary structures. It currently offers prediction of secondary structure from a single sequence, prediction of the consensus secondary structure for a set of aligned sequences and the design of sequences that will fold into a predefined structure. All three services can be accessed via the Vienna RNA web server at http://rna.tbi.univie.ac.at/.

Algorithms↗

On the evolution of primitive genetic codes.

The primordial genetic code probably has been a drastically simplified ancestor of the canonical code that is used by contemporary cells. In order to understand how the present-day code came about we first need to explain how the language of the building plan can change without destroying the encoded information. In this work we introduce a minimal organism model that is based on biophysically reasonable descriptions of RNA and protein, namely secondary structure folding and knowledge based potentials. The evolution of a population of such organism under competition for a common resource is simulated explicitly at the level of individual replication events. Starting with very simple codes, and hence greatly reduced amino acid alphabets, we observe a diversification of the codes in most simulation runs. The driving force behind this effect is the possibility to produce fitter proteins when the repertoire of amino acids is enlarged.

Amino Acid Sequence↗