PubMed Health⌕ Search

Biomedical subjects

Ian Holmes

Publications and source records attributed to Ian Holmes.

5 recordsLinked to original sources

Accelerated probabilistic inference of RNA structure evolution.

BACKGROUND: Pairwise stochastic context-free grammars (Pair SCFGs) are powerful tools for evolutionary analysis of RNA, including simultaneous RNA sequence alignment and secondary structure prediction, but the associated algorithms are intensive in both CPU and memory usage. The same problem is faced by other RNA alignment-and-folding algorithms based on Sankoff's 1985 algorithm. It is therefore desirable to constrain such algorithms, by pre-processing the sequences and using this first pass to limit the range of structures and/or alignments that can be considered. RESULTS: We demonstrate how flexible classes of constraint can be imposed, greatly reducing the computational costs while maintaining a high quality of structural homology prediction. Any score-attributed context-free grammar (e.g. energy-based scoring schemes, or conditionally normalized Pair SCFGs) is amenable to this treatment. It is now possible to combine independent structural and alignment constraints of unprecedented general flexibility in Pair SCFG alignment algorithms. We outline several applications to the bioinformatics of RNA sequence and structure, including Waterman-Eggert N-best alignments and progressive multiple alignment. We evaluate the performance of the algorithm on test examples from the RFAM database. CONCLUSION: A program, Stemloc, that implements these algorithms for efficient RNA sequence alignment and structure prediction is available under the GNU General Public License.

Algorithms↗

Using evolutionary Expectation Maximization to estimate indel rates.

MOTIVATION: The Expectation Maximization (EM) algorithm, in the form of the Baum-Welch algorithm (for hidden Markov models) or the Inside-Outside algorithm (for stochastic context-free grammars), is a powerful way to estimate the parameters of stochastic grammars for biological sequence analysis. To use this algorithm for multiple-sequence evolutionary modelling, it would be useful to apply the EM algorithm to estimate not only the probability parameters of the stochastic grammar, but also the instantaneous mutation rates of the underlying evolutionary model (to facilitate the development of stochastic grammars based on phylogenetic trees, also known as Statistical Alignment). Recently, we showed how to do this for the point substitution component of the evolutionary process; here, we extend these results to the indel process. RESULTS: We present an algorithm for maximum-likelihood estimation of insertion and deletion rates from multiple sequence alignments, using EM, under the single-residue indel model owing to Thorne, Kishino and Felsenstein (the 'TKF91' model). The algorithm converges extremely rapidly, gives accurate results on simulated data that are an improvement over parsimonious estimates (which are shown to underestimate the true indel rate), and gives plausible results on experimental data (coronavirus envelope domains). Owing to the algorithm's close similarity to the Baum-Welch algorithm for training hidden Markov models, it can be used in an 'unsupervised' fashion to estimate rates for unaligned sequences, or estimate several sets of rates for sequences with heterogenous rates. AVAILABILITY: Software implementing the algorithm and the benchmark is available under GPL from http://www.biowiki.org/

Algorithms↗

A probabilistic model for the evolution of RNA structure.

BACKGROUND: For the purposes of finding and aligning noncoding RNA gene- and cis-regulatory elements in multiple-genome datasets, it is useful to be able to derive multi-sequence stochastic grammars (and hence multiple alignment algorithms) systematically, starting from hypotheses about the various kinds of random mutation event and their rates. RESULTS: Here, we consider a highly simplified evolutionary model for RNA, called "The TKF91 Structure Tree" (following Thorne, Kishino and Felsenstein's 1991 model of sequence evolution with indels), which we have implemented for pairwise alignment as proof of principle for such an approach. The model, its strengths and its weaknesses are discussed with reference to four examples of functional ncRNA sequences: a riboswitch (guanine), a zipcode (nanos), a splicing factor (U4) and a ribozyme (RNase P). As shown by our visualisations of posterior probability matrices, the selected examples illustrate three different signatures of natural selection that are highly characteristic of ncRNA: (i) co-ordinated basepair substitutions, (ii) co-ordinated basepair indels and (iii) whole-stem indels. CONCLUSIONS: Although all three types of mutation "event" are built into our model, events of type (i) and (ii) are found to be better modeled than events of type (iii). Nevertheless, we hypothesise from the model's performance on pairwise alignments that it would form an adequate basis for a prototype multiple alignment and genefinding tool.

Evolution, Molecular↗

Sequential partially overlapping gene arrangement in the tricistronic S1 genome segments of avian reovirus and Nelson Bay reovirus: implications for translation initiation.

Previous studies of the avian reovirus strain S1133 (ARV-S1133) S1 genome segment revealed that the open reading frame (ORF) encoding the final sigmaC viral cell attachment protein initiates over 600 nucleotides distal from the 5' end of the S1 mRNA and is preceded by two predicted small nonoverlapping ORFs. To more clearly define the translational properties of this unusual polycistronic RNA, we pursued a comparative analysis of the S1 genome segment of the related Nelson Bay reovirus (NBV). Sequence analysis indicated that the 3'-proximal ORF present on the NBV S1 genome segment also encodes a final sigmaC homolog, as evidenced by the presence of an extended N-terminal heptad repeat characteristic of the coiled-coil region common to the cell attachment proteins of reoviruses. Most importantly, the NBV S1 genome segment contains two conserved ORFs upstream of the final sigmaC coding region that are extended relative to the predicted ORFs of ARV-S1133 and are arranged in a sequential, partially overlapping fashion. Sequence analysis of the S1 genome segments of two additional strains of ARV indicated a similar overlapping tricistronic gene arrangement as predicted for the NBV S1 genome segment. Expression analysis of the ARV S1 genome segment indicated that all three ORFs are functional in vitro and in virus-infected cells. In addition to the previously described p10 and final sigmaC gene products, the S1 genome segment encodes from the central ORF a 17-kDa basic protein (p17) of no known function. Optimizing the translation start site of the ARV p10 ORF lead to an approximately 15-fold increase in p10 expression with little or no effect on translation of the downstream final sigmaC ORF. These results suggest that translation initiation complexes can bypass over 600 nucleotides and two functional overlapping upstream ORFs in order to access the distal final sigmaC start site.

Amino Acid Sequence↗