PubMed Health⌕ Search

Biomedical subjects

L Grate

Publications and source records attributed to L Grate.

8 recordsLinked to original sources

Test of intron predictions reveals novel splice sites, alternatively spliced mRNAs and new introns in meiotically regulated genes of yeast.

Correct identification of all introns is necessary to discern the protein-coding potential of a eukaryotic genome. The existence of most of the spliceosomal introns predicted in the genome of Saccharomyces cerevisiae remains unsupported by molecular evidence. We tested the intron predictions for 87 introns predicted to be present in non-ribosomal protein genes, more than a third of all known or suspected introns in the yeast genome. Evidence supporting 61 of these predictions was obtained, 20 predicted intron sequences were not spliced and six predictions identified an intron-containing region but failed to specify the correct splice sites, yielding a successful prediction rate of <80%. Alternative splicing has not been previously described for this organism, and we identified two genes (YKL186C/ MTR2 and YML034W) which encode alternatively spliced mRNAs; YKL186C/ MTR2 produces at least five different spliced mRNAs. One gene (YGR225W/ SPO70 ) has an intron whose removal is activated during meiosis under control of the MER1 gene. We found eight new introns, suggesting that numerous introns still remain to be discovered. The results show that correct prediction of introns remains a significant barrier to understanding the structure, function and coding capacity of eukaryotic genomes, even in a supposedly simple system like yeast.

Alternative Splicing↗

A homolog of mammalian antizyme is present in fission yeast Schizosaccharomyces pombe but not detected in budding yeast Saccharomyces cerevisiae.

MOTIVATION: The antizymes (AZ) are proteins that regulate cellular polyamine pools in metazoa. To search for remote homologs in single-celled eukaryotes, we used computer software based on hidden Markov models. The most divergent homolog detected was that of the fission yeast Schizosaccharomyces pombe. Sequence identities between S.POMBE: AZ and known AZs are as low as 18-22% in the most conserved C-terminal regions. The authenticity of the S.POMBE: AZ is validated by the presence of a conserved nucleotide sequence that, in metazoa, promotes a +1 programmed ribosomal frameshift required for AZ expression. However, no homolog was detected in the completed genome of the budding yeast Saccharomyces cerevisiae. Procedural details and supplementary information can be found at http://itsa.ucsf.edu/ approximately czhu/AZ.

Amino Acid Sequence↗

Predicting protein structure using only sequence information.

This paper presents results of blind predictions submitted to the CASP3 protein structure prediction experiment. We made predictions using the SAM-T98 method, an iterative hidden Markov model-based method for constructing protein family profiles. The method is purely sequence-based, using no structural information, and yet was able to predict structures as well as all but five of the structure-based methods in CASP3.

Algorithms↗

Genome-wide bioinformatic and molecular analysis of introns in Saccharomyces cerevisiae.

Introns have typically been discovered in an ad hoc fashion: introns are found as a gene is characterized for other reasons. As complete eukaryotic genome sequences become available, better methods for predicting RNA processing signals in raw sequence will be necessary in order to discover genes and predict their expression. Here we present a catalog of 228 yeast introns, arrived at through a combination of bioinformatic and molecular analysis. Introns annotated in the Saccharomyces Genome Database (SGD) were evaluated, questionable introns were removed after failing a test for splicing in vivo, and known introns absent from the SGD annotation were added. A novel branchpoint sequence, AAUUAAC, was identified within an annotated intron that lacks a six-of-seven match to the highly conserved branchpoint consensus UACUAAC. Analysis of the database corroborates many conclusions about pre-mRNA substrate requirements for splicing derived from experimental studies, but indicates that splicing in yeast may not be as rigidly determined by splice-site conservation as had previously been thought. Using this database and a molecular technique that directly displays the lariat intron products of spliced transcripts (intron display), we suggest that the current set of 228 introns is still not complete, and that additional intron-containing genes remain to be discovered in yeast. The database can be accessed at http://www.cse.ucsc.edu/research/compbi o/yeast_introns.html.

Computational Biology↗

Potential SECIS elements in HIV-1 strain HXB2.

It has been proposed on the basis of sequence analysis that HIV-1 encodes a protein containing the amino acid selenocysteine (Sec). Selenocysteine is known to be incorporated into protein in response to a specific RNA secondary structure motif within the mRNA that is being translated. This RNA motif, the selenocysteine insertion sequence (SECIS) element, has not yet been identified in the HIV genome by either biologic or computation methods. This report uses computer-based sequence analysis to identify those locations in HIV-1 strain HXB2 where the current model of the SECIS element could exist. One particularly good match to the SECIS element occurs in an interesting location, spanning the end of env and the start of nef, in a position theoretically capable of directing the previously proposed Sec incorporation.

Algorithms↗

Automatic RNA secondary structure determination with stochastic context-free grammars.

We have developed a method for predicting the common secondary structure of large RNA multiple alignments using only the information in the alignment. It uses a series of progressively more sensitive searches of the data in an iterative manner to discover regions of base pairing; the first pass examines the entire multiple alignment. The searching uses two methods to find base pairings. Mutual information is used to measure covariation between pairs of columns in the multiple alignment and a minimum length encoding method is used to detect column pairs with high potential to base pair. Dynamic programming is used to recover the optimal tree made up of the best potential base pairs and to create a stochastic context-free grammar. The information in the tree guides the next iteration of searching. The method is similar to the traditional comparative sequence analysis technique. The method correctly identifies most of the common secondary structure in 16S and 23S rRNA.

Algorithms↗

RNA modeling using Gibbs sampling and stochastic context free grammars.

A new method of discovering the common secondary structure of a family of homologous RNA sequences using Gibbs sampling and stochastic context-free grammars is proposed. Given an unaligned set of sequences, a Gibbs sampling step simultaneously estimates the secondary structure of each sequence and a set of statistical parameters describing the common secondary structure of the set as a whole. These parameters describe a statistical model of the family. After the Gibbs sampling has produced a crude statistical model for the family, this model is translated into a stochastic context-free grammar, which is then refined by an Expectation Maximization (EM) procedure to produce a more complete model. A prototype implementation of the method is tested on tRNA, pieces of 16S rRNA and on U5 snRNA with good results.

Animals↗