PubMed HealthSearch

Biomedical subjects

S R Eddy

Publications and source records attributed to S R Eddy.

7 recordsLinked to original sources

Maximum discrimination hidden Markov models of sequence consensus.

We introduce a maximum discrimination method for building hidden Markov models (HMMs) of protein or nucleic acid primary sequence consensus. The method compensates for biased representation in sequence data sets, superseding the need for sequence weighting methods. Maximum discrimination HMMs are more sensitive for detecting distant sequence homologs than various other HMM methods or BLAST when tested on globin and protein kinase catalytic domain sequences.

Algorithms

Multiple alignment using hidden Markov models.

A simulated annealing method is described for training hidden Markov models and producing multiple sequence alignments from initially unaligned protein or DNA sequences. Simulated annealing in turn uses a dynamic programming algorithm for correctly sampling suboptimal multiple alignments according to their probability and a Boltzmann temperature factor. The quality of simulated annealing alignments is evaluated on structural alignments of ten different protein families, and compared to the performance of other HMM training methods and the ClustalW program. Simulated annealing is better able to find near-global optima in the multiple alignment probability landscape than the other tested HMM training methods. Neither ClustalW nor simulated annealing produce consistently better alignments compared to each other. Examination of the specific cases in which ClustalW outperforms simulated annealing, and vice versa, provides insight into the strengths and weaknesses of current hidden Markov model approaches.

Algorithms

RNA sequence analysis using covariance models.

We describe a general approach to several RNA sequence analysis problems using probabilistic models that flexibly describe the secondary structure and primary sequence consensus of an RNA sequence family. We call these models 'covariance models'. A covariance model of tRNA sequences is an extremely sensitive and discriminative tool for searching for additional tRNAs and tRNA-related sequences in sequence databases. A model can be built automatically from an existing sequence alignment. We also describe an algorithm for learning a model and hence a consensus secondary structure from initially unaligned example sequences and no prior structural information. Models trained on unaligned tRNA examples correctly predict tRNA secondary structure and produce high-quality multiple alignments. The approach may be applied to any family of small RNA sequences.

Algorithms

Artificial mobile DNA element constructed from the EcoRI endonuclease gene.

There exist several examples of mobile group I introns. These introns appear to use a straightforward mechanism to achieve highly site-specific and efficient insertion into homologous intronless genes. Because the only intron-specific function required by the prevailing model for the mechanism of intron mobility is the introduction of a site-specific double-stranded break in the intronless recipient DNA molecule, we reasoned that it should in principle be possible to construct artificially mobile DNA sequences. We have constructed an artificial mobile element from the gene for the restriction enzyme EcoRI that is capable of site-specific insertion at rates near those of authentic mobile introns. The generality of the mobility mechanism may enable high-efficiency targeted gene replacements or disruptions in a variety of organisms.

Bacteriophage lambda

The phage T4 nrdB intron: a deletion mutant of a version found in the wild.

Bacteriophage T4 possesses three self-splicing group I introns. Two of the three introns are mobile elements; the third, in the gene encoding a subunit of the phage nucleotide reductase (nrdB), is not mobile. Because intron mobility offers a reasonable explanation for the paradoxical occurrence of large intervening sequences in a space-efficient eubacterial phage, it is puzzling that the nrdB intron is not mobile like its compatriots. We have discovered a larger nrdB intron in a closely related phage, and we infer from comparative sequence data that the T4 intron is a deletion mutant derived from this larger intron. This larger nrdB intron encodes an open reading frame of 269 codons, which we have cloned and overexpressed. The overexpressed protein shows a dsDNA endonuclease activity specific for the intronless nrdB gene, typical of mobile introns. Thus, we believe that all three introns of T4 are or were mobile "infectious introns" and that they have entered into and been maintained in the phage population by virtue of this efficient mobility.

Amino Acid Sequence

Nucleotide sequence of yellow fever virus: implications for flavivirus gene expression and evolution.

The sequence of the entire RNA genome of the type flavivirus, yellow fever virus, has been obtained. Inspection of this sequence reveals a single long open reading frame of 10,233 nucleotides, which could encode a polypeptide of 3411 amino acids. The structural proteins are found within the amino-terminal 780 residues of this polyprotein; the remainder of the open reading frame consists of nonstructural viral polypeptides. This genome organization implies that mature viral proteins are produced by posttranslational cleavage of a polyprotein precursor and has implications for flavivirus RNA replication and for the evolutionary relation of this virus family to other RNA viruses.

Base Sequence