PubMed HealthSearch

PubMed · 9485505

Predicting protein structure using hidden Markov models.

Abstract

We discuss how methods based on hidden Markov models performed in the fold-recognition section of the CASP2 experiment. Hidden Markov models were built for a representative set of just over 1,000 structures from the Protein Data Bank (PDB). Each CASP2 target sequence was scored against this library of HMMs. In addition, an HMM was built for each of the target sequences and all of the sequences in PDB were scored against that target model, with a good score on both methods indicating a high probability that the target sequence is homologous to the structure. The method worked well in comparison to other methods used at CASP2 for targets of moderate difficulty, where the closest structure in PDB could be aligned to the target with at least 15% residue identity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

K Karplus, K Sjölander, C Barrett, M Cline, D Haussler, R Hughey, L Holm, C Sander. 1997. Predicting protein structure using hidden Markov models.. https://doi.org/10.1002/(sici)1097-0134(1997)1%2B%3C134%3A%3Aaid-prot18%3E3.3.co%3B2-q

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A free energy analysis by unfolding applied to 125-mers on a cubic lattice.

BACKGROUND: A common approach to the protein folding problem involves computer simulation of folding using lattice models of amino acid sequences. Key factors for good performance in such models are the correct choice of the temperature and the average interaction energy between residues. In order to push the lattice approach to its limit it is important to have a method to adjust these parameters for optimal folding that is not limited by our ability to successfully simulate folding in a reasonable time. RESULTS: In this study, we adopt a simple cubic-lattice model and present a method for calculating the free energy of a chain as a function of the number of native contacts. This does not require that we are able to fold the sequence by simulation and it provides a method of estimating the folding transition temperature. For a given set of parameters, the free energy analysis also allows an estimate of foldability. By applying the method to sequences with 27 and 125 residues, we show that optimal folding occurs near the folding transition temperature and at either zero or small negative average interaction energy. We find ourselves able to fold only 125-mers that have significant short-range native contacts. CONCLUSIONS: A free energy analysis during unfolding is a useful tool for the study of foldability and should be applicable to a variety of folding models. In this way we are able to fold some 125-mer designed sequences and our results confirm the finding that short-range contacts contribute to foldability.

Markov Chains

A covariotide model explains apparent phylogenetic structure of oxygenic photosynthetic lineages.

The aims of the work were (1) to develop statistical tests to identify whether substitution takes place under a covariotide model in sequences used for phylogenetic inference and (2) to determine the influence of covariotide substitution on phylogenetic trees inferred for photosynthetic and other organisms. (Covariotide and covarion models are ones in which sites that are variable in some parts of the underlying tree are invariable in others and vice versa.) Two tests were developed. The first was a contingency test, and the second was an inequality test comparing the expected number of variable sites in two groups with the observed number. Application of these tests to 16S rDNA and tufA sequences from a range of nonphotosynthetic prokaryotes and oxygenic photosynthetic prokaryotes and eukaryotes suggests the occurrence of a covariotide mechanism. The degree of support for partitioning of taxa in reconstructed trees involving these organisms was determined in the presence or absence of sites showing particular substitution patterns. This analysis showed that the support for splits between (1) photosynthetic eukaryotes and prokaryotes and (2) photosynthetic and nonphotosynthetic organisms could be accounted for by patterns arising from covariotide substitution. We show that the additional problem of compositional bias in sequence data needs to be considered in the context of patterns of covariotide/covarion substitution. We argue that while covariotide or covarion substitution may give rise to phylogenetically informative patterns in sequence data, this may not always be so.

Markov Chains

Maximum likelihood estimation of aggregated Markov processes.

We present a maximum likelihood method for the modelling of aggregated Markov processes. The method utilizes the joint probability density of the observed dwell time sequence as likelihood. A forward-backward recursive procedure is developed for efficient computation of the likelihood function and its derivatives with respect to the model parameters. Based on the calculated forward and backward vectors, analytical formulae for the derivatives of the likelihood function are derived. The method exploits the variable metric optimizer for search of the likelihood space. It converges rapidly and is numerically stable. Numerical examples are given to show the effectiveness of the method.

Markov Chains