PubMed Health⌕ Search

Biomedical subjects

B A Reva

Publications and source records attributed to B A Reva.

At least 19 recordsLinked to original sources

Averaging interaction energies over homologs improves protein fold recognition in gapless threading.

Protein structure prediction is limited by the inaccuracy of the simplified energy functions necessary for efficient sorting over many conformations. It was recently suggested (Finkelstein, Phys Rev Lett 1998;80:4823-4825) that these errors can be reduced by energy averaging over a set of homologous sequences. This conclusion is confirmed in this study by testing protein structure recognition in gapless threading. The accuracy of recognition was estimated by the Z-score values obtained in gapless threading tests. For threading, we used 20 target proteins, each having from 20 to 70 homologs taken from the HSSP sequence base. The energy of the native structures was compared with the energy from 34 to 75 thousand of alternative structures generated by threading. The energy calculations were done with our recently developed Calpha atom-based phenomenological potentials. We show that averaging of protein energies over homologs reduces the Z-score from approximately -6.1 (average Z-score for individual chains) to approximately -8.1. This means that a correct fold can be found among 3 x 10(9) random folds in the first case and among 3 x 10(15) in the second. Such increase in selectivity is important for recognition of protein folds.

Cytochrome c Group↗

What is the probability of a chance prediction of a protein structure with an rmsd of 6 A?

BACKGROUND: The root mean square deviation (rmsd) between corresponding atoms of two protein chains is a commonly used measure of similarity between two protein structures. The smaller the rmsd is between two structures, the more similar are these two structures. In protein structure prediction, one needs the rmsd between predicted and experimental structures for which a prediction can be considered to be successful. Success is obvious only when the rmsd is as small as that for closely homologous proteins (< 3 A). To estimate the quality of the prediction in the more general case, one has to compare the native structure not only with the predicted one but also with randomly chosen protein-like folds. One can ask: how many such structures must be considered to find a structure with a given rmsd from the native structure? RESULTS: We calculated the rmsd values between native structures of 142 proteins and all compact structures obtained in the threading of these protein chains over 364 non-homologous structures. The rmsd distributions have a Gaussian form, with the average rmsd approximately proportional to the radius of gyration. CONCLUSIONS: We estimated the number of protein-like structures required to obtain a structure within an rmsd of 6 A to be 10(4)-10(5) for chains of 60-80 residues and 10(11)-10(12) structures for chains of 160-200 residues. The probability of obtaining a 6 A rmsd by chance is so remote that when such structures are obtained from a prediction algorithm, it should be considered quite successful.

Databases, Factual↗

Optimization of protein structure on lattices using a self-consistent field approach.

Lattice modeling of proteins is commonly used to study the protein folding problem. The reduced number of possible conformations of lattice models enormously facilitates exploration of the conformational space. In this work, we suggest a method to search for the optimal lattice models that reproduced the off-lattice structures with minimal errors in geometry and energetics. The method is based on the self-consistent field optimization of a combined pseudoenergy function that includes two force fields: an "interaction field," that drives the residues to optimize the chain energy, and a "geometrical field," that attracts the residues towards their native positions. By varying the contributions of these force fields in the combined pseudoenergy, one can also test the accuracy of potentials: the better the potentials, i.e., the more accurate the "interaction field," and the smaller the contribution of the "geometrical field" required for building accurate lattice models.

Models, Chemical↗

Recognition of protein structure on coarse lattices with residue-residue energy functions.

We suggest and test potentials for the modeling of protein structure on coarse lattices. The coarser the lattice, the more complete and faster is the exploration of the conformational space of a molecule. However, there are inevitable energy errors in lattice modeling caused by distortions in distances between interacting residues; the coarser the lattice, the larger are the energy errors. It is generally believed that an improvement in the accuracy of lattice modelling can be achieved only by reducing the lattice spacing. We reduce the errors on coarse lattices with lattice-adapted potentials. Two methods are used: in the first approach, 'lattice-derived' potentials are obtained directly from a database of lattice models of protein structure; in the second approach, we derive 'lattice-adjusted' potentials using our previously developed method of statistical adjustment of the 'off-lattice' energy functions for lattices. The derivation of off-lattice Calpha atom-based distance-dependent pairwise potentials has been reported previously. The accuracy of 'lattice-derived', 'lattice-adjusted' and 'off-lattice' potentials is estimated in threading tests. It is shown that 'lattice-derived' and 'lattice-adjusted' potentials give virtually the same accuracy and ensure reasonable protein fold recognition on the coarsest considered lattice (spacing 3.8 A), however, the 'off-lattice' potentials, which efficiently recognize off-lattice folds, do not work on this lattice, mainly because of the errors in short-range interactions between neighboring residues.

Models, Chemical↗

Residue-residue mean-force potentials for protein structure recognition.

We present two new sets of energy functions for protein structure recognition, given the primary sequence of amino acids along the polypeptide chain. The first set of potentials is based on the positions of alpha- and the second on positions of beta- and alpha-carbon atoms of amino acid residues. The potentials are derived using a theory of Boltzmann-like statistics of protein structure. The energy terms incorporate both long-range interactions between residues remote along a chain and short-range interactions between near neighbors. Distance dependence is approximated by a piecewise constant function defined on intervals of equal size. The size of the interval is optimized to preserve as much detail as possible without introducing excessive error due to limited statistics. A database of 214 non-homologous proteins was used both for the derivation of the potentials, and for the 'threading' test originally suggested by Hendlich et al. (1990) J. Mol. Biol., 216, 167-180. Special care is taken to avoid systematic error in this test. For threading, we used 100 non-homologous protein chains of 60-205 residues. The energy of each of the native structures was compared with the energy of 43,000 to 19,000 alternative structures generated by threading. Of these 100 native structures, 92 have the lowest energy with alpha-carbon-based potentials and, even more, 98 of these 100 structures, have the lowest energy with the beta- and alpha-carbon based potentials.

Databases, Factual↗

Accurate mean-force pairwise-residue potentials for discrimination of protein folds.

We present two new sets of energy functions for protein structure recognition. The first set of potentials is based on the positions of alpha- and the second on positions of beta-carbon atoms of amino acid residues. The potentials are derived using a theory of Boltzmann-like statistics of protein structure by Finkelstein et al. The energy terms incorporate both long-range interactions between residues remote along a chain and short-range interactions between near neighbors. Distance-dependence is approximated by a piecewise constant function defined on intervals of equal size. The size of this interval is optimized. A database of 222 non-homologous proteins was used both for the derivation of the potentials, and for the "threading" test originally suggested by Hendlich et al. For threading, we used 102 non-homologous protein chains of 60 to 200 residues. The energy of each of the native structures was compared with the energy of 45 to 20 thousand alternative structures generated by threading. Of these 102 native structures 94 have the lowest energy with alpha-carbon-based potentials, and even more, 100 of these 102 structures, have the lowest energy with the beta-carbon-based potentials.

Computer Simulation↗

Adjusting potential energy functions for lattice models of chain molecules.

Lattice models of proteins can approximate off-lattice structure to arbitrary precision with RMS (root mean squared) deviations roughly equal to half the lattice spacing (Rykunov et al., Proteins 22:100-109, 1995; Reva et al., J. Comp. Biol., 1996). However, even small distortions in the positions of chain links lead to significant errors in lattice-based energy calculations (Reva et al., J. Comp. Chem., 1996). These errors arise mainly from rigid interactions (such as steric repulsion) which change their energies considerably at a range which is much smaller than the usual accuracy of lattice modeling (> 1.0 A). To reduce this error, we suggest a procedure of adjusting energy functions to a given lattice. The general approach is illustrated with energy calculations based on pairwise potentials by Kolinski et al. (J. Chem. Phys. 98:1-14, 1993). At all the lattice spacings, from 0.5-3.8 A, the lattice-adjusted potentials improve the accuracy of lattice-based energy calculations and increase the correlations between off-lattice and lattice energies.

Models, Chemical↗

Building self-avoiding lattice models of proteins using a self-consistent field optimization.

We present an algorithm to build self-avoiding lattice models of chain molecules with low RMS deviation from their actual 3D structures. To find the optimal coordinates for the lattice chain model, we minimize a function that consists of three terms: (1) the sum of squared deviations of link coordinates on a lattice from their off-lattice values, (2) the sum of "short-range" terms, penalizing violation of chain connectivity, and (3) the sum of "long-range" repulsive terms, penalizing chain self-intersections. We treat this function as a chain molecule "energy" and minimize it using self-consistent field (SCF) theory to represent the pairwise link repulsions as 3D fields acting on the links. The statistical mechanics of chain molecules enables computation of the chain distribution in this field on the lattice. The field is refined by iteration to become self-consistent with the chain distribution, then dynamic programming is used to find the optimal lattice model as the "lowest-energy" chain pathway in this SCF. We have tested the method on one of the coarsest (and most difficult) lattices used for model building on proteins of all structural types and show that the method is adequate for building self-avoiding models of proteins with low RMS deviations from the actual structures.

Algorithms↗

Search for the most stable folds of protein chains: I. Application of a self-consistent molecular field theory to a problem of protein three-dimensional structure prediction.

We present a general approach to the prediction of 3-D folds of protein chains from their amino acid sequences. The approach is based on the use of the self-consistent molecular field theory for long-range interactions, the use of 1-D statistical mechanics for short-range interactions and on the discovery that there is and should only be a relatively small discrete set of folding patterns. This makes it possible to examine the full variety of 'potentially stable' folds and to determine the thermodynamically stable structure. In this paper, we give the general theoretical background of the approach. The encouraging results of the application of this approach to beta-domains are described in another paper.

Models, Molecular↗

Search for the most stable folds of protein chains: II. Computation of stable architectures of beta-proteins using a self-consistent molecular field theory.

In a preceding paper we presented a novel approach to computation of 3-D folds of protein chains from their amino acid sequences. This approach is a physically correct generalization of the 'threading' methods. It is based on a self-consistent molecular field theory and on a physical theory of protein folding patterns, which make it possible to examine all the variety of 'potentially stable' folding patterns and all the variety of the chain conformations within each of them and to determine the thermodynamically stable structure. In this paper, we apply this approach to single out stable folding patterns and conformations for the chains of beta-sandwich proteins and show that the similarity of the calculated and observed structures is usually rather close.

Amino Acid Sequence↗

Accurate general method for lattice approximation of three-dimensional structure of a chain molecule.

An algorithm based on dynamic programming gives the lattice models having the minimal RMS deviations from the actual folds of protein (RNA, etc.) chains for a given lattice and a given orientation of the macromolecule relative to the lattice. The algorithm is applicable for 3-D lattices of any kind. The accuracy of the lattice approximation increases when the distance between neighbor chain links is not rigidly fixed. Special repulsive potentials facilitate generation of self-avoiding lattice chains. The results of model building show the efficiency and precision of this proposed general method when compared with others.

Algorithms↗

Constructing lattice models of protein chains with side groups.

An algorithm to construct lattice models of polymers with side chains is presented. A search for the global minimum of the error function for a given lattice-to-chain orientation is done by dynamic programming, making the search both fast and complete. Application of the algorithm is illustrated by constructing lattice models for 12 proteins of different sizes and structural types.

Algorithms↗

Search for the stable state of a short chain in a molecular field.

A general approach is developed to search for stable structures of short chain fragments (e.g. of loops or bound oligopeptides) in a given molecular field. This molecular field is produced by the remaining part of a globule or by any other surface with a defined spatial structure. The fragment must be short enough to have no pronounced long-range interactions within itself. The method is illustrated by calculation of the 3-D structures of two loops of bovine pancreatic trypsin inhibitor (BPTI). Computations are based on a lattice model of conformational space and on strict and fast algorithms of 1-D statistical mechanics and dynamic programming (which are very similar in essence). This makes a search of oligopeptide structures only several times (and not several orders of magnitude) longer than that of a dipeptide.

Algorithms↗

A new approach to the design of a sequence with the highest affinity for a molecular surface.

We describe an algorithm to design the primary structures for peptides which must have the strongest binding to a given molecular surface. This problem cannot be solved by a direct combinatorial sorting, because of an enormous number of possible primary and spatial structures. The approach to solve this problem is to describe a state of each residue by two variables: (i) amino acid type and (ii) 3-D coordinate, and to minimize binding energy over all these variables simultaneously. For short chains which have no long-range interactions within themselves, this minimization can be done easily and efficiently by dynamic programming. We also discuss the problem of how to estimate specificity of binding and how to deduce a sequence with maximal specificity for a given surface. We show that this sequence can be deduced by the same algorithm after some modification of energetic parameters.

Algorithms↗

A search for the most stable folds of protein chains.

It is generally believed that it is not sensible to search for a thermodynamically stable structure of a protein because neither a molecule nor a computer can look through all the 3(100) possible (for 100 residues) chain conformations. Here we show that the use of a molecular field theory for the long-range interactions, the use of one-dimensional statistical mechanics for the short-range ones and the discovery that there are and there must be only a small discrete set of folding patterns, make it possible to examine all the variety of 'potentially stable' structures. The general approach and its application is demonstrated here by calculation of stable folds for some beta domains. The most stable of these folds correspond to the observed structures.

Chemical Phenomena↗

[When and how can homologs overcome errors in the energy estimates and make the 3D structure prediction possible].

One still cannot predict the 3D fold of a protein from its amino acid sequence, mainly because of errors in the energy estimates underlying the prediction. However, a recently developed theory [1] shows that having a set of homologs (i.e., the chains with equal, in despite of numerous mutations, 3D folds) one can average the potential of each interaction over the homologs and thus predict the common 3D fold of protein family even when a correct fold prediction for an individual sequence is impossible because the energies are known only approximately. This theoretical conclusion has been verified by simulation of the energy spectra of simplified models of protein chains [2], and the further investigation of these simplified models shows that their true "native" fold can be found by folding of the chain where each interaction potential is averaged over the homologs. In conclusion, the applicability of the "homolog-averaging" approach is tested by recognition of real protein 3D structures. Both the gapless threading of sequences onto the known protein folds [3] and the more practically important gapped threading (which allows to consider not only the known 3D structures, but the more or less similar to them folds as well) shows a significant increase in selectivity of the native chain fold recognition.

Computer Simulation↗

[Determination of the folding of globular protein chain by the self-consistent field method].

It is shown that and how it is possible to single out the chain fold which is thermo-dynamically most stable. The suggested approach is based on two physical ideas: A "molecular field" approximation permits to examine all protein structures which belong to the same "folding pattern". Only a limited set of the "potentially stable" folding patterns have to be examined. The general approach is illustrated by calculations of the stable folds for two beta-domains.

Mathematics↗