PubMed HealthSearch

Biomedical subjects

K T Simons

Publications and source records attributed to K T Simons.

6 recordsLinked to original sources

Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins.

We describe the development of a scoring function based on the decomposition P(structure/sequence) proportional to P(sequence/structure) *P(structure), which outperforms previous scoring functions in correctly identifying native-like protein structures in large ensembles of compact decoys. The first term captures sequence-dependent features of protein structures, such as the burial of hydrophobic residues in the core, the second term, universal sequence-independent features, such as the assembly of beta-strands into beta-sheets. The efficacies of a wide variety of sequence-dependent and sequence-independent features of protein structures for recognizing native-like structures were systematically evaluated using ensembles of approximately 30,000 compact conformations with fixed secondary structure for each of 17 small protein domains. The best results were obtained using a core scoring function with P(sequence/structure) parameterized similarly to our previous work (Simons et al., J Mol Biol 1997;268:209-225] and P(structure) focused on secondary structure packing preferences; while several additional features had some discriminatory power on their own, they did not provide any additional discriminatory power when combined with the core scoring function. Our results, on both the training set and the independent decoy set of Park and Levitt (J Mol Biol 1996;258:367-392), suggest that this scoring function should contribute to the prediction of tertiary structure from knowledge of sequence and secondary structure.

Models, Statistical

Clustering of low-energy conformations near the native structures of small proteins.

Recent experimental studies of the denatured state and theoretical analyses of the folding landscape suggest that there are a large multiplicity of low-energy, partially folded conformations near the native state. In this report, we describe a strategy for predicting protein structure based on the working hypothesis that there are a greater number of low-energy conformations surrounding the correct fold than there are surrounding low-energy incorrect folds. To test this idea, 12 ensembles of 500 to 1,000 low-energy structures for 10 small proteins were analyzed by calculating the rms deviation of the Calpha coordinates between each conformation and every other conformation in the ensemble. In all 12 cases, the conformation with the greatest number of conformations within 4-A rms deviation was closer to the native structure than were the majority of conformations in the ensemble, and in most cases it was among the closest 1 to 5%. These results suggest that, to fold efficiently and retain robustness to changes in amino acid sequence, proteins may have evolved a native structure situated within a broad basin of low-energy conformations, a feature which could facilitate the prediction of protein structure at low resolution.

Computer Simulation

Contact order, transition state placement and the refolding rates of single domain proteins.

Theoretical studies have suggested relationships between the size, stability and topology of a protein fold and the rate and mechanisms by which it is achieved. The recent characterization of the refolding of a number of simple, single domain proteins has provided a means of testing these assertions. Our investigations have revealed statistically significant correlations between the average sequence separation between contacting residues in the native state and the rate and transition state placement of folding for a non-homologous set of simple, single domain proteins. These indicate that proteins featuring primarily sequence-local contacts tend to fold more rapidly and exhibit less compact folding transition states than those characterized by more non-local interactions. No significant relationship is apparent between protein length and folding rates, but a weak correlation is observed between length and the fraction of solvent-exposed surface area buried in the transition state. Anticipated strong relationships between equilibrium folding free energy and folding kinetics, or between chemical denaturant and temperature dependence-derived measures of transition state placement, are not apparent. The observed correlations are consistent with a model of protein folding in which the size and stability of the polypeptide segments organized in the transition state are largely independent of protein length, but are related to the topological complexity of the native state. The correlation between topological complexity and folding rates may reflect chain entropy contributions to the folding barrier.

Animals

Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.

We explore the ability of a simple simulated annealing procedure to assemble native-like structures from fragments of unrelated protein structures with similar local sequences using Bayesian scoring functions. Environment and residue pair specific contributions to the scoring functions appear as the first two terms in a series expansion for the residue probability distributions in the protein database; the decoupling of the distance and environment dependencies of the distributions resolves the major problems with current database-derived scoring functions noted by Thomas and Dill. The simulated annealing procedure rapidly and frequently generates native-like structures for small helical proteins and better than random structures for small beta sheet containing proteins. Most of the simulated structures have native-like solvent accessibility and secondary structure patterns, and thus ensembles of these structures provide a particularly challenging set of decoys for evaluating scoring functions. We investigate the effects of multiple sequence information and different types of conformational constraints on the overall performance of the method, and the ability of a variety of recently developed scoring functions to recognize the native-like conformations in the ensembles of simulated structures.

Bayes Theorem

Characterization of the free energy spectrum of peptostreptococcal protein L.

BACKGROUND: Native state hydrogen/deuterium exchange studies on cytochrome c and RNase H revealed the presence of excited states with partially formed native structure. We set out to determine whether such excited states are populated for a very small and simple protein, the IgG-binding domain of peptostreptococcal protein L. RESULTS: Hydrogen/deuterium exchange data on protein L in 0-1.2 M guanidine fit well to a simple model in which the only contributions to exchange are denaturant-independent local fluctuations and global unfolding. A substantial discrepancy emerged between unfolding free energy estimates from hydrogen/deuterium exchange and linear extrapolation of earlier guanidine denaturation experiments. A better determined estimate of the free energy of unfolding obtained by global analysis of a series of thermal denaturation experiments in the presence of 0-3 M guanidine was in good agreement with the estimate from hydrogen/deuterium exchange. CONCLUSIONS: For protein L under native conditions, there do not appear to be partially folded states with free energies intermediate between that of the folded and unfolded states. The linear extrapolation method significantly underestimates the free energy of folding of protein L due to deviations from linearity in the dependence of the free energy on the denaturant concentration.

Bacterial Proteins

Local sequence-structure correlations in proteins.

Considerable progress has been made in understanding the relationship between local amino acid sequence and local protein structure. Recent highlights include numerous studies of the structures adopted by short peptides, new approaches to correlating sequence patterns with structure patterns, and folding simulations using simple potentials.

Amino Acid Sequence