PubMed HealthSearch

PubMed · 9796822

Optimizing potentials for the inverse protein folding problem.

Abstract

Inverse protein folding, which seeks to identify sequences that fold into a given structure, has been approached by threading candidate sequences onto the structure and scoring them with database-derived potentials. The sequences with the lowest energies are predicted to fold into that structure. It has been argued that the limited success of this type of approach is not due to the discrepancy between the scoring potential and the true potential but is rather due to the fact that sequences choose their lowest-energy structure rather than structures choosing the lowest-energy sequences. Here we develop a non-physical potential scheme optimized for the inverse folding problem. We maximize the average probability of success for a set of lattice proteins to obtain the optimal potential energy function, and show that the potential obtained by our method is more likely to produce successful predictions than the true potential.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

T L Chiu, R A Goldstein. 1998. Optimizing potentials for the inverse protein folding problem.. https://doi.org/10.1093/protein%2F11.9.749

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

NumSimEX: A method using EXX hydrogen exchange mass spectrometry to map the energetics of protein folding landscapes.

Hydrogen exchange mass spectrometry (HXMS) is a powerful tool to understand protein folding pathways and energetics. However, HXMS experiments to date have used exchange conditions termed EX1 or EX2 which limit the information that can be gained compared to the more general EXX exchange regime. If EXX behavior could be understood and analyzed, a single HXMS timecourse on an intact protein could fully map its folding landscape without requiring denaturation. To address this challenge, we developed a numerical simulation method called NumSimEX that models EXX exchange for arbitrarily complex folding pathways. NumSimEx fits protein folding dynamics to experimental HXMS data by iteratively comparing the simulated and experimental timecourses, allowing for determination of both kinetic and thermodynamic protein folding parameters. After analytically verifying NumSimEX's accuracy, we demonstrated its power on HXMS data from beta-2 microglobulin (β2M), a protein involved in dialysis-related amyloidosis. In particular, using NumSimEX, we identified three-state kinetics that near-perfectly matched experimental observation. This proof-of-principle application of NumSimEX sets the stage for harnessing HXMS to expand our understanding of proteins currently excluded from traditional protein folding methods. NumSimEX is freely available at https://github.com/JaswalLab/NumSimEX_Public.

Protein Folding

Favorable domain size in proteins.

BACKGROUND: It has been observed that single-domain proteins and domains in multidomain proteins favor a chain length in the range 100-150 amino acids. To understand the origin of the favored size, we construct an empirical function for the free energy of unfolding versus the chain length. The parameters in the function are derived by fitting to the energy of hydration, entropy and enthalpy of unfolding of nine proteins. Our energy function cannot be used to calculate the energetics accurately for individual proteins because the energetics also depend on other factors, such as the composition and the conformation of the protein. Nevertheless, the energy function statistically characterizes the general relationship between the free energy of unfolding and the size of the protein. RESULTS: The predicted optimal number of residues, which corresponds to the maximum free energy of unfolding, is 100. This is in agreement with a statistical analysis of protein domains derived from their experimental structures. When a chain is too short, our energy function indicates that the change in enthalpy of internal interactions is not favorable enough for folding because of the limited number of inter-residue contacts. A long chain is also unfavorable for a single domain because the cost of configurational entropy increases quadratically as a function of the chain length, whereas the favorable change in enthalpy of internal interactions increases linearly. CONCLUSIONS: Our study shows that the energetic balance is the dominant factor governing protein sizes and it forces a large protein to break into several domains during folding.

Protein Folding