PubMed Health⌕ Search

Biomedical subjects

Stefan M Larson

Publications and source records attributed to Stefan M Larson.

7 recordsLinked to original sources

The family feud: do proteins with similar structures fold via the same pathway?

Theoretical and experimental studies of protein folding have suggested that the topology of the native state may be the most important factor determining the folding pathway of a protein, independent of its specific amino acid sequence. To test this concept, many experimental studies have been carried out with the aim of comparing the folding pathways of proteins that possess similar tertiary structures, but divergent sequences. Many of these studies focus on quantitative comparisons of folding transition state structures, as determined by Phi(f) value analysis of folding kinetic data. In some of these studies, folding transition state structures are found to be highly conserved, whereas in others they are not. We conclude that folds displaying more conserved transition state structures may have the most restricted number of possible folding pathways and that folds displaying low transition state structural conservation possess many potential pathways for reaching the native state.

Binding Sites↗

The relationship between conservation, thermodynamic stability, and function in the SH3 domain hydrophobic core.

To investigate the relationships between sequence conservation, protein stability, and protein function, we have measured the thermodynamic stability, folding kinetics, and in vitro peptide-binding activity of a large number of single-site substitutions in the hydrophobic core of the Fyn SH3 domain. Comparison of these data to that derived from an analysis of a large alignment of SH3 domain sequences revealed a very good correlation between the distinct pattern of conservation observed at each core position and the thermodynamic stability of mutants. Conservation was also found to correlate well with the unfolding rates of mutants, but not to the folding rates, suggesting that evolution selects more strongly for optimal native state packing interactions than for maximal folding rates. Structural analysis suggests that residue-residue core packing interactions are very similar in all SH3 domains, which provides an explanation for the correlation between conservation and mutant stability effects studied in a single SH3 domain. We also demonstrate a correlation between stability and the in vivo activity of mutants, and between conservation and activity. However, the relationship between conservation and activity was very strong only for the three most conserved hydrophobic core positions. The weaker correlation between activity and conservation seen at the other seven core positions indicates that maintenance of protein stability is the dominant selective pressure at these positions. In general, the pattern of conservation at hydrophobic core positions appears to arise from conserved packing constraints, and can be effectively utilized to predict the destabilizing effects of amino acid substitutions.

Amino Acid Sequence↗

Sequence optimization for native state stability determines the evolution and folding kinetics of a small protein.

Investigating the relative importance of protein stability, function, and folding kinetics in driving protein evolution has long been hindered by the fact that we can only compare modern natural proteins, the products of the very process we seek to understand, to each other, with no external references or baselines. Through a large-scale all-atom simulation of protein evolution, we have created a large diverse alignment of SH3 domain sequences which have been selected only for native state stability, with no other influencing factors. Although the average pairwise identity between computationally evolved and natural sequences is only 17%, the residue frequency distributions of the computationally evolved sequences are similar to natural SH3 sequences at 86% of the positions in the domain, suggesting that optimization for the native state structure has dominated the evolution of natural SH3 domains. Additionally, the positions which play a consistent role in the transition state of three well-characterized SH3 domains (by phi-value analysis) are structurally optimized for the native state, and vice versa. Indeed, we see a specific and significant correlation between sequence optimization for native state stability and conservation of transition state structure.

Amino Acid Sequence↗

Increased detection of structural templates using alignments of designed sequences.

Protein structure prediction by comparative modeling benefits greatly from the use of multiple sequence alignment information to improve the accuracy of structural template identification and the alignment of target sequences to structural templates. Unfortunately, this benefit is limited to those protein sequences for which at least several natural sequence homologues exist. We show here that the use of large diverse alignments of computationally designed protein sequences confers many of the same benefits as natural sequences in identifying structural templates for comparative modeling targets. A large-scale massively parallelized application of an all-atom protein design algorithm, including a simple model of peptide backbone flexibility, has allowed us to generate 500 diverse, non-native, high-quality sequences for each of 264 protein structures in our test set. PSI-BLAST searches using the sequence profiles generated from the designed sequences ("reverse" BLAST searches) give near-perfect accuracy in identifying true structural homologues of the parent structure, with 54% coverage. In 41 of 49 genomes scanned using reverse BLAST searches, at least one novel structural template (not found by the standard method of PSI-BLAST against PDB) is identified. Further improvements in coverage, through optimizing the scoring function used to design sequences and continued application to new protein structures beyond the test set, will allow this method to mature into a useful strategy for identifying distantly related structural templates.

Algorithms↗

Atomistic protein folding simulations on the submillisecond time scale using worldwide distributed computing.

Atomistic simulations of protein folding have the potential to be a great complement to experimental studies, but have been severely limited by the time scales accessible with current computer hardware and algorithms. By employing a worldwide distributed computing network of tens of thousands of PCs and algorithms designed to efficiently utilize this new many-processor, highly heterogeneous, loosely coupled distributed computing paradigm, we have been able to simulate hundreds of microseconds of atomistic molecular dynamics. This has allowed us to directly simulate the folding mechanism and to accurately predict the folding rate of several fast-folding proteins and polymers, including a nonbiological helix, polypeptide alpha-helices, a beta-hairpin, and a three-helix bundle protein from the villin headpiece. Our results demonstrate that one can reach the time scales needed to simulate fast folding using distributed computing, and that potential sets used to describe interatomic interactions are sufficiently accurate to reach the folded state with experimentally validated rates, at least for small proteins.

Algorithms↗

Residues participating in the protein folding nucleus do not exhibit preferential evolutionary conservation.

To what extent does natural selection act to optimize the details of protein folding kinetics? In an effort to address this question, the relationship between an amino acid's evolutionary conservation and its role in protein folding kinetics has been investigated intensively. Despite this effort, no consensus has been reached regarding the degree to which residues involved in native-like transition state structure (the folding nucleus) are conserved. Here we report the results of an exhaustive, systematic study of sequence conservation among residues known to participate in the experimentally (Phi-value) defined folding nuclei of all of the appropriately characterized proteins reported to date. We observe no significant evidence that these residues exhibit any anomalous sequence conservation. We do observe, however, a significant bias in the existing kinetic data: the mean sequence conservation of the residues that have been the subject of kinetic characterization is greater than the mean sequence conservation of all residues in 13 of 14 proteins studied. This systematic experimental bias gives rise to the previous observation that the median conservation of residues reported to participate in the folding nucleus is greater than the median conservation of all of the residues in a protein. When this bias is corrected (by comparing, for example, the conservation of residues known to participate in the folding nucleus with that of other, kinetically characterized residues) the previously reported preferential conservation is effectively eliminated. In contrast to well-established theoretical expectations, both poorly and highly conserved residues are apparently equally likely to participate in the protein-folding nucleus.

Bias↗

Thoroughly sampling sequence space: large-scale protein design of structural ensembles.

Modeling the inherent flexibility of the protein backbone as part of computational protein design is necessary to capture the behavior of real proteins and is a prerequisite for the accurate exploration of protein sequence space. We present the results of a broad exploration of sequence space, with backbone flexibility, through a novel approach: large-scale protein design to structural ensembles. A distributed computing architecture has allowed us to generate hundreds of thousands of diverse sequences for a set of 253 naturally occurring proteins, allowing exciting insights into the nature of protein sequence space. Designing to a structural ensemble produces a much greater diversity of sequences than previous studies have reported, and homology searches using profiles derived from the designed sequences against the Protein Data Bank show that the relevance and quality of the sequences is not diminished. The designed sequences have greater overall diversity than corresponding natural sequence alignments, and no direct correlations are seen between the diversity of natural sequence alignments and the diversity of the corresponding designed sequences. For structures in the same fold, the sequence entropies of the designed sequences cluster together tightly. This tight clustering of sequence entropies within a fold and the separation of sequence entropy distributions for different folds suggest that the diversity of designed sequences is primarily determined by a structure's overall fold, and that the designability principle postulated from studies of simple models holds in real proteins. This has important implications for experimental protein design and engineering, as well as providing insight into protein evolution.

Amino Acid Sequence↗