PubMed Health⌕ Search

Biomedical subjects

Amy E Keating

Publications and source records attributed to Amy E Keating.

9 recordsLinked to original sources

Ultra-fast evaluation of protein energies directly from sequence.

The structure, function, stability, and many other properties of a protein in a fixed environment are fully specified by its sequence, but in a manner that is difficult to discern. We present a general approach for rapidly mapping sequences directly to their energies on a pre-specified rigid backbone, an important sub-problem in computational protein design and in some methods for protein structure prediction. The cluster expansion (CE) method that we employ can, in principle, be extended to model any computable or measurable protein property directly as a function of sequence. Here we show how CE can be applied to the problem of computational protein design, and use it to derive excellent approximations of physical potentials. The approach provides several attractive advantages. First, following a one-time derivation of a CE expansion, the amount of time necessary to evaluate the energy of a sequence adopting a specified backbone conformation is reduced by a factor of 10(7) compared to standard full-atom methods for the same task. Second, the agreement between two full-atom methods that we tested and their CE sequence-based expressions is very high (root mean square deviation 1.1-4.7 kcal/mol, R2 = 0.7-1.0). Third, the functional form of the CE energy expression is such that individual terms of the expansion have clear physical interpretations. We derived expressions for the energies of three classic protein design targets-a coiled coil, a zinc finger, and a WW domain-as functions of sequence, and examined the most significant terms. Single-residue and residue-pair interactions are sufficient to accurately capture the energetics of the dimeric coiled coil, whereas higher-order contributions are important for the two more globular folds. For the task of designing novel zinc-finger sequences, a CE-derived energy function provides significantly better solutions than a standard design protocol, in comparable computation time. Given these advantages, CE is likely to find many uses in computational structural modeling.

Algorithms↗

Orientation and oligomerization specificity of the Bcr coiled-coil oligomerization domain.

The Bcr oligomerization domain, from the Bcr-Abl oncoprotein, is an attractive therapeutic target for treating leukemias because it is required for cellular transformation. The domain homodimerizes via an antiparallel coiled coil with an adjacent short, helical swap domain. Inspection of the coiled-coil sequence does not reveal obvious determinants of helix-orientation specificity, raising the possibility that the antiparallel orientation preference and/or the dimeric oligomerization state are due to interactions of the swap domains. To better understand how structural specificity is encoded in Bcr, coiled-coil constructs containing either an N- or C-terminal cysteine were synthesized without the swap domain. When cross-linked to adopt exclusively parallel or antiparallel orientations, these showed similar circular dichroism spectra. Both constructs formed coiled-coil dimers, but the antiparallel construct was approximately 16 degrees C more stable than the parallel to thermal denaturation. Equilibrium disulfide-exchange studies confirmed that the isolated coiled-coil homodimer shows a very strong preference for the antiparallel orientation. We conclude that the orientation and oligomerization preferences of Bcr are not caused by the presence of the swap domains, but rather are directly encoded in the coiled-coil sequence. We further explored possible determinants of structural specificity by mutating residues in the d position of the coiled-coil core. Some of the mutations caused a change in orientation specificity, and all of the mutations led to the formation of higher-order oligomers. In the absence of the swap domain, these residues play an important role in disfavoring alternate states and are especially important for encoding dimeric oligomerization specificity.

Amino Acid Sequence↗

Structure-based prediction of bZIP partnering specificity.

Predicting protein interaction specificity from sequence is an important goal in computational biology. We present a model for predicting the interaction preferences of coiled-coil peptides derived from bZIP transcription factors that performs very well when tested against experimental protein microarray data. We used only sequence information to build atomic-resolution structures for 1711 dimeric complexes, and evaluated these with a variety of functions based on physics, learned empirical weights or experimental coupling energies. A purely physical model, similar to those used for protein design studies, gave reasonable performance. The results were improved significantly when helix propensities were used in place of a structurally explicit model to represent the unfolded reference state. Further improvement resulted upon accounting for residue-residue interactions in competing states in a generic way. Purely physical structure-based methods had difficulty capturing core interactions accurately, especially those involving polar residues such as asparagine. When these terms were replaced with weights from a machine-learning approach, the resulting model was able to correctly order the stabilities of over 6000 pairs of complexes with greater than 90% accuracy. The final model is physically interpretable, and suggests specific pairs of residues that are important for bZIP interaction specificity. Our results illustrate the power and potential of structural modeling as a method for predicting protein interactions and highlight obstacles that must be overcome to reach quantitative accuracy using a de novo approach. Our method shows unprecedented performance in predicting protein-protein interaction specificity accurately using structural modeling and suggests that predicting coiled-coil interactions generally may be within reach.

Basic-Leucine Zipper Transcription Factors↗

Coarse-graining protein energetics in sequence variables.

We show that cluster expansions (CE), previously used to model solid-state materials with binary or ternary configurational disorder, can be extended to the protein design problem. We present a generalized CE framework, in which properties such as energy can be unambiguously expanded in the amino-acid sequence space. The CE coarse grains over nonsequence degrees of freedom (e.g., side-chain conformations) and thereby simplifies the problem of designing proteins, or predicting the compatibility of a sequence with a given structure, by many orders of magnitude. The CE is physically transparent, and can be evaluated through linear regression on the energies of training sequences. We show, as example, that good prediction accuracy is obtained with up to pairwise interactions for a coiled-coil backbone, and that triplet interactions are important in the energetics of a more globular zinc-finger backbone.

Algorithms↗

AVID: an integrative framework for discovering functional relationships among proteins.

BACKGROUND: Determining the functions of uncharacterized proteins is one of the most pressing problems in the post-genomic era. Large scale protein-protein interaction assays, global mRNA expression analyses and systematic protein localization studies provide experimental information that can be used for this purpose. The data from such experiments contain many false positives and false negatives, but can be processed using computational methods to provide reliable information about protein-protein relationships and protein function. An outstanding and important goal is to predict detailed functional annotation for all uncharacterized proteins that is reliable enough to effectively guide experiments. RESULTS: We present AVID, a computational method that uses a multi-stage learning framework to integrate experimental results with sequence information, generating networks reflecting functional similarities among proteins. We illustrate use of the networks by making predictions of detailed Gene Ontology (GO) annotations in three categories: molecular function, biological process, and cellular component. Applied to the yeast Saccharomyces cerevisiae, AVID provides 37,451 pair-wise functional linkages between 4,191 proteins. These relationships are approximately 65-78% accurate, as assessed by cross-validation testing. Assignments of highly detailed functional descriptors to proteins, based on the networks, are estimated to be approximately 67% accurate for GO categories describing molecular function and cellular component and approximately 52% accurate for terms describing biological process. The predictions cover 1,490 proteins with no previous annotation in GO and also assign more detailed functions to many proteins annotated only with less descriptive terms. Predictions made by AVID are largely distinct from those made by other methods. Out of 37,451 predicted pair-wise relationships, the greatest number shared in common with another method is 3,413. CONCLUSION: AVID provides three networks reflecting functional associations among proteins. We use these networks to generate new, highly detailed functional predictions for roughly half of the yeast proteome that are reliable enough to drive targeted experimental investigations. The predictions suggest many specific, testable hypotheses. All of the data are available as downloadable files as well as through an interactive website at http://web.mit.edu/biology/keating/AVID. Thus, AVID will be a valuable resource for experimental biologists.

Algorithms↗

Design of a heterospecific, tetrameric, 21-residue miniprotein with mixed alpha/beta structure.

The study of short, autonomously folding peptides, or "miniproteins," is important for advancing our understanding of protein stability and folding specificity. Although many examples of synthetic alpha-helical structures are known, relatively few mixed alpha/beta structures have been successfully designed. Only one mixed-secondary structure oligomer, an alpha/beta homotetramer, has been reported thus far. In this report, we use structural analysis and computational design to convert this homotetramer into the smallest known alpha/beta-heterotetramer. Computational screening of many possible sequence/structure combinations led efficiently to the design of short, 21-residue peptides that fold cooperatively and autonomously into a specific complex in solution. A 1.95 A crystal structure reveals how steric complementarity and charge patterning encode heterospecificity. The first- and second-generation heterotetrameric miniproteins described here will be useful as simple models for the analysis of protein-protein interaction specificity and as structural platforms for the further elaboration of folding and function.

Amino Acid Sequence↗

Predicting specificity in bZIP coiled-coil protein interactions.

We present a method for predicting protein-protein interactions mediated by the coiled-coil motif. When tested on interactions between nearly all human and yeast bZIP proteins, our method identifies 70% of strong interactions while maintaining that 92% of predictions are correct. Furthermore, cross-validation testing shows that including the bZIP experimental data significantly improves performance. Our method can be used to predict bZIP interactions in other genomes and is a promising approach for predicting coiled-coil interactions more generally.

Amino Acid Motifs↗

Comprehensive identification of human bZIP interactions with coiled-coil arrays.

In eukaryotes, the combinatorial association of sequence-specific DNA binding proteins is essential for transcription. We have used protein arrays to test 492 pairings of a nearly complete set of coiled-coil strands from human basic-region leucine zipper (bZIP) transcription factors. We find considerable partnering selectivity despite the bZIPs' homologous sequences. The interaction data are of high quality, as assessed by their reproducibility, reciprocity, and agreement with previous observations. Biophysical studies in solution support the relative binding strengths observed with the arrays. New associations provide insights into the circadian clock and the unfolded protein response.

Amino Acid Sequence↗

The structure of VibH represents nonribosomal peptide synthetase condensation, cyclization and epimerization domains.

Nonribosomal peptide synthetases (NRPSs) are large, multidomain enzymes that biosynthesize medically important natural products. We report the crystal structure of the free-standing NRPS condensation (C) domain VibH, which catalyzes amide bond formation in the synthesis of vibriobactin, a Vibrio cholerae siderophore. Despite low sequence identity, NRPS condensation enzymes are structurally related to chloramphenicol acetyltransferase (CAT) and dihydrolipoamide acyltransferases. However, although the latter enzymes are homotrimers, VibH is a monomeric pseudodimer. The VibH structure is representative of both NRPS condensation and epimerization domains, as well as the condensation-variant cyclization domains, which are all expected to be monomers. Surprisingly, despite favorable positioning in the active site, a universally conserved histidine important in CAT and in other C domains is not critical for general base catalysis in VibH.

Acyltransferases↗