PubMed Health⌕ Search

Biomedical subjects

P Koehl

Publications and source records attributed to P Koehl.

At least 19 recordsLinked to original sources

Refined crystallographic structure of Pseudomonas aeruginosa exotoxin A and its implications for the molecular mechanism of toxicity.

Exotoxin A of Pseudomonas aeruginosa asserts its cellular toxicity through ADP-ribosylation of translation elongation factor 2, predicated on binding to specific cell surface receptors and intracellular trafficking via a complex pathway that ultimately results in translocation of an enzymatic activity into the cytoplasm. In early work, the crystallographic structure of exotoxin A was determined to 3.0 A resolution, revealing a tertiary fold having three distinct structural domains; subsequent work has shown that the domains are individually responsible for the receptor binding (domain I), transmembrane targeting (domain II), and ADP-ribosyl transferase (domain III) activities, respectively. Here, we report the structures of wild-type and W281A mutant toxin proteins at pH 8.0, refined with data to 1.62 A and 1.45 A resolution, respectively. The refined models clarify several ionic interactions within structural domains I and II that may modulate an obligatory conformational change that is induced by low pH. Proteolytic cleavage by furin is also obligatory for toxicity; the W281A mutant protein is substantially more susceptible to cleavage than the wild-type toxin. The tertiary structures of the furin cleavage sites of the wild-type and W281 mutant toxins are similar; however, the mutant toxin has significantly higher B-factors around the cleavage site, suggesting that the greater susceptibility to furin cleavage is due to increased local disorder/flexibility at the site, rather than to differences in static tertiary structure. Comparison of the refined structures of full-length toxin, which lacks ADP-ribosyl transferase activity, to that of the enzymatic domain alone reveals a salt bridge between Arg467 of the catalytic domain and Glu348 of domain II that restrains the substrate binding cleft in a conformation that precludes NAD+ binding. The refined structures of exotoxin A provide precise models for the design and interpretation of further studies of the mechanism of intoxication.

ADP Ribose Transferases↗

Protein structure similarities.

Comparison of protein structures can reveal distant evolutionary relationships that would not be detected by sequence information alone. This helps to infer functional properties. In recent years, many methods for pairwise protein structure alignment have been proposed and are now available on the World Wide Web. Although these methods have made it possible to compare all available protein structures, they also highlight the remaining difficulties in defining a reliable score for protein structure similarities.

Amino Acid Motifs↗

The ASTRAL compendium for protein structure and sequence analysis.

The ASTRAL compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. The SPACI scores included in the system summarize the overall characteristics of a protein structure. A structural alignments database indicates residue equivalencies in superimposed protein domain structures. The PDB sequence-map files provide a linkage between the amino acid sequence of the molecule studied (SEQRES records in a database entry) and the sequence of the atoms experimentally observed in the structure (ATOM records). These maps are combined with information in the SCOPdatabase to provide sequences of protein domains. Selected subsets of the domain database, with varying degrees of similarity measured in several different ways, are also available. ASTRALmay be accessed at http://astral.stanford.edu/

Amino Acid Sequence↗

Constructing side chains on near-native main chains for ab initio protein structure prediction.

Is there value in constructing side chains while searching protein conformational space during an ab initio simulation? If so, what is the most computationally efficient method for constructing these side chains? To answer these questions, four published approaches were used to construct side chain conformations on a range of near-native main chains generated by ab initio protein structure prediction methods. The accuracy of these approaches was compared with a naive approach that selects the most frequently observed rotamer for a given amino acid to construct side chains. An all-atom conditional probability discriminatory function is useful at selecting conformations with overall low all-atom root mean square deviation (r.m.s.d.) and the discrimination improves on sets that are closer to the native conformation. In addition, the naive approach performs as well as more sophisticated methods in terms of the percentage of chi(1) angles built accurately and the all-atom r. m.s.d., between the native and near-native conformations. The results suggest that the naive method would be extremely useful for fast and efficient side chain construction on vast numbers of conformations for ab initio prediction of protein structure.

Amino Acids↗

De novo protein design. I. In search of stability and specificity.

We have developed a fully automated protein design strategy that works on the entire sequence of the protein and uses a full atom representation. At each step of the procedure, an all-atom model of the protein is built using the template protein structure and the current designed sequence. The energy of the model is used to drive a Monte Carlo optimization in sequence space: random moves are either accepted or rejected based on the Metropolis criterion. We rely on the physical forces that stabilize native protein structures to choose the optimum sequence. Our energy function includes van der Waals interactions, electrostatics and an environment free energy. Successful protein design should be specific and generate a sequence compatible with the template fold and incompatible with competing folds. We impose specificity by maintaining the amino acid composition constant, based on the random energy model. The specificity of the optimized sequence is tested by fold recognition techniques. Successful sequence designs for the B1 domain of protein G, for the lambda repressor and for sperm whale myoglobin are presented. We show that each additional term of the energy function improves the performance of our design procedure: the van der Waals term ensures correct packing, the electrostatics term increases the specificity for the correct native fold, and the environment solvation term ensures a correct pattern of buried hydrophobic and exposed hydrophilic residues. For the globin family, we show that we can design a protein sequence that is stable in the myoglobin fold, yet incompatible with the very similar hemoglobin fold.

Amino Acid Sequence↗

De novo protein design. II. Plasticity in sequence space.

It is generally accepted that many different protein sequences have similar folded structures, and that there is a relatively high probability that a new sequence possesses a previously observed fold. An indirect consequence of this is that protein design should define the sequence space accessible to a given structure, rather than providing a single optimized sequence. We have recently developed a new approach for protein sequence design, which optimizes the complete sequence of a protein based on the knowledge of its backbone structure, its amino acid composition and a physical energy function including van der Waals interactions, electrostatics, and environment free energy. The specificity of the designed sequence for its template backbone is imposed by keeping the amino acid composition fixed. Here, we show that our procedure converges in sequence space, albeit not to the native sequence of the protein. We observe that while polar residues are well conserved in our designed sequences, non-polar amino acids at the surface of a protein are often replaced by polar residues. The designed sequences provide a multiple alignment of sequences that all adopt the same three-dimensional fold. This alignment is used to derive a profile matrix for chicken triose phosphate isomerase, TIM. The matrix is found to recognize significantly the native sequence for TIM, as well as closely related sequences. Possible application of this approach to protein fold recognition is discussed.

Amino Acid Sequence↗

Structure-based conformational preferences of amino acids.

Proteins can be very tolerant to amino acid substitution, even within their core. Understanding the factors responsible for this behavior is of critical importance for protein engineering and design. Mutations in proteins have been quantified in terms of the changes in stability they induce. For example, guest residues in specific secondary structures have been used as probes of conformational preferences of amino acids, yielding propensity scales. Predicting these amino acid propensities would be a good test of any new potential energy functions used to mimic protein stability. We have recently developed a protein design procedure that optimizes whole sequences for a given target conformation based on the knowledge of the template backbone and on a semiempirical potential energy function. This energy function is purely physical, including steric interactions based on a Lennard-Jones potential, electrostatics based on a Coulomb potential, and hydrophobicity in the form of an environment free energy based on accessible surface area and interatomic contact areas. Sequences designed by this procedure for 10 different proteins were analyzed to extract conformational preferences for amino acids. The resulting structure-based propensity scales show significant agreements with experimental propensity scale values, both for alpha-helices and beta-sheets. These results indicate that amino acid conformational preferences are a natural consequence of the potential energy we use. This confirms the accuracy of our potential and indicates that such preferences should not be added as a design criterion.

Amino Acids↗

Accuracy of side-chain prediction upon near-native protein backbones generated by Ab initio folding methods.

The ab initio folding problem can be divided into two sequential tasks of approximately equal computational complexity: the generation of native-like backbone folds and the positioning of side chains upon these backbones. The prediction of side-chain conformation in this context is challenging, because at best only the near-native global fold of the protein is known. To test the effect of displacements in the protein backbones on side-chain prediction for folds generated ab initio, sets of near-native backbones (< or = 4 A C alpha RMS error) for four small proteins were generated by two methods. The steric environment surrounding each residue was probed by placing the side chains in the native conformation on each of these decoys, followed by torsion-space optimization to remove steric clashes on a rigid backbone. We observe that on average 40% of the chi1 angles were displaced by 40 degrees or more, effectively setting the limits in accuracy for side-chain modeling under these conditions. Three different algorithms were subsequently used for prediction of side-chain conformation. The average prediction accuracy for the three methods was remarkably similar: 49% to 51% of the chi1 angles were predicted correctly overall (33% to 36% of the chi1+2 angles). Interestingly, when the inter-side-chain interactions were disregarded, the mean accuracy increased. A consensus approach is described, in which side-chain conformations are defined based on the most frequently predicted chi angles for a given method upon each set of near-native backbones. We find that consensus modeling, which de facto includes backbone flexibility, improves side-chain prediction: chi1 accuracy improved to 51-54% (36-42% of chi1+2). Implications of a consensus method for ab initio protein structure prediction are discussed.

Models, Chemical↗

Influence of protein structure databases on the predictive power of statistical pair potentials.

A long standing goal in protein structure studies is the development of reliable energy functions that can be used both to verify protein models derived from experimental constraints as well as for theoretical protein folding and inverse folding computer experiments. In that respect, knowledge-based statistical pair potentials have attracted considerable interests recently mainly because they include the essential features of protein structures as well as solvent effects at a low computing cost. However, the basis on which statistical potentials are derived have been questioned. In this paper, we investigate statistical pair potentials derived from protein three-dimensional structures, addressing in particular questions related to the form of these potentials, as well as to the content of the database from which they are derived. We have shown that statistical pair potentials depend on the size of the proteins included in the database, and that this dependence can be reduced by considering only pairs of residue close in space (i.e., with a cutoff of 8 A). We have shown also that statistical potentials carry a memory of the quality of the database in terms of the amount and diversity of secondary structure it contains. We find, for example, that potentials derived from a database containing alpha-proteins will only perform best on alpha-proteins in fold recognition computer experiments. We believe that this is an overall weakness of these potentials, which must be kept in mind when constructing a database.

Chemical Phenomena↗

The inverse protein folding problem: self consistent mean field optimisation of a structure specific mutation matrix.

The goal of the inverse folding problem is to supply a list of sequences compatible with a known protein structure. If two-body interactions are taken into account in energy calculations, an exhaustive exploration of the energy landscape in sequence space cannot be achieved because of the huge number of possible combinations. To circumvent this problem, we propose a method in which multiple copies corresponding to every possible side-chain type are attached to each C alpha position in the protein. The weights of each copy (stored in the sequence matrix SM) are refined using mean field theory: each side-chain copy interacts with the mean field generated by all possible side-chain copies at neighbouring positions, weighted by their respective probabilities. The potential energy is simply taken to be amino acid pair potentials of mean force. The method converges in a few cycles to a self-consistent solution. The refined matrix does not depend on the starting point; therefore the method succeeds in removing memory effects. Starting solely from the backbone of the known structure, and without information from the initial sequence, the final sequence matrix SM is shown to be able to retrieve significant sequence information, as observed through a series of structure-recognizes-sequence(s) computer experiments. The issue of specificity is discussed in detail.

Amino Acid Sequence↗

The native sequence determines sidechain packing in a protein, but does optimal sidechain packing determine the native sequence?

Globular proteins have highly compact structures and the corresponding packing interactions are widely considered as the principal determinant of the native structure. It is therefore important that theoretical approaches to protein design explicitly take in account packing, which requires that a full atomic representation of the designed protein is maintained. As a first step towards this goal, we have developed in this report an inverse folding algorithm with the aim of specifically designing amino acid sequences which optimise sidechain packing for a given protein fold. The design is performed by a global Monte Carlo optimisation in sequence space, with constant amino acid composition and a full-atom representation of the various protein models. Packing is defined by a Lennard-Jones potential. The program was tested by designing stable sequence variants for the chymotrypsin inhibitor fold. The final protein models showed an increase in intramolecular atomic contacts and a decrease in the overall volume compared to the native structure. Starting from the backbone only of the target structure, the algorithm did gradually retrieve reliable though limited sequence information. Higher compatibility might be achieved by improving the potential, however our results suggest that packing interactions are an essential element of a yet-to-be-defined successful energy function for protein design.

Algorithms↗

Mean-field minimization methods for biological macromolecules.

Simulations of macromolecular structures involve the minimization of a potential-energy function that presents many local minima. Mean-field theory provides a tool that enables us to escape these minima, by enhancing sampling in conformational space. The number of applications of this technique has increased significantly over the past year, enabling problems with protein-homology modelling and inverted protein structure prediction to be solved.

Amino Acid Sequence↗

Atomic environment energies in proteins defined from statistics of accessible and contact surface areas.

Atomic contact potentials are derived by statistical analysis of atomic surface contact areas versus atom type in a database of non-homologous protein structures. The atomic environment is characterized by the surface area accessible to solvent and the surface of contacts with polar and non-polar atoms. Four types of atoms are considered, namely neutral polar atoms from protein backbones and from protein side-chains, non-polar atoms and charged atoms. Potential energies delta Ej(E) are defined from the preference for an atom of type j to be in a given environment E compared to the expected value if everything was random; Boltzmann's law is then used to transform these preferences into energies. These new potentials very clearly discriminate misfolded from correct structural models. The performance of these potentials are critically assessed by monitoring the recognition of the native fold among a large number of alternative structural folding types (the hide-and-seek procedure), as well as by testing if the native sequence can be recovered from a large number of randomly shuffled sequences for a given 3D fold (a procedure similar to the inverse folding problem). We suggest that these potentials reflect the atomic short range non-local interactions in proteins. To characterise atomic solvation alone, similar potentials were derived as a function of the percentage of solvent-accessible area alone. These energies were found to agree reasonably well with the solvation formalism of Eisenberg and McLachlan.

Animals↗

A self consistent mean field approach to simultaneous gap closure and side-chain positioning in homology modelling.

A new computational procedure which simultaneously provides gap closure and side-chain positioning in homology modelling is described. It uses a database search scheme to generate fragments to model gaps, a rotamer library to define side-chain conformations, and iteratively refines a conformational matrix CM, such that its elements CM(i,j,o) and CM(i,j,k) give the probabilities that the backbone of residue i adopts the conformation described by fragment j and that its side-chain adopts the conformation of its possible rotamer k. Each residue experiences the average of all possible environments, weighted by their respective probabilities. The method converges, thereby deserving the name of 'self consistent mean field' approach.

Algorithms↗

Solution structure of PMP-D2, a 35-residue peptide isolated from the insect Locusta migratoria.

The three-dimensional solution structure of PMP-D2, a 35 amino acid peptide isolated from the insect Locusta migratoria, has been determined from two-dimensional 1H NMR spectroscopy data. The structure calculations were performed from 222 NOE-derived interproton distances and 11 dihedral angles calculated from the JHN-H alpha coupling constants, using either a combination of distance geometry and restrained simulated annealing or by restrained simulated annealing alone. PMP-D2 contains three disulfide bridges that have been assigned from NMR data and structure calculations and independently confirmed using chemical and enzymatic methods. The core region of PMP-D2 adopts a compact globular fold, stabilized by hydrophobic interactions, which consists of a short three-stranded antiparallel beta-sheet involving residues 8-11, 15-19, and 25-29. Back-calculation of the NOESY spectra was used to validate the final structures. Analysis of the CD spectra of PMP-D2 under various conditions of ionic strength and in the presence of organic solvents demonstrates the high stability of this molecule. PMP-D2 was recently shown to inhibit Ca2+ currents. This activity is discussed based on the comparison of PMP-D2 three-dimensional structure with the recently established three-dimensional structure of the Ca2+ channel blocker omega-conotoxin GVIA.

Amino Acid Sequence↗

Application of a self-consistent mean field theory to predict protein side-chains conformation and estimate their conformational entropy.

Understanding the relations between the conformation of the side-chains and the backbone geometry is crucial for structure prediction as well as for homology modelling. To attempt to unravel these rules, we have developed a method which allows us to predict the position of the side-chains from the co-ordinates of the main-chain atoms. This method is based on a rotamer library and refines iteratively a conformational matrix of the side-chains of a protein, CM, such that its current element at each cycle CM (ij) gives the probability that side-chain i of the protein adopts the conformation of its possible rotamer j. Each residue feels the average of all possible environments, weighted by their respective probabilities. The method converges in only a few cycles, thereby deserving the name of self consistent mean field method. Using the rotamer with the highest probability in the optimized conformational matrix to define the conformation of the side-chain leads to the result that on average 72% of chi 1, 75% of chi 2 and 62% of chi 1 + 2 are correctly predicted for a set of 30 proteins. Tests with six pairs of homologous proteins have shown that the method is quite successful even when the protein backbone deviates from the correct conformation. The second application of the optimized conformational matrix was to provide estimates of the conformational entropy of the side-chains in the folded state of the protein. The relevance of this entropy is discussed.

Amino Acids↗