PubMed HealthSearch

Biomedical subjects

S J Wodak

Publications and source records attributed to S J Wodak.

At least 19 recordsLinked to original sources

Extracting information on folding from the amino acid sequence: accurate predictions for protein regions with preferred conformation in the absence of tertiary interactions.

A recently developed procedure to predict backbone structure from the amino acid sequence [Rooman, M., Kocher, J. P., & Wodak, S. (1991) J. Mol. Biol, 221, 961-979] is fine tuned to identify protein segments, of length 5-15 residues, that adopt well-defined conformations in the absence of tertiary interactions. These segments are obtained by requiring that their predicted lowest energy structures have a sizable energy gap relative to other computed conformations. Applying this procedure to 69 proteins of known structure, we find that regions with largest energy gaps--those having highly preferred conformations--are also the most accurately predicted ones. On the basis of previous findings that such regions correlate well with sites that become structured early during folding, our approach provides the means of identifying such sites in proteins without prior knowledge of the tertiary structure. Furthermore, when predictions are performed so as to ignore the influence of residues flanking each segment along the sequence, a situation akin to excising the considered peptide from the rest of the chain, they offer the possibility of identifying protein segments liable to adopt well-defined conformations on their own. The described approach should have useful applications in experimental and theoretical investigations of protein folding and stability, and aid in designing peptide drugs and vaccines.

Amino Acid Sequence

Extracting information on folding from the amino acid sequence: consensus regions with preferred conformation in homologous proteins.

It is investigated whether protein segments predicted to have a well-defined conformational preference in the absence of tertiary interactions are conserved in families of homologous proteins. The prediction method follows the procedures of Rooman, M., Kocher, J.-P., and Wodak, S. (preceding paper in this issue). It uses a knowledge-based force field that incorporates only local interactions along the sequence and identifies segments whose lowest energy structure displays a sizable energy gap relative to other computed conformations. In 13 of the protein families and subfamilies considered that are sufficiently homologous to have similar 3D structures, at least one region is consistently predicted as having the same preferred conformation in virtually all family members. These regions are between 4 and 26 residues long. They are often located at chain ends and correspond primarily to segments of secondary structure heavily involved in interactions with the rest of the protein, suggesting that they could act as nuclei around which other parts of the structure would assemble. Experimental data on early folding intermediates or on protein fragments with appreciable structure in aqueous solution are available for more than half of the protein families. Comparison of our results with these data is quite favorable. They reveal that each of the experimentally identified early formed, or independently stable, substructures harbors at least one of the segments consistently predicted as having a preferred conformation by our procedure. The implications of our findings for the conservation of folding pathways in homologous proteins are discussed.

Adenylate Kinase

Protein engineering of xylose (glucose) isomerase from Actinoplanes missouriensis. 1. Crystallography and site-directed mutagenesis of metal binding sites.

The structure and function of the xylose (glucose) isomerase from Actinoplanes missouriensis have been analyzed by X-ray crystallography and site-directed mutagenesis after cloning and overexpression in Escherichia coli. The crystal structure of wild-type enzyme has been refined to an R factor of 15.2% against diffraction data to 2.2-A resolution. The structures of a number of binary and ternary complexes involving wild-type and mutant enzymes, the divalent cations Mg2+, Co2+, or Mn2+, and either the substrate xylose or substrate analogs have also been determined and refined to comparable R factors. Two metal sites are identified. Metal site 1 is four-coordinated and tetrahedral in the absence of substrate and is six-coordinated and octahedral in its presence; the O2 and O4 atoms of linear inhibitors and substrate bind to metal 1. Metal site 2 is octahedral in all cases; its position changes by 0.7 A when it binds O1 of the substrate and by more than 1 A when it also binds O2; these bonds replace bonds to carboxylate ligands from the protein. Side chains involved in metal binding have been substituted by site-directed mutagenesis. The biochemical properties of the mutant enzymes are presented. Together with structural data, they demonstrate that the two metal ions play an essential part in binding substrates, in stabilizing their open form, and in catalyzing hydride transfer between the C1 and C2 positions.

Actinomycetales

Protein engineering of xylose (glucose) isomerase from Actinoplanes missouriensis. 2. Site-directed mutagenesis of the xylose binding site.

Site-directed mutagenesis in the active site of xylose isomerase derived from Actinoplanes missouriensis is used to investigate the structural and functional role of specific residues. The mutagenesis work together with the crystallographic studies presented in detail in two accompanying papers adds significantly to the understanding of the catalytic mechanism of this enzyme. Changes caused by introduced mutations emphasize the correlation between substrate specificity and cation preference. Mutations in both His 220 and His 54 mainly affect the catalytic rate constant, with catalysis being severely reduced but not abolished, suggesting that both histidines are important, but not essential, for catalysis. Our results thus challenge the hypothesis that His 54 acts as an obligatory catalytic base for ring opening; this residue appears instead to be implicated in governing the anomeric specificity. With none of the active site histidines acting as a catalytic base, the role of the cations in catalyzing proton transfer is confirmed. In addition, Lys 183 appears to play a crucial part in the isomerization step, by assisting the proton shuttle. Other residues also are important but to a lesser extent. The conserved Lys 294 is indirectly involved in binding the activating cations. Among the active site aromatic residues, the tryptophans (16 and 137) play a role in maintaining the general architecture of the substrate binding site while the role of Phe 26 seems to be purely structural.

Actinomycetales

Protein engineering of xylose (glucose) isomerase from Actinoplanes missouriensis. 3. Changing metal specificity and the pH profile by site-directed mutagenesis.

Aldose-ketose isomerization by xylose isomerase requires bivalent cations such as Mg2+, Mn2+, or Co2+. The active site of the enzyme from Actinoplanes missouriensis contains two metal ions that are involved in substrate binding and in catalyzing a hydride shift between the C1 and C2 substrate atoms. Glu 186 is a conserved residue located near the active site but not in contact with the substrate and not with a metal ligand. The E186D and E186Q mutant enzymes were prepared. Both are active, and their metal specificity is different from that of the wild type. The E186Q enzyme is most active with Mn2+ and has a drastically shifted pH optimum. The X-ray analysis of E186Q was performed in the presence of xylose and either Mn2+ or Mg2+. The Mn2+ structure is essentially identical to that of the wild type. In the presence of Mg2+, the carboxylate group of residue Asp 255, which is part of metal site 2 and a metal ligand, turns toward Gln 186 and hydrogen bonds to its side-chain amide. Mg2+ is not bound at metal site 2, explaining the low activity of the mutant with this cation. Movements of Asp 255 also occur in the wild-type enzyme. We propose that they play a role in the O1 to O2 proton relay accompanying the hydride shift.

Actinomycetales

Recurrent alpha beta loop structures in TIM barrel motifs show a distinct pattern of conserved structural features.

A systematic survey of seven parallel alpha/beta barrel protein domains, based on exhaustive structural comparisons, reveals that a sizable proportion of the alpha beta loops in these proteins--20 out of a total of 49--belong to either one of two loop types previously described by Thornton and co-workers. Six loops are of the alpha beta 1 type, with one residue between the alpha-helix and beta-strand, and 13 are of the alpha beta 3 type, with three residues between the helix and the strand. Protein fragments embedding the identified loops, and termed alpha beta connections since they contain parts of the flanking helix and strand, have been analyzed in detail revealing that each type of connection has a distinct set of conserved structural features. The orientation of the beta-strand relative to the helix and loop portions is different owing to a very localized difference in backbone conformation. In alpha beta 1 connections, the chain enters the beta-strand via a residue adopting an extended conformation, while in alpha beta 3 it does so via a residue in a near alpha-helical conformation. Other conserved structural features include distinct patterns of side chain orientation relative to the beta-sheet surface and of main chain H-bonds in the loop and the beta-strand moieties. Significant differences also occur in packing interactions of conserved hydrophobic residues situated in the last turn of the helix. Yet the alpha-helix surface of both types of connections adopts similar orientations relative to the barrel sheet surface. Our results suggest furthermore that conserved hydrophobic residues along the sequence of the connections, may be correlated more with specific patterns of interactions made with neighboring helices and sheet strands than with helix/strand packing within the connection itself. A number of intriguing observations are also made on the distribution of the identified alpha beta 1 and alpha beta 3 loops within the alpha/beta-barrel motifs. They often occur adjacent to each other; alpha beta 3 loops invariably involve even numbered beta-strands, while alpha beta 1 loops involve preferentially odd beta-strands; all the analyzed proteins contain at least one alpha beta 3 loop in the first half of the eightfold alpha/beta barrel. Possible origins of all these observations, and their relevance to the stability and folding of parallel alpha/beta barrel motifs are discussed.

Alcohol Oxidoreductases

Contribution of the hydrophobic effect to protein stability: analysis based on simulations of the Ile-96----Ala mutation in barnase.

Molecular dynamics simulations have been used to compute the difference in the unfolding free energy between wild-type barnase and the mutant in which Ile-96 is replaced by alanine. The simulations yield results (-3.42 and -5.21 kcal/mol) that compare favorably with experimental values (-3.3 and -4.0 kcal/mol). The major contributions to the free energy difference arise from bonding terms involving degrees of freedom of the mutated side chain and from nonbonded interactions of that side chain with its environment in the folded protein. By comparison with simulations of an extended peptide in the absence of solvent, used as a reference state, hydration effects are shown to play a minor role in the overall free energy balance for the Ile----Ala transformation. The implications of these results for our understanding of the hydrophobic effect and its contribution to protein stability are discussed.

Alanine

Prediction of protein backbone conformation based on seven structure assignments. Influence of local interactions.

A method is developed to compute backbone tertiary folds from the amino acid sequence. In this method, the number of degrees of freedom is drastically reduced by neglecting side-chain flexibility, and by describing backbone conformations as combinations of only seven structural states. These are characterized by single values of the dihedral angles phi, psi and omega, representing allowed conformations of the isolated dipeptide. We show that this restrictive model is none the less capable of describing native backbones to within acceptable deviations. Using our backbone description, potentials of mean force are derived from a database of known protein structures, based on statistical influences of single residues and residue pairs on the conformational states in their vicinity along the chain. This yields the force-field component due to local interactions, which is then used to predict lowest-energy conformations from any given amino acid sequence. The prediction algorithm does not require searching conformational space and is therefore extremely fast. Another important asset of our method is that it is able to compute not only the minimum energy conformation, but any number of lowest energy structures, whose relative preferences can be determined from the corresponding computed energy values. The performance of our procedure is tested on short peptides that are likely to be stabilized by local interactions. These include several helical structures and a hexapeptide with a beta-bend conformation, corresponding to peptides shown to have relatively well-defined conformations in aqueous solution, and to protein segments believed to adopt their native conformation early during folding. In addition, several flexible peptides are analysed. Except for the problems encountered in predicting observed disulphide bridges in two of the flexible peptides, and in a somewhat larger fragment comprising residues 30 to 51 of bovine trypsin inhibitor, prediction results compare very favourably with experimental data. Potential applications of our procedure to protein modelling and its extension to protein folding are discussed.

Amino Acid Sequence

Weak correlation between predictive power of individual sequence patterns and overall prediction accuracy in proteins.

Patterns in amino acid properties (polar, hydrophobic, etc.) that characterize secondary structure motifs are derived from a database containing 75 protein structures, with the aim of circumventing the limitations due to data base size so as to increase structure prediction score. Many such sequence-structure associations with high intrinsic predictive power are found, which turn out to be correct 78% of the time when applied individually to proteins outside the learning set. Based on these associations, a prediction method is developed, which reaches the score of 62% on the 3 states alpha-helix, beta-strand, and loop, without using additional constraints. Though this score is quite good compared to that of other available prediction methods, it is much lower than could be expected from the high intrinsic predictive power of the associations used. The reasons underlying this surprising result, which indicate that prediction score and intrinsic predictive power are only weakly coupled, are discussed. It is also shown that the size of the present database still seriously limits prediction scores, even when property patterns are used, and that higher scores are expected in large databases. Clues are provided on the relative influence of neglecting spatial interactions on prediction efficiency, suggesting that, in sufficiently large databases, predicted secondary structures would correspond to those formed early in the folding process. This hypothesis is tested by confronting present predictions with available experimental data on early protein folding intermediates and on small peptides that adopt a relatively stable conformation in water. Although admittedly there are still too few such data, results suggest that the hypothesis might be well founded.

Amino Acid Sequence

SESAM: a relational database for structure and sequence of macromolecules.

A system is described that provides ways of integrating data on protein structure, sequence, and survey results, with molecular graphics and molecular mechanics software. Its major component is the relational database SESAM, presently implemented under the commercial package SYBASE. By design, the database allows full integration--within the same data organization--of raw data on protein structure, sequence, ligands, and heterogroups, obtained from the Brookhaven Protein Databank, with pure sequence information available from other databanks such as SWISS-PROT. It contains in addition higher level descriptions of structural and topological properties, as well as survey results, obtained by executing specialized computer programs. Aside from the very useful attribute of closely combining structural and nonstructural information, other important features distinguish it from analogous systems developed elsewhere. It includes a molecular dictionary with complete description of geometric properties and energy parameters used in modeling and conformational energy calculations. Using this dictionary, structural data are validated by checking for localized inconsistencies in atomic coordinates, atomic symbols, chirality definitions, and flagging errors and incomplete entries. Because of both the dictionary and the validation procedures, SESAM can be readily interfaced with conventional molecular graphics and mechanics software packages, or with other specialized application programs. With the aid of appropriate interfaces, data access is sufficiently fast for SESAM to be interrogated interactively. Prototypes of user interfaces, as well as an interface with the molecular graphics package BRUGEL, are described and the power of the system is illustrated in applications such as homology-based protein modeling, computer-aided protein design, protein structure predictions, analysis of local structure motifs, and of relationships between protein sequence and structure.

Amino Acid Sequence

Automatic definition of recurrent local structure motifs in proteins.

An automatic procedure for defining recurrent folding motifs in proteins of known structure is described. These motifs are formed by short polypeptide fragments of equal size containing between four and seven residues. The method applies a classical clustering algorithm that operates on distances between selected backbone atoms. In one application, we use it to cluster all protein fragments into only four structural classes. This classification is rough considering the observed diversity of local structures, but comparable in homogeneity to the four classes of secondary structure (alpha-helix, beta-strand, turn and coil). Yet, it discriminates between extended and curved coil and distinguishes beta-bulges from beta-strands. In a second application, the clustering procedure is combined with assignment of backbone dihedral angles to allowed regions in the Ramachandran map. This produces an exhaustive repertoire of highly homogeneous families of structural motifs that contains all the beta-hairpins, beta alpha- and alpha beta-loops previously defined by manual procedures, and new structural families of which two examples, a beta alpha-loop and an alpha-helix beginning, are analyzed in detail. The described automatic procedures should be useful in categorizing structure information in proteins, thereby increasing our ability to analyze relations between structure and sequence.

Amino Acid Sequence

Relations between protein sequence and structure and their significance.

The relation between amino acid sequence and local structure in proteins is investigated. The local structures considered are either the four classes of secondary structure (H, E, T and C) or four classes of local conformations defined using measures of conformational similarity based on distances between C alpha atoms. The classes are obtained by applying an automatic clustering procedure to short polypeptide fragments of uniform length from a database of 75 known protein structures. The thrust of our investigation consists of systematically searching the database for simple amino acid patterns of the type Gly-X-Ala-X-X-Val, where X denotes an arbitrary residue. Patterns that are nearly always associated with the same structure are retained. Finding many such associations, we then evaluate by a statistical approach how many among them are non-random and compare the results for different definitions of local structure. A similar comparison is made for the predictive value of retained associations, which is assessed using an internal test based on dividing the database into "learning" and "test" subsets. While we find that local structures defined by conformational similarity are not superior to secondary structure for prediction purposes, they help us gain insight into the factors that influence the predictive value of derived associations. A major conclusion is that the number of retained associations is in large excess over the number expected from a random correlation between sequence and structure, irrespective of how local conformation is defined. However, only a very small number of these associations can be earmarked as reliable using statistical criteria, due to the limited size of the database. We find, for instance, that the pattern Ala-Ala-X-X-Lys reliably characterizes helix, and the pattern Val-X-Val-X-X-X-Ala reliably characterizes extended structure and beta-strand. The possibility is discussed that these and other reliable associations correspond to regions of the polypeptide chain whose conformations are locally determined and that these regions may play a role in folding.

Amino Acid Sequence

The design of idealized alpha/beta-barrels: analysis of beta-sheet closure requirements.

The 8-fold parallel alpha/beta-barrel topology is encountered in proteins that display an impressive variety of functions, suggesting that this topology may be a rather nonspecific and stable folding motif. Consequently, this motif can be considered as an interesting framework to design novel proteins. It has been shown that the shape of the beta-sheet portion of the barrel can be approximated by a hyperboloid. This geometric object may therefore be used as a scaffold to construct an idealized eight-stranded beta-barrel. To facilitate the de novo design of such structures, a collection of modeling tools has been developed allowing secondary structure elements to be mapped onto the scaffold surface and rotation and translation operations to be performed about user defined axes while evaluating their contribution to the conformational energy of the system. These tools have been applied in a systematic study assessing the phi, psi requirements to design symmetric eight stranded beta barrels with optimal hydrogen bonding between adjacent beta-strands. It is observed that: (a) the beta-sheet structure can be closed without introducing irregular stagger between beta-strands and (b) the region of phi, psi dihedral angle space compatible with the formation of regular symmetric eight stranded beta-barrels coincides with the phi, psi region corresponding to average beta-strands in known protein structures, suggesting that barrel closure does not impose gross constraints on beta-strand geometry.

Computer Simulation

Basic design features of the parallel alpha beta barrel, a ubiquitous protein-folding motif.

Basic design features of the beta-sheet portion in parallel alpha beta barrels in known protein structures are analysed in the context of a model of a regular hyperboloid. A formal description of the relationships between beta-sheet twist, number of strands in the sheet and barrel dimensions is derived, and the underlying physical principles are rationalized. Results suggest that the major constraints on the geometry of the beta-sheet portion of the barrel come from the requirements to have optimal H-bonding interactions between beta-strands and to closely pack amino acid side-chains in the barrel interior so as to exclude bulk water. In addition, we show how the hyperboloid model and the ensuing formalism can serve to derive useful geometric and graphic tools for computer-aided protein design de novo. We then illustrate how these tools are used to determine that the requirement to have a closed regular eight-stranded beta-sheet surface imposes no particular constraints on the geometry (phi, psi angles) of the polypeptide backbone. Understanding the role of the amino acid sequence in determining the observed structures remains a major challenge. Detailed comparisons of known alpha beta-barrel structures (and amino acid sequence) with each other, and with polypeptide fragments from other protein crystal structures, reveal only a limited number of common sequence-structure motifs. These belong to characteristic alpha beta 1 and alpha beta 3 loop families previously described in alpha beta proteins, and occur at least once in nearly all the alpha beta-barrel structures examined.

Amino Acid Sequence

Amino acid sequence templates derived from recurrent turn motifs in proteins: critical evaluation of their predictive power.

Amino acid sequence patterns suggested to characterize specific recurrent turn conformation in protein are tested as to their predictive power in a database containing 75 proteins of known structure. Many of these patterns are found to be associated with local structures that differ from the motifs originally used to derive them. It is therefore concluded that, while they could be useful for improving predictions made by other methods, their stand-alone predictive power is poor. The issue of deriving and validating consensus sequence patterns for use in protein structure prediction is raised.

Amino Acid Sequence

Identification of predictive sequence motifs limited by protein structure data base size.

Associations between short amino acid sequence patterns and protein secondary structure classes can be found by searching a data base of known protein structures. Analysis of these associations suggests that secondary structure of proteins can be determined locally by sequence motifs of high predictive value, but at present our ability to find these motifs is limited by the size of the available data bases.

Amino Acid Sequence

Structural principles of parallel beta-barrels in proteins.

Eight-stranded beta-sheets in nine protein structures containing "TIM (triose phosphate isomerase) barrels" are shown to be fitted satisfactorily by hyperboloids, the generating lines of which pass through the beta-strands. Simple parameterizations of the hyperboloid model are then used to determine the constraints that govern key parameters, such as the number of strands in the barrel, and to rationalize the remarkable conservation of strand number, observed to be eight, in nearly all the known examples of parallel beta-barrels. It is shown that the requirement to exclude solvent from the barrel interior, while at the same time keeping an upper limit on strand twist and interstrand distance so as to foster extensive hydrogen bonding interactions within the sheet, imposes strong constraints on barrel geometry. A formal description of the relationships between beta-sheet twist, strand number, and barrel dimensions is given here. It could have important implications for studies of protein folding and design.

Computer Simulation

Calculations of electrostatic properties in proteins. Analysis of contributions from induced protein dipoles.

The calculation of induced dipole moments and of their contribution to electrostatic effects in proteins is implemented following the approach of Warshel. Isotropic polarizabilities are assigned to individual atoms, and the resulting deviation from pairwise interactions is treated by a self-consistent iterative procedure. We give a detailed description of how the formalism is implemented in molecular mechanics and molecular dynamics simulation procedures, and report results based on calculations performed on crystal structures of crambin, liver alcohol dehydrogenase and ribonuclease T1. We focus our analysis on evaluating the contribution of polarizability of the protein matrix to electrostatic energies, local fields, to dipole moments of peptide groups and of secondary structure elements in the polypeptide chain. Our calculations confirm that induced dipole moments in proteins provide important stabilizing contributions to electrostatic energies, and that these contributions cannot be mimicked by the usual approximations where either a continuum dielectric constant, or a distance-dependent dielectric function is used. We find that induced protein dipoles appreciably affect the magnitude and direction of local electrostatic fields in a manner that is strongly influenced by the microscopic environment in the protein. Most strongly affected are fields in charged groups that are involved in close interactions with other charged groups, while the influence on local fields of aliphatic groups is marginal. We find, moreover, that induction effects from surrounding protein atoms tend on average to increase peptide dipoles and helix macro-dipoles by about 16%, again reflecting electrostatic stabilization by the protein matrix, and show that (at least in the alpha/beta domain of alcohol dehydrogenase) the contribution of side-chains to this stabilization is significant.

Alcohol Dehydrogenase