PubMed HealthSearch

Biomedical subjects

A Godzik

Publications and source records attributed to A Godzik.

18 recordsLinked to original sources

Structural diversity in a family of homologous proteins.

An interesting example of a structurally diverse group of sequentially homologous proteins is analyzed at the level of molecular interactions. In this family, the EF-hand calcium-binding proteins, there are examples of at least three distinct mutual positions of the N and C-terminal domains, despite significant sequence homology between all members of this family. Why does a particular protein choose one arrangement over another? To answer this question, detailed models of all proteins in their native structures as well as all alternative sequence/structure combinations are built by comparative modeling. By studying and comparing interactions stabilizing native structures and destabilizing alternative conformations, it is possible to gain insight into how such conformational diversity is achieved. It is shown that some mechanisms used to achieve it are: correlated mutations on the surface of two units and the presence of additional domains/chain fragments stabilizing desired topologies. The implications of these findings, both for structure predictions for other members of this family as well as the general problem of quaternary structure formation, are discussed.

Amino Acid Sequence

Are proteins ideal mixtures of amino acids? Analysis of energy parameter sets.

Various existing derivations of the effective potentials of mean force for the two-body interactions between amino acid side chains in proteins are reviewed and compared to each other. The differences between different parameter sets can be traced to the reference state used to define the zero of energy. Depending on the reference state, the transfer free energy or other pseudo-one-body contributions can be present to various extents in two-body parameter sets. It is, however, possible to compare various derivations directly by concentrating on the "excess" energy-a term that describes the difference between a real protein and an ideal solution of amino acids. Furthermore, the number of protein structures available for analysis allows one to check the consistency of the derivation and the errors by comparing parameters derived from various subsets of the whole database. It is shown that pair interaction preferences are very consistent throughout the database. Independently derived parameter sets have correlation coefficients on the order of 0.8, with the mean difference between equivalent entries of 0.1 kT. Also, the low-quality (low resolution, little or no refinement) structures show similar regularities. There are, however, large differences between interaction parameters derived on the basis of crystallographic structures and structures obtained by the NMR refinement. The origin of the latter difference is not yet understood.

Amino Acid Sequence

In search of the ideal protein sequence.

The inverse of a folding problem is to find the ideal sequence that folds into a particular protein structure. This problem has been addressed using the topology fingerprint-based threading algorithm, capable of calculating a score (energy) of an arbitrary sequence-structure pair. At first, the search is conducted by unconstrained minimization of the energy in sequence space. It is shown that using energy as the only design criterion leads to spurious solutions with incorrect amino acid composition. The problem lies in the general features of the protein energy surface as a function of both structure and sequence. The proposed solution is to design the sequence by maximizing the difference between its energy in the desired structure and in other known protein structures. Depending on the size of the database of structures 'to avoid', sequences bearing significant similarity to the native sequence of the target protein are obtained using this procedure.

Algorithms

Flexible algorithm for direct multiple alignment of protein structures and sequences.

The recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach.

Algorithms

De novo and inverse folding predictions of protein structure and dynamics.

In the last two years, the use of simplified models has facilitated major progress in the globular protein folding problem, viz., the prediction of the three-dimensional (3D) structure of a globular protein from its amino acid sequence. A number of groups have addressed the inverse folding problem where one examines the compatibility of a given sequence with a given (and already determined) structure. A comparison of extant inverse protein-folding algorithms is presented, and methodologies for identifying sequences likely to adopt identical folding topologies, even when they lack sequence homology, are described. Extension to produce structural templates or fingerprints from idealized structures is discussed, and for eight-membered beta-barrel proteins, it is shown that idealized fingerprints constructed from simple topology diagrams can correctly identify sequences having the appropriate topology. Furthermore, this inverse folding algorithm is generalized to predict elements of supersecondary structure including beta-hairpins, helical hairpins and alpha/beta/alpha fragments. Then, we describe a very high coordination number lattice model that can predict the 3D structure of a number of globular proteins de novo; i.e. using just the amino acid sequence. Applications to sequences designed by DeGrado and co-workers [Biophys. J., 61 (1992) A265] predict folding intermediates, native states and relative stabilities in accord with experiment. The methodology has also been applied to the four-helix bundle designed by Richardson and co-workers [Science, 249 (1990) 884] and a redesigned monomeric version of a naturally occurring four-helix dimer, rop. Based on comparison to the rop dimer, the simulations predict conformations with rms values of 3-4 A from native. Furthermore, the de novo algorithms can assess the stability of the folds predicted from the inverse algorithm, while the inverse folding algorithms can assess the quality of the de novo models. Thus, the synergism of the de novo and inverse folding algorithm approaches provides a set of complementary tools that will facilitate further progress on the protein-folding problem.

Algorithms

Regularities in interaction patterns of globular proteins.

The description of protein structure in the language of side chain contact maps is shown to offer many advantages over more traditional approaches. Because it focuses on side chain interactions, it aids in the discovery, study and classification of similarities between interactions defining particular protein folds and offers new insights into the rules of protein structure. For example, there is a small number of characteristic patterns of interactions between protein supersecondary structural fragments, which can be seen in various non-related proteins. Furthermore, the overlap of the side chain contact maps of two proteins provides a new measure of protein structure similarity. As shown in several examples, alignments based on contact map overlaps are a powerful alternative to other structure-based alignments.

Computer Simulation

Sequence-structure matching in globular proteins: application to supersecondary and tertiary structure determination.

A methodology designed to address the inverse globular protein-folding problem (the identification of which sequences are compatible with a given three-dimensional structure) is described. By using a library of protein finger-prints, defined by the side chain interaction pattern, it is possible to match each structure to its own sequence in an exhaustive data base search. It is shown that this is a permissive requirement for the validation of the methodology. To pass the more rigorous test of identifying proteins that are not close sequence homologs, but that have similar structure, the method has been extended to include insertions and deletions in the sequence, which is compared to the fingerprint. This allows for the identification of sequences having little or no sequence homology to the fingerprint. Examples include plastocyanin/azurin/pseudoazurin, the globin family, different families of proteases and cytochromes, including cytochromes c' and b-562, actinidin/papain, and lysozyme/alpha-lactalbumin. Turning to supersecondary structure prediction, we find that alpha/beta/alpha fragments possess sufficient specificity to identify their own and related sequences. By threading a beta-hairpin through a sequence, it is possible to predict the location of such hairpins and turns with remarkable fidelity. Thus, the method greatly extends existing techniques for the prediction of both global structural homology and local supersecondary structure.

Amino Acid Sequence

Topology fingerprint approach to the inverse protein folding problem.

We describe the most general solution to date of the problem of matching globular protein sequences to the appropriate three-dimensional structures. The screening template, against which sequences are tested, is provided by a protein "structural fingerprint" library based on the contact map and the buried/exposed pattern of residues. Then, a lattice Monte Carlo algorithm validates or dismisses the stability of the proposed fold. Examples of known structural similarities between proteins having weakly or unrelated sequences such as the globins and phycocyanins, the eight-member alpha/beta fold of triose phosphate isomerase and even a close structural equivalence between azurin and immunoglobulins are found.

Algorithms

Simulations of the folding pathway of triose phosphate isomerase-type alpha/beta barrel proteins.

Simulations of the folding pathways of two large alpha/beta proteins, the alpha subunit of tryptophan synthase and triose phosphate isomerase, are reported using the knight's walk lattice model of globular proteins and Monte Carlo dynamics. Starting from randomly generated unfolded states and with no assumptions regarding the nature of the folding intermediates, for the tryptophan synthase subunit these simulations predict, in agreement with experiment, the existence and location of a stable equilibrium intermediate comprised of six beta strands on the amino terminus of the molecule. For the case of triose phosphate isomerase, the simulations predict that both amino- and carboxyl-terminal intermediates should be observed. In a significant modification of previous lattice models, this model includes a full heavy atom side chain description and is capable of representing native conformations at the level of 2.5- to 3-A rms deviation for the C alpha positions, as compared to the crystal structure. With a well-balanced compromise between accuracy of the protein description and the computer requirements necessary to perform simulations spanning biologically significant amounts of time, the lattice model described here brings the possibility of studying important biological processes to present-day computers.

Animals

Conservation of residue interactions in a family of Ca-binding proteins.

In the TNC family of Ca-binding proteins (calmodulin, parvalbumin, intestinal calcium binding protein and troponin C) approximately 70 well-conserved amino acid sequences and six crystal structures are known. We find a clear correlation between residue contacts in the structures and residue conservation in the sequences: residues with strong sidechain-sidechain contacts in the three-dimenesional structure tend to be the more conserved in the sequence. This is one way to quantify the intuitive notion of the importance of sidechain interactions for maintaining protein three-dimensional structure in evolution and may usefully be taken into account in planning point mutations in protein engineering.

Amino Acid Sequence

On the interactions of charged side chains with the alpha-helix backbone.

The effects of the position of charged amino acid side chains on the stability of the alpha-helix are investigated. Calculations for the model polyAla 13 residue alpha-helix, with modifications based on experimental work, are performed at three levels of approximation. The observed stabilization of the alpha-helix could be explained by interactions between its macrodipole and charged amino acid side chains. Limitations of the model are discussed.

Amino Acids

Conformational role of His-12 in C-peptide of ribonuclease A.

Possible interactions of the His-12 ring with other side chain and backbone groups of C-peptide lactone (CPL) are discussed. The works published so far are critically reviewed and compared with the latest results obtained by the authors. The main new conclusion is that in the helical conformation of CPL, the Phe-8 and His-12 rings are clustered together. Studies of Phe-8----Ala analogs of CPL and calculations of ring current effects satisfactorily explain the observed environmental shifts of Phe-8 and His-12 protons in NMR spectra of CPL. Interaction between both rings is favorable for alpha-helix formation, but cannot explain an increase in helix stability related with protonation of His-12. This effect arises from favorable interactions of the charged His+-12 ring with the helix backbone.

Histidine

The Monte Carlo simulation of pearl chain formation.

The phenomenon of pearl chain formation (PCF) is investigated by means of a statistical model using the Monte Carlo method. Fifteen particles (cells) interacting with simple dipole-dipole potential are shown to form chains under the influence of an external field with a threshold potential significantly lower than the two particle estimate. A possible overlap between PCF and the thermal effects of an electric field is suggested.

Algorithms

Calculations of the conformational properties of acyclonucleosides. Part I. Stable conformations of acyclovir.

Conformational properties of te antiherpes agent acyclovir (acycloguanosine, ACG) were calculated using molecular mechanics approximation. Eighty two different stable conformations have been determined. The large number of local minima of the total enery, and small differences between them, point to the marked flexibility of the acyclic chain. The barrier to rotation around N9-C1 bond was calculated and found to be asymmetric (the lower equals 17 kJ/mol, the higher 63 kJ/mol). An energetic preference for the compact from of ACG was demonstrated. A comparison of the calculated conformations with the crystallographic structures is presented.

Acyclovir