PubMed HealthSearch

Biomedical subjects

M Gerstein

Publications and source records attributed to M Gerstein.

At least 19 recordsLinked to original sources

A structural census of the current population of protein sequences.

We examine the occurrence of the approximately 300 known protein folds in different groups of organisms. To do this, we characterize a large fraction of the currently known protein sequences ( approximately 140,000) in structural terms, by matching them to known structures via sequence comparison (or by secondary-structure class prediction for those without structural homologues). Overall, we find that an appreciable fraction of the known folds are present in each of the major groups of organisms (e.g., bacteria and eukaryotes share 156 of 275 folds), and most of the common folds are associated with many families of nonhomologous sequences (i.e., >10 sequence families for each common fold). However, different groups of organisms have characteristically distinct distributions of folds. So, for instance, some of the most common folds in vertebrates, such as globins or zinc fingers, are rare or absent in bacteria. Many of these differences in fold usage are biologically reasonable, such as the folds of metabolic enzymes being common in bacteria and those associated with extracellular transport and communication being common in animals. They also have important implications for database-based methods for fold recognition, suggesting that an unknown sequence from a plant is more likely to have a certain fold (e.g., a TIM barrel) than an unknown sequence from an animal.

Amino Acid Sequence

LPFC: an Internet library of protein family core structures.

As the number of protein molecules with known, high-resolution structures increases, it becomes necessary to organize these structures for rapid retrieval, comparison, and analysis. The Protein Data Bank (PDB) currently contains nearly 5,000 entries and is growing exponentially. Most new structures are similar structurally to ones reported previously and can be grouped into families. As the number of members in each family increases, it becomes possible to summarize, statistically, the commonalities and differences within each family. We reported previously a method for finding the atoms in a family alignment that have low spatial variance and those that have higher spatial variance (i.e., the "core" atoms that have the same relative position in all family members and the "non-core" atoms that do not). The core structures we compute have biological significance and provide an excellent quantitative and visual summary of a multiple structural alignment. In order to extend their utility, we have constructed a library of protein family cores, accessible over the World Wide Web at http:/ /www-smi.stanford.edu/projects/helix/LPFC/. This library is generated automatically with publicly available computer programs requiring only a set of multiple alignments as input. It contains quantitative analysis of the spatial variation of atoms within each protein family, the coordinates of the average core structures derived from the families, and display files (in bitmap and VRML formats). Here, we describe the resource and illustrate its applicability by comparing three multiple alignments of the globin family. These three alignments are found to be similar, but with some significant differences related to the diversity of family members and the specific method used for alignment.

Computer Communication Networks

Protein folding: the endgame.

The last stage of protein folding, the "endgame," involves the ordering of amino acid side-chains into a well defined and closely packed configuration. We review a number of topics related to this process. We first describe how the observed packing in protein crystal structures is measured. Such measurements show that the protein interior is packed exceptionally tightly, more so than the protein surface or surrounding solvent and even more efficiently than crystals of simple organic molecules. In vitro protein folding experiments also show that the protein is close-packed in solution and that the tight packing and intercalation of side-chains is a final and essential step in the folding pathway. These experimental observations, in turn, suggest that a folded protein structure can be described as a kind of three-dimensional jigsaw puzzle and that predicting side-chain packing is possible in the sense of solving this puzzle. The major difficulty that must be overcome in predicting side-chain packing is a combinatorial "explosion" in the number of possible configurations. There has been much recent progress towards overcoming this problem, and we survey a variety of the approaches. These approaches differ principally in whether they use ab initio (physical) or more knowledge-based methods, how they divide up and search conformational space, and how they evaluate candidate configurations (using scoring functions). The accuracy of side-chain prediction depends crucially on the (assumed) positioning of the main-chain. Methods for predicting main-chain conformation are, in a sense, not as developed as that for side-chains. We conclude by surveying these methods. As with side-chain prediction, there are a great variety of approaches, which differ in how they divide up and search space and in how they score candidate conformations.

Animals

Packing at the protein-water interface.

We have determined the packing efficiency at the protein-water interface by calculating the volumes of atoms on the protein surface and nearby water molecules in 22 crystal structures. We find that an atom on the protein surface occupies, on average, a volume approximately 7% larger than an atom of equivalent chemical type in the protein core. In these calculations, larger volumes result from voids between atoms and thus imply a looser or less efficient packing. We further find that the volumes of individual atoms are not related to their chemical type but rather to their structural location. More exposed atoms have larger volumes. Moreover, the packing around atoms in locally concave, grooved regions of protein surfaces is looser than that around atoms in locally convex, ridge regions. This as a direct manifestation of surface curvature-dependent hydration. The net volume increase for atoms on the protein surface is compensated by volume decreases in water molecules near the surface. These waters occupy volumes smaller than those in the bulk solvent by up to 20%; the precise amount of this decrease is directly related to the extent of contact with the protein.

Animals

Using iterative dynamic programming to obtain accurate pairwise and multiple alignments of protein structures.

We show how a basic pairwise alignment procedure can be improved to more accurately align conserved structural regions, by using variable, position-dependent gap penalties that depend on secondary structure and by taking the consensus of a number of suboptimal alignments. These improvements, which are novel for structural alignment, are direct analogs of what is possible with normal sequences alignment. They are feasible for us since our basic structural alignment procedure, unlike others, is so similar to normal sequence alignment. We further present preliminary results that show how our procedure can be generalized to produce a multiple alignment of a family of structures. Our approach is based on finding a "median" structure from doing all possible pairwise alignments and then aligning everything to it.

Amino Acid Sequence

Average core structures and variability measures for protein families: application to the immunoglobulins.

A variety of methods are currently available for creating multiple alignments, and these can be used to define and characterize families of related proteins, such as the globins or the immunoglobulins. We have developed a method for using a multiple alignment to identify an average structural "core", a subset of atoms with low structural variation. We show how the means and variances of core-atom positions summarize the commonalities and differences with a family, making them particularly useful in compiling libraries of protein folds. We show further how it is possible to describe the rotation and translation relating two core structures, as in two domains of a multi-domain protein, in a consistent fashion in terms of a "mean" transformation and a deviation about this mean. Once determined, our average core structures (with their implicit measure of structural variation) allow us to define a measure of structural similarity more informative than the usual root-mean-square (RMS) deviation in atomic position, i.e. a "better RMS." Our average structures also permit straightforward comparisons between variation in structure and sequence at each position in a family. We have applied our core-finding methodology in detail to the immunoglobulin family. We find that the structural variability we observe just within the VL and VH domains anticipates the variability that others have observed throughout the whole immunoglobulin superfamily; that a core definition based on sequence conservation, somewhat surprisingly, does not agree with one based on structural similarity; and that the cores of the VL and VH domains vary about 5 degrees in relative orientation across the known structures.

Algorithms

The volume of atoms on the protein surface: calculated from simulation, using Voronoi polyhedra.

We analyze the volume of atoms on the protein surface during a molecular-dynamics simulation of a small protein (pancreatic trypsin inhibitor). To calculate volumes, we use a particular geometric construction, called Voronoi polyhedra, that divides the total volume of the simulation box amongst the atoms, rendering them relatively larger or smaller depending on how tightly they are packed. We find that most of the atoms on the protein surface are larger than those buried in the core (by approximately 6%), except for the charged atoms, which decrease in size, presumably due to electroconstriction. We also find that water molecules are larger near apolar atoms on the protein surface and smaller near charged atoms, in comparison to "bulk" water molecules far from the protein. Taken together, these findings necessarily imply that apolar atoms on the protein surface and their associated water molecules are less tightly packed (than corresponding atoms in the protein core and bulk water) and the opposite is the case for charged atoms. This looser apolar packing and tighter charged packing fundamentally reflects protein-water distances that are larger or smaller than those expected from van der Waals radii. In addition to the calculation of mean volumes, simulations allow us to investigate the volume fluctuations and hence compressibilities of the protein and solvent atoms. The relatively large volume fluctuations of atoms at the protein-water interface indicates that they have a more variable packing than corresponding atoms in the protein core or in bulk water. We try to adhere to traditional conventions throughout our calculations. Nevertheless, we are aware of and discuss three complexities that significantly qualify our calculations: the positioning of the dividing plane between atoms, the problem of vertex error, and the choice of atom radii. In particular, our results highlight how poor a "compromise" the commonly accepted value of 1.4 A is for the radius of a water molecule.

Animals

Binding geometry of alpha-helices that recognize DNA.

Many transcription factors have an alpha-helix that binds to DNA bases in a specific fashion. The DNA-binding geometry of these recognition helices varies substantially. We define a set of parameters to describe the binding geometry of recognition helices and analyze specific stereochemical elements that determine particular geometries. Because the convex surface of the helix must fit into the concave surface of the DNA major groove, the number of degrees of freedom of the recognition helix is reduced from a possible six to a single angle, which we call alpha. The chemically interacting DNA bases and amino acid residues must lie along a common line and have the same spacing along it. This pairing of base positions with residue positions seems to restrict the binding geometry further to a set of discrete values for alpha.

Amino Acid Sequence

Using a measure of structural variation to define a core for the globins.

As the database of three-dimensional protein structures expands, it becomes possible to classify related structures into families. Some of these families, such as the globins, have enough members to allow statistical analysis of conserved features. Previously, we have shown that a probabilistic representation based on means and variances can be useful for defining structural cores for large families. These cores contain the subset of atoms that are in essentially the same relative positions in all members of the family. In addition to defining a core, our method creates an ordered list of atoms, ranked by their structural variation. In applying our core-finding procedure to the globins, we find that helices A, B, G and H form a structural core with low variance. These helices fold early in the folding pathway, and superimpose well with helices in the helix-turn-helix repressor protein family. The non-core helices (F and the parts of other helices that interact with it) are associated with the functional differences among the globins, and are encoded within a separate exon. We have also compared the variability measure implicit in our core structures with measures of sequence variability, using a procedure for measuring sequence variability that helps correct for the biased sampling in the databanks. We find, somewhat surprisingly, that sequence variation does not appear to correlate with structural variation.

Algorithms

DNA recognition and superstructure formation by helix-turn-helix proteins.

The way helix-turn-helix proteins recognize DNA is analysed by comparing their sequences, structures, and binding specificities. Individual recognition helices in these proteins bind to four DNA base pairs with the same geometry. However, pairs of recognition helices in the protein dimers can have different separations and orientations. These differences are used for discriminating between DNAs which have different superstructures, in particular, different numbers of base pairs between sets of the four base pairs.

Amino Acid Sequence

Stereochemical basis of DNA recognition by Zn fingers.

DNA-recognition rules for Zn fingers are discussed in terms of crystal structures. The rules can explain the DNA-binding characteristics of a number of Zn finger proteins for which there are no crystal structures. The rules have two parts: chemical rules, which list the possible pairings between the 4 DNA bases and the 20 amino acid residues, and stereochemical rules, which describe the specific base positions contacted by several amino acid positions in the Zn finger. It is discussed that to maintain the correct binding geometry, in which the N-terminus of the recognition helix is closer to the DNA than the C-terminus, the residues facing the DNA on the helix must be larger near the C-terminus, and that two different types of fingers (A and B) bind to DNA in distinctly different ways and cover different numbers of base pairs.

Amino Acid Sequence

Volume changes on protein folding.

BACKGROUND: Protein volumes change very little on folding at low pressure, but at high pressure the unfolded state is more compact. So far, the molecular origins of this behaviour have not been explained: it is the opposite of that expected from the model of the hydrophobic effect based on the transfer of non-polar solutes from water to organic solvent. RESULTS: We redetermined the mean volumes occupied by residues in the interior of proteins. The new residue volumes are smaller than those given by previous calculations which were based on much more limited data. They show that the packing density in protein interiors is exceptionally high. Comparison of the volumes that residues occupy in proteins with those they occupy in solution shows that aliphatic groups have smaller volumes in protein interiors than in solution, while peptide and charged groups have larger volumes. The cancellation of these volume changes is the reason that the net change on folding is very small. CONCLUSIONS: The exceptionally high density of the protein interior shown here implies that packing forces play a more important role in protein stability than has been believed hitherto.

Amino Acid Sequence

Structural mechanisms for domain movements in proteins.

We survey all the known instances of domain movements in proteins for which there is crystallographic evidence for the movement. We explain these domain movements in terms of the repertoire of low-energy conformation changes that are known to occur in proteins. We first describe the basic elements of this repertoire, hinge and shear motions, and then show how the elements of the repertoire can be combined to produce domain movements. We emphasize that the elements used in particular proteins are determined mainly by the structure of the interfaces between the domains.

Motion

Volume changes in protein evolution.

We have determined the variations in volume that occur during evolution in the buried core of three different families of proteins. The variation of the whole core is very small (approximately 2.5%) compared to the variation at individual sites (approximately 13%). However, by comparing our results to those expected from random sequences with no correlations between sites, we show that the small variation observed may simply be a manifestation of the statistical "law of large numbers" and not reflect any compensating changes in, or global constraints upon, protein sequences. We have also analysed in detail the volume variations at individual sites, both in the core and on the surface, and compared these variations with those expected from random sequences. Individual sites on the surface have nearly the same variation as random sequences (24% versus 28% variation). However, individual sites in the core have about half the variation of random sequences (13% versus 30%). Roughly, half of these core sites strongly conserve their volume (0 to 10% variation); one quarter have moderate variation (10 to 20%); and the remaining quarter vary randomly (20 to 40%). Our results have clear implications for the relationship between protein sequence and structure. For our analysis, we have developed a new and simple method for weighting protein sequences to correct for unequal representation, which we describe in an Appendix.

Algorithms

Solution structure of the DNA binding octapeptide repeat of the K10 gene product.

A putative transcription factor, the Drosophila K10 gene product, contains eight repeats of the octapeptide sequence SPNQQQHP or close variants. The solution structure of the K10 repeat was studied by NMR using a peptide composed of two SPNQQQHP units (referred to here as HP2). To overcome problems caused by degeneracy of backbone amide signals of Gln residues, a series of synthetic peptides containing an 15N-labelled main chain amide at different positions in HP2 were synthesized. In aqueous trifluoroethanol solution, HP2 folds into two structural units; the SPNQ part of each unit folds into a turn structure, while the C-terminal part shows some helical characteristics but is less structured. The N-terminal turn is likely to provide a core that produces a more stable helical structure upon binding to DNA and probably 'caps' the segmented helical unit at its N-terminus. This model is supported by a DNA footprinting study which shows that one SPNQQQHP unit spans four base pairs upon binding to A/T-rich sequences of DNA.

Amino Acid Sequence

Finding an average core structure: application to the globins.

We present a procedure for automatically identifying from a set of aligned protein structures a subset of atoms with only a small amount of structural variation, i.e., a core. We apply this procedure to the globin family of proteins. Based purely on the results of the procedure, we show that the globin fold can be divided into two parts. The part with greater structural variation consists of the residues near the heme (the F helix and parts of the G and H helices), and the part with lesser structural variation (the core) forms a structural framework similar to that of the repressor protein (A, B, and E helices and remainder of the G and H helices). Such a division is consistent with many other structural and biochemical findings. In addition, we find further partitions within the core that may have biological significance. Finally, using the structural core of the globin family as a reference point, we have compared structural variation to sequence variation and shown that a core definition based on sequence conservation does not necessarily agree with one based on structural similarity.

Algorithms