PubMed HealthSearch

Biomedical subjects

C Chothia

Publications and source records attributed to C Chothia.

At least 55 records · Page 3Linked to original sources

Structural mechanisms for domain movements in proteins.

We survey all the known instances of domain movements in proteins for which there is crystallographic evidence for the movement. We explain these domain movements in terms of the repertoire of low-energy conformation changes that are known to occur in proteins. We first describe the basic elements of this repertoire, hinge and shear motions, and then show how the elements of the repertoire can be combined to produce domain movements. We emphasize that the elements used in particular proteins are determined mainly by the structure of the interfaces between the domains.

Motion

Many of the immunoglobulin superfamily domains in cell adhesion molecules and surface receptors belong to a new structural set which is close to that containing variable domains.

On the basis of similarities in sequence and structure, the protein domains that form the immunoglobulin superfamily have been divided into three sets: one with variable-like domains, the V set, and two with different variants of the constant-like domains, the C1 and C2 sets. Examination of a muscle member of the immunoglobulin superfamily, telokin, shows that its structure is closely related to those of the variable domains found in antibodies, CD2, CD4 and CD8. However, it also contains structural features that, previously, have only been found in constant domains. Telokin represents a new structural set in the superfamily which we call the I set. Using the structures of telokin, and variable domains from antibodies, CD4 and CD8, we constructed a profile that describes the sequence characteristics of the structural core common to those proteins. This sequence profile makes a good match to the sequences of many of the immunoglobulin superfamily domains that form the cell adhesion molecules and surface receptors. This match implies that these domains also have structures that belong to the I set.

Amino Acid Sequence

Principles determining the structure of beta-sheet barrels in proteins. I. A theoretical analysis.

The major feature of many proteins is a large beta-sheet that twists and coils to form a closed structure in which the first strand is hydrogen bonded to the last: the beta-sheet barrel. McLachlan classified barrels in terms of two integral parameters: the number of strands in the beta-sheet, n, and the "shear number", S, a measure of the stagger of the strands in the beta-sheet. He showed that the mean radius of a barrel and the extent to which strands are tilted relative to its axis are determined by the values of n and S. Here we show that the (n, S) values determine all the other general structural features of regular beta-sheet barrels, in particular, optimal values of the twist and coiling angles that produce the closed beta-sheet, the hyperboloidal shape and the arrangement of residues in the barrel interior. Consideration of the residue arrangements in the interiors of different potential barrel structures, and of side-chain volumes, suggest that barrels, in which the interiors are close packed by the residues in beta-sheets with good geometries, have structures that correspond to one of only ten different combinations of n and S. In the accompanying paper, we demonstrate, by an analysis of all observed protein structures that contain beta-sheet barrels and for which atomic co-ordinates are available, the validity of these theoretical results.

Hydrogen Bonding

Principles determining the structure of beta-sheet barrels in proteins. II. The observed structures.

In the accompanying paper we derived a set of principles that, we argue, govern the structure of beta-sheet barrels. Barrel structures are classified in terms of two integral parameters: the number of strands in the beta-sheet, n, and a measure of the stagger in the beta-sheet, S. We derived a set of equations that show how the (n, S) values of a barrel structure determine the arrangement of its strands; its general shape; the twist and coiling of the beta-sheet, and the arrangement of residues in the barrel interior. This work suggested that there are ten different combinations of n and S that form barrels with good beta-sheet geometries and interiors close packed by beta-sheet residues. In this paper we demonstrate the validity of these principles. We analyse in detail the observed structures of 39 different beta-sheet barrels. These structures include representatives of all the different barrel structures currently known and for which atomic co-ordinates are available. We show that the observed arrangement of the strands, and the extent of the twist and coiling of the beta-sheets, are very close to those calculated from the (n, S) values for the barrel. Of the 39 structures, 34 have one of the ten (n, S) values that we expect to form barrels with good beta-sheet geometries and interiors close packed by beta-sheet residues. The other five have one of two (n, S) values that give good beta-sheet geometries but radii so large the beta-sheet residues leave cavities at the centre of the barrels. In at least four of these cavities have a functional role.

Amino Acid Sequence

Volume changes in protein evolution.

We have determined the variations in volume that occur during evolution in the buried core of three different families of proteins. The variation of the whole core is very small (approximately 2.5%) compared to the variation at individual sites (approximately 13%). However, by comparing our results to those expected from random sequences with no correlations between sites, we show that the small variation observed may simply be a manifestation of the statistical "law of large numbers" and not reflect any compensating changes in, or global constraints upon, protein sequences. We have also analysed in detail the volume variations at individual sites, both in the core and on the surface, and compared these variations with those expected from random sequences. Individual sites on the surface have nearly the same variation as random sequences (24% versus 28% variation). However, individual sites in the core have about half the variation of random sequences (13% versus 30%). Roughly, half of these core sites strongly conserve their volume (0 to 10% variation); one quarter have moderate variation (10 to 20%); and the remaining quarter vary randomly (20 to 40%). Our results have clear implications for the relationship between protein sequence and structure. For our analysis, we have developed a new and simple method for weighting protein sequences to correct for unequal representation, which we describe in an Appendix.

Algorithms

Structural conservation of hypervariable regions in immunoglobulins evolution.

Analysis of human and mouse immunoglobulins has shown that five of six hypervariable regions that form the antigen binding site have a small repertoire of main chain conformations (canonical structures). Cartilaginous fishes are the most distantly related species to humans known to have an immune system, their evolutionary lines having diverged 450 million years ago. An analysis of VH and V kappa sequences from these fishes shows that all the main chain structures in their L1, L2, H1 and H2 hypervariable regions, and one of those in the L3 region, are the same as those most commonly found in human and mouse. This implies that the canonical structures occurring most commonly in hypervariable regions arose very early in the stages of the evolution of the immune system.

Amino Acid Sequence

Protein families in the metazoan genome.

The evolution of development involves the development of new proteins. Estimates based on the initial results of the genome projects, and on the data banks of protein sequences and structures, suggest that the large majority of proteins come from no more than one thousand families. Members of a family are descended from a common ancestor. Protein families evolve by gene duplication and mutation. Mutations change the conformation of the peripheral regions of proteins; i.e. the regions that are involved, at least in part, in their function. If mutations proceed until only 20% of the residues in related proteins are identical, it is common for the conformational changes to affect half the structure. Most of the proteins involved in the interactions of cells, and in their assembly to form multicellular organisms, are mosaic proteins. These are large and have a modular structure, in that they are built of sets of homologous domains that are drawn from a relatively small number of protein families. Patthy's model for the evolution of mosaic proteins describes how they arose through the insertion of introns into genes, gene duplications and intronic recombination. The rates of progress in the genome sequencing projects, and in protein structure analyses, means that in a few years we will have a fairly complete outline description of the molecules responsible for the structure and function of organisms at several different levels of developmental complexity. This should make a major contribution to our understanding of the evolution of development.

Animals

Domain closure in lactoferrin. Two hinges produce a see-saw motion between alternative close-packed interfaces.

Lactoferrin is an iron transport protein. Upon binding iron, the two domains in the N-terminal half of the molecule move together. Previous work has shown that this domain closure involves two hinges. Using the newly refined structure of the open form, the structural mechanism underlying this motion is analyzed here in detail. Upon closure the domains rotate 54 degrees essentially as rigid bodies. The axis of rotation passes through the two beta-strands linking the domains. These strands contain hinges in the sense that three large torsion angle changes are responsible for the bulk of the motion while smaller torsion angle changes in neighboring residues are responsible for the remainder of the motion. The rotation axes of these three torsion angle changes are nearly parallel to the axis of the overall 54 degrees rotation, so the local motion in the hinges can be directly related to the overall motion. A crucial feature of the hinge residues is that they have very few packing constraints on their main-chain atoms. The domains make different packing contacts with each other in the open and closed forms. These contacts form two interdomain interfaces arranged on either side of the hinges. Pivoting about the hinges produces a see-saw motion between the two interfaces. That is, when the domains close down, residues in the interface on one side of the hinges become buried and close-packed and residues on the other side become exposed. The situation is reversed when the domains open up. Lactoferrin provides a particularly clear example of the general features of hinged domain motion. It is compared to other instances of hinged domain closure and contrasted with instances of shear domain closure, where the overall motion is a summation of many small sliding motions between close-packed segments of polypeptide.

Amino Acid Sequence

Domain closure in adenylate kinase. Joints on either side of two helices close like neighboring fingers.

In large variants of adenylate kinase the AMP and ATP substrates are buried by a domain rotating by 90 degrees. Here conformational changes responsible for this domain closure are determined by an analysis of the open state of beef heart mitochondrial adenylate kinase and the closed state of Escherichia coli adenylate kinase. Although these two proteins have sequence differences, the principal structural changes responsible for the domain movements are large, and can clearly be distinguished from the effects of evolution. The mobile domain is linked to the rest of the protein by two helices packed together in an antiparallel fashion. During the closure, deformations take place in four localized regions, called joints, near the N and C termini of these helices. Three of these joints have simple motions that can be well approximated by rotations of three torsion angles, but the joint that makes contact with the ligand involves motion throughout an extended loop: i.e. two torsions on either side of a reverse turn change significantly. The main chain atoms of the joints have few packing constraints. The first pair of joints is responsible for approximately 30 degrees of the total rotation and the second pair for the remaining approximately 60 degrees. These movements carries along the regions between the joints, the two helices and the rest of the mobile domain, to a first approximation, as rigid bodies. This jointed domain closure mechanism is contrasted with the shear mechanisms found in other enzymes.

Adenylate Kinase

Structural repertoire of the human VH segments.

The VH gene segments produce the part of the VH domains of antibodies that contains the first two hypervariable regions. The sequences of 83 human VH segments with open reading frames, from several individuals, are currently known. It has been shown that these sequences are likely to form a high proportion of the total human repertoire and that an individual's gene repertoire produces about 50 VH segments with different protein sequences. In this paper we present a structural analysis of the amino acid sequences produced by the 83 segments. Particular residue patterns in the sequences of V domains imply particular main-chain conformations, canonical structures, for the hypervariable regions. We show that, in almost all cases, the residue patterns in the VH segments imply that the first hypervariable regions have one of three different canonical structures and that the second hypervariable regions have one of five different canonical structures. The different observed combinations of the canonical structures in the first and second regions means that almost all sequences have one of seven main-chain folds. We describe, in outline, structures of the antigen binding site loops produced by nearly all the VH segments. The exact specificity of the loops is produced by (1) sequence differences in their surface residues, particularly at sites near the centre of the combining site, and (2) sequence differences in the hypervariable and framework regions that modulate the relative positions of the loops.

Amino Acid Sequence

Domain closure in mitochondrial aspartate aminotransferase.

The subunits of the dimeric enzyme aspartate aminotransferase have two domains: one large and one small. The active site lies in a cavity that is close to both the subunit interface and the interface between the two domains. On binding the substrate the domains close together. This closure completely buries the substrate in the active site and moves two arginine side-chains so they form salt bridges with carboxylate groups of the substrate. The salt bridges hold the substrate close to the pyridoxal 5'-phosphate cofactor and in the right position and orientation for the catalysis of the transamination reaction. We describe here the structural changes that produce the domain movements and the closure of the active site. Structural changes occur at the interface between the domains and within the small domain itself. On closure, the core of the small domain rotates by 13 degrees relative to the large domain. Two other regions of the small domain, which form part of the active site, move somewhat differently. A loop, residues 39 to 49, above the active site moves about 1 A less than the core of the small domain. A helix within the small domain forms the "door" of the active site. It moves with the core of the small domain and, in addition, shifts by 1.2 A, rotates by 10 degrees, and switches its first turn from the alpha to the 3(10) conformation. This results in the helix closing the active site. The domain movements are produced by a co-ordinated series of small changes. Within one subunit the polypeptide chain passes twice between the large and small domains. One link involves a peptide in an extended conformation. The second link is in the middle of a long helix that spans both domains. At the interface this helix is kinked and, on closure, the angle of the kink changes to accommodate the movement of the small domain. The interface between the domains is formed by 15 residues in the large domain packing against 12 residues in the small domain and the manner in which these residues pack is essentially the same in the open and closed structures. Domain movements involve changes in the main-chain and side-chain torsion angles in the residues on both sides of the interface. Most of these changes are small; only a few side-chains switch to new conformations.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

beta-Trefoil fold. Patterns of structure and sequence in the Kunitz inhibitors interleukins-1 beta and 1 alpha and fibroblast growth factors.

Previous crystallographic analyses of the Kunitz inhibitors from soybean. Erythrina caffra and wheat, the interleukins-1 beta and 1 alpha and the acidic and basic fibroblast growth factors have shown that they contain a most unusual fold. It is formed by six two-stranded hairpins. Three of these form a barrel structure and the other three are in a triangular array that caps the barrel. The arrangement of the secondary structures gives the molecules a pseudo 3-fold axis. Although the different proteins have very similar structures, many of their sequences have no significant similarities overall. The structural determinants of this fold are described and discussed in this paper. The barrels in the different proteins have the same geometrical features: six strands tilted at 56 degrees to the barrel axis; a barrel diameter of 16 A, and the beta-sheet hydrogen bonded so that it is staggered with a shear number of 12. These features fit McLachlan's equations for ideal barrels formed by beta-sheets. The wide diameter of the barrels is filled by layers of residues that, while not identical in the different proteins, are, in almost all cases, large. The structure of the triangular array of hairpins is determined by the coiling of the strands and the packing of hairpin residues against each other and against residues from the interior of the barrel. The major sequence requirements of this fold are large or medium hydrophobic residues at 18 buried sites. In the different structures the total volume of these residues is 3000 (+/- 120) A. The polyhedron model of protein architecture is used to demonstrate that the main, and in particular the symmetrical, features of this fold arise from the ideal and equal packing of six hairpins, modified only slightly to form hydrogen bonds between the hairpins.

Amino Acid Sequence

Serpin tertiary structure transformation.

Previous crystallographic analyses have demonstrated that proteolytic cleavage of the serpins can result in a dramatic transformation of their tertiary structure. Some 16 residues on the amino terminal side of the cleavage site are inserted into a large beta-sheet to become a central strand, separating the two cleaved residues by about 70 A. We have determined, in outline, the nature of the conformational change responsible for this transformation. After cleavage, a fragment of the protein, consisting of an alpha-helix and three strands of beta-sheet, moves away from the rest of the structure to make the space for the new strand. This movement involves a new type of structural change: sheet residues in the small fragment slide along grooves in an alpha-helix that belongs to the rest of the protein. The general conservation of residues in the regions between the small fragment and the rest of the protein imply that the same mechanism will be found in all serpins that undergo this tertiary structure transformation.

Amino Acid Sequence

Analysis of protein loop closure. Two types of hinges produce one motion in lactate dehydrogenase.

As shown in previous crystallographic investigations, upon binding lactate and NAD, lactate dehydrogenase undergoes a large conformational change that results in a surface loop moving roughly 10 A to cover the active site. In addition, there are appreciable movements (approximately 2 A) of five helices and three other loops. We demonstrate by a new fitting procedure that the loop moves on two hinges separated by a relatively rigid type II turn. The first hinge has few steric constraints on it, and its motion can be well accounted for by large changes in two torsion angles, i.e. as in a classic hinge motion. In contrast, the second hinge, which is part of a helix connected to the end of the loop, has many more constraints on it and distributes its deformation over more torsion angles. This novel motion involves the helix stretching and splitting into alpha-helical and 3(10)-helical components and substantial side-chain repacking in the sense of "cogs hopping between grooves" at its interface with the end of a neighboring helix. The loop is stabilized by five transverse (across loop) hydrogen bonds. These are preserved, through the conformational change and through 17 lactate dehydrogenase sequences, more than the longitudinal hydrogen bonds down the sides of the loop. Through a network of contacts, many of them conserved hydrophobic residues, the motion of the loop is propagated outward to structures that have no direct contact with the ligands. These moving structures are on the surface of the protein, and the whole protein can be subdivided into concentric shells of increasing mobility.

Amino Acid Sequence

The structural homology of amicyanin from Thiobacillus versutus to plant plastocyanins.

The complete amino acid sequence of the blue copper protein amicyanin of Thiobacillus versutus, induced when the bacterium is grown on methylamine, has been determined as follows: QDKITVTSEKPVAAADVPADAVVVGIEKMKYLTPEVTIKAGETVYWVNGEVMPHNVA FKKGIVGEDAFRGEMMTKDQAYAITFNEAGSYDYFCTPHPFMRGKVIVE. The four copper ligand residues in this 106-residue-containing polypeptide chain are His54, Cys93, His96, and Met99. The Thiobacillus amicyanin is 52% similar to the amicyanin of Pseudomonas AM1, the only other copper protein known with the same spacing between the second histidine ligand and the methionine ligand. T. versutus amicyanin contains no cysteine bridge and is more closely related to the plant copper protein plastocyanin than to the bacterial copper protein azurin. Alignment of the two known amicyanin sequences with the consensus sequence of the plastocyanins and comparison with the known three-dimensional structure of poplar leaves plastocyanin reveals that the bacterial proteins have the same overall structure with two beta-sheets packed face to face. The major structural differences between the amicyanins and the plastocyanins appear to be located in two of the five loops that connect the six identified beta-strands of the amicyanins. The first of these two loops, connecting strands F and G, contains a ligand histidine and must have a different conformation from the same loop in the plastocyanins because it is shorter by two amino acids. Further differences occur in the loop connecting the strands D and E. This loop contains only 17 residues in amicyanin whereas the corresponding loop of plastocyanin contains 25 residues. Despite these differences the amicyanins appear much closer related to the plastocyanins than to the azurins. The present findings demonstrate that the occurrence of blue copper proteins with clearly plastocyanin-like features is not restricted to photosynthetic redox chains.

Amino Acid Sequence

Asymmetry in protein structures.

The asymmetry of L-amino acids determines the asymmetrical features of alpha-helices and beta-sheets. These in turn determine two principal aspects of the three-dimensional structure of proteins: the preferred ways in which alpha-helices and beta-sheets pack together, and certain topological features of the paths followed by polypeptide chains through structures. Though the asymmetrical nature of amino acids plays the central role in determining the asymmetrical aspects of protein structures, it has little or no influence on the next level of biological structures--assemblies of protein molecules.

Amino Acids