PubMed HealthSearch

Biomedical subjects

J M Thornton

Publications and source records attributed to J M Thornton.

At least 19 recordsLinked to original sources

Depicting topology and handedness in jellyroll structures.

The jellyroll structure is a special case of the Greek key topology and, to date, has only been observed in complete form in one of its four possible arrangements. Like other elements of super-secondary structure involving the beta-strand (e.g. the beta alpha beta unit) the known structure forms a right-handed superhelix. The possibility of losing such tertiary information and other problems associated with representing these structures by two-dimensional topology diagrams are discussed. A series of rules are presented which allow this three-dimensional information to be represented in two-dimensional topology diagrams from which the handedness of a jellyroll structure can be determined.

Protein Conformation

Towards an understanding of the arginine-aspartate interaction.

We have made a comparison of the geometries of intra- and intermolecular arginine-aspartate interactions by extracting orientation information from protein co-ordinate data. The results show a pronounced difference, with both types of interaction preferring to form twin N-H . . . O = C hydrogen bonds, but involving different nitrogen atoms. In intramolecular interactions, the aspartate favours a "side on" geometry, forming hydrogen bonds with N epsilon and N eta 2; in the intermolecular case, however, "end on" contacts involving N eta 1 and N eta 2 of the arginine are preferred. We have used Distributed Multipole Analysis of the methylguanidinium-acetate system to model the electrostatic component of the arginine-aspartate ion pair interaction in vacuo. We find, in agreement with the experimental arginine-aspartate distribution, that side on and end on doubly N-H . . . O = C hydrogen-bonded configurations are clearly the most favourable, with the side on being marginally lower in energy. Thus, despite the many competing side-chain interactions in proteins, many arginine-aspartate pairs adopt one of the minimum electrostatic energy conformations, or one close to a minimum. Within each of the two regions (side on and end on) we find only a small energy gap between the "symmetric" doubly hydrogen-bonded and slightly displaced "staggered" structures, again in agreement with the crystal structure data. Further calculations of the total ab initio interaction energy show that this follows the electrostatic term in its orientational variation, this phenomenon of "electrostatic domination" being well known in hydrogen-bonded systems. The end on arginine nitrogen atoms are observed to be more surface-exposed than N epsilon, as demonstrated by their greater accessibilities over a large sample of proteins. This helps explain the side on and end on preferences of intra- and intermolecular interactions, respectively. We also note the effect of short sequence intervals, particularly i in equilibrium with i + 2 relationships, in forcing many intramolecular contacts to be side on.

Arginine

A new approach to protein fold recognition.

The prediction of protein tertiary structure from sequence using molecular energy calculations has not yet been successful; an alternative strategy of recognizing known motifs or folds in sequences looks more promising. We present here a new approach to fold recognition, whereby sequences are fitted directly onto the backbone coordinates of known protein structures. Our method for protein fold recognition involves automatic modelling of protein structures using a given sequence, and is based on the frameworks of known protein folds. The plausibility of each model, and hence the degree of compatibility between the sequence and the proposed structure, is evaluated by means of a set of empirical potentials derived from proteins of known structure. The novel aspect of our approach is that the matching of sequences to backbone coordinates is performed in full three-dimensional space, incorporating specific pair interactions explicitly.

Algorithms

Stereochemical quality of protein structure coordinates.

Methods have been developed to assess the stereochemical quality of any protein structure both globally and locally using various criteria. Several parameters can be derived from the coordinates of a given structure. Global parameters include the distribution of phi, psi and chi 1 torsion angles, and hydrogen bond energies. There are clear correlations between these parameters and resolution; as the resolution improves, the distribution of the parameters becomes more clustered. These features show a broad distribution about ideal values derived from high-resolution structures. Some structures have tightly clustered distributions even at relatively low resolutions, while others show abnormal scatter though the data go to high resolution. Additional indicators of local irregularity include proline phi angles, peptide bond planarities, disulfide bond lengths, and their chi 3 torsion angles. These stereochemical parameters have been used to generate measures of stereochemical quality which provide a simple guide as to the reliability of a structure, in addition to the most important measures, resolution and R-factor. The parameters used in this evaluation are not novel, and are easily calculated from structure coordinates. A program suite is currently being developed which will quickly check a given structure, highlighting unusual stereochemistry and possible errors.

Databases, Bibliographic

Substrate recognition by proteinases.

The molecular recognition of limited proteolytic site substrates by serine proteinases has been compared and contrasted to the recognition of serine proteinase inhibitors, utilising the coordinate sets contained in the Brookhaven Protein Databank. Most families of these inhibitors are known to possess a structurally conserved recognition motif at their reactive site-binding loops. Structural comparisons with trypsin limited proteolytic sites revealed that the in situ conformation of these substrates bears little resemblance to the inhibitor-binding loops. Assuming that both inhibitors and substrates bind to the proteinase in the same manner, segmental mobility would be required to permit substrates to adopt an 'inhibitor-like' binding conformation, which is presumed to be necessary for proteolysis. Modelling experiments have been conducted to attempt to introduce such a conformation into tryptic limited proteolytic segments of the native proteins, to test the ability of the limited proteolytic sites to alter their geometry. Further to this, the conformational parameters of accessibility, protrusion, mobility and secondary structure have been analysed and incorporated into a predictive algorithm to assign likely limited proteolytic sites within native protein structures.

Binding Sites

The rapid generation of mutation data matrices from protein sequences.

An efficient means for generating mutation data matrices from large numbers of protein sequences is presented here. By means of an approximate peptide-based sequence comparison algorithm, the set sequences are clustered at the 85% identity level. The closest relating pairs of sequences are aligned, and observed amino acid exchanges tallied in a matrix. The raw mutation frequency matrix is processed in a similar way to that described by Dayhoff et al. (1978), and so the resulting matrices may be easily used in current sequence analysis applications, in place of the standard mutation data matrices, which have not been updated for 13 years. The method is fast enough to process the entire SWISS-PROT databank in 20 h on a Sun SPARCstation 1, and is fast enough to generate a matrix from a specific family or class of proteins in minutes. Differences observed between our 250 PAM mutation data matrix and the matrix calculated by Dayhoff et al. are briefly discussed.

Algorithms

A topological model for hepatitis B surface antigen.

A model of hepatitis B surface antigen has been derived, based on extensive sequence analysis and biochemical data. The surface antigen sequences of the human, woodchuck, ground squirrel and duck hepadnaviruses were examined using hydrophobicity, hydrophobic moments, flexibility and secondary structure prediction. The helix phase diagram, which is a modified version of Eisenberg's hydrophobic moment plots and which specifically addresses the problem of transmembrane helices, was used to examine the predicted helices. In this model four transmembrane helices are predicted. The N and C termini and the second hydrophilic region, which bears the major B-cell antigenic determinants, are external. It is suggested that the transmembrane helices may pack to form a channel through the membrane and may also be involved in the mechanisms of cell entry. A significant difference between the duck hepadnavirus and the mammalian HBsAg sequences was found, hence care must be taken when extrapolating data between the duck and the human surface antigen.

Amino Acid Sequence

Analysis of protein main-chain solvation as a function of secondary structure.

We have analysed the hydration of main-chain carbonyl and amide groups in 24 high-resolution well-refined protein structures as a function of the secondary structure in which these polar groups occur. We find that main-chain atoms in beta-sheets are as hydrated as those in alpha-helices, with most interactions involving "free" amide and carbonyl groups that do not participate in secondary structure hydrogen bonds. The distributions of water molecules around these non-bonded carbonyl groups reflect specific steric interactions due to the local secondary structure. Approximately 20% and 4%, respectively of bonded carbonyl and amide groups interact with solvent. These include interactions with carbonyl groups on the exposed faces of alpha-helices that have been correlated previously with bending of the helix. Water molecules interacting with alpha-helices occur mainly at the amino and carbonyl termini of the helices, in which case the solvent sites maintain the hydrogen bonding by bridging between residues i and i-3 or i-4 at the amino terminus and between i and i+3 or i+4 at the carbonyl terminus. We also see a number of solvent-mediated Ncap and Ccap interactions. The water molecules interacting with beta-sheets occur mainly at the edges, in which case they extend the sheet structure, or at the ends of strands, in which case they extend the beta-ladder. In summary, the solvent networks appear to extend the hydrogen-bonding structure of the secondary structures. In beta-turns, which usually occur at the surface of a protein, exposed amide and carbonyl groups are often hydrated, especially close to glycine residues. Occasionally water molecules form a bridge between residues i and i+3 in the turn and this may provide extra stabilization.

Amides

Molecular recognition. Conformational analysis of limited proteolytic sites and serine proteinase protein inhibitors.

The conformations of known tryptic limited proteolytic sites have been analysed and compared to the structures of the binding regions of serine proteinase inhibitors, as they are found when complexed to a serine proteinase. Conformational parameters studied include main-chain torsion angles, root-mean-square fits, accessibility, mobility and protrusion indices. As observed before, the inhibitors share a common main-chain conformation at the binding loop from P3-P'3 (Schechter & Berger notation), which is maintained throughout all the serine proteinase inhibitor families for which X-ray data is available, despite lack of similarity in the rest of the protein. This canonical structure is not found amongst the limited proteolytic sites (or nicksites), which differ markedly from the inhibitor binding loop conformation, and also amongst themselves. The experimentally determined nicksites are in general both accessible and protruding; as are the inhibitor binding loops, as well as being typically flexible regions of structure, as denoted by elevated temperature factors from crystallographic determinations. For cleavage by serine proteinases these loops must radically alter their local conformations and a large motion of the loop relative to the structure, in some cases, would be required to orientate these sites for cleavage.

Amino Acid Sequence

Pi-pi interactions: the geometry and energetics of phenylalanine-phenylalanine interactions in proteins.

The geometries of aromatic-aromatic interactions between phenylalanine residues in proteins are analysed in detail and correlated with energy calculations. A new definition of the interplanar angle is important for distinguishing favourable edge-to-face and unfavourable face-to-face orientations. The experimental observations are scattered over a wide range of conformational space, with no strongly preferred single orientation. However, Phe-Phe interactions occur almost exclusively in electrostatically attractive geometries: electrostatically unfavourable regions are only sparsely populated. Electrostatics dominate the geometry of interaction, while van der Waals' interactions are less significant, probably due to the hydrophobic environment of the protein core. The observations on proteins support the Hunter-Sanders rules for pi-pi interactions. In particular, offset stacked geometries, which theory predicts to be favourable, are observed experimentally. For monocyclic aromatics, use of a C-H dipole, the approach used in molecular mechanics calculations, accounts well for these aromatic-aromatic interactions. Comparison with the results obtained from the small molecules database indicates that the protein and small molecule crystal environments are very different.

Databases, Factual

Influence of proline residues on protein conformation.

To study the influence of proline residues on three-dimensional structure, an analysis has been made of all proline residues and their local conformations extracted from the Brookhaven Protein Data bank. We have considered the conformation of the proline itself, the relative occurrence of cis and trans peptides preceding proline residues, the influence of proline on the conformation of the preceding residue and the conformations of various proline patterns (Pro-Pro, Pro-X-Pro, etc.). The results highlight the unique role of proline in determining local conformation.

Amino Acid Sequence

An extension of secondary structure prediction towards the prediction of tertiary structure.

Secondary structure prediction parameters and optimised decision constants for use with the method of Garnier et al. [(1978) J. Mol. Biol. 120, 97-120] have been derived for two new and distinct substates of beta-structure. These we term internal and external on the basis of their hydrogen bonding patterns. The profiles of the amino acids for several of the parameters are considerably different in the two substates. Predictions using the new parameters attempt to distinguish the strands at the core of the beta-sheet from those at its edges and so restrict the possible topologies in tertiary structure prediction. The potential application of these parameters is illustrated for the class of beta/alpha proteins.

Adenylate Kinase

Modelling antibody combining sites: a review.

The combining site in an antibody is built up from six hypervariable loop regions, three from the light chain and three from the heavy chain. Three-dimensional structures have been elucidated by X-ray crystallographic studies for a variety of combining sites and for four protein-antibody complexes. Since sequence determination is relatively straightforward, whilst structure determination is often difficult and time consuming, it would be useful to be able to predict the structure of a combining site from its sequence. Knowledge of the structure would facilitate modifications of antibodies for specific aims. The structure of an antibody combining site for which only the sequence is known can be modelled on the basis of homology with a protein of known structure. The most demanding steps are modelling of the hypervariable loops and inclusion of the amino acid side chains. Final energy refinement of the model is achieved by conventional energy minimization techniques, but simulated annealing promises to be a powerful way of predicting side chain conformation. Four different groups have modelled antibody combining sites and compared their predictions with observed structures: for the main chain conformations resolutions of less than 1 A have been achieved. Future developments will increase the efficiency of the modelling procedures and permit accurate predictions of side chain conformations.

Binding Sites, Antibody

Structure prediction and modelling.

Protein structure prediction from sequence remains a major goal in molecular biology. The methods described in this review concentrate on deriving structural information through the detection of similarities between a test sequence and a database of known structures. Such methods are often referred to as knowledge-based strategies reflecting the use of a structural database in the analyses. The past year has seen considerable advances in both the development of automated procedures and their application to protein sequences of outstanding biological interest.

Algorithms

A novel method for the modelling of peptide ligands to their receptors.

A knowledge-based approach to the modelling of enzyme-peptide inhibitor complexes is described. Given the structure of an enzyme, and knowledge of its binding site, the method seeks to predict the binding geometry of a peptide ligand. This novel method involves using examples of side-chain packing derived from proteins of known three-dimensional structure to define possible packing arrangements of a peptide inhibitor group to its binding site. A suite of programs, GEMINI, was written and used to predict the packing of pairs of amino acid groups from three inhibitors complexed to their enzymes for which the X-ray structures were available. These included the Phe group of the inhibitor H142 bound to endothiapepsin, the Leu group of CLT complexed to thermolysin and the C-terminus of Gly-L-Tyr bound to carboxypeptidase A. A detailed comparison of the modelled and observed inhibitor coordinates was made. This approach may be extended to modelling other types of protein interactions.

Aspartic Acid Endopeptidases

Influence of secondary structure on the hydration of serine, threonine and tyrosine residues in proteins.

Previous analysis of experimental data on the solvation of high resolution protein structures has shown that preferred interaction sites for water molecules exist around most amino acid side chains. We have extended this analysis to look in more detail at the distributions around serine, threonine and tyrosine. We find that for serine and threonine side chains the preferred interaction sites of solvent molecules with the hydroxyl group depends on secondary structure and the chi 1 torsion angle of the side chain. For tyrosine side chains the hydroxyl group is too far from the main chain to reflect secondary structure influences. Specific patterns of hydration are observed in which water molecules 'bridge' between the hydroxyl side-chain atom and another main chain or side-chain atom.

Amino Acids