PubMed Health⌕ Search

Biomedical subjects

M W MacArthur

Publications and source records attributed to M W MacArthur.

14 recordsLinked to original sources

Surface beta-strands in proteins: identification using an hydropathy technique.

From a representative set of monomeric globular proteins with known three-dimensional structures, beta-strands with lengths > or = 5 amino acids have been identified and catalogued. By ascertaining the accessible surface areas of the constituent residues in these strands, and by checking whether the exposed/buried pattern is 80% or more similar to that in an idealized surface strand, a subset of structures can be delineated in which the beta-strands are all sited on the surface of the protein. The corresponding sequence data show that about 50% of the residues are apolar (Val, Ile, Leu, Phe, Tyr, Ala) and that the common occurrence of valine (14.3%), isoleucine (9.6%), and threonine (8.1%) is a characteristic feature. The frequencies of occurrence of those amino acids in the strands that face the aqueous environment and the interior have also been determined separately and show that most surface strands have a substructure of the form (apolar-X)(n), where X is approximately equally divided between apolar, charged, and hydrophilic residues. Using the frequency data thus obtained, allied with an algorithm to delineate potential surface beta-strands from characteristic hydropathy profiles, it is now possible to search through the sequences of proteins with unknown tertiary structures and make realistic predictions of the presence of this element of structure on the protein surface. In addition, new data are presented on the distribution of the various types of residues on the surface of proteins and in their interior. Significant differences were observed, not all of which have been identified previously. Furthermore, the distribution of the types of residue in a surface beta-strand was compared to that corresponding to the surfaces of all of the proteins in our database. Again, very characteristic differences were observed. These are helpful in recognizing the presence of surface beta-strands.

Algorithms↗

Protein side-chain conformation: a systematic variation of chi 1 mean values with resolution - a consequence of multiple rotameric states?

A systematic variation with resolution of the mean values of the gauche-, trans and gauche+ chi1 rotamers in protein structures determined by X-ray crystallography has been observed. Further analysis revealed that these correlations differ considerably between residue types, being highly significant for some residue types (e.g. Ser, Thr, Leu, Lys) and absent for others (e.g. aromatics). For the individual residue types which exhibited the trend most strongly, these changes were accompanied by corresponding systematic variations in the percentage relative populations in the three energy wells. Examination of a uniformly sized subset of monomers showed that this effect, while attenuated, was still present, and was thus not entirely a consequence of the change in size and surface area which also correlates with resolution. An analysis of B values in the disfavoured high-energy barrier region between the rotameric wells showed a pronounced tendency towards larger than average values. As a plausible hypothesis, it is suggested here that these observations can be accounted for by the presence of multiple rotameric states. The averaged electron density produced by dual occupancy at low resolution giving an averaged conformation is resolved at high resolution into its individual components.

Asparagine↗

Validation of protein models derived from experiment.

The growing number of protein structures solved at atomic resolution holds the promise of further improvements in geometry-based validation parameters. Additionally, the estimated standard uncertainties of the atomic coordinates have been computed for a number of X-ray structures, providing a measure of the coordinate precision. In NMR spectroscopy, a measure analogous to the crystallographic R-factor has been developed.

Crystallography, X-Ray↗

Assessment of comparative modeling in CASP2.

An assessment is presented for all submissions to the comparative modeling challenge in the 1996 Critical Assessment of Structure Prediction (CASP2). Of the original 12 target structures, 9 were solved prior to the meeting: 8 by X-ray crystallography and 1 by NMR spectroscopy. These targets varied over a large range of difficulty, as assessed by the percentage sequence identity with the principal parent structure, which ranged from 20% up to 85%. The overall quality of the models reflected the similarity of the principal parent. As expected, when the sequence alignment was correct, the core was accurately modeled, with the largest deviations occurring in the loops. Models were built which gave C alpha root-mean-square deviations (RMSDs) compared with the observed structure of < 1 A for targets with high parental similarity; even at 26% sequence identity, the best model structures had C alpha deviations of only 2.2 A. Overall, these deviations are comparable with those observed between the parent structure and the target, but locally there are several examples where the model approaches closer to the target than does the parent. There were three targets below 25% sequence identity, and the models generated for these targets were, in general, significantly less accurate. This principally reflects errors in the alignment which, if systematically shifted, can generate C alpha RMSDs > 18 A. Compared with CASP1, the geometry of the models was significantly improved with no D-amino acids. By far the major contribution to RMSD error was the alignment accuracy, which varied from 100% down to 7% over the range of targets. In the structurally variable regions, global shifts, caused by hinge bending, were the major source of error, giving significantly lower local RMSDs than global RMSDs. In over 50% of these noncore regions, the difference between global and local RMSDs was more than 3 A, and was as high as 10 A for one structurally variable region. For the side chains, the chi 1 RMSDs are strongly correlated with the C alpha RMSDs. For models with C alpha deviations less than 1 A, on average 78.5% of side chains are placed in the correct rotamer, although the chi 1 RMSDs, though clearly better than random, were disappointing at around 46 degrees. As the backbone deviations increased, the side chain placement became less accurate, with an average chi 1 RMSD of 75 degrees on a 1.5-2.5 A C alpha backbone (average 51.4% correct rotamer). Refinement by energy minimization or molecular dynamics made only minor adjustments to improve local geometry and generally made small, but not significant, improvements to the RMSD. In total, 19 groups submitted 62 models (89 coordinate sets) that could be assessed. Most modelers used manual adjustments to sequence alignments and, in general, good alignments were obtained down to 25% sequence identity. The modeling methods ranged from "classical" modeling, involving core building followed by loop and side chain addition, to more sophisticated approaches based on probability distributions, Monte Carlo sampling or distance geometry. For each target, several groups produced equally good models, given the expected errors in the structures (about 0.5 A). No one method came out as clearly superior, although the approaches that inherit directly from the parents generally performed better than the more radical techniques. However, for each target there were some poor models, usually reflecting a poor sequence alignment, and the range of accuracy for each target is therefore large. Fully automated methods are able to perform very well for "easy" targets (85% sequence identity with parent), but when modeling using a distantly related parent, care and expertise, especially in performing the alignment, still appear to be important factors in generating accurate models.

Models, Molecular↗

Structures of N-termini of helices in proteins.

We have surveyed 393 N-termini of alpha-helices and 156 N-termini of 3(10)-helices in 85 high resolution, non-homologous protein crystal structures for N-cap side-chain rotamer preferences, hydrogen bonding patterns, and solvent accessibilities. We find very strong rotamer preferences that are unique to N-cap sites. The following rules are generally observed for N-capping in alpha-helices: Thr and Ser N-cap side chains adopt the gauche - rotamer, hydrogen bond to the N3 NH and have psi restricted to 164 +/- 8 degrees. Asp and Asn N-cap side chains either adopt the gauche - rotamer and hydrogen bond to the N3 NH with psi = 172 +/- 10 degrees, or adopt the trans rotamer and hydrogen bond to both the N2 and N3 NH groups with psi = 1-7 +/- 19 degrees. With all other N-caps, the side chain is found in the gauche + rotamer so that the side chain does not interact unfavorably with the N-terminus by blocking solvation and psi is unrestricted. An i, i + 3 hydrogen bond from N3 NH to the N-cap backbone C = O in more likely to form at the N-terminus when an unfavorable N-cap is present. In the 3(10)-helix Asn and Asp remain favorable N-caps as they can hydrogen bond to the N2 NH while in the trans rotamer; in contrast, Ser and Thr are disfavored as their preferred hydrogen bonding partner (N3 NH) is inaccessible. This suggests that Ser is the optimum choice of N-cap when alpha-helix formation is to be encouraged while 3(10)-helix formation discouraged. The strong energetic and structural preferences found for N-caps, which differ greatly from positions within helix interiors, suggest that N-caps should be treated explicitly in any consideration of helical structure in peptides or proteins.

Hydrogen Bonding↗

Deviations from planarity of the peptide bond in peptides and proteins.

The work described here is the result of a survey of the peptide omega angles in the Cambridge Structural Database of small molecules, which was carried out to establish "ideal" or "target" values for their distribution. We have shown that substantial deviations from planarity can be tolerated with a standard deviation in the angle of up to 6 degrees about a mean value for the trans peptide that is less than 180 degrees . The distortion can arise from pyramidalization at the amino nitrogen atom as well as simple twist about the peptide bond. We include an analysis of omega angles in the existing database of protein structure (PDB) and show that their distributions can depend on the refinement method used, but no correlation with resolution is evident. A surprising finding was a systematic variation of omega in phi,psi space in proteins as well as in the linear and cyclic peptides. This is particularly manifest as a consistent difference between the mean omega values in chains of left and right-hand chirality. This dichotomy is observed for all the standard amino acids and is especially striking in the absence of secondary structure. The phenomenon is discussed in the context of theoretical work on peptide analogues, and the implications for protein conformation and structure are briefly considered.

Amino Acids↗

Analysis of main chain torsion angles in proteins: prediction of NMR coupling constants for native and random coil conformations.

Using a data base of 85 high resolution protein crystal structures the distributions of main chain torsion angles, both in secondary structure and in coil regions where no secondary structure is present, have been analysed. These torsion angle distributions have been used to predict NMR homonuclear and heteronuclear coupling constants for residues in secondary structure using known Karplus relationships. For alpha helices, 3(10) helices and beta strands mean predicted 3JHN alpha coupling constants are 4.8, 5.6 and 8.5 Hz, respectively. These values differ significantly from those expected for the ideal phi angles (3.9, 3.0 and 8.9 Hz; phi = -57 degrees, -49 degrees, -139 degrees for alpha and 3(10) helices and beta strands (antiparallel), respectively) in regular secondary structure, but agree well with available experimental NMR data for nine proteins. The crystallographic data set has also been used to provide a basis for interpreting coupling constants measured for peptides and denatured proteins. Using a model for a random coil, in which all residues adopt distributions of phi, psi angles equivalent to those seen for residues in the coil regions of native folded proteins, predicted 3JHN alpha values for different residue types have been found to range from 5.9 Hz and 6.1 Hz for glycine and alanine, respectively, to 7.7 Hz for valine. A good correlation has been found between the predicted 3JHN alpha coupling constants for this model and experimental values for a set of peptides that other evidence suggest are highly unstructured. For other peptides, however, deviations from the predictions of the model are clear and provide evidence for additional interactions within otherwise disordered states. The values of homonuclear and heteronuclear coupling constants derived from the protein data base listed here therefore provide a basis not only for analysing the secondary structure of native proteins in solution but for assessing and interpreting the extent of structure present in peptides and non-native states of proteins.

Databases, Factual↗

AQUA and PROCHECK-NMR: programs for checking the quality of protein structures solved by NMR.

The AQUA and PROCHECK-NMR programs provide a means of validating the geometry and restraint violations of an ensemble of protein structures solved by solution NMR. The outputs include a detailed breakdown of the restraint violations, a number of plots in PostScript format and summary statistics. These various analyses indicate both the degree of agreement of the model structures with the experimental dat, and the quality of their geometrical properties. They are intended to be of use both to support ongoing NMR structure determination and in the validation of the final results.

Amino Acid Sequence↗

Protein folds: towards understanding folding from inspection of native structures.

Following a short summary of some of the principal features of folded proteins, the results of two complementary studies of protein structure are presented, the first concerned with the factors which influence secondary structure propensity and the second an analysis of protein topology. In an attempt to deconvolute the physical contributions to secondary structure propensities, we have calculated intrinsic phi, psi propensities, derived from the coil regions of proteins. Comparison of intrinsic phi, psi propensities with their equivalent secondary structure values show correlations for both helix and strand. This suggests that the local dipeptide, steric and electrostatic interactions have a major influence on secondary structure propensity. We then proceed to inspect the distribution of protein domain folds observed to date. Several folds occur very commonly, so that 46% of the current non-homologous database comprises only nine folds. The implications of these results for protein folding are discussed.

Computer Simulation↗

Intrinsic phi, psi propensities of amino acids, derived from the coil regions of known structures.

Many different factors contribute to secondary structure propensities, including phi, psi preferences, side-chain interactions, steric effects and hydrophobic tertiary contacts. To deconvolute these competing factors, we have adopted a novel approach which quantifies the intrinsic phi, psi propensities for residues in coil regions (that is, residues not in alpha-helix and not in beta-strand). Comparisons of intrinsic phi, psi propensities with their equivalent secondary structure propensities show that while correlations for helix are relatively weak, those for strand are much stronger. This paper describes our new phi, psi propensities and provides an explanation for the variations observed.

Amino Acids↗

NMR and crystallography--complementary approaches to structure determination.

A knowledge of the three-dimensional structure of a protein is essential to understand how a protein performs its functions. It is also a prime requirement for the rational design of novel sequences with specific structural, chemical or catalytic properties. Until recently, such information could only be obtained from X-ray diffraction studies on protein crystals. In the past few years, however, nuclear magnetic resonance (NMR) spectroscopy in solution has rapidly become established as an effective alternative method. This review briefly examines the two techniques and their relevance to protein engineering and design.

Animals↗

Conformational analysis of protein structures derived from NMR data.

A study is presented of the conformational characteristics of NMR-derived protein structures in the Protein Data Bank compared to X-ray structures. Both ensemble and energy-minimized average structures are analyzed. We have addressed the problem using the methods developed for crystal structures by examining the distribution of phi, psi, and chi angles as indicators of global conformational irregularity. All these features in NMR structures occur to varying degrees in multiple conformational states. Some measures of local geometry are very tightly constrained by the methods used to generate the structure, e.g., proline phi angles, alpha-helix phi,psi angles, omega angles, and C alpha chirality. The more lightly restrained torsion angles do show increased clustering as the number of overall experimental observations increases. phi, psi, and chi 1 angle conformational heterogeneity is strongly correlated with accessibility but shows additional differences which reflect the differing number of observations possible in NMR for the various side chains (e.g., many for Trp, few for Ser). In general, we find that the core is defined to a notional resolution of 2.0 to 2.3 A. Of real interest is the behavior of surface residues and in particular the side chains where multiple rotameric states in different structures can vary from 10% to 88%. Later generation structures show a much tighter definition which correlates with increasing use of J-coupling information, stereospecific assignments, and heteronuclear techniques. A suite of programs is being developed to address the special needs of NMR-derived structures which will take into account the existence of increased mobility in solution.

Analysis of Variance↗

Stereochemical quality of protein structure coordinates.

Methods have been developed to assess the stereochemical quality of any protein structure both globally and locally using various criteria. Several parameters can be derived from the coordinates of a given structure. Global parameters include the distribution of phi, psi and chi 1 torsion angles, and hydrogen bond energies. There are clear correlations between these parameters and resolution; as the resolution improves, the distribution of the parameters becomes more clustered. These features show a broad distribution about ideal values derived from high-resolution structures. Some structures have tightly clustered distributions even at relatively low resolutions, while others show abnormal scatter though the data go to high resolution. Additional indicators of local irregularity include proline phi angles, peptide bond planarities, disulfide bond lengths, and their chi 3 torsion angles. These stereochemical parameters have been used to generate measures of stereochemical quality which provide a simple guide as to the reliability of a structure, in addition to the most important measures, resolution and R-factor. The parameters used in this evaluation are not novel, and are easily calculated from structure coordinates. A program suite is currently being developed which will quickly check a given structure, highlighting unusual stereochemistry and possible errors.

Databases, Bibliographic↗

Influence of proline residues on protein conformation.

To study the influence of proline residues on three-dimensional structure, an analysis has been made of all proline residues and their local conformations extracted from the Brookhaven Protein Data bank. We have considered the conformation of the proline itself, the relative occurrence of cis and trans peptides preceding proline residues, the influence of proline on the conformation of the preceding residue and the conformations of various proline patterns (Pro-Pro, Pro-X-Pro, etc.). The results highlight the unique role of proline in determining local conformation.

Amino Acid Sequence↗