PubMed Health⌕ Search

Biomedical subjects

Andrew J Doig

Publications and source records attributed to Andrew J Doig.

7 recordsLinked to original sources

Distinguishing enzyme structures from non-enzymes without alignments.

The ability to predict protein function from structure is becoming increasingly important as the number of structures resolved is growing more rapidly than our capacity to study function. Current methods for predicting protein function are mostly reliant on identifying a similar protein of known function. For proteins that are highly dissimilar or are only similar to proteins also lacking functional annotations, these methods fail. Here, we show that protein function can be predicted as enzymatic or not without resorting to alignments. We describe 1178 high-resolution proteins in a structurally non-redundant subset of the Protein Data Bank using simple features such as secondary-structure content, amino acid propensities, surface properties and ligands. The subset is split into two functional groupings, enzymes and non-enzymes. We use the support vector machine-learning algorithm to develop models that are capable of assigning the protein class. Validation of the method shows that the function can be predicted to an accuracy of 77% using 52 features to describe each protein. An adaptive search of possible subsets of features produces a simplified model based on 36 features that predicts at an accuracy of 80%. We compare the method to sequence-based methods that also avoid calculating alignments and predict a recently released set of unrelated proteins. The most useful features for distinguishing enzymes from non-enzymes are secondary-structure content, amino acid frequencies, number of disulphide bonds and size of the largest cleft. This method is applicable to any structure as it does not require the identification of sequence or structural similarity to a protein of known function.

Algorithms↗

Design strategies for anti-amyloid agents.

Numerous diseases have been linked to a common pathogenic process called amyloidosis, whereby proteins or peptides clump together in the brain or body to form toxic soluble oligomers and/or insoluble fibres. An attractive strategy to develop therapies for these diseases is therefore to inhibit or reverse protein/peptide aggregation. A diverse range of small organic ligands have been found to act as aggregation inhibitors. Alternatively, the wild-type peptide can be derivatised so that it still binds to the amyloid target, but prevents further aggregation. This can be achieved by adding a bulky group or charged amino acid to either end of the peptide, or by incorporating proline residues or N-methylated amide groups.

Amyloid beta-Peptides↗

Recent advances in helix-coil theory.

Peptide helices in solution form a complex mixture of all helix, all coil or, most frequently, central helices with frayed coil ends. In order to interpret experiments on helical peptides and make theoretical predictions on helices, it is therefore essential to use a helix-coil theory that takes account of this equilibrium. The original Zimm-Bragg and Lifson-Roig helix-coil theories have been greatly extended in the last 10 years to include additional interactions. These include preferences for the N-cap, N1, N2, N3 and C-cap positions, capping motifs, helix dipoles, side chain interactions and 3(10)-helix formation. These have been applied to determine energies for these preferences from experimental data and to predict the helix contents of peptides. This review discusses these newly recognised structural features of helices and how they have been included in helix-coil models.

Peptides↗

Stabilizing interactions between aromatic and basic side chains in alpha-helical peptides and proteins. Tyrosine effects on helix circular dichroism.

Here we investigate the structures and energetics of interactions between aromatic (Phe or Tyr) and basic (Lys or Arg) amino acids in alpha-helices. Side chain interaction energies are measured using helical peptides, by quantifying their helicities with circular dichroism at 222 nm and interpreting the results with Lifson-Roig-based helix/coil theory. A difficulty in working with Tyr is that the aromatic ring perturbs the CD spectrum, giving an incorrect helicity. We calculated the effect of Tyr on the CD at 222 nm by deriving the intensities of the bands directly from the electronic and magnetic transition dipole moments through the rotational strengths corresponding to each excited state of the polypeptide. This gives an improved value of the helix preference of Tyr (from 0.48 to 0.35) and a correction to the helicity for the peptides containing Tyr. We find that Phe-Lys, Lys-Phe, Phe-Arg, Arg-Phe, and Tyr-Lys are all stabilizing by -0.10 to -0.18 kcal.mol-1 when placed i, i + 4 on the surface of a helix in aqueous solution, despite the great difference in polarity between these residues. Interactions between these side chains have previously been attributed to cation-pi bonds. A survey of protein structures shows that they are in fact predominantly hydrophobic interactions between the CH2 groups of Lys or Arg and the aromatic rings.

Amino Acid Sequence↗

Information-theoretic analysis of protein sequences shows that amino acids self-cluster.

We analyse for each of 20 amino acids X the statistics of spacings between consecutive occurrences of X within the well-characterized Saccharomyces cerevisiae genome. The occurrences of amino acids may exhibit near random, clustered or smoothed out behaviour, like one-dimensional stochastic processes along the protein chain. If amino acids are distributed randomly within a sequence, then they follow a Poisson process, and a histogram of the number of observations of each gap size would asymptotically follow a negative exponential distribution. The novelty of the present approach lies in the use of differential geometric methods to quantify information on sequencing of amino acids and groups of amino acids, via the sequences of intervals between their occurrences. The differential geometry arises from an information-theoretic distance function on the two-dimensional space of stochastic processes subordinate to gamma distributions-which latter include the random process as a special case. We find that maximum-likelihood estimates of parametric statistics show that all 20 amino acids tend to cluster, some substantially. In other words, the frequencies of short gap lengths tend to be higher and the variance of the gap lengths is greater than expected by chance. This may be because localizing amino acids with the same properties may favour secondary structure formation or transmembrane domains. Gap sizes of 1 or 2 are generally disfavoured, 1 strongly so. The only exceptions to this are Gln and Ser, as a result of poly(Gln) or poly(Ser) sequences. There are preferences for gaps of 4 and 7 that can be attributed to alpha -helices. In particular, a favoured gap of 7 for Leu is found in coiled coils. Our method contributes to the characterization of whole sequences by extracting and quantifying stable stochastic features.

Amino Acids↗

Effect of phosphorylation on alpha-helix stability as a function of position.

We have investigated the effect of placing phosphoserine at the N-cap, N1, N2, N3, and interior position in alanine-based alpha-helical peptides. Helix contents of each peptide were measured by CD spectroscopy and titrations performed to determine pK(a) values. Data were analyzed with modified Lifson-Roig theory to determine helix-coil parameters (n, n(1), n(2), n(3), and w) and free energy changes for phosphoserine at each helical position. Results are given for a -1 and -2 phosphoserine charge state. Results show that phosphoserine stabilizes at the N-terminal positions by as much as 2.3 kcal.mol(-1), while destabilizes in the helix interior by 1.2 kcal.mol(-1), relative to serine. The rank order of free energies relative to serine at each position is N2 > N3 > N1 > N-cap > interior. Moreover, -2 phosphoserine is the most preferred residue known at each of these N-terminal positions. Experimental pK(a) values for the -1 to -2 phosphoserine transition are in the order N2 < N-cap < N1 < N3 < interior. This order agrees well with electrostatics calculations carried out with phosphoserine at the N-terminal positions and interior positions. Combining these with calculations at the C3, C2, C1, and C-cap positions gives results for phosphoserine along the length of the helix. We see a transition from phosphoserine stabilization at the N-terminus to destabilization at the C-terminus and can explain this in terms of the balance of protein solvation, favorable interactions, and dehydration. These results give insight into the phosphorylatable control of biological systems through positive or negative changes in stability.

Amino Acid Sequence↗

A critical assessment of the secondary structure alpha-helices and their termini in proteins.

Secondary structure prediction from amino acid sequence is a key component of protein structure prediction, with current accuracy at approximately 75%. We analysed two state-of-the-art secondary structure prediction methods, PHD and JPRED, comparing predictions with secondary structure assigned by the algorithms DSSP and STRIDE. The specific focus of our study was alpha-helix N-termini, as empirical free energy scales are available for residue preferences at N-terminal positions. Although these prediction methods perform well in general at predicting the alpha-helical locations and length distributions in proteins, they perform less well at predicting the correct helical termini. For example, although most predicted alpha-helices overlap a real alpha-helix (with relatively few completely missed or extra predicted helices), only one-third of JPRED and PHD predictions correctly identify the N-terminus. Analysis of neighbouring N-terminal sequences to predicted helical N-termini shows that the correct N-terminus is often within one or two residues. More importantly, the true N-terminal motif is, on average, more favourable as judged by our experimentally measured free energies. This suggests a simple, but powerful, strategy to improve secondary structure prediction using empirically derived energies to adjust the predicted output to a more favourable N-terminal sequence.

Algorithms↗