Percent sequence identity; the need to be explicit.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Alex C W May.
Explore the source record for details and available documents.
The iron- and manganese-containing superoxide dismutases (Fe/Mn-SOD) share the same chemical function and spatial structure but can be distinguished according to their modes of oligomerization and their metal ion specificity. They appear as homodimers or homotetramers and usually require a specific metal for activity. On the basis of 261 aligned SOD sequences and 12 superimposed x-ray structures, two phenetic trees were constructed, one sequence-based and the other structure-based. Their comparison reveals the imperfect correlation of sequence and structural changes; hyperthermophilicity requires the largest sequence alterations, whereas dimer/tetramer and manganese/iron specificities are induced by the most sizable structural differences within the monomers. A systematic investigation of sequence and structure characteristics conserved in all aligned SOD sequences or in subsets sharing common oligomeric and/or metal specificities was performed. Several residues were identified as guaranteeing the common function and dimeric conformation, others as determining the tetramer formation, and yet others as potentially responsible for metal specificity. Some form cation-pi interactions between an aromatic ring and a fully or partially positively charged group, suggesting that these interactions play a significant role in the structure and function of SOD enzymes. Dimer/tetramer- and iron/manganese-specific fingerprints were derived from the set of conserved residues; they can be used to propose selected residue substitutions in view of the experimental validation of our in silico derived hypotheses.
Bioinformatic software has used various numerical encoding schemes to describe amino acid sequences. Orthogonal encoding, employing 20 numbers to describe the amino acid type of one protein residue, is often used with artificial neural network (ANN) models. However, this can increase the model complexity, thus leading to difficulty in implementation and poor performance. Here, we use ANNs to derive encoding schemes for the amino acid types from protein three-dimensional structure alignments. Each of the 20 amino acid types is characterized with a few real numbers. Our schemes are tested on the simulation of amino acid substitution matrices. These simplified schemes outperform the orthogonal encoding on small data sets. Using one of these encoding schemes, we generate a colouring scheme for the amino acids in which comparable amino acids are in similar colours. We expect it to be useful for visual inspection and manual editing of protein multiple sequence alignments.
MOTIVATION: Fold recognition programs align a probe protein sequence onto protein three-dimensional (3D) structure templates. The alignment between the probe sequence and the most suitable template can be used to predict the 3D structure and often biological function of the probe. Here we present a new threading scoring function of protein sequence-structure compatibility. An artificial neural network model is trained to predict compatibility of amino acid side-chains with structural environments. Log-odds scores of predicted probabilities from this model can then be used to construct protein sequence-structure alignments. RESULTS: Our model is tested on discrimination of native and decoy protein 3D structures. With a residue level structural description, its performance is comparable to those of pseudo-energy functions with atom level structural descriptions, better than the two functions with residue level structural descriptions. AVAILABILITY: The C++ source code of our neural network model is available at http://mathbio.nimr.mrc.ac.uk/~kxlin.
It is often possible to identify sequence motifs that characterize a protein family in terms of its fold and/or function from aligned protein sequences. Such motifs can be used to search for new family members. Partitioning of sequence alignments into regions of similar amino acid variability is usually done by hand. Here, I present a completely automatic method for this purpose: one that is guaranteed to produce globally optimal solutions at all levels of partition granularity. The method is used to compare the tempo of sequence diversity across reliable three-dimensional (3D) structure-based alignments of 209 protein families (HOMSTRAD) and that for 69 superfamilies (CAMPASS). (The mean alignment length for HOMSTRAD and CAMPASS are very similar.) Surprisingly, the optimal segmentation distributions for the closely related proteins and distantly related ones are found to be very similar. Also, optimal segmentation identifies an unusual protein superfamily. Finally, protein 3D structure clues from the tempo of sequence diversity across alignments are examined. The method is general, and could be applied to any area of comparative biological sequence and 3D structure analysis where the constraint of the inherent linear organization of the data imposes an ordering on the set of objects to be clustered.