PubMed Health⌕ Search

Biomedical subjects

S D Rufino

Publications and source records attributed to S D Rufino.

6 recordsLinked to original sources

Predicting the conformational class of short and medium size loops connecting regular secondary structures: application to comparative modelling.

Loops are regions of non-repetitive conformation connecting regular secondary structures. They are both the most difficult and error prone regions of a protein to solve by X-ray crystallography and the hardest regions to model using comparative procedures. Although a loop can sometimes be modelled from a homologue, very often it must be selected from outside the family. The loop prediction procedure, SLoop, attempts to identify the conformational class of the loop rather than to select a specific loop from a set of fragments extracted from known structures or generated ab initio. Templates are constructed for each of the 161 loop conformational classes that have been identified from the clustering of the structures of some 2024 loops of one to eight residues in length. A class template describes both sequence preferences and relative disposition of bounding secondary structures. During comparative modelling, the conformation of a loop can be predicted by identifying a loop class with which its sequence and disposition of bounding secondary structures are compatible. The procedure is tested on an unrelated non-redundant set of 1785 loops under stringent and lax evaluation schemes. Optimal sequence score cut-offs are identified such that the prediction rate is equal to the percentage of loops assigned to acceptable classes. Under the stringent evaluation, at the optimal sequence score cut-off, a conformation is predicted for 50% of loops of which 47% are correct, while under the lax evaluation a conformation is predicted for 63% of loops of which 54% are correct. Sequence score is shown to be a good indicator of the probability of a prediction being correct. Loop length also has a strong affect on prediction outcomes. Considering only loops of two to five residues in length, under the stringent evaluation 62% of loops are predicted with 52% of these predictions being correct while under the lax evaluation predictions are provided for 75% of loops of which 57% are correct.

Computer Simulation↗

Conformational analysis and clustering of short and medium size loops connecting regular secondary structures: a database for modeling and prediction.

Loops are regions of nonrepetitive conformation connecting regular secondary structures. We identified 2,024 loops of one to eight residues in length, with acceptable main-chain bond lengths and peptide bond angles, from a database of 223 protein and protein-domain structures. Each loop is characterized by its sequence, main-chain conformation, and relative disposition of its bounding secondary structures as described by the separation between the tips of their axes and the angle between them. Loops, grouped according to their length and type of their bounding secondary structures, were superposed and clustered into 161 conformational classes, corresponding to 63% of all loops. Of these, 109 (51% of the loops) were populated by at least four nonhomologous loops or four loops sharing a low sequence identity. Another 52 classes, including 12% of the loops, were populated by at least three loops of low sequence similarity from three or fewer nonhomologous groups. Loop class suprafamilies resulting from variations in the termini of secondary structures are discussed in this article. Most previously described loop conformations were found among the classes. New classes included a 2:4 type IV hairpin, a helix-capping loop, and a loop that mediates dinucleotide-binding. The relative disposition of bounding secondary structures varies among loop classes, with some classes such as beta-hairpins being very restrictive. For each class, sequence preferences as key residues were identified; those most frequently at these conserved positions than in proteins were Gly, Asp, Pro, Phe, and Cys. Most of these residues are involved in stabilizing loop conformation, often through a positive phi conformation or secondary structure capping. Identification of helix-capping residues and beta-breakers among the highly conserved positions supported our decision to group loops according to their bounding secondary structures. Several of the identified loop classes were associated with specific functions, and all of the member loops had the same function; key residues were conserved for this purpose, as is the case for the parvalbumin-like calcium-binding loops. A significant number, but not all, of the member loops of other loop classes had the same function, as is the case for the helix-turn-helix DNA-binding loops. This article provides a systematic and coherent conformational classification of loops, covering a broad range of lengths and all four combinations of bounding secondary structure types, and supplies a useful basis for modelling of loop conformations where the bounding secondary structures are known or reliably predicted.

Databases, Factual↗

A database of globular protein structural domains: clustering of representative family members into similar folds.

BACKGROUND: A database of globular domains, derived from a non-redundant set of proteins, is useful for the sequence analysis of aligned domains, for structural comparisons, for understanding domain stability and flexibility and for fold recognition procedures. Domains are defined by the program DIAL and classified structurally using the procedure SEA. RESULTS: The DIAL-derived domain database (DDBASE) consists of 436 protein chains involving 695 protein domains. Of these, 206 are alpha-class, 191 are beta-class and 294 alpha and beta class. The domains, 63% from multidomain proteins and 73% less than 150 residues in length, were clustered automatically using both single-link cluster analysis and hierarchical clustering to give a quantitative estimate of similarity in the domain-fold space. CONCLUSIONS: Highly populated and well described folds (doubly wound alpha/beta, singly wound alpha/beta barrels, globins alpha, large Greek-key beta and flavin-binding alpha/beta) are recognized at a SEA cut-off score of 0.55 in single-link clustering and at 0.65 in hierarchical clustering, although functionally related families are usually clearly distinguished at more stringent values.

Cluster Analysis↗

Analysis, clustering and prediction of the conformation of short and medium size loops connecting regular secondary structures.

Loops are regions of non-repetitive conformation connecting regular secondary structures. They are both the most difficult and error prone regions of a protein to solve by X-ray crystallography and the hardest regions to model using knowledge-based procedures. While the core of a protein can be straight forwardly modelled from the structurally conserved regions of homologues of known structure, loops must be modelled from a selected homologue or from a loop chosen from outside the family. Here we present a loop prediction procedure that attempts to identify the conformational class of the loop rather than to select a specific loop from a database of fragments. The structures of some 2083 loops of one to eight residues in length were extracted from a database of 225 protein and protein domain structures. For each loop, the relative disposition of its bounding secondary structures is described by the separation between the tips of their axes, the angle and dihedral angle between their axes. From the clustering of the loops according to the root mean square deviation of their spatial fit, a total of 162 loop conformational classes, including 79% of loops, were identified. One-hundred and eight of these, involving 66% of the loops, were populated by at least four non-homologous loops or four loops sharing a low sequence identity. Another 54 classes, including 13% of the loops, were populated by at least three loops of low sequence similarity from three or fewer non-homologous groups. Most of the previously described loop conformations were found among the populated classes. For each class a template was constructed containing both sequence preferences and the relative disposition of bounding secondary structures among member loops. During comparative modelling, the conformation of a loop can be predicted by identifying a loop class with which its sequence and disposition of bounding secondary structures are compatible.

Amino Acid Sequence↗

The recognition of protein structure and function from sequence: adding value to genome data.

The explosion of DNA sequence data from genome projects presents many challenges. For instance, we must extend our current knowledge of protein structure and function so that it can be applied to these new sequences. The derivation of rules for the relationships between sequence and structure allow us to recognize a common fold by the use of tertiary templates. New techniques enable us to begin to meet the challenge of rule-based modelling of distantly related proteins. This paper describes an integrated and knowledge-based approach to the prediction of protein structure and function which can maximize the value of sequence information.

Amino Acid Sequence↗

Structure-based identification and clustering of protein families and superfamilies.

We describe an approach to protein structure comparison designed to detect distantly related proteins of similar fold, where the procedure must be sufficiently flexible to take into account the elasticity of protein folds without losing specificity. Protein structures are represented as a series of secondary structure elements, where for each element a local environment describes its relations with the elements that surround it. Secondary structures are then aligned by comparing their features and local environments. The procedure is illustrated with searches of a database of 468 protein structures in order to identify proteins of similar topology to porcine pepsin, porphobilinogen deaminase and serum amyloid P-component. In all cases the searches correctly identify protein structures of similar fold as the search proteins. Multiple cross-comparisons of protein structures allow the clustering of proteins of similar fold. This is exemplified with a clustering of alpha/beta- and beta-class protein structures. We discuss applications of the comparison and clustering of three-dimensional protein structures to comparative modelling and structure-based protein design.

Algorithms↗