PubMed Health⌕ Search

Biomedical subjects

P Tufféry

Publications and source records attributed to P Tufféry.

11 recordsLinked to original sources

Hidden Markov model-derived structural alphabet for proteins: the learning of protein local shapes captures sequence specificity.

Understanding and predicting protein structures depend on the complexity and the accuracy of the models used to represent them. We have recently set up a Hidden Markov Model to optimally compress protein three-dimensional conformations into a one-dimensional series of letters of a structural alphabet. Such a model learns simultaneously the shape of representative structural letters describing the local conformation and the logic of their connections, i.e. the transition matrix between the letters. Here, we move one step further and report some evidence that such a model of protein local architecture also captures some accurate amino acid features. All the letters have specific and distinct amino acid distributions. Moreover, we show that words of amino acids can have significant propensities for some letters. Perspectives point towards the prediction of the series of letters describing the structure of a protein from its amino acid sequence.

Amino Acid Sequence↗

RPBS: a web resource for structural bioinformatics.

RPBS (Ressource Parisienne en Bioinformatique Structurale) is a resource dedicated primarily to structural bioinformatics. It is the result of a joint effort by several teams to set up an interface that offers original and powerful methods in the field. As an illustration, we focus here on three such methods uniquely available at RPBS: AUTOMAT for sequence databank scanning, YAKUSA for structure databank scanning and WLOOP for homology loop modelling. The RPBS server can be accessed at http://bioserv.rpbs.jussieu.fr/ and the specific services at http://bioserv.rpbs.jussieu.fr/SpecificServices.html.

Computational Biology↗

SCit: web tools for protein side chain conformation analysis.

SCit is a web server providing services for protein side chain conformation analysis and side chain positioning. Specific services use the dependence of the side chain conformations on the local backbone conformation, which is described using a structural alphabet that describes the conformation of fragments of four-residue length in a limited library of structural prototypes. Based on this concept, SCit uses sets of rotameric conformations dependent on the local backbone conformation of each protein for side chain positioning and the identification of side chains with unlikely conformations. The SCit web server is accessible at http://bioserv.rpbs.jussieu.fr/SCit.

Amino Acids↗

A hidden markov model derived structural alphabet for proteins.

Understanding and predicting protein structures depends on the complexity and the accuracy of the models used to represent them. We have set up a hidden Markov model that discretizes protein backbone conformation as series of overlapping fragments (states) of four residues length. This approach learns simultaneously the geometry of the states and their connections. We obtain, using a statistical criterion, an optimal systematic decomposition of the conformational variability of the protein peptidic chain in 27 states with strong connection logic. This result is stable over different protein sets. Our model fits well the previous knowledge related to protein architecture organisation and seems able to grab some subtle details of protein organisation, such as helix sub-level organisation schemes. Taking into account the dependence between the states results in a description of local protein structure of low complexity. On an average, the model makes use of only 8.3 states among 27 to describe each position of a protein structure. Although we use short fragments, the learning process on entire protein conformations captures the logic of the assembly on a larger scale. Using such a model, the structure of proteins can be reconstructed with an average accuracy close to 1.1A root-mean-square deviation and for a complexity of only 3. Finally, we also observe that sequence specificity increases with the number of states of the structural alphabet. Such models can constitute a very relevant approach to the analysis of protein architecture in particular for protein structure prediction.

Algorithms↗

Critical assessment of side-chain conformational space sampling procedures designed for quantifying the effect of side-chain environment.

We introduce a family of procedures designed to sample side-chain conformational space at particular locations in protein structures. These procedures (CRSP) use intensive cycles of random assignment of side-chain conformations followed by minimization to determine all the conformations that a group of side-chains can adopt simultaneously. First, we consider a procedure evolving in the dihedral space (dCRSP). Our results suggest that it can accurately map low-energy conformations adopted by clusters of side-chains of a protein. dCRSP is relatively insensitive to various important parameters, and it is sufficiently accurate to capture efficiently the constraint induced by the environment on the conformations a particular side-chain can adopt. Our results show that dCRSP, compared with molecular dynamics (MD), can overcome the problem of the limited set of conformations reached in a reasonable amount of simulations. Next, we introduce procedures (vCRSP) in which valence angles are relaxed, and we assess how efficiently they quantify the conformational entropy of side-chains in the protein native state. For simple peptides, entropies obtained with vCRSP are fully compatible with those obtained with a Monte Carlo procedure. For side-chains in a protein environment, however, vCRSP appears of limited use. Finally, we consider a two-step procedure that combines dCRSP and vCRSP. Our tests suggest that it is able to overcome the limitations of vCRSP. We also note that dCRSP provides a reasonable initial approximation. This family of procedures offers promise in quantifying the contribution of conformational entropy to the energetics of protein structures.

Algorithms↗

Predicting the disulfide bonding state of cysteines using protein descriptors.

Knowledge of the disulfide bonding state of the cysteines of proteins is of major interest in designing numerous molecular biology experiments, or in predicting their three-dimensional structure. Previous methods using the information gained from aligned sets of sequences have reached up to 82% of success in predicting the oxidation state of cysteines. In the present study, we assess the relative efficiency of different descriptors in predicting the cysteine disulfide bonding states. Our results suggest that the information on the residues flanking the cysteines is less informative about the disulfide bonding state than about the amino acid content of the whole protein. Using a combination of logistic functions learned with subsets of proteins homogeneous in terms of their amino acid content, we propose a simple prediction approach, starting from a single sequence, that reaches success rates close to 84%. This score can be improved by avoiding predictions regarding cysteines for which the decision is not well marked. For example, we obtain a score close to 87% correct prediction when we exclude predicting 10% of the cysteines.

Amino Acids↗

CS-PSeq-Gen: simulating the evolution of protein sequence under constraints.

UNLABELLED: CS-PSeq-Gen is a program derived from PSeq-Gen, designed to perform simulations of the evolution of protein sequences under the constraints of a reconstructed phylogeny. It also provides a basis for the investigation of the correlated evolution of sites. AVAILABILITY: http://condor.urbb.jussieu.fr/CS-PSeq-Gen.html

Amino Acid Sequence↗

PredAcc: prediction of solvent accessibility.

UNLABELLED: PredAcc is a tool for predicting the solvent accessibility of protein residues from the sequence at different relative accessibility levels (0-55%). The prediction rate varies between 70. 7% (for 25% relative accessibility) and 85.7% (for 0% relative accessibility). Amino acids are predicted in four categories: almost certainly hidden and almost certainly exposed with a given a posteriori prediction error, probably hidden and probably exposed otherwise. AVAILABILITY: http://condor.urbb.jussieu.fr/PredAccCfg.html CONTACT: tuffery@urbb.jussieu.fr

Amino Acids↗

Prediction of protein side chain conformations: a study on the influence of backbone accuracy on conformation stability in the rotamer space.

We have studied the effect of backbone inaccuracy on the efficiency of protein side chain conformation prediction using rotamer libraries. The backbones were generated by randomly perturbing the crystallographic conformation of 12 proteins and exhibit C alpha r.m.s.d.s of up to 2 A. Our results show that, even for a perturbation of the backbone fully compatible with the temperature factors of the proteins, the predicted side chain conformations of approximately 10% of the buried side chains remain variable. This fraction increases further for larger backbone deviations. However, for backbone deviations of up to 2 A r.m.s.d., the predicted side chain r.m.s.d. varies only in a ratio of < 1.4. Moreover, a possible strategy for obtaining side chain conformations close to the experimental ones consists of extracting the consensus conformations of the side chains from a series of backbone conformations. Such a procedure allows the computation of the side chain conformations with no loss of accuracy for backbones exhibiting r.m.s.d.s of up to 1 A from the crystallographic coordinates. For larger backbone deviations (up to 2 A r.m.s.d.) the r.m.s.d. of the buried side chains increases from 1.33 up to 1.60 A. We also discuss the influence of the size of the rotamer library on the quality of the prediction.

Crystallography, X-Ray↗

XmMol: an X11 and motif program for macromolecular visualization and modeling.

XmMol is a desktop tool designed to provide both interactive molecular graphics on X11 displays and easy interface with external applications. A kernel provides an interactive wire-frame display of macromolecules. It supports depth cueing, 3D clipping, and stereo. Various representations, coloring, and labeling modes are proposed. Docking and interactive backbone deformation tools are also supported. Communication protocols allow the user to develop new external features or to use XmMol as a visualization tool for external numerical programs.

Computer Graphics↗

Packing and recognition of protein structural elements: a new approach applied to the 4-helix bundle of myohemerythrin.

We present a novel search strategy for determining the optimal packing of protein secondary structure elements. The approach is based on conformational energy optimization using a predetermined set of side chain rotamers and appropriate methods for sampling the conformational space of peptide fragments having fixed backbone geometries. An application to the 4-helix bundle of myohemerythrin is presented. It is shown that the conformations of the amino acid side chains are largely determined at the level of helix pairs and that superposition of these results can be used to construct the full bundle. The final solution obtained, taking into account restrictions due to the lateral amphiphilicity of the helices, differs from the native structure by only a 20 degrees rotation of a single helix.

Amino Acid Sequence↗