PubMed Health⌕ Search

Biomedical subjects

R Casadio

Publications and source records attributed to R Casadio.

At least 19 recordsLinked to original sources

K-Fold: a tool for the prediction of the protein folding kinetic order and rate.

UNLABELLED: K-Fold is a tool for the automatic prediction of the protein folding kinetic order and rate. The tool is based on a support vector machine (SVM) that was trained on a data set of 63 proteins, whose 3D structure and folding mechanism are known from experiments already described in the literature. The method predicts whether a protein of known atomic structure folds according to a two-state or a multi-state kinetics and correctly classifies 81% of the folding mechanisms when tested over the training set of the 63 proteins. It also predicts as a further option the logarithm of the folding rate. To the best of our knowledge, the tool discriminates for the first time whether a protein is characterized by a two state or a multiple state kinetics, during the folding process, and concomitantly estimates also the value of the constant rate of the process. When used to predict the logarithm of the folding rate, K-Fold scores with a correlation value to the experimental data of 0.74 (with a SE of 1.2). AVAILABILITY: http://gpcr.biocomp.unibo.it/cgi/predictors/K-Fold/K-Fold.cgi. SUPPLEMENTARY INFORMATION: http://gpcr.biocomp.unibo.it/~emidio/K-Fold/K-Fold_help.html.

Algorithms↗

Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.

MOTIVATION: Human single nucleotide polymorphisms (SNPs) are the most frequent type of genetic variation in human population. One of the most important goals of SNP projects is to understand which human genotype variations are related to Mendelian and complex diseases. Great interest is focused on non-synonymous coding SNPs (nsSNPs) that are responsible of protein single point mutation. nsSNPs can be neutral or disease associated. It is known that the mutation of only one residue in a protein sequence can be related to a number of pathological conditions of dramatic social impact such as Alzheimer's, Parkinson's and Creutzfeldt-Jakob's diseases. The quality and completeness of presently available SNPs databases allows the application of machine learning techniques to predict the insurgence of human diseases due to single point protein mutation starting from the protein sequence. RESULTS: In this paper, we develop a method based on support vector machines (SVMs) that starting from the protein sequence information can predict whether a new phenotype derived from a nsSNP can be related to a genetic disease in humans. Using a dataset of 21 185 single point mutations, 61% of which are disease-related, out of 3587 proteins, we show that our predictor can reach more than 74% accuracy in the specific task of predicting whether a single point mutation can be disease related or not. Our method, although based on less information, outperforms other web-available predictors implementing different approaches. AVAILABILITY: A beta version of the web tool is available at http://gpcr.biocomp.unibo.it/cgi/predictors/PhD-SNP/PhD-SNP.cgi

Algorithms↗

A novel RGDS-analog inhibits angiogenesis in vitro and in vivo.

In this study the anti-angiogenic action of a novel non-peptide RGDS-analog named RAM was tested in vitro and in vivo. RAM inhibited FGF-2-induced chemotaxis by 80% in an adhesion-independent way. Further, it induced HUVEC-apoptosis in collagen-seeded HUVEC, indicating that such pro-apoptotic effect was adhesion-independent. In vivo studies revealed that RAM inhibited FGF-2 induced angiogenesis by 60% in the mouse Matrigel-assay and in the chicken-egg chorion-allantoic membrane assay. Finally, RAM was markedly more stable in serum as compared to the template RGDS and after 24 h incubation in 100% serum was significantly more active than RGDS. Taken together these results show that RAM exerts anti-chemotactic and pro-apoptotic effects, by an unexpected adhesion-independent mechanism, as we have recently shown for the template RGDS molecule [Blood 103 (2004) 4180], and has in vivo relevant anti-angiogenic properties, with marked stability in serum; therefore, RAM represents a novel promising anti-angiogenic molecule.

Angiogenesis Inhibitors↗

Dynamics of the minimally frustrated helices determine the hierarchical folding of small helical proteins.

In this paper we aim at determining the key residues of small helical proteins in order to build up reduced models of the folding dynamics. We start by arguing that the folding process can be dissected into concurrent fast and slow dynamics. The fast events are the quasiautonomous coil-to-helix transitions occurring in the minimally frustrated initiation sites of folding in the early stages of the process. The slow processes consist in the docking of the fluctuating helices formed in these critical sites. We show that a neural network devised to predict native secondary structures from sequence can be used to estimate the probabilities of formation of these helical traits as they are embedded in the protein. The resulting probabilities are shown to correlate well with the experimental helicities measured in the same isolated peptides. The relevance of this finding to the hierarchical character of folding is confirmed within the framework of a diffusion-collision-like mechanism. We demonstrate that thermodynamic and topological features of these critical helices allow accurate estimation of the folding times of five proteins that have been kinetically studied. This suggests that these critical helices determine the fundamental events of the whole folding process. A remarkable feature of our model is that not all of the native helices are eligible as critical helices, whereas the whole set of the native helices has been used so far in other reconstructions of the folding mechanism. This stresses that the minimally frustrated helices of these helical proteins comprise the minimal set of determinants of the folding process.

Binding Sites↗

A neural network approach to evaluate fold recognition results.

Fold recognition techniques assist the exploration of protein structures, and web-based servers are part of the standard set of tools used in the analysis of biochemical problems. Despite their success, current methods are only able to predict the correct fold in a relatively small number of cases. We propose an approach that improves the selection of correct folds from among the results of two methods implemented as web servers (SAMT99 and 3DPSSM). Our approach is based on the training of a system of neural networks with models generated by the servers and a set of associated characteristics such as the quality of the sequence-structure alignment, distribution of sequence features (sequence-conserved positions and apolar residues), and compactness of the resulting models. Our results show that it is possible to detect adequate folds to model 80% of the sequences with a high level of confidence. The improvements achieved by taking into account sequence characteristics open the door to future improvements by directly including such factors in the step of model generation. This approach has been implemented as an automatic system LIBELLULA, available as a public web server at http://www.pdg.cnb.uam.es/servers/libellula.html.

Internet↗

Characterisation of immunodominant protein encoded by the F1L gene of orf virus strains isolated in Italy.

We analysed the molecular properties of the immunodominant protein of different orf virus strains isolated in Italy. The F1L encoding genes and the deduced amino acid sequences of all strains were determined and compared, and they showed several mutations. Structural analysis was carried out in order to assess the influence of amino acid variations on protein structure demonstrating a conservation of the secondary structure. Western blot analysis and immunogold electron microscopy showed that all orf virus strains were antigenically identical. The results of our study confirmed the immunogenicity of the F1L protein; furthermore, our data suggest a possible involvement of the protein in the virus cycle.

Amino Acid Sequence↗

Model of interaction of the IL-1 receptor accessory protein IL-1RAcP with the IL-1beta/IL-1R(I) complex.

A preliminary model has been calculated for the activating interaction of the interleukin 1 receptor (IL-1R) accessory protein IL-1RAcP with the ligand/receptor complex IL-1beta/IL-1R(I). First, IL-1RAcP was modeled on the crystal structure of IL-1R(I) bound to IL-1beta. Then, the IL-1RAcP model was docked using specific programs to the crystal structure of the IL-1beta/IL-1R(I) complex. Two types of models were predicted, with comparable probability. Experimental data obtained with the use of IL-1beta peptides and antibodies, and with mutated IL-1beta proteins, support the BACK model, in which IL-1RAcP establishes contacts with the back of IL-1R(I) wrapped around IL-1beta.

Animals↗

Progress in predicting inter-residue contacts of proteins with neural networks and correlated mutations.

This article presents recent progress in predicting inter-residue contacts of proteins with a neural network-based method. Improvement over the results obtained at the previous CASP3 competition is attained by using as input to the network a complex code, which includes evolutionary information, sequence conservation, correlated mutations, and predicted secondary structures. The predictor was trained and cross-validated on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well-resolved three-dimensional structures. The method could assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor greater than 6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Although far from being ideal, these scores are the highest reported so far for predicting protein contact maps. On 29 targets automatically predicted by the server (CORNET) the average accuracy is 0.14. The predictor is poorly performing on all alpha proteins, not represented in the training set. On all beta and mixed proteins (22 targets) the average accuracy is 0.16. This set comprises proteins of different complexity and different chain length, suggesting that the predictor is capable of generalization over a broad number of features.

Mutation↗

Structure-based computational study of the catalytic and inhibition mechanisms of urease.

The viability of different mechanisms of catalysis and inhibition of the nickel-containing enzyme urease was explored using the available high-resolution structures of the enzyme isolated from Bacillus pasteurii in the native form and inhibited with several substrates. The structures and charge distribution of urea, its catalytic transition state, and three enzyme inhibitors were calculated using ab initio and density functional theory methods. The DOCK program suite was employed to determine families of structures of urease complexes characterized by docking energy scores indicative of their relative stability according to steric and electrostatic criteria. Adjustment of the parameters used by DOCK, in order to account for the presence of the metal ion in the active site, resulted in the calculation of best energy structures for the nickel-bound inhibitors beta-mercaptoethanol, acetohydroxamic acid, and diamidophosphoric acid. These calculated structures are in good agreement with the experimentally determined structures, and provide hints on the reactivity and mobility of the inhibitors in the active site. The same docking protocol was applied to the substrate urea and its catalytic transition state, in order to shed light onto the possible catalytic steps occurring at the binuclear nickel active site. These calculations suggest that the most viable pathway for urea hydrolysis involve a nucleophilic attack by the bridging, and not the terminal, nickel-bound hydroxide onto a urea molecule, with active site residues playing important roles in orienting and activating the substrate, and stabilizing the catalytic transition state.

Algorithms↗

Prediction of disulfide connectivity in proteins.

MOTIVATION: A major problem in protein structure prediction is the correct location of disulfide bridges in cysteine-rich proteins. In protein-folding prediction, the location of disulfide bridges can strongly reduce the search in the conformational space. Therefore the correct prediction of the disulfide connectivity starting from the protein residue sequence may also help in predicting its 3D structure. RESULTS: In this paper we equate the problem of predicting the disulfide connectivity in proteins to a problem of finding the graph matching with the maximum weight. The graph vertices are the residues of cysteine-forming disulfide bridges, and the weight edges are contact potentials. In order to solve this problem we develop and test different residue contact potentials. The best performing one, based on the Edmonds-Gabow algorithm and Monte-Carlo simulated annealing reaches an accuracy significantly higher than that obtained with a general mean force contact potential. Significantly, in the case of proteins with four disulfide bonds in the structure, the accuracy is 17 times higher than that of a random predictor. The method presented here can be used to locate putative disulfide bridges in protein-folding. AVAILABILITY: The program is available upon request from the authors. CONTACT: Casadio@alma.unibo.it; Piero@biocomp.unibo.it.

Algorithms↗

RCNPRED: prediction of the residue co-ordination numbers in proteins.

UNLABELLED: The RCNPRED server implements a neural network-based method to predict the co-ordination numbers of residues starting from the protein sequence. Using evolutionary information as input, RCNPRED predicts the residue states of the proteins in the database with 69% accuracy and scores 12 percentage points higher than a simple statistical method. Moreover the server implements a neural network to predict the relative solvent accessibility of each residue. A protein sequence can be directly submitted to RCNPRED: residue co-ordination numbers and solvent accessibility for each chain are returned via e-mail. AVAILABILITY: Freely available to non-commercial users at http://prion.biocomp.unibo.it/rcnpred.html.

Databases, Factual↗

Improved prediction of the number of residue contacts in proteins by recurrent neural networks.

Knowing the number of residue contacts in a protein is crucial for deriving constraints useful in modeling protein folding, protein structure, and/or scoring remote homology searches. Here we use an ensemble of bi-directional recurrent neural network architectures and evolutionary information to improve the state-of-the-art in contact prediction using a large corpus of curated data. The ensemble is used to discriminate between two different states of residue contacts, characterized by a contact number higher or lower than the average value of the residue distribution. The ensemble achieves performances ranging from 70.1% to 73.1% depending on the radius adopted to discriminate contacts (6Ato 12A). These performances represent gains of 15% to 20% over the base line statistical predictors always assigning an aminoacid to the most numerous state, 3% to 7% better than any previous method. Combination of different radius predictors further improves the performance. SERVER: http://promoter.ics.uci.edu/BRNN-PRED/.

Amino Acid Sequence↗

Prediction of contact maps with neural networks and correlated mutations.

Contact maps of proteins are predicted with neural network-based methods, using as input codings of increasing complexity including evolutionary information, sequence conservation, correlated mutations and predicted secondary structures. Neural networks are trained on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well resolved three-dimensional structures. Proteins are selected from the Protein Data Bank database provided that they align with at least 15 similar sequences in the corresponding families. The predictors are trained to learn the association rules between the covalent structure of each protein and its contact map with a standard back propagation algorithm and tested on the same protein set with a cross-validation procedure. Our results indicate that the method can assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor >6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Furthermore, filtering the network outputs with a procedure based on the residue coordination numbers, the accuracy of predictions increases up to 0.25 for all the proteins, with an 8-fold deviation from a random predictor. These scores are the highest reported so far for predicting protein contact maps.

Algorithms↗

Prediction of the transmembrane regions of beta-barrel membrane proteins with a neural network-based predictor.

A method based on neural networks is trained and tested on a nonredundant set of beta-barrel membrane proteins known at atomic resolution with a jackknife procedure. The method predicts the topography of transmembrane beta strands with residue accuracy as high as 78% when evolutionary information is used as input to the network. Of the transmembrane beta-strands included in the training set, 93% are correctly assigned. The predictor includes an algorithm of model optimization, based on dynamic programming, that correctly models eight out of the 11 proteins present in the training/testing set. In addition, protein topology is assigned on the basis of the location of the longest loops in the models. We propose this as a general method to fill the gap of the prediction of beta-barrel membrane proteins.

Algorithms↗

Predictions of protein segments with the same aminoacid sequence and different secondary structure: a benchmark for predictive methods.

The most stringent test for predictive methods of protein secondary structure is whether identical short sequences that are known to be present with different conformations in different proteins known at atomic resolution can be correctly discriminated. In this study, we show that the prediction efficiency of this type of segments in unrelated proteins reaches an average accuracy per residue ranging from about 72 to 75% (depending on the alignment method used to generate the input sequence profile) only when methods of the third generation are used. A comparison of different methods based on segment statistics (2nd generation methods) and/or including also evolutionary information (3rd generation methods) indicate that the discrimination of the different conformations of identical segments is dependent on the method used for the prediction. Accuracy is similar when methods similarly performing on the secondary structure prediction are tested. When evolutionary information is taken into account as compared to single sequence input, the number of correctly discriminated pairs is increased twofold. The results also highlight the predictive capability of neural networks for identical segments whose conformation differs in different proteins.

Algorithms↗

Ligand-induced conformational changes in tissue transglutaminase: Monte Carlo analysis of small-angle scattering data.

Small-angle neutron and x-ray scattering experiments have been performed on type 2 tissular transglutaminase to characterize the conformational changes that bring about Ca(2+) activation and guanosine triphosphate (GTP) inhibition. The native and a proteolyzed form of the enzyme, in the presence and in the absence of the two effectors, were considered. To describe the shape of transglutaminase in the different conformations, a Monte Carlo method for calculating small-angle neutron scattering profiles was developed by taking into account the computer-designed structure of the native transglutaminase, the results of the Guinier analysis, and the essential role played by the solvent-exposed peptide loop for the conformational changes of the protein after activation. Although the range of the neutron scattering data is rather limited, by using the Monte Carlo analysis, and because the structure of the native protein is available, the distribution of the protein conformations after ligand interaction was obtained. Calcium activation promotes a rotation of the C-terminal with respect to the N-terminal domain around the solvent-exposed peptide loop that connects the two regions. The psi angle between the longest axes of the two pairs of domains is found to be above 50 degrees, larger than the psi value of 35 degrees calculated for the native transglutaminase. On the other hand, the addition of GTP makes possible conformations characterized by psi angles lower than 34 degrees. These results are in good agreement with the proposed enzyme activity regulation: in the presence of GTP, the catalytic site is shielded by the more compact protein structure, while the conformational changes induced by Ca(2+) make the active site accessible to the substrate.

Calcium↗

The effect of tryptophanyl substitution on folding and structure of myoglobin.

Mammalian myoglobins contain two tryptophanyl residues at the invariant positions 7 (A-5) and 14 (A-12) in the N-terminal region (A helix) of the protein molecule. The crucial role of tryptophanyl residues has been investigated by site-directed mutagenesis and molecular dynamics simulation. The apomyoglobin mutants with a double W-->F substitution were found to be not correctly folded and therefore not expressed as holoprotein. The introduction of a tyrosyl residue at position 7, that is, W7YW14F, resulted in the expression of a correctly folded myoglobin. Not correctly folded apomyoglobins were found with the following mutants: W7FW14Y, W7EW14F, W7FW14E, W7KW14F, W7FW14K. Moreover, in all these cases, very low levels of expression were observed. The acid-induced denaturation curves of wild-type and folded mutant W7YW14F, obtained following the fluorescence variation of the extrinsic fluorophore 1-anilino-8-naphthalenesulfonate, revealed that the stability of the native state of mutant apoprotein is decreased, thus indicating that the replacement W-->Y in position 7 is able to restore a correct folding but not the same stability. Molecular dynamics simulation indicated that both tryptophans are involved in forming favorable, specific tertiary interactions in the native apomyoglobin structure. The lack of some of these interactions caused by tryptophanyl replacement affects the overall protein structure and may provide an explanation for the observed stability decrease. In the case of the double W-->F substitution, the simulated structure shows conclusively the domain formed by helices A, G and H to be not correctly folded. This effect is attenuated if at least one of the two residues is conserved or a tyrosyl residue replaces W7.

Myoglobin↗

Neural networks predict protein folding and structure: artificial intelligence faces biomolecular complexity.

In the genomic era DNA sequencing is increasing our knowledge of the molecular structure of genetic codes from bacteria to man at a hyperbolic rate. Billions of nucleotides and millions of aminoacids are already filling the electronic files of the data bases presently available, which contain a tremendous amount of information on the most biologically relevant macromolecules, such as DNA, RNA and proteins. The most urgent problem originates from the need to single out the relevant information amidst a wealth of general features. Intelligent tools are therefore needed to optimise the search. Data mining for sequence analysis in biotechnology has been substantially aided by the development of new powerful methods borrowed from the machine learning approach. In this paper we discuss the application of artificial feedforward neural networks to deal with some fundamental problems tied with the folding process and the structure-function relationship in proteins.

Databases, Factual↗