PubMed Health⌕ Search

Biomedical subjects

P Fariselli

Publications and source records attributed to P Fariselli.

At least 19 recordsLinked to original sources

A neural network approach to evaluate fold recognition results.

Fold recognition techniques assist the exploration of protein structures, and web-based servers are part of the standard set of tools used in the analysis of biochemical problems. Despite their success, current methods are only able to predict the correct fold in a relatively small number of cases. We propose an approach that improves the selection of correct folds from among the results of two methods implemented as web servers (SAMT99 and 3DPSSM). Our approach is based on the training of a system of neural networks with models generated by the servers and a set of associated characteristics such as the quality of the sequence-structure alignment, distribution of sequence features (sequence-conserved positions and apolar residues), and compactness of the resulting models. Our results show that it is possible to detect adequate folds to model 80% of the sequences with a high level of confidence. The improvements achieved by taking into account sequence characteristics open the door to future improvements by directly including such factors in the step of model generation. This approach has been implemented as an automatic system LIBELLULA, available as a public web server at http://www.pdg.cnb.uam.es/servers/libellula.html.

Internet↗

Progress in predicting inter-residue contacts of proteins with neural networks and correlated mutations.

This article presents recent progress in predicting inter-residue contacts of proteins with a neural network-based method. Improvement over the results obtained at the previous CASP3 competition is attained by using as input to the network a complex code, which includes evolutionary information, sequence conservation, correlated mutations, and predicted secondary structures. The predictor was trained and cross-validated on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well-resolved three-dimensional structures. The method could assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor greater than 6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Although far from being ideal, these scores are the highest reported so far for predicting protein contact maps. On 29 targets automatically predicted by the server (CORNET) the average accuracy is 0.14. The predictor is poorly performing on all alpha proteins, not represented in the training set. On all beta and mixed proteins (22 targets) the average accuracy is 0.16. This set comprises proteins of different complexity and different chain length, suggesting that the predictor is capable of generalization over a broad number of features.

Mutation↗

Prediction of disulfide connectivity in proteins.

MOTIVATION: A major problem in protein structure prediction is the correct location of disulfide bridges in cysteine-rich proteins. In protein-folding prediction, the location of disulfide bridges can strongly reduce the search in the conformational space. Therefore the correct prediction of the disulfide connectivity starting from the protein residue sequence may also help in predicting its 3D structure. RESULTS: In this paper we equate the problem of predicting the disulfide connectivity in proteins to a problem of finding the graph matching with the maximum weight. The graph vertices are the residues of cysteine-forming disulfide bridges, and the weight edges are contact potentials. In order to solve this problem we develop and test different residue contact potentials. The best performing one, based on the Edmonds-Gabow algorithm and Monte-Carlo simulated annealing reaches an accuracy significantly higher than that obtained with a general mean force contact potential. Significantly, in the case of proteins with four disulfide bonds in the structure, the accuracy is 17 times higher than that of a random predictor. The method presented here can be used to locate putative disulfide bridges in protein-folding. AVAILABILITY: The program is available upon request from the authors. CONTACT: Casadio@alma.unibo.it; Piero@biocomp.unibo.it.

Algorithms↗

RCNPRED: prediction of the residue co-ordination numbers in proteins.

UNLABELLED: The RCNPRED server implements a neural network-based method to predict the co-ordination numbers of residues starting from the protein sequence. Using evolutionary information as input, RCNPRED predicts the residue states of the proteins in the database with 69% accuracy and scores 12 percentage points higher than a simple statistical method. Moreover the server implements a neural network to predict the relative solvent accessibility of each residue. A protein sequence can be directly submitted to RCNPRED: residue co-ordination numbers and solvent accessibility for each chain are returned via e-mail. AVAILABILITY: Freely available to non-commercial users at http://prion.biocomp.unibo.it/rcnpred.html.

Databases, Factual↗

Improved prediction of the number of residue contacts in proteins by recurrent neural networks.

Knowing the number of residue contacts in a protein is crucial for deriving constraints useful in modeling protein folding, protein structure, and/or scoring remote homology searches. Here we use an ensemble of bi-directional recurrent neural network architectures and evolutionary information to improve the state-of-the-art in contact prediction using a large corpus of curated data. The ensemble is used to discriminate between two different states of residue contacts, characterized by a contact number higher or lower than the average value of the residue distribution. The ensemble achieves performances ranging from 70.1% to 73.1% depending on the radius adopted to discriminate contacts (6Ato 12A). These performances represent gains of 15% to 20% over the base line statistical predictors always assigning an aminoacid to the most numerous state, 3% to 7% better than any previous method. Combination of different radius predictors further improves the performance. SERVER: http://promoter.ics.uci.edu/BRNN-PRED/.

Amino Acid Sequence↗

Prediction of contact maps with neural networks and correlated mutations.

Contact maps of proteins are predicted with neural network-based methods, using as input codings of increasing complexity including evolutionary information, sequence conservation, correlated mutations and predicted secondary structures. Neural networks are trained on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well resolved three-dimensional structures. Proteins are selected from the Protein Data Bank database provided that they align with at least 15 similar sequences in the corresponding families. The predictors are trained to learn the association rules between the covalent structure of each protein and its contact map with a standard back propagation algorithm and tested on the same protein set with a cross-validation procedure. Our results indicate that the method can assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor >6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Furthermore, filtering the network outputs with a procedure based on the residue coordination numbers, the accuracy of predictions increases up to 0.25 for all the proteins, with an 8-fold deviation from a random predictor. These scores are the highest reported so far for predicting protein contact maps.

Algorithms↗

Prediction of the transmembrane regions of beta-barrel membrane proteins with a neural network-based predictor.

A method based on neural networks is trained and tested on a nonredundant set of beta-barrel membrane proteins known at atomic resolution with a jackknife procedure. The method predicts the topography of transmembrane beta strands with residue accuracy as high as 78% when evolutionary information is used as input to the network. Of the transmembrane beta-strands included in the training set, 93% are correctly assigned. The predictor includes an algorithm of model optimization, based on dynamic programming, that correctly models eight out of the 11 proteins present in the training/testing set. In addition, protein topology is assigned on the basis of the location of the longest loops in the models. We propose this as a general method to fill the gap of the prediction of beta-barrel membrane proteins.

Algorithms↗

Predictions of protein segments with the same aminoacid sequence and different secondary structure: a benchmark for predictive methods.

The most stringent test for predictive methods of protein secondary structure is whether identical short sequences that are known to be present with different conformations in different proteins known at atomic resolution can be correctly discriminated. In this study, we show that the prediction efficiency of this type of segments in unrelated proteins reaches an average accuracy per residue ranging from about 72 to 75% (depending on the alignment method used to generate the input sequence profile) only when methods of the third generation are used. A comparison of different methods based on segment statistics (2nd generation methods) and/or including also evolutionary information (3rd generation methods) indicate that the discrimination of the different conformations of identical segments is dependent on the method used for the prediction. Accuracy is similar when methods similarly performing on the secondary structure prediction are tested. When evolutionary information is taken into account as compared to single sequence input, the number of correctly discriminated pairs is increased twofold. The results also highlight the predictive capability of neural networks for identical segments whose conformation differs in different proteins.

Algorithms↗

Neural networks predict protein folding and structure: artificial intelligence faces biomolecular complexity.

In the genomic era DNA sequencing is increasing our knowledge of the molecular structure of genetic codes from bacteria to man at a hyperbolic rate. Billions of nucleotides and millions of aminoacids are already filling the electronic files of the data bases presently available, which contain a tremendous amount of information on the most biologically relevant macromolecules, such as DNA, RNA and proteins. The most urgent problem originates from the need to single out the relevant information amidst a wealth of general features. Intelligent tools are therefore needed to optimise the search. Data mining for sequence analysis in biotechnology has been substantially aided by the development of new powerful methods borrowed from the machine learning approach. In this paper we discuss the application of artificial feedforward neural networks to deal with some fundamental problems tied with the folding process and the structure-function relationship in proteins.

Databases, Factual↗

Prediction of the number of residue contacts in proteins.

Knowing the number of residue contacts in a protein is crucial for deriving constraints useful in modeling protein folding and/or scoring remote homology search. Here we focus on the prediction of residue contacts and show that this figure can be predicted with a neural network based method. The accuracy of the prediction is 12 percentage points higher than that of a simple statistical method. The neural network is used to discriminate between two different states of residue contacts, characterized by a contact number higher or lower than the average value of the residue distribution. When evolutionary information is taken into account, our method correctly predicts 69% of the residue states in the data base and it adds to the prediction of residue solvent accessibility. The predictor is available at htpp://www.biocomp.unibo.it

Animals↗

Role of evolutionary information in predicting the disulfide-bonding state of cysteine in proteins.

A neural network-based predictor is trained to distinguish the bonding states of cysteine in proteins starting from the residue chain. Training is performed by using 2,452 cysteine-containing segments extracted from 641 nonhomologous proteins of well-resolved three-dimensional structure. After a cross-validation procedure, efficiency of the prediction scores were as high as 72% when the predictor is trained by using protein single sequences. The addition of evolutionary information in the form of multiple sequence alignment and a jury of neural networks increases the prediction efficiency up to 81%. Assessment of the goodness of the prediction with a reliability index indicates that more than 60% of the predictions have an accuracy level greater than 90%. A comparison with a statistical method previously described and tested on the same database shows that the neural network-based predictor is performing with the highest efficiency. Proteins 1999;36:340-346.

Binding Sites↗

A neural network based predictor of residue contacts in proteins.

We describe a method based on neural networks for predicting contact maps of proteins using as input chemicophysical and evolutionary information. Neural networks are trained on a data set comprising the contact maps of 200 non-homologous proteins of well resolved three-dimensional structures. The systems learn the association rules between the covalent structure of each protein and its correspondent contact map by means of a standard back propagation algorithm. Validation of the predictor on the training set and on 408 proteins of known structure which are not homologous to those contained in the training set indicate that this method scores higher than statistical approaches previously described and based on correlated mutations and sequence information.

Animals↗

Quantum mechanical analysis of oxygenated and deoxygenated states of hemocyanin: theoretical clues for a plausible allosteric model of oxygen binding.

In this work with ab initio computations, we describe relevant interactions between protein active sites and ligands, using as a test case arthropod hemocyanins. A computational analysis of models corresponding to the oxygenated and deoxygenated forms of the hemocyanin active site is performed using the Density Functional Theory approach. We characterize the electron density distribution of the binding site with and without bound oxygen in relation to the geometry, which stems out of the crystals of three hemocyanin proteins, namely the oxygenated form from the horseshoe crab Limulus polyphemus, and the deoxygenated forms, respectively, from the same source and from another arthropod, the spiny lobster Panulirus interruplus. Comparison of the three available crystals indicate structural differences at the oxygen binding site, which cannot be explained only by the presence and absence of the oxygen ligand, since the geometry of the ligand site of the deoxygenated Panulirus hemocyanin is rather similar to that of the oxygenated Limulus protein. This finding was interpreted in the frame of a mechanism of allosteric regulation for oxygen binding. However, the cooperative mechanism, which is experimentally well documented, is only partially supported by crystallographic data, since no oxygenated crystal of Panulirus hemocyanin is presently available. We address the following question: is the local ligand geometry responsible for the difference of the dicopper distance observed in the two deoxygenated forms of hemocyanin or is it necessary to advocate the allosteric regulation of the active site conformations in order to reconcile the different crystal forms? We find that the difference of the dicopper distance between the two deoxygenated hemocyanins is not due to the small differences of ligand geometry found in the crystals and conclude that it must be therefore stabilized by the whole protein tertiary structure.

Allosteric Regulation↗

A data base of minimally frustrated alpha helical segments extracted from proteins according to an entropy criterion.

A data base of minimally frustrated alpha helical segments is defined by filtering a set comprising 822 non redundant proteins, which contain 4783 alpha helical structures. The data base definition is performed using a neural network-based alpha helix predictor, whose outputs are rated according to an entropy criterion. A comparison with the presently available experimental results indicates that a subset of the data base contains the initiation sites of protein folding experimentally detected and also protein fragments which fold into stable isolated alpha helices. This suggests the usage of the data base (and/or of the predictor) to highlight patterns which govern the stability of alpha helices in proteins and the helical behavior of isolated protein fragments.

Algorithms↗

Can functional regions of proteins be predicted from their coding sequences? The case study of G-protein coupled receptors.

A filter based on a set of unsupervised neural networks trained with a winner-take-all strategy discloses signals along the coding sequences of G-protein coupled receptors. By comparing with the existing experimental data it appears that these signals correlate with putative functional domains of the proteins. After protein alignment within subfamilies, signals cluster in protein regions which, according to the presently available experimental results, are described as possible functional domains of the folded proteins. The mapping procedure reveals characteristic regions in the coding sequences common and/or characteristic of the receptor subtype. This is particularly noticeable for the third cytoplasmic loop, which is likely to be involved in the molecular coupling of all the subfamilies with G-proteins. The results indicate that our mapping can highlight intrinsic representative features of the coding sequences which, in the case of G-protein coupled receptors, are characteristic of protein functional regions and suggest a possible application of the filter for predicting functional determinants in proteins starting from the coding sequence.

Amino Acid Sequence↗

An entropy criterion to detect minimally frustrated intermediates in native proteins.

The analysis of the information flow in a feed-forward neural network suggests that the output of the network can be used to compute a structural entropy for the sequence-to-secondary structure mapping. On this basis, we formulate a minimum entropy criterion for the identification of minimally frustrated traits with helical conformation that correspond to initiation sites of protein folding. The entropy of protein segments can be viewed as a nucleation propensity that is useful to characterize putative regions where folding is likely to be initiated with the formation of stretches of alpha-helices under the predominant influence of local interactions. Our procedure is successfully tested in the search for initiation sites of protein folding for which independent experimental and computational evidence exists. Our results lend support to the view that folding is a hierarchical event in which, in harmony with the minimal frustration principle, the final conformation preserves structural modules formed in the early stages of the process.

Proteins↗

A high diffusion coefficient for coenzyme Q10 might be related to a folded structure.

We measured the lateral diffusion of different coenzyme Q homologues and analogues in model lipid vesicles using the fluorescence collisional quenching technique with pyrene derivatives and found diffusion coefficients in the range of 10(-6) cm2/s. Theoretical diffusion coefficients for these highly hydrophobic components were calculated according to the free volume theory. An important parameter in the free volume theory is the relative dimension between diffusant and solvent: a molecular dynamics computer simulation of the coenzymes yielded their most probable geometries and volumes and revealed surprisingly similar sizes of the short and long homologues, due to a folded structure of the isoprenoid chain in the latter, with a length for coenzyme Q10 of 21 A. Using this information we were able to calculate diffusion coefficients in the range of 10(-6) cm2/s, in good agreement with those found experimentally.

Coenzymes↗

Self-organizing neural maps of the coding sequences of G-protein-coupled receptors reveal local domains associated with potentially functional determinants in the proteins.

Mapping of the coding sequences of the best characterized subfamilies of G-protein-coupled receptors is performed with unsupervised neural networks based on a winner-take-all strategy. High order features therefrom extracted originate signals along the aligned protein sequences of the different subfamilies. These plots reveal characteristic domains common and/or characteristic of the receptor subfamily. By comparison with the existing experimental results, it is obtained that most of the regions signalled by clustering overlap with possible functional regions in the folded proteins. This is particularly noticeable for the third cytoplasmic loop, which is likely to be involved in the molecular coupling with the G-proteins. The results suggest that functional regions in proteins may be characterized by intrinsic representative features in the coding sequences which can be enlighted by high order mapping.

Binding Sites↗