PubMed Health⌕ Search

Biomedical subjects

M Reczko

Publications and source records attributed to M Reczko.

8 recordsLinked to original sources

DIANA-EST: a statistical analysis.

MOTIVATION: Expressed Sequence Tags (ESTs) are next to cDNA sequences as the most direct way to locate in silico the genes of the genome and determine their structure. Currently ESTs make up more than 60% of all the database entries. The goal of this work is the development of a new program called DNA Intelligent Analysis for ESTs (DIANA-EST) based on a combination of Artificial Neural Networks (ANN) and statistics for the characterization of the coding regions within ESTs and the reconstruction of the encoded protein. RESULTS: 89.7% of the nucleotides from an independent test set with 127 ESTs were predicted correctly as to whether they are coding or non coding. AVAILABILITY: The program is available upon request from the author. CONTACT: Present address: Department of Genetics, University of Pennsylvania, School of Medicine, 475 Clinical Research Building, 415 Curie Boulevard, Philadelphia, PA 19104-6145, USA. artemis@pcbi.upenn.edu.

Computational Biology↗

Protein fold class prediction: new methods of statistical classification.

Feed forward neural networks are compared with standard and new statistical classification procedures for the classification of proteins. We applied logistic regression, an additive model and projection pursuit regression from the methods based on a posterior probabilities; linear, quadratic and a flexible discriminant analysis from the methods based on class conditional probabilities, and the K-nearest-neighbors classification rule. Both, the apparent error rate obtained with the training sample (n = 143) and the test error rate obtained with the test sample (n = 125) and the 10-fold cross validation error were calculated. We conclude that some of the standard statistical methods are potent competitors to the more flexible tools of machine learning.

Algorithms↗

Prediction of protein hydration sites from sequence by modular neural networks.

The hydration properties of a protein are important determinants of its structure and function. Here, modular neural networks are employed to predict ordered hydration sites using protein sequence information. First, secondary structure and solvent accessibility are predicted from sequence with two separate neural networks. These predictions are used as input together with protein sequences for networks predicting hydration of residues, backbone atoms and sidechains. These networks are trained with protein crystal structures. The prediction of hydration is improved by adding information on secondary structure and solvent accessibility and, using actual values of these properties, residue hydration can be predicted to 77% accuracy with a Matthews coefficient of 0.43. However, predicted property data with an accuracy of 60-70% result in less than half the improvement in predictive performance observed using the actual values. The inclusion of property information allows a smaller sequence window to be used in the networks to predict hydration. It has a greater impact on the accuracy of hydration site prediction for backbone atoms than for sidechains and for non-polar than polar residues. The networks provide insight into the mutual interdependencies between the location of ordered water sites and the structural and chemical characteristics of the protein residues.

Amino Acid Sequence↗

An update of the DEF database of protein fold class predictions.

An update is given on the Database of Expected Fold classes (DEF) that contains a collection of fold-class predictions made from protein sequences and a mail server that provides new predictions for new sequences. To any given sequence one of 49 fold-classes is chosen to classify the structure related to the sequence with high accuracy. The updated prediction system is developed using data from the new version of the 3D-ALI database of aligned protein structures and thus is giving more reliable and more detailed predictions than the previous DEF system.

Amino Acid Sequence↗

A parallel neural network simulator on the connection machine CM-5.

We here present a parallel implementation of artificial neural networks on the connection machine CM-5 and compare it with other parallel implementations on SIMD and MIMD architectures. This parallel implementation was developed with the goal of efficiently training large neural networks with huge training pattern sets for applications in molecular biology, in particular the prediction of coding regions in DNA sequences. The implementation uses training pattern parallelism and makes use of the parallel I/O facilities of the CM-5 and its efficient reduction operations available within the control network to achieve a high scalability. The parallel simulator obtains a maximum speed of 149.25 MCUPS for training feedforward networks with backpropagation on a 512 processor CM-5 system without using the CM-5 vector facility. The implementation poses no restriction on the type of network topology and works with different batch training algorithms like BP. Quickprop and Rprop.

Algorithms↗

Prediction of hypervariable CDR-H3 loop structures in antibodies.

The structure of the most variable antibody hypervariable loop, CDR-H3, has been predicted from amino acid sequence alone. In contrast to other approaches predictions are made for loop lengths up to 17 residues. The predictions have been achieved using artificial neural networks which are trained on a large set of loops from the Brookhaven Protein Databank which have structures similar to CDR-H3. The loop structures are described by the two backbone dihedral angles phi and psi for each residue. For 21 CDR-H3 loops unique to the neural network, the prediction of dihedral angles leads to an average root mean square deviation in the Cartesian coordinates of 2.65 A. The present method, when combined with existing modelling protocols, provides an important addition to the structural prediction of the complementarity determining regions of antibodies.

Amino Acid Sequence↗

The DEF data base of sequence based protein fold class predictions.

A new method for predicting protein fold-classes and protein domains from sequence data is constructed and used for generating a data base of protein fold-class assignments. Any given sequence of amino acids is assigned a specific prediction of one out of 45 typical protein fold-classes, a prediction of one out of 4 super fold-classes for the content of secondary structures and a profile of fold-class predictions along the sequence. The prediction accuracy for the super fold-classes is around 91% correct and 82% correct for the specific fold-classes. This accuracy is maintained down to a few percent of sequence identity.

Amino Acid Sequence↗

Protein secondary structure prediction with partially recurrent neural networks.

Partially recurrent neural networks with different topologies are applied for secondary structure prediction of proteins. The state of some activations in the network is available after a pattern presentation via feedback connections as additional input during the processing of the next pattern in a sequence. A reference data set containing 91 proteins in the training set and 15 non-homologous proteins in the test set is used for training and testing a network with a modified, hierarchical Elman architecture. The network predicts the secondary structures alpha-helix, beta-sheet, and "coil" for each amino acid. The percentage of correctly classified amino acids is 67.83% on the training set and 63.98% on the test set. The best performance of a three-layer feedforward network is 62.7% on the same test set. A cascaded network, where the outputs of the recurrent network are processed by a second net with 13 x 3 inputs, four hidden and three output units has a predictive performance of 64.49%. The best corresponding feedforward net has a performance of 64.3%.

Amino Acid Sequence↗