PubMed HealthSearch

Biomedical subjects

B Pflugfelder

Publications and source records attributed to B Pflugfelder.

3 recordsLinked to original sources

Self-organized neural maps of human protein sequences.

We have recently described a method based on artificial neural networks to cluster protein sequences into families. The network was trained with Kohonen's unsupervised learning algorithm using, as inputs, the matrix patterns derived from the dipeptide composition of the proteins. We present here a large-scale application of that method to classify the 1,758 human protein sequences stored in the SwissProt database (release 19.0), whose lengths are greater than 50 amino acids. In the final 2-dimensional topologically ordered map of 15 x 15 neurons, proteins belonging to known families were associated with the same neuron or with neighboring ones. Also, as an attempt to reduce the time-consuming learning procedure, we compared 2 learning protocols: one of 500 epochs (100 SUN CPU-hours [CPU-h]), and another one of 30 epochs (6.7 CPU-h). A further reduction of learning-computing time, by a factor of about 3.3, with similar protein clustering results, was achieved using a matrix of 11 x 11 components to represent the sequences. Although network training is time consuming, the classification of a new protein in the final ordered map is very fast (14.6 CPU-seconds). We also show a comparison between the artificial neural network approach and conventional methods of biosequence analysis.

Algorithms

A hybrid method to cluster protein sequences based on statistics and artificial neural networks.

We have recently proposed a method, based on artificial neural networks (ANNs) to cluster protein sequences into families according to their degree of sequence similarity. The network was trained with an unsupervised learning algorithm, using, as inputs, matrix patterns derived from the bipeptide composition of the protein sequences. We describe here some further improvements to that approach. First, we propose a statistical method to cluster a set of bipeptidic matrices into families. It consists of three stages: (i) principal component analysis, (ii) determination of the optimal number M of clusters and (iii) final classification of the bipeptidic matrices into M clusters. Using a set of 444 protein sequences, we show that the classification given by the statistical method is in agreement with biological knowledge. We also show that the resulting classification is very similar to the one previously obtained with the ANN approach. Finally, we propose a new hybrid method of the statistical and ANN approaches, in which the results of the statistical method are used to choose the number of neurons and inputs of the network. We show that a network built in this way, and fed with a few principal components of the set of bipeptidic matrices as input signals, can be trained in an extremely short computing time. The resulting topological maps do not essentially differ from the ones obtained with the initial ANN approach.

Algorithms

Protein classification using neural networks.

We have recently described a method based on Artificial Neural Networks to cluster protein sequences into families. The network was trained with Kohonen's unsupervised-learning algorithm using, as inputs, matrix patterns derived from the bipeptide composition of the proteins. We show here the application of that method to classify 1758 protein sequences, using as inputs a limited number of principal components of the bipeptidic matrices. As a result of training, the network self-organized the activation of its neurons into a topologically ordered map, in which proteins belonging to a known family (immunoglobulins, actins, interferons, myosins, HLA histocompatibility antigens, hemoglobins, etc.) were usually associated with the same neuron or with neighboring ones. Once the topological map has been obtained, the classification of new sequences is very fast.

Algorithms