PubMed HealthSearch

Biomedical subjects

E A Ferrán

Publications and source records attributed to E A Ferrán.

6 recordsLinked to original sources

Improved differential screening approach to analyse transcriptional variations in organized cDNA libraries.

A cDNA library was generated from rat brain tissues and organized into 1536-well plates, using a fluorescence activated cell sorter (FACS), acting as a single cell deposition system. The organized library containing 10,000 clones, with 60% full-length cDNA inserts, allowed the generation of multiple identical membrane replicas. Each replica was hybridized with a complex probe obtained from a particular brain tissue or a given cultured cell. The signal intensity for each of the clones present on the membrane, quantified with a standard image-analysis software, is proportional both to the abundance of the corresponding mRNA in the probe and to the amount of plasmid template on the membrane. The latter value was thus used to normalize the signals produced with complex probes, to optimize the comparison of mRNA expression levels for the different systems under study. The construction of high-quality cDNA libraries, the generation of identical membrane replicas and comparable probes, and the utilization of an image-analysis software package, coupled with the normalization of the spot intensity by assaying plasmid quantity, significantly improves the differential screening approach. Altogether, these technical improvements open the possibility to compare a great number of different probes and, in consequence, to accumulate biological information for each clone present in an organized cDNA library. The functional information obtained should complement data from DNA sequencing projects.

Animals

Self-organized neural maps of human protein sequences.

We have recently described a method based on artificial neural networks to cluster protein sequences into families. The network was trained with Kohonen's unsupervised learning algorithm using, as inputs, the matrix patterns derived from the dipeptide composition of the proteins. We present here a large-scale application of that method to classify the 1,758 human protein sequences stored in the SwissProt database (release 19.0), whose lengths are greater than 50 amino acids. In the final 2-dimensional topologically ordered map of 15 x 15 neurons, proteins belonging to known families were associated with the same neuron or with neighboring ones. Also, as an attempt to reduce the time-consuming learning procedure, we compared 2 learning protocols: one of 500 epochs (100 SUN CPU-hours [CPU-h]), and another one of 30 epochs (6.7 CPU-h). A further reduction of learning-computing time, by a factor of about 3.3, with similar protein clustering results, was achieved using a matrix of 11 x 11 components to represent the sequences. Although network training is time consuming, the classification of a new protein in the final ordered map is very fast (14.6 CPU-seconds). We also show a comparison between the artificial neural network approach and conventional methods of biosequence analysis.

Algorithms

A hybrid method to cluster protein sequences based on statistics and artificial neural networks.

We have recently proposed a method, based on artificial neural networks (ANNs) to cluster protein sequences into families according to their degree of sequence similarity. The network was trained with an unsupervised learning algorithm, using, as inputs, matrix patterns derived from the bipeptide composition of the protein sequences. We describe here some further improvements to that approach. First, we propose a statistical method to cluster a set of bipeptidic matrices into families. It consists of three stages: (i) principal component analysis, (ii) determination of the optimal number M of clusters and (iii) final classification of the bipeptidic matrices into M clusters. Using a set of 444 protein sequences, we show that the classification given by the statistical method is in agreement with biological knowledge. We also show that the resulting classification is very similar to the one previously obtained with the ANN approach. Finally, we propose a new hybrid method of the statistical and ANN approaches, in which the results of the statistical method are used to choose the number of neurons and inputs of the network. We show that a network built in this way, and fed with a few principal components of the set of bipeptidic matrices as input signals, can be trained in an extremely short computing time. The resulting topological maps do not essentially differ from the ones obtained with the initial ANN approach.

Algorithms

Protein classification using neural networks.

We have recently described a method based on Artificial Neural Networks to cluster protein sequences into families. The network was trained with Kohonen's unsupervised-learning algorithm using, as inputs, matrix patterns derived from the bipeptide composition of the proteins. We show here the application of that method to classify 1758 protein sequences, using as inputs a limited number of principal components of the bipeptidic matrices. As a result of training, the network self-organized the activation of its neurons into a topologically ordered map, in which proteins belonging to a known family (immunoglobulins, actins, interferons, myosins, HLA histocompatibility antigens, hemoglobins, etc.) were usually associated with the same neuron or with neighboring ones. Once the topological map has been obtained, the classification of new sequences is very fast.

Algorithms

Clustering proteins into families using artificial neural networks.

An artificial neural network was used to cluster proteins into families. The network, composed of 7 x 7 neurons, was trained with the Kohonen unsupervised learning algorithm using, as inputs, matrix patterns derived from the bipeptide composition of 447 proteins, belonging to 13 different families. As a result of the training, and without any a priori indication of the number or composition of the expected families, the network self-organized the activation of its neurons into topologically ordered maps in which almost all the proteins (96.7%) were correctly clustered into the corresponding families. In a second computational experiment, a similar network was trained with one family of the previous learning set (76 cytochrome c sequences). The new neural map clustered these proteins into 25 different neurons (five in the first experiment), wherein phylogenetically related sequences were positioned close to each other. This result shows that the network can adapt the clustering resolution to the complexity of the learning set, a useful feature when working with an unknown number of clusters. Although the learning stage is time consuming, once the topological map is obtained, the classification of new proteins is very fast. Altogether, our results suggest that this novel approach may be a useful tool to organize the search for homologies in large macromolecular databases.

Algorithms

Topological maps of protein sequences.

A new method based on neural networks to cluster proteins into families is described. The network is trained with the Kohonen unsupervised learning algorithm, using matrix pattern representations of the protein sequences as inputs. The components (x, y) of these 20 x 20 matrix patterns are the normalized frequencies of all pairs xy of amino acids in each sequence. We investigate the influence of different learning parameters in the final topological maps obtained with a learning set of ten proteins belonging to three established families. In all cases, except in those where the synaptic vectors remains nearly unchanged during learning, the ten proteins are correctly classified into the expected families. The classification by the trained network of mutated or incomplete sequences of the learned proteins is also analysed. The neural network gives a correct classification for a sequence mutated in 21.5% +/- 7% of its amino acids and for fragments representing 7.5% +/- 3% of the original sequence. Similar results were obtained with a learning set of 32 proteins belonging to 15 families. These results show that a neural network can be trained following the Kohonen algorithm to obtain topological maps of protein sequences, where related proteins are finally associated to the same winner neuron or to neighboring ones, and that the trained network can be applied to rapidly classify new sequences. This approach opens new possibilities to find rapid and efficient algorithms to organize and search for homologies in the whole protein database.

Amino Acid Sequence