PubMed HealthSearch

Biomedical subjects

H Soldano

Publications and source records attributed to H Soldano.

5 recordsLinked to original sources

'Multifrequency' location and clustering of sequence patterns from proteins.

In previous work, we have shown that a set of characteristics, defined as (code frequency) pairs, can be derived from a protein family by the use of a signal-processing method. This method enables the location and extraction of sequence patterns by taking into account each (code frequency) pair individually. In the present paper, we propose to extend this method in order to detect and visualize patterns by taking into account several pairs simultaneously. Two 'multifrequency' methods are described. The first one is based on a rewriting of the sequences with new symbols which summarize the frequency information. The second method is based on a clustering of the patterns associated with each pair. Both methods lead to the definition of significant consensus sequences. Some results obtained with calcium-binding proteins and serine proteases are also discussed.

Amino Acid Sequence

Sequence analysis of cell cycle control (cdc2) protein kinases among protein serine/threonine kinases.

Among protein serine/threonine kinases, the CDC2 proteins are both well characterized as protein serine/threonine kinases and are functionally involved in the control of cell division. Protein serine/threonine kinase sequences were analysed using Fourier transform of the coded sequences. Characteristic code/frequency pairs were extracted from a set of well defined protein serine/threonine kinases. The characteristic frequencies 0.179, 0.250 and 0.408 distinguished protein serine/threonine kinases from proteins which did not have the biological activity. Pertinent patterns in the sequence, responsible for the code/frequency pairs detection were searched and found to be correlated with the putative catalytic domain of the proteins. Protein serine/threonine kinases involved in cell division control, CDC2 protein kinases, were compared to the other protein serine/threonine kinases. Specific code/frequency pairs were extracted from the sequences and could be related to the function or regulation of the kinases in cell division. Two CDC2 related proteins CDC2(Mm) from mice and CDC2(Gg) from chicken were shown to fit well with the CDC2 proteins, whereas KIN28, PHO85 and PSKJ3, which share sequence homology but not functional activity with the CDC2 proteins, were clearly excluded from the CDC2 proteins by the characteristic code/frequency pairs. Pertinent patterns in the CDC2 proteins were analysed and mapped on the CDC2 related protein sequences. Four patterns were correlated with the code/frequency detection and therefore, could be associated to the regulation of the CDC2-related proteins.

Amino Acid Sequence

A scale-independent signal processing method for sequence analysis.

In this paper, we present methods to detect and localize patterns in biologically related protein sequences (family). The patterns common to the sequences of the family are detected by using Fourier analysis. No previous scales (codes) are needed, they are actually produced as a result of the analysis procedure, together with the frequencies of the Fourier decompositions. Characteristic features of the family are thus expressed as (code-frequency) pairs. Various tools are proposed in order to localize the patterns, to compare the codes, and to evaluate the proximity of an arbitrary sequence to the investigated family. The general strategy is illustrated on a family composed of calcium-binding proteins.

Amino Acid Sequence

Statistico-syntactic learning techniques.

The methods of "learning from examples" enable the solving of problems of classification: discrimination between two classes of objects, assimilation of an object to a class of objects representing a property. They are used in a situation where we don't know a priori a procedure in order to decide, but we have examples (in sufficient amount). After a learning stage with the examples, a procedure to solve the problem is built. In the exposed methodology the description of an object is a list of attributes, the acquired knowledge is sets of "rules" considered as arguments in favour of a particular decision.

Computers

From data banks to data bases.

The information collected in national and international libraries on nucleotide and protein sequences cannot be directly treated for proper handling by existing software. Therefore we evaluated the feasibility of constructing a data base for Escherichia coli using the data present in the banks. The knowhow thus acquired was applied to Bacillus subtilis. Specific examples of the general procedure are given.

Bacillus subtilis