PubMed HealthSearch

Biomedical subjects

I Muchnik

Publications and source records attributed to I Muchnik.

7 recordsLinked to original sources

Protein folding class predictor for SCOP: approach based on global descriptors.

This work demonstrates new techniques developed for the prediction of protein folding class in the context of the most comprehensive Structural Classification of Proteins (SCOP). The prediction method uses global descriptors of a protein in terms of the physical, chemical and structural properties of its constituent amino acids. Neural networks are utilized to combine these descriptors in a specific way to discriminate members of a given folding class from members of all other classes. It is shown that a specific amino acid's properties work completely differently on different folding classes. This creates the possibility of finding an individual set of descriptors that works best on a particular folding class.

Algorithms

Reconstruction of ancient molecular phylogeny.

Support for contradictory phylogenies is often obtained when molecular sequence data from different genes is used to reconstruct phylogenetic histories. Contradictory phylogenies can result from many data anomalies including unrecognized paralogy. Paralogy, defined as the reconstruction of a phylogenetic tree from a mixture of genes generated by duplications, has generally not been formally included in phylogenetic reconstructions. Here we undertake the task of reconstructing a single most likely evolutionary relationship among a range of taxa from a large set of apparently inconsistent gene trees. Under the assumption that differences among gene trees can be explained by gene duplications, and consequent losses, we have developed a method to obtain the global phylogeny minimizing the total number of postulated duplications and losses and to trace back such individual gene duplications to global genome duplications. We have used this method to infer the most likely phylogenetic relationship among 16 major higher eukaryotic taxa from the sequences of 53 different genes. Only five independent genome duplication events need to be postulated in order to explain the inconsistencies among these trees.

Algorithms

Prediction of protein folding class using global description of amino acid sequence.

We present a method for predicting protein folding class based on global protein chain description and a voting process. Selection of the best descriptors was achieved by a computer-simulated neural network trained on a data base consisting of 83 folding classes. Protein-chain descriptors include overall composition, transition, and distribution of amino acid attributes, such as relative hydrophobicity, predicted secondary structure, and predicted solvent exposure. Cross-validation testing was performed on 15 of the largest classes. The test shows that proteins were assigned to the correct class (correct positive prediction) with an average accuracy of 71.7%, whereas the inverse prediction of proteins as not belonging to a particular class (correct negative prediction) was 90-95% accurate. When tested on 254 structures used in this study, the top two predictions contained the correct class in 91% of the cases.

Amino Acid Sequence

A biologically consistent model for comparing molecular phylogenies.

In the framework of the problem of combining different gene trees into a unique species phylogeny, a model for duplication/speciation/loss events along the evolutionary tree is introduced. The model is employed for embedding a phylogeny tree into another one via the so-called duplication/speciation principle requiring that the gene duplicated evolves in such a way that any of the contemporary species involved bears only one of the gene copies diverged. The number of biologically meaningful elements in the embedding result (duplications, losses, information gaps) is considered a (asymmetric) dissimilarity measure between the trees. The model duplication concept is compared with that one defined previously in terms of a mapping procedure for the trees. A graph-theoretic reformulation of the measure is derived.

Evolution, Molecular

Relation between protein structure, sequence homology and composition of amino acids.

A method of quantitative comparison of two classifications rules applied to protein folding problem is presented. Classification of proteins based on sequence homology and based on amino acid composition were compared and analyzed according to this approach. The coefficient of correlation between these classification methods and the procedure of estimation of robustness of the coefficient are discussed.

Amino Acid Sequence

Modeling protein cores with Markov random fields.

A mathematical formalism is introduced that has general applicability to many protein structure models used in the various approaches to the "inverse protein folding problem." The inverse nature of the problem arises from the fact that one begins with a set of assumed tertiary structures and searches for those most compatible with a new sequence, rather than attempting to predict the structure directly from the new sequence. The formalism is based on the well-known theory of Markov random fields (MRFs). Our MRF formulation provides explicit representations for the relevant amino acid position environments and the physical topologies of the structural contacts. In particular, MRF models can readily be constructed for the secondary structure packing topologies found in protein domain cores, or other structural motifs, that are anticipated to be common among large sets of both homologous and nonhomologous proteins. MRF models are probabilistic and can exploit the statistical data from the limited number of proteins having known domain structures. The MRF approach leads to a new scoring function for comparing different threadings (placements) of a sequence through different structure models. The scoring function is very important, because comparing alternative structure models with each other is a key step in the inverse folding problem. Unlike previously published scoring functions, the one derived in this paper is based on a comprehensive probabilistic formulation of the threading problem.

Amino Acid Sequence

Efficient pattern comparative method for selecting functionally important motifs in protein sequences: application to zinc enzymes.

We have developed a pattern comparative method for identifying functionally important motifs in protein sequences. The essence of most standard pattern comparative methods is a comparison of patterns occurring in different sequences using an optimized weight matrix. In contrast, our approach is based on a measure of similarity among all the candidate motifs within the same sequence. This method may prove to be particularly efficient for proteins encoding the same biochemical function, but with different primary sequences, and when tertiary structure information from one or more sequences is available. We have applied this method to a special class of zinc-binding enzymes known as endopeptidases.

Amino Acid Sequence