PubMed Health⌕ Search

Biomedical subjects

Y D Cai

Publications and source records attributed to Y D Cai.

16 recordsLinked to original sources

Support vector machines for predicting protein structural class.

BACKGROUND: We apply a new machine learning method, the so-called Support Vector Machine method, to predict the protein structural class. Support Vector Machine method is performed based on the database derived from SCOP, in which protein domains are classified based on known structures and the evolutionary relationships and the principles that govern their 3-D structure. RESULTS: High rates of both self-consistency and jackknife tests are obtained. The good results indicate that the structural class of a protein is considerably correlated with its amino acid composition. CONCLUSIONS: It is expected that the Support Vector Machine method and the elegant component-coupled method, also named as the covariant discrimination algorithm, if complemented with each other, can provide a powerful computational tool for predicting the structural classes of proteins.

Algorithms↗

Is it a paradox or misinterpretation?

The paradox recently raised by Wang and Yuan (Proteins 2000;38:165-175) in protein structural class prediction is actually a misinterpretation of the data reported in the literature. The Bayes decision rule, which was deemed by Wang and Yuan to be the most powerful method for predicting protein structural classes based on the amino acid composition, and applied by these investigators to derive the upper limit of prediction rate for structural classes, is actually completely the same as the component-coupled algorithm proposed by previous investigators (Chou et al., Proteins 1998;31:97-103). Owing to lack of a complete or near-complete training data set, the upper limit rate thus derived by these investigators might be both invalid and misleading. Clarification of these points will further stimulate investigation of this interesting area.

Algorithms↗

Artificial neural network model for predicting membrane protein types.

Membrane proteins can be classified among the following five types: (1) type I membrane protein. (2) type II membrane protein. (3) multipass transmembrane proteins. (4) lipid chain-anchored membrane proteins, and (5) GPI-anchored membrane proteins. T. Kohonen's self-organization model which is a typical neural network is applied for predicting the type of a given membrane protein based on its amino acid composition. As a result, the high rates of self-consistency (94.80%) and cross-validation (77.76%), and stronger fault-tolerant ability were obtained.

Algorithms↗

Using neural networks for prediction of subcellular location of prokaryotic and eukaryotic proteins.

T. Kohonen's self-organization model, a typical neural network model, was applied to predict the subcellular location of proteins from their amino acid composition. The Reinhardt and Hubbard database was used to examine the performance of the neural network method. The rates of correct prediction for the three possible subcellular location of prokaryotic proteins were 96.1% by the self-consistency test and 84.4% by the jackknife test. The rates of correct prediction for the four possible subcellular location of eukaryotic proteins were 95.6% by the self-consistency test and 70.6% by the jackknife test.

Algorithms↗

Support vector machines for prediction of protein subcellular location.

Support Vector Machine (SVM), which is one kind of learning machines, was applied to predict the subcellular location of proteins from their amino acid composition. In this research, the proteins are classified into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracall, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane, and (12) vacuole, which have covered almost all the organelles and subcellular compartments in an animal or plant cell. The examination for the self-consistency and the jackknife test of the SVMs method was tested for the three sets: 2022 proteins, 2161 proteins, and 2319 proteins. As a result, the correct rate of self-consistency and jackknife test reaches 91 and 82% for 2022 proteins, 89 and 75% for 2161 proteins, and 85 and 73% for 2319 proteins, respectively. Furthermore, the predicting rate was tested by the three independent testing datasets containing 2240 proteins, 2513 proteins, and 2591 proteins. The correct prediction rates reach 82, 75, and 73% for 2240 proteins, 2513 proteins, and 2591 proteins, respectively.

Algorithms↗

Artificial neural network method for predicting HIV protease cleavage sites in protein.

Knowledge of the polyprotein cleavage sites by HIV protease will refine our understanding of its specificity, and the information thus acquired will be useful for designing specific and efficient HIV protease inhibitors. The search for inhibitors of HIV protease will be greatly expedited if one can find an accurate, robust, and rapid method for predicting the cleavage sites in proteins by HIV protease. In this paper, Kohonen's self-organization model, which uses typical artificial neural networks, is applied to predict the cleavability of oligopeptides by proteases with multiple and extended specificity subsites. We selected HIV-1 protease as the subject of study. We chose 299 oligopeptides for the training set, and another 63 oligopeptides for the test set. Because of its high rate of correct prediction (58/63 = 92.06%) and stronger fault-tolerant ability, the neural network method should be a useful technique for finding effective inhibitors of HIV protease, which is one of the targets in designing potential drugs against AIDS. The principle of the artificial neural network method can also be applied to analyzing the specificity of any multisubsite enzyme.

Amino Acid Sequence↗

Prediction of beta-turns.

Kohonen's self-organization model, a neural network model, is applied to predict the beta-turns in proteins. There are 455 beta-turn tetrapeptides and 3807 non-beta-turn tetrapeptides in the training database. The rates of correct prediction for the 110 beta-turn tetrapeptides and 30,229 non-beta-turn tetrapeptides in the testing database are 81.8% and 90.7%, respectively. The high quality of prediction of neural network model implies that the residue-coupled effect along a polypeptide chain is important for the formation of reversal turns, such as beta-turns, during the process of protein folding.

Amino Acid Sequence↗

Artificial neural network method for predicting the specificity of GalNAc-transferase.

The specificity of GalNAc-transferase is consistent with the existence of an extended site composed of nine subsites, denoted by R4, R3, R2, R1, R0, R1', R2', R3', and R4', where the acceptor at R0 is either Ser or Thr to which the reducing monosaccharide is anchored. To predict whether a peptide will react with the enzyme to form a Ser- or Thr-conjugated glycopeptide, a neural network method--Kohonen's self-organization model is proposed in this paper. Three hundred five oligopeptides are chosen for the training site, with another 30 oligopeptides for the test set. Because of its high correct prediction rate (26/30 = 86.7%) and stronger fault-tolerant ability, it is expected that the neural network method can be used as a technique for predicting O-glycosylation and designing effective inhibitors of GalNAc-transferase. It might also be useful for targeting drugs to specific sites in the body and for enzyme replacement therapy for the treatment of genetic disorders.

Glycosylation↗

The breast cancer gene product TSG101: a regulator of ubiquitination?

Sequence analysis is a powerful tool to obtain structural and functional information about genes and their products. Here we show that TSG101, a gene subjected to somatic mutations in breast cancer, contains an amino terminal domain that is a homologue of ubiquitin conjugating enzymes (UBCs) and not, as previously proposed, DNA-binding domains. As the UBC active site residue is replaced in the TSG101 sequence in a similar manner to several other members of the UBC family, we propose a role for TSG101 in regulating the ubiquitination of short-lived gene products.

Amino Acid Sequence↗