PubMed Health⌕ Search

Biomedical subjects

Kuo-Chen Chou

Publications and source records attributed to Kuo-Chen Chou.

At least 19 recordsLinked to original sources

Predicting 22 protein localizations in budding yeast.

According to the recent experiments, proteins in budding yeast can be distinctly classified into 22 subcellular locations. Of these proteins, some bear the multi-locational feature, i.e., occur in more than one location. However, so far all the existing methods in predicting protein subcellular location were developed to deal with only the mono-locational case where a query protein is assumed to belong to one, and only one, subcellular location. To stimulate the development of subcellular location prediction, an augmentation procedure is formulated that will enable the existing methods to tackle the multi-locational problem as well. It has been observed thru a jackknife cross-validation test that the success rate obtained by the augmented GO-FnD-PseAA algorithm [BBRC 320 (2004) 1236] is overwhelmingly higher than those by the other augmented methods. It is anticipated that the augmented GO-FunD-PseAA predictor will become a very useful tool in predicting protein subcellular localization for both basic research and practical application.

Algorithms↗

Predicting protein structural class by functional domain composition.

The functional domain composition is introduced to predict the structural class of a protein or domain according to the following classification: all-alpha, all-beta, alpha/beta, alpha+beta, micro (multi-domain), sigma (small protein), and rho (peptide). The advantage by doing so is that both the sequence-order-related features and the function-related features are naturally incorporated in the predictor. As a demonstration, the jackknife cross-validation test was performed on a dataset that consists of proteins and domains with only less than 20% sequence identity to each other in order to get rid of any homologous bias. The overall success rate thus obtained was 98%. In contrast to this, the corresponding rates obtained by the simple geometry approaches based on the amino acid composition were only 36-39%. This indicates that using the functional domain composition to represent the sample of a protein for statistical prediction is very promising, and that the functional type of a domain is closely correlated with its structural class.

Databases, Protein↗

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes.

MOTIVATION: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation. RESULTS: To incorporate the sequence-order effects into the predictor, the 'amphiphilic pseudo amino acid composition' is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2lambda discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2lambda numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function.

Algorithms↗

Prediction of protein subcellular locations by GO-FunD-PseAA predictor.

The localization of a protein in a cell is closely correlated with its biological function. With the explosion of protein sequences entering into DataBanks, it is highly desired to develop an automated method that can fast identify their subcellular location. This will expedite the annotation process, providing timely useful information for both basic research and industrial application. In view of this, a powerful predictor has been developed by hybridizing the gene ontology approach [Nat. Genet. 25 (2000) 25], functional domain composition approach [J. Biol. Chem. 277 (2002) 45765], and the pseudo-amino acid composition approach [Proteins Struct. Funct. Genet. 43 (2001) 246; Erratum: ibid. 44 (2001) 60]. As a showcase, the recently constructed dataset [Bioinformatics 19 (2003) 1656] was used for demonstration. The dataset contains 7589 proteins classified into 12 subcellular locations: chloroplast, cytoplasmic, cytoskeleton, endoplasmic reticulum, extracellular, Golgi apparatus, lysosomal, mitochondrial, nuclear, peroxisomal, plasma membrane, and vacuolar. The overall success rate of prediction obtained by the jackknife cross-validation was 92%. This is so far the highest success rate performed on this dataset by following an objective and rigorous cross-validation procedure.

Algorithms↗

Insights from modelling the 3D structure of the extracellular domain of alpha7 nicotinic acetylcholine receptor.

Based on the crystal structure of acetylcholine-binding protein, the three-dimensional structures of the extracellular domain, or the ligand-binding domains, of the monomer, homodimer, and homopentamer of the alpha7 nicotinic acetylcholine receptor were derived. The interface between two subunits, where the ligand-binding site is located, was investigated. Furthermore, an explicit definition of the ligand-binding pocket was illustrated that might provide useful clues for conducting various mutagenesis studies for finding drugs against schizophrenia and Alzheimer's disease.

Amino Acid Sequence↗

Identify catalytic triads of serine hydrolases by support vector machines.

The core of an enzyme molecule is its active site from the viewpoints of both academic research and industrial application. To reveal the structural and functional mechanism of an enzyme, one needs to know its active site; to conduct structure-based drug design by regulating the function of an enzyme, one needs to know the active site and its microenvironment as well. Given the atomic coordinates of an enzyme molecule, how can we predict its active site? To tackle such a problem, a distance group approach was proposed and the support vector machine algorithm applied to predict the catalytic triad of serine hydrolase family. The success rate by jackknife test for the 139 serine hydrolases was 85%, implying that the method is quite promising and may become a useful tool in structural bioinformatics.

Algorithms↗

Predicting subcellular localization of proteins by hybridizing functional domain composition and pseudo-amino acid composition.

Recent advances in large-scale genome sequencing have led to the rapid accumulation of amino acid sequences of proteins whose functions are unknown. Since the functions of these proteins are closely correlated with their subcellular localizations, many efforts have been made to develop a variety of methods for predicting protein subcellular location. In this study, based on the strategy by hybridizing the functional domain composition and the pseudo-amino acid composition (Cai and Chou [2003]: Biochem. Biophys. Res. Commun. 305:407-411), the Intimate Sorting Algorithm (ISort predictor) was developed for predicting the protein subcellular location. As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate by the jackknife test for the plant protein dataset was 85.4%, and that for the non-plant protein dataset 91.9%. These are so far the highest success rates achieved for the two datasets by following a rigorous cross validation test procedure, further confirming that such a hybrid approach may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology.

Algorithms↗

Modelling extracellular domains of GABA-A receptors: subtypes 1, 2, 3, and 5.

GABA is the main inhibitory neurotransmitter in the mammalian central nervous system. When GABA binds to the ubiquitous GABA-A receptors on neurons, chloride channels are activated leading to a rapid increase in chloride conductance that depresses excitatory depolarization. The GABA-A receptors are targets for many clinically important drugs, such as the benzodiazepines, general anaesthetics, and barbiturates. All of these drugs enhance the chloride current activated by GABA. Of the GABA-A receptor family, the subtype 2 is critical for the treatment of anxiety spectrum disorders. To avoid unwanted side effects, it is necessary to find highly selective drugs that interact only with subtype 2 but not with the related receptors such as subtypes 1, 3, and 5. To realize such a goal, it is important to have not only the 3D (dimensional) structure of subtype 2 but also the 3D structures of subtypes 1, 3, and 5. In this study, the 3D structures of all the four subtypes of GABA-A receptors have been derived. The computer-modeled heteropentameric structures bear the following features: (1) each of the five subunits in the pentamer has an intrachain disulfide bond, a hallmark of ligand-gated pentameric channels; (2) those residues which are sensitive to the binding of the benzodiazepine site ligands are grouped around the alpha1,2,3,5/gamma2 interfaces; and (3) those residues which are sensitive to the binding of GABA molecules are grouped around the alpha1,2,3,5/beta2 interfaces. All these findings are fully consistent with experimental observations. Meanwhile, for those sensitive or key residues, a close look at their subtle difference among the four subtypes has been provided through a highlighted superposition picture. In addition to providing the atomic coordinates, the predicted structures have further clarified some ambiguities that could not been uniquely determined by the existing experimental data, such as the directionality of the subunit arrangement in the heteropentamers. The 3D models may provide a reasonable structural frame or footing for designing highly selective drugs. The present models might be also useful in understanding the basic mechanism of operation of the GABA-A receptors, stimulating novel strategies for developing more specific drugs and better treatments.

Amino Acid Sequence↗

A novel approach to predict active sites of enzyme molecules.

Enzymes are critical in many cellular signaling cascades. With many enzyme structures being solved, there is an increasing need to develop an automated method for identifying their active sites. However, given the atomic coordinates of an enzyme molecule, how can we predict its active site? This is a vitally important problem because the core of an enzyme molecule is its active site from the viewpoints of both pure scientific research and industrial application. In this article, a topological entity was introduced to characterize the enzymatic active site. Based on such a concept, the covariant discriminant algorithm was formulated for identifying the active site. As a paradigm, the serine hydrolase family was demonstrated. The overall success rate by jackknife test for a data set of 88 enzyme molecules was 99.92%, and that for a data set of 50 independent enzyme molecules was 99.91%. Meanwhile, it was shown through an example that the prediction algorithm can also be used to find any typographic error of a PDB file in annotating the constituent amino acids of catalytic triad and to suggest a possible correction. The very high success rates are due to the introduction of a covariance matrix in the prediction algorithm that makes allowance for taking into account the coupling effects among the key constituent atoms of active site. It is anticipated that the novel approach is quite promising and may become a useful high throughput tool in enzymology, proteomics, and structural bioinformatics.

Algorithms↗

Application of SVM to predict membrane protein types.

As a continuous effort to develop automated methods for predicting membrane protein types that was initiated by Chou and Elrod (PROTEINS: Structure, Function, and Genetics, 1999, 34, 137-153), the support vector machine (SVM) is introduced. Results obtained through re-substitution, jackknife, and independent data set tests, respectively, have indicated that the SVM approach is quite a promising one, suggesting that the covariant discriminant algorithm (Chou and Elrod, Protein Eng. 12 (1999) 107) and SVM, if effectively complemented with each other, will become a powerful tool for predicting membrane protein types and the other protein attributes as well.

Algorithms↗

Predicting subcellular localization of proteins in a hybridization space.

MOTIVATION: The localization of a protein in a cell is closely correlated with its biological function. With the number of sequences entering into databanks rapidly increasing, the importance of developing a powerful high-throughput tool to determine protein subcellular location has become self-evident. In view of this, the Nearest Neighbour Algorithm was developed for predicting the protein subcellular location using the strategy of hybridizing the information derived from the recent development in gene ontology with that from the functional domain composition as well as the pseudo amino acid composition. RESULTS: As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate of the jackknife test for the plant protein dataset was 86%, and that for the non-plant protein dataset 91.2%. These are the highest success rates achieved so far for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach (particularly by incorporating the knowledge of gene ontology) may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. AVAILABILITY: The software would be made available on sending a request to the authors.

Algorithms↗

Bio-support vector machines for computational proteomics.

MOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness.

Algorithms↗

Predicting the linkage sites in glycoproteins using bio-basis function neural network.

MOTIVATION: Although, it is known that O-glycosidically linked oligosaccharides are commonly conjugated to a serine, threonine or hydroxylysine residue of the polypeptide, the chemical nature of the anchoring monosaccharide and the size of the oligosaccharide unit varies. Among different types, O-linked or mucin-type oligosaccharides are intimately involved in the secretion of proteins, be they enzymes, hormones or structural glycoproteins. Knowledge of the linkage sites in glycoproteins is critical to the design of specific and efficient inhibitors against the enzyme to catalyse the formation of the carbohydrate-peptide linkage. RESULTS: We present a method for predicting the linkage sites in O-linked glycoproteins using bio-basis function neural networks. The mean prediction accuracy of this method is 91.15 +/- 2.75% while it is 82.28 +/- 6.45% using back-propagation neural networks. Importantly, this method has significantly reduced the CPU time for modelling.

Algorithms↗

Structural bioinformatics and its impact to biomedical science.

During the last two decades, the number of sequence-known proteins has increased rapidly. In contrast, the corresponding increment for structure-known proteins is much slower. The unbalanced situation has critically limited our ability to understand the molecular mechanism of proteins and conduct structure-based drug design by timely using the updated information of newly found sequences. Therefore, it is highly desired to develop an automated method for fast deriving the 3D (3-dimensional) structure of a protein from its sequence. Under such a circumstance, the structural bioinformatics was emerging naturally as the time required. In this review, three main strategies developed in structural bioinformatics, i.e., pure energetic approach, heuristic approach, and homology modeling approach, as well as their underlying principles, are briefly introduced. Meanwhile, a series of demonstrations are presented to show how the structural bioinformatics has been applied to timely derive the 3D structures of some functionally important proteins, helping to understand their action mechanisms and stimulating the course of drug discovery. Also, the limitation of these approaches and the future challenges of structural bioinformatics are briefly addressed.

Amino Acid Sequence↗

P-selectin cell adhesion molecule in inflammation, thrombosis, cancer growth and metastasis.

P-selectin (CD62P) is a member of the selectin family of cell adhesion molecules. It is expressed on stimulated endothelial cells and activated platelets and mediates leukocyte rolling on stimulated endothelial cells and heterotypic aggregation of activated platelets onto leukocytes. It also mediates heterotypic aggregation of activated platelets to cancer cells and adhesion of cancer cells to stimulated endothelial cells. Using P-selectin knockout mice, the importance of P-selectin-mediated cell adhesive interactions in the pathogeneses of inflammation, thrombosis, growth and metastasis of cancers has been clearly demonstrated. Here we will summarize the current knowledge and highlight the important progress in the biomedical research of P-selectin biology, providing a suitable target for therapeutic interventions developed through both experimental and bioinformatic approaches.

Animals↗

Prediction and classification of protein subcellular location-sequence-order effect and pseudo amino acid composition.

Given a protein sequence, how to identify its subcellular location? With the rapid increase in newly found protein sequences entering into databanks, the problem has become more and more important because the function of a protein is closely correlated with its localization. To practically deal with the challenge, a dataset has been established that allows the identification performed among the following 14 subcellular locations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cytoplasm, (5) cytoskeleton, (6) endoplasmic reticulum, (7) extracellular, (8) Golgi apparatus, (9) lysosome, (10) mitochondria, (11) nucleus, (12) peroxisome, (13) plasma membrane, and (14) vacuole. Compared with the datasets constructed by the previous investigators, the current one represents the largest in the scope of localizations covered, and hence many proteins which were totally out of picture in the previous treatments, can now be investigated. Meanwhile, to enhance the potential and flexibility in taking into account the sequence-order effect, the series-mode pseudo-amino-acid-composition has been introduced as a representation for a protein. High success rates are obtained by the re-substitution test, jackknife test, and independent dataset test, respectively. It is anticipated that the current automated method can be developed to a high throughput tool for practical usage in both basic research and pharmaceutical industry.

Algorithms↗

A new hybrid approach to predict subcellular localization of proteins by incorporating gene ontology.

Based on the recent development in the gene ontology and functional domain databases, a new hybridization approach is developed for predicting protein subcellular location by combining the gene product, functional domain, and quasi-sequence-order effects. As a showcase, the same prokaryotic and eukaryotic datasets, which were studied by many previous investigators, are used for demonstration. The overall success rate by the jackknife test for the prokaryotic set is 94.7% and that for the eukaryotic set 92.9%. These are so far the highest success rates achieved for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. The very high success rates also reflect the fact that the subcellular localization of a protein is closely correlated with: (1). the biological objective to which the gene or gene product contributes, (2). the biochemical activity of a gene product, and (3). the place in the cell where a gene product is active.

Algorithms↗