PubMed Health⌕ Search

Biomedical subjects

K-C Chou

Publications and source records attributed to K-C Chou.

9 recordsLinked to original sources

Predicting secretory protein signal sequence cleavage sites by fusing the marks of global alignments.

A newly synthesized secretory protein in cells bears a special sequence, called signal peptide or sequence, which plays the role of "address tag" in guiding the protein to wherever it is needed. Such a unique function of signal sequences has stimulated novel strategies for drug design or reprogramming cells for gene therapy. To realize these new ideas and plans, however, it is important to develop an automated method for fast and accurately identifying the signal sequences or their cleavage sites. In this paper, a new method is developed for predicting the signal sequence of a query secretory protein by fusing the results from a series of global alignments through a voting system. The very high success rates thus obtained suggest that the novel approach is very promising, and that the new method may become a useful vehicle in identifying signal sequence, or at least serve as a complementary tool to the existing algorithms of this field.

Algorithms↗

Using ensemble classifier to identify membrane protein types.

Predicting membrane protein type is both an important and challenging topic in current molecular and cellular biology. This is because knowledge of membrane protein type often provides useful clues for determining, or sheds light upon, the function of an uncharacterized membrane protein. With the explosion of newly-found protein sequences in the post-genomic era, it is in a great demand to develop a computational method for fast and reliably identifying the types of membrane proteins according to their primary sequences. In this paper, a novel classifier, the so-called "ensemble classifier", was introduced. It is formed by fusing a set of nearest neighbor (NN) classifiers, each of which is defined in a different pseudo amino acid composition space. The type for a query protein is determined by the outcome of voting among these constituent individual classifiers. It was demonstrated through the self-consistency test, jackknife test, and independent dataset test that the ensemble classifier outperformed other existing classifiers widely used in biological literatures. It is anticipated that the idea of ensemble classifier can also be used to improve the prediction quality in classifying other attributes of proteins according to their sequences.

Algorithms↗

Virtual screening for finding natural inhibitor against cathepsin-L for SARS therapy.

Recently Simmons et al. reported a new mechanism for SARS virus entry into target cells, where MDL28170 was identified as an efficient inhibitor of CTSL-meditated substrate cleavage with IC(50) of 2.5 nmol/l. Based on the molecule fingerprint searching method, 11 natural molecules were found in the Traditional Chinese Medicines Database (TCMD). Molecular simulation indicates that the MOL376 (a compound derived from a Chinese medicine herb with the therapeutic efficacy on the human body such as relieving cough, removing the phlegm, and relieving asthma) has not only the highest binding energy with the receptor but also the good match in geometric conformation. It was observed through docking studies that the van der Waals interactions made substantial contributions to the affinity, and that the receptor active pocket was too large for MDL21870 but more suitable for MOL736. Accordingly, MOL736 might possibly become a promising lead compound for CTSL inhibition for SARS therapy.

Algorithms↗

Anti-SARS drug screening by molecular docking.

Starting from a collection of 1386 druggable compounds obtained from the 3D pharmacophore search, we performed a similarity search to narrow down the scope of docking studies. The template molecule is KZ7088 (Chou et al., 2003, Biochem Biophys Res Commun 308: 148-151). The MDL MACCS keys were used to fingerprint the molecules. The Tanimoto coefficient is taken as the metric to compare fingerprints. If the similarity threshold was 0.8, a set of 50 unique hits and 103 conformers were retrieved as a result of similarity search. The AutoDock 3.011 was used to carry out molecular docking of 50 ligands to their macromolecular protein receptors. Three compounds, i.e., C(28)H(34)O(4)N(7)Cl, C(21)H(36)O(5)N(6), and C(21)H(36)O(5)N(6), were found that may be promising candidates for further investigation. The main feature shared by these three potential inhibitors as well as the information of the involved side chains of SARS Cov Mpro may provide useful insights for the development of potent inhibitors against SARS enzyme.

Antiviral Agents↗

Using cellular automata images and pseudo amino acid composition to predict protein subcellular location.

The avalanche of newly found protein sequences in the post-genomic era has motivated and challenged us to develop an automated method that can rapidly and accurately predict the localization of an uncharacterized protein in cells because the knowledge thus obtained can greatly speed up the process in finding its biological functions. However, it is very difficult to establish such a desired predictor by acquiring the key statistical information buried in a pile of extremely complicated and highly variable sequences. In this paper, based on the concept of the pseudo amino acid composition (Chou, K. C. PROTEINS: Structure, Function, and Genetics, 2001, 43: 246-255), the approach of cellular automata image is introduced to cope with this problem. Many important features, which are originally hidden in the long amino acid sequences, can be clearly displayed through their cellular automata images. One of the remarkable merits by doing so is that many image recognition tools can be straightforwardly applied to the target aimed here. High success rates were observed through the self-consistency, jackknife, and independent dataset tests, respectively.

Algorithms↗

Using pseudo amino acid composition to predict protein subcellular location: approached with Lyapunov index, Bessel function, and Chebyshev filter.

With the avalanche of new protein sequences we are facing in the post-genomic era, it is vitally important to develop an automated method for fast and accurately determining the subcellular location of uncharacterized proteins. In this article, based on the concept of pseudo amino acid composition (Chou, K.C. Proteins: Structure, Function, and Genetics, 2001, 43: 246-255), three pseudo amino acid components are introduced via Lyapunov index, Bessel function, Chebyshev filter that can be more efficiently used to deal with the chaos and complexity in protein sequences, leading to a higher success rate in predicting protein subcellular location.

Amino Acid Sequence↗

Using string kernel to predict signal peptide cleavage site based on subsite coupling model.

Owing to the importance of signal peptides for studying the molecular mechanisms of genetic diseases, reprogramming cells for gene therapy, and finding new drugs for healing a specific defect, it is in great demand to develop a fast and accurate method to identify the signal peptides. Introduction of the so-called {-3,-1, +1} coupling model (Chou, K. C.: Protein Engineering, 2001, 14-2, 75-79) has made it possible to take into account the coupling effect among some key subsites and hence can significantly enhance the prediction quality of peptide cleavage site. Based on the subsite coupling model, a kind of string kernels for protein sequence is introduced. Integrating the biologically relevant prior knowledge, the constructed string kernels can thus be used by any kernel-based method. A Support vector machines (SVM) is thus built to predict the cleavage site of signal peptides from the protein sequences. The current approach is compared with the classical weight matrix method. At small false positive ratios, our method outperforms the classical weight matrix method, indicating the current approach may at least serve as a powerful complemental tool to other existing methods for predicting the signal peptide cleavage site. The software that generated the results reported in this paper is available upon requirement, and will appear at http://www.pami.sjtu.edu.cn/wm.

Models, Genetic↗

Using cellular automata to generate image representation for biological sequences.

A novel approach to visualize biological sequences is developed based on cellular automata (Wolfram, S. Nature 1984, 311, 419-424), a set of discrete dynamical systems in which space and time are discrete. By transforming the symbolic sequence codes into the digital codes, and using some optimal space-time evolvement rules of cellular automata, a biological sequence can be represented by a unique image, the so-called cellular automata image. Many important features, which are originally hidden in a long and complicated biological sequence, can be clearly revealed thru its cellular automata image. With biological sequences entering into databanks rapidly increasing in the post-genomic era, it is anticipated that the cellular automata image will become a very useful vehicle for investigation into their key features, identification of their function, as well as revelation of their "fingerprint". It is anticipated that by using the concept of the pseudo amino acid composition (Chou, K.C. Proteins: Structure, Function, and Genetics, 2001, 43, 246-255), the cellular automata image approach can also be used to improve the quality of predicting protein attributes, such as structural class and subcellular location.

Animals↗

Using complexity measure factor to predict protein subcellular location.

Recent advances in large-scale genome sequencing have led to the rapid accumulation of amino acid sequences of proteins whose functions are unknown. Because the functions of these proteins are closely correlated with their subcellular localizations, it is vitally important to develop an automated method as a high-throughput tool to timely identify their subcellular location. Based on the concept of the pseudo amino acid composition by which a considerable amount of sequence-order effects can be incorporated into a set of discrete numbers (Chou, K. C., Proteins: Structure, Function, and Genetics, 2001, 43: 246-255), the complexity measure approach is introduced. The advantage by incorporating the complexity measure factor as one of the pseudo amino acid components for a protein is that it can more effectively reflect its overall sequence-order feature than the conventional correlation factors. With such a formulation frame to represent the samples of protein sequences, the covariant-discriminant predictor (Chou, K. C. and Elrod, D. W., Protein Engineering, 1999, 12: 107-118) was adopted to conduct prediction. High success rates were obtained by both the jackknife cross-validation test and independent dataset test, suggesting that introduction of the concept of the complexity measure into prediction of protein subcellular location is quite promising, and might also hold a great potential as a useful vehicle for the other areas of molecular biology.

Algorithms↗