PubMed Health⌕ Search

Biomedical subjects

C J Zheng

Publications and source records attributed to C J Zheng.

At least 19 recordsLinked to original sources

PharmGED: Pharmacogenetic Effect Database.

Prediction and elucidation of pharmacogenetic effects is important for facilitating the development of personalized medicines. Knowledge of polymorphism-induced and other types of drug-response variations is needed for facilitating such studies. Although databases of pharmacogenetic knowledge, polymorphism and toxicogenomic information have appeared, some of the relevant data are provided in separate web-pages and in terms of relatively long descriptions quoted from literatures. To facilitate easy and quick assessment of the relevant information, it is helpful to develop databases that provide all of the information related to a pharmacogenetic effect in the same web-page and in brief descriptions. We developed a database, Pharmacogenetic Effect Database (PharmGED), for providing sequence, function, polymorphism, affected drugs and pharmacogenetic effects. PharmGED can be accessed at http://bidd.cz3.nus.edu.sg/phg/ free of charge for academic use. It currently contains 1825 entries covering 108 disease conditions, 266 distinct proteins, 693 polymorphisms, 414 drugs/ligands cited from 856 references.

Animals↗

Prediction of MHC-binding peptides of flexible lengths from sequence-derived structural and physicochemical properties.

Peptide binding to MHC is critical for antigen recognition by T-cells. To facilitate vaccine design, computational methods have been developed for predicting MHC-binding peptides, which achieve impressive prediction accuracies of 70-90% for binders and 40-80% for non-binders. These methods have been developed for peptides of fixed lengths, for a limited number of alleles, trained from small number of non-binders, and in some cases based straightforwardly on sequence. These limit prediction coverage and accuracy particularly for non-binders. It is desirable to explore methods that predict binders of flexible lengths from sequence-derived physicochemical properties and trained from diverse sets of non-binders. This work explores support vector machines (SVM) as such a method for developing prediction systems of 18 MHC class I and 12 class II alleles by using 4208-3252 binders and 234,333-168,793 non-binders, and evaluated by an independent set of 545-476 binders and 110,564-84,430 non-binders. Binder accuracies are 86-99% for 25 and 70-80% for 5 alleles, non-binder accuracies are 96-99% for 30 alleles. Binder accuracies are comparable and non-binder accuracies substantially improved against other results. Our method correctly predicts 73.3% of the 15 newly-published epitopes in the last 4 months of 2005. Of the 251 recently-published HLA-A*0201 non-epitopes predicted as binders by other methods, 63 are predicted as binders by our method. Screening of HIV-1 genome shows that, compared to other methods, a comparable percentage (75-100%) of its known epitopes is correctly predicted, while a lower percentage (0.01-5% for 24 and 5-8% for 6 alleles) of its constituent peptides are predicted as binders. Our software can be accessed at .

Alleles↗

Prediction of the functional class of lipid binding proteins from sequence-derived properties irrespective of sequence similarity.

Lipid binding proteins play important roles in signaling, regulation, membrane trafficking, immune response, lipid metabolism, and transport. Because of their functional and sequence diversity, it is desirable to explore additional methods for predicting lipid binding proteins irrespective of sequence similarity. This work explores the use of support vector machines (SVMs) as such a method. SVM prediction systems are developed using 14,776 lipid binding and 133,441 nonlipid binding proteins and are evaluated by an independent set of 6,768 lipid binding and 64,761 nonlipid binding proteins. The computed prediction accuracy is 78.9, 79.5, 82.2, 79.5, 84.4, 76.6, 90.6, 79.0, and 89.9% for lipid degradation, lipid metabolism, lipid synthesis, lipid transport, lipid binding, lipopolysaccharide biosynthesis, lipoprotein, lipoyl, and all lipid binding proteins, respectively. The accuracy for the nonmember proteins of each class is 99.9, 99.2, 99.6, 99.8, 99.9, 99.8, 98.5, 99.9, and 97.0%, respectively. Comparable accuracies are obtained when homologous proteins are considered as one, or by using a different SVM kernel function. Our method predicts 86.8% of the 76 lipid binding proteins nonhomologous to any protein in the Swiss-Prot database and 89.0% of the 73 known lipid binding domains as lipid binding. These findings suggest the usefulness of SVMs for facilitating the prediction of lipid binding proteins. Our software can be accessed at the SVMProt server (http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi).

Algorithms↗

Therapeutic targets: progress of their exploration and investigation of their characteristics.

Modern drug discovery is primarily based on the search and subsequent testing of drug candidates acting on a preselected therapeutic target. Progress in genomics, protein structure, proteomics, and disease mechanisms has led to a growing interest in and effort for finding new targets and more effective exploration of existing targets. The number of reported targets of marketed and investigational drugs has significantly increased in the past 8 years. There are 1535 targets collected in the therapeutic target database compared with approximately 500 targets reported in a 1996 review. Knowledge of these targets is helpful for molecular dissection of the mechanism of action of drugs and for predicting features that guide new drug design and the search for new targets. This article summarizes the progress of target exploration and investigates the characteristics of the currently explored targets to analyze their sequence, structure, family representation, pathway association, tissue distribution, and genome location features for finding clues useful for searching for new targets. Possible "rules" to guide the search for druggable proteins and the feasibility of using a statistical learning method for predicting druggable proteins directly from their sequences are discussed.

Adrenergic beta-Antagonists↗

Prediction of compounds with specific pharmacodynamic, pharmacokinetic or toxicological property by statistical learning methods.

Computational methods for predicting compounds of specific pharmacodynamic, pharmacokinetic, or toxicological property are useful for facilitating drug discovery and drug safety evaluation. The quantitative structure-activity relationship (QSAR) and quantitative structure-property relationship (QSPR) methods are the most successfully used statistical learning methods for predicting compounds of specific property. More recently, other statistical learning methods such as neural networks and support vector machines have been explored for predicting compounds of higher structural diversity than those covered by QSAR and QSPR. These methods have shown promising potential in a number of studies. This article is intended to review the strategies, current progresses and underlying difficulties in using statistical learning methods for predicting compounds of specific property. It also evaluates algorithms commonly used for representing structural and physicochemical properties of compounds.

Pharmacokinetics↗

Prediction of functional class of novel plant proteins by a statistical learning method.

In plant genomes, the function of a substantial percentage of the putative protein-coding open reading frames (ORFs) is unknown. These ORFs have no significant sequence similarity to known proteins, which complicates the task of functional study of these proteins. Efforts are being made to explore methods that are complementary to, or may be used in combination with, sequence alignment and clustering methods. A web-based protein functional class prediction software, SVMProt, has shown some capability for predicting functional class of distantly related proteins. Here the usefulness of SVMProt for functional study of novel plant proteins is evaluated. To test SVMProt, 49 plant proteins (without a sequence homolog in the Swiss-Prot protein database, not in the SVMProt training set, and with functional indications provided in the literature) were selected from a comprehensive search of MEDLINE abstracts and Swiss-Prot databases in 1999-2004. These represent unique proteins the function of which, at present, cannot be confidently predicted by sequence alignment and clustering methods. The predicted functional class of 31 proteins was consistent, and that of four other proteins was weakly consistent, with published functions. Overall, the functional class of 71.4% of these proteins was consistent, or weakly consistent, with functional indications described in the literature. SVMProt shows a certain level of ability to provide useful hints about the functions of novel plant proteins with no similarity to known proteins.

Artificial Intelligence↗

Prediction of functional class of novel bacterial proteins without the use of sequence similarity by a statistical learning method.

A substantial percentage of the putative protein-encoding open reading frames (ORFs) in bacterial genomes have no homolog of known function, and their function cannot be confidently assigned on the basis of sequence similarity. Methods not based on sequence similarity are needed and being developed. One method, SVMProt (http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi), predicts protein functional family irrespective of sequence similarity (Nucleic Acids Res. 2003;31:3692-3697). While it has been tested on a large number of proteins, its capability for non-homologous proteins has so far been evaluated for a relatively small number of proteins, and additional tests are needed to more fully assess SVMProt. In this work, 90 novel bacterial proteins (non-homologous to known proteins) are used to evaluate the capability of SVMProt. These proteins are such that none of their homologs are in the Swiss-Prot database, their functions not clearly described in the literature, and they themselves and their homologs are not included in the training sets of SVMProt. They represent proteins whose function cannot be confidently predicted by sequence similarity methods at present. The predicted functional class of 76.7% of each of these proteins shows various levels of consistency with the literature-described function, compared to the overall accuracy of 87% for the SVMProt functional class assignment of 34,582 proteins that have at least one homolog of known function. Our study suggests that SVMProt is capable of assigning functional class for novel bacterial proteins at a level not too much lower than that of sequence alignment methods for homologous proteins.

Artificial Intelligence↗

Trends in exploration of therapeutic targets.

Lead discovery against a preselected therapeutic target is a key component in modern drug development. Continuous effort and increasing interest has been directed at the search for new targets, which has led to the identification of a growing number of them. Data from the therapeutic target database, at http://bidd.nus.edu.sg/group/cjttd/ttd.asp, show that, as of July 2004, the number of documented targets of marketed and investigational drugs has reached 1,174 distinct proteins (including subtypes) and 27 nucleic acids, 239 of which are targets of the marketed drugs. Analysis of these targets, particularly those of recently approved drugs and patented investigational agents, provide useful hints about general trends of target exploration and current focus in drug discovery for the treatment of high impact diseases needing effective or more treatment options.

Databases, Factual↗

MoViES: molecular vibrations evaluation server for analysis of fluctuational dynamics of proteins and nucleic acids.

Analysis of vibrational motions and thermal fluctuational dynamics is a widely used approach for studying structural, dynamic and functional properties of proteins and nucleic acids. Development of a freely accessible web server for computation of vibrational and thermal fluctuational dynamics of biomolecules is thus useful for facilitating the relevant studies. We have developed a computer program for computing vibrational normal modes and thermal fluctuational properties of proteins and nucleic acids and applied it in several studies. In our program, vibrational normal modes are computed by using modified AMBER molecular mechanics force fields, and thermal fluctuational properties are computed by means of a self-consistent harmonic approximation method. A web version of our program, MoViES (Molecular Vibrations Evaluation Server), was set up to facilitate the use of our program to study vibrational dynamics of proteins and nucleic acids. This software was tested on selected proteins, which show that the computed normal modes and thermal fluctuational bond disruption probabilities are consistent with experimental findings and other normal mode computations. MoViES can be accessed at http://ang.cz3.nus.edu.sg/cgi-bin/prog/norm.pl.

Computational Biology↗

TRMP: a database of therapeutically relevant multiple pathways.

UNLABELLED: Disease processes often involve crosstalks between proteins in different pathways. Different proteins have been used as separate therapeutic targets for the same disease. Synergetic targeting of multiple targets has been explored in combination therapy of a number of diseases. Potential harmful interactions of multiple targeting have also been closely studied. To facilitate mechanistic study of drug actions and a more comprehensive understanding the relationship between different targets of the same disease, it is useful to develop a database of known therapeutically relevant multiple pathways (TRMPs). Information about non-target proteins and natural small molecules involved in these pathways also provides useful hint for searching new therapeutic targets and facilitate the understanding of how therapeutic targets interact with other molecules in performing specific tasks. The TRMPs database is designed to provide information about such multiple pathways along with related therapeutic targets, corresponding drugs/ligands, targeted disease conditions, constituent individual pathways, structural and functional information about each protein in the pathways. Cross links to other databases are also introduced to facilitate the access of information about individual pathways and proteins. AVAILABILITY: This database can be accessed at http://bidd.nus.edu.sg/group/trmp/trmp.asp and it currently contains 11 entries of multiple pathways, 97 entries of individual pathways, 120 targets covering 72 disease conditions together with 120 sets of drugs directed at each of these targets. Each entry can be retrieved through multiple methods including multiple pathway name, individual pathway name and disease name. SUPPLEMENTARY INFORMATION: http://bidd.nus.edu.sg/group/trmp/sm.pdf

Biomarkers↗

Threshold distributions of phenylthiocarbamide (PTC) in the Chinese population.

The ability to taste phenylthiocarbamide (PTC) is a well-documented Mendelian trait. Mapping and cloning the gene(s) responsible for the PTC tasting ability would help to delineate the molecular basis for the variations in PTC tasting ability in humans and to shed new light on taste chemosensory functions. In view of the spectacular successes in genome science, the positional cloning strategy seems to be a feasible approach to the isolation of the gene(s) underlying the PTC tasting ability. As a first step toward mapping the gene(s), we collected PTC taste threshold data on 106 individuals, most of them being university students, in Shanghai, China. Using various parametric and nonparametric statistical methods, we have found that the data set is best described by a bimodal distribution. The frequency of PTC nontasters is estimated to be 10%. This is consistent with the view that the PTC nontasting ability follows a recessive mode of inheritance. Several authors had previously reported PTC data on Chinese living outside China. Our data are, to our knowledge, the first ever collected from the Chinese population within China.

Adolescent↗

Comparison of the Green scale versus magnitude estimation for taste perception.

The Green scale is a new psychophysical method that is simple for subjects to use, but its relation with magnitude estimation has yet to be fully characterized. In comparing the consistency between the Green scale and magnitude estimation, we found that the former seems to provide a psychological oral sensation measurement that is different from the latter method. A simple correction formula can be derived.

Humans↗

Development of 124 sequence-tagged sites and cytogenetic localization of 217 cosmids for human chromosome 10.

A total of 124 new chromosome 10-specific sequence-tagged sites (STSs) were derived from two sources: (1) DNA sequences obtained from anonymous clones in new libraries enriched for human chromosome 10 inserts, and (2) published sequences of genes and other loci already known to map to chromosome 10. Libraries were constructed from a somatic cell hybrid carrying human chromosomes 10 and Y. A cosmid library was made from total DNA of the hybrid and probed with labeled total human DNA to identify clones with human DNA inserts. Two hundred seventeen cosmids were mapped to regions of human chromosome 10 by fluorescence in situ hybridization. Twenty-five cosmids represent probes that have been placed on the genetic map previously. One hundred ninety-two cosmids represent new probes that have not been mapped previously. Cosmids carrying inserts with CA repeats were identified by hybridization with a labeled poly(dC-dA)-poly(dG-dT) probe and subcloned to yield microsatellite STS markers. Two small insert plasmid libraries were made, the first by subcloning inserts from a chromosome 10-enriched lambda phage library (LL10NS01) and the second by cloning Alu element-mediated PCR products amplified from hybrid DNA. STSs were generated from the DNA sequences of clone inserts. Chromosome 10-specific STSs were distinguished from Y chromosome STSs by one or both of the following criteria: (1) successful PCR amplification from a template consisting of DNA from another chromosome 10-containing cell line, NA10926B, or (2) FISH localization to chromosome 10 of the source cosmid or of YACs isolated by PCR screening with the STS. These libraries were the source of 90 new chromosome 10-specific STSs, 42 of which contain CA repeats.

Base Sequence↗