PubMed Health⌕ Search

Biomedical subjects

Cheng-Yan Kao

Publications and source records attributed to Cheng-Yan Kao.

8 recordsLinked to original sources

Integrated minimum-set primers and unique probe design algorithms for differential detection on symptom-related pathogens.

MOTIVATION: Differential detection on symptom-related pathogens (SRP) is critical for fast identification and accurate control against epidemic diseases. Conventional polymerase chain reaction (PCR) requires a large number of unique primers to amplify selected SRP target sequences. With multiple-use primers (mu-primers), multiple targets can be amplified and detected in one PCR experiment under standard reaction condition and reduced detection complexity. However, the time complexity of designing mu-primers with the best heuristic method available is too vast. We have formulated minimum-set mu-primer design problem as a set covering problem (SCP), and used modified compact genetic algorithm (MCGA) to solve this problem optimally and efficiently. We have also proposed new strategies of primer/probe design algorithm (PDA) on combining both minimum-set (MS) mu-primers and unique (UniQ) probes. Designed primer/probe set by PDA-MS/UniQ can amplify multiple genes simultaneously upon physical presence with minimum-set mu-primer amplification (MMA) before intended differential detection with probes-array hybridization (PAH) on the selected target set of SRP. RESULTS: The proposed PDA-MS/UniQ method pursues a much smaller number of primers set compared with conventional PCR. In the simulation experiment for amplifying 12 669 target sequences, the performance of our method with 68% reduction on required mu-primers number seems to be superior to the compared heuristic approaches in both computation efficiency and reduction percentage. Our integrated PDA-MS/UniQ method is applied to the differential detection on 9 plant viruses from 4 genera with MMA and PAH of 11 mu-primers instead of 18 unique ones in conventional PCR while amplifying overall 9 target sequences. The results of wet lab experiments with integrated MMA-PAH system have successfully validated the specificity and sensitivity of the primers/probes designed with our integrated PDA-MS/UniQ method.

Algorithms↗

Improving disulfide connectivity prediction with sequential distance between oxidized cysteines.

SUMMARY: Predicting disulfide connectivity precisely helps towards the solution of protein structure prediction. In this study, a descriptor derived from the sequential distance between oxidized cysteines (denoted as DOC) is proposed. An approach using support vector machine (SVM) method based on weighted graph matching was further developed to predict the disulfide connectivity pattern in proteins. When DOC was applied, prediction accuracy of 63% for our SVM models could be achieved, which is significantly higher than those obtained from previous approaches. The results show that using the non-local descriptor DOC coupled with local sequence profiles significantly improves the prediction accuracy. These improvements demonstrate that DOC, with a proper scaling scheme, is an effective feature for the prediction of disulfide connectivity. The method developed in this work is available at the web server PreCys (prediction of cys-cys linkages of proteins).

Chymotrypsinogen↗

A stochastic differential equation model for quantifying transcriptional regulatory network in Saccharomyces cerevisiae.

MOTIVATION: The explosion of microarray studies has promised to shed light on the temporal expression patterns of thousands of genes simultaneously. However, available methods are far from adequate in efficiently extracting useful information to aid in a greater understanding of transcriptional regulatory network. Biological systems have been modeled as dynamic systems for a long history, such as genetic networks and cell regulatory network. This study evaluated if the stochastic differential equation (SDE), which is prominent for modeling dynamic diffusion process originating from the irregular Brownian motion, can be applied in modeling the transcriptional regulatory network in Saccharomyces cerevisiae. RESULTS: To model the time-continuous gene-expression datasets, a model of SDE is applied to depict irregular patterns. Our goal is to fit a generalized linear model by combining putative regulators to estimate the transcriptional pattern of a target gene. Goodness-of-fit is evaluated by log-likelihood and Akaike Information Criterion. Moreover, estimations of the contribution of regulators and inference of transcriptional pattern are implemented by statistical approaches. Our SDE model is basic but the test results agree well with the observed dynamic expression patterns. It implies that advanced SDE model might be perfectly suited to portray transcriptional regulatory networks. AVAILABILITY: The R code is available on request. CONTACT: cykao@csie.ntu.edu.tw SUPPLEMENTARY INFORMATION: http://www.csie.ntu.edu.tw/~b89x035/yeast/

Gene Expression Regulation↗

Cysteine separations profiles on protein sequences infer disulfide connectivity.

MOTIVATION: Disulfide bonds play an important role in protein folding. A precise prediction of disulfide connectivity can strongly reduce the conformational search space and increase the accuracy in protein structure prediction. Conventional disulfide connectivity predictions use sequence information, and prediction accuracy is limited. Here, by using an alternative scheme with global information for disulfide connectivity prediction, higher performance is obtained with respect to other approaches. RESULT: Cysteine separation profiles have been used to predict the disulfide connectivity of proteins. The separations among oxidized cysteine residues on a protein sequence have been encoded into vectors named cysteine separation profiles (CSPs). Through comparisons of their CSPs, the disulfide connectivity of a test protein is inferred from a non-redundant template set. For non-redundant proteins in SwissProt 39 (SP39) sharing less than 30% sequence identity, the prediction accuracy of a fourfold cross-validation is 49%. The prediction accuracy of disulfide connectivity for proteins in SwissProt 43 (SP43) is even higher (53%). The relationship between the similarity of CSPs and the prediction accuracy is also discussed. The method proposed in this work is relatively simple and can generate higher accuracies compared to conventional methods. It may be also combined with other algorithms for further improvements in protein structure prediction. AVAILABILITY: The program and datasets are available from the authors upon request. CONTACT: cykao@csie.ntu.edu.tw.

Algorithms↗

POINT: a database for the prediction of protein-protein interactions based on the orthologous interactome.

One possible path towards understanding the biological function of a target protein is through the discovery of how it interfaces within protein-protein interaction networks. The goal of this study was to create a virtual protein-protein interaction model using the concepts of orthologous conservation (or interologs) to elucidate the interacting networks of a particular target protein. POINT (the prediction of interactome database) is a functional database for the prediction of the human protein-protein interactome based on available orthologous interactome datasets. POINT integrates several publicly accessible databases, with emphasis placed on the extraction of a large quantity of mouse, fruit fly, worm and yeast protein-protein interactions datasets from the Database of Interacting Proteins (DIP), followed by conversion of them into a predicted human interactome. In addition, protein-protein interactions require both temporal synchronicity and precise spatial proximity. POINT therefore also incorporates correlated mRNA expression clusters obtained from cell cycle microarray databases and subcellular localization from Gene Ontology to further pinpoint the likelihood of biological relevance of each predicted interacting sets of protein partners.

Animals↗

An evolutionary approach for gene expression patterns.

This study presents an evolutionary algorithm, called a heterogeneous selection genetic algorithm (HeSGA), for analyzing the patterns of gene expression on microarray data. Microarray technologies have provided the means to monitor the expression levels of a large number of genes simultaneously. Gene clustering and gene ordering are important in analyzing a large body of microarray expression data. The proposed method simultaneously solves gene clustering and gene-ordering problems by integrating global and local search mechanisms. Clustering and ordering information is used to identify functionally related genes and to infer genetic networks from immense microarray expression data. HeSGA was tested on eight test microarray datasets, ranging in size from 147 to 6221 genes. The experimental clustering and visual results indicate that HeSGA not only ordered genes smoothly but also grouped genes with similar gene expressions. Visualized results and a new scoring function that references predefined functional categories were employed to confirm the biological interpretations of results yielded using HeSGA and other methods. These results indicate that HeSGA has potential in analyzing gene expression patterns.

Algorithms↗

An evolutionary algorithm for large traveling salesman problems.

This work proposes an evolutionary algorithm, called the heterogeneous selection evolutionary algorithm (HeSEA), for solving large traveling salesman problems (TSP). The strengths and limitations of numerous well-known genetic operators are first analyzed, along with local search methods for TSPs from their solution qualities and mechanisms for preserving and adding edges. Based on this analysis, a new approach, HeSEA is proposed which integrates edge assembly crossover (EAX) and Lin-Kernighan (LK) local search, through family competition and heterogeneous pairing selection. This study demonstrates experimentally that EAX and LK can compensate for each other's disadvantages. Family competition and heterogeneous pairing selections are used to maintain the diversity of the population, which is especially useful for evolutionary algorithms in solving large TSPs. The proposed method was evaluated on 16 well-known TSPs in which the numbers of cities range from 318 to 13509. Experimental results indicate that HeSEA performs well and is very competitive with other approaches. The proposed method can determine the optimum path when the number of cities is under 10,000 and the mean solution quality is within 0.0074% above the optimum for each test problem. These findings imply that the proposed method can find tours robustly with a fixed small population and a limited family competition length in reasonable time, when used to solve large TSPs.

Algorithms↗

GEM: a Gaussian Evolutionary Method for predicting protein side-chain conformations.

We have developed an evolutionary approach to predicting protein side-chain conformations. This approach, referred to as the Gaussian Evolutionary Method (GEM), combines both discrete and continuous global search mechanisms. The former helps speed up convergence by reducing the size of rotamer space, whereas the latter, integrating decreasing-based Gaussian mutations and self-adaptive Gaussian mutations, continuously adapts dihedrals to optimal conformations. We tested our approach on 38 proteins ranging in size from 46 to 325 residues and showed that the results were comparable to those using other methods. The average accuracies of our predictions were 80% for chi(1), 66% for chi(1 + 2), and 1.36 A for the root mean square deviation of side-chain positions. We found that if our scoring function was perfect, the prediction accuracy was also essentially perfect. However, perfect prediction could not be achieved if only a discrete search mechanism was applied. These results suggest that GEM is robust and can be used to examine the factors limiting the accuracy of protein side-chain prediction methods. Furthermore, it can be used to systematically evaluate and thus improve scoring functions.

Algorithms↗