PubMed Health⌕ Search

Biomedical subjects

Ueng-Cheng Yang

Publications and source records attributed to Ueng-Cheng Yang.

3 recordsLinked to original sources

Multi-class clustering and prediction in the analysis of microarray data.

DNA microarray technology provides tools for studying the expression profiles of a large number of distinct genes simultaneously. This technology has been applied to sample clustering and sample prediction. Because of a large number of genes measured, many of the genes in the original data set are irrelevant to the analysis. Selection of discriminatory genes is critical to the accuracy of clustering and prediction. This paper considers statistical significance testing approach to selecting discriminatory gene sets for multi-class clustering and prediction of experimental samples. A toxicogenomic data set with nine treatments (a control and eight metals, As, Cd, Ni, Cr, Sb, Pb, Cu, and AsV with a total of 55 samples) is used to illustrate a general framework of the approach. Among four selected gene sets, a gene set omega(I) formed by the intersection of the F-test and the set of the union of one-versus-all t-tests performs the best in terms of clustering as well as prediction. Hierarchical and two modified partition (k-means) methods all show that the set omega(I) is able to group the 55 samples into seven clusters reasonably well, in which the As and AsV samples are considered as one cluster (the same group) as are the Cd and Cu samples. With respect to prediction, the overall accuracy for the gene set omega(I) using the nearest neighbors algorithm to predict 55 samples into one of the nine treatments is 85%.

Algorithms↗

Gene selection for sample classifications in microarray experiments.

DNA microarray technology provides useful tools for profiling global gene expression patterns in different cell/tissue samples. One major challenge is the large number of genes relative to the number of samples. The use of all genes can suppress or reduce the performance of a classification rule due to the noise of nondiscriminatory genes. Selection of an optimal subset from the original gene set becomes an important prestep in sample classification. In this study, we propose a family-wise error (FWE) rate approach to selection of discriminatory genes for two-sample or multiple-sample classification. The FWE approach controls the probability of the number of one or more false positives at a prespecified level. A public colon cancer data set is used to evaluate the performance of the proposed approach for the two classification methods: k nearest neighbors (k-NN) and support vector machine (SVM). The selected gene sets from the proposed procedure appears to perform better than or comparable to several results reported in the literature using the univariate analysis without performing multivariate search. In addition, we apply the FWE approach to a toxicogenomic data set with nine treatments (a control and eight metals, As, Cd, Ni, Cr, Sb, Pb, Cu, and AsV) for a total of 55 samples for a multisample classification. Two gene sets are considered: the gene set omegaF formed by the ANOVA F-test, and a gene set omegaT formed by the union of one-versus-all t-tests. The predicted accuracies are evaluated using the internal and external crossvalidation. Using the SVM classification, the overall accuracies to predict 55 samples into one of the nine treatments are above 80% for internal crossvalidation. OmegaF has slightly higher accuracy rates than omegaT. The overall predicted accuracies are above 70% for the external crossvalidation; the two gene sets omegaT and omegaF performed equally well.

Colonic Neoplasms↗

Description of the transcriptomes of immune response-activated hemocytes from the mosquito vectors Aedes aegypti and Armigeres subalbatus.

Mosquito-borne diseases, including dengue, malaria, and lymphatic filariasis, exact a devastating toll on global health and economics, killing or debilitating millions every year (54). Mosquito innate immune responses are at the forefront of concerted research efforts aimed at defining potential target genes that could be manipulated to engineer pathogen resistance in vector populations. We aimed to describe the pivotal role that circulating blood cells (called hemocytes) play in immunity by generating a total of 11,952 Aedes aegypti and 12,790 Armigeres subalbatus expressed sequence tag (EST) sequences from immune response-activated hemocyte libraries. These ESTs collapsed into 2,686 and 2,107 EST clusters, respectively. The clusters were used to adapt the web-based interface for annotating bacterial genomes called A Systematic Annotation Package for Community Analysis of Genomes (ASAP) for analysis of ESTs. Each cluster was categorically characterized and annotated in ASAP based on sequence similarity to five sequence databases. The sequence data and annotations can be viewed in ASAP at https://asap.ahabs.wisc.edu/annotation/php/ASAP1.htm. The data presented here represent the results of the first high-throughput in vivo analysis of the transcriptome of immunocytes from an invertebrate. Among the sequences are those for numerous immunity-related genes, many of which parallel those employed in vertebrate innate immunity, that have never been described for these mosquitoes. The sequences and annotations presented in this paper have been submitted to GenBank under accession numbers AY 431103 to AY 433788 (Aedes aegypti) and AY 439334 to AY 441440 (Armigeres subalbatus).

Aedes↗