PubMed Health⌕ Search

Biomedical subjects

Sihua Peng

Publications and source records attributed to Sihua Peng.

3 recordsLinked to original sources

An exploration of 3'-end processing signals and their tissue distribution in Oryza sativa.

The 3' untranslated regions deeply affect many properties of eukaryotic mRNA. In plants, the polyadenine control signals contained in these regions seem to be more variable than of mammals. Three cDNA libraries derived from the leaf, endosperm and stem tissues of rice were sequenced from the 3'-end. Of the 9911 transcripts analyzed, 5723 unique transcripts were identified from the leaf sequences, 2934 from the endosperm and 1254 from the stem. The information entropy and two statistical methods were used to compile a list of rice poly(A) control signals. Based on their distribution, these signals can be roughly grouped into far-upstream element (FUE), near-upstream element (NUE), T-rich region (TRE) and downstream element (DE). The distribution of rice conserved regions is similar to the previous model from Arabidopsis and yeast, with a few differences in word constructions. Interestingly, we also found the word distributions were diverse in the cleavage site of downstream sequences of different rice tissues. The signal bias in downstream sequences may lead mRNA to be differently cleaved in different rice tissues.

3' Untranslated Regions↗

Multiclass cancer classification and biomarker discovery using GA-based algorithms.

MOTIVATION: The development of microarray-based high-throughput gene profiling has led to the hope that this technology could provide an efficient and accurate means of diagnosing and classifying tumors, as well as predicting prognoses and effective treatments. However, the large amount of data generated by microarrays requires effective reduction of discriminant gene features into reliable sets of tumor biomarkers for such multiclass tumor discrimination. The availability of reliable sets of biomarkers, especially serum biomarkers, should have a major impact on our understanding and treatment of cancer. RESULTS: We have combined genetic algorithm (GA) and all paired (AP) support vector machine (SVM) methods for multiclass cancer categorization. Predictive features can be automatically determined through iterative GA/SVM, leading to very compact sets of non-redundant cancer-relevant genes with the best classification performance reported to date. Interestingly, these different classifier sets harbor only modest overlapping gene features but have similar levels of accuracy in leave-one-out cross-validations (LOOCV). Further characterization of these optimal tumor discriminant features, including the use of nearest shrunken centroids (NSC), analysis of annotations and literature text mining, reveals previously unappreciated tumor subclasses and a series of genes that could be used as cancer biomarkers. With this approach, we believe that microarray-based multiclass molecular analysis can be an effective tool for cancer biomarker discovery and subsequent molecular cancer diagnosis.

Algorithms↗

Molecular classification of cancer types from microarray data using the combination of genetic algorithms and support vector machines.

Simultaneous multiclass classification of tumor types is essential for future clinical implementations of microarray-based cancer diagnosis. In this study, we have combined genetic algorithms (GAs) and all paired support vector machines (SVMs) for multiclass cancer identification. The predictive features have been selected through iterative SVMs/GAs, and recursive feature elimination post-processing steps, leading to a very compact cancer-related predictive gene set. Leave-one-out cross-validations yielded accuracies of 87.93% for the eight-class and 85.19% for the fourteen-class cancer classifications, outperforming the results derived from previously published methods.

Algorithms↗