PubMed Health⌕ Search

Biomedical subjects

Yuanyuan Ding

Publications and source records attributed to Yuanyuan Ding.

2 recordsLinked to original sources

Improving the performance of SVM-RFE to select genes in microarray data.

BACKGROUND: Recursive Feature Elimination is a common and well-studied method for reducing the number of attributes used for further analysis or development of prediction models. The effectiveness of the RFE algorithm is generally considered excellent, but the primary obstacle in using it is the amount of computational power required. RESULTS: Here we introduce a variant of RFE which employs ideas from simulated annealing. The goal of the algorithm is to improve the computational performance of recursive feature elimination by eliminating chunks of features at a time with as little effect on the quality of the reduced feature set as possible. The algorithm has been tested on several large gene expression data sets. The RFE algorithm is implemented using a Support Vector Machine to assist in identifying the least useful gene(s) to eliminate. CONCLUSION: The algorithm is simple and efficient and generates a set of attributes that is very similar to the set produced by RFE.

Algorithms↗

The effect of normalization on microarray data analysis.

This paper contains a description of several common normalization methods used in microarray analysis, and compares the effect of these methods on microarray data. The importance of background subtraction is also addressed. The research focuses on three parts. The first uses three statistical methods: t-test, Wilcoxon signed rank test, and sign test to measure the difference between background subtracted data and nonbackground subtracted data. The second part of the study uses the same three statistical methods to compare whether data normalized with different normalization methods yield similar results. The third part of the study focuses on whether these differently normalized data will influence the result of gene selection (dimension reduction). The comparisons are done for several data sets to help identify similarity patterns. The conclusion of this study is that background subtraction can make a difference, especially for some data sets with poorer quality data. The choice of normalization method, for the most part, makes little difference in the sense that the methods produce similarly normalized data. But, based on the third part of analysis, we found that when gene selection is performed on these differently normalized data, somewhat different gene sets are obtained. Thus, the choice of normalization method will likely have some effect on the final analysis.

Oligonucleotide Array Sequence Analysis↗