PubMed Health⌕ Search

PubMed · 15585121

The effect of normalization on microarray data analysis.

Abstract

This paper contains a description of several common normalization methods used in microarray analysis, and compares the effect of these methods on microarray data. The importance of background subtraction is also addressed. The research focuses on three parts. The first uses three statistical methods: t-test, Wilcoxon signed rank test, and sign test to measure the difference between background subtracted data and nonbackground subtracted data. The second part of the study uses the same three statistical methods to compare whether data normalized with different normalization methods yield similar results. The third part of the study focuses on whether these differently normalized data will influence the result of gene selection (dimension reduction). The comparisons are done for several data sets to help identify similarity patterns. The conclusion of this study is that background subtraction can make a difference, especially for some data sets with poorer quality data. The choice of normalization method, for the most part, makes little difference in the sense that the methods produce similarly normalized data. But, based on the third part of analysis, we found that when gene selection is performed on these differently normalized data, somewhat different gene sets are obtained. Thus, the choice of normalization method will likely have some effect on the final analysis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuanyuan Ding, Dawn Wilkins. 2004. The effect of normalization on microarray data analysis.. https://doi.org/10.1089/dna.2004.23.635

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Identification of Lhp1p-associated RNAs by microarray analysis in Saccharomyces cerevisiae reveals association with coding and noncoding RNAs.

La is a conserved eukaryotic RNA-binding protein best known for its role in the biogenesis of noncoding RNAs transcribed by RNA polymerase III. To broaden our understanding of the function of the La homologous protein (Lhp1) in Saccharomyces cerevisiae, we have taken a genomics approach. Lhp1 ribonucleoprotein complexes were immunoprecipitated and bound RNAs were examined by hybridization to whole-genome microarrays that include >6,000 ORFs, documented noncoding RNAs, and the intervening intergenic regions. Demonstrating the validity of this approach, associations with previously known Lhp1p-associated RNAs were detected and associations with additional noncoding RNAs, including multiple tRNAs and small nucleolar RNAs, were revealed. Indicating that this approach provides a robust method for discovering RNAs, the data also identify associations between Lhp1p and several intergenic regions, three of which encode the recently annotated putative snoRNAs: RUF1, RUF2, and RUF3. Unexpectedly, we find that Lhp1p is also associated with a subset of coding mRNAs. These mRNAs include many ribosomal protein transcripts as well as the mRNA encoding Hac1p, a transcription factor required during the unfolded protein stress response. In cells lacking LHP1, Hac1p levels are decreased 2- to 3-fold, whereas no changes are detected in the levels of spliced or unspliced HAC1 mRNA or in the stability of Hac1p. Finally, although LHP1 is dispensable for growth under standard conditions, we find that it is required when the unfolded protein response is induced at elevated temperatures. These results suggest that Lhp1p may play a novel role in the translation of one or more cellular mRNAs.

Oligonucleotide Array Sequence Analysis↗

Simulated annealing of microarray data reduces noise and enables cross-experimental comparisons.

Microarrays are a powerful tool for assessing the genome-wide induction of a transcriptional response to internal or external stimuli, but are not considered quantitatively rigorous (i.e., the signal intensity of hybridized probe is normally used to quantify relative transcript abundance). Thus, it is difficult, if not impossible, to accurately compare separate microarray experiments without a reference standard. However, even among replicated microarray experiments, each gene varies significantly in the amount of signal detected, suggesting no single gene would be appropriate as a standard. We propose and test a method to "align" experimental transcription profiles to a set of reference experiments using simulated annealing (SA), essentially using the relative positions of all genes as a reference standard. SA attempts to find a globally optimal adjustment factor for the relative expression level of each experimental gene expression signal, given a previously observed range of gene expression measurements. By defining a relative dynamic range of gene expression under control conditions for all genes, we can more accurately compare transcription profiles between separate experiments and, potentially, between species--enabling comparative transcriptomics. Testing SA on a published dataset, we find that it significantly reduces interexperimental variation, suggesting it holds promise to accomplish this goal.

Oligonucleotide Array Sequence Analysis↗

Microarray truths and consequences.

For many, analysis of a microarray experiment starts with a spreadsheet of expression levels. While great attention is duly paid to RNA extraction, preparation and hybridization, relatively little care is devoted to extraction of expression levels from the fluorescent image. By delegating this step to a click of the mouse the exact extraction process is masked and researchers may be unwittingly compromising their data. In this review, we describe the most common mistakes committed on the path from the image to the spreadsheet and their impact on data quality. Remedies are further proposed for most of the popular microarray platforms in use today.

Oligonucleotide Array Sequence Analysis↗