PubMed Health⌕ Search

PubMed · 14597310

Normalization of cDNA microarray data.

Abstract

Normalization means to adjust microarray data for effects which arise from variation in the technology rather than from biological differences between the RNA samples or between the printed probes. This paper describes normalization methods based on the fact that dye balance typically varies with spot intensity and with spatial position on the array. Print-tip loess normalization provides a well-tested general purpose normalization method which has given good results on a wide range of arrays. The method may be refined by using quality weights for individual spots. The method is best combined with diagnostic plots of the data which display the spatial and intensity trends. When diagnostic plots show that biases still remain in the data after normalization, further normalization steps such as plate-order normalization or scale-normalization between the arrays may be undertaken. Composite normalization may be used when control spots are available which are known to be not differentially expressed. Variations on loess normalization include global loess normalization and two-dimensional normalization. Detailed commands are given to implement the normalization techniques using freely available software.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gordon K Smyth, Terry Speed. 2003. Normalization of cDNA microarray data.. https://doi.org/10.1016/s1046-2023(03)00155-5

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Sequence evaluation of four specific cDNA libraries for developmental genomics of sunflower.

Four different cDNA libraries were constructed from sunflower protoplasts growing under embryogenic and non-embryogenic conditions: one standard library from each condition and two subtractive libraries in opposite sense. A total of 22,876 cDNA clones were obtained and 4800 ESTs were sequenced, giving rise to 2479 high quality ESTs representing an unigene set of 1502 sequences. This set was compared with ESTs represented in public databases using the programs BLASTN and BLASTX, and its members were classified according to putative function using the catalog in the Kyoto Encyclopedia of Genes and Genomes (KEGG). Some 33% of sequences failed to align with existing plant ESTs and therefore represent putative novel genes. The libraries show a low level of redundancy and, on average, 50% of the present ESTs have not been previously reported for sunflower. Several potentially interesting genes were identified, based on their homology with genes involved in animal zygotic division or plant embryogenesis. We also identified two ESTs that show significantly different levels of expression under embryogenic and non-embryogenic conditions. The libraries described here represent an original and valuable resource for the discovery of yet unknown genes putatively involved in dicot embryogenesis and improving our knowledge of the mechanisms involved in polarity acquisition by plant embryos.

DNA, Complementary↗

Improving the statistical detection of regulated genes from microarray data using intensity-based variance estimation.

BACKGROUND: Gene microarray technology provides the ability to study the regulation of thousands of genes simultaneously, but its potential is limited without an estimate of the statistical significance of the observed changes in gene expression. Due to the large number of genes being tested and the comparatively small number of array replicates (e.g., N = 3), standard statistical methods such as the Student's t-test fail to produce reliable results. Two other statistical approaches commonly used to improve significance estimates are a penalized t-test and a Z-test using intensity-dependent variance estimates. RESULTS: The performance of these approaches is compared using a dataset of 23 replicates, and a new implementation of the Z-test is introduced that pools together variance estimates of genes with similar minimum intensity. Significance estimates based on 3 replicate arrays are calculated using each statistical technique, and their accuracy is evaluated by comparing them to a reliable estimate based on the remaining 20 replicates. The reproducibility of each test statistic is evaluated by applying it to multiple, independent sets of 3 replicate arrays. Two implementations of a Z-test using intensity-dependent variance produce more reproducible results than two implementations of a penalized t-test. Furthermore, the minimum intensity-based Z-statistic demonstrates higher accuracy and higher or equal precision than all other statistical techniques tested. CONCLUSION: An intensity-based variance estimation technique provides one simple, effective approach that can improve p-value estimates for differentially regulated genes derived from replicated microarray datasets. Implementations of the Z-test algorithms are available at http://vessels.bwh.harvard.edu/software/papers/bmcg2004.

DNA, Complementary↗

Application of the split-ubiquitin membrane yeast two-hybrid system to investigate membrane protein interactions.

The characterization of protein-protein interactions provides the foundation for further studies concerning protein complex function and regulation. Since the advent of the yeast two-hybrid assay, many additional genetic systems based upon the principle of protein fragment complementation have been designed. One such system, the split-ubiquitin membrane yeast two-hybrid system (MbYTH), is able to analyze the interaction status between two integral membrane proteins. This ability of the MbYTH system augments genetic analysis of protein interactions by covering for the inherent limitation of the yeast two-hybrid system when studying membrane protein interactions. Herein, we provide a description of the MbYTH method and detailed protocols in order to monitor protein interactions and discover novel interacting partners using the MbYTH system.

DNA, Complementary↗