PubMed Health⌕ Search

PubMed · 11262977

Determining significant fold differences in gene expression analysis.

Abstract

A typical use for RNA expression microarrays is comparing the measurement of gene expression of two groups. There has not been a study reproducing an entire experiment and modeling the distribution of reproducibility of fold differences. Our goal was to create a model of significance for fold differences, then maximize the number of ESTs above that threshold. Multiple strategies were tested to filter out those ESTs contributing to noise, thus decreasing the requirements of what was needed for significance. We found that even though RNA expression levels appears consistent in duplicate measurements, when entire experiments are duplicated, the calculated fold differences are not as consistent. Thus, it is critically important to repeat as many data points as possible, to ensure that genes and ESTs labeled as significant are truly so. We were successfully able to use duplicated expression measurements to model the duplicated fold differences, and to calculate the levels of fold difference needed to reach significance. This approach can be applied to many other experiments to ascertain significance without a priori assumptions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A J Butte, J Ye, H U Häring, M Stumvoll, M F White, I S Kohane. 2001. Determining significant fold differences in gene expression analysis.. https://doi.org/10.1142/9789814447362_0002

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

EST-SSR based genetic polymorphism among Lablab (Lablab purpureus L. Sweet) accessions contrasting for drought stress at seedling stage.

Lablab is a multipurpose and the most drought-tolerant (DT) crop compared with its relatives. Despite its potential, Lablab is still an underutilized crop with a lack of improved varieties in many countries. The DT (D349, D147, HA4, D363, D352, D359, D348, D311, D55 and D250) and drought-susceptible (DS) (D271, D66, D106, D6, D26, D255, D28, D186, D95, and D258) accessions were earlier identified according to their morphological and biochemical responses to moisture stress at the seedling stage. These accessions were used to establish genetic polymorphism among the accessions contrasting for drought stress based on the Expressed Sequence Tag-Simple Sequence Repeats (EST-SSR) markers. The CTAB protocol was employed for the genomic DNA extraction. After DNA quality and quantity verification, the PCR was conducted using 16 EST-SSR primer pairs specific to the Lablab. The products were separated through the horizontal polyacrylamide gel electrophoresis (hPAGE). Discriminating ability of the markers and primers' efficiency were evaluated based on various genetic parameters. Principal Coordinate Analysis (PCoA) was performed to estimate the distance matrix among the population and among the accessions. While cluster analysis was processed to trace the genetic relationship among the accessions, dendrogram was constructed to decipher their genetic relationship. Analysis of Molecular Variance (AMOVA) was finally computed to quantify the diversity level and genetic relationship among the population, and among the accessions. A low polymorphism (GD = 0.19) was observed between the DT and DS accessions, likely due to limited discriminatory power of the EST-SSR markers. However, the PCoA, cluster analysis and AMOVA identified DT (D147, HA4, and D349) and DS (D106, D95, and D271) accessions as strongly contrasting populations under drought stress, with D147, HA4, D349, D363, D359, D352, and D348 further recommended as DT accessions. Given the low polymorphism observed, further validation using more informative molecular markers and advanced genomic approaches is recommended to improve the identification of drought-tolerance genes and related QTLs to support Lablab breeding programs.

Expressed Sequence Tags↗

Differential gene expression during seed germination in barley (Hordeum vulgare L.).

A barley cDNA macroarray comprising 1,440 unique genes was used to analyze the spatial and temporal patterns of gene expression in embryo, scutellum and endosperm tissue during different stages of germination. Among the set of expressed genes, 69 displayed the highest mRNA level in endosperm tissue, 58 were up-regulated in both embryo and scutellum, 11 were specifically expressed in the embryo and 16 in scutellum tissue. Based on Blast X analyses, 70% of the differentially expressed genes could be assigned a putative function. One set of genes, expressed in both embryo and scutellum tissue, included functions in cell division, protein translation, nucleotide metabolism, carbohydrate metabolism and some transporters. The other set of genes expressed in endosperm encodes several metabolic pathways including carbohydrate and amino acid metabolism as well as protease inhibitors and storage proteins. As shown for a storage protein and a trypsin inhibitor, the endosperm of the germinating barley grain contains a considerable amount of residual mRNA which was produced during seed development and which is degraded during early stages of germination. Based on similar expression patterns in the endosperm tissue, we identified 29 genes which may undergo the same degradation process.

Expressed Sequence Tags↗

Correct identification of genes from serial analysis of gene expression tag sequences.

SAGE (serial analysis of gene expression) is a remarkable technique for genome-wide analysis of gene expression. It is crucial to understand the extent to which SAGE can accurately indicate a gene or expressed sequence tag (EST) with a single tag. We analyzed the effect of the size of SAGE tag on gene identification. Our observation indicates that SAGE tags are in general not long enough to achieve the degree of uniqueness of identification originally envisaged. Our observations also indicate that the limitation of using SAGE tag to identify a gene can be overcome by converting SAGE tags into longer 3' EST sequences with the generation of longer cDNA fragments from SAGE tages for gene identification (GLGI) method.

Expressed Sequence Tags↗