PubMed Health⌕ Search

PubMed · 12097349

Interactive exploration of microarray gene expression patterns in a reduced dimensional space.

Abstract

The very high dimensional space of gene expression measurements obtained by DNA microarrays impedes the detection of underlying patterns in gene expression data and the identification of discriminatory genes. In this paper we show the use of projection methods such as principal components analysis (PCA) to obtain a direct link between patterns in the genes and patterns in samples. This feature is useful in the initial interactive pattern exploration of gene expression data and data-driven learning of the nature and types of samples. Using oligonucleotide microarray measurements of 40 samples from different normal human tissues, we show that distinct patterns are obtained when the genes are projected on a two-dimensional plane spanned by the loadings of the two major principal components. These patterns define the particular genes associated with a sample class (i.e., tissue). When used separately from the other genes, these class-specific (i.e., tissue-specific) genes in turn define distinct tissue patterns in the projection space spanned by the scores of the two major principal components. In this study, PCA projection facilitated discriminatory gene selection for different tissues and identified tissue-specific gene expression signatures for liver, skeletal muscle, and brain samples. Furthermore, it allowed the classification of nine new samples belonging to these three types using the linear combination of the expression levels of the tissue-specific genes determined from the first set of samples. The application of the technique to other published data sets is also discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jatin Misra, William Schmitt, Daehee Hwang, Li-Li Hsiao, Steve Gullans, George Stephanopoulos, Gregory Stephanopoulos. 2002. Interactive exploration of microarray gene expression patterns in a reduced dimensional space.. https://doi.org/10.1101/gr.225302

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling↗

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling↗

PoweREST: Statistical power estimation for spatial transcriptomics experiments to detect differentially expressed genes between two conditions.

Recent advancements in spatial transcriptomics (ST) have significantly enhanced biological research in various domains. However, the high cost for current ST data generation techniques restricts the large-scale application of ST. Consequently, maximization of the use of available resources to achieve robust statistical power for ST data is a pressing need. One fundamental question in ST analysis is detection of differentially expressed genes (DEGs) under different conditions using ST data. Such DEG analyses are performed frequently, but their power calculations are rarely discussed in the literature. To address this gap, we developed PoweREST, a power estimation tool designed to support the power calculation for DEG detection with 10X Genomics Visium data. PoweREST enables power estimation both before any ST experiments and after preliminary data are collected, making it suitable for a wide variety of power analyses in ST studies. We also provide a user-friendly, program-free web application that allows users to interactively calculate and visualize study power along with relevant parameters.

Gene Expression Profiling↗