PubMed Health⌕ Search

PubMed · 12507787

Partitioning large-sample microarray-based gene expression profiles using principal components analysis.

Abstract

Principal components analysis (PCA) is useful for reproducing the total variation among hundreds or thousands of continuously-scaled variables with a much smaller number of unobservable variables called 'latent factors'. The CLUSFAVOR computer program was used to implement PCA for identifying groups of genes with similar expression profiles from a large number of genes used on DNA microarrays. This paper describes the principal components solution to the factor model of the correlation matrix R, calculation of eigenvalues and eigenvectors of R, extraction of factors, and calculation of factor loadings and identification of genes with similar loading patterns to construct groups of genes with similar expression profiles. With regard to extraction of factors, it was found that more than 90% of the total variance in input data could be accounted for by extracting factors whose eigenvalues exceed unity. Bipolar factors containing strong positive and negative loadings can also be used for identifying two unique groups of genes, since expression profiles of genes that load positive are unlike expression profiles of genes that load negative on the same factor. While PCA does not provide the absolute answer to a multidimensional problem, it nevertheless can provide a heuristic with which natural groupings of genes with similar expression profiles can be assembled. While cluster analysis essentially generates a single dendogram (tree branch) containing every gene in the input data, PCA can be used to assemble gene expression profiles that strongly correlate with the latent factors accounting for a majority of total variance. Example results for CLUSFAVOR computer program runs are provided.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Leif E Peterson. 2003. Partitioning large-sample microarray-based gene expression profiles using principal components analysis.. https://doi.org/10.1016/s0169-2607(02)00009-3

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling↗

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling↗

PoweREST: Statistical power estimation for spatial transcriptomics experiments to detect differentially expressed genes between two conditions.

Recent advancements in spatial transcriptomics (ST) have significantly enhanced biological research in various domains. However, the high cost for current ST data generation techniques restricts the large-scale application of ST. Consequently, maximization of the use of available resources to achieve robust statistical power for ST data is a pressing need. One fundamental question in ST analysis is detection of differentially expressed genes (DEGs) under different conditions using ST data. Such DEG analyses are performed frequently, but their power calculations are rarely discussed in the literature. To address this gap, we developed PoweREST, a power estimation tool designed to support the power calculation for DEG detection with 10X Genomics Visium data. PoweREST enables power estimation both before any ST experiments and after preliminary data are collected, making it suitable for a wide variety of power analyses in ST studies. We also provide a user-friendly, program-free web application that allows users to interactively calculate and visualize study power along with relevant parameters.

Gene Expression Profiling↗