PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “transcriptomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

An isoform-resolution transcriptomic atlas of colorectal cancer from long-read single-cell sequencing.

Colorectal cancer (CRC) ranks as the second leading cause of cancer deaths globally. In recent years, short-read single-cell RNA sequencing (scRNA-seq) has been instrumental in deciphering tumor heterogeneities. However, these studies only enable gene-level quantification but neglect alterations in transcript structures arising from alternative end processing or splicing. In this study, we integrated short- and long-read scRNA-seq of CRC samples to build an isoform-resolution CRC transcriptomic atlas. We identified 394 dysregulated transcript structures in tumor epithelial cells, including 299 resulting from various combinations of splicing events. Second, we characterized genes and isoforms associated with epithelial lineages and subpopulations exhibiting distinct prognoses. Among 31,935 isoforms with novel junctions, 330 were supported by The Cancer Genome Atlas RNA-seq and mass spectrometry data. Finally, we built an algorithm that integrated novel peptides derived from open reading frames of recurrent tumor-specific transcripts with mass spectrometry data and identified recurring neoepitopes that may aid the development of cancer vaccines.

Colorectal Neoplasms↗

The extracellular vesicle transcriptome provides tissue-specific functional genomic annotation relevant to disease susceptibility in obesity.

We characterized circulating extracellular vesicles (EVs) in obese and lean humans, identifying transcriptional cargo differentially expressed in obesity (277 unique genes; false discovery rate < 10%). Since circulating EVs may have broad origin, we compared this obesity EV transcriptome with expression from human visceral-adipose-tissue-derived EVs from freshly collected and cultured biopsies from the same obese individuals, observing high concordance. Using a comprehensive set of adipose-specific epigenomic and chromatin conformation assays, we found that the differentially expressed transcripts from the EVs were those regulated in adipose by body mass index-associated SNPs (p < 5 &#xd7; 10-8) from a large-scale genome-wide association study (GWAS). Using a phenome-wide association study of the regulatory SNPs for the EV-derived transcripts, we identified a substantial enrichment for inflammatory phenotypes, including type 2 diabetes. Collectively, these findings represent the convergence of the GWAS (genetics), epigenomics (transcript regulation), and EV (liquid biopsy) fields, enabling powerful future genomic studies of complex diseases.

Humans↗

Comparative transcriptomic and physiological analyses uncover key regulatory pathways associated with drought tolerance in wheat.

Drought severely limits wheat yield, yet its molecular basis remains incompletely understood. We compared a drought-tolerant line (A25) and a drought-sensitive line (A8) under water deficit across three developmental stages using physiological assays and transcriptomics. A25 exhibited stronger osmotic adjustment and antioxidant defense, with higher proline accumulation and enhanced activities of ascorbate peroxidase, catalase, and other ROS-scavenging enzymes. RNA-seq revealed distinct drought-responsive expression patterns, with differentially expressed genes enriched in MAPK signaling and ABA-dependent pathways. ABA-responsive genes were more abundant and strongly induced in A25, suggesting enhanced ABA signal transduction as a key mechanism. Weighted gene co-expression network analysis identified a drought-associated purple module positively correlated with physiological resistance, from which six hub genes (MAPKKK17, Avr9/Cf-9, RPPL1, RGA1, UBC28, AGPs5) were highlighted as potential regulators. Collectively, coordinated activation of ABA signaling and MAPK cascades, mediated by these hub genes, underlies the robust drought tolerance of A25, providing promising molecular targets for wheat breeding and improvement.

Triticum↗

Integrated transcriptomic and metabolomic analyses reveal key regulators associated with lipid metabolic differences between subcutaneous and visceral adipose tissues in sheep.

The location of fat deposition has a significant impact on meat quality and body health, and different adipose tissues exhibit significant differences in lipid metabolism and immune regulation. This study aimed to systematically compare the phenotypic characteristics, transcriptome, and metabolome of subcutaneous adipose tissue (SAT) and two types of visceral adipose tissue (VAT) in sheep, in order to reveal the metabolic differences between SAT and VAT and their potential regulatory mechanisms. The results showed that compared with VAT, SAT had stronger triglyceride deposition ability and obvious cellular hypertrophy. Through integrative analysis, 15 key lipid metabolism genes and 12 differential metabolites were identified. Among them, ACACA, FASN, ELOVL6, SCD, as well as metabolites palmitic acid and glycerol-3-phosphate, may play a central role in SAT lipid synthesis and storage; whereas IGFBP2, ADRB3, LTA4H, and metabolites arachidonic acid and leukotriene B4 may be involved in the lipolysis regulation and inflammatory response of VAT. These findings may provide deeper insights into the regulatory mechanisms of fat deposition in sheep.

Animals↗

Transcriptomic and metabolomic analyses revealed the action mechanism of nesfatin-1 gene on glucolipid metabolism during early development stage of largemouth bass.

Nesfatin-1 has biological roles including the suppression of food intake and the regulation of glucose and lipid metabolism. However, the information available regarding nesfatin-1 in the glycolipid metabolism in the early development stage of fish is still limited. In order to investigate the role of the nesfatin-1 gene in the early development stage of the largemouth bass (Micropterus salmoides), the nesfatin-1 gene was knocked down using siRNA interference technology. Then, we evaluated its mRNA expression levels, transcriptomes and metabolomes. The mRNA expression levels of nesfatin-1 gene were appreciably decreased at 48&#xa0;h, 72&#xa0;h and 96&#xa0;h after injection of nesfatin-1 siRNA in the early development stage. The omics results revealed that knockdown of the nesfatin-1 gene induced 1833 differentially expressed genes (DEGs) and 2370 differentially expressed metabolites (DEMs). Bioinformatic analysis enriched the most affected molecular pathways (sphingolipid metabolism, fatty acid elongation, amino sugar and nucleotide sugar metabolism and biosynthesis of unsaturated fatty acids) and metabolic pathways (biosynthesis of unsaturated fatty acids, sphingolipid metabolism and amino sugar and nucleotide sugar metabolism) in early development stage of largemouth bass. In amino sugar and nucleotide sugar metabolism, increased expression levels of genes such as chic, chs1, and gck genes, alongside decreased expression levels of the chia.1 gene, resulted in significantly elevated concentrations of N-Acetyl-D-glucosamine, &#x3b2;-d-fructose 6-phosphate, &#x3b2;-d-Fructose, D-mannose 6-phosphate, d-glucose, d-glucose 1-phosphate, UDP-glucose, and UDP-glucuronate, whilst the concentration of UDP-N-acetyl-&#x3b1;-D-glucosamine was markedly reduced. Therefore, the nesfatin-1 gene may influence the early development stage of largemouth bass by affecting signaling pathways associated with glycolipid metabolism. Our findings further expand the understanding of&#xa0;molecular mechanisms of the nesfatin-1 gene, and provide further theoretical support for the initial breeding and feed adaptation of largemouth bass.

Animals↗

A chromosome-level genome assembly and developmental transcriptome profiling reveal stage-specific remodeling of the molecular chaperone system in Helicoverpa armigera.

Helicoverpa armigera is one of the most destructive lepidopteran pests worldwide owing to its remarkable polyphagy, long-distance migration, and rapid adaptation to insecticides. Here, we present a chromosome-level genome assembly of H. armigera generated from a field-collected individual in southwestern China, providing a valuable resource for future population genomic and pangenome studies. Developmental transcriptome analyses of first-instar larvae, fifth-instar larvae, and adults identified 6817, 3519, and 5518 differentially expressed genes, respectively, including 797 shared among all developmental transitions. Functional enrichment and co-expression network analyses revealed extensive transcriptional reprogramming, characterized by coordinated regulation of glycolysis, the tricarboxylic acid (TCA) cycle, and fatty acid &#x3b2;-oxidation, indicating dynamic metabolic remodeling during development. Genome-wide analysis identified 77 heat shock protein (HSP) genes belonging to six subfamilies. These genes were unevenly distributed across chromosomes, with HSP20 members exhibiting extensive tandem duplication. Expression profiling revealed pronounced stage specificity, suggesting progressive remodeling of molecular chaperone networks during development. Early larvae primarily relied on HSP40/HSP60/HSP70 and HSP10/HSP60 chaperone systems; fifth-instar larvae exhibited HSP20-centered proteostasis; and adults predominantly expressed HSP40 together with multiple HSP70 members, accompanied by enrichment of stress response and metamorphosis-related functions. This study provides new insights into developmental transcriptional regulation, metabolic remodeling, and stage-specific specialization of molecular chaperone networks in H. armigera, establishing a foundation for future studies of stress adaptation, population genomic variation, and developmental mechanisms.

Cotton bollworm↗

Integrating GWAS and Transcriptome Analysis Identifies Candidate Genes for Kernel Starch Quality Traits in Maize.

Maize (Zea mays L.) starch quality is a complex trait with significant implications for grain processing and industrial applications. However, the genetic basis underlying starch quality, particularly for gelatinization and thermodynamic properties, remains poorly understood. In this study, we evaluated 12 starch quality traits, including seven gelatinization characteristics, four thermodynamic traits, and kernel starch content (KSC) in a diverse panel of 335 maize inbred lines. Considerable phenotypic variation was observed for all traits. A total of 228 quantitative trait loci (QTLs) were significantly associated with 12 starch quality traits through genome-wide association studies (GWAS). By integrating a dynamic transcriptome analysis of two maize inbred lines with contrasting starch quality, we identified 60 candidate genes. One gene, waxy1, encoding a starch synthase, was found to be associated with enthalpy of gelatinization (&#x394;Hgel) and pasting temperature (Ptemp). Six variants in waxy1 contributed to natural variation in &#x394;Hgel and Ptemp, and a cost-effective InDel and two PARMS-based molecular markers were developed and validated in 144 maize inbred lines, enabling efficient marker-assisted selection. Our findings provide key genes and molecular markers for high-quality maize breeding with improved starch properties.

Zea mays↗

Unbiased screen of human transcriptome reveals an unexpected role of 3'UTRs in translation initiation.

Although most eukaryotic mRNAs require a 5'-cap for translation initiation, some can also be translated through a poorly studied cap-independent pathway. Here we develop a circRNA-based system and unbiasedly identify more than 10,000 sequences in the human transcriptome that contain Cap-independent Translation Initiators (CiTIs). Surprisingly, most of the identified CiTIs are located in 3'UTRs, which mainly promote translation initiation in mRNAs bearing highly structured 5'UTR. Mechanistically, CiTI recruits several translation initiation factors including eIF3 and DHX29, which in turn unwind 5'UTR structures and facilitate ribosome scanning. Functionally, we show that the translation of HIF1A mRNA, an endogenous DHX29 target, is antagonistically regulated by its 5'UTR structure and a new 3'-CiTI in response to hypoxia. Consistently, deletion of 3'-CiTI suppresses cell growth in hypoxia and tumor progression in vivo. Collectively, our study uncovers a new regulatory mode for translation where the 3'UTR actively participate in the translation initiation.

Humans↗

Species-wide quantitative transcriptomes and proteomes reveal distinct genetic control of gene expression variation in yeast.

Gene expression varies between individuals and corresponds to a key step linking genotypes to phenotypes. However, our knowledge regarding the species-wide genetic control of protein abundance, including its dependency on transcript levels, is very limited. Here, we have determined quantitative proteomes of a large population of 942 diverse natural Saccharomyces cerevisiae yeast isolates. We found that mRNA and protein abundances are weakly correlated at the population gene level. While the protein coexpression network recapitulates major biological functions, differential expression patterns reveal proteomic signatures related to specific populations. Comprehensive genetic association analyses highlight that genetic variants associated with variation in protein (pQTL) and transcript (eQTL) levels poorly overlap (3%). Our results demonstrate that transcriptome and proteome are governed by distinct genetic bases, likely explained by protein turnover. It also highlights the importance of integrating these different levels of gene expression to better understand the genotype-phenotype relationship.

Saccharomyces cerevisiae↗

Changes in the transcriptome and synthetic lethal dependencies following KRAS mutant expression reveal profound tissue specificity.

Oncogenic KRAS mutations exhibit a striking tissue-restricted tropism, occurring with high frequency in pancreatic, colorectal, and lung adenocarcinomas while remaining rare in other lineages. The molecular basis for why these specific tissues are uniquely permissive to KRAS transformation, and how this context shapes therapeutic vulnerabilities, remains poorly defined. Here, we utilized CRISPR-mediated genome engineering to generate endogenous, conditional KRAS-mutant isogenic cell line models across three primary permissive lineages (lung, colon, and pancreas) and the nonpermissive breast lineage. Integrated genome-wide CRISPR fitness screens and comparative transcriptome analyses revealed that KRAS-driven synthetic lethal (SL) dependencies are profoundly shaped by their tissue of origin. Strikingly, we observed minimal overlap in SL hits across lineages, with only three genes shared among the permissive lines, suggesting that the KRAS oncogene operates through divergent, context-specific genetic networks. Mechanistically, we show that KRAS activation induces a universal MYC-driven metabolic signature, but the specific machinery required to sustain this state is lineage-restricted. We identified a dependency on the diphthamide synthesis pathway to maintain translational fidelity amid a KRAS-induced hypertranslational state. These findings demonstrate that even when driven by the same oncogene, tumors exhibit distinct regulatory landscapes and unique genetic vulnerabilities. Our results provide a framework for developing lineage-aware therapeutic strategies, moving beyond universal KRAS inhibition toward targeted interventions tailored to a tumor's specific tissue context.

Proto-Oncogene Proteins p21(ras)↗

Ionizing radiation induces bidirectional transcriptomic reprogramming and dynamic NOS2/TREM2 regulation in triple-negative breast cancer cells.

PURPOSE: To characterize irradiation-associated transcriptomic changes in murine triple-negative breast cancer cells and examine dose- and time-response patterns of selected radiation-responsive candidates. MATERIALS AND METHODS: RNA sequencing (RNA-seq) was performed in 4T1 cells collected 24&#x2009;h after 4&#x2009;Gy irradiation, followed by Reactome and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment and gene set enrichment analyses. Representative RNA-seq-derived genes were examined by reverse transcription quantitative PCR (RT-qPCR), and selected immune- and inflammation-related transcripts were further assessed across additional radiation doses and post-irradiation time points. Inducible nitric oxide synthase (NOS2) and triggering receptor expressed on myeloid cells 2 (TREM2) protein abundance was assessed by Western blotting, and nitrite accumulation in culture supernatants was measured using a Griess reagent-based assay as an indirect readout of nitric oxide production. RESULTS: RNA sequencing identified 757 differentially expressed genes, including 285 upregulated and 472 downregulated genes. Irradiation was associated with enrichment of inflammatory, interferon-related, immune-system, and cell-adhesion transcriptional signatures, whereas downregulated genes were enriched in cell-cycle-, chromosome-cohesion-, DNA-damage-response-, DNA-repair-, and SUMOylation-related pathways. Selected immune- and inflammation-related transcripts showed distinct temporal patterns. Nos2 mRNA increased across the examined 0-6&#x2009;Gy dose range and at later post-irradiation time points, whereas NOS2 protein showed different kinetics, with an early peak after 4&#x2009;Gy irradiation and no clear further increase above 6&#x2009;Gy. Nitrite accumulation increased after irradiation. Trem2 showed the largest fold increase among strongly upregulated transcripts identified by RNA-seq, but RT-qPCR detected a significant increase only at 24&#x2009;h, and TREM2 protein abundance remained unchanged across the examined doses and time points. CONCLUSIONS: Ionizing radiation was associated with broad bidirectional transcriptional remodeling in 4T1 cells, involving immune-, inflammatory-, and interferon-related signatures together with reduced representation of cell-cycle- and DNA-repair-related gene sets. The discordant mRNA and protein patterns of NOS2 and TREM2 indicate that transcript-level responses do not necessarily translate into corresponding protein-level changes. These findings define irradiation-associated molecular responses requiring further functional investigation.

Triple-negative breast cancer↗

Integrated transcriptomic and functional characterization of Claudin-1 reveals its oncogenic and immunomodulatory roles in pancreatic ductal adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) remains among the deadliest malignancies, driven by its invasive nature and lack of effective biomarkers. Disruption of the epithelial barrier, mediated by tight junction components, is a critical yet underexplored contributor to PDAC progression. Claudins, integral regulators of tight junction integrity, display altered expression across cancers, but their prognostic and immunomodulatory roles in PDAC remain unclear. We performed an integrative analysis of 177 RNA-Seq datasets from TCGA and GTEx to characterize Claudin family alterations in PDAC. Differential expression, copy number variation, methylation, and co-expression networks were analyzed alongside clinical and survival data. Prognostic significance was assessed using Kaplan - Meier and Cox regression analyses, while immune cell infiltration was examined using deconvolution algorithms. Functional validation of Claudin-1 was conducted in Capan-1 cells using CRISPR/Cas9 knockout, followed by proliferation, wound-healing, and Western blot assays. Ten Claudin genes were significantly dysregulated, with Claudin-1 and Claudin-4 frequently amplified and associated with advanced stage and poor survival. High Claudin-1 expression correlated with reduced immune infiltration, indicating an immune-excluded phenotype characterized by immune cells retained in the tumor stroma but largely absent from the tumor parenchyma. Claudin-1 knockout markedly inhibited proliferation, migration, and EMT, evidenced by downregulation of Snail and Slug and restoration of E-cadherin expression. This integrative transcriptomic and functional study identifies Claudin-1 as a key driver of PDAC aggressiveness and immune modulation. These findings establish Claudin-1 as a promising prognostic biomarker and therapeutic target for restoring epithelial integrity and counteracting immune evasion in pancreatic cancer.

Humans↗

Integrating single-cell transcriptomics to construct an oncogene-driven prognostic model and elucidate metabolic-immune crosstalk in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related deaths, its progression and treatment heterogeneity are mainly influenced by driver gene and tumor micro-environment (TME) interactions. Nevertheless, the mechanisms of this process at the single-cell level remain unclear. This study integrated TCGA and multi-center single-cell transcriptome data to identify a 575 genes HCC-specific core set, developing a single-cell "oncogene scoring" system to quantify individual carcinogenic activity. This score is significantly elevated in malignant and proliferative T cells and is closely associated with metabolic reprogramming, aberrant cell&#x2012;cell communication, and immunosuppressive phenotypes. Based on these characteristics, we constructed a machine learning-based Random Survival Forest (RSF) prognostic model validated in multiple independent cohorts, which classifies patients into distinct risk subtypes. The high-risk group exhibits genomic instability, increased tumor stemness, and immune evasion, while the low-risk group was more sensitive to drugs such as sorafenib. This study highlights the potential pathways by which high oncogenic activity is associated with HCC progression, suggesting a profound link with single-cell metabolic&#x2012;immune crosstalk. The constructed RSF model offers a promising computational framework for risk stratification and provides hypothesis-generating insights that may inform future personalized treatment strategies for HCC patients.

Hepatocellular carcinoma↗

SpatialRNA: a Python package for easy application of Graph Neural Network models on single-molecule spatial transcriptomics dataset.

SUMMARY: Image-based spatial transcriptomics (iST) deliver gene expression measurements of RNA transcripts in tissue slices with single-molecule resolution and spatial context preserved. Modern Graph Neural Network (GNN) models are promising methods for capturing the complex molecular and cellular phenotypes in tissues at single-transcript and single-cell levels. A key application of GNNs is the detection of spatial domains or niches, that is, groups of molecules and/or cells that collaboratively work together to produce complex phenotypes. Due to the vast number of detected transcripts in (iST) dataset, applying GNNs on RNA molecule graphs is not trivial. We present a Python package, SpatialRNA, for easy (sub)graph generation from tissue samples and provide comprehensive tutorials for convenient and efficient application of Graph Neural Network models under the PyG framework. This highly scalable tool comprehensively segments tissue into spatial domains, aiding in biological interpretation of iST data and its underlying molecular microenvironments. AVAILABILITY AND IMPLEMENTATION: The SpatialRNA package is freely accessible from online repository https://github.com/ruqianl/spatialrna and can be installed via pip. Comprehensive tutorials, guidance on parameter selection, and complete workflows of case studies are available from the documentation website https://ruqianl.github.io/spatialrna_docs/, and uploaded on Zenodo with a DOI 10.5281/zenodo.17339575.

Neural Networks, Computer↗

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.

MOTIVATION: Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP technologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. RESULTS: Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis. AVAILABILITY AND IMPLEMENTATION: The source code of stDyer-image and detailed tutorials are available at https://github.com/ericcombiolab/stDyer-image.

Proteomics↗

Identification of immune cell type-specific susceptibility genes in multiple cancers using transcriptome-wide association studies.

BACKGROUND: Transcriptome-wide association studies (TWAS) integrate gene expression and genome-wide association studies (GWAS) to identify disease susceptibility genes. Because gene expression varies substantially across cell types within tissues, cell type-specific prediction models may enhance the power of TWAS. METHODS: We conducted cell type-specific TWAS leveraging single-cell RNA sequencing data from the OneK1K cohort (14 immune cell types, 1.27 million cells) and GWAS summary statistics for 7 cancers (>290&#x200a;000 cases in total). To improve prediction accuracy, we developed a modeling framework that incorporates shared gene expression effects across cell types. RESULTS: At a false discovery rate of 5%, we identified 106 (Bonferroni 5%: 13) previously unreported loci for breast cancer, 51 (4) loci for prostate cancer, 11 (4) loci for lung cancer, 39 (5) loci for melanoma, 9 (1) loci for ovarian cancer, and 2 (1) loci for diffuse large B-cell lymphoma, with most genes exhibiting cell type specificity. Gene set analyses confirmed joint associations of unreported genes with breast and prostate cancer risk in UK Biobank data. Additional lung tissue single-cell RNA sequencing data with 113 individuals validated 18 of 32 (56.3%) statistically significant genes for lung cancer. Across cancers, 139 statistically significant genes were shared by at least 2 cancer types and were primarily enriched in specific immune cell types. CONCLUSION: Cell type-specific TWAS improve the identification of novel cancer susceptibility loci and provide insights into the immune landscape of cancer etiology.

Humans↗

Precision-Based Filtering Facilitates Cross-Referencing of Conventional and Single-Nucleus Transcriptomes to Identify Time- and Temperature-Sensitive Cell Populations.

Transcriptome analysis via RNA sequencing (RNAseq) has become a ubiquitous method of molecular characterization from whole organisms, dissected tissues, and single cells. These experiments continue to provide an extraordinary volume of data describing molecular states and responses to many conditions. However, standard approaches to RNAseq analysis commonly use expression level filters that eliminate potentially useful data in the service of decreasing noise. Here we describe the implementation of a coefficient of variation-based filter for RNAseq gene expression data. This filter prioritizes consistent data across replicates, allowing lowly-expressed genes with low-variation measurements to be retained for downstream analysis. We show, using two independent Arabidopsis RNAseq datasets, that this filter allows for the inclusion of many more transcription factors than even a low-stringency expression level filter. This effect is independent of sequencing depth. We find that these lowly-expressed genes mark specific cell clusters in our single-nucleus (sn)RNAseq dataset and may facilitate future characterization of currently unknown cell types or states. We further characterize communities of co-expressed genes, sampled across the day at two growth temperatures, in relation to snRNAseq cell clusters, finding evidence for a highly photosynthetic cell population, and a cell state marked by high cell division and translation. These methods can be expanded to RNAseq analysis in many systems, facilitating the construction of more detailed models of tissue-specific gene regulatory networks.

Transcriptome analysis↗