PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Single-cell RNA-seq”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

106 records · Page 6Linked to original sources

OmicsTweezer: A distribution-independent cell deconvolution model for multi-omics Data.

Cell deconvolution estimates cell type proportions from bulk omics data, enabling insights into tissue microenvironments and disease. However, practical applications are often hindered by batch effects between bulk data and referenced single-cell data, a challenge that is frequently overlooked. To address this discrepancy, we developed OmicsTweezer, a distribution-independent cell deconvolution model. By integrating optimal transport with deep learning, OmicsTweezer aligns simulated and real data in a shared latent space, effectively mitigating data shifts and inter-omics distribution differences. OmicsTweezer is versatile, capable of deconvolving bulk RNA-seq, bulk proteomics, and spatial transcriptomics. Extensive evaluations on simulated and real-world datasets demonstrate its robustness and accuracy. Furthermore, applications in prostate and colon cancer showcase OmicsTweezer's ability to identify biologically meaningful cell types. As a unified deconvolution framework for multi-omics data, OmicsTweezer offers an efficient and powerful tool for studying disease microenvironments.

Humans↗

Polygenic enrichment analysis in multi-omics levels identifies cell/tissue specific associations with schizophrenia based on single-cell RNA sequencing data.

OBJECTIVE: Understanding the specific cellular origin and tissue heterogeneity in schizophrenia is critically important for exploring the disease etiology. This study aims to investigate these aspects by performing multiple analyses based on omics data. METHOD: We performed single-cell disease relevance score (scDRS) algorithm to link brain single-cell RNA sequencing (scRNA-seq) with schizophrenia risk across multi-omics scales at single-cell resolution. This approach identified cell types with overexpression of schizophrenia-related genes implicated by multi-omics panels (ATAC-seq, RNA-seq, TWAS, and GWAS). Schizophrenia-related genes from these multi-omics panels were extracted and combined with scRNA-seq data to calculate scDRS. Subsequently, the cell-type vs. disease association and tissue heterogeneity were assessed using scDRS for each omics panel. RESULTS: We identified two novel cell subpopulations in the brain that differentially express SCUBE3 (59 cells, 7.0 %) and FN1 (21 cells, 2.5 %). At the individual cell level, schizophrenia-associated cell subpopulations included microglial cell associated with ATAC-seq panel (Passociation = 0.002, Pheterogeneity = 0.009) and deep layer neuron suggestively associated with GWAS panel (Passociation = 0.033, Pheterogeneity = 0.017). At the brain tissue level, microglial cell was significantly associated with cortical plate in ATAC-seq panel (Passociation = 0.002, Pheterogeneity = 0.011). Gene level analysis identified several genes associated with schizophrenia across multi-omics panels. CONCLUSIONS: Our study outlines the signature of cell subpopulations, brain regions, and disease risk genes in schizophrenia at single-cell resolution across multi-omics scales. These findings provide a reference for future precision medicine approaches targeting specific cell types and brain regions in schizophrenia.

Schizophrenia↗

Integrated single-cell and bulk transcriptomic analysis identifies a novel senescent fibroblast subtype associated with poor prognosis in acral melanoma.

BACKGROUND: Acral melanoma (AM) exhibits significant intratumoral heterogeneity, but its tumor microenvironment (TME) and immune regulation remain unclear. This study aims to dissect TME heterogeneity and establish a prognostic model based on key cell subpopulations. METHODS: We collected AM single-cell RNA sequencing (scRNA-seq) and bulk RNA-seq data from the Gene Expression Omnibus (GEO) and the Cancer Genome Atlas (TCGA). Unsupervised clustering, CellChat, and Scissor analysis were performed to characterize cellular heterogeneity, cell-cell communication, and prognosis-related cell subpopulations. Kaplan-Meier analysis was used to assess the prognostic value of key genes, which were further validated by multiplex immunohistochemistry (mIHC). RESULTS: In AM, Mel_C2, C7, and C9 with high SEMA6A and KIT expression were strongly linked to poor prognosis. We further identified a senescent fibroblast subpopulation (sCAF_CDKN2A) characterized by high fibroblast senescence signature (FSS) scores. Integrating Scissor analysis of fibroblast subtypes with bulk prognostic data, we identified COL3A1, VCAN, and KIT as prognosis-associated genes upregulated in poor-outcome-related fibroblast subsets. Cell-cell communication analysis revealed that sCAF_CDKN2A engages in an immunosuppressive network, interacting with regulatory T cells (Tregs) via MIF signaling and receiving signals from exhausted CD8+ T cells through PPIA-BSG interactions. Using transcription factor expression patterns from these fibroblast subtypes, we constructed a prognostic model that effectively stratified patients into distinct risk groups with significant differences in overall survival (OS). mIHC confirmed significantly higher protein levels of SEMA6A and COL3A1 in tumor tissues compared to matched normal tissues. CONCLUSIONS: We established a novel prognostic model for AM and identified sCAF_CDKN2A as an immunosuppressive senescent fibroblast subpopulation driving poor prognosis.

Acral melanoma↗

Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model.

MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.

Algorithms↗

Nested co-expression network analysis identifies compact gene clusters in a black box.

MOTIVATION: Digital analysis of biological systems requires methods capable of identifying both broad and nested gene modules reflecting complex biological processes. Existing transcriptomic methods often miss compact gene sets corresponding to subprocesses in specialized cell types, limiting insights into functional heterogeneity. RESULTS: We present Nested-WGCNA, a two-stage unsupervised network analysis algorithm designed to identify coarse-grained and fine-grained gene modules. Applied to bulk RNA-Seq data, Nested-WGCNA reveals stable modules reproducible across datasets. When validated against scRNA-Seq data, these modules correspond to both major and minor immune cell subtypes. Application to immunotherapy response datasets uncovers predictive and prognostic biomarkers, highlighting its utility in treatment stratification and biomarker discovery. AVAILABILITY: The NestedWGCNA source code and analysis pipeline are available on GitHub (https://github.com/ilyada/NestedWGCNA) and archived on Zenodo (https://doi.org/10.5281/zenodo.18959244).

Algorithms↗

Predicting gene-specific regulation with transcriptomic and epigenetic single-cell data.

MOTIVATION: Analysis of single cell ATAC-seq and RNA-seq data has allowed to gain unprecedented insights into gene regulation by allowing to define cell type-specific regulatory regions and their effects on gene expression. While powerful, such analysis is challenging due to the inherent sparsity of single cell data. RESULTS: We present a new approach, MetaFR, to learn gene-specific models that link open-chromatin variation from scATAC-seq data to gene expression from scRNA-seq. Using efficient regression trees, we illustrate that accurate expression prediction models can be learned on the single-cell or meta-cell level. Validation was done using fine-mapped eQTLs. Meta-cell models were found to outperform single-cell models for most genes. Comparison to the SOTA method SCARlink revealed advantages of MetaFR in terms of runtime and prediction performance. MetaFR thus allows time-efficient analysis and obtains reliable models of gene expression prediction, which can be used to study gene regulation in any organism for which scRNA-seq and scATAC-seq data is available. AVAILABILITY AND IMPLEMENTATION: MetaFR is available under https://github.com/SchulzLab/MetaFR.

Single-Cell Analysis↗

Unsupervised multiscale clustering of single-cell transcriptomes to identify hierarchical structures of cell subtypes.

BACKGROUND: Cell clustering is an essential step in uncovering cellular architectures in single-cell RNA sequencing (scRNA-seq) data. However, the existing cell clustering approaches are not well designed to dissect complex structures of cellular landscapes at a finer resolution. RESULTS: Here, we develop a multiscale clustering (MSC) approach to construct a sparse cell-cell correlation network for unsupervised identification of de novo cell types and subtypes across multiple resolutions. Based upon simulated silver- and gold-standard data as well as real scRNA-seq data in diseases, MSC demonstrates significantly improved performance compared to established benchmark methods and reveals a biologically meaningful cell hierarchy to facilitate the discovery of novel disease-associated cell subtypes and mechanisms. CONCLUSIONS: We present MSC as a new single-cell multiscale clustering framework as a powerful tool for advancing discoveries in disease-associated cell populations using single-cell sequencing data.

Single-Cell Analysis↗

Cell-type-specific response to silicon treatment in soybean leaves revealed by single-nucleus RNA sequencing and targeted gene editing.

Mineral nutrient uptake and deposition profoundly influence plant development, stress resilience, and productivity. Silicon (Si), though classified as a non-essential element, significantly influences a plant's physiology, particularly in fortifying defense responses and mitigating stress. While the genetic and molecular mechanisms of Si uptake and transport are well studied in monocots, particularly rice, their role in dicot species, such as soybean, remains unclear at the cellular and molecular levels. In this study, we utilized single-nucleus RNA sequencing (snRNA-seq) to dissect cellular responses to Si accumulation in soybean leaves. We identified distinct cellular populations, including a unique Si-induced or Si-associated cell cluster within vascular cells, suggesting a specialized mechanism of Si distribution. Si treatment notably induced the expression of defense-related genes, with a pronounced enrichment in vascular cells, underscoring their pivotal role in activating plant defense mechanisms. Moreover, Si modulated the expression of genes involved in phytoalexin biosynthesis, salicylic acid, and immune receptor signaling, suggesting transcriptional priming of genes involved in defense responses. Further investigation of Si transporters revealed precise expression of an Si efflux gene in epidermal cells in response to Si treatment. We also validated the role of efflux Si transporters using a Xenopus oocyte assay and CRISPR/Cas9 genome editing of composite soybean plant roots. This study provides critical insights into the biotic stress regulatory networks influenced by Si treatment in soybean leaves at the single-cell level, thus laying the foundation for enhancing stress tolerance through optimized mineral nutrient uptake.

Glycine max↗

Multi-omics characterization of a GPRC5A+ epithelial subpopulation associated with malignant features in colorectal cancer.

BACKGROUND: Colorectal cancer (CRC) exhibits marked cellular heterogeneity, and the cellular context of malignancy-associated epithelial programs remains incompletely defined. METHODS: We integrated 2,993 CRC samples spanning bulk RNA-seq (n = 2,568; two OS/RFS cohorts), scRNA-seq (281,961 cells/152 specimens), spatial transcriptomics (n = 6), and proteomics (n = 267). Analyses included single-cell integration/annotation, GSVA/HALLMARK, interactome, pseudotime, and ligand-receptor mapping; functional CRISPR assays, EMT immunoblotting, and xenografts; TF profiling (SCENIC/JASPAR/ChIP-qPCR); and exploratory drug-response prediction (OncoPredict), cell-sensitivity assays, and docking/MD modeling. RESULTS: We constructed a stage-stratified single-cell atlas and resolved eleven malignant epithelial subsets, characterizing Epi_4 as late-stage-enriched with EMT, hypoxia, and inflammatory programs and adverse OS/RFS. GPRC5A marked this subset, which we define as GPRC5A+Epi; its expression rose from stage I→IV and was associated with poor outcomes across cohorts, with concordant spatial/proteomic observations. GPRC5A perturbation affected CRC proliferation, migration/invasion, EMT, and xenograft tumorigenicity, supporting a functionally important role in the tested models. SCENIC and ChIP-qPCR supported FOSL1 as an upstream regulator that occupies the GPRC5A promoter. Spatial and ligand-receptor analyses predicted close association and potentially reciprocal signaling between GPRC5A+Epi and POSTN+fibroblasts (COL1A1-SDC4, COL1A1/1A2-ITGA2/ITGB1, PPIA-BSG); concurrent high GPRC5A+Epi/POSTN+Fib signatures were associated with inferior OS/RFS. Drug-response analyses identified an association between GPRC5A status and trametinib sensitivity. Docking/MD produced a computational model of a possible trametinib-GPRC5A interaction, which remains experimentally unvalidated. CONCLUSIONS: GPRC5A⁺Epi is a malignancy-associated epithelial state in CRC, and GPRC5A is functionally important for malignant phenotypes in the tested models. Its inferred relationships with POSTN⁺ fibroblasts and the trametinib findings should be regarded as hypothesis-generating pending functional crosstalk, direct-binding, and therapeutic validation.

Humans↗

Malignant epithelial states drive immune dysfunction in ampulla of Vater carcinoma.

BACKGROUND: Ampulla of Vater (AoV) carcinoma is a rare malignancy arising at the junction of intestinal and pancreatobiliary epithelium. Its heterogeneous clinical behavior and histological diversity have hindered therapeutic advances, and the cellular basis of this heterogeneity remains unclear. We aimed to construct a single-cell transcriptomic atlas of AoV carcinoma, with a focus on identifying epithelial subtypes and their interactions with the tumor microenvironment (TME). METHODS: We performed single-cell RNA sequencing on eight primary AoV tumors and four matched normal tissues. Comprehensive clustering and transcriptomic analyses identified cell-type composition, epithelial heterogeneity, and tumor-immune interactions. Findings were validated using deconvolution of bulk RNA-seq data from 62 AoV carcinoma patients. Results Malignant epithelial cells were categorized into four distinct subtypes: Int-Wnt, PB-KRAS, Int-Hypoxia, and Cycling stage. PB-KRAS cells exhibited stem-like transcriptional programs and high genomic instability. Deconvolution analysis of bulk RNA-seq data from the independent AoV cohort revealed that enrichment of the PB-KRAS subtype correlated with tumor recurrence and poor survival. Our immune profiling analysis discovered a significant association between PB-KRAS subtype and GZMK+ CD8+ T cells, which are in a pre-dysfunctional state, alongside SPP1+ macrophages exhibiting immunosuppressive traits. Spatial transcriptome data further supports the immunosuppressive natures of TME around PB-KRAS subtype malignant epithelial cells in AoV carcinoma. CONCLUSIONS: Our study presents a single-cell atlas of AoV carcinoma, highlighting the molecular diversity of malignant epithelium and its association with the immune microenvironment. The PB-KRAS subtype emerges as a stem-like, immunosuppressive tumor state associated with poor prognosis, providing insights for future therapeutic targeting.

Ampulla of Vater carcinoma↗

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans↗

Integrated single-cell transcriptomics, Mendelian randomization, and machine learning identify CEBPZ as an immune-related biomarker in oral lichen planus.

BACKGROUND: Oral lichen planus (OLP) is a chronic, immune-mediated oral mucosal disease with complex pathophysiology and potential for malignant transformation. Understanding its molecular basis is critical for the development of precise diagnostic and therapeutic strategies. OBJECTIVES: We aimed to identify key immune-related biomarkers and characterize cellular dynamics in OLP, with a particular focus on the role of CEBPZ in disease pathogenesis. MATERIAL AND METHODS: We analyzed single-cell RNA sequencing (scRNA-seq) data from OLP lamina propria samples (GSE211630) to identify disease-specific T-cell subpopulations using high-dimensional weighted gene co-expression network analysis (hdWGCNA) for oxidative stress-related gene modules.-data-based Mendelian randomization (SMR) integrated FinnGen genome-wide association study (GWAS; 342,499 Europeans) data with Genotype-Tissue Expression (GTEx) expression quantitative trait loci (eQTL) data to identify causal genes. Machine learning (ML) models (least absolute shrinkage and selection operator (LASSO) and convolutional neural network (CNN)) were developed using bulk RNA-seq datasets (GSE52130 and GSE38616) for diagnostic purposes. RESULTS: We identified OLP-specific T-cell populations (clusters 0, 3, 5, 7, 13, and 15) with enhanced migration inhibition factor (MIF) pathway signaling toward B cells and monocytes. Two oxidative stress-associated modules contained hub genes, including CEBPZ. Summary-data-based Mendelian randomization analysis identified 231 OLP-associated genes, with CEBPZ uniquely intersecting LASSO-selected markers (odds ratio (OR) = 1.057, 95% confidence interval (95% CI) = 1.013-1.102, p = 0.010). Machine learning models achieved area under the curve (AUC) values ranging from 0.653 to 0.745, with the CNN model reaching a validation accuracy of 0.735. CEBPZ showed elevated expression in OLP T cells and correlated with enhanced MIF-(CD74+CXCR4) signaling. CONCLUSIONS: This integrative approach identifies CEBPZ as a pivotal biomarker linking genetic susceptibility, oxidative stress, and immune dysregulation in OLP. Our diagnostic models offer promising tools for OLP management.

CEBPZ↗

Non-structural maintenance of chromosome condensin I complex subunit H knockdown suppresses malignant progression of esophageal squamous cell carcinoma via the Wnt/β-catenin signaling pathway.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) remains a major cause of cancer-related mortality, and effective therapeutic targets are still limited. Non-structural maintenance of chromosome condensin I complex subunit H (NCAPH) has been implicated in tumorigenesis; however, its clinical relevance, functional roles, and underlying mechanisms in ESCC are not fully defined. We aimed to characterize the expression pattern, prognostic value, biological functions, and mechanistic basis of NCAPH in ESCC. METHODS: Public datasets from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) were analyzed to evaluate NCAPH expression and clinical associations. Single-cell RNA sequencing (scRNA-seq) data were used to map cell-type-specific distribution of NCAPH in tumor and adjacent tissues. NCAPH was silenced in KYSE150 and KYSE510 cells using lentiviral short hairpin RNAs (shRNAs), followed by Cell Counting Kit-8 (CCK-8), colony formation, wound-healing, and Transwell migration/invasion assays. A nude mouse xenograft model was established to assess the effect of NCAPH knockdown in vivo. RNA sequencing (RNA-seq), quantitative polymerase chain reaction (qPCR), western blotting, and enzyme-linked immunosorbent assay (ELISA) were performed to explore potential mechanisms. RESULTS: NCAPH was consistently upregulated in ESCC across multiple cohorts and was associated with unfavorable clinicopathological features and poorer survival. Functional assays demonstrated that NCAPH knockdown significantly inhibited ESCC cell proliferation, migration, invasion, and clonogenic growth. In vivo, NCAPH silencing suppressed xenograft tumor growth. Mechanistically, transcriptomic profiling and molecular validation indicated attenuation of Wnt/β-catenin signaling following NCAPH depletion, accompanied by reduced β-catenin and downstream targets. CONCLUSIONS: NCAPH promotes malignant progression of ESCC, at least in part through activation of the Wnt/β-catenin pathway, and may serve as a potential biomarker and therapeutic target.

Esophageal squamous cell carcinoma (ESCC)↗

FANCI promotes esophageal squamous cell carcinoma progression and cell cycle regulation and interacts with FANCD2.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) is an aggressive malignancy with poor clinical outcomes, and reliable molecular biomarkers and therapeutic targets remain limited. Fanconi anemia group I protein (FANCI) is a core component of the Fanconi anemia (FA) pathway, but its expression pattern, clinical significance, and functional role in ESCC have not been comprehensively defined. This study aimed to investigate FANCI expression and prognostic value in ESCC, assess its effects on malignant cellular phenotypes and tumor growth, and explore its potential mechanistic relationship with Fanconi anemia group D2 protein (FANCD2) and cell-cycle regulation. METHODS: Multi-cohort analyses were performed using The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets, together with ESCC single-cell RNA sequencing (RNA-seq) data. FANCI functions were assessed by bidirectional gain- and loss-of-function experiments in vitro (proliferation, colony formation, migration, invasion, apoptosis, and cell-cycle assays) and by xenograft models in vivo. Mechanistic studies included protein-protein interaction (PPI) analyses, co-immunoprecipitation (Co-IP), and immunofluorescence (IF) colocalization. RESULTS: FANCI was consistently upregulated in ESCC across bulk transcriptomic datasets and was further supported by quantitative polymerase chain reaction (qPCR), Western blotting, and immunohistochemistry (IHC). FANCI discriminated ESCC from normal tissues in TCGA-ESCC and was independently validated in GSE53624 [area under the curve (AUC) =0.940 and 0.975, respectively]. FANCI was associated with poorer overall survival (OS) and shorter disease-free interval (DFI), and these findings were validated in an independent GEO cohort. Functionally, FANCI promoted ESCC cell proliferation, migration, and invasion, while inhibiting apoptosis; FANCI knockdown suppressed tumor growth in vivo and induced G2/M cell-cycle arrest. Mechanistically, FANCI physically interacted with FANCD2, colocalized with FANCD2 in the nucleus, and was associated with altered FANCD2 protein abundance, consistent with cell-cycle and DNA repair-related programs. Single-cell analysis indicated that FANCI was enriched in epithelial cells and associated with higher activity of malignant functional programs. In TCGA-ESCC, FANCI-high tumors showed distinct mutation profiles, a trend toward increased tumor mutation burden (TMB), and altered immune-associated signatures. CONCLUSIONS: FANCI is upregulated in ESCC and is associated with diagnostic and prognostic value. It promotes malignant phenotypes and tumor growth, potentially through a FANCI-FANCD2-linked cell-cycle/DNA repair program, supporting FANCI as a candidate biomarker and therapeutic target in ESCC.

Esophageal squamous cell carcinoma (ESCC)↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-κB, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-κB, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Big data analytics for CLEC5A dynamics based on single cell genomics and proteomics reveal its diverse functions in human diseases.

BACKGROUND: CLEC5A (C-type lectin domain family 5 member A) is an innate immune receptor implicated in inflammatory signaling, contributing to hyperinflammatory responses in infections and sterile inflammation. However, CLEC5A dynamics in human diseases remain to be identified. Here, we systematically characterized CLEC5A dynamics in humans across cells, tissues, and disease states, and to explore the functional significance of CLEC5A in macrophage activation based on single-cell genomics. METHODS: With multi-omics (scRNA-seq, proteomics and big data analytics), we analyzed extensive human transcriptomic datasets (>42,000 samples) to profile CLEC5A expression by cell type, tissue, and disease. Single-nucleus RNA-seq (snRNA-seq) from pediatric congenital heart disease and a virtual CLEC5A gene knockout were also performed to characterize CLEC5A dynamics in humans. RESULTS: CLEC5A is highly enriched in innate immune cells, particularly in macrophages and neutrophils. Baseline CLEC5A in most tissues is low, but it is markedly upregulated in inflammatory and infectious diseases. CLEC5A expression has sex-specific differences in certain organs. Single-cell analysis showed that CLEC5A can be considered novel marker of proinflammatory macrophages with elevated cytokine production, antigen presentation, and impaired phagocytosis. Virtual CLEC5A knockout analysis identified coordinated perturbation of immune-regulatory pathways and overlapping genes linking CLEC5A to macrophage activation networks. CONCLUSION: CLEC5A is predominantly expressed in myeloid cells and acts as a key amplifier of inflammation in human diseases. Our findings highlight CLEC5A as a potential biomarker and therapeutic target in myeloid-driven hyperinflammatory conditions, warranting further experimental and translational validation.

Humans↗