PubMed HealthSearch

PubMed · 42635210

ARCADIA reveals spatially dependent transcriptional programs through integration of scRNA-seq and spatial proteomics.

Abstract

MOTIVATION: Cellular states are strongly influenced by spatial context, but single-cell RNA sequencing (scRNA-seq) loses information about local tissue organization, while spatial proteomic assays capture limited marker panels that constrain transcriptomic inference. Integrating these modalities can elucidate how spatial niches shape transcriptional programs, yet existing approaches depend on either feature-level correspondence such as gene-protein linkage or cell-level barcode pairing, which is often unavailable. RESULTS: We present ARCADIA (ARchetype-based Clustering and Alignment with Dual Integrative Autoencoders), a generative framework for cross-modal integration that operates without cell barcode pairing and does not assume direct feature-to-feature correspondence. ARCADIA identifies modality-specific archetypes, that is, convex combinations of cells representing extreme phenotypic states, and aligns these anchors across modalities by minimizing the discrepancy between their cell-type composition profiles. The aligned archetypes define a shared coordinate system that anchors dual variational autoencoders (VAEs) trained with cross-modal geometric regularization, preserving archetype structure and spatial neighborhood information while enabling bidirectional translation between modalities. On semi-synthetic CITE-seq data, ARCADIA outperforms existing weak-linkage methods. Applied to independent human tonsil scRNA-seq and CODEX data, ARCADIA reconstructs known tissue architecture and reveals spatially dependent transcriptional programs linking B-cell maturation and T-cell activation or exhaustion to microenvironmental niches. AVAILABILITY AND IMPLEMENTATION: Source code is accessible at https://github.com/azizilab/ARCADIA_public. Reproducibility scripts and data are available at https://github.com/azizilab/arcadia_reproducibility.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bar Rozenman, Kevin Hoffer-Hawlik, Nicholas Djedjos, Elham Azizi. 2026-08-01. ARCADIA reveals spatially dependent transcriptional programs through integration of scRNA-seq and spatial proteomics.. https://doi.org/10.1093/bioinformatics%2Fbtag454

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Multi-omic analyses of the same sample using metabolomics, lipidomics, proteomics, phosphoproteomics, and glycoproteomics.

Mass spectrometry (MS)-based multi-omics offers powerful tools to comprehensively characterize proteins, post-translational modifications, metabolites, and lipids. However, these measurements are typically performed using separate sample preparation workflows and modality-specific liquid chromatography mass spectrometry (LC-MS) platforms, limiting integration and constraining applications to small amounts of sample materials, especially scarce clinical specimens. Here, we describe a unified nano-LC-MS framework that enables metabolomic, lipidomic, proteomic, phosphoproteomic, and glycoproteomic analyses from the same starting material using a single nano-LC-MS platform, with only the chromatographic conditions, acquisition methods, and enrichment procedures tailored to each omics. This integrated strategy reduces workflow complexity and sample consumption while improves analytical continuity across molecular layers. By enabling deep multi-omics characterization from the same sample, this platform provides a practical foundation for comprehensive analysis of precious clinical samples.

Proteomics

4D-DIA proteomics reveals distinct proteolytic landscapes induced by mechanical stress, Agrobacterium, and a viral capsid precursor.

Nicotiana benthamiana is a widely used platform for plant molecular farming, yet recombinant protein yields are frequently compromised by the host's innate defense mechanisms, particularly proteolytic degradation. While the general effects of Agroinfiltration are known, the distinct contributions of mechanical injury, bacterial perception, and product-specific stress remain poorly resolved. Here we utilized high-depth 4D-DIA proteomics to dissect the host response across three dimensions: physical stress (buffer infiltration), pathogen-associated stress (Agrobacterium), and product-associated stress (GFP vs. the FMDV capsid precursor P1_2A). We demonstrate that buffer infiltration is not a neutral event but an independent inducer of cell wall remodeling and oxidative stress. By filtering out these background effects, we defined a core Agrobacterium-responsive proteome characterized by a growth-defense trade-off. We also expanded the known protease repertoire of N. benthamiana to 1,505 enzymes through improved genomic annotation. We found that the expression of the FMDV capsid precursor P1_2A was associated with a distinct and more pronounced protease profile compared to soluble GFP, characterized by the upregulation of subtilases and cysteine proteases. These findings suggest that host proteolytic responses vary with the recombinant cargo, a factor worth considering when designing engineering strategies for the production of complex biopharmaceuticals in plants.

Proteomics

An Instrumental Optimization of a Label-Free Proteomic Method for Trace Protein Input.

Liquid chromatography-mass spectrometry (LC-MS)-based proteomics of trace-level samples, such as tens of cells or spatially resolved tissue regions, offers unique biological insights but is often constrained by the requirement for specialized, costly instrumentation. In this study, we developed a scalable workflow for the deep proteomic analysis of low- to ultralow-input samples by systematically optimizing a widely adopted Orbitrap and UHPLC platform to maximize sensitivity, precision, and throughput. This optimized workflow identified over 5600 proteins from 5 ng of peptides and 3400 proteins from 20 sorted cells, achieving a throughput of 30 analyses per day while maintaining deep proteome coverage and high quantitative reproducibility. Furthermore, by applying this method to spatially resolved proteomics, we identified over 6100 proteins from microscale regions of interest (ROIs) within a formalin-fixed, paraffin-embedded (FFPE) tissue. A data-driven normalization strategy was employed to correct for variable cellularity across tissue regions, effectively revealing intratumor heterogeneity and distinct molecular and functional signatures, including pathway activations not apparent in parallel spatial transcriptomic analysis. Ultimately, this accessible, high-performance method substantially lowers the instrumentation barrier for the deep proteomic profiling of trace-level biological samples.

Proteomics