PubMed HealthSearch

SEARCH · PubMed Health

Results for “barcode”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Plant species identification by genome skimming across the vascular plant tree of life.

Accurate species identification is essential for biodiversity conservation and sustainable use, yet standard plant DNA barcoding often fails to achieve species-level resolution. We present a large-scale empirical evaluation of genome skimming as a tool to improve plant species discrimination. Using standardised data from 1969 individuals representing 475 species from 32 genera across major lineages of the vascular plant tree of life, we compare conventional plastid + internal transcribed spacer (ITS) barcodes with genome skimming approaches. Standard barcoding using rbcL, matK, trnH-psbA and ITS resolved about half of species (49.3%), with six genera showing <&#x2009;25% species discrimination. By contrast, genome skimming enabled the recovery of complete plastid genomes, yielding 57.6% species discrimination. It also generated sufficient nuclear genomic data for additional resolution from k-mer analysis, achieving 66.8% species discrimination - an average gain of 17.5% over standard barcodes - while eliminating cases of extreme failure (<&#x2009;25% resolution). The recovery of complete plastomes and ribosomal DNAs from genome skims also ensures backward compatibility with existing barcode datasets. Our results demonstrate that genome skimming provides data that substantially improves species-level resolution across diverse plant lineages and offers a scalable, high-throughput approach for building comprehensive reference resources to support global biodiversity initiatives.

DNA Barcoding, Taxonomic

Homologous chromosome recognition via nonspecific interactions.

In many organisms, most notably Drosophila, homologous chromosomes in somatic cells associate with each other, a phenomenon known as somatic homolog pairing. Unlike in meiosis, where homology is read out at the level of DNA sequence complementarity, somatic homolog pairing takes place without double strand breaks or strand invasion, thus requiring some other mechanism for homologs to recognize each other. Several studies have suggested a "specific button" model, in which a series of distinct regions in the genome, known as buttons, can associate with each other, presumably mediated by different proteins that bind to these different regions. Here we consider an alternative model, which we term the "button barcode" model, in which there is only one type of recognition site or adhesion button, present in many copies in the genome, each of which can associate with any of the others with equal affinity. An important component of this model is that the buttons are non-uniformly distributed, such that alignment of a chromosome with its correct homolog, compared with a non-homolog, is energetically favored; since to achieve nonhomologous alignment, chromosomes would be required to mechanically deform in order to bring their buttons into mutual register. We investigated several types of barcodes and examined their effect on pairing fidelity. We found that high fidelity homolog recognition can be achieved by arranging chromosome pairing buttons according to an actual industrial barcode used for warehouse sorting. By simulating randomly generated non-uniform button distributions, many highly effective button barcodes can be easily found, some of which achieve virtually perfect pairing fidelity. This model is consistent with existing literature on the effect of translocations of different sizes on homolog pairing. We conclude that a button barcode model can attain highly specific homolog recognition, comparable to that seen in actual cells undergoing somatic homolog pairing, without the need for specific interactions. This model may have implications for how meiotic pairing is achieved.

Preprint

Molecular identification of Hymenopteran insects collected by using Malaise traps from Hazarganji Chiltan National Park Quetta, Pakistan.

The order Hymenoptera holds great significance for humans, particularly in tropical and subtropical regions, due to its role as a pollinator of wild and cultivated flowering plants, parasites of destructive insects and honey producers. Despite this importance, limited attention has been given to the genetic diversity and molecular identification of Hymenopteran insects in most protected areas. This study provides insights into the first DNA barcode of Hymenopteran insects collected from Hazarganji Chiltan National Park (HCNP) and contributes to the global reference library of DNA barcodes. A total of 784 insect specimens were collected using Malaise traps, out of which 538 (68.62%) specimens were morphologically identified as Hymenopteran insects. The highest abundance of species of Hymenoptera (133/538, 24.72%) was observed during August and least in November (16/538, 2.97%). Genomic DNA extraction was performed individually from 90/538 (16.73%) morphologically identified specimens using the standard phenol-chloroform method, which were subjected separately to the PCR for their molecular confirmation via the amplification of cytochrome c oxidase subunit 1 (cox1) gene. The BLAST analyses of obtained sequences showed 91.64% to 100% identities with related sequences and clustered phylogenetically with their corresponding sequences that were reported from Australia, Bulgaria, Canada, Finland, Germany, India, Israel, and Pakistan. Additionally, total of 13 barcode index numbers (BINs) were assigned by Barcode of Life Data Systems (BOLD), out of which 12 were un-unique and one was unique (BOLD: AEU1239) which was assigned for Anthidium punctatum. This indicates the potential geographical variation of Hymenopteran population in HCNP. Further comprehensive studies are needed to molecularly confirm the existing insect species in HCNP and evaluate their impacts on the environment, both as beneficial (for example, pollination, honey producers and natural enemies) and detrimental (for example, venomous stings, crop damage, and pathogens transmission).

Humans

SpaceBar enables clone tracing in spatial transcriptomic data.

We report a cellular barcoding strategy, SpaceBar, that enables simultaneous clone tracing and spatial transcriptomics profiling. Our approach uses a library of 96 synthetic barcode sequences that can be robustly detected by imaging based spatial transcriptomics (seqFISH), delivered such that each cell is labeled with a combination of barcodes. We used these barcodes to label melanoma cells in a tumor xenograft model and profiled both clone identity and spatial gene expression in situ. We developed a gene scoring metric that quantifies how strongly gene expression is driven by intrinsic cellular cues or extrinsic environmental signals. Our framework distinguishes between clonal dynamics and environmentally-driven transcriptional regulation in complex tissue contexts.

Journal Article

Deciphering Cell Fate and Clonal Dynamics via Integrative Single-Cell Lineage Modeling.

Through natural or synthetic lineage barcodes, single-cell technologies now enable the joint measurement of molecular states and clonal identities, providing an unprecedented opportunity to study cell fate and dynamics. Yet, most computational methods for inferring cell development and differentiation rely exclusively on transcriptional similarity, overlooking the lineage information encoded by lineage barcodes. This limitation is exemplified by T cells, where subtle transcriptional differences mark divergent fates with distinct biological activity. Single-cell RNA and matched TCR sequencing is now ubiquitous in the analysis of clinical samples, where the TCR sequence provides an endogenous clonal barcode and could reveal clonal T cell responses. We present Clonotrace, a computational framework that jointly models gene expression and clonotype information to infer cell state transitions and fate biases with higher fidelity. While motivated by challenges in analyzing T cell populations, especially in the tumor microenvironment and immunotherapy settings, Clonotrace is broadly applicable to any lineage-barcoded single-cell dataset. Across diverse systems including T cells, hematopoietic differentiation, and cancer therapy resistance models, Clonotrace reveals differentiation hierarchies, distinguishes unipotent from multipotent states, and identifies candidate fate-determining genes driving lineage commitment.

Journal Article

Pooled PPIseq: Screening the SARS-CoV-2 and human interface with a scalable multiplexed protein-protein interaction assay platform.

Protein-Protein Interactions (PPIs) are a key interface between virus and host, and these interactions are important to both viral reprogramming of the host and to host restriction of viral infection. In particular, viral-host PPI networks can be used to further our understanding of the molecular mechanisms of tissue specificity, host range, and virulence. At higher scales, viral-host PPI screening could also be used to screen for small-molecule antivirals that interfere with essential viral-host interactions, or to explore how the PPI networks between interacting viral and host genomes co-evolve. Current high-throughput PPI assays have screened entire viral-host PPI networks. However, these studies are time consuming, often require specialized equipment, and are difficult to further scale. Here, we develop methods that make larger-scale viral-host PPI screening more accessible. This approach combines the mDHFR split-tag reporter with the iSeq2 interaction-barcoding system to permit massively-multiplexed PPI quantification by simple pooled engineering of barcoded constructs, integration of these constructs into budding yeast, and fitness measurements by pooled cell competitions and barcode-sequencing. We applied this method to screen for PPIs between SARS-CoV-2 proteins and human proteins, screening in triplicate >180,000 ORF-ORF combinations represented by >1,000,000 barcoded lineages. Our results complement previous screens by identifying 74 putative PPIs, including interactions between ORF7A with the taste receptors TAS2R41 and TAS2R7, and between NSP4 with the transmembrane KDELR2 and KDELR3. We show that this PPI screening method is highly scalable, enabling larger studies aimed at generating a broad understanding of how viral effector proteins converge on cellular targets to effect replication.

Humans

Deciphering Cell Fate and Clonal Dynamics via Integrative Single-Cell Lineage Modeling.

Through natural or synthetic lineage barcodes, single-cell technologies now enable the joint measurement of molecular states and clonal identities, providing an unprecedented opportunity to study cell fate and dynamics. Yet, most computational methods for inferring cell development and differentiation rely exclusively on transcriptional similarity, overlooking the lineage information encoded by lineage barcodes. This limitation is exemplified by T cells, where subtle transcriptional differences mark divergent fates with distinct biological activity. Single-cell RNA and matched TCR sequencing is now ubiquitous in the analysis of clinical samples, where the TCR sequence provides an endogenous clonal barcode and could reveal clonal T cell responses. We present Clonotrace, a computational framework that jointly models gene expression and clonotype information to infer cell state transitions and fate biases with higher fidelity. While motivated by challenges in analyzing T cell populations, especially in the tumor microenvironment and immunotherapy settings, Clonotrace is broadly applicable to any lineage-barcoded single-cell dataset. Across diverse systems including T cells, hematopoietic differentiation, and cancer therapy resistance models, Clonotrace reveals differentiation hierarchies, distinguishes unipotent from multipotent states, and identifies candidate fate-determining genes driving lineage commitment.

Journal Article

A systematic capsid evolution approach performed in vivo for the design of AAV vectors with tailored properties and tropism.

Adeno-associated virus (AAV) capsid modification enables the generation of recombinant vectors with tailored properties and tropism. Most approaches to date depend on random screening, enrichment, and serendipity. The approach explored here, called BRAVE (barcoded rational AAV vector evolution), enables efficient selection of engineered capsid structures on a large scale using only a single screening round in vivo. The approach stands in contrast to previous methods that require multiple generations of enrichment. With the BRAVE approach, each virus particle displays a peptide, derived from a protein, of known function on the AAV capsid surface, and a unique molecular barcode in the packaged genome. The sequencing of RNA-expressed barcodes from a single-generation in vivo screen allows the mapping of putative binding sequences from hundreds of proteins simultaneously. Using the BRAVE approach and hidden Markov model-based clustering, we present 25 synthetic capsid variants with refined properties, such as retrograde axonal transport in specific subtypes of neurons, as shown for both rodent and human dopaminergic neurons.

barcoding

A genetic atlas for the butterflies of continental Canada and United States.

Multi-locus genetic data for phylogeographic studies is generally limited in geographic and taxonomic scope as most studies only examine a few related species. The strong adoption of DNA barcoding has generated large datasets of mtDNA COI sequences. This work examines the butterfly fauna of Canada and United States based on 13,236 COI barcode records derived from 619 species. It compiles i) geographic maps depicting the spatial distribution of haplotypes, ii) haplotype networks (minimum spanning trees), and iii) standard indices of genetic diversity such as nucleotide diversity (&#x3c0;), haplotype richness (H), and a measure of spatial genetic structure (GST). High intraspecific genetic diversity and marked spatial structure were observed in the northwestern and southern North America, as well as in proximity to mountain chains. While species generally displayed concordance between genetic diversity and spatial structure, some revealed incongruence between these two metrics. Interestingly, most species falling in this category shared their barcode sequences with one at least other species. Aside from revealing large-scale phylogeographic patterns and shedding light on the processes underlying these patterns, this work also exposed cases of potential synonymy and hybridization.

Animals

Modeling homologous chromosome recognition via nonspecific interactions.

In many organisms, most notably Drosophila, homologous chromosomes associate in somatic cells, a phenomenon known as somatic pairing, which takes place without double strand breaks or strand invasion, thus requiring some other mechanism for homologs to recognize each other. Several studies have suggested a "specific button" model, in which a series of distinct regions in the genome, known as buttons, can associate with each other, mediated by different proteins that bind to these different regions. Here, we use computational modeling to evaluate an alternative "button barcode" model, in which there is only one type of recognition site or adhesion button, present in many copies in the genome, each of which can associate with any of the others with equal affinity. In this model, buttons are nonuniformly distributed, such that alignment of a chromosome with its correct homolog, compared with a nonhomolog, is energetically favored; since to achieve nonhomologous alignment, chromosomes would be required to mechanically deform in order to bring their buttons into mutual register. By simulating randomly generated nonuniform button distributions, many highly effective button barcodes can be easily found, some of which achieve virtually perfect pairing fidelity. This model is consistent with existing literature on the effect of translocations of different sizes on homolog pairing. We conclude that a button barcode model can attain highly specific homolog recognition, comparable to that seen in actual cells undergoing somatic homolog pairing, without the need for specific interactions. This model may have implications for how meiotic pairing is achieved.

Animals

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100&#x2009;000 and 1&#x2009;000&#x2009;000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

DNA Extraction Optimisation for Minute Land Snails of Vertigo M&#xfc;ller, 1773 (Gastropoda: Vertiginidae): A Comparative Evaluation of Six Methods, Including a Non-Destructive Shell-Preserving Protocol.

No systematic comparison of DNA extraction strategies exists for minute Vertiginidae (shell height <&#x2009;3&#x2009;mm), a group posing a dual analytical challenge: extremely low tissue input and co-purified PCR-inhibitory mucus. For legally protected species, an additional requirement to preserve the shell voucher further constrains available protocols. Using Vertigo antivertigo as the model species, we compared six approaches applied to specimens preserved in 96% ethanol (n&#x2009;=&#x2009;10 per method): two HotSHOT alkaline-lysis protocols (destructive and non-destructive shell-preserving variants), a modified CTAB protocol supplemented with PVP-40 and DTT, and three commercial silica-column kits (GeneJET Genomic, DNeasy Blood & Tissue, QIAamp DNA Micro). DNA yields were quantified by QuantiFluor fluorometry, and PCR performance was subsequently assessed across four loci (COI barcode, COI mini-barcode, ITS1, ITS2). DNeasy Blood & Tissue produced the highest fluorometric concentrations; QIAamp DNA Micro and CTAB&#x2009;+&#x2009;PVP-40 gave intermediate values. The shell-preserving HotSHOT variant yielded lower concentrations but improved A260/230 ratios. BSA and trehalose supplementation increased PCR success in inhibition-prone HotSHOT extracts from 70% to 100%. ITS1 Sanger sequencing of three Vertigo species listed in Annex II of the EU Habitats Directive, all extracted with the shell-preserving protocol, confirmed species-level identification (99.8%-100% BLASTn identity; mean Phred Q&#x2009;>&#x2009;51). The shell-preserving non-destructive HotSHOT protocol yields sequenceable DNA from protected Vertiginidae while retaining the morphological voucher, making it the preferred option for conservation-genetic monitoring. The practical decision framework documented here-integrating voucher preservation, amplification robustness and per-sample cost-has broad applicability to other minute terrestrial gastropods processed in large-scale biodiversity surveys.

Habitats Directive

In vivo genome editing of central nervous system SIV reservoirs in ART-suppressed rhesus macaques.

Latent human immunodeficiency virus type 1 (HIV-1) reservoirs in the central nervous system (CNS) may sustain viral persistence and neuroinflammation contributing to HIV-associated neurocognitive disorders (HAND) despite suppressive ART. AAV9-delivered CRISPR has successfully edited SIV proviral DNA in peripheral tissues with acceptable safety profiles, but the extent of in vivo genome editing in the brain remains unclear. Using SIV-infected rhesus macaques, we mapped intact proviral DNA across CNS regions and tested systemic AAV9-CRISPR-Cas9 targeting conserved sites within &#x3a8; packaging signal and Gag region. Ten adult rhesus macaques were infected with genetically barcoded SIVmac239, suppressed with ART, then randomized to receive intravenous AAV9-SaCas9 with dual gRNAs (&#x3a8; + Gag) or a Cas9-only control. At necropsy after viral rebound, SIV genomes were detected in multiple brain regions as well as lymphoid tissues, confirming the CNS as a persistent reservoir during ART. Barcode analysis revealed region-specific patterns consistent with compartmentalized CNS persistence. In CRISPR-treated animals, proviral editing was measurable across anatomically distinct CNS sites. These findings demonstrate that intact and potentially replication-competent virus persists in the primate brain under ART and that systemic AAV9-CRISPR can reach and edit proviral DNA in this sanctuary, supporting genome editing as a strategy toward durable remission of CNS reservoirs.

ART

Directed evolution of engineered virus-like particles with improved production and transduction efficiencies.

Engineered virus-like particles (eVLPs) are promising vehicles for transient delivery of proteins and RNAs, including gene editing agents. We report a system for the laboratory evolution of eVLPs that enables the discovery of eVLP variants with improved properties. The system uses barcoded guide RNAs loaded within DNA-free eVLP-packaged cargos to uniquely label each eVLP variant in a library, enabling the identification of desired variants following selections for desired properties. We applied this system to mutate and select eVLP capsids with improved eVLP production properties or transduction efficiencies in human cells. By combining beneficial capsid mutations, we developed fifth-generation (v5) eVLPs, which exhibit a 2-4-fold increase in cultured mammalian cell delivery potency compared to previous-best v4 eVLPs. Analyses of v5 eVLPs suggest that these capsid mutations optimize packaging and delivery of desired ribonucleoprotein cargos rather than native viral genomes and substantially alter eVLP capsid structure. These findings suggest the potential of barcoded eVLP evolution to support the development of improved eVLPs.

Humans

scnanoseq: an nf-core pipeline for Oxford Nanopore single-cell RNA-sequencing.

MOTIVATION: Recent advancements in long-read single-cell RNA sequencing (scRNA-seq) have facilitated the quantification of full-length transcripts and isoforms at the single-cell level. Historically, long-read data would need to be complemented with short-read single-cell data in order to overcome the higher sequencing errors to correctly identify cellular barcodes and unique molecular identifiers. Improvements in Oxford Nanopore sequencing, and development of novel computational methods have removed this requirement. Though these methods now exist, the limited availability of modular and portable workflows remains a challenge. RESULTS: Here, we present, nf-core/scnanoseq, a secondary analysis pipeline for long-read single-cell and single-nuclei RNA that delivers gene and transcript-level quantification. The scnanoseq pipeline is implemented using Nextflow and is built upon the nf-core framework, enabling portability across computational environments, scalability and reproducibility of results across pipeline runs. The nf-core/scnanoseq workflow follows best practices for analyzing single-cell and single-nuclei data, performing barcode detection and correction, genome and transcriptome read alignment, unique molecular identifier deduplication, gene and transcript quantification, and extensive quality control reporting. AVAILABILITY AND IMPLEMENTATION: The source code, and detailed documentation are freely available at https://github.com/nf-core/scnanoseq and https://nf-co.re/scnanoseq under the MIT License. Documentation for the version of nf-core/scnanoseq used for this paper, including default parameters and descriptions of output files are available at https://nf-co.re/scnanoseq/1.1.0.

Single-Cell Analysis

ARCADIA reveals spatially dependent transcriptional programs through integration of scRNA-seq and spatial proteomics.

MOTIVATION: Cellular states are strongly influenced by spatial context, but single-cell RNA sequencing (scRNA-seq) loses information about local tissue organization, while spatial proteomic assays capture limited marker panels that constrain transcriptomic inference. Integrating these modalities can elucidate how spatial niches shape transcriptional programs, yet existing approaches depend on either feature-level correspondence such as gene-protein linkage or cell-level barcode pairing, which is often unavailable. RESULTS: We present ARCADIA (ARchetype-based Clustering and Alignment with Dual Integrative Autoencoders), a generative framework for cross-modal integration that operates without cell barcode pairing and does not assume direct feature-to-feature correspondence. ARCADIA identifies modality-specific archetypes, that is, convex combinations of cells representing extreme phenotypic states, and aligns these anchors across modalities by minimizing the discrepancy between their cell-type composition profiles. The aligned archetypes define a shared coordinate system that anchors dual variational autoencoders (VAEs) trained with cross-modal geometric regularization, preserving archetype structure and spatial neighborhood information while enabling bidirectional translation between modalities. On semi-synthetic CITE-seq data, ARCADIA outperforms existing weak-linkage methods. Applied to independent human tonsil scRNA-seq and CODEX data, ARCADIA reconstructs known tissue architecture and reveals spatially dependent transcriptional programs linking B-cell maturation and T-cell activation or exhaustion to microenvironmental niches. AVAILABILITY AND IMPLEMENTATION: Source code is accessible at https://github.com/azizilab/ARCADIA_public. Reproducibility scripts and data are available at https://github.com/azizilab/arcadia_reproducibility.

Proteomics