PubMed HealthSearch

SEARCH · PubMed Health

Results for “Cross-modal alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

8 recordsLinked to original sources

A multi-modal survival prediction framework with group-based batch training and structural consistency alignment.

OBJECTIVE: Integrating whole-slide images (WSIs) with transcriptomic profiles is pivotal for enhancing cancer survival prediction. However, the intrinsic gigapixel resolution and variable sequence lengths of WSIs create a fundamental trade-off between training efficiency and the preservation of data heterogeneity in existing frameworks. Furthermore, substantial statistical and structural discrepancies between histological and genomic modalities often impede effective cross-modal alignment and fusion, thereby limiting prognostic accuracy. METHODS: We propose PRISM, an efficient multi-modal learning framework for integrating WSIs with transcriptomic profiles. To reconcile training efficiency with full data heterogeneity, PRISM first stochastically partitions variable-length WSI sequences into a main subset and a complementary residual subset, both of which are packed into fixed-length groups for batch training. The main subset is processed in the main branch, utilizing isolation masking to maintain intra-group sequence independence. Simultaneously, the residual subset is consolidated into "hyperslides" within a residual branch that leverages tailored supervision, effectively capturing inter-slide correlations. Furthermore, PRISM integrates an Informative Token Aggregation (ITA) module to reduce redundancy in WSIs and employs Cross-batch Structural Consistency Alignment (CBSCA) mechanism to enhance inter-modal structural connectivity. Finally, efficient cross-modal feature interaction is achieved through a Low-rank Bilinear Gated Fusion (LBGF) module. Code is available at https://github.com/Alisa2080/PRISM. RESULTS: Compared with existing methods, PRISM achieves the best overall C-index across five TCGA cohorts. On the larger TCGA-BRCA dataset, PRISM requires only 6 hours of training time, substantially reducing computational cost relative to strong multimodal baselines. Furthermore, comprehensive evaluations demonstrate that PRISM achieves the best overall IBS ranking and favorable time-dependent AUC performance at 1, 3, and 5 years, thereby delivering a more favorable trade-off between prognostic performance and computational efficiency. CONCLUSION: PRISM provides a favorable balance between predictive performance, calibration quality, and computational efficiency, highlighting its potential for practical deployment in multimodal survival modeling for computational pathology.

Humans

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans

SIVA: diagonal integration of spatial multi-omics data via spatially informed variational autoencoders and anchor guidance.

MOTIVATION: Understanding cellular states and regulatory programs requires integrative analysis of multiple omics layers. Although recent spatial sequencing technologies allow molecular profiling of cells within their tissue context, paired spatial multi-omics assays are still limited by technical complexity and cost. This creates a pressing need for diagonal integration methods that enable joint analysis of unpaired spatial omics datasets. RESULTS: We propose SIVA, a deep generative framework based on Spatially-Informed Variational Autoencoders with Anchor Guidance, for diagonal integration of spatial multi-modal data. SIVA employs modality-specific variational autoencoders (VAEs) with a hybrid latent embedding that integrates Gaussian process and standard Gaussian priors, enabling joint modeling of spatially structured variation and dominant underlying data distributions across modalities. To facilitate cross-modal alignment in the absence of one-to-one cell correspondence, SIVA adopts a dual integration strategy combining global distribution alignment via Maximum Mean Discrepancy and local correspondence guidance using mutual nearest neighbor anchors. Extensive experiments across multiple cross-slice integration scenarios demonstrate that SIVA achieves robust and accurate integration of unpaired spatial omics datasets, consistently outperforming existing methods. AVAILABILITY AND IMPLEMENTATION: The source codes are available at https://github.com/PelenJiang/SIVA.

Autoencoder

ARISE: RNA-anchored shared-edge topology and hierarchical fusion for spatial multi-omics integration.

MOTIVATION: Spatial multi-omics technologies jointly profile transcriptomes, proteins and chromatin accessibility in situ, enabling integrative analysis of tissue organization across molecular layers. However, most existing graph-based integration methods rely on independently constructed modality-specific k-nearest-neighbor graphs. When auxiliary modalities are sparse or noisy, these graphs can become topologically discordant, propagate spurious edges, weaken cross-modal alignment, and reduce spatial domain resolution. RESULTS: We present Anchored RNA for Integrated Spatial Embedding (ARISE), an RNA expression anchored framework for spatial multi-omics integration. ARISE defines a shared-edge topology by intersecting RNA feature-similarity and spatial-proximity graphs, encodes auxiliary modalities on this common scaffold, and integrates them through inside-out hierarchical fusion. We further show theoretically that graph intersection minimizes false-positive edges within a broad class of k-of-r graph fusion rules, providing a principled basis for topology anchoring. Across various spatial multi-omics benchmarks spanning simulated and real datasets in bi-modal and tri-modal settings, ARISE improves spatial domain identification, cross-modal consistency, and preservation of tissue structure relative to existing methods. Furthermore, the learned representation supports biologically meaningful downstream analyses, including marker-based domain annotation, pathway enrichment, and cis-regulatory inference, indicating that ARISE yields a robust and interpretable framework for spatial multi-omics integration. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XiangxiangWang-code/ARISE. The archived version used in this study is available at https://doi.org/10.6084/m9.figshare.32686137.v2.

Multiomics

ARCADIA reveals spatially dependent transcriptional programs through integration of scRNA-seq and spatial proteomics.

MOTIVATION: Cellular states are strongly influenced by spatial context, but single-cell RNA sequencing (scRNA-seq) loses information about local tissue organization, while spatial proteomic assays capture limited marker panels that constrain transcriptomic inference. Integrating these modalities can elucidate how spatial niches shape transcriptional programs, yet existing approaches depend on either feature-level correspondence such as gene-protein linkage or cell-level barcode pairing, which is often unavailable. RESULTS: We present ARCADIA (ARchetype-based Clustering and Alignment with Dual Integrative Autoencoders), a generative framework for cross-modal integration that operates without cell barcode pairing and does not assume direct feature-to-feature correspondence. ARCADIA identifies modality-specific archetypes, that is, convex combinations of cells representing extreme phenotypic states, and aligns these anchors across modalities by minimizing the discrepancy between their cell-type composition profiles. The aligned archetypes define a shared coordinate system that anchors dual variational autoencoders (VAEs) trained with cross-modal geometric regularization, preserving archetype structure and spatial neighborhood information while enabling bidirectional translation between modalities. On semi-synthetic CITE-seq data, ARCADIA outperforms existing weak-linkage methods. Applied to independent human tonsil scRNA-seq and CODEX data, ARCADIA reconstructs known tissue architecture and reveals spatially dependent transcriptional programs linking B-cell maturation and T-cell activation or exhaustion to microenvironmental niches. AVAILABILITY AND IMPLEMENTATION: Source code is accessible at https://github.com/azizilab/ARCADIA_public. Reproducibility scripts and data are available at https://github.com/azizilab/arcadia_reproducibility.

Proteomics

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans

scMultiNODE: Integrative and Scalable Framework for Multi-Modal Temporal Single-Cell Data.

Measuring single-cell genomic profiles at different timepoints enables our understanding of cell development. This understanding is more comprehensive when we perform an integrative analysis of multiple measurements (or modalities) across various developmental stages. However, obtaining such measurements from the same set of single cells is resource-intensive, restricting our ability to study them jointly. We introduce scMultiNODE, an unsupervised integration model that combines gene expression and chromatin accessibility measurements in developing single cells, while preserving cell type variations and cellular dynamics. First, scMultiNODE uses a scalable, Quantized Gromov-Wasserstein optimal transport to align a large number of cells across different measurements. Next, it utilizes neural ordinary differential equations to explicitly model cell development with a regularization term to learn a dynamic latent space. Experiments on six real-world developmental single-cell datasets demonstrate that scMultiNODE can integrate temporally profiled multi-modal single-cell measurements more effectively than existing methods that focus on cell type variations and often overlook cellular dynamics. We also demonstrate that scMultiNODE's joint latent space facilitates several insightful downstream analyses of single-cell development, including the investigation of complex cell trajectories and the enabling of cross-modal label transfer. The data and code are publicly available at https://github.com/rsinghlab/scMultiNODE.

autoencoders