PubMed HealthSearch

SEARCH · PubMed Health

Results for “multi-modal data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

11 recordsLinked to original sources

Social disconnection integrates genetic and proteomic risks in suicidal ideation and depression.

Suicidal ideation (SI) and major depressive disorder (MDD) are complex psychiatric conditions arising from the interplay of genetic liability, molecular processes, and psychosocial factors. While these dimensions have been extensively studied in isolation, their joint contribution to SI and MDD remains unclear. This study integrates multi-modal data to elucidate these synergistic effects and develop robust models for individual-level risk stratification. Leveraging longitudinal multi-modal data from 13,085 UK Biobank participants, we integrated genomic, proteomic, and social connection profiles. We developed interpretable risk scores using a rigorous supervised machine learning framework encompassing diverse linear and ensemble classifiers. Permutation importance was employed to quantify feature contributions and derive transparent, weighted risk metrics across diverse classifiers. These scores were validated through association, interaction, and mediation analyses. Social connection-based risk scores significantly differentiated cases and controls across the two suicidal ideation phenotypes at 2017 and 2023 with cross-sectional analyses (AUCs: 0.70 - 0.73), outperforming proteomic-only models. Functional dimensions of social connection emerged as the most informative predictors. Longitudinal analyses revealed that social risk scores at baseline predicted suicidal ideation onset six years later, independent of demographic covariates. Interaction analyses demonstrated that polygenic risk for suicide attempt significantly interacted with both social and proteomic risk features in relation to depression. Structural equation models further confirmed that social disconnection acts as a key mediator linking genetic predisposition to MDD and SI. Social disconnection is a critical risk factor mediating the impact of genetic vulnerability on psychiatric outcomes. Integrating social, genetic, and molecular data supports a multilevel framework for risk stratification and highlights the potential of socially oriented interventions to mitigate biological risk.

Humans

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

scMultiNODE: Integrative and Scalable Framework for Multi-Modal Temporal Single-Cell Data.

Measuring single-cell genomic profiles at different timepoints enables our understanding of cell development. This understanding is more comprehensive when we perform an integrative analysis of multiple measurements (or modalities) across various developmental stages. However, obtaining such measurements from the same set of single cells is resource-intensive, restricting our ability to study them jointly. We introduce scMultiNODE, an unsupervised integration model that combines gene expression and chromatin accessibility measurements in developing single cells, while preserving cell type variations and cellular dynamics. First, scMultiNODE uses a scalable, Quantized Gromov-Wasserstein optimal transport to align a large number of cells across different measurements. Next, it utilizes neural ordinary differential equations to explicitly model cell development with a regularization term to learn a dynamic latent space. Experiments on six real-world developmental single-cell datasets demonstrate that scMultiNODE can integrate temporally profiled multi-modal single-cell measurements more effectively than existing methods that focus on cell type variations and often overlook cellular dynamics. We also demonstrate that scMultiNODE's joint latent space facilitates several insightful downstream analyses of single-cell development, including the investigation of complex cell trajectories and the enabling of cross-modal label transfer. The data and code are publicly available at https://github.com/rsinghlab/scMultiNODE.

autoencoders

Representation learning for multi-modal spatially resolved transcriptomics data.

MOTIVATION: Spatial transcriptomics enables in-depth molecular characterization of samples on a morphology and RNA level while preserving spatial location. Integrating the resulting multi-modal data is an unsolved problem, and developing new solutions in precision medicine depends on improved methodologies. RESULTS: We introduce AESTETIK, a convolutional deep learning model that jointly integrates spatial, transcriptomics, and morphology information to learn accurate spot representations. AESTETIK yielded substantially improved cluster assignments on widely adopted technology platforms (e.g. 10x Genomics™, NanoString™) across multiple datasets. We achieved performance enhancement on structured tissues (e.g. brain) with a 21% increase in median ARI over previous state-of-the-art methods. Notably, AESTETIK also demonstrated superior performance on cancer tissues with heterogeneous cell populations, showing a 2-fold increase in breast cancer, 79% in melanoma, and 21% in liver cancer. We expect that these advances will enable a multi-modal understanding of key biological processes. AVAILABILITY AND IMPLEMENTATION: AESTETIK is implemented in Python 3 and is available as open source software at http://www.github.com/ratschlab/aestetik. The Snakemake pipeline for reproducing the results is available at http://www.github.com/ratschlab/st-rep.

Spatial Transcriptomics

SIVA: diagonal integration of spatial multi-omics data via spatially informed variational autoencoders and anchor guidance.

MOTIVATION: Understanding cellular states and regulatory programs requires integrative analysis of multiple omics layers. Although recent spatial sequencing technologies allow molecular profiling of cells within their tissue context, paired spatial multi-omics assays are still limited by technical complexity and cost. This creates a pressing need for diagonal integration methods that enable joint analysis of unpaired spatial omics datasets. RESULTS: We propose SIVA, a deep generative framework based on Spatially-Informed Variational Autoencoders with Anchor Guidance, for diagonal integration of spatial multi-modal data. SIVA employs modality-specific variational autoencoders (VAEs) with a hybrid latent embedding that integrates Gaussian process and standard Gaussian priors, enabling joint modeling of spatially structured variation and dominant underlying data distributions across modalities. To facilitate cross-modal alignment in the absence of one-to-one cell correspondence, SIVA adopts a dual integration strategy combining global distribution alignment via Maximum Mean Discrepancy and local correspondence guidance using mutual nearest neighbor anchors. Extensive experiments across multiple cross-slice integration scenarios demonstrate that SIVA achieves robust and accurate integration of unpaired spatial omics datasets, consistently outperforming existing methods. AVAILABILITY AND IMPLEMENTATION: The source codes are available at https://github.com/PelenJiang/SIVA.

Autoencoder

A multi-modal survival prediction framework with group-based batch training and structural consistency alignment.

OBJECTIVE: Integrating whole-slide images (WSIs) with transcriptomic profiles is pivotal for enhancing cancer survival prediction. However, the intrinsic gigapixel resolution and variable sequence lengths of WSIs create a fundamental trade-off between training efficiency and the preservation of data heterogeneity in existing frameworks. Furthermore, substantial statistical and structural discrepancies between histological and genomic modalities often impede effective cross-modal alignment and fusion, thereby limiting prognostic accuracy. METHODS: We propose PRISM, an efficient multi-modal learning framework for integrating WSIs with transcriptomic profiles. To reconcile training efficiency with full data heterogeneity, PRISM first stochastically partitions variable-length WSI sequences into a main subset and a complementary residual subset, both of which are packed into fixed-length groups for batch training. The main subset is processed in the main branch, utilizing isolation masking to maintain intra-group sequence independence. Simultaneously, the residual subset is consolidated into "hyperslides" within a residual branch that leverages tailored supervision, effectively capturing inter-slide correlations. Furthermore, PRISM integrates an Informative Token Aggregation (ITA) module to reduce redundancy in WSIs and employs Cross-batch Structural Consistency Alignment (CBSCA) mechanism to enhance inter-modal structural connectivity. Finally, efficient cross-modal feature interaction is achieved through a Low-rank Bilinear Gated Fusion (LBGF) module. Code is available at https://github.com/Alisa2080/PRISM. RESULTS: Compared with existing methods, PRISM achieves the best overall C-index across five TCGA cohorts. On the larger TCGA-BRCA dataset, PRISM requires only 6 hours of training time, substantially reducing computational cost relative to strong multimodal baselines. Furthermore, comprehensive evaluations demonstrate that PRISM achieves the best overall IBS ranking and favorable time-dependent AUC performance at 1, 3, and 5 years, thereby delivering a more favorable trade-off between prognostic performance and computational efficiency. CONCLUSION: PRISM provides a favorable balance between predictive performance, calibration quality, and computational efficiency, highlighting its potential for practical deployment in multimodal survival modeling for computational pathology.

Humans

AI-genomics synergy for drug repurposing in breast cancer: an interpretability-driven framework.

Breast cancer's genomic heterogeneity complicates drug discovery, making repurposing an attractive but challenging strategy. Advances in artificial intelligence now enable integration of multi-omics data to reveal drug-gene-disease relationships and generate subtype-specific repurposing hypotheses. In this Review, we examine AI-driven computational approaches from signature-based to multi-modal frameworks and propose an integrated interpretability-driven framework linking mechanistic validation with clinical translation toward more transparent and actionable precision oncology.

Journal Article

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

Associations Between Short Video Exposure, Empathy and Attitudes Toward End-Of-Life Care Among Nursing Students: A Cross-Sectional Study.

AIM: This cross-sectional study examined the associations between short video exposure, nursing students' empathy, and attitudes toward end-of-life (EOL) care, and tested whether perceived impact is statistically consistent with an indirect pathway in these relationships. DESIGN: A descriptive cross-sectional study. METHODS: In total, 534 undergraduate nursing students were included. Data were collected using a self-designed questionnaire, including the Attitudes Toward Care of the Dying Scale and the Jefferson Scale of Empathy-Health Professions Student version for empathy assessment. Statistical analysis for correlation and mediation analysis (PROCESS macro) was performed. RESULTS: 85.96% of students watch short videos for more than 30&#x2009;min daily, with more than 60% of them viewing EOL-related content. Students with prior caregiving experience or formal palliative care education showed significantly higher empathy and more positive attitudes (p&#x2009;<&#x2009;0.05). Exposure to medical and EOL-related short videos was positively correlated with perceived impact, empathy, and positive EOL attitudes, with effect sizes ranging from very weak to modest (r&#x2009;=&#x2009;0.10 to 0.27). The data were consistent with an indirect pathway between short video exposure and empathy via perceived impact (indirect effect&#x2009;=&#x2009;0.04; 95% bootstrap CI [0.01, 0.08]). However, for EOL attitudes, short video exposure showed a direct association rather than an indirect pathway via perceived impact (direct effect&#x2009;=&#x2009;0.09, p&#x2009;<&#x2009;0.01). CONCLUSION: In this cross-sectional study, short video exposure was modestly associated with nursing students' empathy, with data consistent with an indirect pathway via perceived impact; the observed associations explained only approximately 1% to 7% of the variance in the outcome variables. However, reshaping EOL attitudes may require more systematic education beyond brief video exposure. These findings are hypothesis-generating and await validation through longitudinal and experimental research using standardized video content. IMPLICATIONS FOR NURSING PRACTICE: Nursing educators should consider integrating curated short video content into palliative care curricula to enhance students' empathy and perceived impact of end-of-life education. However, brief video exposure alone may be insufficient to reshape deeper end-of-life attitudes, suggesting the need for comprehensive, multi-modal educational strategies.

Humans

DIVAS: an R package for identifying shared and individual variations of multiomics data.

MOTIVATION: Multiomics data integration aims to identify biological patterns shared across molecular modalities. Most existing methods detect either jointly shared variation, across all modalities, or individual variation, unique to a single modality, but overlook partially shared variation, shared by only a subset of modalities. This is a critical limitation, because many biological mechanisms manifest in some but not all molecular modalities. RESULTS: We present an open-source R package implementing data integration via analysis of subspaces (DIVAS), a framework for systematically identifying jointly shared, partially shared and individual variations across multiple data types. DIVAS combines angle-based subspace analysis with inference through rotational bootstrap, hierarchically searching all combinations of modalities to decompose multiomics data into interpretable components with scores and loadings. In simulations with a known sharing structure, DIVAS recovered every component across a wide range of noise levels, whereas existing methods did not. Applied to multi-modal COVID-19 data, it reveals partially shared immune and metabolic dysregulation patterns underpinning disease severity that conventional approaches would miss. AVAILABILITY AND IMPLEMENTATION: DIVAS is available at https://github.com/ByronSyun/DIVAS, with documentation and vignettes. The COVID-19 case study vignette is available at https://byronsyun.github.io/DIVAS_COVID19_CaseStudy/.

Multiomics