PubMed HealthSearch

SEARCH · PubMed Health

Results for “multimodal data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique.

The advent of single-cell multi-omics sequencing technology makes it possible for researchers to leverage multiple modalities for individual cells. Here, we present a protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique called moETM. We describe steps for data preprocessing, multi-omics integration, inclusion of prior pathway knowledge, and cross-omics imputation. As a demonstration, we used the single-cell multi-omics data collected from bone marrow mononuclear cells (GSE194122) as in our original study. For complete details on the use and execution of this protocol, please refer to Zhou et al.1.

Deep Learning

GICPIdb: an archival repository of multimodal data focusing on pathological images for gastrointestinal cancers.

INTRODUCTION: Deep learning (DL) shows great potential for predicting biomarkers from routine histopathological slides of gastrointestinal (GI) cancers. Yet most existing models are validated on limited patient cohorts, while pathological image annotation and molecular marker standardization demand substantial professional expertise. To address these gaps, we constructed the Gastrointestinal Cancer Pathological Image Archive (GICPIdb, gicpidb.shubuzuo.top), a dedicated database and web platform covering seven major GI cancer types. METHODS: High-quality hematoxylin and eosin (H&E)-stained whole-slide images were collected from multiple sources and uniformly processed. Image annotations were performed by board-certified pathologists following standardized protocols. GICPIdb offers five interactive web modules for data uploading, quality control, feature extraction, online annotation and AI-based prediction. Its intuitive interface supports data browsing, retrieval, visualization and downloading. RESULTS: The database houses 2,863 pathologist-annotated, uniformly processed, high-quality H&E stained images collected from 2,655 patients. Of these, 1,699 patients were sourced from The Cancer Genome Atlas (TCGA), 182 from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), and 424 from China-Japan Friendship Hospital and 350 from Chifeng Municipal Hospital in Inner Mongolia, China. It also integrates data on over 50 key molecular markers (e.g., MSI, TMB) and prognostic labels related to survival, recurrence and metastasis. DISCUSSION: GICPIdb aims to promote the development of DL-driven AI tools for cancer research and clinical translation. The multi-institutional data collection and standardized annotation pipeline are expected to enhance the generalizability and reproducibility of AI-based prediction models across diverse patient populations.

deep learning

The GSA Family in 2025: A Broadened Sharing Platform for Multi-omics and Multimodal Data.

The Genome Sequence Archive family (GSA family) provides a comprehensive suite of database resources for archiving, retrieving, and sharing multi-omics data for the global academic and industrial communities. It currently comprises four distinct database members: the Genome Sequence Archive (GSA, https://ngdc.cncb.ac.cn/gsa), the Genome Sequence Archive for Human (GSA-Human, https://ngdc.cncb.ac.cn/gsa-human), the Open Archive for Miscellaneous Data (OMIX, https://ngdc.cncb.ac.cn/omix), and the Open Biomedical Imaging Archive (OBIA, https://ngdc.cncb.ac.cn/obia). Compared to its 2021 version, the GSA family has expanded significantly by introducing a new repository, the OBIA, and by comprehensively upgrading the existing databases. Notable enhancements to the existing members include broadening the range of accepted data types, strengthening quality control systems, improving the data retrieval system, and refining data-sharing management mechanisms.

Humans

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 ± 0.0994, with a Log-rank testp-value of 1.6553×10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 ± 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 ± 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 ± 0.1211) and discrete-time survival models such as DeepHit (0.7655 ± 0.1041) and Nnet-surv (0.7694 ± 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 ± 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 ± 0.0818) and Multimodal Co-Attention Transformer (0.8102 ± 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell

Building a Digital Health Research Platform to Enable Recruitment, Enrollment, Data Collection, and Follow-Up for a Highly Diverse Longitudinal US Cohort of 1 Million People in the All of Us Research Program: Design and Implementation Study.

BACKGROUND: Longitudinal cohort studies have traditionally relied on clinic-based recruitment models, which limit cohort diversity and the generalizability of research outcomes. Digital research platforms can be used to increase participant access, improve study engagement, streamline data collection, and increase data quality; however, the efficacy and sustainability of digitally enabled studies rely heavily on the design, implementation, and management of the digital platform being used. OBJECTIVE: We sought to design and build a secure, privacy-preserving, validated, participant-centric digital health research platform (DHRP) to recruit and enroll participants, collect multimodal data, and engage participants from diverse backgrounds in the National Institutes of Health's (NIH) All of Us Research Program (AOU). AOU is an ongoing national, multiyear study aimed to build a research cohort of 1 million participants that reflects the diversity of the United States, including minority, health-disparate, and other populations underrepresented in biomedical research (UBR). METHODS: We collaborated with community members, health care provider organizations (HPOs), and NIH leadership to design, build, and validate a secure, feature-rich digital platform to facilitate multisite, hybrid, and remote study participation and multimodal data collection in AOU. Participants were recruited by in-person, print, and online digital campaigns. Participants securely accessed the DHRP via web and mobile apps, either independently or with research staff support. The participant-facing tool facilitated electronic informed consent (eConsent), multisource data collection (eg, surveys, genomic results, wearables, and electronic health records [EHRs]), and ongoing participant engagement. We also built tools for research staff to conduct remote participant support, study workflow management, participant tracking, data analytics, data harmonization, and data management. RESULTS: We built a secure, participant-centric DHRP with engaging functionality used to recruit, engage, and collect data from 705,719 diverse participants throughout the United States. As of April 2024, 87% (n=613,976) of the participants enrolled via the platform were from UBR groups, including racial and ethnic minorities (n=282,429, 46%), rural dwelling individuals (n=49,118, 8%), those over the age of 65 years (n=190,333, 31%), and individuals with low socioeconomic status (n=122,795, 20%). CONCLUSIONS: We built a participant-centric digital platform with tools to enable engagement with individuals from different racial, ethnic, and socioeconomic backgrounds and other UBR groups. This DHRP demonstrated successful use among diverse participants. These findings could be used as best practices for the effective use of digital platforms to build and sustain cohorts of various study designs and increase engagement with diverse populations in health research.

Humans

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans

transFusion: a novel comprehensive platform for integration analysis of single-cell and spatial transcriptomics.

MOTIVATION: Understanding spatial organization, intercellular interactions, and regulatory networks within the spatial context of tissues is crucial for uncovering complex biological processes and disease mechanisms. Spatial transcriptomics technologies have revolutionized this field by enabling the spatially resolved profiling of gene expression. 10× Visium has emerged as the predominant spatial technology, but its low resolution and the complexity of integrating multimodal datasets present significant analytical challenges, particularly for researchers with limited computational and statistical expertise. Current spatial transcriptomics analysis platforms generally fall short of effectively integrating multimodal data and maximizing the utility of spatial information-such as uncovering complex cellular spatial dependencies, multimodal gradient patterns, and spatial coexpression of ligand-receptor pairs and regulatory networks related to disease or biological states-thereby limiting their ability to provide comprehensive end-to-end analytical workflows when analyzing 10× Visium data. RESULTS: To address these limitations, we developed transFusion, a novel, advanced web-based platform specializing in the most comprehensive and effective integration analysis of scRNA-seq and 10× Visium spatial transcriptomics data. transFusion offers 12 key functions, from basic visualization to advanced analyses, including intercellular dependency analysis, ligand-receptor coexpression identification and visualization, and spatial multimodal gradient variation patterns. Two case studies were used to demonstrate transFusion's capabilities in exploring tissue architecture, intercellular communication, dependency networks, and multimodal gradient variation patterns with minimal computational skills and statistical expertise. transFusion provides a flexible and powerful framework for multimodal data integration analysis. AVAILABILITY AND IMPLEMENTATION: transFusion is freely available at https://github.com/WQLin8/transFusion.

Spatial Transcriptomics

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence

Next-Generation Disease Profiling by Integrating Histopathology with Spatial Multi-Omics Data.

The field of pathology has experienced several transformative changes in recent years with the advent of digital pathology and spatial multi-omics. These technologies have enhanced every aspect of pathology practice, from streamlining daily workflows to generating high-fidelity multi-omics data that provide pathologists with novel tools to refine disease profiling and clinical diagnosis. Each layer of multimodal data (genomic, metabolomic, proteomic, or transcriptomic) has uncovered a distinct facet of disease pathologies, and combined with machine learning/artificial intelligence-based data analysis and pattern recognition models, has provided holistic understanding of regulatory mechanisms underpinning them. However, high-dimensional data have far exceeded the volume, scale, and complexity of immunostaining methods implemented by pathologists and, thus, have generated significant challenges related to deconvolution, interpretation, and clinical translation. Furthermore, these multimodal studies have predominantly relied on computational methods to process data and extract disease-relevant insights, thus raising questions around relevance or role of a pathologist in this new era of multi-omics. This review will provide a perspective on the evolving fields of molecular histopathology and spatial -omics, leveraging them to approach disease profiling, and redefining the role of a pathologist during this process.

Humans

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans