PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multimodal”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 ± 0.0994, with a Log-rank testp-value of 1.6553×10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 ± 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 ± 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 ± 0.1211) and discrete-time survival models such as DeepHit (0.7655 ± 0.1041) and Nnet-surv (0.7694 ± 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 ± 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 ± 0.0818) and Multimodal Co-Attention Transformer (0.8102 ± 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence

Effects of testosterone-augmented multimodal exercise intervention in spinal cord injury: a randomized controlled trial.

CONTEXT: Spinal cord injury (SCI) leads to profound muscle atrophy, aerobic deconditioning, and metabolic dysfunction. Exercise-based interventions alone produce modest benefits. Whether testosterone can augment physiologic responses to exercise in this population remains untested. OBJECTIVE: To evaluate efficacy and safety of home-based intervention combining functional electrical stimulation-assisted leg cycling (FES-LC), arm ergometry (AE), and testosterone compared with FES-LC, AE plus placebo in adults with SCI. METHODS: This randomized, placebo-controlled, double-blind trial enrolled 84 adults (76 males and 8 females) aged 19-70 years with SCI (neurologic levels C4-T12; AIS grades A-D). Participants were randomized to multimodality intervention (home-based FES-LC, AE and intramuscular testosterone undecanoate) (n = 38) or control intervention (FES-LC, AE plus placebo) (n = 46) for 16 weeks. The primary outcome was change in aerobic capacity (peak VO2) during AE cardiopulmonary exercise testing. Secondary outcomes included lean mass, hemoglobin, cardiometabolic markers, and safety. RESULTS: Mean (SD) age was 44 (13) years and time since injury was 13.9 (13) years). Between-group changes in peak VO2 were not statistically significant. Within-group improvements were larger in multimodality (∼19% increase; 0.10 L/min; 95% CI, 0.02-0.18 L/min) compared to controls (∼6% increase; 0.06 L/min; 95% CI, -0.01-0.13). The multimodality group gained significantly more lean mass (whole-body:1.84 kg, 95% CI: 0.52-3.16, P = .007; lower extremity 0.92 kg, 95% CI: 0.38-1.45, P = .001), and anemia was corrected in a greater proportion of participants. Adverse event rates were similar between groups. CONCLUSION: A home-based multimodality intervention combining FES-LC, AE, and testosterone was safe and associated with greater improvements in lean mass and hemoglobin. Although between-group differences in aerobic capacity were not statistically significant, greater within-group increases were observed in the multimodality group. These findings may inform future studies of testosterone-augmented exercise interventions for individuals living with SCI.

Humans

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans

Dual-Reporter Gene-Based Multimodal Imaging for Tracking Mesenchymal Stem Cells in Diabetic Skin Wound Repair.

BACKGROUND: Diabetic foot ulcer (DFU) is a clinically challenging complication characterized by poor healing outcomes, and conventional therapies provide limited benefit. Mesenchymal stem cell (MSC) transplantation offers a promising strategy for DFU repair. However, the low survival of transplanted MSCs in the hostile wound microenvironment, coupled with the lack of real-time, non-invasive methods to track these cells in vivo, severely hampers their therapeutic efficacy and clinical translation. METHODS: We engineered MSCs to co-express a dual reporter system comprising near-infrared fluorescent protein (iRFP) and ferritin heavy chain (FTH1). These modified cells were then integrated with a fibrin glue (FG) scaffold to create a unified platform that supports both multimodal imaging and therapeutic function within skin wounds. First, FTH1 overexpression enhances the antioxidant capacity of MSCs, while the FG scaffold provides structural support; this combination enhances cell survival and retention. Second, the iRFP/FTH1 dual reporter enables near-infrared fluorescence imaging and MRI-based localization, establishing a multimodal platform for real-time cell tracking. RESULTS: In a full-thickness skin defect model in diabetic mice, multimodal imaging revealed that transplanted cells persisted in the wound area for approximately seven days. Treatment with iRFP/FTH1-MSCs/FG significantly accelerated wound closure and promoted hair follicle regeneration and angiogenesis. Additionally, local iron deposition resulting from FTH1 expression enhanced fibroblast migration and collagen synthesis, further facilitating extracellular matrix remodeling. Mechanistic studies demonstrated that this therapy drives macrophage polarization toward the anti-inflammatory M2 phenotype and activates the PI3K-AKT-VEGF signaling pathway. These complementary effects synergistically enhance tissue regeneration and systematically improve diabetic wound healing. CONCLUSIONS: Collectively, this multimodal stem cell-scaffold system effectively integrates dynamic cell tracking with stem cell therapy during skin wound repair. It addresses a critical technical gap in visualizing stem cells within the wound microenvironment and provides valuable methodological and theoretical foundations for optimizing regenerative strategies for diabetic skin wounds.

Animals

Nociception-guided opioid administration within multimodal analgesia for laparoscopic endometriosis surgery: a randomized controlled trial.

Women with endometriosis are at increased risk of severe postoperative pain due to nociceptive sensitization. While multimodal analgesia reduces opioid use, the added value of objective nociception monitoring remains unclear. This study evaluated whether NOL&#xae;-guided opioid titration improves perioperative outcomes within a standardized multimodal regimen. In this prospective, randomized, single-blinded trial, premenopausal women undergoing laparoscopic surgery for suspected endometriosis or adenomyosis were assigned to NOL&#xae;-guided analgesia or standard care based on clinical assessment. All patients received a standardized multimodal protocol. The primary outcome was total perioperative opioid consumption. Secondary outcomes included postoperative pain scores (NRS) and PACU length of stay. Exploratory analyses assessed the association between preoperative pain (Mankoski Pain Scale, MPS) and postoperative outcomes. A total of 111 patients were analyzed (NOL&#xae;: n&#x2009;=&#x2009;54; control: n&#x2009;=&#x2009;57). Total perioperative opioid consumption did not differ significantly between groups (adjusted mean difference&#x2009;=&#x2009;14&#xa0;&#x3bc;g for Fentanyl and 52&#xa0;&#x3bc;g for Remifentanil; p&#x2009;=&#x2009;0.8). Surgery duration was an independent predictor of opioid use (p&#x2009;<&#x2009;0.001) and PACU length of stay (p&#x2009;=&#x2009;0.01), whereas treatment group had no significant effect. Postoperative pain scores were comparable between groups at all time points. NOL&#xae;-derived metrics were not associated with opioid consumption or pain. Higher preoperative MPS scores independently predicted higher pain scores in the late PACU phase. NOL&#xae;-guided opioid titration did not reduce perioperative opioid consumption or improve early postoperative outcomes compared with standard multimodal analgesia in women undergoing laparoscopic surgery for endometriosis.

Humans

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans

Assessment of the role and effectiveness of nurse-led multimodal intervention in the rehabilitation of dysphagia in patients with brain tumors.

BACKGROUND: Dysphagia is a common complication in patients with brain tumors, which has a profound adverse impact on patients' health status and quality of life. However, there is a relative lack of research on the rehabilitation of dysphagia in brain tumor patients, especially regarding the role and effectiveness of nurse-led multimodal interventions in the rehabilitation of dysphagia in brain tumor patients, which lacks systematic assessment and in-depth discussion. AIM: This study aimed to evaluate the role and effectiveness of a nurse-led multimodal intervention in improving swallowing function and quality of life in brain tumor patients with dysphagia. METHODS: In this study, a randomized controlled trial (RCT) design was used to select 120 dysphagia patients among brain tumor patients admitted to our hospital during the period of January 2024 to May 2024 as the study subjects, and they were stratified and randomly divided into an intervention group (n&#x2009;=&#x2009;60) and a control group (n&#x2009;=&#x2009;60). While the control group received conventional nursing care and treatment protocols, the intervention group received a nurse-led multimodal intervention program, including personalized swallowing training, nutritional support, psychological care, and a family-participatory rehabilitation program, which was developed and dynamically adjusted by nurses, rehabilitation therapists, and dietitians. Differences in data before and after the intervention were analyzed using the paired t-test or Wilcoxon signed-rank test, and between-group comparisons were made using the independent samples t-test or Mann-Whitney U test. RESULTS: Both the intervention and control groups showed improvement in swallowing function among the patients. The Kubota drinking test score, Saito's swallowing function grading, and the quality of life scores for patients in the intervention group showed a significant enhancement compared to those in the control group (P&#x2009;<&#x2009;0.05), indicating that the intervention was more effective than the control. When compared within groups, all scores in both the intervention and control groups improved gradually with the time of intervention (P&#x2009;<&#x2009;0.05). The improvement was significantly higher in the intervention group than in the control group. CONCLUSION: This study demonstrates that a nurse-led multimodal intervention is significantly effective in improving swallowing function and quality of life in patients with brain tumors. The intervention provides comprehensive rehabilitation support for patients through multidisciplinary collaboration and personalized care and has certain clinical promotion value.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Impact of a multimodal prehabilitation program on postoperative cognitive dysfunction: a single-center randomized controlled trial.

BACKGROUND: Postoperative cognitive dysfunction (POCD) is a frequent complication after cardiac surgery. Exercise-based prehabilitation may enhance functional reserve and reduce vulnerability to perioperative cerebral insults. We hypothesized that multimodal prehabilitation reduces POCD 3&#xa0;months after cardiac surgery. METHODS: This prespecified substudy of a single-center randomized controlled trial (NCT03466606) included patients aged &#x2265;50&#xa0;years undergoing elective coronary artery bypass grafting and/or valve surgery. Participants were randomized 1:1 to 4-6&#xa0;weeks of multimodal prehabilitation (exercise training, nutritional support, and psychological support) or standard preoperative care. Cognitive function was assessed at baseline and 3&#xa0;months postoperatively using an age- and education-adjusted neuropsychological battery. POCD was defined as performance &#x2265;1.5 standard deviations below normative values in at least 2 cognitive tests, excluding the Mini-Mental State Examination. Logistic regression analyses were performed to evaluate factors associated with POCD. RESULTS: Of 160 participants screened from the parent trial, 134 met eligibility criteria for the substudy and were randomized; 116 completed 3-month follow-up (prehabilitation n&#xa0;=&#xa0;53; control n&#xa0;=&#xa0;63). POCD occurred in 29 patients (25%), including 15/53 (28%) in the prehabilitation group and 14/63 (22%) in controls (odds ratio [OR] 1.37, 95% confidence interval [CI] 0.54-3.50, P&#xa0;=&#xa0;0.52). In multivariable analysis, preoperative cognitive impairment was independently associated with POCD (OR 13.28, 95% CI 4.06-43.41, P&#xa0;<&#xa0;0.001), whereas prehabilitation was not (OR 1.09, 95% CI 0.35-3.45, P&#xa0;=&#xa0;0.877). Higher physical activity levels at 3&#xa0;months were associated with lower odds of POCD (OR 0.97, 95% CI 0.95-1.00, P&#xa0;=&#xa0;0.047). CONCLUSIONS: In this randomized controlled trial, a 4-6-week multimodal prehabilitation program did not reduce postoperative cognitive dysfunction 3&#xa0;months after cardiac surgery. Although the intervention did not achieve measurable cognitive protection, the observed association between postoperative physical activity levels and postoperative cognitive dysfunction warrants further investigation.

Humans

Foundation model based multimodal transformer framework for survival analysis in HER2 stratified breast cancer.

Objective. To improve survival prediction for HER2-positive breast cancer by integrating histopathological, molecular, and clinical data using a multimodal transformer framework.Approach. We propose a multimodal transformer framework for breast cancer survival prediction using HER2 stratified (SurvMBC), a foundation model-enhanced architecture that fuses three data modalities: whole-slide images, clinical narratives, and molecular features. Tumor microenvironment features are extracted using a pathology language and image pre-training (PLIP), clinical narratives are processed with BioBERT, and miRNA expression plus DNA methylation data are embedded using Gen2Vec. These representations are integrated through a cross-modal transformer with attention mechanisms for survival prediction.Main results. The model was evaluated on 1,095 HER2-positive breast cancer patients from The Cancer Genome Atlas. SurvMBC achieved a concordance index (C-index) of 0.857 (95% CI: 0.834, 0.880), a low integrated Brier score, and a strong inverse negative binomial log-likelihood. Risk stratification based on model outputs significantly separated high- and low-risk groups (log-rankp< 0.01) and showed strong associations with tumor stage, grade, and hormone receptor status (allp< 0.05).Significance. SurvMBC demonstrates the effectiveness of multimodal fusion in addressing tumor heterogeneity and improving prognostic accuracy. The attention-based integration enables context-aware learning of survival-relevant features across modalities, supporting individualized risk stratification and risk-adaptive treatment planning for HER2 stratified breast cancer patients.

Breast Neoplasms

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology