PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multimodal Fusion”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks↗

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R² of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans↗

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans↗

CT and SPECT image registration and fusion for spatial localization of metastatic processes using radiolabeled monoclonals.

The fusion of computed tomography (CT) and single-photon emission computerized tomography (SPECT) antibody images can enhance the information provided by either single modality by providing precise anatomical-functional correlation. Functional abnormalities seen on low resolution SPECT antibody images can be precisely located with specific anatomic structures seen in high resolution CT images. External fiducials located on each image modality aid in the automated registration, alignment and matching of CT and SPECT antibody images. The potential benefits of multimodal fusion include (A) the discrimination of more subtle activity peaks using anatomic organ segmentation, (B) temporal discrimination of recurrent disease, (C) assessment of residual activity post-surgery and (D) automated localization of significant focal activity. In addition, the correlation of function with anatomy may be used to establish the physiologic status of ambiguously identified objects in the anatomic image.

Aged↗

Foundation model based multimodal transformer framework for survival analysis in HER2 stratified breast cancer.

Objective. To improve survival prediction for HER2-positive breast cancer by integrating histopathological, molecular, and clinical data using a multimodal transformer framework.Approach. We propose a multimodal transformer framework for breast cancer survival prediction using HER2 stratified (SurvMBC), a foundation model-enhanced architecture that fuses three data modalities: whole-slide images, clinical narratives, and molecular features. Tumor microenvironment features are extracted using a pathology language and image pre-training (PLIP), clinical narratives are processed with BioBERT, and miRNA expression plus DNA methylation data are embedded using Gen2Vec. These representations are integrated through a cross-modal transformer with attention mechanisms for survival prediction.Main results. The model was evaluated on 1,095 HER2-positive breast cancer patients from The Cancer Genome Atlas. SurvMBC achieved a concordance index (C-index) of 0.857 (95% CI: 0.834, 0.880), a low integrated Brier score, and a strong inverse negative binomial log-likelihood. Risk stratification based on model outputs significantly separated high- and low-risk groups (log-rankp< 0.01) and showed strong associations with tumor stage, grade, and hormone receptor status (allp< 0.05).Significance. SurvMBC demonstrates the effectiveness of multimodal fusion in addressing tumor heterogeneity and improving prognostic accuracy. The attention-based integration enables context-aware learning of survival-relevant features across modalities, supporting individualized risk stratification and risk-adaptive treatment planning for HER2 stratified breast cancer patients.

Breast Neoplasms↗

Peptide molecular lock-engineered nanobodies enable an oriented dual-modal immunoassay for reliable detection of Cronobacter sakazakii.

Conventional nanobody ELISAs for trace Cronobacter sakazakii in powdered infant formula suffer from random orientation and low signal output. We developed an oriented dual-modal immunoassay that combines site-specific biotinylation via a C-terminal AviTag and a peptide molecular lock, enabling controlled surface orientation while preserving nanobody structural integrity. This strategy was further integrated with phage-displayed nanobodies for multivalent amplification and both fluorescent and colorimetric readouts. The assay exhibited a broad linear range of 103-106&#xa0;CFU/mL, with limits of detection (LODs) of 6.70&#xa0;&#xd7;&#xa0;102&#xa0;CFU/mL for fluorescence and 1.55&#xa0;&#xd7;&#xa0;103&#xa0;CFU/mL for colorimetry, showing improved sensitivity compared with the conventional passive adsorption-based Nb-ELISA evaluated in this study. XGBoost-based multimodal fusion improved quantitative accuracy, and SHAP analysis elucidated modality contributions. In spiked powdered infant formula samples, recoveries ranged from 92.1% to 118% with coefficients of variation below 5.98%, confirming acceptable matrix tolerance and analytical reliability.

Cronobacter sakazakii↗

Automatic matching of homologous histological sections.

The role of neuroanatomical atlases is undergoing a significant redefinition as digital atlases become available. These have the potential to serve as more than passive guides and to hold the role of directing segmentation and multimodal fusion of experimental data. Key elements needed to support these new tasks are registration algorithms. For images derived from histological procedures, the need is for techniques to map the two-dimensional (2-D) images of the sectional material into the reference atlas which may be a full three-dimensional (3-D) data set or one consisting of a series of 2-D images. A variety of 2-D-2-D registration methods are available to align experimental images with the atlas once the corresponding plane of section through the atlas has been identified. Methods to automate the identification of the homologous plane, however, have not been previously reported. In this paper we use the external section contour to drive the identification and registration procedure. For this purpose, we model the contours by B-splines because of their attractive properties the most important of which are: 1) smoothness and continuity; 2) local controllability which implies that local changes in shape are confined to the B-spline parameters local to that change; 3) shape invariance under affine transformation, which means that the affine transformed curve is still a B-spline whose control points are related to the object control points through the transformation. In this paper we present a fast algorithm for estimating the control points of the B-spline which is robust to nonuniform sampling, noise, and local deformations. Curve matching is achieved by using a similarity measure that depends directly on the parameters of the B-spline. Performance tests are reported using histological material from rat brains.

Animals↗

Registering coronal histological 2-D sections of a rat brain with coronal sections of a 3-D brain atlas using geometric curve invariants and B-spline representation.

A new approach is proposed for registering a set of histological coronal two-dimensional images of a rat brain sectional material with coronal sections of a three-dimensional brain atlas, an intrinsic step and a significant challenge to current efforts in brain mapping and multimodal fusion of experimental data. The alignment problem is based on matching external contours of the brain sections, and operates in the presence of tissue distortion and tears which are routinely encountered, and possible scale, rotation, and shear changes (the affine and weak perspective groups). It is based on a novel set of local absolute affine invariants derived from the set of ordered inflection points on the external contour represented by a cubic B-spline curve. The inflection points are local intrinsic geometric features, which are preserved under both the affine and the weak perspective transformations. The invariants are constructed from the sequence of area patches bounded by the contour and the line connecting two consecutive inflection points, and hence do make direct use of the area (volume) invariance property associated with the affine transformation. These local absolute invariants are very well suited to handle the tissue distortion and tears (occlusion problem).

Algorithms↗

SPECT in the year 2000: basic principles.

OBJECTIVE: SPECT has become a routine procedure in most nuclear medicine departments. SPECT provides significant technical challenges for the nuclear medicine technologist, as compared with planar imaging, in the areas of SPECT acquisition, image reconstruction, and data processing. Many new advances in SPECT methodology are becoming available, such as iterative reconstruction, multimodality fusion, and advanced gated cardiac SPECT. SPECT imaging is demanding and requires careful attention to proper acquisition protocols, whether circular or noncircular orbits, and postprocessing is becoming more complex with the addition of iterative reconstruction and attenuation correction algorithms, among others. Understanding the principles of SPECT is essential not only to produce the highest quality scans but also to identify image artifacts. After reading this article, the nuclear medicine technologist should be able to: (a) describe the historical development and benefits of SPECT imaging; (b) state the impact of image matrix size, number of projections, and arc of rotation on final SPECT image quality; (c) discuss the trade-offs between image noise content and spatial and contrast resolution in SPECT reconstruction; (d) discuss SPECT filters and their impact on image quality; (e) explain the differences between filtered backprojection and iterative reconstruction; and (f) describe the impact of attenuation and scatter in SPECT imaging and the advantages and pitfalls of attenuation correction methods.

Humans↗

Imaging prostate cancer with 11C-choline PET/CT.

UNLABELLED: The ability of 11C-choline and multimodality fusion imaging with integrated PET and contrast-enhanced CT (PET/CT) was investigated to delineate prostate carcinoma (PCa) within the prostate and to differentiate cancer tissue from normal prostate, benign prostate hyperplasia, and focal chronic prostatitis. METHODS: All patients with PCa gave written informed consent. Twenty-six patients with clinical stage T1, T2, or T3 and biopsy-proven PCa underwent 11C-choline PET/CT after intravenous injection of 1,112 +/- 131 MBq 11C-choline, radical retropubic prostatovesiculectomy, and standardized prostate tissue sampling. Maximal standardized uptake values (SUVs) of 11C-choline within 36 segments of the prostate were determined. PET/CT results were correlated with histopathologic results, prostate-specific antigen (PSA), Gleason score, and pT stage. RESULTS: The SUV of 11C-choline in PCa tissue was 3.5 +/- 1.3 (mean +/- SD) and significantly higher than that in prostate tissue with benign histopathologic lesions (2.0 +/- 0.6; P < 0.001 benign histopathology vs. cancer). Visual and quantitative analyses of segmental 11C-choline uptake of each patient unambiguously located PCa in 26 of 26 patients and 25 of 26 patients, respectively. A threshold SUV of 2.65 yielded an area under the receiver-operating-characteristic (ROC) curve of 0.89 +/- 0.01 for correctly locating PCa. The maximal 11C-choline SUV did not correlate significantly with PSA or Gleason score but did correlate with T stage (P = 0.01; Spearman r = 0.49). CONCLUSION: 11C-Choline PET/CT can accurately detect and locate major areas with PCa and differentiate segments with PCa from those with benign hyperplasia, chronic prostatitis, or normal prostate tissue. The maximal tumoral 11C-choline uptake is related to pT stage.

Aged↗

Computer applications in radiology.

Computer applications in radiology are evolving rapidly, tied to incremental improvements in hardware, software, and methods. In computer hardware, the emergence of dramatically improved graphic and computational performance for engineering workstations enables their use for visualization. Major changes in networking, storage, and display technology play a major role in influencing applications. The use of three-dimensional digitizers to perform localization of real three-dimensional points in conjunction with images and the rendering of objects using rapid prototyping methods, such as stereolithography, were recently reported. Major software advances have taken place through the availability of applications packages that are operated with menu-driven or point-and-click user interfaces, data flow languages, or complete turnkey applications. Imaging methods including CT, MR imaging, digital radiography, biomagnetism, and optical range sensing, which take advantage of advanced computer technology, are new this year. Image processing for multimodality fusion or image registration, visualization, reconstruction, and quantification of images, have been reported at a wide variety of conferences and in key publications. New computer methods to fabricate custom orthopaedic implants, and to improve imaging technology assessment were introduced.

Brain↗

A data fusion environment for multimodal and multi-informational neuronavigation.

OBJECTIVE: Part of the planning and performance of neurosurgery consists of determining target areas, areas to be avoided, landmark areas, and trajectories, all of which are components of the surgical script. Nowadays, neurosurgeons have access to multimodal medical imaging to support the definition of the surgical script. The purpose of this paper is to present a software environment developed by the authors that allows full multimodal and multi-informational planning as well as neuronavigation for epilepsy and tumor surgery. MATERIALS AND METHODS: We have developed a data fusion environment dedicated to neuronavigation around the Surgical Microscope Neuronavigator system (Carl Zeiss, Oberkochen, Germany). This environment includes registration, segmentation, 3D visualization, and interaction-applied tools. It provides the neuronavigation system with the multimodal information involved in the definition of the surgical script: lesional areas, sulci, ventricles segmented from magnetic resonance imaging (MRI), vessels segmented from magnetic resonance angiography (MRA), functional areas from magneto-encephalography (MEG), and functional magnetic resonance imaging (fMRI) for somatosensory, motor, or language activation. These data are considered to be relevant for the performance of the surgical procedure. The definition of each entity results from the same procedure: registration to the anatomical MRI data set (defined as the reference data set), segmentation, fused 3D display, selection of the relevant entities for the surgical step, encoding in 3D surface-based representation, and storage of the 3D surfaces in a file recognized by the neuronavigation software (STP 3.4, Leibinger; Freiburg, Germany). RESULTS: Multimodal neuronavigation is illustrated with two clinical cases for which multimodal information was introduced into the neuronavigation system. Lesional areas were used to define and follow the surgical path, sulci and vessels helped identify the anatomical environment of the surgical field, and, finally, MEG and fMRI functional information helped determine the position of functional high-risk areas. CONCLUSION: In this short evaluation, the ability to access preoperative multi-functional and anatomical data within the neuronavigation system was a valuable support for the surgical procedure.

Adult↗

Probing plasma membrane microdomains in cowpea protoplasts using lipidated GFP-fusion proteins and multimode FRET microscopy.

Summary Multimode fluorescence resonance energy transfer (FRET) microscopy was applied to study the plasma membrane organization using different lipidated green fluorescent protein (GFP)-fusion proteins co-expressed in cowpea protoplasts. Cyan fluorescent protein (CFP) was fused to the hyper variable region of a small maize GTPase (ROP7) and yellow fluorescent protein (YFP) was fused to the N-myristoylation motif of the calcium-dependent protein kinase 1 (LeCPK1) of tomato. Upon co-expressing in cowpea protoplasts a perfect co-localization at the plasma membrane of the constructs was observed. Acceptor-photobleaching FRET microscopy indicated a FRET efficiency of 58% in protoplasts co-expressing CFP-Zm7hvr and myrLeCPK1-YFP, whereas no FRET was apparent in protoplasts co-expressing CFP-Zm7hvr and YFP. Fluorescence spectral imaging microscopy (FSPIM) revealed, upon excitation at 435 nm, strong YFP emission in the fluorescence spectra of the protoplasts expressing CFP-Zm7hvr and myrLeCPK1-YFP. Also, fluorescence lifetime imaging microscopy (FLIM) analysis indicated FRET because the CFP fluorescence lifetime of CFP-Zm7hvr was reduced in the presence of myrLeCPK1-YFP. A FRET fluorescence recovery after photobleaching (FRAP) analysis on a partially acceptor-bleached protoplast co-expressing CFP-Zm7hvr and myrLeCPK1-YFP revealed slow requenching of the CFP fluorescence in the acceptor-bleached area upon diffusion of unbleached acceptors into this area. The slow exchange of myrLeCPK1-YFP in the complex with CFP-Zm7hvr reflects a relatively high stability of the complex. Together, the FRET data suggest the existence of plasma membrane lipid microdomains in cowpea protoplasts.

Base Sequence↗

[Fusion of data. Multimodality in neurological imaging].

The development of tomographic imaging methods, which provide anatomical and functional information in a digital form, has transformed the approach to the central nervous system. Patient management is now frequently based on fusion of data from the same or different modalities, which potentiates the performance of each technique. These fusions were initially performed manually on film supports, but are now increasingly performed by consultation stations which process digital data bases. In view of the increasing demands for data fusion, it has become necessary to develop, in parallel with these new techniques, image transmission networks and visualisation stations on which reconstructions and fusions are performed. Improvements in software should facilitate acquisition techniques which should increasingly resemble standard techniques. The logistic applied must also be as simple as possible in order to be widely implanted and easy to use.

Central Nervous System Diseases↗

Multimodality cranial image fusion using external markers applied via a vacuum mouthpiece and a case report.

PURPOSE: To present a simple and precise method of combining functional information of cranial SPECT and PET images with CT and MRI, in any combination. MATERIAL AND METHODS: Imaging is performed with a hockey mask-like reference frame with image modality-specific markers in precisely defined positions. This frame is reproducibly connected to the VBH vacuum mouthpiece, granting objectively identical repositioning of the frame with respect to the cranium. Using these markers, the desired 3-D imaging modalities can then be manually or automatically registered. This information can be used for diagnosis, treatment planning, and evaluation of follow-up, while the same vacuum mouthpiece allows precisely reproducible stereotactic head fixation during radiotherapy. RESULTS: 244 CT and MR data sets of 49 patients were registered to a root square mean error (RSME) of 0.9 mm (mean). 64 SPECT-CT fusions on 18 of these patients gave an RMSE of 1.4 mm, and 40 PET-CT data sets of eight patients were registered to 1.3 mm. An example of the method is given by means of a case report of a 52-year-old patient with bilateral optic nerve meningioma. CONCLUSION: This technique is a simple, objective and accurate registration tool to combine diagnosis, treatment planning, treatment, and follow-up, all via an individualized vacuum mouthpiece. Especially for low-resolution PET and even more so for some very diffuse SPECT data sets, activity can now be accurately correlated to anatomic structures.

Equipment Design↗

A novel method for the 3-dimensional simulation of orthognathic surgery by using a multimodal image-fusion technique.

INTRODUCTION: The aim of this study was to establish a novel method for simulating orthognathic surgery in 3-dimensional (3D) space. METHODS: This system mainly consists of 6 procedures: (1) reconstruction of a virtual skull model (VS) from presurgical computed tomography scans; (2) reconstruction of virtual dentition models from 3D surface scanning of dental casts occluded at presurgical and postsurgical intercuspal positions (VD1 and VD2, respectively); (3) reconstruction of a preliminary fusion model of VS and VD1 by an initial intermodality registration; (4) reconstruction of another preliminary fusion model of VS, VD1, and VD2 by a second intramodality registration; (5) repositioning of bony segments by a third intramodality registration and reconstruction of final fusion models at presurgery and postsurgery; and (6) 3D analysis of the movement of bony segments. To test this system, 2 patients with severe skeletal deformities, who had undergone presurgical orthodontic treatment, were used as models. Registration accuracy was determined by the root mean squared distance between the corresponding fiducial markers in a set of 2 images. RESULTS AND CONCLUSIONS: The sum of the root mean squared error of the 3 registration processes was less than 0.4 mm in both patients. This simulation system could be used to precisely realize the presurgical and postsurgical occlusal relationships and craniofacial morphology of a patient with severe skeletal deformities, and to quantitatively describe the movement of a given anatomical point of bony segments. It is assumed that there could be significant benefits in sharing visual and quantitative 3D information from this simulation system among orthodontists and surgeons.

Adult↗

Combination of hyperbolic functions for multimodal biometrics data fusion.

In this paper, we treat the problem of combining fingerprint and speech biometric decisions as a classifier fusion problem. By exploiting the specialist capabilities of each classifier, a combined classifier may yield results which would not be possible in a single classifier. The Feedforward Neural Network provides a natural choice for such data fusion as it has been shown to be a universal approximator. However, the training process remains much to be a trial-and-error effort since no learning algorithm can guarantee convergence to optimal solution within finite iterations. In this work, we propose a network model to generate different combinations of the hyperbolic functions to achieve some approximation and classification properties. This is to circumvent the iterative training problem as seen in neural networks learning. In many decision data fusion applications, since individual classifiers or estimators to be combined would have attained a certain level of classification or approximation accuracy, this hyperbolic functions network can be used to combine these classifiers taking their decision outputs as the inputs to the network. The proposed hyperbolic functions network model is first applied to a function approximation problem to illustrate its approximation capability. This is followed by some case studies on pattern classification problems. The model is finally applied to combine the fingerprint and speaker verification decisions which show either better or comparable results with respect to several commonly used methods.

Algorithms↗

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 &#xb1; 0.0994, with a Log-rank testp-value of 1.6553&#xd7;10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 &#xb1; 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 &#xb1; 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 &#xb1; 0.1211) and discrete-time survival models such as DeepHit (0.7655 &#xb1; 0.1041) and Nnet-surv (0.7694 &#xb1; 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 &#xb1; 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 &#xb1; 0.0818) and Multimodal Co-Attention Transformer (0.8102 &#xb1; 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell↗