PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multimodal learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Multimodal classical conditioning of fear: contributions of direct, observational, and verbal experiences to current fears.

The authors propose that a multimodal classical conditioning model be considered when clinicians or clinical researchers study the etiology of fears and anxieties learned by human beings. They argue that fears can be built through the combined effects of direct, observed, and verbally presented classical conditioning trials. Multimodal classical conditioning is offered as an alternative to the three pathways to fear argument prominent in the human fear literature. In contrast to the three pathways position, the authors present theoretical arguments for why "learning by observation" and "learning through the receipt of verbal information" should be considered classical conditioning through observational and verbal modes. The paper includes a demonstration of how data, commonly collected in research on the three pathways to fear, would be studied differently using a multimodal classical conditioning perspective. Finally, the authors discuss implications for assessment, treatment, and prevention of learned fears in humans.

Adolescent↗

The role of circulating tumor DNA (ctDNA) to detect minimal residual disease in locally advanced gastroesophageal carcinoma: the BUTTERFLY study.

BACKGROUND: Despite advances in perioperative and neoadjuvant strategies, patients with locally advanced gastroesophageal cancers remain at high risk of recurrence after curative intent treatment. No validated biomarkers are available to detect minimal residual disease (MRD) or to guide post-operative risk-adapted management. Circulating tumor DNA (ctDNA) has emerged as a noninvasive tool for disease monitoring; single-parameter or tumor-informed assays, however, may lack sensitivity in low-tumor burden settings. Multimodal, tumor-agnostic approaches may overcome these limitations. METHODS: The BUTTERFLY study is a prospective, multicenter observational study enrolling patients with stage II-III gastric, gastroesophageal junction, or esophageal cancer treated with perioperative chemotherapy or neoadjuvant chemoradiotherapy followed by surgery. It evaluates the diagnostic performance and prognostic value of an academic, tumor-agnostic, multimodal ctDNA assay for MRD detection and prognostic stratification. Serial plasma samples are collected from baseline through post-operative follow-up and at relapse. Cell-free DNA is analyzed using the Agnostic Liquid Biopsy Multimodal Advancement (ALMA) platform, integrating tumor fraction estimation, somatic copy number alterations, fragmentomic features, single-nucleotide variants, and whole-genome methylation profiling. Multimodal features are combined with clinical variables using machine learning-based models to enhance MRD detection and relapse risk stratification. The primary endpoint includes sensitivity and specificity of ALMA-defined ctDNA/MRD status at the 4-8 weeks after surgery landmark, whereas secondary endpoints assess diagnostic performance at other time points and associations between ctDNA status and dynamics with disease-free survival, overall survival, treatment response, and lead time to recurrence. FUTURE PERSPECTIVES: If validated, this tumor-agnostic, multimodal ctDNA approach may enable earlier molecular relapse detection and support personalized post-operative management strategies.

circulating tumor DNA (ctDNA)↗

Implicit multisensory associations influence voice recognition.

Natural objects provide partially redundant information to the brain through different sensory modalities. For example, voices and faces both give information about the speech content, age, and gender of a person. Thanks to this redundancy, multimodal recognition is fast, robust, and automatic. In unimodal perception, however, only part of the information about an object is available. Here, we addressed whether, even under conditions of unimodal sensory input, crossmodal neural circuits that have been shaped by previous associative learning become activated and underpin a performance benefit. We measured brain activity with functional magnetic resonance imaging before, while, and after participants learned to associate either sensory redundant stimuli, i.e. voices and faces, or arbitrary multimodal combinations, i.e. voices and written names, ring tones, and cell phones or brand names of these cell phones. After learning, participants were better at recognizing unimodal auditory voices that had been paired with faces than those paired with written names, and association of voices with faces resulted in an increased functional coupling between voice and face areas. No such effects were observed for ring tones that had been paired with cell phones or names. These findings demonstrate that brief exposure to ecologically valid and sensory redundant stimulus pairs, such as voices and faces, induces specific multisensory associations. Consistent with predictive coding theories, associative representations become thereafter available for unimodal perception and facilitate object recognition. These data suggest that for natural objects effective predictive signals can be generated across sensory systems and proceed by optimization of functional connectivity between specialized cortical sensory modules.

Acoustic Stimulation↗

Neuronal responsiveness to various sensory stimuli, and associative learning in the rat amygdala.

Neuronal activities were recorded from the amygdala and amygdalostriatal transition area of behaving rats during discrimination of conditioned auditory, visual, olfactory, and somatosensory stimuli associated with positive and/or negative reinforcements. Neurons were also tested with taste solution and various sensory stimuli that were not associated with reinforcement. Of the 1195 neurons tested, 475 responded to one or more sensory stimuli. Of these, 256 neurons responded exclusively to a unimodal sensory stimulus, 128 to multimodal sensory stimuli, and the remaining 91 could not be classified. Distribution of unimodal neurons was correlated with anatomical projections to the amygdala from sensory thalamus or sensory cortices. Multimodal neurons were located mainly in the basolateral and central nuclei of the amgydala. Response latencies of neurons in the basolateral nucleus were longer than those in other nuclei and neurons in the central nucleus had both short and long latencies. Neurons responsive to a given stimulus were more frequently encountered in the amygdalas of the trained rats than in those of the rats not trained to associate that stimulus with a reinforcement. Multimodal neurons that responded to conditioned and/or unconditioned stimuli used in the associative learned tasks were concentrated in the basolateral and central nuclei. The results indicate that some amygdalar neurons receive exclusive single sensory information, and the others receive information from two or more sensory inputs. Considering the long latencies and multimodal responsiveness, the basolateral and central nuclei of the amygdala might be foci where various kinds of sensory information converge. It is also suggested that the basolateral and central nuclei of the amygdala have critical roles in associative learning to relate sensory information to reinforcement or affective significance.

Acoustic Stimulation↗

Heterarchical reinforcement-learning model for integration of multiple cortico-striatal loops: fMRI examination in stimulus-action-reward association learning.

The brain's most difficult computation in decision-making learning is searching for essential information related to rewards among vast multimodal inputs and then integrating it into beneficial behaviors. Contextual cues consisting of limbic, cognitive, visual, auditory, somatosensory, and motor signals need to be associated with both rewards and actions by utilizing an internal representation such as reward prediction and reward prediction error. Previous studies have suggested that a suitable brain structure for such integration is the neural circuitry associated with multiple cortico-striatal loops. However, computational exploration still remains into how the information in and around these multiple closed loops can be shared and transferred. Here, we propose a "heterarchical reinforcement learning" model, where reward prediction made by more limbic and cognitive loops is propagated to motor loops by spiral projections between the striatum and substantia nigra, assisted by cortical projections to the pedunculopontine tegmental nucleus, which sends excitatory input to the substantia nigra. The model makes several fMRI-testable predictions of brain activity during stimulus-action-reward association learning. The caudate nucleus and the cognitive cortical areas are correlated with reward prediction error, while the putamen and motor-related areas are correlated with stimulus-action-dependent reward prediction. Furthermore, a heterogeneous activity pattern within the striatum is predicted depending on learning difficulty, i.e., the anterior medial caudate nucleus will be correlated more with reward prediction error when learning becomes difficult, while the posterior putamen will be correlated more with stimulus-action-dependent reward prediction in easy learning. Our fMRI results revealed that different cortico-striatal loops are operating, as suggested by the proposed model.

Association Learning↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

Cognitive therapy for multiple sclerosis: a preliminary study.

BACKGROUND: Drug treatments for multiple sclerosis are expensive, may cause side effects, and do not have demonstrated efficacy for cognitive deficits associated with this disease. OBJECTIVE: To test the effectiveness of a multimodal cognitive therapy on cognitive and physical measures known to be affected in multiple sclerosis. DESIGN: Quasi-experimental wait-list control. SETTING: Alternative medicine clinic. PATIENTS: 27 persons with clinically definite multiple sclerosis. INTERVENTION: Multimodal cognitive therapy. MAIN OUTCOME MEASURES: Neuropsychological measures of verbal learning and memory, abstraction, vocabulary, and information processing speed; Beck Depression Inventory; tactile sensitivity of the hands; grip strength; and visual acuity. MAIN RESULTS: 12 of 14 patients in the therapy group and 10 of 13 patients in the control group completed 24 weeks of treatment and all assessments. Patients who received therapy showed significantly greater improvement in verbal learning, verbal abstraction, depression, and some measures of grip strength and tactile sensitivity than did patients in the untreated control group. The groups did not differ in the magnitude of change on vocabulary, information processing speed, or visual acuity. CONCLUSION: Cognitive therapy appears to be a promising treatment for ameliorating some symptoms of multiple sclerosis. A larger study with a randomized design and additional outcome measures is warranted.

Adult↗

Research agenda to advance anhedonia assessment, understanding and treatment: an ECNP-GALENOS expert meeting report.

Anhedonia, broadly defined as a reduced ability to experience interest or pleasure, represents an important transdiagnostic neuropsychiatric symptom dimension which may benefit from targeted diagnostics and treatments. Different lines of research have proposed that it comprises multiple facets, including deficits in anticipatory ('wanting') and consummatory ('liking') reward processing as well as reward learning and affects different aspects of life (eg, social, physical, cognitive). Certain facets-more specifically anticipation, motivation and reward learning-likely involve blunted phasic dopaminergic signalling. However, recent meta-analytical evidence of human depression studies indicates that prodopaminergic antidepressants produce relatively small improvements in anhedonia symptoms and suggest that mechanisms beyond dopamine likely contribute to anhedonia. This stimulated an expert meeting to review the literature and define priorities for future research in anhedonia. A central key priority is developing a translational biologically-informed nomenclature and consensus that solves the current mismatch between constructs, paradigms and measures, and mechanisms, which separates discrete reward-related processes such as effort allocation, reward learning and anticipatory interest versus consummatory pleasure. Clinical research priorities are improved multimodal measurement tools, integrating neurobiological frameworks (eg, neuroimaging, electrophysiology and liquid biomarkers capturing dopaminergic, glutamatergic, opioid and immunometabolic pathways) and transdiagnostic studies across neuropsychiatric disorders and developmental stages. Innovative trial designs that explicitly target anhedonic phenotypes as a primary outcome and test mechanism-based interventions are also needed. Translational research recommendations include back-translation strategies that begin with patient-relevant phenotypes followed by the development of comparable human and animal tasks that target reward-related processes, such as effort allocation, reward learning and anticipatory interest versus consummatory pleasure, improve cross-species behavioural paradigms and enhance methodological rigour and reproducibility. Collectively, these recommendations will help refine the conceptualisation of anhedonia and advance its role within precision psychiatry as a mechanistically grounded target across multiple disorders.

Humans↗

[Significance of palliative resection of gastrointestinal tumors].

Before any palliative tumor resection, the morbidity and mortality risks must be carefully weighed against the continued prognosis (including quick and lasting relief of discomfort from the tumor) and alternative strategies such as bypass, chemotherapy, and radiotherapy. Multimodal concepts have seen considerable progress in recent years, and endoscopic and interventional methods have expanded the instrumentarium for palliative tumor therapy. Thus the value of palliative resection must be reassessed. The most important criteria and study results are described here, as they have resulted in increased interest in palliative tumor resection within a multimodal treatment for most gastrointestinal tumors. More studies are needed to learn how much can realistically be expected of these new approaches.

Chemotherapy, Adjuvant↗

Efforts towards a precision medicine approach in juvenile idiopathic arthritis.

Juvenile idiopathic arthritis (JIA) is the commonest group of childhood arthritides. Despite the availability of advanced therapeutics, many children and young people (CYP) with JIA experience disease flares, and in some, chronic joint damage. Tailoring treatment based on unique biological profiles would benefit CYP with JIA given their variable clinical presentation and disease course. To date, biomarkers to predict treatment response are lacking. With advances in single cell technologies, we are now able to profile the genes and proteins of target tissues at unprecedented resolution to define the biological basis of disease and guide novel treatment approaches. The complex analyses and combination of biological and clinical outcome data from large datasets across disease phenotypes have become possible with the development of computational and machine learning methods. Here, we summarize the strategies to integrate data through multimodal based approaches to maximize precision medicine and research priorities for CYP with JIA.

Humans↗

Electroencephalographic biofeedback for the treatment of attention-deficit hyperactivity disorder in childhood and adolescence.

Considerable scientific effort has been directed at developing effective treatments for attention-deficit hyperactivity disorder (ADHD). Among alternative treatment approaches, electroencephalographic (EEG) biofeedback has gained promising empirical support in recent years. Short-term effects were shown to be comparable to those of stimulant medication at the behavioral and neuropsychological level, leading to significant decreases of inattention, hyperactivity and impulsivity. In addition, EEG biofeedback results in concomitant improvement of neurophysiological patterns. EEG biofeedback may already be used within a multimodal setting, providing affected children and adolescents with a means of learning to counterbalance their ADHD symptoms without side effects. However, there is still a strong need for more empirically and methodologically sound evaluation studies.

Adolescent↗

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer↗

[New trends in the search for nootropic preparations].

The paper describes the effects of the new nootropic agents nooglutyl and glycine N-phenyl-L-prolyl ethyl ester (GVS-111). Nooglutyl, a derivative of L-glutamic and oxynicotinic acids, that has glutamatergic effects is a highly active drug in treating disturbances of memory and learning, protecting against ischemic neuronal damage and brain injury. GVS-111 is a substituted prolyl dipeptide that has the properties of enhancing cognitive functions and is able to prevent the learning impairment provoked by shock, scopolamine, brain injury, and other damages. Multimodal mechanisms responsible for the nootropic effects of nooglutyl and GVS-111.

Amnesia↗

Sound improves visual discrimination learning in avian predators.

Aposematic insects use warning colours to deter predators, but many also produce odours or sounds when attacked by a predator. One possible role for these additional components is that they promote the association between the warning colour and the non-profitability it signals, thus reducing the chance of future attacks from visually hunting predators. This experiment explicitly tests this idea by looking at the effects of sound on a visual discrimination task. Young domestic chicks were trained to look for food rewards under coloured paper cones scattered in an experimental arena. In a subsequent visual discrimination task, they learned to discriminate between rewarded and non-rewarded hats on the basis of colour. Half the chicks performed this task in silence, whilst the other half had a tone played when they attacked non-rewarded hats. The presence of the tone improved the speed of colour discrimination learning. This demonstrates that there could be a selective advantage for aposematic coloured insects to emit sounds when attacked, since avian predators will learn to avoid their coloration more quickly. The role of psychological interactions between signal components in receivers is discussed in relation to the evolution of multimodal displays.

Acoustic Stimulation↗

Multimodal signals: enhancement and constraint of song motor patterns by visual display.

Many birds perform visual signals during their learned songs, but little is known about the interrelationship between visual and vocal displays. We show here that male brown-headed cowbirds (Molothrus ater) synchronize the most elaborate wing movements of their display with atypically long silent periods in their song, potentially avoiding adverse biomechanical effects on sound production. Furthermore, expiratory effort for song is significantly reduced when cowbirds perform their wing display. These results show a close integration between vocal and visual displays and suggest that constraints and synergistic interactions between the motor patterns of multimodal signals influence the evolution of birdsong.

Abdominal Muscles↗

Scalable, generalizable and uncertainty-aware integration of spatial multiomics across diverse modalities and platforms with SCIGMA.

Recent advances in spatial omics technologies have enabled simultaneous profiling of transcriptomic, proteomic, epigenomic, metabolomic and imaging data at high spatial resolution, offering unprecedented opportunities to dissect tissue complexity. However, integrating these diverse and large-scale spatial multimodal datasets remains a major computational challenge. We present SCIGMA, a scalable and generalizable deep learning framework for spatial multiomics integration. SCIGMA introduces an uncertainty-aware contrastive learning objective and multiview graph neural networks to preserve modality-specific signals while learning biologically meaningful joint representations. Unlike previous methods, SCIGMA provides spatially resolved uncertainty estimates, interpretably identifying regions of biological or technical heterogeneity. SCIGMA supports integration of up to five modalities, and its modular framework is extensible to future technologies with even more modalities. It also scales to more than 1 million spatial locations, enabling analysis of high-resolution datasets such as Visium HD and Xenium Prime. We evaluated SCIGMA across 19 datasets spanning 8 modalities, 10 tissues and 9 platforms. On benchmarkable datasets, SCIGMA outperformed other methods in spatial domain detection, modality preservation, feature reconstruction and reproducibility. SCIGMA identifies biologically meaningful structures, refined spatial domains and modality-specific regulatory programs, providing a robust, flexible and future-ready solution for scalable spatial multimodal integration.

Multiomics↗

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans↗

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis↗