PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multimodal learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans↗

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans↗

Learning the thyroid examination--a multimodality intervention for internal medicine residents.

BACKGROUND: Many physicians have inadequate physical diagnosis skills and cannot detect thyroid abnormalities on physical examination. PURPOSE: To evaluate a multimodality intervention to improve thyroid examination skills using a prospective controlled trial in first-year residents enrolled in an academic internal medicine program. METHODS: The intervention group received a 60-minute educational session during which an endocrinologist described anatomical landmarks, thyroid abnormalities, and examination techniques using a slide show, computerized animation, videotape, and live demonstration on a volunteer with goiter. Residents examined a normal and a goitrous thyroid under the observation of a preceptor and received an evidence-based handout on the thyroid examination. The control group received no specific intervention. Examination technique and identification of thyroid abnormalities were blindly assessed in 2 stations of an objective structured clinical examination (OSCE). RESULTS: Of the 19 residents in the intervention group and the 20 in the control group, 6 (32%) and 3 (15%), respectively, observed the neck for thyroid abnormalities (P = 0.3), 17 (90%) and 16 (80%) used proper hand position (P = 0.7), and 13 (68%) and 15 (75%) had the patient swallow while the neck was palpated (P = 0.7). There was a significant difference in the mean scores based on thyroid physical findings during the OSCE between the intervention and control groups (100 vs. 52.5 [maximal possible score = 200], P = 0.047). CONCLUSION: A 1-hour multimodality learning session furthered the ability of first-year internal medicine residents to detect thyroid abnormalities.

Clinical Competence↗

Category learning through multimodality sensing.

Humans and other animals learn to form complex categories without receiving a target output, or teaching signal, with each input pattern. In contrast, most computer algorithms that emulate such performance assume the brain is provided with the correct output at the neuronal level or require grossly unphysiological methods of information propagation. Natural environments do not contain explicit labeling signals, but they do contain important information in the form of temporal correlations between sensations to different sensory modalities, and humans are affected by this correlational structure (Howells, 1944; McGurk & MacDonald, 1976; MacDonald & McGurk, 1978; Zellner & Kautz, 1990; Durgin & Proffitt, 1996). In this article we describe a simple, unsupervised neural network algorithm that also uses this natural structure. Using only the co-occurring patterns of lip motion and sound signals from a human speaker, the network learns separate visual and auditory speech classifiers that perform comparably to supervised networks.

Algorithms↗

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks↗

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans↗

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning↗

[Learning curve--calculation and value in laparoscopic surgery].

The learning curve shows the progress in mastering a new method. It is completed when the monitored parameters reach a steady state and when the final results can be compared with literature. The earlier used analysis of the performance-improvement with its "on the spots" appraisals at certain time-intervals is replaced by a continuous assessment. The multimode learning curve is particularly useful for it, because not only one parameter (f.e. operation-time), but also several important factors can be put together into one single graphic. For the operation-time, the Moving Average Method is useful. For incidents, which may happen or not like a conversion from laparoscopy to laparotomy as well as complications, the Cusum-method is of practical use. The learning curves of the technique of laparoscopic cholecystectomy, colo-rectal surgery, fundoplicatio and hernia surgery have been completed. Also, the learning curve of the industry is well advanced. Reliable data for the learning curves of individual surgeons for certain operations cannot be given, as, only now, young doctors are being trained on a large scale in laparoscopic technique as used to be the case in the open abdominal surgery. This will influence greatly the learning curves and will shorten the time till their completion. Different bias concerning the individual surgeons and their clinics prohibit the production of comparable curves. Several factors like the patient respectively his abdomen are complicating all this. That's why the learning curves cannot be used as benchmarks to compare different surgeons or clinics, as long as no valid scoring system concerning the complexity of a surgical intervention exists. Learning curves which become quality curves after reaching a steady state, can be used for the individual monitoring of a surgeon's performance and serve as a quality measurement of a clinic. The learning curves of the laparoscopic cholecystectomy, fundoplicatio, colo-rectal surgery and hernia surgery are discussed in particular The mandatory number of operations needed to learn a new method cannot yet be established today, even if all the existing data are consulted. Therefore, the learning curve is a useful instrument to monitor the individual progress and the results of a clinic in the meaning of an individual quality-management. After completion of the learning curve, a quality curve using the same parameters will be given, which shows the deviations of its own standard.

Cholecystectomy, Laparoscopic↗

Decoding cancer with artificial intelligence: Transforming research, diagnosis, and therapy with future insights.

Cancer remains one of the leading global health burdens, with increasing complexity in genomic, imaging, and clinical datasets presenting significant challenges for effective management. Artificial intelligence (AI) has emerged as a powerful tool to address these challenges by enabling pattern recognition, knowledge integration, and data-driven decision-making. This review highlights recent advances in the application of AI across cancer research, diagnosis, and therapy. In research, AI accelerates drug discovery and repurposing, enhances genomic data interpretation, and facilitates biomarker identification through multi-omics integration. In diagnosis, AI has demonstrated high technical performance in radiology for lesion detection and image segmentation, in pathology for tumour grading and molecular prediction, and in liquid biopsy for non-invasive biomarker analysis. In therapy, AI supports precision medicine by predicting treatment responses, monitoring disease progression, and optimizing clinical trial design. Despite these advances, barriers such as data heterogeneity, algorithmic bias, interpretability, and regulatory challenges remain. Future directions, including explainable AI, federated learning, multimodal modelling, and digital twins, hold promise for translating AI-driven innovations into routine oncology practice. Significance Statement This review provides a timely synthesis of recent (2020-2025) advances in artificial intelligence across cancer research, diagnosis, and therapy, highlighting applications in drug discovery, genomics, multi-omics biomarker identification, and clinical decision-making. By integrating technological progress with translational and clinical relevance, this work serves as a valuable resource for bridging AI innovation with precision oncology practice. As a narrative review, the literature was identified through targeted PubMed, Scopus, and Google Scholar searches, combining terms for artificial intelligence, machine learning, and deep learning with cancer-related keywords, with priority given to peer-reviewed studies published between 2020 and 2025, seminal earlier works, and official regulatory or guideline documents. Within each domain, representative studies were selected to illustrate methodological diversity, clinical context, and current translational readiness rather than to provide exhaustive coverage of an extremely rapidly evolving field.

Artificial intelligence↗

Which senses play a role in nonhuman primate food selection? A comparison between squirrel monkeys and spider monkeys.

In order to optimize foraging efficiency and avoid toxicosis, animals must be able to detect, discriminate, and learn about the predictive signals of potential food. Primates are typically regarded as animals that rely mainly on their highly developed visual systems, and little is known about the role that the other senses may play in food selection. It was therefore the aim of the present study to assess which senses are involved in the evaluation of food by two species of New World primates: the squirrel monkey and the spider monkey. To this end, six animals per species were repeatedly presented with both familiar and novel food items, and their behavior was videotaped and analyzed. To obtain a further indication of the relative importance of visual and chemosensory cues, the animals were also presented with familiar food items that were experimentally modified in color, odor, or both color and odor. The results demonstrate that squirrel monkeys and spider monkeys use olfactory, gustatory, and tactile cues in addition to visual information to evaluate novel food, whereas they mainly inspect familiar food items visually prior to consumption. Our findings also show that in both species the use of nonvisual cues decreased rapidly with repeated presentations of novel food, suggesting a fast multimodal learning process. Further, the two species clearly differ in their relative use of nonvisual cues when evaluating novel or modified food, with spider monkeys relying more on olfactory cues than squirrel monkeys, and squirrel monkeys relying more on tactile cues compared to spider monkeys.

Animals↗

The effects of posterior parietal and posterior temporal cortical lesions on multimodal spatial and nonspatial competencies in rats.

The functional consequences of posterior parietal (PPC) and posterior temporal (Te2/3) cortical lesions on rat spatial and nonspatial multimodal learning, and memory were assessed using three behavioral paradigms. In the first, a stimulus-elicited object-place recognition task, PPC-lesioned animals were found to habituate to repeated presentation and dishabituate to changes in the visual and auditory properties of the objects but they fail to respond to changes in their location. The Te2/3-lesioned animals recognized changes in spatial location, but not changes in auditory or visual characteristics of the objects. Sham controls recognized both. In water maze-based auditory and visual place object conditional learning tasks, Te2/3 and Sham controls learned both discriminations, whereas the performance of PPC animals was significantly retarded. In the third paradigm, all three groups learned the visual discriminations. Although PPC-lesioned animals subsequently demonstrated recognition of the amodal property of duration in a visual/auditory cross-modal transfer (CMT) test, they were unable to do so on two CMT tasks involving the property of space. In all three tests, the Sham controls consistently displayed CMT and the Te2/3-lesioned animals did not. The present study extends the description of somewhat distinctive roles played by two association cortical regions (PPC and Te2/3) in the perceptual/cognitive functioning, particularly with respect to auditory stimuli and correspondences between auditory and visual events.

Acoustic Stimulation↗

A qualitative evaluation of a faith-based breast and cervical cancer screening intervention for African American women.

This article presents a formative evaluation of a CDC Racial and Ethnic Approaches to Community Health (REACH) 2010 faith-based breast and cervical cancer early detection and prevention intervention for African American women living in urban communities. Focus groups were conducted with a sample of women (N=94) recruited from each church participating in the intervention. One focus group was conducted in each of the nine participating churches following completion of the 6-month REACH 2010 intervention. Transcribed data were coded to identify relevant themes. Key findings included (a) the acceptability of receiving cancer education within the context of a faith community, (b) the importance of pastoral input, (c) the effectiveness of personal testimonies and lay health advocates, (d) the saliency of biblical scripture in reinforcing health messages, (e) the effectiveness of multimodal learning aids, and (f) the relationship between cervical cancer and social stigma. Study findings have implications for enhancing faith-based breast and cervical cancer prevention efforts in African American communities.

Black or African American↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗

Order-dependent timing of unimodal and multimodal stimulation affects prenatal auditory learning in bobwhite quail embryos.

This study examined the relationship between unimodal and multimodal sensory stimulation and their effects on prenatal auditory learning in bobwhite quail embryos. Embryos exposed to a maternal call in the 24 hr prior to hatching (unimodal condition) significantly preferred this familiar call over an unfamiliar call in postnatal testing, but failed to demonstrate this preference when the maternal call was presented concurrently with non-synchronized patterned light (multimodal condition). To further explore this interference effect, we provided one group of embryos concurrent exposure to a maternal call and patterned light for 12 hr followed by 12 hr exposure to the call alone (multimodal-->unimodal call). This group failed to prefer the familiar call during postnatal testing. In contrast, reversing the order of presentation during prenatal exposure (unimodal call-->multimodal) led a second group of subjects to significantly prefer the familiar call, suggesting that the order-dependent timing of sensory stimulation can significantly impact prenatal auditory learning. Experiment 3 examined the influence of modality versus timing of sensory stimulation on prenatal auditory learning by providing three groups of embryos with exposure to a maternal call during the 12 hr prior to hatching and by varying the duration of visual stimulation. Results indicate that 12 hr unimodal exposure to patterned light does not support prenatal auditory learning when it is followed by 12 hr exposure to multimodal stimulation (light-->multimodal), but can facilitate prenatal auditory learning when it is followed by unimodal exposure to the call alone (light-->call). Results are discussed in terms of intersensory relationships during perinatal development.

Acoustic Stimulation↗

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics↗