PubMed HealthSearch

SEARCH · PubMed Health

Results for “Explainable AI”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology

Decoding cancer with artificial intelligence: Transforming research, diagnosis, and therapy with future insights.

Cancer remains one of the leading global health burdens, with increasing complexity in genomic, imaging, and clinical datasets presenting significant challenges for effective management. Artificial intelligence (AI) has emerged as a powerful tool to address these challenges by enabling pattern recognition, knowledge integration, and data-driven decision-making. This review highlights recent advances in the application of AI across cancer research, diagnosis, and therapy. In research, AI accelerates drug discovery and repurposing, enhances genomic data interpretation, and facilitates biomarker identification through multi-omics integration. In diagnosis, AI has demonstrated high technical performance in radiology for lesion detection and image segmentation, in pathology for tumour grading and molecular prediction, and in liquid biopsy for non-invasive biomarker analysis. In therapy, AI supports precision medicine by predicting treatment responses, monitoring disease progression, and optimizing clinical trial design. Despite these advances, barriers such as data heterogeneity, algorithmic bias, interpretability, and regulatory challenges remain. Future directions, including explainable AI, federated learning, multimodal modelling, and digital twins, hold promise for translating AI-driven innovations into routine oncology practice. Significance Statement This review provides a timely synthesis of recent (2020-2025) advances in artificial intelligence across cancer research, diagnosis, and therapy, highlighting applications in drug discovery, genomics, multi-omics biomarker identification, and clinical decision-making. By integrating technological progress with translational and clinical relevance, this work serves as a valuable resource for bridging AI innovation with precision oncology practice. As a narrative review, the literature was identified through targeted PubMed, Scopus, and Google Scholar searches, combining terms for artificial intelligence, machine learning, and deep learning with cancer-related keywords, with priority given to peer-reviewed studies published between 2020 and 2025, seminal earlier works, and official regulatory or guideline documents. Within each domain, representative studies were selected to illustrate methodological diversity, clinical context, and current translational readiness rather than to provide exhaustive coverage of an extremely rapidly evolving field.

Artificial intelligence

Artificial Intelligence and Machine Learning Applications in Fibromuscular Dysplasia: Transforming Diagnosis, Risk Stratification, and Clinical Decision-Making.

Fibromuscular dysplasia (FMD) is a non-atherosclerotic vascular disorder with heterogeneous presentations, making diagnosis and management highly dependent on imaging and clinical expertise. This narrative review examines how artificial intelligence (AI) and machine learning (ML) are transforming FMD care. AI-enhanced imaging, particularly convolutional neural network-based analysis, improves detection of the characteristic "string-of-beads" pattern on CT angiography, magnetic resonance angiography, and ultrasound, although FMD-specific validation remains limited. ML models facilitate risk stratification, prediction of disease progression, and early identification of complications such as aneurysms and stroke by integrating clinical, imaging, and genomic data. AI-driven clinical decision support systems further enable personalized treatment selection through pharmacogenomic insights and robot-assisted interventions. Despite promising real-world applications, challenges persist, including limited large-scale datasets, workflow integration, regulatory barriers, and algorithmic bias affecting underrepresented populations. Future advances in explainable AI, federated learning, and digital health integration may enable a shift toward predictive, patient-centered FMD management.

Humans

AI-driven multi-omics modeling of myalgic encephalomyelitis/chronic fatigue syndrome.

Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a chronic illness with a multifactorial etiology and heterogeneous symptomatology, posing major challenges for diagnosis and treatment. Here we present BioMapAI, a supervised deep neural network trained on a 4-year, longitudinal, multi-omics dataset from 249 participants, which integrates gut metagenomics, plasma metabolomics, immune cell profiling, blood laboratory data and detailed clinical symptoms. By simultaneously modeling these diverse data types to predict clinical severity, BioMapAI identifies disease- and symptom-specific biomarkers and classifies ME/CFS in both held-out and independent external cohorts. Using an explainable AI approach, we construct a unique connectivity map spanning the microbiome, immune system and plasma metabolome in health and ME/CFS adjusted for age, gender and additional clinical factors. This map uncovers altered associations between microbial metabolism (for example, short-chain fatty acids, branched-chain amino acids, tryptophan, benzoate), plasma lipids and bile acids, and heightened inflammatory responses in mucosal and inflammatory T cell subsets (MAIT, &#x3b3;&#x3b4;T) secreting IFN-&#x3b3; and GzA. Overall, BioMapAI provides unprecedented systems-level insights into ME/CFS, refining existing hypotheses and hypothesizing unique mechanisms-specifically, how multi-omics dynamics are associated to the disease's heterogeneous symptoms.

Humans

Esketamine multi-omic biomarker evaluation in major depressive disorder (EMBER-MDD): concept, objectives and methodologies of a non-clinical investigator-initiated study.

Treatment resistance (TR) in major depressive disorder (MDD) affects a substantial minority of patients and is hard to recognize early, delaying intensified care. The Esketamine multi-omic biomarker evaluation in MDD (EMBER-MDD) is a non-interventional, investigator-initiated, in-vitro study within the EU Psych-STRATA programme, analyzing biospecimens collected in the randomized INTENSIFY study and the mirror OBS-TR cohort after participants complete treatment. EMBER-MDD aims to discover individual-omic and integrated multi-omic (hypothesis-free) biomarkers and signatures associated with TR risk, and molecular correlates of clinical response to esketamine nasal spray versus treatment as usual (TAU). Biomaterials will derive from approximately 420 adults with MDD (estimated n&#x2009;=&#x2009;210 esketamine; n&#x2009;=&#x2009;210 TAU) and include whole blood, RNA-stabilized whole blood, plasma and serum, sampled at baseline and, when feasible, during and after treatment (up to ~&#x2009;5,040 aliquots stored at -&#x2009;80&#xa0;&#xb0;C). Genomics will use baseline DNA genotyping on Illumina Infinium GSA v3.0+MD arrays; epigenomics will profile genome-wide DNA methylation across time points using MethylationEPIC v2.0; transcriptomics will employ mRNA-seq (NovaSeq X/ X Plus); and proteomics/ metabolomics will be generated using high-throughput Olink and/ or Biocrates platforms. Each layer will undergo state-of-the-art preprocessing and analyses (e.g., GWAS/ PRS, EWAS, differential expression, WGCNA, pathway and network analyses), followed by integrative strategies including QTL mapping (meQTL/ eQTL/ pQTL/ mQTL) and intermediate-fusion machine learning with nested cross-validation, explainable AI (SHAP/ LIME) and treatment-effect modelling. All outputs are research-only and will not support individual efficacy, tolerability, or clinical decision-making. The study will deliver robust biosignatures and mechanistic hypotheses to guide future validation and inform stratified, molecularly guided intervention strategies in subsequent prospective trials. Trial registration number: 2023-506617-21-00 and 2025-178-f-S.

Humans

Adversarial attack of sequence-free enhancer prediction identifies chromatin architecture.

MOTIVATION: The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements "enhance" specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. RESULTS: We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks [adversarial particle swarm optimization (APSO)] to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference. AVAILABILITY AND IMPLEMENTATION: All software and code for data downloading, processing, enhancer inference, eXplainable AI (XAI), and complete figure generation are publicly available on GitHub at https://github.com/EpiGenomicsCode/ChromEnhancer and Zenodo at https://doi.org/10.5281/zenodo.15652797.

Enhancer Elements, Genetic

Graph neural network-based risk stratification of prostate cancer using gene expression and SHAP interpretability.

Accurate risk stratification is essential for guiding treatment decisions and preventing over treatment of prostate cancer, which remains one of the most prevalent cancers among adult men. While the Gleason score, obtained from prostate biopsies, is routinely used to assess tumor aggressiveness, the biopsy procedure carries risks such as pain, infection, and, in some cases, serious complications such as sepsis. In this study, we proposed an artificial intelligence-based framework that integrates mRNA expression profiles with functional interaction networks to classify prostate cancer patients into low-, medium-, and high-risk groups defined by Gleason scores. The pipeline comprised five steps: (1) data collection from The Cancer Genome Atlas (TCGA), (2) preprocessing of gene expression data, (3) two-stage feature selection to identify informative biomarkers, (4) risk classification using a dual-branch graph neural network (GNN) that combines gene-gene interaction graphs with sample-level expression features, and (5) model interpretation using SHAP to quantify feature contributions. Differentially expressed genes were identified in the High (ASPN, GMNN, PEBP4, C2, KNCK17), Medium (C2, IGSF1, ASPN, CDKN3, AMH), and Low (TNMD, VWA5B2, ST6GALNAC5, CYP3A5, PHGR1) risk groups, underscoring the molecular heterogeneity of disease progression. On an independent held-out test set, the model achieved AUCs of 0.86, 0.88, and 0.95 for the low-, medium-, and high-risk groups, respectively, with an overall accuracy of 80%. These results suggest that combining GNN-based modeling with explainable AI can capture both global and local molecular patterns relevant to tumor aggressiveness. However, as the model was developed and evaluated solely on the TCGA cohort, the findings should be regarded as exploratory, and external validation will be required to establish generalizability. Within these limitations, the proposed framework highlights the potential of molecular profiling and graph-based deep learning to support more precise, potentially less invasive, risk assessment and individualized treatment planning in prostate cancer.

Prostatic Neoplasms

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

Single night studies in obstructive sleep apnea.

The role of single night studies and the determinants of effective nasal continuous positive airway (CPAP) pressures were determined in 412 consecutive patients between 1984 and 1989. Patients chosen for analysis had an apnea index (AI) of greater than or equal to 20 hr-1 prior to CPAP. The AI was 67 +/- 30 hr-1, the body mass index (BMI) was 36 +/- 9 kg/m2, the age was 51 +/- 13 yr and the lowest oxygen saturation was 72 +/- 14%. Effective CPAP (9 +/- 3 cm H2O) was documented in 320 patients on single night studies and resulted in a 99% reduction in the frequency of obstructive events and improvement in the lowest O2 saturation to 94 +/- 5%. Only 18% of the variability in effective CPAP could be explained by AI and BMI. Single night studies are sufficient to establish effective CPAP in 78% of patients and offer considerable conservation of resources compared to routine multiple night studies. Effective CPAP pressures are variable and must be determined by incremental CPAP trials.

Arousal

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals

A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.

BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.

Humans

Uncovering heterogeneous effects via localized feature selection.

Identifying features that interact to trigger disease, while accounting for heterogeneity across diverse populations, is essential for the development of precision and targeted medicine. Despite the availability of vast and complex health-related datasets, most existing works focus on identifying disease-associated features at the population level or within a few subpopulations, often overlooking individual-level heterogeneity within these groups. To address this limitation, we propose a framework that utilizes localized test statistics to identify disease-associated features tailored to individual profiles. Our method leverages the recently developed knockoffs methodology to control the noise level of the selection set so that the results are replicable. Moreover, it allows for the discovery of hidden heterogeneous effects within the data, as demonstrated in an application to single-cell RNA sequencing data for Alzheimer's disease. By aggregating localized feature selection results, our framework also enables powerful population-level feature selection. Our framework provides a powerful tool for exploratory studies of precision medicine, offering the potential to generate novel hypotheses for confirmatory biological experiments.

Alzheimer Disease

A pan-cancer multi-omic SuperLearner for regulated cell death survival topologies.

INTRODUCTION: Regulated cell death (RCD) pathways influence tumor progression and immune modulation. We previously constructed a signature database mapping 25 RCD forms across seven multi-omic layers and 33 tumor types (CancerRCDShiny). Despite their ability to identify risk populations, translating these signatures into personalized clinical workflows requires a shift from cohort stratification to individualized risk mapping by modeling patient risk (survival topologies) to capture the non-linear dynamics of RCD signatures. METHODS: We engineered a pan-cancer multi-omic SuperLearner pipeline across 33 cancer types. Phase I performed zero-leakage harmonization and groupwise imputation to prevent cross-cohort amalgamation. Phase II deployed Elastic Net-regularized Cox regression as a CANARY diagnostic to map proportional hazards failures. Strata with a 35% missingness barrier entered Phase III, deploying a Quadripartite ensemble: Random Survival Forests, XGBoost, Survival-Boruta, and Multi-Task Logistic Regression, fused within an Elastic Net Multi-View Meta-Learner (MVL), with post-hoc TreeSHAP and LIME interpretability. RESULTS: The CANARY diagnostic demonstrated the structural invalidity of pan-cancer geometric proportional hazards. Across 96 admissible strata, Phase III executed algorithmic displacement: continuous multi-omic topologies suppressed static genomic mutations and copy number variations (85.7% vs. 0.0% apex retention). The MVL stabilized predictions against extreme variance; LIME surrogate validations (R 2&#x202f;<&#x202f;0.10) confirmed the systematic failure of linear interpretative proxies. N-dimensional TreeSHAP interaction mapping exposed synergistic and antagonistic rescue trajectories defining individualized Survival Topologies, which were invisible to additive models. The architecture was deployed as CancerRCDPredictor, a digital molecular tumor board with integrated LLM capabilities. The MVL SuperLearner achieved a median C-index of 0.749 (IQR: 0.722-0.836) across 96 modelable strata, with 95% bootstrap confidence intervals confirming precision (median width: 0.052) and permutation significance in 93.8% of strata (p&#x202f;<&#x202f;0.001). External CPTAC validation across ten cancer types demonstrated significant cross-cohort generalizability in clear cell renal carcinoma (KIRC; C-index 0.675, p&#x202f;=&#x202f;0.017) and modest performance across the remaining adequately powered cancers (median 0.582), underscoring the need for larger multi-institutional validation cohorts. CONCLUSION: This pan-cancer multi-omic SuperLearner bypasses linear topological failures, advancing beyond generalized stratification to establish a deterministically mapped architecture for predicting RCD-related survival topologies. Through the CancerRCDPredictor interface, multi-omic insights translate into individualized survival topology exploration, providing a foundation for future precision oncology validation.

SuperLearner

AI-integrated digital breeding for crop improvement.

Crop breeding increasingly depends on the effective integration and interpretation of large, heterogeneous datasets spanning genomic, phenotypic, multi-omics, and environmental layers. Conventional breeding approaches are often insufficient to capture the complex relationships among these data or to support timely selection decisions. Digital breeding can help address this limitation by complementing field experimentation, mixed models, and genomic prediction with the integration of biological data and computational prediction throughout the breeding process. In particular, the rapid advancement of artificial intelligence (AI) has improved the analysis of high-dimensional datasets and broadened its application to trait prediction, selection, and breeding design. Here, we review recent developments in AI-enabled digital breeding, encompassing genomic, phenomic, and multi-omics data generation and analysis, predictive modeling, explainable and generative AI, and data-driven breeding decision support. We further discuss emerging AI applications, their current contributions to crop research and breeding, and the major considerations affecting their reliable and practical implementation. Collectively, this review provides a structured understanding of the roles of AI across the digital breeding process and offers guidance for future methodological development and practical application in crop improvement.

artificial intelligence

AI-driven diagnostic and prognostic models for metabolic dysfunction-associated steatotic liver disease: insights from clinical, imaging, and multi-omics studies-a scoping review.

Metabolic dysfunction-associated steatotic liver disease (MASLD), formerly known as non-alcoholic fatty liver disease (NAFLD), is the most common chronic liver disease around the world, affecting 33.6% of the adult population (95% CI: 28.1%-39.5%; I 2&#x2009;=&#x2009;99.9%), or roughly one in three. The extent of the liver damage is variable, from simple steatosis to metabolic dysfunction-associated steatohepatitis (MASH, formerly NASH), cirrhosis and hepatocellular carcinoma (HCC). Early diagnosis is essential to prevent serious liver damage. Traditional diagnostic techniques such as liver biopsy, imaging, and biomarker testing are all invasive, costly, reduced sensitive to early-stage disease, and they also have variability among observers. Modern diagnostic and prognostic approaches based on the principles of Artificial Intelligence (AI) and specifically on machine learning (ML) and deep learning (DL) have enabled multimodal approaches integrating clinical, imaging and molecular data. This scoping review conducted per PRISMA-ScR guidelines, synthesizes findings from 73 studies (search window 2020-2026) across three dimensions: clinical data driven models, imaging-based classifiers (ultrasound, CT and MRI), and multi-omics (genomics, transcriptomics and proteomics) techniques. Moreover, emergence of models such as U-Net and LiverNet 2.x, classification models like DeepLiverNet and BiLSTM models, as well as transformer frameworks and the identification of biomarkers models are also described. This study also investigates challenges such as data heterogeneity, data interpretability, fairness and real-world clinical application. Finally, important areas of research opportunities and future directions are highlighted to present a developing clinically applicable, explainable and ethical AI solutions to manage MASLD.

MASLD

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning