PubMed HealthSearch

SEARCH · PubMed Health

Results for “PROBAST+AI”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

6 recordsLinked to original sources

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis. A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST + AI for prediction model studies. Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST + AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection. AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95 % CI 0.85-0.94; 95 % prediction interval 0.62-0.98), with sensitivity of 0.80 (95 % CI 0.77-0.83) and specificity of 0.87 (95 % CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Artificial intelligence-derived myocardial fibrosis on cardiac magnetic resonance for prognosis in cardiomyopathy: A systematic review of a sparse evidence base.

BACKGROUND: Myocardial fibrosis on cardiovascular magnetic resonance (CMR), assessed by late gadolinium enhancement (LGE) and parametric mapping, is an established predictor of adverse events in cardiomyopathy. We assessed whether artificial intelligence (AI) quantification of fibrosis adds independent prognostic value. METHODS: We searched six databases, a clinical-trials register, and a preprint server from inception to 13 June 2026. Eligible studies used AI to generate a fibrosis marker in adults with ischemic or nonischemic cardiomyopathy, with covariate-adjusted outcomes over ≥12 months. Risk of bias was assessed using PROBAST, PROBAST+AI, and QUIPS. Fewer than three comparable studies precluded meta-analysis; certainty was rated using GRADE. RESULTS: Of 448 records (381 after de-duplication), 18 full texts were reviewed and two included, one peer-reviewed and one preprint. In an ischemic-cardiomyopathy registry (Ghanbari et al.; n = 216 analytic, 26 events), AI-derived dense LGE scar predicted arrhythmic events (univariable hazard ratio [HR] 2.35, 95% CI 1.33-4.15), and AI-derived but not manual scar improved discrimination beyond guideline criteria (area under the curve 0.63 to 0.68; p = 0.02). In a nonischemic dilated-cardiomyopathy preprint (Kim et al.; n = 347, 119 events), automated extracellular volume ≥30% predicted cardiovascular death or heart-failure hospitalization (adjusted HR 2.00, 95% CI 1.32-3.03). Both were at high risk of bias, with data-derived thresholds and no external validation. CONCLUSIONS: Across only two studies, AI-derived fibrosis was independently associated with adverse cardiovascular events, but its added value over manual quantification remains unproven. Certainty was very low. The evidence base is sparse and not yet ready for clinical use.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75 161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et al., Nanda et al., Naylor et al., and Van Leeuwen et al., each showing fair discrimination. The Teede et al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et al. and van Leeuwen et al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Foundations of Artificial Intelligence in Hepatology: What a Clinician Needs to Know.

This review focuses on foundational knowledge about artificial intelligence (AI) in hepatology, exploring how AI, including machine learning and deep learning, leverages large-scale clinical data to transform the diagnosis, risk assessment, prognostication, and management of liver diseases. Online resources are described to offer fundamental AI knowledge and essential technical skills and to facilitate clinician participation across the entire AI lifecycle, ensuring they contribute not only as end users but also in development and deployment. Unlike traditional statistical approaches that prioritize interpretable parameters and clinical insight, AI focuses on maximizing predictive accuracy by identifying complex, often non-linear patterns using high-dimensional data, albeit often at the cost of model interpretability. AI is demonstrating clinical utility in liver histopathology and radiological imaging, significantly improving detection accuracy for cirrhosis, clinically significant portal hypertension, and hepatocellular carcinoma. Beyond diagnostics, AI-driven prediction models are emerging to provide personalized risk stratification for the development of liver-related complications and treatment guidance, based on complex data including longitudinal laboratory results, comorbidities, and co-medication use to monitor disease progression and therapy response. The field is rapidly expanding into novel areas such as analyzing patient-reported outcomes, genomic data, and real-time liver function monitoring, offering deeper mechanistic insights alongside clinical tools. Despite the potential to revolutionize hepatology practice and research, successful integration into routine care faces challenges. These include seamless workflow integration with existing electronic health records, establishing clear liability frameworks, and guaranteeing protection of patient privacy. Addressing these hurdles requires collaborative efforts from clinicians, researchers, and regulators to develop best practices and governance. Understanding the transformative capabilities, current applications, emerging frontiers, and essential implementation considerations is crucial for clinicians navigating the evolving AI landscape and responsibly utilizing its power for improved patient outcomes.

PROBAST+AI