PubMed HealthSearch

SEARCH · PubMed Health

Results for “Risk prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Predicting risk of ischemic stroke: A transformer model using genomic data.

BACKGROUND AND OBJECTIVE: Ischemic stroke is a leading cause of mortality and long-term disability worldwide. Genetic factors contribute to IS susceptibility, yet conventional polygenic risk score approaches are primarily based on additive effects and may not fully capture non-linear relationships or positional context and interactions among genetic variants. This study aimed to develop and evaluate a transformer-based genomic model incorporating position-wise genotype embedding for IS risk prediction. METHODS: We conducted a genome-wide association study using the UK Biobank dataset to identify IS-associated loci. Gene prioritisation was subsequently performed using tissue-specific expression quantitative trait locus-based Mendelian randomisation and colocalization analyses in whole blood and brain cortex. We then developed a transformer-based model that encoded genotype and SNP-position information using a position-wise embedding layer. Model performance was evaluated across three UK Biobank control definitions and externally assessed in the independent All of Us cohort. Performance metrics included the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score. RESULTS: Across the three UK Biobank control definitions, the proposed method achieved the numerically highest discrimination among the evaluated models, with AUROCs of 0.8109, 0.7843, and 0.7468 using MRF-negative, combined, and MRF-positive controls, respectively. In the external All of Us cohort, the proposed method achieved an AUROC of 0.7251 and retained the highest AUROC among the evaluated models. In a separate incident-stroke survival analysis, medium- and high-score groups had hazard ratios of 1.13 and 1.21, respectively, relative to the low-score group. A total of 18 IS-associated loci were identified. Among the tissue-specific MR results, EDEM2 in the brain cortex remained significant after Bonferroni correction, while DCHS2 showed a nominal association. CONCLUSIONS: The proposed transformer-based framework provides a genomic modelling approach that achieved the highest discrimination among the evaluated models in this study and retained comparative performance in an independent external cohort. In further applications, integrating this genomic framework with conventional clinical, lifestyle, and environmental risk factors may support more comprehensive and personalised IS risk assessment. Prospective, population-representative, and multi-ancestry validation will be important to establish its potential role in future prevention-oriented risk management.

Genomics and bioinformatics

Impact of personalised risk predictions on breast cancer risk perceptions: insights from the BREATHE study.

OBJECTIVE: Biennial mammography screening is well-established for women aged 50 and above, but guidelines for younger women are less clear. Risk-based screening may provide women with key information to make informed decisions about their breast cancer risk and screening. This study examines how predicted breast cancer (BC) risk shapes women's perception and confidence in risk prediction. METHODS: Women aged 35 to 59 years were recruited for a prospective multi-centre cohort and stratified into above-average, average, or below-average BC risk categories based on genetic and non-genetic risk factors. Perceived risk was assessed at enrolment and after participants were informed of their predicted risk. We used ordinal models to identify predictors of perceived risk and logistic regression to examine the relationship between changes in perceived risk and confidence in the risk prediction. RESULTS: At enrolment, 43% and 47% of 4112 participants perceived their BC risk pre-result as low or average, respectively. Thirty-five percent adjusted their perceived risk to align more closely with their predicted risk. Predictors of perceived risk post-result: perceived risk pre-result, predicted risk, ethnicity and having regular menstruation. Participants who underestimated their BC risk were nearly eight times more likely to have low confidence in the accuracy of their predicted risk (OR for underestimation vs. accurate perception: 7.94 [95% CI 5.60-11.28]). Predictors of perceived risk post-result: perceived risk pre-result, predicted risk, ethnicity and having regular menstruation. Confidence in risk prediction was lowest when women's perceived risk pre-result was lower than their predicted risk (OR-2 vs 0 [95%CI] 5.06 [3.67 to 6.97]). CONCLUSION: Many women underestimated their BC risk, and their initial perceptions were influenced by the knowledge of their predicted risk. Women who underestimated their risk had less confidence in their predicted risk scores.

Humans

Large-scale multi-omics enhance risk prediction for type 2 diabetes.

BACKGROUND: Polygenic risk scores (PRS), metabolomics, and proteomics have each shown promise in improving type 2 diabetes risk prediction, but their combined utility beyond established clinical models remains unclear. We aimed to evaluate whether integrating multi-omics biomarkers enhances 10-year type 2 diabetes risk prediction beyond single-omics extensions and the clinical Cambridge Diabetes Risk Score (CDRS), which includes HbA1c measurements. METHODS: We analysed data from 42,840 UK Biobank participants without diagnosed diabetes at baseline. The study population was split into a derivation set (Phase 1 metabolomics release, N&#x2009;=&#x2009;23,108) to fit models and an independent validation set (Phase 2 release, N&#x2009;=&#x2009;19,732) to evaluate performance. Data for a PRS for type 2 diabetes, 11 metabolites, and 15 proteins were added to the CDRS to develop multi-omics prediction models. Model performance was evaluated using Harrell's C-index and the net reclassification index (NRI). RESULTS: During 10 years of follow-up, 1090 participants developed incident type 2 diabetes. Among individual omics layers, proteomics contributed the greatest improvement in predictive performance, increasing the C-index from 0.862 (clinical CDRS) to 0.884 (&#x394;C-index; + 0.022; P&#x2009;<&#x2009;0.001), with a continuous NRI of 42.0%. The full multi-omics model further significantly increased the C-index compared to a model combining the clinical CDRS with proteomics data (C-index, 0.891; &#x394;C-index; + 0.007; P&#x2009;<&#x2009;0.001). CONCLUSION: Integrating proteomics, metabolomics, and a diabetes-PRS into a clinical model substantially improves type 2 diabetes risk prediction beyond single-omics extensions. Several of the selected proteins and metabolites are on cardiovascular disease pathways, highlighting the link between diabetes and cardiovascular risk. However, the C-index difference between the proteomics extended and full multi-omics extended models is small, and the clinical models extended with proteomics data would be easier to translate into routine care because it needs only the measurement of 15 proteins. External validation and cost-effectiveness analyses are needed to support clinical adoption.

Humans

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75&#x2009;161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et&#xa0;al., Nanda et&#xa0;al., Naylor et&#xa0;al., and Van Leeuwen et&#xa0;al., each showing fair discrimination. The Teede et&#xa0;al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et&#xa0;al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et&#xa0;al. and van Leeuwen et&#xa0;al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article

Cross-Device Adaptation of Mirai for Mammography-Based Breast Cancer Risk Prediction.

Fine-tuning can adapt pretrained medical imaging models to new clinical datasets, but device-specific domain shifts may limit generalizability. We evaluated Mirai, a mammography-based deep learning model for breast cancer risk prediction, in a large screening cohort containing Hologic and General Electric (GE) full-field digital mammography systems, including GE Premium View (GE PV) and Tissue Equalization (GE TE) post-processing software. Native Mirai showed lower performance on TE images than on Hologic or PV images. Fine-tuning on TE images improved TE performance, particularly for short-term risk prediction, but substantially reduced performance on Hologic images, consistent with catastrophic forgetting. To mitigate this effect, we developed a device-invariant model using interleaved multi-device sampling and conditional adversarial training. This approach largely restored Hologic performance while maintaining improved TE performance, providing better robustness across heterogeneous imaging platforms. Comparison of cumulative and annual risk AUCs over a five-year time horizon further showed that performance gains were driven mainly by short- and intermediate-term predictions. These findings highlight both the value and dangers of device-specific fine-tuning and support balanced domain-adaptation strategies for deploying mammography-based risk models across diverse clinical imaging environments.

Journal Article

Risk prediction models for blood transfusion in patients undergoing total hip and knee arthroplasty: a systematic review and meta-analysis.

OBJECTIVE: To systematically review and evaluate published risk prediction models for perioperative blood transfusion in patients undergoing total hip or knee arthroplasty (THA/TKA). METHODS: We systematically searched PubMed, Web of Science, the Cochrane Library, and Embase from inception to May 31, 2025. Two researchers independently screened the literature, extracted data, and assessed the risk of bias and applicability using the Prediction model Risk Of Bias Assessment Tool (PROBAST). The area under the receiver operating characteristic curve (AUC) values were pooled via a meta-analysis using Stata 18.0. RESULTS: d Fourteen studies containing 36 prediction models were included. The incidence of blood transfusion among THA/TKA patients ranged from 3.2% to 30.8%. Preoperative hemoglobin (Hb) level, tranexamic acid (TXA) use, operative duration, intraoperative blood loss, and age were the most frequently incorporated predictors. Model sensitivity ranged from 58% to 94.5%, and specificity ranged from 71.3% to 94%. Meta-analysis showed that the pooled AUC value of the 13 validated models was 0.87 (95% CI: 0.85-0.90), suggesting good discriminatory performance. All models were rated as having a high risk of bias. The applicability of four studies was rated as unclear. CONCLUSION: Although the included studies demonstrated promising discriminative ability of prediction models for blood transfusion in THA/TKA, all were assessed as having a high risk of bias using the PROBAST tool. Therefore, future research should prioritize the development of models with larger sample sizes, rigorous study designs, and multicenter external validation.

Humans

Large-Scale Plasma Proteomics Identifies Early Molecular Deviations and Improves Risk Prediction for Heart Failure Among Individuals With Obesity.

AIMS: Heart failure (HF) is a major global public health challenge, with obesity being one of its key risk factors. Although several HF risk prediction models have been developed in the general population, few are specifically tailored to individuals with obesity. This underscores the urgent need for precise biomarkers to improve individual risk stratification and enable personalized prevention strategies. We aimed to develop and validate a plasma proteomics-based protein risk score (PRS) to predict incident HF among individuals with obesity. MATERIALS AND METHODS: We analysed 9831 participants with obesity (BMI &#x2265;&#x2009;30&#x2009;kg/m2) from the UK Biobank with baseline measurements of 2911 circulating proteins and up to 16&#x2009;years of follow-up. Multivariable Cox regression identified proteins associated with incident HF after comprehensive covariate adjustment. A PRS was constructed using LASSO regression and evaluated in a held-out test set. Protein trajectories before HF onset were reconstructed using LOESS modelling. To enhance clinical feasibility, a minimal protein panel was identified using LightGBM with forward feature selection. RESULTS: A total of 727 participants developed HF during follow-up. Multivariable cox analyses identified 578 proteins significantly associated with HF. LASSO regression further selected 81 proteins to build the PRS, which showed a strong association with HF risk in both training (HR 3.57; 95% CI 3.19-4.00) and test cohorts (HR 2.45; 95% CI 2.20-2.74). Adding the PRS improved prediction beyond age and sex (&#x394;C&#x2009;=&#x2009;0.091) and beyond the Pooled Cohort Equations to Prevent Heart Failure (PCP-HF) model (&#x394;C&#x2009;=&#x2009;0.052), with consistent gains in NRI and IDI. Proteomic deviations were detectable up to 16&#x2009;years before diagnosis. A four-protein panel (GDF15, NT-proBNP, TNFRSF10B, CTHRC1) achieved robust discrimination (AUC 0.789), outperforming NT-proBNP alone (AUC 0.695) and complementing the PCP-HF model (combined AUC 0.803). DISCUSSION: Large-scale plasma proteomics substantially improves HF risk prediction in individuals with obesity and reveals long-standing molecular alterations preceding clinical onset. A simplified four-protein panel maintains robust predictive accuracy and provides a practical approach for the early detection and targeted prevention of obesity-related HF.

Humans

Risk prediction in patients with heart failure with preserved ejection fraction: the LIFE-Preserved model.

BACKGROUND AND AIMS: Heart failure (HF) with preserved ejection fraction (HFpEF) constitutes a heterogeneous disease with varying prognosis. Given the rising incidence of HFpEF, accurate risk prediction for these patients is needed to identify high-risk individuals, who may benefit the most from preventive treatments. The LIFE-Preserved model was developed and validated for the prediction of individual short-term and lifetime risk for HF hospitalization or cardiovascular (CV) death in patients with HFpEF. METHODS: LIFE-Preserved was derived in 20 332 patients aged 40-90 years with a left ventricular ejection fraction &#x2265; 50% from the Swedish HF Registry. Cause- and sex-specific Cox models were derived to predict the risk of HF hospitalization or CV death using 14 routinely available predictors. Use of age as the timescale allowed for predictions beyond the maximum follow-up duration in the derivation data, adjusted for competing risks. External validation was performed in two trials (EMPEROR-Preserved and TOPCAT-Americas) and three registries (NHS England Secure Data Environment, Veterans Affairs, and HF-Particles). Model performance was assessed by discrimination and calibration. RESULTS: During a median follow-up of 1.8 years (interquartile range .6-4.2, maximum 19 years), 9341 first HF hospitalizations or CV deaths (46%) were observed in Swedish HF Registry. External validation included data from 28 062 patients with HFpEF [9930 (35%) first HF hospitalizations or CV deaths]. Pooled C-statistics were .714 (95% confidence interval .652-.775) in trials and .658 (95% confidence interval .599-.717 in registries, with adequate calibration in all external validation sources. Performance was similar in men and women. An interactive calculator of the LIFE-Preserved model has been made available here. CONCLUSIONS: The LIFE-Preserved model enables prediction of short-term and lifetime risk of HF hospitalization or CV death in patients with HFpEF. The model could serve as a tool to identify high-risk HFpEF patients, guiding clinical management and shared decision-making.

Humans

Assessing comorbidities and predicting risk: A primer for APRNs.

Today's clinical environments are rife with tools designed to comprehensively account for medical complexity and comorbidities while predicting risk for a host of adverse health-related outcomes. Therefore, it is imperative that advanced practice registered nurses (APRNs) understand the structure and function of these tools, their similarities and differences, their limitations, and strategies for appropriate incorporation into practice. This article offers a practical overview for APRNs, emphasizing clinical implications and guidance for aligning assessment tools with the clinical population of interest to improve care delivery, quality, and patient outcomes.

Humans

MyGeneRisk Colon: A Web-Based Tool for Personalized Colorectal Cancer Risk Prediction Based on Genetics and Lifestyle.

Colorectal cancer (CRC) is a leading cause of cancer-related death, with incidence rising substantially among individuals under 50 years of age. Polygenic risk scores (PRS) hold promise for identifying high-risk individuals; when combined with lifestyle factors, they substantially improve prediction accuracy compared with models based on lifestyle factors alone. However, few clinical tools currently exist that facilitate this integrated, PRS-enhanced risk assessment. To bridge this gap, we developed MyGeneRisk Colo n, a publicly accessible web portal that delivers individualized CRC risk prediction by incorporating genetic, demographic, family history, and lifestyle factors. This paper details the development of the underlying risk prediction model, the portal's architecture and data security, our reporting framework, and engagement with a community advisory panel. Designed as a user-friendly platform, MyGeneRisk Colon aims to effectively communicate personalized CRC risk profiles and educate users and healthcare providers about prevention strategies.

Journal Article

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Cardiac Troponins and Cardiovascular Disease Risk Prediction: An Individual-Participant-Data Meta-Analysis.

BACKGROUND: The extent to which high-sensitivity cardiac troponin can predict cardiovascular disease (CVD) is uncertain. OBJECTIVES: We aimed to quantify the potential advantage of adding information on cardiac troponins to conventional risk factors in the prevention of CVD. METHODS: We meta-analyzed individual-participant data from 15 cohorts, comprising 62,150 participants without prior CVD. We calculated HRs, measures of risk discrimination, and reclassification after adding cardiac troponin T (cTnT) or I (cTnI) to conventional risk factors. The primary outcome was first-onset CVD (ie, coronary heart disease or stroke). We then modeled the implications of initiating statin therapy using incidence rates from 2.1 million individuals from the United Kingdom. RESULTS: Among participants with cTnT or cTnI measurements, 8,133 and 3,749 incident CVD events occurred during a median follow-up of 11.8 and 9.8 years, respectively. HRs for CVD per 1-SD higher concentration were 1.31 (95%&#xa0;CI: 1.25-1.37) for cTnT and 1.26 (95%&#xa0;CI: 1.19-1.33) for cTnI. Addition of cTnT or cTnI to conventional risk factors was associated with C-index increases of 0.015 (95%&#xa0;CI: 0.012-0.018) and 0.012 (95%&#xa0;CI: 0.009-0.015) and continuous net reclassification improvements of 6% and 5% in cases and 22% and 17% in noncases. One additional CVD event would be prevented for every 408 and 473 individuals screened based on statin therapy in those whose CVD risk is reclassified from intermediate to high risk after cTnT or cTnI measurement, respectively. CONCLUSIONS: Measurement of cardiac troponin results in a modest improvement in the prediction of first-onset CVD that may translate into population health benefits if used at scale.

Humans

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-&#x3ba;B, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-&#x3ba;B, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4

Osteoporosis genetic risk prediction using bone mineral density polygenic scores in Japanese: TMM CommCohort study.

Osteoporosis and fractures are major health concerns. We developed and validated a polygenic score (PGS) for quantitative ultrasound (QUS)-defined osteoporosis risk in Japanese individuals using heel QUS-derived T-scores. Genome-wide association study summary statistics from up to 10,794 participants in the Tohoku Medical Megabank Community-Based Cohort identified genome-wide significant loci, including MBL2, TMEM135, and WNT16. PGS models were constructed and evaluated using independent datasets for model selection (n&#x2009;=&#x2009;1419) and validation (n&#x2009;=&#x2009;8711). Adding the PGS to age and sex yielded only modest improvements in discrimination, whereas PGS quintiles supported genetic risk stratification. Compared with the intermediate group, individuals in the lowest PGS quintile had higher odds of the outcome (1.22, 95% confidence interval [CI]: 1.07-1.40), whereas those in the highest quintile had lower odds (0.85, 95% CI: 0.74-0.98). During prospective follow-up (mean 3.5 years), a similar gradient was observed, with higher incidence rate ratios in the lowest quintile (1.42, 95% CI: 1.17-1.73) and lower incidence rate ratios in the highest quintile (0.70, 95% CI: 0.54-0.89). No statistically significant interaction between age and PGS was observed, and age-T-score regression analyses showed no differences in age-related T-score decline across genetic risk groups. However, analyses in young adults (20-44 years) and extrapolation to age 20 suggested lower bone status around peak bone mass in individuals at high genetic risk for QUS-defined osteoporosis. These findings suggest that a Japanese-specific PGS may help identify individuals at elevated genetic risk earlier in adulthood.

Journal Article