PubMed HealthSearch

SEARCH · PubMed Health

Results for “Predictive Learning Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged ≥65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans

Pathomics-based machine learning models for predicting METTL5 expression and prognosis in lung adenocarcinoma.

BACKGROUND: METTL5, an N6-methyladenosine (m6A) RNA methyltransferase, has been implicated in tumor progression, but its prognostic value and non-invasive prediction in lung adenocarcinoma (LUAD) remain unclear. This study aimed to develop a pathomics-based machine learning model to predict METTL5 expression from histopathological images and evaluate its prognostic significance in LUAD. METHODS: A total of 327 LUAD patients from The Cancer Genome Atlas (TCGA) with matched hematoxylin and eosin (H&E) slides, transcriptomic, and clinical data were included and randomly divided into training and validation sets (7:3). Quantitative histopathological features were extracted using PyRadiomics. Feature selection was performed via maximum relevance minimum redundancy (mRMR) and recursive feature elimination (RFE), followed by construction of a Gradient Boosting Machine (GBM) model. A pathomics score (PS) was generated to assess prognostic relevance. Survival analyses, gene set variation analysis (GSVA), tumor mutational burden (TMB), immune infiltration analysis, and in vitro functional assays were conducted. RESULTS: METTL5 overexpression was independently associated with poor overall survival [hazard ratio (HR) =1.637, P=0.007]. The model achieved good predictive performance [area under the curve (AUC) =0.847 in the training set and 0.752 in the validation set]. High PS was significantly associated with worse survival and remained an independent prognostic factor (HR =1.563, P=0.03). Elevated PS correlated with altered metabolic pathways, increased TMB, and immune microenvironment changes. METTL5 knockdown reduced proliferation, migration, invasion, and epithelial-mesenchymal transition (EMT) in A549 cells. CONCLUSIONS: The pathomics-based model accurately predicts METTL5 expression and provides prognostic stratification in LUAD, supporting its potential as a practical imaging-derived biomarker.

Methyltransferase-like 5

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure

Soffritto: a deep learning model for predicting high-resolution replication timing.

MOTIVATION: Replication timing (RT) refers to the order in which DNA loci are replicated during S phase. RT is cell-type specific and implicated in cellular processes including transcription, differentiation, and disease. RT is typically quantified genome-wide using two-fraction assays (e.g. Repli-Seq) which sort cells into early and late S phase fractions followed by DNA sequencing, yielding a ratio as the RT signal. While two-fraction RT data are widely available in multiple cell lines, it is limited in its ability to capture high-resolution RT features. To address this, high-resolution Repli-Seq, which quantifies RT across 16 fractions, was developed, but it is costly and technically challenging with very limited data generated to date. RESULTS: Here, we developed Soffritto, a deep learning model that predicts high-resolution RT data using two-fraction RT data, histone ChIP-seq data, GC content, and gene density as input. Soffritto is composed of a Long Short-Term Memory (LSTM) module and a prediction module. The LSTM module learns long- and short-range interactions between genomic bins, while the prediction module is composed of a fully connected layer that outputs a 16-fraction probability vector for each bin using the LSTM module's embeddings as input. By performing both within cell line and cross-cell line training and testing for five human and mouse cell lines, we show that Soffritto is able to capture experimental 16-fraction RT signals with high accuracy, and the predicted signals allow detection of high-resolution RT patterns. AVAILABILITY AND IMPLEMENTATION: Soffritto is available at https://github.com/ay-lab/Soffritto.

Deep Learning

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve–based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC = 0.844). The NODM cohort was stratified into high- (n = 2,362) and low-risk (n = 5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

G4STAB: a multi-input deep learning model to predict G-quadruplex thermodynamic stability based on sequence and salt concentration.

MOTIVATION: G-quadruplexes (G4s) are non-canonical nucleic acid structures formed in guanine-rich regions that modulate gene regulation and genomic stability. The thermodynamic stability of G4s directly influences their biological functions and potential as therapeutic targets. However, current quantitative frameworks for predicting G4 stability rely on predetermined structural features, limiting their effectiveness for diverse G4 topologies, and fail to account for environmental factors such as ion concentration and pH that significantly modulate G4 stability in cellular contexts. RESULTS: We present G4STAB, a multi-input deep learning neural network that accurately predicts DNA G4 melting temperatures based on sequence features, salt concentration, and pH. Trained on 2382 diverse DNA G4 sequences, our model achieves high accuracy (R 2=0.8) without relying on predetermined G4 structural features. G4STAB successfully captures established G4 stability determinants and proposes previously unobserved sequence-stability relationships. Analysis of 391 502 experimentally validated G4s reveals that cancer-like ionic environments alter G4 stability profiles, with a 13.5-fold increase in the number of structures exhibiting physiological melting temperatures (36-42°C). These findings suggest systematic genomic patterns in G4 stability responses across chromosomes and gene types. AVAILABILITY AND IMPLEMENTATION: G4STAB is available at https://github.com/donn-liew/G4STAB; G4STAB web database interface is available at https://donn-liew.github.io/g4stab-web-database/.

G-Quadruplexes

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins.

INTRODUCTION: The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. METHODS: We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. RESULTS: We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. DISCUSSION: The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

DNA-binding proteins

RP3Net: a deep learning model for predicting recombinant protein production in Escherichia coli.

MOTIVATION: Recombinant protein expression can be a limiting step in the production of protein reagents for drug discovery and other biotechnology applications. We introduce RP3Net (Recombinant Protein Production Prediction Network), an AI model of small-scale heterologous soluble protein expression in Escherichia coli. RP3Net utilizes the most recent protein and genomic foundational models. A curated dataset of internal experimental results from AstraZeneca and publicly available data from the Structural Genomics Consortium was used for training, validation and testing of RP3Net. RESULTS: RP3Net achieves an increase in area under the receiver operator curve (AUROC) of 0.15, compared to a baseline model. When experimentally validated on an independent, prospective, manually selected set of 97 constructs, RP3Net outperformed currently available models, with an AUROC of 0.83, delivering accurate predictions in 77% of the cases, and correctly identifying successfully expressing constructs in 92% of cases. AVAILABILITY AND IMPLEMENTATION: The model, along with installation and running instructions, is available under an MIT licence at https://github.com/RP3Net/RP3Net, DOI 10.5281/zenodo.17243498.

Escherichia coli

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n = 5), support vector machines (n = 4), k-nearest neighbor (n = 3), decision trees (n = 3), random forests (n = 5), neural networks (n = 2), linear discriminant analysis (n = 1), and pre-trained CNNs (n = 1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n = 12 to n = 39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans

Machine learning-based drug susceptibility prediction from Candida genomic data.

OBJECTIVES: Invasive Candida infection is an increasing clinical concern, with antifungal resistance rising across multiple species. However, rapid and accurate antifungal susceptibility testing (AFST) remains limited in routine practice. The study evaluated species distribution and antifungal susceptibility of invasive Candida isolates in China and assessed the feasibility of combining whole-genome sequencing (WGS) with machine learning to predict minimum inhibitory concentrations (MICs). METHODS: Consecutive non-repetitive isolates were collected from 20 hospitals in 13 provinces during 2022-2023. MICs of nine antifungal agents were determined by broth microdilution, and WGS was performed for species accounting for >5% of the total isolates. Genomic 11-mer features were extracted and used to train random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost) models, followed by optimization of the best-performing algorithm. RESULTS: A total of 337 isolates were obtained from blood (n = 232) and sterile body fluids (n = 105), comprising C. albicans (n = 103), C. tropicalis (n = 71), C. parapsilosis (n = 67), and C. glabrata (n = 63). Non-albicans Candida showed higher azole and echinocandin resistance, with C. tropicalis notably resistant to azoles and C. glabrata to echinocandins. Among the three models, RF demonstrated the best performance on 304 sequenced isolates. The optimized RF model was evaluated by the receiver operating characteristic (ROC) curve analysis and achieved an average area under the ROC curve (AUC) of 0.979 (95% CI: 0.974-0.984), essential agreement over 90.1%, and categorical agreement over 93.2% across species. CONCLUSIONS: These findings underscore the clinical challenge posed by non-albicans Candida resistance, and indicate that WGS-based MIC prediction may offer a highly accurate reference for earlier antifungal therapy.

Antifungal Agents

Proteomics-driven discovery of intervention windows and risk subtypes in osteoporosis: A prospective cohort study.

Given the limited feasibility of population-wide bone mineral density screening and the infrequency of long-term monitoring in healthy individuals, identifying the window for early intervention and the populations to be prioritized for screening is critical. This study aimed to identify intervention windows for osteoporosis and to determine potential high-risk subtypes within the healthy population. Based on proteomic data from 41,408 healthy adults, we conducted the DE-SWAN method to identify change peaks in plasma protein during the pre-diagnostic osteoporosis phase, and employed finite Gaussian mixture model-based clustering to delineate high-risk subtypes of osteoporosis. We identified 122 protein biomarkers significantly associated with osteoporosis risk throughout the follow-up period. Importantly, we identified two critical peaks occurring approximately 10 and 6 years before diagnosis, with the former enriched in immune-related pathways and the latter prominently involving responses to retinoic acid and glucocorticoids. Furthermore, one high-risk subtype for osteoporosis was identified in both males and females, termed the Frailty and Obesity Subtype. This subtype is characterized by a high degree of frailty and obesity, accompanied by a significantly elevated risk of both osteoporosis and fractures. Finally, we developed a predictive model comprising 10 proteins for identifying high-risk subtypes of osteoporosis, which demonstrated better performance than the traditional risk factor model (AUC: 0.743 vs. 0.680). Our findings demonstrate that proteomic profiling can reveal early molecular changes and identify high-risk subtypes years before clinical onset, providing a foundation for screening and precision prevention of osteoporosis.

Proteomics

Predicting 5-Year Mortality in Non-Small-Cell Lung Cancer Using the Korean Central Cancer Registry: Model Development and Validation Study.

BACKGROUND: Non-small-cell lung cancer (NSCLC) is one of the most common cancers and a leading cause of cancer-related mortality, making prognostic prediction clinically essential. Machine learning models are increasingly used to assess prognosis; however, developing systems that combine high discrimination with clear, clinically interpretable reasoning remains challenging. OBJECTIVE: This study aimed to develop deep learning models that predict 5-year mortality in NSCLC using data from the Korea Central Cancer Registry and quantify feature importance through permutation testing. METHODS: We identified 3144 patients diagnosed between 2014 and 2017 who had complete clinical data, pulmonary function test results, histological information, genomic data, and staging details. After preprocessing, the cohort was divided into stratified training, validation, and test sets in a 70%-15%-15% ratio. Five models were tuned using Hyperband across 10 predefined feature groups. The primary evaluation metric was the area under the receiver operating characteristic curve (AUC); additional metrics included accuracy, F1-score, precision, and recall. Groupwise permutation importance was calculated for each model, and the concordance of importance rankings was assessed using the Friedman test. RESULTS: All 5 models yielded comparable discrimination values on the test set (AUC=0.875-0.879). Model A was selected as the primary model and achieved an AUC of 0.879, an accuracy of 0.806, an F1-score of 0.824, and a Brier score of 0.142. Permuting the stage resulted in the largest decrease in AUC (0.217), followed by the pulmonary function test (0.016). Gene mutation had a modest overall impact but became more influential within the adenocarcinoma subset. The Friedman test showed no statistically significant differences in importance rankings across the models (P=.93). CONCLUSIONS: A grouped-input deep learning framework achieved discrimination comparable to a conventional Cox proportional hazards model using the same routine clinical variables for 5-year mortality prediction in NSCLC. Group-level permutation importance provided stable and reproducible insights into the clinical factors influencing risk, which may guide future model refinement and clinical decision-making.

Humans

The Progress of Gout Prediction Models Based on Multi-source Data.

INTRODUCTION: Gout, a highly serious inflammatory disease that is caused by monosodium urate crystals, is becoming an increasingly significant health concern. Artificial Intelligence and multi-omics-based research have made significant gains for the early detection and prevention of gout based on diverse approaches. This review intends to summarize current advances in forecasting gout susceptibility and gout-related symptoms, evaluate the predictive efficacy of different features, and ascertain which clinical and omics characteristics are most effective in these prediction models. METHODS: We explored the PubMed database after 2010 using keywords such as "gout", "predictive model", "risk prediction", and "machine learning", and confined our search to Englishlanguage articles. The original peer-reviewed research articles that developed gout models were selected. Research that was not original or lacked internal validation was excluded. RESULTS: Clinical features, genomics, microbiomics, radiomics, and metabolomics have been utilized to construct models related to gout and have demonstrated excellent predictive performance. Multisource data prediction models usually exhibit better effectiveness. DISCUSSION: Gout-oriented models performed excellently in predictive performance but present limitations in certain clinical and omics domains. However, if they are to affect actual patient care, they must overcome some external confirmation roadblocks and the fiscal and practical implications they will face ahead of time. CONCLUSION: This review indicates that clinical and multi-omics models of gout are significant instruments for clinical decision-making. The models constructed in these studies may be crucial for the treatment of gout and its practical benefits.

Gout

ESMpHLA: Evolutionary Scale Model-Based Deep Learning Prediction of HLA Class I Binding Peptides.

The recognition of endogenous peptides by HLA class I plays a crucial role in CD8+ T cell immune responses and human adaptive cell immune. Thus, the prediction of HLA class I-peptide binding affinities is always the core issue for the research of immune recognition and vaccine development. In this study, an evolutionary scale model (ESM) combined with parallel CNN blocks and a cross attention mechanism was used to construct a novel ESMpHLA model for predicting HLA class I binding peptides. Based on the 91,560 binding peptides of 41 HLA-A alleles, 56,731 of 50 HLA-B alleles and 2444 of 10 HLA-C alleles, the ESMpHLA model was successfully established and achieved satisfying prediction performances with the overall accuracy and AUC values of 0.874 and 0.938 for the test dataset. The results indicate that the ESMpHLA model performs well in dealing with different HLA class I 2-field alleles as well as the peptides with different lengths. Then, the generalisation ability of the ESMpHLA model was validated by an independent test dataset compiled from recent IEDB weekly benchmark datasets. The results showed that the ESMpHLA model achieved the highest ROC-AUC and PR-AUC values when compared with the latest BVMHC, CapsNet-MHC, STMHCpan and BVLSTM models. In addition, two ensemble models were also established by integrating the above 5 deep learning models using soft-voting and hard-voting strategies.

Humans

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans