PubMed HealthSearch

SEARCH · PubMed Health

Results for “Boosting Machine Learning Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

18 recordsLinked to original sources

Machine Learning and Metabolomics to Characterize Warburg-Like Metabolic Subtypes in Human Retinal Endothelial Cells Exposed to Risk Factors Associated With Proliferative Diabetic Retinopathy.

PURPOSE: High glucose (HG), hypoxia (Hyp), and their combination are major risk factors for proliferative diabetic retinopathy (PDR). Although these conditions induce features of the Warburg-like metabolic reprogramming in human retinal endothelial cells (HRECs), it remains unclear whether they produce distinct metabolic and angiogenic subtypes. This study aimed to characterize the Warburg-like-associated metabolic heterogeneity induced by these PDR-related risk factors and evaluate the ability of supervised machine-learning models to distinguish these subtypes. METHODS: HRECs were cultured under normoglycemic, HG, Hyp (2% O2), and combined HG-Hyp conditions. Untargeted LC-MS/MS metabolomics quantified metabolites spanning carbohydrates, amino acids, nucleotides, and lipids. Principal component analysis (PCA) assessed overall metabolic variation, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis identified metabolic pathways associated with angiogenesis. In vitro angiogenesis assays measured endothelial tube formation and branching. Nine supervised classifiers (decision tree, logistic regression, naïve Bayes, random forest, K-Nearest Neighbors, neural network, gradient boosting, AdaBoost, and Support Vector Machine) were trained on the highest-ranked metabolites selected by the Information Gain Ratio feature-ranking approach. Model performance was evaluated using 10-fold cross-validation, leave-one-out cross-validation (LOOCV), permutation testing, and a classifier stability analysis under biologically meaningful distributional shift using an independent chemically induced hypoxia model (CoCl2). RESULTS: PCA revealed partial separation of metabolic profiles across conditions, indicating different Warburg-like metabolic subtypes. The combined HG-Hyp condition exhibited enhanced angiogenic potential relative to either HG or Hyp alone. KEGG pathway enrichment analysis identified fatty acid biosynthesis and elongation among the most significantly enriched pathways in HRECs under combined HG-Hyp conditions, alongside amino sugar and nucleotide sugar metabolism, glycerophospholipid metabolism, the pentose phosphate pathway, and glycolysis/gluconeogenesis. Supervised machine-learning classifiers distinguished these metabolic subtypes, with AdaBoost and gradient Boosting showing the most balanced, reproducible performance across 10-fold cross-validation, LOOCV, and permutation testing, and remaining the most reliable classifiers under domain-shift testing (area under the curve = 0.88, P = 0.0061). CONCLUSIONS: In this exploratory analysis, HG, Hyp, and their combination drive metabolically and functionally distinct subtypes of Warburg-like metabolic reprogramming in HRECs, with HG-Hyp in combination producing a highly angiogenic phenotype. Boosting-based ensemble classifiers provide a promising framework for detecting these subtypes even under domain-shift conditions, warranting validation in larger independent datasets. TRANSLATIONAL RELEVANCE: Integrating metabolomics with machine-learning classification offers a strategy to identify Warburg-like metabolic subtypes in retinal endothelial cells, providing insights into angiogenic mechanisms and guiding the development of targeted diagnostics or therapeutics for PDR.

Humans

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral

Machine learning-based drug susceptibility prediction from Candida genomic data.

OBJECTIVES: Invasive Candida infection is an increasing clinical concern, with antifungal resistance rising across multiple species. However, rapid and accurate antifungal susceptibility testing (AFST) remains limited in routine practice. The study evaluated species distribution and antifungal susceptibility of invasive Candida isolates in China and assessed the feasibility of combining whole-genome sequencing (WGS) with machine learning to predict minimum inhibitory concentrations (MICs). METHODS: Consecutive non-repetitive isolates were collected from 20 hospitals in 13 provinces during 2022-2023. MICs of nine antifungal agents were determined by broth microdilution, and WGS was performed for species accounting for >5% of the total isolates. Genomic 11-mer features were extracted and used to train random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost) models, followed by optimization of the best-performing algorithm. RESULTS: A total of 337 isolates were obtained from blood (n = 232) and sterile body fluids (n = 105), comprising C. albicans (n = 103), C. tropicalis (n = 71), C. parapsilosis (n = 67), and C. glabrata (n = 63). Non-albicans Candida showed higher azole and echinocandin resistance, with C. tropicalis notably resistant to azoles and C. glabrata to echinocandins. Among the three models, RF demonstrated the best performance on 304 sequenced isolates. The optimized RF model was evaluated by the receiver operating characteristic (ROC) curve analysis and achieved an average area under the ROC curve (AUC) of 0.979 (95% CI: 0.974-0.984), essential agreement over 90.1%, and categorical agreement over 93.2% across species. CONCLUSIONS: These findings underscore the clinical challenge posed by non-albicans Candida resistance, and indicate that WGS-based MIC prediction may offer a highly accurate reference for earlier antifungal therapy.

Antifungal Agents

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n = 83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Predicting natural variation in the yeast phenotypic landscape with machine learning.

Most organismal traits result from the complex interplay of many genetic and environmental factors, making their prediction difficult. Here, we used machine learning (ML) models to explore phenotype predictions for 223 traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains isolated worldwide. We benchmarked a ML pipeline with multiple linear and non-linear models to predict phenotypes from genotypes and gene expression, and determined gradient boosting machines as the best-performing model. Gene function disruption scores and gene presence/absence emerged as best predictors, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The prediction accuracy broadly varied among phenotypes, with stress resistance being easier to predict compared to growth across nutrients. ML identified relevant genomic features linked to phenotypes, including high-impact variants with established relationships to phenotypes, despite these being rare in the population. Near-perfect accuracies were achieved when other phenomics data mostly in similar conditions were used, suggesting that useful information can be conveyed across phenotypes. Overall, our study underscores the power of ML to interpret the functional outcome of genetic variants.

Genetic Variation

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in&#xa0;gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n&#x2009;=&#x2009;42, Healthy: n&#x2009;=&#x2009;54), Franzosa et al. (IBD: n&#x2009;=&#x2009;164, Healthy: n&#x2009;=&#x2009;56), and Yachida et al. (CRC: n&#x2009;=&#x2009;150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans

Dissecting genetic architecture and improving machine learning&#x2011;based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC&#xa0;>&#xa0;0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

Stratifying lung adenocarcinoma: a novel prognostic model based on mitochondrial outer membrane permeabilization activity.

UNLABELLED: Mitochondrial outer membrane permeabilization (MOMP) is a core apoptotic regulatory event that dictates mitochondrial integrity, where full activation drives cell death and sublethal dysregulation contributes to tumor genomic instability. We used the Cancer Genome Atlas lung adenocarcinoma cohort (TCGA-LUAD) as the training cohort and the Gene Expression Omnibus dataset GSE42127 as the validation cohort to identify prognostic genes related to MOMP activity in lung adenocarcinoma (LUAD) and to evaluate their potential biological significance. By intersecting MOMP-related genes with differentially expressed genes, combined with survival analysis, Mendelian randomization analysis, and 101 machine-learning algorithm combinations, seven prognostic genes, namely BIRC5, PSMD11, TNFRSF13C, YWHAZ, YWHAG, CYCS, and LTB, were identified. Next, an optimal prognostic model was constructed based on the gradient boosting machine (GBM) algorithm. Based on the risk score, LUAD patients were stratified into high- and low-risk groups, and patients in the high-risk group exhibited poorer overall survival in both the training and validation cohorts. Furthermore, a nomogram integrating the risk score and clinicopathological factors was developed and showed favorable predictive performance for 1-, 3-, and 5-year survival. Meanwhile, functional and immune analyses revealed that the high-risk group was enriched in DNA replication-related pathways and demonstrated a higher tumor mutation burden (TMB). Correlation analysis indicated that TNFRSF13C was positively correlated with activated B cells, whereas BIRC5 was negatively correlated with eosinophils, suggesting that MOMP-related genes might be involved in remodeling the immune microenvironment of LUAD. Drug sensitivity analysis showed differences in predicted half-maximal inhibitory concentration (IC50) values between the risk groups, suggesting the potential value of this model in assisting therapeutic stratification. Single-cell RNA sequencing (scRNA-seq) further identified T lymphocytes as a key cell type, with numerous prognostic genes exhibiting differential expression in T cells or dynamic changes during differentiation. We suggest that the MOMP-related signature established in this study may provide a reference for prognostic stratification in LUAD and offers candidate prognostic genes for subsequent experimental and clinical validation. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05058-6.

Lung adenocarcinoma

Immunoinformatics Approach for Optimization of Targeted Vaccine Design: New Paradigm in Clinical Trials and Healthcare Management.

INTRODUCTION: The immunoinformatics approach combines bioinformatics and computational tools, offering a revolutionary method for improving vaccine development by analyzing immune responses at the molecular level. Immunoinformatics enables the creation of customized vaccines designed for specific infections or cancer cells. OBJECTIVE: The primary objective of immunoinformatics is to enhance the vaccine development process by predicting and boosting the body's immune response. It aims to identify potential immunogenic epitopes and biomarkers that are important for creating vaccines with greater specificity and efficacy, especially when dealing with large-scale data. METHODS: Immunoinformatics utilizes a combination of proteomic, genomic, and epigenomic data, as well as machine learning algorithms and artificial intelligence techniques. These tools predict how various immunological components, e.g., T-cell and B-cell epitopes, interact with the immune system. This approach allows researchers to avoid traditional trial-and-error methods, enabling the efficient identification of potential vaccine candidates. Additionally, personalized vaccines can be developed by considering individual genetic and immunological characteristics. RESULTS: The use of immunoinformatics techniques accelerates the screening of vaccine candidates, enhances patient stratification, and optimizes formulations for clinical trials. This approach has been shown to improve vaccine safety, efficacy, and development speed. It also holds promise for managing healthcare on a large scale by producing vaccines tailored to specific populations, thereby improving the overall effectiveness of vaccination programs. CONCLUSION: Immunoinformatics represents a transformative approach to vaccine research, improving clinical trial efficiency and enabling the development of more reliable, flexible, and personalized vaccines. This approach has the potential to significantly enhance global healthcare outcomes by accelerating the vaccine development process and optimizing vaccination strategies.

Immunoinformatics

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans