PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve–based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC = 0.844). The NODM cohort was stratified into high- (n = 2,362) and low-risk (n = 5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)↗

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n&#x202f;=&#x202f;26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette &#x2248; 0.16) that remained unassociated with overall survival (log-rank p&#x202f;=&#x202f;0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) =&#x202f;0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p&#x202f;=&#x202f;0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans↗

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n&#xa0;=&#xa0;5), support vector machines (n&#xa0;=&#xa0;4), k-nearest neighbor (n&#xa0;=&#xa0;3), decision trees (n&#xa0;=&#xa0;3), random forests (n&#xa0;=&#xa0;5), neural networks (n&#xa0;=&#xa0;2), linear discriminant analysis (n&#xa0;=&#xa0;1), and pre-trained CNNs (n&#xa0;=&#xa0;1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n&#xa0;=&#xa0;12 to n&#xa0;=&#xa0;39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-&#x3ba;B, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-&#x3ba;B, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Machine learning-based drug susceptibility prediction from Candida genomic data.

OBJECTIVES: Invasive Candida infection is an increasing clinical concern, with antifungal resistance rising across multiple species. However, rapid and accurate antifungal susceptibility testing (AFST) remains limited in routine practice. The study evaluated species distribution and antifungal susceptibility of invasive Candida isolates in China and assessed the feasibility of combining whole-genome sequencing (WGS) with machine learning to predict minimum inhibitory concentrations (MICs). METHODS: Consecutive non-repetitive isolates were collected from 20 hospitals in 13 provinces during 2022-2023. MICs of nine antifungal agents were determined by broth microdilution, and WGS was performed for species accounting for >5% of the total isolates. Genomic 11-mer features were extracted and used to train random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost) models, followed by optimization of the best-performing algorithm. RESULTS: A total of 337 isolates were obtained from blood (n = 232) and sterile body fluids (n = 105), comprising C. albicans (n = 103), C. tropicalis (n = 71), C. parapsilosis (n = 67), and C. glabrata (n = 63). Non-albicans Candida showed higher azole and echinocandin resistance, with C. tropicalis notably resistant to azoles and C. glabrata to echinocandins. Among the three models, RF demonstrated the best performance on 304 sequenced isolates. The optimized RF model was evaluated by the receiver operating characteristic (ROC) curve analysis and achieved an average area under the ROC curve (AUC) of 0.979 (95% CI: 0.974-0.984), essential agreement over 90.1%, and categorical agreement over 93.2% across species. CONCLUSIONS: These findings underscore the clinical challenge posed by non-albicans Candida resistance, and indicate that WGS-based MIC prediction may offer a highly accurate reference for earlier antifungal therapy.

Antifungal Agents↗

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning↗

Predicting 5-Year Mortality in Non-Small-Cell Lung Cancer Using the Korean Central Cancer Registry: Model Development and Validation Study.

BACKGROUND: Non-small-cell lung cancer (NSCLC) is one of the most common cancers and a leading cause of cancer-related mortality, making prognostic prediction clinically essential. Machine learning models are increasingly used to assess prognosis; however, developing systems that combine high discrimination with clear, clinically interpretable reasoning remains challenging. OBJECTIVE: This study aimed to develop deep learning models that predict 5-year mortality in NSCLC using data from the Korea Central Cancer Registry and quantify feature importance through permutation testing. METHODS: We identified 3144 patients diagnosed between 2014 and 2017 who had complete clinical data, pulmonary function test results, histological information, genomic data, and staging details. After preprocessing, the cohort was divided into stratified training, validation, and test sets in a 70%-15%-15% ratio. Five models were tuned using Hyperband across 10 predefined feature groups. The primary evaluation metric was the area under the receiver operating characteristic curve (AUC); additional metrics included accuracy, F1-score, precision, and recall. Groupwise permutation importance was calculated for each model, and the concordance of importance rankings was assessed using the Friedman test. RESULTS: All 5 models yielded comparable discrimination values on the test set (AUC=0.875-0.879). Model A was selected as the primary model and achieved an AUC of 0.879, an accuracy of 0.806, an F1-score of 0.824, and a Brier score of 0.142. Permuting the stage resulted in the largest decrease in AUC (0.217), followed by the pulmonary function test (0.016). Gene mutation had a modest overall impact but became more influential within the adenocarcinoma subset. The Friedman test showed no statistically significant differences in importance rankings across the models (P=.93). CONCLUSIONS: A grouped-input deep learning framework achieved discrimination comparable to a conventional Cox proportional hazards model using the same routine clinical variables for 5-year mortality prediction in NSCLC. Group-level permutation importance provided stable and reproducible insights into the clinical factors influencing risk, which may guide future model refinement and clinical decision-making.

Humans↗

Machine learning-based analysis of oral rinse samples to identify candidate proteomic signatures for severe periodontitis: a pilot study.

This pilot study investigated whether candidate protein signatures from oral rinse samples can distinguish patients with severe periodontitis (stage III/IV) and its subtypes, generalized and localized periodontitis, from non-periodontitis controls. Participants rinsed with phosphate-buffered saline, and samples were analyzed using a Proximity Extension Assay targeting 92 inflammatory and 92 immuno-oncology proteins. A machine learning approach using repeated nested cross-validation and SHAP was implemented to identify protein signatures. The study included 38 patients (18 with localized periodontitis and 20 with generalized periodontitis) and 16 controls. After data preprocessing, 54 samples and 141 proteins were retained. Proteins Gal-1, HGF, TNFSF14, CD27, and ARG1 distinguished periodontitis from controls (ROC-AUC&#x2009;=&#x2009;0.85, 95% CI 0.82, 0.87). For generalized periodontitis, we found a protein signature including TNFSF14, Gal-1, STAMBP, MUC-16, S100A12, HGF, CASP-8, CD27, LAP TGF-&#x3b2;1, TNFRSF9, and uPA (ROC-AUC&#x2009;=&#x2009;0.92, 95% CI 0.90, 0.94). For localized periodontitis, we identified ARG1 (ROC-AUC&#x2009;=&#x2009;0.72, 95% CI 0.68, 0.76). No proteomic signature distinguishing generalized periodontitis from localized periodontitis was identified. This pilot study indicated that oral rinses are suitable for proteomic profiling, and there was a putative protein signature that could differentiate periodontitis, generalized periodontitis, and localized periodontitis from controls. These findings warrant validation in larger independent cohorts, including a clearly defined gingivitis group, before real-world non-invasive screening applications can be considered.

Humans↗

Developing a machine learning-based prognosis and immunotherapeutic response signature in colorectal cancer: insights from ferroptosis, fatty acid dynamics, and the tumor microenvironment.

INSTRUCTION: Colorectal cancer (CRC) poses a challenge to public health and is characterized by a high incidence rate. This study explored the relationship between ferroptosis and fatty acid metabolism in the tumor microenvironment (TME) of patients with CRC to identify how these interactions impact the prognosis and effectiveness of immunotherapy, focusing on patient outcomes and the potential for predicting treatment response. METHODS: Using datasets from multiple cohorts, including The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO), we conducted an in-depth multi-omics study to uncover the relationship between ferroptosis regulators and fatty acid metabolism in CRC. Through unsupervised clustering, we discovered unique patterns that link ferroptosis and fatty acid metabolism, and further investigated them in the context of immune cell infiltration and pathway analysis. We developed the FeFAMscore, a prognostic model created using a combination of machine learning algorithms, and assessed its predictive power for patient outcomes and responsiveness to treatment. The FeFAMscore signature expression level was confirmed using RT-PCR, and ACAA2 progression in cancer was further verified. RESULTS: This study revealed significant correlations between ferroptosis regulators and fatty acid metabolism-related genes with respect to tumor progression. Three distinct patient clusters with varied prognoses and immune cell infiltration were identified. The FeFAMscore demonstrated superior prognostic accuracy over existing models, with a C-index of 0.689 in the training cohort and values ranging from 0.648 to 0.720 in four independent validation cohorts. It also responses to immunotherapy and chemotherapy, indicating a sensitive response of special therapies (e.g., anti-PD-1, anti-CTLA4, osimertinib) in high FeFAMscore patients. CONCLUSION: Ferroptosis regulators and fatty acid metabolism-related genes not only enhance immune activation, but also contribute to immune escape. Thus, the FeFAMscore, a novel prognostic tool, is promising for predicting both the prognosis and efficacy of immunotherapeutic strategies in patients with CRC.

Ferroptosis↗

Metagenomic polymorphic toxin effector and immunity profiling predicts microbiome development and disease-related dysbiosis.

Bacteria use antagonistic interbacterial weapons, such as polymorphic toxin secretion systems (TSS), to compete for niches in the human gut microbiome. We hypothesized that TSS influence gut microbiome development and disease-related dysbiosis. We developed a bioinformatic marker gene approach (PolyProf) to quantify TSS including ~200 effector and immunity genes and applied it to ~15,000 publicly available human metagenomes. PolyProf alpha and beta diversity readily distinguished 12 different human disease states and enabled the construction of highly accurate linear regression classifier machine learning models. Elastic net machine learning models integrating bacterial taxonomy with PolyProf had strong predictive value for 12 disease states, outperforming models utilizing taxonomy alone. During microbiome development in the first year of life, PolyProf alpha diversity increases, and beta diversity becomes increasingly like the maternal microbiome, influenced by vertical transfer, delivery mode, and breastfeeding. PolyProf is related to strain sharing among adults through social interactions. In summary, TSS genes strongly correlate with microbiome development and interpersonal strain sharing, suggesting roles for interbacterial antagonism. Since PolyProf distinguishes diverse adult disease statuses, these dynamics may contribute to non-genetic inheritance.IMPORTANCEPrevious research has demonstrated that bacteria compete within the gut microbiome using toxin secretion systems (TSS). How TSS contribute to human microbiome development and the microbiome alterations observed in human diseases is not known. This study develops a new bioinformatic tool for profiling TSS-related genes in metagenomic data. Application of this approach to large-scale human fecal metagenomic data demonstrates the dynamic association of TSS during microbiome development, including the exchange of strains among social contacts. TSS gene abundance patterns are highly predictive of 12 disease states. This study advances the field by enabling TSS profiling in metagenomes and by identifying disease and microbiome development biomarkers that provide hypotheses for future mechanistic studies and may be useful for disease diagnosis.

Dysbiosis↗

Machine learning-based analysis of the impact of 5'&#xa0;untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions↗

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN↗

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning↗

Rule induction and instance-based learning applied in medical diagnosis.

Machine learning methods have been applied in a variety of medical domains in order to improve medical decision making. Improved medical diagnosis and prognosis can be achieved through automatic analysis of patient data stored in medical records, i.e., by learning from past experience. Given patient records with corresponding diagnoses, machine learning methods are able to classify new cases either through constructing explicit rules that generalize the training cases (e.g., rule induction) or by storing (some of) the training cases for reference (instance-based learning). This paper presents the methodologies of rule induction and instance-based learning and their application to medical diagnosis, in particular, the problem of early diagnosis of rheumatic diseases. It also discusses the possibility to use existing expert knowledge to support the learning process and the utility of such knowledge.

Algorithms↗

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans↗

Unveiling the power of TIIC: A prognostic tool for esophageal adenocarcinoma.

BACKGROUND: Esophageal adenocarcinoma (EAC) remains a lethal malignancy with limited prognostic tools for guiding immunotherapy. Tumor-infiltrating immune cells (TIICs) play a critical role in EAC prognosis and treatment response. METHODS: We integrated single-cell RNA sequencing and bulk transcriptome data from TCGA and GEO databases. TIIC-specific RNAs were identified via tissue specificity index calculation combined with machine learning feature selection. Twenty machine learning algorithms were benchmarked to construct an optimal TIIC signature score (TIIC-Score) based on the comprehensive C-index. Immunotherapy response, genomic mutation, and copy number variation were analyzed. Summary-data-based Mendelian randomization (SMR) and two-sample Mendelian randomization (MR) were performed to explore genetic associations. Core prognostic TIIC-related genes were functionally validated in esophageal cancer cell lines through loss-of-function assays. RESULTS: The TIIC-Score demonstrated robust prognostic value for 1-, 2-, and 3-year overall survival across multiple cohorts, outperforming 22 published models. High TIIC-Score was associated with poor survival and increased chromosomal instability. Mutation profiling revealed high frequencies of TP53 (78.2%), TTN (48.7%), and SYNE1 (30.8%). MR analysis identified a significant association between gastro-oesophageal reflux and EAC risk at SNP rs8130507. Functionally, CCNI was upregulated in esophageal cancer cells, and its knockdown suppressed malignant phenotypes while promoting apoptosis, supporting its pro-tumorigenic role. CONCLUSION: The TIIC-Score provides a novel prognostic framework for EAC that effectively stratifies patient risk and may help identify individuals most likely to benefit from immunotherapy.

Esophageal adenocarcinoma↗

MyESL: A Software for Evolutionary Sparse Learning in Molecular Phylogenetics and Genomics.

Evolutionary sparse learning uses supervised machine learning to build evolutionary models where genomic sites loci are parameters. It uses the Least Absolute Shrinkage and Selection Operator with bi-level sparsity to connect a specific phylogenetic hypothesis with sequence variation across genomic loci. The MyESL software addresses the need for open-source tools to perform evolutionary sparse learning analyses, offering features to preprocess input phylogenomic alignments, post-process output models to generate molecular evolutionary metrics, and make Least Absolute Shrinkage and Selection Operator regression adaptable and efficient for phylogenetic trees and alignments. The core of MyESL, which constructs models with logistic regressions using bi-level sparsity, is written in C++. Its input data preprocessing and result post-processing tools are developed in Python. Compared to other tools, MyESL is more computationally efficient and provides evolution-friendly inputs and outputs. These features have already enabled the use of MyESL in two phylogenomic applications, one to identify outlier sequences and fragile clades in inferred phylogenies and another to build genetic models of convergent traits. In addition to the use in a Python environment, MyESL is available as a standalone executable compatible across multiple platforms, which can be directly integrated into scripts and third-party software. The source code, executable, and documentation for MyESL are openly accessible at https://github.com/kumarlabgit/MyESL.

Phylogeny↗