PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans

Machine learning-assisted plasma PEA proteomics enables differential diagnosis of melancholic depression and bipolar disorder.

Differentiating bipolar disorder (BD) from major depressive disorder (MDD) remains a critical unmet need in psychiatry due to overlapping clinical presentations and the absence of reliable biological markers. In this study, we assessed the capacity of multivariate machine learning models to accurately differentiate BD from MDD with melancholic features using plasma proteomic profiles obtained via Proximity Extension Assay (PEA) technology. A total of 67 participants were included (23 BD, 20 MDD, and 24 HC), and plasma protein expression was assessed using the Olink Target 96 Neurology panel. Differential proteomic analysis revealed distinct disorder-specific expression patterns, identifying 21 differentially expressed proteins in BD versus MDD, 18 in BD versus healthy controls, and 7 in MDD versus healthy controls. Using a stepwise feature reduction strategy, machine learning models were trained on three feature sets comprising all proteins, the top 20 most informative proteins, and the top 5 most beneficial proteins, and evaluated across BD-MDD, BD-HC, and MDD-HC classification tasks using five algorithms. For BD-MDD discrimination, the Random Forest model achieved the highest performance when trained on the top 5 protein set (LXN, HAGH, MATN3, PLXNB1, and CTSC), yielding an AUC of 0.905, with similarly strong performance observed using the top 20 protein set. Feature importance analysis highlighted proteins involved in neurodevelopmental processes, immune regulation, and extracellular matrix organization. Overall, these findings demonstrate that integrating plasma proteomics with machine learning enables robust differentiation between BD and MDD with melancholic features, supporting the development of scalable and biologically informed diagnostic tools for precision psychiatry.

Bipolar disorder

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n&#x202f;=&#x202f;26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette &#x2248; 0.16) that remained unassociated with overall survival (log-rank p&#x202f;=&#x202f;0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) =&#x202f;0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p&#x202f;=&#x202f;0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n&#xa0;=&#xa0;5), support vector machines (n&#xa0;=&#xa0;4), k-nearest neighbor (n&#xa0;=&#xa0;3), decision trees (n&#xa0;=&#xa0;3), random forests (n&#xa0;=&#xa0;5), neural networks (n&#xa0;=&#xa0;2), linear discriminant analysis (n&#xa0;=&#xa0;1), and pre-trained CNNs (n&#xa0;=&#xa0;1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n&#xa0;=&#xa0;12 to n&#xa0;=&#xa0;39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-&#x3ba;B, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-&#x3ba;B, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4

Machine learning-based drug susceptibility prediction from Candida genomic data.

OBJECTIVES: Invasive Candida infection is an increasing clinical concern, with antifungal resistance rising across multiple species. However, rapid and accurate antifungal susceptibility testing (AFST) remains limited in routine practice. The study evaluated species distribution and antifungal susceptibility of invasive Candida isolates in China and assessed the feasibility of combining whole-genome sequencing (WGS) with machine learning to predict minimum inhibitory concentrations (MICs). METHODS: Consecutive non-repetitive isolates were collected from 20 hospitals in 13 provinces during 2022-2023. MICs of nine antifungal agents were determined by broth microdilution, and WGS was performed for species accounting for >5% of the total isolates. Genomic 11-mer features were extracted and used to train random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost) models, followed by optimization of the best-performing algorithm. RESULTS: A total of 337 isolates were obtained from blood (n = 232) and sterile body fluids (n = 105), comprising C. albicans (n = 103), C. tropicalis (n = 71), C. parapsilosis (n = 67), and C. glabrata (n = 63). Non-albicans Candida showed higher azole and echinocandin resistance, with C. tropicalis notably resistant to azoles and C. glabrata to echinocandins. Among the three models, RF demonstrated the best performance on 304 sequenced isolates. The optimized RF model was evaluated by the receiver operating characteristic (ROC) curve analysis and achieved an average area under the ROC curve (AUC) of 0.979 (95% CI: 0.974-0.984), essential agreement over 90.1%, and categorical agreement over 93.2% across species. CONCLUSIONS: These findings underscore the clinical challenge posed by non-albicans Candida resistance, and indicate that WGS-based MIC prediction may offer a highly accurate reference for earlier antifungal therapy.

Antifungal Agents

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning

Predicting 5-Year Mortality in Non-Small-Cell Lung Cancer Using the Korean Central Cancer Registry: Model Development and Validation Study.

BACKGROUND: Non-small-cell lung cancer (NSCLC) is one of the most common cancers and a leading cause of cancer-related mortality, making prognostic prediction clinically essential. Machine learning models are increasingly used to assess prognosis; however, developing systems that combine high discrimination with clear, clinically interpretable reasoning remains challenging. OBJECTIVE: This study aimed to develop deep learning models that predict 5-year mortality in NSCLC using data from the Korea Central Cancer Registry and quantify feature importance through permutation testing. METHODS: We identified 3144 patients diagnosed between 2014 and 2017 who had complete clinical data, pulmonary function test results, histological information, genomic data, and staging details. After preprocessing, the cohort was divided into stratified training, validation, and test sets in a 70%-15%-15% ratio. Five models were tuned using Hyperband across 10 predefined feature groups. The primary evaluation metric was the area under the receiver operating characteristic curve (AUC); additional metrics included accuracy, F1-score, precision, and recall. Groupwise permutation importance was calculated for each model, and the concordance of importance rankings was assessed using the Friedman test. RESULTS: All 5 models yielded comparable discrimination values on the test set (AUC=0.875-0.879). Model A was selected as the primary model and achieved an AUC of 0.879, an accuracy of 0.806, an F1-score of 0.824, and a Brier score of 0.142. Permuting the stage resulted in the largest decrease in AUC (0.217), followed by the pulmonary function test (0.016). Gene mutation had a modest overall impact but became more influential within the adenocarcinoma subset. The Friedman test showed no statistically significant differences in importance rankings across the models (P=.93). CONCLUSIONS: A grouped-input deep learning framework achieved discrimination comparable to a conventional Cox proportional hazards model using the same routine clinical variables for 5-year mortality prediction in NSCLC. Group-level permutation importance provided stable and reproducible insights into the clinical factors influencing risk, which may guide future model refinement and clinical decision-making.

Humans

Machine learning-based analysis of oral rinse samples to identify candidate proteomic signatures for severe periodontitis: a pilot study.

This pilot study investigated whether candidate protein signatures from oral rinse samples can distinguish patients with severe periodontitis (stage III/IV) and its subtypes, generalized and localized periodontitis, from non-periodontitis controls. Participants rinsed with phosphate-buffered saline, and samples were analyzed using a Proximity Extension Assay targeting 92 inflammatory and 92 immuno-oncology proteins. A machine learning approach using repeated nested cross-validation and SHAP was implemented to identify protein signatures. The study included 38 patients (18 with localized periodontitis and 20 with generalized periodontitis) and 16 controls. After data preprocessing, 54 samples and 141 proteins were retained. Proteins Gal-1, HGF, TNFSF14, CD27, and ARG1 distinguished periodontitis from controls (ROC-AUC&#x2009;=&#x2009;0.85, 95% CI 0.82, 0.87). For generalized periodontitis, we found a protein signature including TNFSF14, Gal-1, STAMBP, MUC-16, S100A12, HGF, CASP-8, CD27, LAP TGF-&#x3b2;1, TNFRSF9, and uPA (ROC-AUC&#x2009;=&#x2009;0.92, 95% CI 0.90, 0.94). For localized periodontitis, we identified ARG1 (ROC-AUC&#x2009;=&#x2009;0.72, 95% CI 0.68, 0.76). No proteomic signature distinguishing generalized periodontitis from localized periodontitis was identified. This pilot study indicated that oral rinses are suitable for proteomic profiling, and there was a putative protein signature that could differentiate periodontitis, generalized periodontitis, and localized periodontitis from controls. These findings warrant validation in larger independent cohorts, including a clearly defined gingivitis group, before real-world non-invasive screening applications can be considered.

Humans

Developing a machine learning-based prognosis and immunotherapeutic response signature in colorectal cancer: insights from ferroptosis, fatty acid dynamics, and the tumor microenvironment.

INSTRUCTION: Colorectal cancer (CRC) poses a challenge to public health and is characterized by a high incidence rate. This study explored the relationship between ferroptosis and fatty acid metabolism in the tumor microenvironment (TME) of patients with CRC to identify how these interactions impact the prognosis and effectiveness of immunotherapy, focusing on patient outcomes and the potential for predicting treatment response. METHODS: Using datasets from multiple cohorts, including The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO), we conducted an in-depth multi-omics study to uncover the relationship between ferroptosis regulators and fatty acid metabolism in CRC. Through unsupervised clustering, we discovered unique patterns that link ferroptosis and fatty acid metabolism, and further investigated them in the context of immune cell infiltration and pathway analysis. We developed the FeFAMscore, a prognostic model created using a combination of machine learning algorithms, and assessed its predictive power for patient outcomes and responsiveness to treatment. The FeFAMscore signature expression level was confirmed using RT-PCR, and ACAA2 progression in cancer was further verified. RESULTS: This study revealed significant correlations between ferroptosis regulators and fatty acid metabolism-related genes with respect to tumor progression. Three distinct patient clusters with varied prognoses and immune cell infiltration were identified. The FeFAMscore demonstrated superior prognostic accuracy over existing models, with a C-index of 0.689 in the training cohort and values ranging from 0.648 to 0.720 in four independent validation cohorts. It also responses to immunotherapy and chemotherapy, indicating a sensitive response of special therapies (e.g., anti-PD-1, anti-CTLA4, osimertinib) in high FeFAMscore patients. CONCLUSION: Ferroptosis regulators and fatty acid metabolism-related genes not only enhance immune activation, but also contribute to immune escape. Thus, the FeFAMscore, a novel prognostic tool, is promising for predicting both the prognosis and efficacy of immunotherapeutic strategies in patients with CRC.

Ferroptosis

Metagenomic polymorphic toxin effector and immunity profiling predicts microbiome development and disease-related dysbiosis.

Bacteria use antagonistic interbacterial weapons, such as polymorphic toxin secretion systems (TSS), to compete for niches in the human gut microbiome. We hypothesized that TSS influence gut microbiome development and disease-related dysbiosis. We developed a bioinformatic marker gene approach (PolyProf) to quantify TSS including ~200 effector and immunity genes and applied it to ~15,000 publicly available human metagenomes. PolyProf alpha and beta diversity readily distinguished 12 different human disease states and enabled the construction of highly accurate linear regression classifier machine learning models. Elastic net machine learning models integrating bacterial taxonomy with PolyProf had strong predictive value for 12 disease states, outperforming models utilizing taxonomy alone. During microbiome development in the first year of life, PolyProf alpha diversity increases, and beta diversity becomes increasingly like the maternal microbiome, influenced by vertical transfer, delivery mode, and breastfeeding. PolyProf is related to strain sharing among adults through social interactions. In summary, TSS genes strongly correlate with microbiome development and interpersonal strain sharing, suggesting roles for interbacterial antagonism. Since PolyProf distinguishes diverse adult disease statuses, these dynamics may contribute to non-genetic inheritance.IMPORTANCEPrevious research has demonstrated that bacteria compete within the gut microbiome using toxin secretion systems (TSS). How TSS contribute to human microbiome development and the microbiome alterations observed in human diseases is not known. This study develops a new bioinformatic tool for profiling TSS-related genes in metagenomic data. Application of this approach to large-scale human fecal metagenomic data demonstrates the dynamic association of TSS during microbiome development, including the exchange of strains among social contacts. TSS gene abundance patterns are highly predictive of 12 disease states. This study advances the field by enabling TSS profiling in metagenomes and by identifying disease and microbiome development biomarkers that provide hypotheses for future mechanistic studies and may be useful for disease diagnosis.

Dysbiosis

Machine learning-based analysis of the impact of 5'&#xa0;untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans