PubMed HealthSearch

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning

Proteomic and machine learning analysis predicts treatment response signatures in Myasthenia Gravis.

BACKGROUND: Myasthenia gravis (MG) is a prototypical antibody-mediated autoimmune disease with variable treatment responses with a need for biomarkers to guide therapeutic decision making. Proteomic profiling, coupled with machine learning, offers a hypothesis-free approach to identify multi-protein signatures associated with treatment response. METHODS: We analyzed sera collected at entry (baseline) from participants in a phase 3 trial randomized trial comparing thymectomy plus prednisone versus prednisone alone, along with matched controls using liquid chromatography-mass spectrometry. We derived disease-specific proteomic signatures and evaluated associations between baseline proteins and 6-month clinical outcomes using multiple machine-learning approaches with internal validation. RESULTS: Baseline serum proteomes distinguished MG from controls, with pathway enrichment implicating complement activation, immunoglobulin production, and T-cell receptor signaling. Distinct protein panels predicted 6-month clinical improvement within each treatment arm. In the thymectomy-plus-prednisone group, models captured non-linear relationships of predictive proteins in contrast with the predominant additive patterns observed in the prednisone-alone group. Predictive proteins were enriched for T-cell signaling and leukocyte trafficking functions, providing insight into treatment-specific biology. CONCLUSIONS: Baseline serum proteomics captures core disease characteristics of MG and predicts short-term clinical response in a treatment-specific manner. While our results require validation in independent cohorts, these findings could enable biomarker-guided selection of thymectomy, refine risk stratification, and furnish mechanistic readouts for future MG trials and clinical care. We aim to conduct future studies using -omic approaches to validate these baseline predictive biomarkers and pathways of treatment response in patients with MG.

Adult

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus

Machine learning-based drug susceptibility prediction from Candida genomic data.

OBJECTIVES: Invasive Candida infection is an increasing clinical concern, with antifungal resistance rising across multiple species. However, rapid and accurate antifungal susceptibility testing (AFST) remains limited in routine practice. The study evaluated species distribution and antifungal susceptibility of invasive Candida isolates in China and assessed the feasibility of combining whole-genome sequencing (WGS) with machine learning to predict minimum inhibitory concentrations (MICs). METHODS: Consecutive non-repetitive isolates were collected from 20 hospitals in 13 provinces during 2022-2023. MICs of nine antifungal agents were determined by broth microdilution, and WGS was performed for species accounting for >5% of the total isolates. Genomic 11-mer features were extracted and used to train random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost) models, followed by optimization of the best-performing algorithm. RESULTS: A total of 337 isolates were obtained from blood (n = 232) and sterile body fluids (n = 105), comprising C. albicans (n = 103), C. tropicalis (n = 71), C. parapsilosis (n = 67), and C. glabrata (n = 63). Non-albicans Candida showed higher azole and echinocandin resistance, with C. tropicalis notably resistant to azoles and C. glabrata to echinocandins. Among the three models, RF demonstrated the best performance on 304 sequenced isolates. The optimized RF model was evaluated by the receiver operating characteristic (ROC) curve analysis and achieved an average area under the ROC curve (AUC) of 0.979 (95% CI: 0.974-0.984), essential agreement over 90.1%, and categorical agreement over 93.2% across species. CONCLUSIONS: These findings underscore the clinical challenge posed by non-albicans Candida resistance, and indicate that WGS-based MIC prediction may offer a highly accurate reference for earlier antifungal therapy.

Antifungal Agents

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics&#x2013;based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics&#x2013;based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

Machine learning analysis of the human initiator region reveals key features of different types of core promoters.

The initiator (Inr) is the starting point for the transcription of many genes. Here, we generated highly predictive machine learning models of the human Inr region, and determined that the Inr is present in &#x223c;60% of focused human promoters, identified a novel TATA-specific Inr, and detected the overlapping but functionally distinct TCT motif. Quantitative genome-wide analyses revealed a strict and synergistic interaction between the Inr and DPR, an inverse relationship between the TATA and DPR, a flexible and sometimes independent function of the TATA box in relation to the Inr, and different properties of the TCT motif in humans versus Drosophila.

Humans

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans

Dissecting genetic architecture and improving machine learning&#x2011;based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC&#xa0;>&#xa0;0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

Pathomics-based machine learning models for predicting METTL5 expression and prognosis in lung adenocarcinoma.

BACKGROUND: METTL5, an N6-methyladenosine (m6A) RNA methyltransferase, has been implicated in tumor progression, but its prognostic value and non-invasive prediction in lung adenocarcinoma (LUAD) remain unclear. This study aimed to develop a pathomics-based machine learning model to predict METTL5 expression from histopathological images and evaluate its prognostic significance in LUAD. METHODS: A total of 327 LUAD patients from The Cancer Genome Atlas (TCGA) with matched hematoxylin and eosin (H&E) slides, transcriptomic, and clinical data were included and randomly divided into training and validation sets (7:3). Quantitative histopathological features were extracted using PyRadiomics. Feature selection was performed via maximum relevance minimum redundancy (mRMR) and recursive feature elimination (RFE), followed by construction of a Gradient Boosting Machine (GBM) model. A pathomics score (PS) was generated to assess prognostic relevance. Survival analyses, gene set variation analysis (GSVA), tumor mutational burden (TMB), immune infiltration analysis, and in vitro functional assays were conducted. RESULTS: METTL5 overexpression was independently associated with poor overall survival [hazard ratio (HR) =1.637, P=0.007]. The model achieved good predictive performance [area under the curve (AUC) =0.847 in the training set and 0.752 in the validation set]. High PS was significantly associated with worse survival and remained an independent prognostic factor (HR =1.563, P=0.03). Elevated PS correlated with altered metabolic pathways, increased TMB, and immune microenvironment changes. METTL5 knockdown reduced proliferation, migration, invasion, and epithelial-mesenchymal transition (EMT) in A549 cells. CONCLUSIONS: The pathomics-based model accurately predicts METTL5 expression and provides prognostic stratification in LUAD, supporting its potential as a practical imaging-derived biomarker.

Methyltransferase-like 5

Uncovering encrypted antimicrobial peptides in health-associated Lactobacillaceae by large-scale genomics and machine learning.

BACKGROUND: Antimicrobial peptides (AMPs) are well known for their broad-spectrum activity and have shown great promise in addressing the antibiotic-resistant crisis. The Lactobacillaceae family, recognized for its health-promoting effects in humans, represents a valuable source of novel AMPs. However, the global prevalence and distribution of AMPs within Lactobacillaceae remains largely unknown, which limits the efficient discovery and development of novel AMPs. RESULTS: We analyzed all available genomes (10,327 genomes), encompassing 38 genera and 515 species, to investigate the biosynthetic potential (indicated by the number of AMP sequences in the genome) of AMP in the Lactobacillaceae family. We demonstrated Lactobacillaceae species had ubiquitous (69.90%) biosynthetic potential of AMPs. Overall, 9601 AMPs were identified, clustering into 2092 gene cluster families (GCFs), which showed strong interspecies specificity (95.27%), intraspecies heterogeneity (93.31%), and habitat uniqueness (95.83%), that greatly expanded on the AMP sequence landscape. Novelty assessment indicated that 1516 GCFs (72.47%) had no similarity to any known AMPs in existing databases. Machine learning predictions suggested that novel AMPs from Lactobacillaceae possessed strong antimicrobial potential, with 664 GCFs having an additive minimum inhibitory concentration (MIC) below 100&#xa0;&#x3bc;M. We randomly synthesized 16 AMPs (with predicted MIC&#x2009;<&#x2009;100&#xa0;&#x3bc;M) and identified 10 AMPs exhibiting varied-spectrum activity against 11 common pathogens. Finally, we identified one Lactobacillus delbrueckii-originated AMP (delbruin_1) having broad-spectrum (all 11 pathogens) and high antimicrobial activity (average MIC&#x2009;=&#x2009;38.56 &#xb5;M), which proved its potential as a clinically viable antimicrobial agent. CONCLUSIONS: We uncovered the global prevalence of AMPs in Lactobacillaceae and proved that Lactobacillaceae is an untapped and invaluable source of novel AMPs to combat the antibiotic-resistance crisis. Meanwhile, we provided a machine learning-guided framework for AMP discovery, offering a scalable roadmap for identifying novel AMPs not only in Lactobacillaceae but also in other organisms. Video Abstract.

Machine Learning

Bioactive peptides for meat quality and preservation: Integrating peptidomics and computational screening.

Bioactive peptides generated from meat proteins, fermented meat products, and slaughter by-products have attracted increasing attention as functional molecules for improving meat quality and preservation. In meat systems, peptides can be produced through endogenous postmortem proteolysis, microbial fermentation, gastrointestinal digestion, or controlled enzymatic hydrolysis of underutilized animal by-products. These peptides are closely associated with key meat science endpoints, including postmortem tenderization, oxidative stability, color retention, flavor development, microbial inhibition, and the valorization of processing by-products. However, although high-resolution peptidomics has greatly expanded the identification of meat-derived peptide sequences, their translation into practical meat applications remains limited by matrix interactions, processing stability, sensory constraints, safety concerns, and insufficient validation in real meat systems. This review synthesizes recent advances in meat-related peptidomics and computational screening, including sequence-based prediction, machine learning, molecular docking, molecular dynamics, stability assessment, and safety-oriented filtering. Particular attention is given to how these approaches can prioritize peptides with antioxidant, antimicrobial, flavor-modulating, and preservation-related functions under meat-specific technological constraints. By integrating peptide generation pathways, mass spectrometry-based identification, in silico prioritization, and meat quality endpoints, this review proposes a stage-gated framework for translating meat-derived bioactive peptides from discovery to application. Future research should strengthen matrix-specific validation, standardized peptidomic reporting, and safety assessment to support the use of bioactive peptides in meat quality improvement, clean-label preservation, and circular utilization of meat industry by-products.

Animals

Ecological Restoration of the Soil-Like Function in the Bauxite Residue: Natural Microbiomes Mediated Molecular Transformation of Dissolved Organic Matter.

Soilization of bauxite residues offers a scalable route for long-term carbon management and ecological restoration. However, the microbial processes that transform exogenous organic inputs into stable soil-like carbon pools remain poorly resolved. Here, we combined cross-ecosystem meta-analysis, machine-learning prediction, native synthetic community (SynCom) construction, 13C-labeled straw microcosms, field validation, Fourier transform ion cyclotron resonance mass spectrometry, and genome-resolved metagenomics to unravel microbiome-mediated carbon transformation at the dissolved organic matter (DOM) molecular scale. Our meta-analysis revealed that alkaline industrial wastes retained soil-like DOM signatures but were enriched in microbial humic- and protein-like components, indicating active yet incomplete carbon processing. Guided by these patterns, native SynCom inoculation increased 13C incorporation into total organic carbon (TOC) and dissolved organic carbon (DOC), enlarged biodegradable and adsorbable DOC fractions, and shifted DOM from recalcitrant aromatic pools toward oxygenated carbohydrate-, tannin-, and phenolic-like molecular classes. Genome-resolved analyses linked this transformation to complementary polymer degradation and nutrient-cycling functions across fungal and bacterial guilds, including enriched carbohydrate-active enzymes in straw-carbon-utilizing metagenome-assembled genomes. Null model and thermodynamic analyses further showed that microbial communities were constrained by homogeneous selection, whereas DOM molecules were diversified through variable selection and redox-dependent transformation. Field-scale validation confirmed that SynCom promoted TOC and DOC accumulation and humic-like, high-density DOM fractions under alkaline conditions. Together, these findings establish a mechanistic framework in which functional microbiomes couple plant carbon depolymerization, DOM molecular diversification, and mineral-interactive carbon stabilization, providing a microbiome-guided strategy for carbon sequestration and soilization in the bauxite residue.

Soil

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of &#x2265;40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged &#x2265;65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans