PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Proteomics-enabled learning machine algorithms enhance the prediction of cardiovascular diseases in patients with type 2 diabetes mellitus.

BACKGROUND AND AIMS: Estimating the risk of cardiovascular disease (CVD) complications in type 2 diabetes mellitus (T2DM) patients is critical in the medical decision-making process. This study aimed to use a machine learning technique combined with proteomics to develop personalized models for predicting CVD in patients with T2DM. METHODS AND RESULTS: In total, 874 patients with T2DM and 2,920 Olink proteins obtained from the UK Biobank were used in this study. Proteins were screened using Cox regression and LASSO regression. A basic model containing clinical features and a full model combining proteome and clinical features were constructed using the random survival forest algorithm. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictive performance of the models and compare them with other CVD predictive models. Compared with the basic model, the full model performed better in predicting CVD, with time-dependent AUCs of 0.81 (3 years), 0.74 (5 years) and 0.74 (10 years) (0.77, 0.69 and 0.67). We calculated the risk scores of the Framingham, ASCVD and Score2-Diabetes models. The results revealed that the prediction performance of the full model was also better than that of the abovementioned models. In terms of differentiation accuracy, the results of the net reclassification improvement index and integrated discrimination improvement index showed that the full model can identify high-risk individuals more accurately (accuracy rate: 79% vs. 69%). CONCLUSIONS: Proteomics can be used to predict cardiovascular complications in diabetic patients. It is also necessary to consider the applicability of the model due to the limitations of the sample size and the constraints of proteomics in clinical applications.

Humans

Prediction of antimicrobial minimum inhibitory concentration from bacterial genomes using a scalable and interpretable machine learning approach.

Although machine learning models can predict antimicrobial susceptibility from bacterial whole genome sequencing (WGS), state-of-the-art approaches are computationally demanding or dependent on knowledge of genetic resistance determinants. Here, we describe an efficient data-driven approach to predicting minimum inhibitory concentration (MIC) by progressively extending and refining predictive genome segments, independent of prior knowledge of resistance determinants. Resultant models had high interpretability - known and potentially novel resistance determinants were captured. Using 762 clinical E. coli strains, 71.6% of predictions were within one dilution of the measured MIC. Models trained with this algorithm generalised better onto external data (F1 score = 0.85) compared with alternative models trained on annotated resistance determinants (F1 = 0.82) or k-mer counts (F1 = 0.74). Computational demands were low (RAM usage 23.6GB vs 38.8GB for k-mer model). These advantages represent an important advance in predicting antimicrobial susceptibility from WGS, with potential applications for clinical diagnostics, drug development, and surveillance.

Journal Article

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE: This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS: A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS: The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION: The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Morus

Modeling unknowns: A vision for uncertainty-aware machine learning in healthcare.

The integration of machine learning (ML) into healthcare is accelerating, driven by the proliferation of biomedical data and the promise of data-driven clinical support. A key challenge in this context is managing the pervasive uncertainty inherent in medical reasoning and decision-making. Despite its recognized importance, uncertainty is often underrepresented in the design and evaluation of clinical AI systems. Here we report an editorial overview of a special issue dedicated to uncertainty modeling in medical AI, which gathers theoretical, methodological, and practical contributions addressing this critical gap. Across these works, authors reveal that fewer than 4% of studies address uncertainty explicitly, and propose alternative design principles-such as optimizing for clinical net benefit or embedding explainability with confidence estimates. Notable contributions include the RelAI system for real-time prediction reliability, empirical findings on how uncertainty communication shapes clinical interpretation, and benchmarks for out-of-distribution detection in tabular data. Furthermore, this issue highlights the use of causal reasoning and anomaly detection to enhance system robustness and accountability. Together, these studies argue that representing, communicating, and operationalizing uncertainty are essential not only for clinical safety but also for building trust in AI-driven care. This special issue thus repositions uncertainty from a limitation to a foundational asset in the responsible deployment of ML in healthcare.

Machine Learning

Machine learning to differentiate colonization from infection in multidrug-resistant Gram-negative bacteria: implications for further research.

PURPOSE OF REVIEW: Machine learning has emerged as a promising tool to support antimicrobial decision-making in infectious diseases. In colonized patients, distinguishing multidrug-resistant Gram-negative bacteria (MDR-GNB) colonization from true infection remains a major clinical challenge, as both delayed appropriate therapy in severe infections and unnecessary broad-spectrum antimicrobial use may adversely affect patient outcomes and antimicrobial stewardship. This review discusses the current evidence on machine learning models for predicting or detecting MDR-GNB infection in colonized patients, highlights key methodological limitations of the available literature, and outlines future research priorities. RECENT FINDINGS: Current evidence specifically evaluating machine learning models beyond logistic regression in MDR-GNB-colonized patients remains limited. Overall, while machine learning may achieve encouraging discriminatory performance, important methodological limitations persist. Most notably, predictive models are frequently developed in heterogeneous populations that do not reflect the clinically relevant populations of colonized patients in which treatment decisions are made. Furthermore, improvements in predictive performance remain modest, possibly reflecting limited sample sizes and data granularity rather than insufficient algorithmic complexity. In our opinion, future advances could require multicenter datasets enriched with longitudinal clinical, microbiological, and genomic information, together with automated feature extraction from electronic health records. SUMMARY: The main challenge for machine learning in predicting MDR-GNB infection in colonized patients may lie not in developing increasingly sophisticated algorithms, but in generating clinically representative datasets and adopting rigorous methodological standards for model development, validation, calibration, and implementation. Future research should prioritize clinically meaningful target populations and demonstrate improvements in patient outcomes and antimicrobial stewardship beyond conventional measures of predictive performance.

antimicrobial resistance

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis.

PURPOSE: Given the existing uncertainties regarding the link between Di(2-ethylhexyl) phthalate (DEHP) exposure and gastric cancer (GC) progression, this study aimed to clarify their association, identify the toxic targets of DEHP, and elucidate the underlying molecular mechanisms. METHODS: Multiple integrated approaches were employed, including Gene Expression Omnibus (GEO) data analysis, network toxicology, molecular docking, and machine learning. STRING and Cytoscape tools were utilized to identify key targets, while Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the functional enrichment of intersecting targets. Machine learning and SHAP analysis were applied to screen core targets in GC. Molecular docking was performed to evaluate the binding affinity of DEHP toward core targets, and 200 ns molecular dynamics simulations were further conducted for representative complexes to validate their dynamic stability. RESULTS: A total of 18 key targets were identified using STRING and Cytoscape. GO and KEGG enrichment analyses demonstrated that these intersecting targets were primarily enriched in the extracellular region, as well as the Calcium signaling pathway and cAMP signaling pathway. Through machine learning analyses, 7 key genes (ADRB2, ESRRG, GRIA4, IL13RA2, NR3C2, PLA2G1B, and SULT2A1) were identified as core targets in GC through machine learning analyses. Molecular docking simulations revealed strong binding specificity between DEHP and the target proteins. Among them, NR3C2 and ADRB2 exhibited relatively high predictive importance in the machine learning models. DEHP showed favorable binding affinity toward these core targets, and molecular dynamics simulations further confirmed that ADRB2-DEHP and NR3C2-DEHP complexes maintained stable conformations throughout the simulation. CONCLUSIONS: Our findings identified GC associated genes that were computationally predicted as potential targets of DEHP. These results indicated structural compatibility between DEHP and its target proteins but did not prove that DEHP exposure accounts for the gene expression changes in GC.

Molecular Docking Simulation

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning

DNA methylation and machine learning: challenges and perspective toward enhanced clinical diagnostics.

DNA methylation is an epigenetic modification that regulates gene expression by adding methyl groups to DNA, affecting cellular function and disease development. Machine learning, a subset of artificial intelligence, analyzes large datasets to identify patterns and make predictions. Over the past two decades, advances in bioinformatics technologies for arrays and sequencing have generated vast amounts of data, leading to the widespread adoption of machine learning methods for analyzing complex biological information for medical problems. This review explores recent advancements in DNA methylation studies that leverage emerging machine learning techniques for more precise, comprehensive, and rapid patient diagnostics based on DNA methylation markers. We present a general workflow for researchers, from clinical research questions to result interpretation and monitoring. Additionally, we showcase successful examples in diagnosing cancer, neurodevelopmental disorders, and multifactorial diseases. Some of these studies have led to the development of diagnostic platforms that have entered the global healthcare market, highlighting the promising future of this field.

Humans

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient = - 1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10 years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung

Integrative machine learning models to unravel gut microbial dysbiosis and functional disruption in polycystic ovary syndrome.

OBJECTIVE: To study gut microbial diversity and metabolic pathway disruptions in women with PolyCystic Ovary Syndrome (PCOS) compared with healthy controls, and to evaluate the diagnostic potential of microbiome-driven machine learning models. DESIGN: Case-controlled metagenomic data analysis SUBJECTS: Gut metagenomic data from women diagnosed with PCOS and age-matched healthy female controls EXPOSURE: Presence of PCOS MAIN OUTCOME MEASURES: The primary outcome measures will include gut microbial alpha and beta diversity indices, microbial taxon abundance, functional pathway profiles, predicted metabolite levels, microbe-functional pathway-metabolite interaction networks, and the diagnostic accuracy of microbiome-based machine learning models. RESULTS: Alpha and beta diversity analyses revealed marked gut microbial dysbiosis in women with PCOS, despite comparable species richness to healthy controls. Differential abundance analysis identified 41 significantly altered microbial species, including enrichment of proinflammatory taxa, such as Bacteroides vulgatus and Ruminococcus gnavus, and depletion of beneficial commensals, including Roseburia hominis and Prevotella copri. These compositional shifts indicate a proinflammatory microbial community structure in PCOS. Functional profiling demonstrated the upregulation of pathways involved in nucleotide turnover, lipid and carbohydrate metabolism, and neurotransmitter synthesis, potentially contributing to metabolic and neuroendocrine disruption. Network analysis revealed fragmented and unstable microbial-metabolite associations in PCOS compared with cohesive networks in controls. Microbiome-based machine learning models achieved a diagnostic accuracy of 84.25% (area under the curve 0.93), underscoring their predictive potential. CONCLUSION: The gut microbiome in PCOS is characterized by a proinflammatory community structure and disrupted metabolic pathways. These findings demonstrate the diagnostic potential of microbiome-based models and underscore the gut microbiome as a promising target for therapeutic interventions in the management of PCOS.

Polycystic Ovary Syndrome

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

Gut microbiota-derived metabolites target C5AR1/KDM2A/HCAR3 axis in inflammatory bowel disease: a multi-machine learning algorithms and molecular docking study.

BACKGROUND: Inflammatory bowel disease (IBD) is a chronic recurrent disorder. Gut microbiota-derived metabolites regulate intestinal homeostasis, but their molecular mechanisms in IBD remain unclear. Current studies lack systematic "microbiota-metabolite-target" network mining with multi-method validation. This study integrates network pharmacology, three machine learning algorithms, and molecular docking to construct this regulatory network in IBD. METHODS: Transcriptome data were obtained from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified using limma (p < 0.05, |log2FC| > 0.5). Weighted gene co-expression network analysis (WGCNA) with an optimal soft threshold of &#x3b2; = 7 was performed to identify key module genes. Candidate genes were obtained by intersecting DEGs, gut microbiota-associated genes from the gutMGene database, and WGCNA module genes. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to explore the functional roles of candidate genes. Core genes were identified using three machine learning algorithms (LASSO, Boruta, and SVM-RFE), followed by protein-protein interaction (PPI) network analysis. Molecular docking was performed to assess the binding affinities between hub proteins and gut microbiota-derived metabolites. RESULTS: A total of 885 DEGs were identified between the IBD and control groups, including 463 upregulated and 422 downregulated genes. WGCNA identified 280 key module genes from the purple and yellow modules. The intersection of DEGs, gut microbiota-associated genes, and WGCNA module genes yielded 19 core candidate genes. PPI network analysis combined with three machine learning algorithms jointly identified C5AR1, KDM2A, and HCAR3 as core hub genes. ROC curve analysis demonstrated that all three hub genes achieved AUC values greater than 0.7 in both the training and validation sets, indicating excellent diagnostic performance for IBD. Enrichment analysis revealed significant associations with the TNF, NF-&#x3ba;B, and IL-17 signaling pathways. Molecular docking confirmed stable binding of C5AR1 with 1,3-Diphenylpropan-2-Ol (-7.87 &#xb1; 0.83 kcal&#xb7;mol-&#xb9;) and HCAR3 with 3-Indolepropionic Acid (-6.35 &#xb1; 0.70 kcal&#xb7;mol-&#xb9;), both below -5.0 kcal&#xb7;mol-&#xb9;. CONCLUSION: This study first constructs a "gut microbiota-metabolite-hub gene" axis in IBD, providing a computational framework for microbiota-targeted precision therapy, and identifying C5AR1/KDM2A/HCAR3 as computationally predicted diagnostic biomarkers and 1,3-Diphenylpropan-2-Ol/3-Indolepropionic Acid as candidate intervention molecules that warrant further experimental validation.

Molecular Docking Simulation

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning

Machine Learning and Metabolomics to Characterize Warburg-Like Metabolic Subtypes in Human Retinal Endothelial Cells Exposed to Risk Factors Associated With Proliferative Diabetic Retinopathy.

PURPOSE: High glucose (HG), hypoxia (Hyp), and their combination are major risk factors for proliferative diabetic retinopathy (PDR). Although these conditions induce features of the Warburg-like metabolic reprogramming in human retinal endothelial cells (HRECs), it remains unclear whether they produce distinct metabolic and angiogenic subtypes. This study aimed to characterize the Warburg-like-associated metabolic heterogeneity induced by these PDR-related risk factors and evaluate the ability of supervised machine-learning models to distinguish these subtypes. METHODS: HRECs were cultured under normoglycemic, HG, Hyp (2% O2), and combined HG-Hyp conditions. Untargeted LC-MS/MS metabolomics quantified metabolites spanning carbohydrates, amino acids, nucleotides, and lipids. Principal component analysis (PCA) assessed overall metabolic variation, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis identified metabolic pathways associated with angiogenesis. In vitro angiogenesis assays measured endothelial tube formation and branching. Nine supervised classifiers (decision tree, logistic regression, na&#xef;ve Bayes, random forest, K-Nearest Neighbors, neural network, gradient boosting, AdaBoost, and Support Vector Machine) were trained on the highest-ranked metabolites selected by the Information Gain Ratio feature-ranking approach. Model performance was evaluated using 10-fold cross-validation, leave-one-out cross-validation (LOOCV), permutation testing, and a classifier stability analysis under biologically meaningful distributional shift using an independent chemically induced hypoxia model (CoCl2). RESULTS: PCA revealed partial separation of metabolic profiles across conditions, indicating different Warburg-like metabolic subtypes. The combined HG-Hyp condition exhibited enhanced angiogenic potential relative to either HG or Hyp alone. KEGG pathway enrichment analysis identified fatty acid biosynthesis and elongation among the most significantly enriched pathways in HRECs under combined HG-Hyp conditions, alongside amino sugar and nucleotide sugar metabolism, glycerophospholipid metabolism, the pentose phosphate pathway, and glycolysis/gluconeogenesis. Supervised machine-learning classifiers distinguished these metabolic subtypes, with AdaBoost and gradient Boosting showing the most balanced, reproducible performance across 10-fold cross-validation, LOOCV, and permutation testing, and remaining the most reliable classifiers under domain-shift testing (area under the curve = 0.88, P = 0.0061). CONCLUSIONS: In this exploratory analysis, HG, Hyp, and their combination drive metabolically and functionally distinct subtypes of Warburg-like metabolic reprogramming in HRECs, with HG-Hyp in combination producing a highly angiogenic phenotype. Boosting-based ensemble classifiers provide a promising framework for detecting these subtypes even under domain-shift conditions, warranting validation in larger independent datasets. TRANSLATIONAL RELEVANCE: Integrating metabolomics with machine-learning classification offers a strategy to identify Warburg-like metabolic subtypes in retinal endothelial cells, providing insights into angiogenic mechanisms and guiding the development of targeted diagnostics or therapeutics for PDR.

Humans

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq