PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning (ML)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans↗

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI↗

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans↗

HINN: Hierarchical Input Neural Network identifies multi-omics biomarker for cognitive decline.

Understanding complex diseases requires models that can integrate diverse layers of biological data while yielding insights that are biologically interpretable. Although multi-omics integration with machine learning (ML) has advanced disease prediction and biomarker discovery, most existing approaches overlook the hierarchical and regulatory relationships that connect these molecular layers. Here, we present the Hierarchical Input Neural Network (HINN), a deep learning framework that incorporates known cross-omics relationships directly into its architecture, capturing the flow of information from genomics to epigenomics, transcriptomics, and downstream biological processes. By embedding these relationships, HINN improves both predictive performance and biological interpretability. We applied HINN to blood-derived multi-omics data from individuals with Alzheimer's disease or mild cognitive impairment to predict cognitive scores from standardized assessments. HINN outperformed both baseline and state-of-the-art models and pinpointed multi-omics biomarkers-including SNPs and promoter-region CpG sites in ATP6V1C1 and RCHY1 -that were significantly correlated with plasma p-Tau181 levels. These features map to biologically relevant processes with potential implications for cognitive decline. Our findings demonstrate how combining deep learning with biological knowledge can uncover interpretable, blood-based biomarkers for cognitive decline due to complex diseases such as Alzheimer's. All code and data are openly available at https://github.com/bozdaglab/HINN.

Alzheimer’s disease↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

A support for decision-making: cost-sensitive learning system.

This paper investigates a machine learning (ML) algorithm for supporting a decision-making system that is able to handle diagnostic problems. The input data are expressed by solved cases of patients' diagnoses, and the output is formed by a set of decision rules which may be directly exploited for a decision support. We have chosen the methodology of covering ML algorithms, namely the CN2 algorithm, as a starting point, and designed and implemented a certain extension of CN2 that comprises: advanced discretizing numerical attributes and incorporating attribute cost to economize the classification.

Algorithms↗

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry↗

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products↗

Data-driven approaches in green microbiology: strategies for plant growth-promoting bacteria.

Plant growth-promoting bacteria (PGPB) are gaining attention as scalable biological solutions to enhance crop productivity and resilience. However, accurately identifying and characterizing PGPB remains challenging, particularly under variable environmental conditions where microbial functions are context-dependent and shaped by complex plant-microbe interactions. Advances in high-throughput sequencing have shifted the field from culture-dependent approaches to genome-informed strategies, enabling large-scale taxonomic and functional profiling. Although trait-based databases support the prediction of plant-beneficial genes, they capture only a fraction of the underlying biological complexity and often require labor-intensive analyses. Machine learning (ML) and deep learning (DL) have emerged as powerful tools to integrate genomic, physiological, and ecological data, enabling the prioritization of candidate strains with plant growth-promoting potential. To evaluate advances in the field, we conducted a systematic review of studies integrating ML and DL with PGPB characterization, assessing algorithm selection, performance, and target plant systems. Across 248 observations, only 6.0% of studies directly addressed PGPB screening, whereas the majority (77.4%) focused on plant disease detection, revealing a substantial gap in the application of AI to beneficial microorganisms for plant growth. Convolutional neural networks (CNNs) were the most frequently applied algorithms, largely driven by image-based phenotyping tasks. Overall, the field is constrained by limited datasets, high computational demands, and challenges in modeling multispecies and host-associated interactions. We highlight the need for integrative and interpretable ML and DL frameworks that bridge genomic data and functional validation. Such approaches represent a promising path toward scalable, data-driven discovery and deployment of bioinoculants in sustainable agriculture.

Agriculture↗

Machine learning in sedimentation modelling.

The paper presents machine learning (ML) models that predict sedimentation in the harbour basin of the Port of Rotterdam. The important factors affecting the sedimentation process such as waves, wind, tides, surge, river discharge, etc. are studied, the corresponding time series data is analysed, missing values are estimated and the most important variables behind the process are chosen as the inputs. Two ML methods are used: MLP ANN and M5 model tree. The latter is a collection of piece-wise linear regression models, each being an expert for a particular region of the input space. The models are trained on the data collected during 1992-1998 and tested by the data of 1999-2000. The predictive accuracy of the models is found to be adequate for the potential use in the operational decision making.

Algorithms↗

Precision medicine in combating antimicrobial resistance: A comprehensive review.

Antimicrobial resistance (AMR) represents one of the most pressing threats to global public health, undermining the effectiveness of modern antimicrobial therapy and challenging decades of medical progress. This comprehensive review examines the transition from broad-spectrum empirical therapy toward precision medicine as an integrated framework for improving antimicrobial use and combating AMR. Precision medicine seeks to tailor treatment decisions by combining pathogen-specific genomic and resistance data with relevant host characteristics to optimize therapy while limiting unnecessary antimicrobial exposure and the selective pressures that drive resistance. The review synthesizes advances reported from 2020, highlighting established and emerging approaches including rapid molecular diagnostics, next-generation sequencing, CRISPR-based detection, machine learning (ML)-assisted decision support, precision dosing, and targeted therapeutics such as bacteriophage therapy, antimicrobial peptides, and bacterial proteolysis-targeting chimeras. Rather than functioning as isolated technologies, these approaches achieve their greatest clinical value when integrated within antimicrobial stewardship programs and a One Health framework that recognizes the interconnected human, animal, and environmental drivers of resistance. Despite considerable progress, important challenges remain, including equitable access to advanced technologies, interpretation of increasingly complex datasets, workforce and infrastructure limitations, and evolving regulatory pathways for novel diagnostics and therapeutics. This review concludes that while precision medicine is not a standalone solution, its successful implementation will depend on coordinated integration of diagnostics, host factors, computational tools, pharmacological optimization, and stewardship strategies to improve patient outcomes while preserving the long-term effectiveness of existing antimicrobials.

Antimicrobial resistance↗

A method for automated temporal knowledge acquisition applied to sleep-related breathing disorders.

This paper presents a method for the discovery of temporal patterns in multivariate time series and their conversion into a linguistic knowledge representation applied to sleep-related breathing disorders. The main idea lies in introducing several abstraction levels that allow a step-wise identification of temporal patterns. Self-organizing neural networks are used to discover elementary patterns in the time series. Machine learning (ML) algorithms use the results of the neural networks to automatically generate a rule-based description. At the next levels, temporal grammatical rules are inferred. This method covers one of the main "bottlenecks" in the design of knowledge-based systems, namely, the knowledge acquisition problem. An evaluation of the rules lead to an overall sensitivity of 0.762, and a specificity of 0.758.

Artificial Intelligence↗

Explaining the output of ensembles in medical decision support on a case by case basis.

The use of ensembles in machine learning (ML) has had a considerable impact in increasing the accuracy and stability of predictors. This increase in accuracy has come at the cost of comprehensibility as, by definition, an ensemble model is considerably more complex than its component models. This is of significance for decision support systems in medicine because of the reluctance to use models that are essentially black boxes. Work on making ensembles comprehensible has so far focused on global models that mirror the behaviour of the ensemble as closely as possible. With such global models there is a clear tradeoff between comprehensibility and fidelity. In this paper, we pursue another tack, looking at local comprehensibility where the output of the ensemble is explained on a case-by-case basis. We argue that this meets the requirements of medical decision support systems. The approach presented here identifies the ensemble members that best fit the case in question and presents the behaviour of these in explanation.

Anticoagulants↗

Deep assessment of machine learning techniques using patient treatment in acute abdominal pain in children.

Learning from patient records may aid knowledge acquisition and decision making. Existing inductive machine learning (ML) systems such us NewId, CN2, C4.5 and AQ15 learn from past case histories using symbolic and/or numeric values. These systems learn symbolic rules (IF... THEN like) which link an antecedent set of clinical factors to a consequent class or decision. This paper compares the learning performance of alternative ML systems with each other and with respect to a novel approach using logic minimization, called LML, to learn from data. Patient cases were taken from the archives of the Paediatric Surgery Clinic of the University Hospital of Crete, Heraklion, Greece. Comparison of ML system performance is based both on classification accuracy and on informal expert assessment of learned knowledge.

Abdomen, Acute↗

A genomic catalog of Earth's bacterial and archaeal symbionts.

Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15-23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host-microorganism associations.

Journal Article↗

A new perspective on V3 phenotype prediction.

The particular coreceptor used by a strain of HIV-1 to enter a host cell is highly indicative of its pathology. HIV-1 coreceptor usage is primarily determined by the amino add sequences of the V3 loop region of the viral envelope glycoprotein. The canonical approach to sequence-based prediction of coreceptor usage was derived via statistical analysis of a less reliable and significantly smaller data set than is presently available. We aimed to produce a superior phenotypic classifier by applying modern machine learning (ML) techniques to the current database of V3 loop sequences with known phenotype. The trained classifiers along with the sequence data are available for public use at the supplementary website: http://genomiac2.ucsd.edu:8080/wetcat/v3.html and http://www.cs.waikato.ac.nz/ml/weka[corrected].

Amino Acid Sequence↗

Digital pathology and spatial omics in steatohepatitis: Clinical applications and discovery potentials.

Steatohepatitis with diverse etiologies is the most common histological manifestation in patients with liver disease. However, there are currently no specific histopathological features pathognomonic for metabolic dysfunction-associated steatotic liver disease, alcohol-associated liver disease, or metabolic dysfunction-associated steatotic liver disease with increased alcohol intake. Digitizing traditional pathology slides has created an emerging field of digital pathology, allowing for easier access, storage, sharing, and analysis of whole-slide images. Artificial intelligence (AI) algorithms have been developed for whole-slide images to enhance the accuracy and speed of the histological interpretation of steatohepatitis and are currently employed in biomarker development. Spatial biology is a novel field that enables investigators to map gene and protein expression within a specific region of interest on liver histological sections, examine disease heterogeneity within tissues, and understand the relationship between molecular changes and distinct tissue morphology. Here, we review the utility of digital pathology (using linear and nonlinear microscopy) augmented with AI analysis to improve the accuracy of histological interpretation. We will also discuss the spatial omics landscape with special emphasis on the strengths and limitations of established spatial transcriptomics and proteomics technologies and their application in steatohepatitis. We then highlight the power of multimodal integration of digital pathology augmented by machine learning (ML)algorithms with spatial biology. The review concludes with a discussion of the current gaps in knowledge, the limitations and premises of these tools and technologies, and the areas of future research.

Humans↗

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article↗