PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Study Protocol for HeartMagic: A Prospective Observational Cohort Characterizing Subtypes of Heart Failure With Preserved Ejection Fraction.

BACKGROUND: Heart failure (HF) is a life-threatening syndrome with significant morbidity and mortality. Although evidence-based drug treatments have effectively reduced morbidity and mortality in HF with reduced ejection fraction (EF), few therapies have been demonstrated to improve outcomes in HF with preserved EF. This may be caused by the existence of several HF with preserved EF subtypes that each need different treatments. There is therefore an unmet need for a comprehensive approach to subtype patients with HF with preserved EF. This protocol details the approach employed in the HeartMagic (Heart Failure Studied With a Machine Learning, Genomics, and Imaging Combination) study to address this gap. METHODS: This prospective multicenter observational cohort study will include 500 consecutive patients with HF with preserved EF at 2 Swiss university hospitals, along with 50 age-matched patients with HF with reduced EF and 50 healthy controls. In addition to routine clinical workup, participants undergo genomic, transcriptomic, and metabolomic analyses, and the anatomy, composition, and function of the heart are quantified by comprehensive echocardiography and magnetic resonance imaging. Quantitative magnetic resonance imaging is also applied to characterize the kidney. The primary outcome is a composite of 1-year cardiovascular mortality or rehospitalization. Machine learning-based multimodal clustering will be employed to identify distinct HF with preserved EF subtypes. Statistical analysis will include group comparisons, survival analysis, and integrative multimodal clustering combining clinical, imaging, ECG, genomic, transcriptomic, and metabolomic data to identify and validate HF with preserved EF subtypes. CONCLUSIONS: The integration of comprehensive magnetic resonance imaging with extensive genomic and metabolomic profiling in this study will result in an unprecedented panoramic view of HF with preserved EF and help distinguish functional subgroups, which may provide a basis for personalized therapies.

Aged↗

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly popular approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n&#x2009;=&#x2009;1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy&#x2009;=&#x2009;89%) using 3728 features and MoRSAE (accuracy&#x2009;=&#x2009;84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta&#x2009;=&#x2009;0.6839, p=0.006), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta&#x2009;=&#x2009;1.92; MoRS, beta&#x2009;=&#x2009;1.99 and MoRSAE, beta&#x2009;=&#x2009;1.77) displayed a significant (p&#x2009;<&#x2009;0.001) predictive power for post-deployment PTSD. CONCLUSION: The inclusion of exposure variables adds to the predictive power of MRS. Classification-based MRS may be useful in predicting risk of future PTSD in populations with anticipated trauma exposure. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting PTSD and, relatedly, improve their performance in independent cohorts.

Humans↗

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts &#x223c;50&#x2009;000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship↗

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic↗

Identification of autophagy-related genes as potential biomarkers correlated with immune infiltration in bipolar disorder: a bioinformatics analysis.

BACKGROUND: Bipolar disorder (BPD) is a kind of manic and depressive phase alternate episodes of serious mental illness, and it is correlated with well-documented cortical brain abnormalities. Emerging evidence supports that autophagy dysfunction in neuronal system contributes to pathophysiological changes in neurological disease. However, the role of autophagy in bipolar disorder has rarely been elucidated. This study aimed to identify the autophagy-related gene as a potential biomarker Correlated to immune infiltration in BPD. METHODS: The microarray dataset GSE23848 and autophagy-related genes (ARGs) were downloaded. Differentially expressed genes (DEGs) between normal and BPD samples were screened using the R software. Machine learning algorithms were performed to screen the significant candidate biomarker from autophagy-related differentially expressed genes (ARDEGs). The correlation between the screened ARDEGs and infiltrating immune cells was explored through correlation analysis. RESULTS: In this study, the autophagy pathway was abundantly enriched and activated in BPD, as indicated by Pathway enrichment analysis. We identified 16 ARDEGs in BPD compared to the normal group. A signature of 4 ARDEGs (ERN1, ATG3, CTSB, and EIF2AK3) was screened. ROC analysis showed that the above genes have good diagnostic performance. In addition, immune correlation analysis considered that the above four genes significantly correlated with immune cells in BPD. CONCLUSIONS: Autophagy - immune cell axis mediates pathophysiological changes in BPD. Four important ARDEGs are prospective to be potential biomarkers associated with immune infiltration in BPD and helpful for the prediction or diagnosis of BPD.

Bipolar Disorder↗

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans↗

Rapid glycomic analysis of serum EVs reveals altered N-glycosylation patterns in ASD.

Objective laboratory diagnostics for autism spectrum disorder (ASD) are lacking, necessitating rapid clinical screening tools. Because serum extracellular vesicle (EV) N-glycosylation captures critical neurodevelopmental signatures, we developed a fast, biologically interpretable diagnostic strategy. EVs from ASD patients with language impairment and neurotypical controls were isolated using a rapid extra-polyethylene glycol precipitation/filtration (EPF) workflow, benchmarked against ultracentrifugation. Following MALDI-TOF/MS profiling, machine learning was re-evaluated using repeated nested cross-validation to reduce optimistic bias and potential information leakage. Among five classifiers, Random Forest (RF) showed the best overall balance across discrimination, calibration, and classification metrics. RF-based SHAP analysis provided transparent interpretation, highlighting key discriminative glycans, including H4N3S1F1, H5N5S1F1, and H3N5F1. To elucidate molecular mechanisms, we integrated public EV transcriptomic data. This revealed significant dysregulation of N-glycosylation machinery genes (e.g., MAN1A1, NEU1, OSTC, RPN2), whose expression directionally aligned with observed glycan shifts in synaptic pathways. Collectively, this rapid serum EV N-glycomic workflow, combined with leakage-controlled RF-based interpretation, provides a promising foundation for non-invasive ASD biomarker discovery and future multicenter validation.

Humans↗

miRNA Target Prediction: An Overview of the Past and Current Tools.

MicroRNAs (miRNAs) are among the most studied molecules in recent years, and since their discovery, many miRNAs have been identified across various species. As members of the non-coding RNA family, miRNAs are key players in post-transcriptional gene regulation. These molecules can inhibit translation or promote degradation of messenger RNA (mRNA) by binding to the 3' untranslated region (UTR) of mRNA, thereby influencing almost all biological processes. To identify a miRNA's biological role, it is essential to predict the target sites to which it binds, a goal made possible through bioinformatics tools. This chapter discusses the bioinformatics tools commonly used for this purpose. Also, it analyzes the main factors considered in target prediction, such as seed match, free energy, conservation, site accessibility, multiple binding site contribution, and machine learning and deep learning approaches. Understanding the principles underlying these predictive methodologies is crucial for advancing one's biological research on miRNAs.

MicroRNAs↗

AI-driven diagnostic and prognostic models for metabolic dysfunction-associated steatotic liver disease: insights from clinical, imaging, and multi-omics studies-a scoping review.

Metabolic dysfunction-associated steatotic liver disease (MASLD), formerly known as non-alcoholic fatty liver disease (NAFLD), is the most common chronic liver disease around the world, affecting 33.6% of the adult population (95% CI: 28.1%-39.5%; I 2&#x2009;=&#x2009;99.9%), or roughly one in three. The extent of the liver damage is variable, from simple steatosis to metabolic dysfunction-associated steatohepatitis (MASH, formerly NASH), cirrhosis and hepatocellular carcinoma (HCC). Early diagnosis is essential to prevent serious liver damage. Traditional diagnostic techniques such as liver biopsy, imaging, and biomarker testing are all invasive, costly, reduced sensitive to early-stage disease, and they also have variability among observers. Modern diagnostic and prognostic approaches based on the principles of Artificial Intelligence (AI) and specifically on machine learning (ML) and deep learning (DL) have enabled multimodal approaches integrating clinical, imaging and molecular data. This scoping review conducted per PRISMA-ScR guidelines, synthesizes findings from 73 studies (search window 2020-2026) across three dimensions: clinical data driven models, imaging-based classifiers (ultrasound, CT and MRI), and multi-omics (genomics, transcriptomics and proteomics) techniques. Moreover, emergence of models such as U-Net and LiverNet 2.x, classification models like DeepLiverNet and BiLSTM models, as well as transformer frameworks and the identification of biomarkers models are also described. This study also investigates challenges such as data heterogeneity, data interpretability, fairness and real-world clinical application. Finally, important areas of research opportunities and future directions are highlighted to present a developing clinically applicable, explainable and ethical AI solutions to manage MASLD.

MASLD↗

PGS-GS: a framework integrating polygenic scores and genomic selection in animal breeding.

Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.

Animals↗

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV&#xa0;>&#xa0;100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation↗

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms↗

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals↗

DNA Methylation-Based Classification of Kidney Neoplasms.

Renal neoplasms are morphologically and molecularly heterogeneous, with their diagnosis often hindered by interobserver variability and overlapping microscopic features. A subset of cases is unclassifiable despite immunohistochemical, mutation, and cytogenetic-based diagnostic workup. Through examination of the genome-wide DNA methylation signatures of over 2000 renal neoplasms, we identified 23 coherent groups that correlate with known neoplasm types and identified novel clinically relevant subtypes of existing neoplasm types. We used machine learning models to develop and validate a classifier trained on DNA methylation profiles of 1284 samples. The classifier was tested on an external data set of 287 renal neoplasms with >90% concordance between expected neoplasm type and high-score DNA methylation-based classification. Discordance between the original histologic label and methylation class led to potential reclassification of some cases. This work demonstrates proof of principle for the feasibility of a DNA methylation classifier as a clinically useful tool to assist in the diagnosis of renal neoplasms.

Humans↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

Minimizing Off-Target Effects of CRISPR-Cas9 With Optimized sgRNA: Evaluation of Efficiency and Specificity in the Tumor Protein 53 (TP53) Region.

CRISPR-Cas9 is a widely used genetic tool with therapeutic potential in molecular biology. CRISPR-Cas9 enables precise genome editing by its ability to target specific DNA sequence. After off-target and on-target regions are identified, CRISPR-Cas9 is applied to these regions based on the match between the guide RNA (gRNA) and target DNA sequence. This study points to the off-target impact of mismatches between the gRNA and target DNA on exon regions of the TP53 gene, which are involved in regulating multiple genes and cellular functions. Off-target positions are typically evaluated using scoring methods. In this study, we have used latent class analysis to reveal subclasses of off-target positions. Thus, we have created the levels of off-target positions and evaluated the effects of mismatching positions within these classes using machine learning classifiers. The results revealed that mismatching positions could be categorized into three levels: low, middle, and high off-target positions. We have improved a computational framework to minimize off-target effects and to identify the PAM sequences in the gRNA design. Thus, carefully designed gRNAs will ensure that desired genetic edits are performed and target variants are achieved. This work will avail the future research aimed at optimizing genome editing by customizing CRISPR-Cas9 to target specific protospacer DNA through gRNA.

CRISPR-Cas Systems↗

Genetic mapping and predictive modeling of paralog synthetic lethality.

Paralogs are abundant in the human genome and thought to be a primary source of synthetic lethality, yet the vast paralogome remains largely uncharacterized. A digenic screen of 36,648 paralogous pairs in the human genome revealed that synthetic lethalities were infrequent and varied in penetrance in different tumor backgrounds. We hypothesized that the variable penetrance of synthetic lethalities resulted from complex polygenic interactions with different cellular contexts. A machine learning classifier of a subset of paralog pairs tested across 49 cancer models revealed that endogenous perturbations in related pathways predicted paralog synthetic lethality. Further, predictive modeling of paralog synthetic lethality showed that the strength of synthetic lethal interactions was largely due to the overlap and essentiality of the protein-protein interaction networks shared by the paralog pairs. Collectively, this study tested 36,648 digenic paralog interactions and delineated the key feature classes that underlie the heterogeneity of paralog synthetic lethalities.

Humans↗