PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV > 100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation↗

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms↗

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals↗

DNA Methylation-Based Classification of Kidney Neoplasms.

Renal neoplasms are morphologically and molecularly heterogeneous, with their diagnosis often hindered by interobserver variability and overlapping microscopic features. A subset of cases is unclassifiable despite immunohistochemical, mutation, and cytogenetic-based diagnostic workup. Through examination of the genome-wide DNA methylation signatures of over 2000 renal neoplasms, we identified 23 coherent groups that correlate with known neoplasm types and identified novel clinically relevant subtypes of existing neoplasm types. We used machine learning models to develop and validate a classifier trained on DNA methylation profiles of 1284 samples. The classifier was tested on an external data set of 287 renal neoplasms with >90% concordance between expected neoplasm type and high-score DNA methylation-based classification. Discordance between the original histologic label and methylation class led to potential reclassification of some cases. This work demonstrates proof of principle for the feasibility of a DNA methylation classifier as a clinically useful tool to assist in the diagnosis of renal neoplasms.

Humans↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

Minimizing Off-Target Effects of CRISPR-Cas9 With Optimized sgRNA: Evaluation of Efficiency and Specificity in the Tumor Protein 53 (TP53) Region.

CRISPR-Cas9 is a widely used genetic tool with therapeutic potential in molecular biology. CRISPR-Cas9 enables precise genome editing by its ability to target specific DNA sequence. After off-target and on-target regions are identified, CRISPR-Cas9 is applied to these regions based on the match between the guide RNA (gRNA) and target DNA sequence. This study points to the off-target impact of mismatches between the gRNA and target DNA on exon regions of the TP53 gene, which are involved in regulating multiple genes and cellular functions. Off-target positions are typically evaluated using scoring methods. In this study, we have used latent class analysis to reveal subclasses of off-target positions. Thus, we have created the levels of off-target positions and evaluated the effects of mismatching positions within these classes using machine learning classifiers. The results revealed that mismatching positions could be categorized into three levels: low, middle, and high off-target positions. We have improved a computational framework to minimize off-target effects and to identify the PAM sequences in the gRNA design. Thus, carefully designed gRNAs will ensure that desired genetic edits are performed and target variants are achieved. This work will avail the future research aimed at optimizing genome editing by customizing CRISPR-Cas9 to target specific protospacer DNA through gRNA.

CRISPR-Cas Systems↗

Genetic mapping and predictive modeling of paralog synthetic lethality.

Paralogs are abundant in the human genome and thought to be a primary source of synthetic lethality, yet the vast paralogome remains largely uncharacterized. A digenic screen of 36,648 paralogous pairs in the human genome revealed that synthetic lethalities were infrequent and varied in penetrance in different tumor backgrounds. We hypothesized that the variable penetrance of synthetic lethalities resulted from complex polygenic interactions with different cellular contexts. A machine learning classifier of a subset of paralog pairs tested across 49 cancer models revealed that endogenous perturbations in related pathways predicted paralog synthetic lethality. Further, predictive modeling of paralog synthetic lethality showed that the strength of synthetic lethal interactions was largely due to the overlap and essentiality of the protein-protein interaction networks shared by the paralog pairs. Collectively, this study tested 36,648 digenic paralog interactions and delineated the key feature classes that underlie the heterogeneity of paralog synthetic lethalities.

Humans↗

Identifying genes related to drug anticancer mechanisms using support vector machine.

In an effort to identify genes related to the cell line chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine was used to label genes into any of the five predefined anticancer drug mechanistic categories. Among dozens of unequivocally categorized genes, many were known to be causally related to the drug mechanisms. For example, a few genes were found to be involved in the biological process triggered by the drugs (e.g. DNA polymerase epsilon was the direct target for the drugs from DNA antimetabolites category). DNA repair-related genes were found to be enriched for about eight-fold in the resulting gene set relative to the entire gene set. Some uncharacterized transcripts might be of interest in future studies. This method of correlating the drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Antineoplastic Agents↗

Symbiogenesis in learning classifier systems.

Symbiosis is the phenomenon in which organisms of different species live together in close association, resulting in a raised level of fitness for one or more of the organisms. Symbiogenesis is the name given to the process by which symbiotic partners combine and unify, that is, become genetically linked, giving rise to new morphologies and physiologies evolutionarily more advanced than their constituents. The importance of this process in the evolution of complexity is now well established. Learning classifier systems are a machine learning technique that uses both evolutionary computing techniques and reinforcement learning to develop a population of cooperative rules to solve a given task. In this article we examine the use of symbiogenesis within the classifier system rule base to improve their performance. Results show that incorporating simple rule linkage does not give any benefits. The concept of (temporal) encapsulation is then added to the symbiotic rules and shown to improve performance in ambiguous/non-Markov environments.

Algorithms↗

Identification of Immune Response-Related Proteomic Biomarkers in Moyamoya Disease Using Serum Olink Proteomics.

Moyamoya disease, a rare chronic cerebrovascular disorder, requires invasive digital subtraction angiography (DSA) for diagnosis. This study employed high-throughput proteomics to identify plasma biomarkers for Moyamoya disease diagnosis. We conducted immunopanel analysis using the Olink platform to evaluate 92 immune-related proteins in plasma samples from 88 Moyamoya disease patients and 88 healthy controls. Key proteins were identified through differential expression analysis, GO, and KEGG enrichment analysis. A diagnostic model was constructed using LASSO regression, Boruta algorithm, and machine learning models including random forest and XGBoost. Validation of these proteins was performed using GEO external data sets, followed by prediction of potential therapeutic drugs and molecular docking validation through pharmacogenomic databases. A total of 44 differentially expressed proteins were identified through the Olink immunopanel, with 12 downregulated and 32 upregulated. GO and KEGG analyses revealed significant enrichment of these proteins in innate immune responses and signaling pathways such as NF-kB and MAPK. Through LASSO, random forest, and protein under-area analysis, four potential biomarkers for Moyamoya disease (MGMT, SIT1, PRDX1, TRAF2) were identified. A diagnostic model using these proteins showed the highest AUC value with the XGBoost model. Additionally, TRAF2 and PRDX1 exhibited significant expression differences in Moyamoya disease patients within the GEO data set. Our study revealed the immune landscape of Moyamoya disease, identified four biomarkers, and established a variety of diagnostic models.

Humans↗

Synthetic DNA barcodes identify singlets in scRNA-seq datasets and evaluate doublet algorithms.

Single-cell RNA sequencing (scRNA-seq) datasets contain true single cells, or singlets, in addition to cells that coalesce during the protocol, or doublets. Identifying singlets with high fidelity in scRNA-seq is necessary to avoid false negative and false positive discoveries. Although several methodologies have been proposed, they are typically tested on highly heterogeneous datasets and lack a priori knowledge of true singlets. Here, we leveraged datasets with synthetically introduced DNA barcodes for a hitherto unexplored application: to extract ground-truth singlets. We demonstrated the feasibility of our framework, "singletCode," to evaluate existing doublet detection methods across a range of contexts. We also leveraged our ground-truth singlets to train a proof-of-concept machine learning classifier, which outperformed other doublet detection algorithms. Our integrative framework can identify ground-truth singlets and enable robust doublet detection in non-barcoded datasets.

Algorithms↗

Support vector machine classification of 18F-FDG PET scans across subtypes of amyotrophic lateral sclerosis.

PURPOSE: While 18F-FDG PET imaging has demonstrated diagnostic value in people with Amyotrophic Lateral Sclerosis (PwALS) and group-level differences were identified between different disease subtypes (e.g., genetic and clinical variants), refining and validating a machine-learning-based subject-level diagnostic algorithm may improve the general applicability and reliability of 18F-FDG PET as a diagnostic tool in ALS. In this study, we employed support vector machines (SVM) to further explore the diagnostic potential of 18F-FDG PET in ALS, alongside its ability to classify between different genetic subtypes or clinical phenotypes. METHODS: 18F-FDG PET data of 36 healthy volunteers (HV), 25 people with ALS-mimicking diseases (Mimics), and 167 PwALS, grouped by genetic status (e.g., sporadic (sALS) or carrying a C9orf72 hexanucleotide repeat expansion (ALSC9orf72RE) and onset (bulbar or spinal) type, acquired with Biograph 'TruePoint' PET/CT scanner, were included in the study (Dataset 1). A second dataset of 183 PwALS and 31 Mimics acquired with Biograph 'HiRez' scanner was included as an independent cross-validation set (Dataset 2). PET images were spatially normalised to MNI space to fit linear SVMs with cross-validation. Only age-matched groups were considered to eliminate age-related effects. RESULTS: For Dataset 1, the linear SVM resulted in an average accuracy of 0.86 for the classification of ALS vs. HV, 0.53 for ALS vs. Mimics, 0.83 for ALSC9orf72RE vs. sALS, and 0.58 for bulbar vs. spinal onset. These findings were corroborated with Dataset2, with an accuracy of up to 0.76 for ALSC9orf72RE vs. sALS, and 0.59 for bulbar vs. spinal. CONCLUSION: 18F-FDG brain PET imaging, combined with SVM and age-matching, can distinguish between ALSC9orf72RE and sALS with good accuracy, but lacks sufficient discriminative power to differentiate between ALS and Mimics and between different sites of onset.

Humans↗

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding↗

Inferring metabolic objectives and trade-offs in single cells during embryogenesis.

While proliferating cells optimize their metabolism to produce biomass, the metabolic objectives of cells that perform non-proliferative tasks are unclear. The opposing requirements for optimizing each objective result in a trade-off that forces single cells to prioritize their metabolic needs and optimally allocate limited resources. Here, we present single-cell optimization objective and trade-off inference (SCOOTI), which infers metabolic objectives and trade-offs in biological systems by integrating bulk and single-cell omics data, using metabolic modeling and machine learning. We validated SCOOTI by identifying essential genes from CRISPR-Cas9 screens in embryonic stem cells, and by inferring the metabolic objectives of quiescent cells, during different cell-cycle phases. Applying this to embryonic cell states, we observed a decrease in metabolic entropy upon development. We further uncovered a trade-off between glutathione and biosynthetic precursors in one-cell zygote, two-cell embryo, and blastocyst cells, potentially representing a trade-off between pluripotency and proliferation. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis↗

Unlocking the Circulating Proteome: Toward Clinical Translation.

Blood-based proteomics is approaching a translational inflection point. Driven by advances in measurement technologies, rapid expansion of analytical capabilities, and growing adoption across research and medical communities, there is increasing demand for clinically actionable biomarkers. As the field transitions away from purely large-scale discovery-oriented studies toward more informed, targeted, application-driven analyses, the generation of proteomic data is no longer the bottleneck. Instead, the central challenge is to translate these measurements into robust, reproducible, and clinically meaningful insights. In this Review, we assess recent technological and methodological developments, evaluate persistent preanalytical and interpretative limitations, and outline the key steps required for clinical translation. We focus on three deeply interconnected dimensions: the capabilities and constraints of current measurement platforms, the role of computational and machine learning approaches in extracting biological and clinical signals, and the emergence of large-scale population studies that create new opportunities for validation and generalization. Finally, we discuss a forward-looking vision in which proteomics plays a central role in dynamic, multilayered omics frameworks, where integration with genomics, temporal profiling, and imaging can deepen our understanding of health, disease, and therapeutic response.

Humans↗

Antimicrobial resistance analysis of Klebsiella pneumoniae bloodstream infections based on a random forest algorithm: a longitudinal study based on data from tertiary hospitals in China from 2012 to 2023.

BACKGROUND: Bloodstream infections (BSIs) caused by Klebsiella pneumoniae pose a significant global health burden, complicated by rising antimicrobial resistance (AMR). This study aimed to characterize resistance patterns, identify predictors of carbapenem resistance, and develop a machine learning model to predict patient outcomes. METHODS: In a retrospective analysis of 109 279 K. pneumoniae BSIs from tertiary hospitals in China (2012-2023), 11&#x2009;000 isolates underwent whole-genome sequencing (WGS) and antimicrobial susceptibility testing. Cox proportional hazards and logistic regression models identified predictors of 30-day mortality and carbapenem-resistant K. pneumoniae (CRKP), respectively. A random forest model predicted AMR trends and outcomes, evaluated by accuracy, precision, recall, and ROC-AUC using R Studio (R Studio, Inc., Boston, MA, USA). RESULTS: Carbapenem resistance occurred in 32.3% of isolates, with rates of 41.9% for third-generation cephalosporins and 41.2% for fluoroquinolones. Among sequenced isolates, ST11 with blaKPC was the dominant CRKP genotype (12.0%). blaKPC (OR 3.97, 95% CI 3.10-5.11) and blaNDM (OR 2.80, 95% CI 2.07-3.71) strongly predicted carbapenem resistance; ICU admission predicted 30-day mortality (HR 2.10, 95% CI 1.80-2.46, p<0.001). Mortality was higher in CRKP (40.2%) vs. susceptible cases (21.5%). The random forest model achieved 89.2% accuracy and 0.92 ROC-AUC, with drug share, age, and CRKP status as top predictors. CONCLUSIONS: CRKP, especially ST11-blaKPC, drives excess mortality. Key predictors highlight the urgency for enhanced AMR surveillance and targeted therapy.

Humans↗

AI-driven CRISPR screening: optimizing gene editing through automation and intelligent decision support.

BACKGROUND: CRISPR-based genetic screening has become a central methodology in functional genomics, enabling systematic interrogation of gene function, genetic interactions and context-dependent vulnerabilities at scale. However, the rapid expansion of screening modalities-including multi-condition designs, combinatorial perturbations, in vivo applications and single-cell readouts-has exposed fundamental limitations of heuristic-driven experimental design and post hoc statistical analysis. MAIN BODY: This Review synthesizes how artificial intelligence is reshaping CRISPR screening by introducing predictive, adaptive and system-level intelligence across the experimental lifecycle. We organize recent advances into two tightly coupled modules. First, machine learning and deep learning (ML/DL) methods optimize experimental design by learning context-dependent perturbation behavior, anticipating confounding effects and enabling iterative, information-efficient screening strategies. Second, large language model-agent (LLM-agent) systems complement these advances by externalizing scientific reasoning, integrating biological knowledge at scale and coordinating analysis and decision-making in human-in-the-loop workflows. CONCLUSIONS: Together, ML/DL and LLM-agent approaches reframe CRISPR screening from a static analytical pipeline into an intelligent experimental system, with important implications for robustness, scalability and biological discovery.

Artificial Intelligence↗