PubMed HealthSearch

PubMed · 42477691

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

Abstract

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Salvatore D'Antona, Mawada Elmagboul Abdalla Abakar, Daniele Ramazzotti, Marco Antoniotti, Alex Graudenzi. 2026-07-20. Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.. https://doi.org/10.1186/s13040-026-00584-8

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Decoding gene regulation in plant genomes with artificial intelligence.

One of the central goals of plant functional genomics is to uncover regulatory mechanisms that shape agriculturally important traits to inform crop improvement. Recent advances in machine learning (ML) and artificial intelligence (AI), especially Large Language Models (LLMs), have greatly transformed our ability to derive regulatory information from complex genomics data. This review starts with a brief introduction of recent advances in AI and ML. We then present a plant-focused synthesis of emerging applications of AI- and LLM tools to: (i) predict epigenomic features, regulatory DNA elements, and gene expressions; (ii) infer gene regulatory network; and (iii) estimate post-transcriptional regulation.

Artificial intelligence

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

Development of a Computational Histology Artificial Intelligence-Powered Prognostic Biomarker in Colorectal Cancer in The Cancer Genome Atlas.

BACKGROUND: Risk stratification in colorectal cancer (CRC) plays an important role in treatment decision-making. As such, prognostic biomarkers that can augment risk stratification have clinical value. Quantitative histologic features from routine hematoxylin and eosin (H&E)-stained whole slide images (WSIs) provide a novel avenue for biomarker discovery. In this study, we explored the potential for a computational histology artificial intelligence (CHAI) platform to develop and validate a prognostic biomarker in CRC. METHODS: The Cancer Genome Atlas Colorectal Adenocarcinoma project was utilized for this study, with inclusion of all subjects (stage I-IV) with available digitized H&E specimens. The cohort was split into development and validation cohorts by a stratified random split. The previously developed CHAI platform was applied in the development cohort to construct a continuous risk score from histologic features associated with progression-free interval (PFI) that was dichotomized based on an optimized cutpoint for distinguishing PFI into a high risk CHAI (+) and lower risk CHAI (-). PFI was compared between CHAI (+) and CHAI (-) patients in the validation cohort in multivariable Cox proportional hazards models. Time-dependent area under the curve (tdAUC) and C-indices were also calculated for PFI. RESULTS: A total of 583 participants were included in the study, with 409 assigned to the validation cohort. The CHAI biomarker classified 229 participants (56%) as CHAI (+) and 180 (44%) as CHAI (-) in the validation set. CHAI (+) participants had worse PFI in a multivariable analysis adjusting for available clinicopathologic variables (hazard ratio (HR) = 2.65; 95% confidence interval (CI), 1.63-4.30). TdAUC for the CHAI biomarker was 0.60 (95% CI, 0.53-0.67) at 12 months, 0.62 (0.55-0.69) at 36 months, and 0.67 (0.55-0.79) at 60 months; the C-index was 0.62 (95% CI, 0.58-0.67). CONCLUSIONS: The CHAI platform was used to develop a prognostic digital pathology biomarker in CRC. This demonstrates the feasibility and potential to apply this artificial intelligence-based digital pathology biomarker platform for risk stratification in CRC and supports its further study.

Artificial intelligence