PubMed HealthSearch

PubMed · 42556124

Computational metabolomics at scale: from open data to insight.

Abstract

Metabolomics data are currently generated at scale thanks to the evolution of technologies that have led to marked improvements in the number of metabolites detected, spanning all chemical classes. These data are increasingly submitted to public repositories for data reuse, integration, and interpretation. Despite the availability of public resources and associated computational tools, the field still lacks a widely adopted, consistent data and analytics infrastructure capable of transforming this wealth of information into scientific insight. Indeed, the metabolomics field is just now scratching the surface of being able to harness the power of new computational technologies. In this review, we summarize discussions from the "Dagstuhl-Seminar 24181 Computational Metabolomics: Towards Molecules, Models, and their Meaning" with a focus on public data availability, open data standards, data and knowledge integration, and education. Our goal is to raise awareness and adoption of the latest open science resources while highlighting key areas needing further development.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ewy A Mathé, Justin Jj van der Hooft, Haley Chatelaine, Louis-Félix Nothias, Stacey N Reinke, Juan Antonio Vizcaíno, Egon L Willighagen, Timothy Md Ebbels, Soha Hassoun. 2026-08-05. Computational metabolomics at scale: from open data to insight.. https://doi.org/10.1016/j.copbio.2026.103557

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Column switching liquid chromatography dual mass spectrometry system for simultaneous untargeted metabolomics and targeted exposomics.

Exposome-wide association studies (ExWAS) require the detection of metabolites and exposures with diverse chemical properties across wide concentration ranges, a task that typically demands multiple analytical methods. To address this challenge, we develop an integrated column-switching two-dimensional liquid chromatography-dual mass spectrometry (2DLC-dual-MS) system. This system employs a 2DLC setup to sequentially separate polar and non-polar compounds with log P ranging from -8 to 15. The separated fractions are directed via a three-way valve to a high-resolution MS (HRMS) and a triple quadrupole MS (TQMS), enabling simultaneous untargeted metabolome analysis and targeted quantification of 601 exposures. The method is particularly suited for the concurrent analysis of metabolome and exposome in human blood, where their concentrations typically differ by 2-3 orders of magnitude. In a demonstration application on lung adenocarcinoma ExWAS, the system exhibits good stability over more than 300 consecutive injections for both metabolome and exposome analysis, confirming its robustness for ExWAS applications.

Metabolomics

hypeR-GEM: connecting metabolite signatures to enzyme-coding genes via genome-scale metabolic models.

MOTIVATION: Enrichment analysis is a cornerstone of "omics" data interpretation, enabling researchers to connect analysis results to biological processes and generate testable hypotheses. Enrichment analysis in metabolomics poses distinct challenges for interpretation and multi-omics integration due to the lack of well-defined and consistent connections to well-curated gene-centered biological knowledge repositories. To address these challenges, we developed hypeR-GEM, a methodology and associated R package that adapts gene set enrichment analysis to metabolomics. hypeR-GEM leverages genome-scale metabolic models (GEMs) to infer reaction-based links between metabolites and enzyme-coding genes, enabling the mapping of metabolite signatures to gene signatures and their subsequent annotation via gene set enrichment analysis. RESULTS: We validated hypeR-GEM using paired metabolomics-proteomics and metabolomics-transcriptomics datasets by assessing whether genes mapped from metabolites significantly overlapped with differentially expressed proteins or transcripts. We further evaluated whether pathways enriched via hypeR-GEM-mapped genes corresponded to those derived from paired proteomic or transcriptomic data. In most datasets analyzed, both the predicted enzyme-coding genes and the associated enriched pathways showed significant concordance with independently derived omics signatures, supporting the utility and robustness of hypeR-GEM. Finally, we applied hypeR-GEM to the analysis of age-associated metabolic signatures from the New England Centenarian Study. The results revealed consistent enrichment of lipid-related pathways, aligning with the well-established role of lipid metabolism in aging, and highlighted additional pathways not captured in the metabolites' annotation, demonstrating hypeR-GEM's practical utility in a real-world use case. AVAILABILITY AND IMPLEMENTATION: The hypeR-GEM R package, documentation, and workflow examples are freely available at https://github.com/montilab/hypeR-GEM and archived at https://doi.org/10.5281/zenodo.20586748.

Metabolomics

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics