PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome scale models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets.

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Escherichia coli

gMISpy: integration of complex regulatory networks and genome scale metabolic models.

MOTIVATION: Genome-scale metabolic models lack explicit regulatory mechanisms, limiting their predictive accuracy for genetic interventions. Current methods for computing genetic Minimal Cut Sets either ignore regulatory networks entirely or use simplified acyclic representations that cannot capture regulatory feedback loops, ubiquitous features critical in cellular modeling. RESULTS: We developed gMISpy, a Python package that that enables efficient computation of genetic Minimal Intervention Sets (gMISs) in integrated genome-scale metabolic and regulatory networks. gMISpy incorporates cyclic regulatory logic into our previous computational framework using layered Boolean networks and BoNesis framework, resulting in a more accurate modeling of how regulatory interactions affect metabolic genes. Benchmarking across four different regulatory networks with Human-GEM showed consistent improvements in prediction accuracy, with Matthews correlation coefficient gains ranging from 2.50% to 14.42%. Validation against cancer data from DepMap and Project Score confirmed that cyclic integration reduces false positives and better captures biological vulnerabilities compared to acyclic approaches. AVAILABILITY AND IMPLEMENTATION: https://github.com/PlanesLab/cyclic-gMISpy.

Software

Evolution and applications of genome-scale metabolic models in yeast systems biology studies.

Genome-scale metabolic models (GEMs) can be used to simulate the metabolic network of an organism in a systematic and holistic way. Different yeast species, including Saccharomyces cerevisiae, have emerged as powerful cell factories for bioproduction. Recently, with the dedicated efforts from the scientific community, significant progress has been made in the development of yeast GEMs. Numerous versions of yeast GEMs and the derived multiscale models have been released, facilitating integrative omics analysis and rational strain design for different types of yeast cell factories. These advancements reflected the evolution and maturation of yeast GEMs together with a model ecosystem around them. This review will summarize the development and expansion of yeast GEMs and discuss their applications in yeast systems biology studies. It is anticipated that yeast GEMs will continue to play an increasingly important role in pioneering yeast physiological and metabolic studies in coming years.

Systems Biology

MiNEApy: enhancing enrichment network analysis in metabolic networks.

MOTIVATION: Modeling genome-scale metabolic networks (GEMs) helps understand metabolic fluxes in cells at a specific state under defined environmental conditions or perturbations. Elementary flux modes (EFMs) are powerful tools for simplifying complex metabolic networks into smaller, more manageable pathways. However, the enumeration of all EFMs, especially within GEMs, poses significant challenges due to computational complexity. Additionally, traditional EFM approaches often fail to capture essential aspects of metabolism, such as co-factor balancing and by-product generation. The previously developed Minimum Network Enrichment Analysis (MiNEA) method addresses these limitations by enumerating alternative minimal networks for given biomass building blocks and metabolic tasks. MiNEA facilitates a deeper understanding of metabolic task flexibility and context-specific metabolic routes by integrating condition-specific transcriptomics, proteomics, and metabolomics data. This approach offers significant improvements in the analysis of metabolic pathways, providing more comprehensive insights into cellular metabolism. RESULTS: Here, I present MiNEApy, a Python package reimplementation of MiNEA, which computes minimal networks and performs enrichment analysis. I demonstrate the application of MiNEApy on both a small-scale and a genome-scale model of the bacterium Escherichia coli, showcasing its ability to conduct minimal network enrichment analysis using minimal networks and context-specific data. AVAILABILITY AND IMPLEMENTATION: MiNEApy can be accessed at: https://github.com/vpandey-om/mineapy.

Metabolic Networks and Pathways

hypeR-GEM: connecting metabolite signatures to enzyme-coding genes via genome-scale metabolic models.

MOTIVATION: Enrichment analysis is a cornerstone of "omics" data interpretation, enabling researchers to connect analysis results to biological processes and generate testable hypotheses. Enrichment analysis in metabolomics poses distinct challenges for interpretation and multi-omics integration due to the lack of well-defined and consistent connections to well-curated gene-centered biological knowledge repositories. To address these challenges, we developed hypeR-GEM, a methodology and associated R package that adapts gene set enrichment analysis to metabolomics. hypeR-GEM leverages genome-scale metabolic models (GEMs) to infer reaction-based links between metabolites and enzyme-coding genes, enabling the mapping of metabolite signatures to gene signatures and their subsequent annotation via gene set enrichment analysis. RESULTS: We validated hypeR-GEM using paired metabolomics-proteomics and metabolomics-transcriptomics datasets by assessing whether genes mapped from metabolites significantly overlapped with differentially expressed proteins or transcripts. We further evaluated whether pathways enriched via hypeR-GEM-mapped genes corresponded to those derived from paired proteomic or transcriptomic data. In most datasets analyzed, both the predicted enzyme-coding genes and the associated enriched pathways showed significant concordance with independently derived omics signatures, supporting the utility and robustness of hypeR-GEM. Finally, we applied hypeR-GEM to the analysis of age-associated metabolic signatures from the New England Centenarian Study. The results revealed consistent enrichment of lipid-related pathways, aligning with the well-established role of lipid metabolism in aging, and highlighted additional pathways not captured in the metabolites' annotation, demonstrating hypeR-GEM's practical utility in a real-world use case. AVAILABILITY AND IMPLEMENTATION: The hypeR-GEM R package, documentation, and workflow examples are freely available at https://github.com/montilab/hypeR-GEM and archived at https://doi.org/10.5281/zenodo.20586748.

Metabolomics

RBC-GEM: A genome-scale metabolic model for systems biology of the human red blood cell.

Advancements with cost-effective, high-throughput omics technologies have had a transformative effect on both fundamental and translational research in the medical sciences. These advancements have facilitated a departure from the traditional view of human red blood cells (RBCs) as mere carriers of hemoglobin, devoid of significant biological complexity. Over the past decade, proteomic analyses have identified a growing number of different proteins present within RBCs, enabling systems biology analysis of their physiological functions. Here, we introduce RBC-GEM, one of the most comprehensive, curated genome-scale metabolic reconstructions of a specific human cell type to-date. It was developed through meta-analysis of proteomic data from 29 studies published over the past two decades resulting in an RBC proteome composed of more than 4,600 distinct proteins. Through workflow-guided manual curation, we have compiled the metabolic reactions carried out by this proteome to form a genome-scale metabolic model (GEM) of the RBC. RBC-GEM is hosted on a version-controlled GitHub repository, ensuring adherence to the standardized protocols for metabolic reconstruction quality control and data stewardship principles. RBC-GEM represents a metabolic network is a consisting of 820 genes encoding proteins acting on 1,685 unique metabolites through 2,723 biochemical reactions: a 740% size expansion over its predecessor. We demonstrated the utility of RBC-GEM by creating context-specific proteome-constrained models derived from proteomic data of stored RBCs for 616 blood donors, and classified reactions based on their simulated abundance dependence. This reconstruction as an up-to-date curated GEM can be used for contextualization of data and for the construction of a computational whole-cell models of the human RBC.

Humans

A genome-scale metabolic model for the denitrifying bacterium Thauera sp. MZ1T accurately predicts degradation of pollutants and production of polymers.

The denitrifying bacterium Thauera sp. MZ1T, a common member of microbial communities in wastewater treatment facilities, can produce different compounds from a range of carbon (C) and nitrogen (N) sources under aerobic and anaerobic conditions. In these different conditions, Thauera modifies its metabolism to produce different compounds that influence the microbial community. In particular, Thauera sp. MZ1T produces different exopolysaccharides with floc-forming properties, impacting the physical disposition of wastewater consortia and the efficiency of nutrient assimilation by the microbial community. Under N-limiting conditions, Thauera sp. MZ1T decreases its growth rate and accelerates the accumulation of polyhydroxyalkanoate-related (PHA) compounds including polyhydroxybutyrate (PHB), which plays a fundamental role as C and energy storage in this β-proteobacterium. However, the metabolic mechanisms employed by Thauera sp. MZ1T to assimilate and catabolize many of the different C and N sources under aerobic and anaerobic conditions remain unknown. Systems biology approaches such as genome-scale metabolic modeling have been successfully used to unveil complex metabolic mechanisms for various microorganisms. Here, we developed a comprehensive metabolic model (M-model) for Thauera sp. MZ1T (iThauera861), consisting of 1,744 metabolites, 2,384 reactions, and 861 genes. We validated the model experimentally using over 70 different C and N sources under both aerobic and anaerobic conditions. iThauera861 achieved a prediction accuracy of 95% for growth on various C and N sources and close to 85% for assimilation of aromatic compounds under denitrifying conditions. The M-model was subsequently deployed to determine the effects of substrates, oxygen presence, and the C:N ratio on the production of PHB and exopolysaccharides (EPS), showing the highest polymer yields are achieved with nucleotides and amino acids under aerobic conditions. This comprehensive M-model will help reveal the metabolic processes by which this ubiquitous species influences communities in wastewater treatment systems and natural environments.

Thauera

Genome-scale metabolic modeling reveals increased reliance on valine catabolism in clinical isolates of Klebsiella pneumoniae.

Infections due to carbapenem-resistant Enterobacteriaceae have recently emerged as one of the most urgent threats to hospitalized patients within the United States and Europe. By far the most common etiological agent of these infections is Klebsiella pneumoniae, frequently manifesting in hospital-acquired pneumonia with a mortality rate of ~50% even with antimicrobial intervention. We performed transcriptomic analysis of data collected previously from in vitro characterization of both laboratory and clinical isolates which revealed shifts in expression of multiple master metabolic regulators across isolate types. Metabolism has been previously shown to be an effective target for antibacterial therapy, and genome-scale metabolic network reconstructions (GENREs) have provided a powerful means to accelerate identification of potential targets in silico. Combining these techniques with the transcriptome meta-analysis, we generated context-specific models of metabolism utilizing a well-curated GENRE of K. pneumoniae (iYL1228) to identify novel therapeutic targets. Functional metabolic analyses revealed that both composition and metabolic activity of clinical isolate-associated context-specific models significantly differs from laboratory isolate-associated models of the bacterium. Additionally, we identified increased catabolism of L-valine in clinical isolate-specific growth simulations. These findings warrant future studies for potential efficacy of valine transaminase inhibition as a target against K. pneumoniae infection.

Humans

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans

Simple scaling laws control the genetic architectures of human complex traits.

Genome-wide association studies have revealed that the genetic architectures of complex traits vary widely, including in terms of the numbers, effect sizes, and allele frequencies of significant hits. However, at present we lack a principled way of understanding the similarities and differences among traits. Here, we describe a probabilistic model that combines the effects of mutation, drift, and stabilizing selection at individual sites with a genome-scale model of phenotypic variation. In this model, the architecture of a trait arises from the distribution of selection coefficients of mutations and from two scaling parameters. We fit this model for 95 highly polygenic quantitative traits of different kinds from the UK Biobank. Notably, we infer that all these traits have fairly similar, though not identical, distributions of selection coefficients. This similarity suggests that differences in architectures of highly polygenic traits arise mainly from the two scaling parameters: the mutational target size and heritability per site, which vary by orders of magnitude among traits. When these two scale factors are accounted for, we find that the architectures of all 95 traits are very similar.

Humans

Dynamic metabolic modelling of ATP allocation during viral infection.

Viral pathogens, like SARS-CoV-2, hijack the host's macromolecular production machinery, imposing an energetic burden that is distributed across cellular metabolism. To explore the dynamic metabolic tension between the host's survival and viral replication, we developed a computational framework that uses genome-scale models to perform dynamic flux balance analysis of human cell metabolism during virus infections. Relative to previous models, our framework addresses the physiology of viral infections of non-proliferating host cells through two new features. First, by incorporating the lipid content of SARS-CoV-2 biomass, we discovered activation of previously overlooked pathways giving rise to new predictions of possible drug targets. Furthermore, we introduce a dynamic model that simulates the partitioning of resources between the virus and the host cell, capturing the extent to which the competition depletes the human cells from essential ATP. By incorporating viral dynamics into our COMETS framework for spatio-temporal modelling of metabolism, we provide a mechanistic, dynamic and generalizable starting point for bridging systems biology modelling with viral pathogenesis. This framework could be extended to broadly incorporate phage dynamics in microbial systems and ecosystems.

Humans

GENKI: A generative framework for scalable and robust metabolic kinetic modeling.

GENKI (Generative ENsemble KPI-Informed) is a variational autoencoder-based framework for large-scale kinetic modeling of metabolism. Developed for metabolic engineering applications, GENKI is designed to improve the recovery of kinetically feasible models that reproduce experimentally observed phenotypes under genetic and environmental perturbations. The framework is trained on feasible kinetic model ensembles and uses phenotype-based key performance indicators (KPIs), derived from multi-omics and bioprocess data, to label and enrich models according to their agreement with mutant and condition-specific observations. This enables targeted generation of biologically relevant parameter sets with improved predictive performance. Crucially, GENKI recovers kinetic parameter sets that jointly reproduce wild-type and multiple perturbed physiologies within a single model. We apply GENKI to large-scale kinetic models of Escherichia coli and Saccharomyces cerevisiae under enzyme perturbations and oxygen shifts. In both systems, GENKI enriches kinetic ensembles with models that more accurately reproduce experimentally observed physiologies across multiple perturbations and conditions. GENKI therefore provides a practical framework for perturbation-aware kinetic model refinement within iterative Design-Build-Test-Learn workflows.

DBTL

Model-driven analysis reveals oxidative stress adaptation enabling efficient energy utilization in a Crabtree-negative Saccharomyces cerevisiae.

Although abolishing the Crabtree effect in Saccharomyces cerevisiae through a pyruvate dehydrogenase bypass eliminates carbon loss through ethanol overflow metabolism, it compromises growth rates. While the Crabtree effect has been a valuable natural adaptation, it is energetically inferior to respiration and is generally undesirable in cell factories engineered to produce assimilatory compounds. Restoring growth efficiency in Crabtree-negative strains remains a central challenge. Through adaptive laboratory evolution of the engineered strain (sZJD23) and subsequent reverse engineering, a variant (sZJD28) with markedly improved growth was identified. This improvement is driven primarily by a mutation in MED2 (encoding a Mediator complex subunit) and, to a lesser extent, a mutation in GPD1 (encoding glycerol-3-phosphate dehydrogenase). By integrating quantitative proteomics with enzyme-constrained genome-scale modelling, we demonstrate that these mutations jointly enable a more efficient mode of oxidative stress adaptation and energy utilization. The GPD1 mutation suppresses a protein-costly, suboptimal NAD⁺-recycling strategy reliant on glycerol synthesis, while the MED2 mutation reshapes the oxidative stress response towards peroxisomal detoxification. Collectively, these adjustments optimize metabolic flux distribution and reduce protein costs in energy metabolism, thereby increasing ATP availability. Our findings reveal how coordinated mutations in regulatory and metabolic genes restore growth fitness in engineered Crabtree-negative yeast.

Saccharomyces cerevisiae

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning

In silico encounters: harnessing metabolic modelling to understand plant-microbe interactions.

Understanding plant-microbe interactions is vital for developing sustainable agricultural practices and mitigating the consequences of climate change on food security. Plant-microbe interactions can improve nutrient acquisition, reduce dependency on chemical fertilizers, affect plant health, growth, and yield, and impact plants' resistance to biotic and abiotic stresses. These interactions are largely driven by metabolic exchanges and can thus be understood through metabolic network modelling. Recent developments in genomics, metagenomics, phenotyping, and synthetic biology now enable researchers to harness the potential of metabolic modelling at the genome scale. Here, we review studies that utilize genome-scale metabolic modelling to study plant-microbe interactions in symbiotic, pathogenic, and microbial community systems. This review catalogues how metabolic modelling has advanced our understanding of the plant host and its associated microorganisms as a holobiont. We showcase how these models can contextualize heterogeneous datasets and serve as valuable tools to dissect and quantify underlying mechanisms. Finally, we consider studies that employ metabolic models as a testbed for in silico design of synthetic microbial communities with predefined traits. We conclude by discussing broader implications of the presented studies, future perspectives, and outstanding challenges.

Plants

WILDkCAT: extract, retrieve, and predict enzyme turnover numbers of constraint-based metabolic models.

SUMMARY: Accurate enzyme turnover numbers are essential for building enzyme-constrained genome-scale metabolic models. However, collecting and curating these parameters remains a major bottleneck. Indeed, kcat values are scattered across multiple databases, reported under varying experimental conditions, and often missing for many enzymes. To address this challenge, we present WILDkCAT, a Python-based pipeline that enables the retrieval of kcat values from wild-type enzyme measured under user-specified pH and temperature ranges for a given metabolic model. The application to Escherichia coli (iML1515) and Homo sapiens (Human-GEM) models demonstrated the ability of WILDkCAT to retrieve substantial kcat coverage and its applicability across diverse genome-scale models. AVAILABILITY AND IMPLEMENTATION: WILDkCAT is available at https://github.com/sysbiolux/WILDkCAT and from PyPI. WILDkCAT works on all major operating systems and computer architectures. The documentation is available at https://sysbiolux.github.io/WILDkCAT.

Software

In vivo and in silico models of Drosophila for Parkinson's disease.

The fruit fly Drosophila melanogaster has emerged as an important model organism to shed light on neurodegeneration. Parkinson's disease (PD) is the second most prevalent neurodegenerative disorder, the cause of which is still mostly unclear. The long-term use of available PD drugs may have major side effects, and they only target the symptoms without providing any effective cure for the disease. Therefore, in vivo and in silico approaches are extensively used to model PD-like phenotypes in Drosophila and investigate cellular alterations underlying PD pathogenesis. In vivo models are particularly crucial to provide insight into the PD-related molecular processes. It has been a preferred approach to investigate these models by collecting omics datasets, which can be further analysed using in silico modeling such as genome-scale metabolic models and artificial intelligence applications. This review aims to summarise in vivo and in silico modeling studies in the literature to illustrate the potential of the Drosophila in the characterisation of PD-related biological mechanisms towards providing early biomarkers and novel treatment options for PD.

Humans

dAMN: a genome-scale neural-mechanistic hybrid model to predict bacterial growth dynamics.

SUMMARY: This study presents dAMN, a genome-scale neural-mechanistic hybrid model that combines neural networks with dynamic flux balance analysis to predict bacterial growth dynamics across diverse nutrient environments. Using a residual network architecture, dAMN predicts reaction fluxes and lag-phase parameters from initial medium composition, then integrates these predictions under stoichiometric constraints derived from genome-scale metabolic models. Trained on Escherichia coli and Pseudomonas putida growth datasets across combinatorial media, dAMN accurately forecasts temporal growth dynamics and generalizes to unseen media conditions, with mean R² ≥ 0.9. The model also reproduces biologically relevant behaviors including substrate depletion, acetate overflow, and diauxic shifts, while explicitly modeling lag phases usually absent from standard dFBA. AVAILABILITY AND IMPLEMENTATION: The dAMN software, associated models, and datasets are available at https://github.com/brsynth/dAMN-main-release and via Zenodo DOI: 10.5281/zenodo.17908125.

Escherichia coli