PubMed HealthSearch

SEARCH · PubMed Health

Results for “coarse-grained modeling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

9 recordsLinked to original sources

Coarse-grained model of serial dilution dynamics in synthetic human gut microbiome.

Many microbial communities in nature are complex, with hundreds of coexisting strains and the resources they consume. We currently lack the ability to assemble and manipulate such communities in a predictable manner in the lab. Here, we take a first step in this direction by introducing and studying a simplified consumer resource model of such complex communities in serial dilution experiments. The main assumption of our model is that during the growth phase of the cycle, strains share resources and produce metabolic byproducts in proportion to their average abundances and strain-specific consumption/production fluxes. We fit the model to describe serial dilution experiments in hCom2, a defined synthetic human gut microbiome with a steady-state diversity of 63 species growing on a rich media, using consumption and production fluxes inferred from metabolomics experiments. The model predicts serial dilution dynamics reasonably well, with a correlation coefficient between predicted and observed strain abundances as high as 0.8. We applied our model to: (i) calculate steady-state abundances of leave-one-out communities and use these results to infer the interaction network between strains; (ii) explore direct and indirect interactions between strains and resources by increasing concentrations of individual resources and monitoring changes in strain abundances; (iii) construct a resource supplementation protocol to maximally equalize steady-state strain abundances.

Gastrointestinal Microbiome

Coarse-grained resource allocation modeling for decoding and rewiring microbial metabolism.

Microbial metabolism is a complex, emergent system driven by the coordinated interplay of intricate and dynamic molecular processes. To elucidate cellular behavior and enable biotechnological applications, quantitative models that address the inherent complexity of metabolism have been developed from a resource allocation perspective. Here, we synthesize recent advances in coarse-grained resource allocation frameworks and their applications in understanding microbial physiology and guiding gene circuit design. These frameworks reveal global regulatory constraints and predict cellular adaptation to nutrient and environmental changes. In addition, they enable the quantification of metabolic costs, the dissection of circuit-host interactions, and the development of strategies for burden mitigation. Collectively, these modeling frameworks provide a powerful platform for uncovering quantitative principles of microbial growth and engineering robust synthetic biological systems.

coarse-grained modeling

Molecular dynamics simulations reveal subtle consequences of H3K9 and H3K27 tri-methylation on chromatin constituents.

Epigenetic modifications of histone tails are key mechanisms of genome regulation. In particular, tri-methylation of lysines (K) 9 and K27 of the histone H3 tail is important for genome silencing. In this work, we explore, using all-atom molecular dynamics simulations, the effect of these two epigenetic marks on the structure and interactions of the H3 tail in several contexts: isolated tails, nucleosomes, chromatosomes, and stacked nucleosomes. Overall, we find that although the isolated tails do not show significant conformational changes upon methylation, a more flexible and extended H3 tail compared to the native tail results in the nucleosome systems, with K9 methylation effects more pronounced. This change could facilitate the interaction of the tail with protein readers like heterochromatin protein 1 or Polycomb group. We also observe that both methylations increase the interactions of the H3 tail with the linker DNA in the context of the chromatosome, producing a chromatosome with tighter linker DNA, which could favor chromatin compaction. For stacked nucleosomes mimicking i±2 zigzag interactions, methylation of either K9 or K27 reduces the interactions of one of the H3 tails with its parental nucleosome and increases its interactions with the nonparental nucleosome, which could also help compact the chromatin fiber. In the three nucleosome-containing systems, we observe an asymmetry between the two tails, especially in the chromatosome, where one tail extends to interact with the linker DNA. This asymmetry modulates the effect that methylation has on each tail. Thus, overall, methylations of K9 and K27 have a subtle but notable impact on the H3 tail structure and its interactions within the chromatin fiber. These results help explain how this epigenetic modification compacts chromatin fibers and promotes longer-range interactions; these changes also guide how to approximate these effects in coarse-grained chromatin models.

Histones

One chromatin, many structures: From ensemble contact maps to single-cell 3D organization.

Understanding how chromatin folds in three dimensions remains challenging because most experimental assays capture low-dimensional projections of an underlying, highly heterogeneous polymer. Here, we present an ensemble-based interpretive framework built on the previously introduced Self-Returning Excluded Volume (SR-EV) model, a minimal generator of chromatin conformations using a nucleosome-indexed coarse-grained representation based on stochastic return rules and excluded-volume geometry. Despite its simplicity, SR-EV recapitulates key experimental signatures across scales: heterogeneous nanoscale packing domains resembling ChromEMT and ChromSTEM observations, sparse and highly variable single-configuration contact patterns analogous to single-cell chromosome conformation capture (Hi-C), and robust ensemble-level contact enrichment consistent with topologically associating domains (TADs). In this framework, Hi-C loop and TAD signatures are interpreted as ensemble-level statistical enrichments rather than invariant features of single-cell conformations. SR-EV is explicitly designed to generate large ensembles of complete three-dimensional chromatin configurations that can be projected consistently onto two-dimensional contact maps and one-dimensional genomic profiles. By introducing architectural-protein effects only through ensemble selection rather than explicit forces, SR-EV supports a separation between intrinsic polymer geometry and regulatory bias and suggests that TAD-like features can emerge as statistical enrichments rather than deterministic three-dimensional structures. Coordination number and probe-based accessibility computed directly from SR-EV provide a unified link between three-dimensional packing, two-dimensional contact maps, and one-dimensional genomic profiles. The main contribution of this work is to show, within a single coarse-grained framework, how these multimodal observables arise as linked projections of the same heterogeneous chromatin ensemble through averaging and conditional sampling. Together, these results establish SR-EV as a minimal and geometrically grounded mesoscale reference framework for interpreting how heterogeneous chromatin ensembles give rise to multimodal experimental observables while remaining consistent with the fact that chromatin organization is realized in individual cells.

Chromatin

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Pervasive fitness trade-offs revealed by rapid adaptation to shifting population densities in large experimental populations of Drosophila melanogaster.

Trade-offs are an inherent feature of organismal biology that are expected play a fundamental role in the evolution of natural populations. Efforts to quantify trade-offs are largely confined to phenotypic measurements and the identification of negative genetic-correlations among fitness-relevant traits. Here, we use time-series genomic data collected during experimental evolution in large, genetically diverse populations of Drosophila melanogaster to directly measure the manifestation of trade-offs in response to fluctuating selection on ecological timescales. Specifically, we first conducted a lab-based selection experiment to quantify a genome-wide signal of antagonistic pleiotropy elicited in response to shifting population densities and associated with reproduction and stress tolerance selection. In doing so, we identified a putative role of two cosmopolitan inversions in these trade-offs. We then conducted an independent experiment to show that a simple manipulation of increasing population density under controlled lab-based conditions identified loci that are relevant to selection during population expansion and collapse in a complex, semi-natural setting. In concert, our results reveal how adaptation in complex, natural environments can be coarse-grained in such a manner to drive repeatable and predictable patterns of genomic variation, and further add credence to models positing a role of generic fitness trade-offs in the maintenance of variation in natural populations.

Drosophila melanogaster

Coarse-grained chromatin dynamics by tracking multiple similarly labeled gene loci.

The "holy grail" of chromatin research would be to follow the chromatin configuration in individual live cells over time. One way to achieve this goal would be to track the positions of multiple loci arranged along the chromatin polymer with fluorescent labels. Using distinguishable labels would define each locus uniquely in a microscopic image but would restrict the number of loci that could be observed simultaneously due to experimental limits to the number of distinguishable labels. Using the same label for all loci circumvents this limitation but requires a (currently lacking) framework for how to establish each observed locus identity, i.e., to which genomic position it corresponds. Here, we analyze theoretically, using simulations of Rouse model polymers, how single-particle tracking of multiple identically labeled loci enables the determination of loci identity. We show that the probability of correctly assigning observed loci to genomic positions converges exponentially to unity as the number of observed loci configurations increases. The convergence rate depends only weakly on the number of labeled loci, so that even large numbers of loci can be identified with high fidelity by tracking them across about eight independent chromatin configurations. In the case of two distinct labels that alternate along the chromatin polymer, we find that the probability of the correct assignment converges faster than for same-labeled loci, requiring observation of fewer independent chromatin configurations to establish loci identities. Finally, for a modified Rouse model polymer, which realizes a population of dynamic loops, we find that the success probability also converges to unity exponentially as the number of observed loci configurations increases, albeit slightly more slowly than for a classical Rouse model polymer. Altogether, these results establish particle tracking of multiple identically or alternately labeled loci over time as a feasible way to infer temporal dynamics of the coarse-grained configuration of the chromatin polymer in individual living cells.

Chromatin

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family