PubMed HealthSearch

SEARCH · PubMed Health

Results for “microbial dark matter”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

6 recordsLinked to original sources

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes

Seed2LP: seed inference in metabolic networks for reverse ecology applications.

MOTIVATION: A challenging problem in microbiology is to determine nutritional requirements of microorganisms and culture them, especially for the microbial dark matter detected solely with culture-independent methods. The latter foster an increasing amount of genomic sequences that can be explored with reverse ecology approaches to raise hypotheses on the corresponding populations. Building upon genome-scale metabolic networks (GSMNs) obtained from genome annotations, metabolic models predict contextualized phenotypes using nutrient information. RESULTS: We developed the tool Seed2LP, addressing the inverse problem of predicting source nutrients, or seeds, from a GSMN and a metabolic objective. The originality of Seed2LP is its hybrid model, combining a scalable and discrete Boolean approximation of metabolic activity, with the numerically accurate flux balance analysis (FBA). Seed inference is highly customizable, with multiple search and solving modes, exploring the search space of external and internal metabolites combinations. Application to a benchmark of 107 curated GSMNs highlights the usefulness of a logic modelling method over a graph-based approach to predict seeds, and the relevance of hybrid solving to satisfy FBA constraints. Focusing on the dependency between metabolism and environment, Seed2LP is a computational support contributing to address the multifactorial challenge of culturing possibly uncultured microorganisms. AVAILABILITY AND IMPLEMENTATION: Seed2LP is available on https://github.com/bioasp/seed2lp.

Metabolic Networks and Pathways

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans

Microbial Phototrophic, Heterotrophic, and Diazotrophic Activities Associated with Aggregates in the Permanent Ice Cover of Lake Bonney, Antarctica.

Abstract The McMurdo Dry Valley lakes, Antarctica, one of the Earth's southernmost ecosystems containing liquid water, harbor some of the most environmentally extreme (cold, nutrient-deprived) conditions on the planet. Lake Bonney has a permanent ice cover that supports a unique microbial habitat, provided by soil particles blown onto the lake surface from the surrounding, ice-free valley floor. During continuous sunlight summers (Nov.-Feb.), the dark soil particles are heated by solar radiation and melt their way into the ice matrix. Layers and patches of aggregates and liquid water are formed. Aggregates contain a complex cyanobacterial-bacterial community, concurrently conducting photosynthesis (CO2 fixation), nitrogen (N2) fixation, decomposition, and biogeochemical zonation needed to complete essential nutrient cycles. Aggregate-associated CO2- and N2-fixation rates were low and confined to liquid water (i.e., no detectable activities in the ice phase). CO2 fixation was mediated by cyanobacteria; both cyanobacteria and eubacteria appeared responsible for N2 fixation. CO2 fixation was stimulated primarily by nitrogen (NO3-), but also by phosphorus (PO43-). PO43- and iron (FeCl3 + EDTA) enrichment stimulated of N2 fixation. Microautoradiographic and physiological studies indicate a morphologically and metabolically diverse microbial community, exhibiting different cell-specific photosynthetic and heterotrophic activities. The microbial community is involved in physical (particle aggregation) and chemical (establishing redox gradients) modification of a nutrient- and organic matter-enriched microbial "oasis," embedded in the desertlike (i.e., nutrient depleted) lake ice cover. Aggregate-associated production and nutrient cycling represent microbial self-sustenance in a microenvironment supporting "life at the edge," as it is known on Earth.

Journal Article

Novelty, diversity, and genetic dark matter in enterococci of invertebrates.

Enterococci appear to have originated in the guts of early terrestrializing arthropods and invertebrates over 425 million years ago-hosts that are now highly diverse and widespread in nature today. Yet most knowledge of the genus comes from human infection-associated lineages with genomes swollen by the recent accretion of foreign DNA conveyed by mobile elements. Because invertebrates dominate terrestrial animal diversity and biomass, they would be predicted to constitute a major but little-explored reservoir of enterococcal diversity. We therefore systematically examined Enterococcus association and species diversification in invertebrate hosts of the comparatively natural, isolated, but well-characterized environment of the Azorean island of Terceira. Over 100 invertebrate specimens were examined for associated enterococci, which were taxonomically classified by whole-genome sequencing. Supporting the existence of a large pool of uncharacterized enterococci and Enterococcus-adapted genes, 40% (eight of 20) of the Enterococcus species identified were either undescribed, including four candidate new species described here, or very recently discovered. In contrast, control isolates from vertebrates were exclusively of known species typical of sampling elsewhere, discounting geographic isolation as a main driver of the novelty observed. Further, because of the abundance of E. casseliflavus and E. flavescens in this collection, we obtained the resolution necessary to quantify the divergence and decipher the drivers of speciation in the controversial division between these naturally vancomycin-resistant species. These findings provide robust support for the existence of a large pool of new species and unexplored adaptive traits in invertebrate-associated enterococci-diverse environmental survival traits optimized for expression in an enterococcal background, and well positioned for transmission into human-associated enterococcal strains.IMPORTANCEEnterococci are auxotrophic gut-associated bacteria that co-evolved with their terrestrial hosts over many eons. In the last 75 years-the "antibiotic era"-E. faecalis and E. faecium gained genes for antibiotic resistance and enhanced virulence, emerging as leading causes of multidrug-resistant infection. Little is known about the source of those genes or the pathway by which they entered human-associated strains. A recent global survey suggested a potentially large repository of uncharacterized genetic diversity in the enterococci of invertebrates. We directly tested this prospect by examining enterococci of invertebrate hosts in a largely natural and pastoral environment. Our findings provide clear evidence that invertebrates naturally harbor vast unexplored enterococcal diversity. Moreover, associations are likely driven by intrinsic host selection factors rather than geographic isolation. This expands our knowledge of Enterococcus biodiversity, including the identification of four novel species, identifying a vast reservoir of enterococcal genes available to species that colonize and infect humans.

Animals