PubMed HealthSearch

SEARCH · PubMed Health

Results for “metagenome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Rapid pan-microbial metagenomics for pathogen detection and personalised therapy in the intensive care unit: a single-centre prospective observational study.

BACKGROUND: Most clinical metagenomic studies do not provide rapid results, detect pathogens from all microbial kingdoms, or measure clinical impacts. We aimed to evaluate the feasibility, performance, and clinical impacts of a rapid pan-microbial respiratory metagenomic service for patients admitted to intensive care units (ICUs). METHODS: This was a single-centre observational study of a rapid metagenomics service that tests respiratory samples from ICU patients at Guy's and St Thomas' hospitals, London, UK, between Dec 5, 2023, and April 12, 2024. Testing used a previously published pan-microbial metagenomics workflow, which simultaneously detects bacteria, fungi, and DNA and RNA viruses; provides same-day preliminary results after 2 h; and provides final results after 24 h. Patients were included if they were aged 18 years or older, admitted to the ICU, had confirmed respiratory failure requiring supplemental oxygen or advanced airway support, and had at least one of the following: (1) clinical suspicion of lower respiratory tract infection based on clinical, biochemical, or radiological findings, (2) sepsis of unknown origin, and (3) concern from an intensive care physician regarding inflammatory pathology. Patients with a suspected or confirmed containment level three organism were excluded. The outcome was performance characteristics of the metagenomic test compared with routine diagnostic testing, detection of additional pathogens by metagenomics, change in antimicrobial prescribing within 24 h of testing, and initiation of immunomodulation. FINDINGS: We processed 114 samples (1-5 per day) from 74 patients (39 [53%] female and 35 [47%] male). 107 (94%) of 114 samples passed quality control, of which 101 (94%) provided same-day preliminary results. Bacteria were detected in 45 (43%) of 104 tested specimens, fungal organisms in 17 (16%) of 104 tested specimens, and viruses in 28 (34%) of 83 tested specimens. Sensitivity in lower respiratory tract samples after 24 h was 97% (95% CI 87-100) for bacteria, 89% (65-99) for fungi, and 89% (71-98) for viruses, with only one false positive for bacteria. Metagenomics identified 42 pathogens not detected by other tests in 32 (30%) of 107 samples. Antimicrobial therapy was changed after metagenomic results from 30 (28%) of 107 samples: 22 (21%) were de-escalated and eight (7%) were escalated. Metagenomics contributed to the initiation of immunomodulation in 15 (20%) of 74 patients for a range of inflammatory conditions. Pathogens with clinical significance to local infection control or national public health were found in ten (14%) of 74 patients, including three invasive Group A streptococci, two parvovirus B19, and one each of HIV-1, measles virus, Mycobacterium tuberculosis, Neisseria meningitidis, and Mycoplasma pneumoniae. INTERPRETATION: Respiratory metagenomics for ICU patients showed good performance and turnaround time, and diverse clinical and public health benefits. This ability to inform both personalised patient therapy and infectious disease surveillance needs evaluation in multicentre studies. FUNDING: None.

Humans

Computed tomography-guided precision biopsy combined with metagenomic next-generation sequencing for etiological diagnosis in patients with blood culture-negative systemic infections.

ObjectiveTo evaluate the diagnostic efficacy of computed tomography-guided percutaneous biopsy combined with metagenomic next-generation sequencing in patients with blood culture-negative systemic infections and to assess the clinical impact of using this combined strategy for etiological confirmation and guidance of targeted antimicrobial therapy.MethodsThis single-center retrospective observational cohort study enrolled 78 patients who met the Sepsis-3 consensus criteria for suspected systemic infection and had negative conventional microbiological work-ups (at least two sets of blood cultures) between April 2022 and March 2025. All patients underwent computed tomography-guided biopsy of radiologically identified infectious foci, with specimens processed concurrently for conventional culture and metagenomic next-generation sequencing. Diagnostic performance was benchmarked against the final comprehensive clinical diagnosis, and the influence of metagenomic next-generation sequencing findings on antimicrobial therapy modification was analyzed. Sample size calculation, based on a prior study estimating an metagenomic next-generation sequencing detection rate of 85% (&#x3b1;&#x2009;=&#x2009;0.05, &#x3b2;&#x2009;=&#x2009;0.2), indicated a minimum of 68 cases; accordingly, 78 patients were enrolled.ResultsComputed tomography-guided biopsy was technically successful in all 78 patients (100%). The pathogen detection rate of metagenomic next-generation sequencing (91.0%, 71/78) was significantly higher than that of conventional culture (55.1%, 43/78; p&#x2009;<&#x2009;0.001). Using the final clinical diagnosis as the reference standard, metagenomic next-generation sequencing achieved a sensitivity of 94.7% (95% confidence interval: 86.9-98.5), specificity of 100.0% (95% confidence interval: 29.2-100.0), positive predictive value of 100.0% (95% confidence interval: 94.9-100.0), and negative predictive value of 42.9% (95% confidence interval: 9.9-81.6). Among the 35 culture-negative specimens, metagenomic next-generation sequencing established a definitive microbiological diagnosis in 28 cases (80.0%) and detected polymicrobial infections in 11 cases (14.1% of the cohort). Antimicrobial therapy was rationally adjusted based on metagenomic next-generation sequencing results in 69.2% (54/78) of the patients.ConclusionsThe integration of computed tomography-guided precision biopsy with metagenomic next-generation sequencing offers a highly effective diagnostic approach for blood culture-negative systemic infections. This synergistic strategy improves etiological diagnosis by providing high-yield target specimens that enable comprehensive, unbiased pathogen screening, facilitates differentiation between infectious and non-infectious etiologies, and supplies critical evidence for guiding precision antimicrobial therapy. These findings highlight the growing role of interventional radiology in the contemporary framework of precision infectious disease management.

Humans

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics

Optimizing a culture-enriched hybrid metagenomics pipeline to assess the AMR footprint of livestock manure in anaerobic digestate.

The role of environmental samples from livestock production systems, including manure and anaerobic digestate, as reservoirs of antimicrobial resistance genes (ARGs) is likely underestimated because conventional metagenomic approaches can overlook low-abundance ARGs and often lack the resolution to associate these genes with their microbial hosts and co-localized mobile genetic elements (MGEs). We evaluated whether culture-enriched metagenomics (CEMG), with and without antibiotic selection, enhances ARG detection in anaerobic digestate and improves the resolution of ARG-MGE-host associations using hybrid short- and long-read metagenomic assembly. CEMG increased ARG recovery; mean ARG abundance rose from 15.4 counts per million (CPM) in metagenomic fresh digestate (FD) to 124 CPM in CEMG without antibiotics and 160 CPM in antibiotic-selective CEMG. In FD, only 9 unique ARGs were detected, whereas CEMG recovered 112, including ARGs of clinical importance, such as glycopeptide resistance, beta-lactamase genes, and the cfr 23S rRNA methyltransferase conferring cross-resistance to multiple antibiotic classes. Antibiotic selection induced targeted, class-specific shifts in ARG profiles, with ARGs associated with tetracycline resistance consistently enriched across treatments. Hybrid metagenomic assembly resolved the genomic context of 784 ARGs, of which 59.3% were co-localized with at least one class of MGEs, predominantly plasmids and integrative conjugative elements/integrative mobilizable elements. Biocide and metal resistance genes frequently co-occurred with ARGs on the same contigs. Together, these findings demonstrate that antibiotic-selective culture enrichment enhances resistome surveillance by improving detection of low-abundance ARGs, while hybrid assembly provides critical genomic context for assessing their mobility and host associations.IMPORTANCELivestock manure and its byproducts, such as anaerobic digestate, are recognized as important environmental reservoirs of antimicrobial resistance genes (ARGs) and resistant bacteria, yet current metagenomic approaches may underestimate this risk by failing to detect low-abundance but clinically relevant ARGs. Here, we show that integrating culture enrichment with hybrid metagenomics improves ARG recovery and reveals ARG co-localization with mobile genetic elements and putative bacterial hosts. This approach captures a cultivable and condition-responsive fraction of the resistome that is not readily accessible through direct metagenomic sequencing alone, providing a more informative framework for environmental AMR surveillance.

anaerobic digestion

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 &#xd7; 150 bp and 2 &#xd7; 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 &#xd7; 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 &#xd7; 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 &#xd7; 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 &#xd7; 150 bp and 2 &#xd7; 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 &#xd7; 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics

High-resolution metagenome assembly for modern long reads with myloasm.

Long-read metagenome assembly promises complete genomic recovery from microbiomes. However, the complexity of metagenomes poses challenges. We present myloasm, a metagenome assembler for PacBio HiFi and Oxford Nanopore Technologies (ONT) R10.4 long reads. Myloasm uses polymorphic k-mers to construct a high-resolution string graph and then leverages differential abundance for graph simplification. On real-world ONT metagenomes, myloasm assembled three times more complete circular contigs than the next-best assembler. Myloasm can make ONT and HiFi comparable for assembly: for a jointly sequenced gut metagenome, myloasm with ONT assembled more complete circular genomes than any assembler with HiFi. Myloasm recovers previously inaccessible within-species diversity; we recovered six complete Prevotella copri single-contig genomes from a gut metagenome and eight complete TM7 (Saccharibacteria) contigs with > 93% similarity from an oral metagenome. With this improved resolution, we resolved two 98% similar ermF antibiotic resistance genes spreading through distinct strain-specific mobile genetic elements in a human gut.

Journal Article

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome

De novo assembly and authentication of ancient DNA metagenomes with nf-core/mag.

Ancient DNA provides a direct window into the evolutionary processes that have shaped living microbial species today, as well as their now extinct relatives. Advances in both sequencing methods and de novo assembly techniques have not only resulted in a flood of modern metagenomic sequencing data, but they have also allowed palaeogenomicists to retrieve vast amounts of ancient DNA from past microorganisms, including species and strains without modern reference genomes. However, the degraded nature of ancient DNA means that the standard techniques of genome assembly developed for modern DNA are unlikely to perform effectively, unless heavily modified. This hinders the incorporation of ancient data into broader metagenomic studies that would otherwise benefit from having deep time information on the evolution of different microbial species. In this primer and protocol paper, we provide guidance on ways to adapt existing metagenomic de novo assembly processes, including data input, tools, and settings, in order to perform more robustly and effectively on ancient DNA. After assembly, we then further describe how ancient DNA contigs can be identified and validated. The key steps of ancient metagenomic assembly are now integrated in a dedicated ancient DNA mode in the established pipeline nf-core/mag. By introducing support for ancient DNA data in nf-core/mag, we aim to improve the ability of researchers to more regularly integrate de novo assembled ancient microbial data into broader metagenomics studies of microbial ecology and evolution.

DNA, Ancient

MADCAP: isolation of novel nAb-na&#xef;ve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV

Metagenomic insights and biosynthetic potential of Candidatus Entotheonella symbiont associated with Halichondria marine sponges.

Korea, being surrounded by the sea, provides a rich habitat for marine sponges, which have been a prolific source of bioactive natural products. Although a diverse array of structurally novel natural products has been isolated from Korean marine sponges, their biosynthetic origins remain largely unknown. To explore the biosynthetic potential of Korean marine sponges, we conducted metagenomic analyses of sponges inhabiting the East Sea of Korea. This analysis revealed a symbiotic association of Candidatus Entotheonella bacteria with Halichondria sponges. Here, we report a new chemically rich Entotheonella variant, which we named Ca. Entotheonella halido. Remarkably, this symbiont makes up 69% of the microbial community in the sponge Halichondira dokdoensis. Genome-resolved metagenomics enabled us to obtain a high-quality Ca. E. halido genome, which represents the largest (12 Mb) and highest quality among previously reported Entotheonella genomes. We also identified the biosynthetic gene cluster (BGC) of the known sponge-derived Halicylindramides from the Ca. E. halido genome, enabling us to determine their biosynthetic origin. This new symbiotic association expands the host diversity and biosynthetic potential of metabolically talented bacterial genus Ca. Entotheonella symbionts.IMPORTANCEOur study reports the discovery of a new bacterial symbiont Ca. Entotheonella halido associated with the Korean marine sponge Halichondria dokdoensis. Using genome-resolved metagenomics, we recovered a high-quality Ca. E. halido MAG (Metagenome-Assembled Genome), which represents the largest and most complete Ca. Entotheonella MAG reported to date. Pangenome and BGC network analyses revealed a remarkably high BGC diversity within the Ca. Entotheonella pangenome, with almost no overlapping BGCs between different MAGs. The cryptic and genetically unique BGCs present in the Ca. Entotheonella pangenome represents a promising source of new bioactive natural products.

Animals

Microbial metagenomes from Lake Soyang, the largest freshwater reservoir in South Korea.

Lake ecosystems play a fundamental role in the global biogeochemical cycling of essential elements such as carbon, nitrogen, and phosphorus. Microorganisms within these ecosystems mediate key processes that regulate these cycles. Metagenomic analyses provide valuable insights into the taxonomic and functional diversity of microbial communities in various environments, including freshwater habitats. Here, we present a comprehensive metagenomic dataset derived from Lake Soyang, the largest freshwater reservoir in South Korea. A total of 28 metagenomes were generated from water samples collected across two distinct sampling periods: the first set (n&#x2009;=&#x2009;8) was obtained between April 2014 and January 2015 from two depths (1&#x2009;m and 50&#x2009;m) in four different seasons, while the second set (n&#x2009;=&#x2009;20) was collected between January 2019 and November 2019 from five depths (1, 10, 20, 40, and 90&#x2009;m) over four seasons. Metagenomic sequencing yielded 9.3-21.8 Gbp per sample. This dataset provides a valuable resource for future studies exploring the ecophysiological characteristics of microbial communities in pelagic freshwater environments.

Republic of Korea

GiantHunter: accurate detection of giant virus in metagenomic data using reinforcement-learning and Monte Carlo tree search.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) are notable for their large genomes and extensive gene repertoires, which contribute to their widespread environmental presence and critical roles in processes such as host metabolic reprogramming and nutrient cycling. Metagenomic sequencing has emerged as a powerful tool for uncovering novel NCLDVs in environmental samples. However, identifying NCLDV sequences in metagenomic data remains challenging due to their high genomic diversity, limited reference genomes, and shared regions with other microbes. Existing alignment-based and machine learning methods struggle with achieving optimal trade-offs between sensitivity and precision. RESULTS: In this work, we present GiantHunter, a reinforcement learning-based tool for identifying NCLDVs from metagenomic data. By employing a Monte Carlo tree search strategy, GiantHunter dynamically selects representative non-NCLDV sequences as the negative training data, enabling the model to establish a robust decision boundary. Benchmarking on rigorously designed experiments shows that GiantHunter achieves high precision while maintaining competitive sensitivity, improving the F1-score by 10% and reducing computational cost by 90% compared to the second-best method. To demonstrate its real-world utility, we applied GiantHunter to 60 metagenomic datasets collected from six cities along the Yangtze River, located both upstream and downstream of the Three Gorges Dam. The results reveal significant differences in NCLDV diversity correlated with proximity to the dam, likely influenced by reduced flow velocity caused by the dam. These findings highlight GiantHunter's potential to advance our understanding of NCLDVs and their ecological roles in diverse environments. AVAILABILITY AND IMPLEMENTATION: The source code of GiantHunter is available via: https://github.com/FuchuanQu/GiantHunter.

Metagenomics

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9&#xb1;10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics

VirBinn improves viral genome binning from metagenomic Hi-C through graph diffusion.

MOTIVATION: Metagenomic Hi-C provides in situ proximity signals that can improve genome binning and enable virus-host-association analysis. However, viral genome recovery remains difficult because virus-virus Hi-C contact matrices are extremely sparse. Viral genomes are small, often low-abundance, and frequently assemble into short contigs, leaving many true within-genome links unobserved and causing viral bins to fragment. RESULTS: We present VirBinn, a graph-diffusion framework for viral binning from metagenomic Hi-C. VirBinn enhances virus-virus connectivity through two complementary mechanisms: random-walk-with-restart enhancement on the sparse virus-virus contact graph and host-guided diffusion that propagates viral seeds through the host network to infer indirect virus-virus associations. The enhanced views are integrated and clustered using Leiden community detection to produce viral metagenome-assembled genomes (vMAGs). On dataset-specific simulation benchmarks with ground truth, VirBinn consistently recovers more high-quality vMAGs than Hi-C-based and shotgun-based baselines and substantially increases the number of near-complete genomes. On four real metagenomic Hi-C datasets spanning human gut, pig gut, sheep gut (long-read assembly), and wastewater, VirBinn yields more high-completeness vMAGs under CheckV and produces bins with strong within-cluster contact support. Finally, host linkage analysis using reconstructed host MAGs reveals habitat-specific host-association patterns and plausible host taxonomic profiles. AVAILABILITY AND IMPLEMENTATION: VirBinn is available at https://github.com/dyxstat/VirBinn. The scripts to reproduce the results and figures in this article are available at https://github.com/dyxstat/Reproduce_VirBinn.

Genome, Viral