PubMed HealthSearch

PubMed · 41913056

FALCON2: compression-based metagenomic classification of ancient viruses.

Abstract

MOTIVATION: Ancient DNA (aDNA) sequences present unique challenges for taxonomic classification due to extreme fragmentation (reads 20-100 bp), end-biased cytosine deamination, and high contamination rates. Conventional metagenomic classifiers based on exact k-mer matching or alignment lose discriminative power on such short and damaged reads, limiting the analysis of paleogenomic samples. RESULTS: We present FALCON2, a compression-based metagenomic classifier that leverages position-aware finite-context models to maintain high accuracy on degraded viral ancient viruses. FALCON2 consolidates the capabilities of its predecessor, FALCON-meta, into a unified executable with enhanced features including model persistence, direct processing of compressed inputs, multiple file handling, and optional pre-filtering methodologies for contaminated samples. Under controlled benchmarking with database, taxonomy, and thread parity on simulated viral datasets, FALCON2 achieved an Area Under the Curve of Receiver Operating Characteristic (AUC-ROC) of 0.999, an Area Under Precision-Recall Curve (AUPRC) of 0.968, and an F1-score of 0.918, substantially outperforming Centrifuge (AUPRC = 0.625), Kraken2 (AUPRC = 0.184), and CLARK-S (AUPRC = 0.013) on pooled micro-averaged metrics. FALCON2's advantage is most pronounced on ultra-short reads (20-40 bp), where exact k-mers become sparse. FALCON2 pre-filtering at threshold 0.7 improved precision by 10 percentage points with negligible recall loss. FALCON2 runs on systems with 4-8 GB RAM for typical analyses. AVAILABILITY AND IMPLEMENTATION: FALCON2 is freely available at https://github.com/cobilab/FALCON2 under GPL v3 license. Benchmarking data and scripts are archived at DOI: https://doi.org/10.5281/zenodo.17291214.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Luís L Marques, Armando J Pinho, Diogo Pratas. 2026-05-03. FALCON2: compression-based metagenomic classification of ancient viruses.. https://doi.org/10.1093/bioinformatics%2Fbtag155

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20×2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30× coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N = 75), stool (N = 242), and saliva (N = 190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics

Functional convergence of rTCA-related carbon-fixation potential and biochemical residue accumulation in seagrass sediments.

Seagrass meadows are globally significant blue carbon ecosystems, yet the microbial and biochemical mechanisms driving sediment organic carbon (SOC) accumulation remain poorly understood. To address this, we employed an integrated approach combining metagenomic sequencing, biochemical assays, and structural equation modeling to investigate carbon cycling in the seagrass and adjacent unvegetated sediments of Swan Lake, China. A total of 115,179 carbon fixation genes and 119,615 decomposition genes were identified, revealing distinct microbial community structures among the habitats. Seagrass sediments harbored more diverse carbon-fixing (CFMs) and decomposing microorganisms (CDMs), with 83 medium-to high-quality metagenome-assembled genomes (MAGs) recovered. While neutral community model analysis indicated that stochastic processes predominantly governed community assembly, functional analyses highlighted specific drivers of sequestration. The reductive tricarboxylic acid (rTCA) cycle emerged as the dominant carbon fixation pathway, with key genes (e.g., aclA, korA) showing strong positive correlations with SOC. Conversely, decomposition pathways for starch and lignin were negatively associated with SOC. Furthermore, seagrass sediments exhibited elevated concentrations of total amino sugars (TAS) and lignin phenols (TLP), which linked significantly to carbon fixation rather than decomposition. PLS-SEM revealed statistically significant associations among seagrass traits, environmental variables, microbial carbon-fixation potential, biochemical residue pools, and SOC, supporting a mechanistic pathway in which enhanced microbial functional potential drives the accumulation of recalcitrant biochemical residues, thereby facilitating long-term carbon retention in sediments. These findings emphasize the pivotal role of microbial anabolism and the accumulation of biosynthetic residues in sediment carbon storage, suggesting a functional convergence in seagrass-driven carbon sinks.

Metagenomics

Community-driven updates for comprehensive long-read metagenomics and enhanced binning in nf-core/mag v5.

SUMMARY: nf-core/mag is a reproducible Nextflow pipeline for best-practice metagenomic de novo assembly and binning within the nf-core framework. Here we present a major update that adds support for long-read-only assembly and bin refinement, includes five new binning tools, expands taxonomic classification to viruses and eukaryotes, and improves bin quality evaluation with new tools and latest databases. Through sustained community-driven development spanning seven years and four primary curator teams, nf-core/mag remains actively developed as an open source workflow for metagenomic analysis, benefiting from contributions from across the broader metagenomics, nf-core, and Nextflow ecosystem. AVAILABILITY AND IMPLEMENTATION: The source code of nf-core/mag v5 is available on GitHub (https://github.com/nf-core/mag) under the open source MIT license, with v5.5.0 source code archived on Zenodo (https://zenodo.org/records/21735731). Documentation is viewable on the nf-core website (https://nf-co.re/mag).

Metagenomics