PubMed HealthSearch

Biomedical subjects

Michael C Schatz

Publications and source records attributed to Michael C Schatz.

4 recordsLinked to original sources

Perseus: Lineage-Aware Refinement of Kraken2 Taxonomic Classification for Long Read Metagenomes.

MOTIVATION: Long-read metagenomic sequencing improves assembly contiguity and enables genome-resolved analysis of complex microbial communities, but accurate taxonomic classification of long reads and assembled contigs remains challenging. Highly scalable k-mer-based classifiers such as Kraken2 frequently over-assign fine-rank taxonomic labels when applied to long-read data, producing high false positive classification rates driven by sparse or localized k-mer matches, particularly in microbiomes with extensive taxonomic novelty. RESULTS: We present Perseus, a lineage-aware confidence estimation framework for taxonomic classification that models the spatial distribution and hierarchical consistency of k-mer evidence along sequences. This formulation reframes taxonomic classification as a hierarchical confidence estimation problem rather than a single-rank prediction task. Perseus refines k-mer-level taxonomic signals from Kraken2 using a multi-headed convolutional neural network that estimates calibrated confidence scores for taxonomic correctness at each canonical rank. Using these estimates, Perseus confirms assignments, backs off to higher taxonomic ranks, or abstains when evidence is insufficient, prioritizing correctness and lineage consistency over overly specific assignments. Across simulations of taxonomic novelty and real-world metagenomic datasets, Perseus consistently and substantially reduces the false assignment rate while improving precision and lineage-consistent accuracy. These improvements are most pronounced for long reads and assembled contigs, where spatial context enables reliable discrimination between consistent taxonomic signal and spurious matches. AVAILABILITY AND IMPLEMENTATION: Perseus integrates with existing Kraken2 workflows and is available at https://github.com/matnguyen/perseus.

Journal Article

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Microbiology Galaxy Lab: The first community-driven gateway for reproducible and FAIR analysis of microbial data.

The explosion of microbial omics data has outpaced the ability of many researchers to analyze it, with complex tools and limited computational resources creating barriers to discovery. To address this gap, we present the Microbiology Galaxy Lab: a free, globally accessible, community-supported platform that combines state-of-the-art analytical power with user-friendly accessibility. Supported by the Galaxy and global microbiology communities, this platform integrates over 315 tool suites and 115 curated workflows, enabling comprehensive metabarcoding, (meta)genomic, (meta)transcriptomic, and (meta)proteomic data analysis within a FAIR-aligned environment. It also supports research in the health and infectious disease sectors, as well as in environmental microbiology. The platform's utility is exemplified through various use cases, including antimicrobial resistance tracking, biomarker prediction, microbiome classification, and functional annotation of key microbes. Built on reproducibility and community engagement, it supports creation, sharing, and updating of best-practice workflows. Over 35 tutorials and learning paths empower scientists, fostering an ecosystem that keeps resources at the forefront of microbial science. The Microbiology Galaxy Lab enables collective analysis, democratising research, thereby accelerating discovery across the global microbiology community (microbiology.usegalaxy.org, .eu, .org.au, .fr).

Journal Article