PubMed HealthSearch

PubMed · 41967838

Accelerated long-read variant calling with Clair3 for whole-genome sequencing.

Abstract

SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30× whole-genome sequence in 12-20 minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15 minutes on an Apple Mac Studio (32 threads), which is ∼10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30× ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhenxian Zheng, Minggao He, Xian Yu, Junzhe Li, Lei Chen, Angel On Ki Wong, Jingcheng Zhang, Yekai Zhou, Ruibang Luo. 2026-05-03. Accelerated long-read variant calling with Clair3 for whole-genome sequencing.. https://doi.org/10.1093/bioinformatics%2Fbtag181

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Whole genome sequencing analysis and functional characterization of Lacticaseibacillus rhamnosus HP-B1083.

Lacticaseibacillus rhamnosus is an important strain for the biotransformation of natural products, and its crude extract exhibits biotransformation effect on glycosidic compounds such as baicalin. To further explore the potential of this strain, particularly given its previously demonstrated high-efficiency β-glucuronidase activity for baicalin conversion, whole-genome sequencing and functional annotation of Lacticaseibacillus rhamnosus HP-B1083 were performed in this study, and its acid tolerance, bile salt tolerance, short-term heat resistance and antibacterial activity were evaluated. The results showed that the strain possessed a circular chromosome with a full length of 3,090,505 bp and a GC content of 46.69%. Gene annotation revealed that the genome contained 2941 coding sequences (CDS) and 112 non-coding RNA genes, including 60 tRNA genes, 1 tmRNA gene, 36 misc_RNA genes and 15 rRNA genes. The functional annotations further reveal that this genome is rich in genes related to carbohydrate metabolism, hydrolases, and transferases, which is highly consistent with its phenotypic characteristics in glycoside transformation and the synthesis of antibacterial substances. In addition, acid tolerance, bile salt tolerance and short-term heat resistance experiments verified that HP-B1083 had acid resistance, bile salt resistance and short-term heat resistance. Antibacterial activity tests confirmed that HP-B1083 produced inhibition zone diameters over 10 mm against common foodborne pathogenic bacteria such as Escherichia coli and Bacillus cereus. Therefore, Lacticaseibacillus rhamnosus HP-B1083 has important application prospects in the development of functional foods, preparation of enzyme preparations and pharmaceutical industry.

Whole Genome Sequencing

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3× coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30× coverage further enhances sensitivity, enabling detection of TFs as low as 1 × 10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing

Whole Genome Sequencing Reveals How Plasticity and Genetic Differentiation Underlie Sympatric Morphs of Arctic Charr.

Salmonids have a remarkable ability to form sympatric morphs after postglacial colonisation of freshwater lakes. These morphs often differ in morphology, feeding and spawning behaviour. Here, we explored the genetic basis of morph differentiation in Arctic charr (n = 283) by first establishing a high-quality reference genome and then using this in whole genome sequencing of distinct morphs present in two Norwegian and two Icelandic lakes. The four lakes represent the spectrum of genetic differentiation between morphs from one lake with no genetic differentiation between morphs, implying phenotypic plasticity, to two lakes with locus-specific genetic differentiation, implying incomplete reproductive isolation, and one lake with strong genome-wide divergence consistent with complete reproductive isolation. As many as 12 putative inversions ranging from 0.45 to 3.25 Mbp in size segregated among the four morphs present in one lake, Thingvallavatn, and these contributed significantly to the genetic differentiation among morphs. None of the putative inversions were found in any of the other lakes, but there were cases of partial haplotype sharing in similar morph contrasts in other lakes. Our findings are consistent with a highly polygenic basis of morph differentiation with population-specific selection on alleles linked to the development of similar morph phenotypes. The results support a model where morph differentiation is first established through phenotypic plasticity, leading to niche expansion and separation. This may be followed by gradual development of reproductive isolation, locus-specific differentiation and eventually complete reproductive isolation and genome-wide divergence.

Whole Genome Sequencing