PubMed HealthSearch

SEARCH · PubMed Health

Results for “variant calling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

CBIcall: a configuration-driven framework for variant calling in large sequencing cohorts.

MOTIVATION: Variant calling for next-generation sequencing (NGS) data relies on a diverse ecosystem of tools and workflows. Large-scale collaborative studies increasingly adopt federated analysis, where each institution processes sensitive data locally using standardized pipelines. Deploying identical pipelines across multiple centers remains challenging because heterogeneous software environments and computing policies can cause workflow divergence and inconsistent results. RESULTS: We developed CBIcall, a workflow backend-flexible, configuration-driven framework that runs standardized variant-calling pipelines from raw FASTQ files to analysis-ready VCFs. Users define each analysis in a single YAML parameters file, which CBIcall resolves against a controlled workflow registry and resource catalog. The execution driver validates parameters and checks compatibility among pipelines, analysis modes, workflow backends, genome builds, tool versions, and resource bundles. CBIcall supports reproducibility auditing by comparing executions using recorded provenance and output fingerprints. CBIcall dispatches validated workflows natively through Bash, Cromwell, Nextflow and Snakemake backends and provides production-ready pipelines for germline WES, WGS (single-sample or cohort joint genotyping following GATK Best Practices), and mitochondrial DNA analysis. We evaluated analytical performance using public benchmark datasets and validated reproducibility across four computing environments. We further deployed CBIcall in the EU HEREDITARY project, where it processed 1102 samples with both WES and mtDNA pipelines on an institutional HPC system, supporting its suitability for reproducible cohort-scale genomic analyses. AVAILABILITY AND IMPLEMENTATION: CBIcall is open source (GPLv3) and distributed with ready-to-run pipelines; full dependency and installation documentation is available at https://github.com/CNAG-Biomedical-Informatics/cbicall.

Journal Article

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Accelerated long-read variant calling with Clair3 for whole-genome sequencing.

SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30× whole-genome sequence in 12-20 minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15 minutes on an Apple Mac Studio (32 threads), which is ∼10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30× ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.

Whole Genome Sequencing

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Diploid genome assembly of human fibroblast cell lines enables clone specific variant calling, improved read mapping and accurate phasing.

Human cell lines are fundamental tools in biomedical research and are widely used in disease modeling, drug development, and many other domains. Here, we present chromosome-level, phased diploid genome assemblies of two popular human cell lines: the BJ foreskin fibroblast line and the IMR-90 fetal lung fibroblast line. Our high-quality assemblies, generated using long-read and Hi-C sequencing data, reveal substantial structural variation, including more than 50,000 insertions, deletions, duplications, and inversions compared to the recent T2T-CHM13v2.0 reference. Our assemblies provide detailed maps of genetic variation, enabling more accurate variant calling and the ability to phase reads when using newly generated or historical sequencing data on these cell lines or their derivatives. All assemblies and associated data have been made available as a resource for the research community. We envision that diploid genome assembly will become a cornerstone approach for personalized medicine in the near future.

Journal Article

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

Whole-Exome and Whole-Genome Sequencing of Candidate Pharmacogenomic and Schizophrenia-Related Genes in Sudanese Families with Schizophrenia.

BACKGROUND: Schizophrenia is considered a neuro-developmental disorder leading to disastrous lifelong disability of the patients and their families. There is a lack of data regarding pharmacogenomics of schizophrenia in Sudan. This study aimed to identify different genes affecting the treatment outcomes in Sudanese patients with schizophrenia. METHODS: A case-control study was conducted on seven families having more than one member diagnosed with schizophrenia. This was a small exploratory family-based sequencing study involving 18 affected individuals and 8 controls from seven families. Ethical clearance and informed consent were obtained. Demographic data were collected using a standardized data collection sheet. DNA was extracted from blood samples collected from patients and control groups. Then, whole-exome and genome sequencing were performed. Sixty-six genes associated with schizophrenia, treatment, and treatment resistance were selected from the variant calling file. Variants showing single-nucleotide polymorphisms (SNPs) were identified. These variants were then classified based on their impact on the protein-coding sequence into high- and moderate-impact. Moreover, indel mutations were also identified. RESULTS: Twelve variants of seven genes (COMT, FMO1, LPL, CYP2E1, ABCC1, GRM3, CYP2C9) were identified as genes with impact and potential association with schizophrenia (p-value=0.006632). Forty-three genes had a moderate impact, and they showed a potential association with schizophrenia (p-value=0.0004436). Two variants were indel mutations (CYP2D6, DTNBP1) and showed association with schizophrenia (p-value=0.004741). The p-values were generated from different databases. CONCLUSION: This exploratory family-based sequencing study identified several potentially relevant pharmacogenomic and schizophrenia-associated variants in Sudanese families, warranting validation in larger and ethnically diverse cohorts.

antipsychotics

Disruption of GAD1 protein architecture by a novel missense variant in a consanguineous family with autosomal recessive intellectual disability.

BACKGROUND: Intellectual disability represents a heterogeneous group of neurodevelopmental disorders marked by significant impairments in intellectual functioning and adaptive behavior. Among the various causes, genetic factors play a major role, with autosomal recessive intellectual disability (ARID) constituting a genetically diverse subgroup. ARID is prevalent in consanguineous families and arises from homozygous mutations that disrupt critical genes involved in brain development and function. OBJECTIVE: This study aimed to identify disease-causing genetic variants responsible for ARID in a consanguineous Pakistani family and to evaluate the structural and functional impact of a novel variant identified in GAD1 through protein modeling. METHODS: A consanguineous family affected with intellectual disability was enrolled. Whole-exome sequencing was performed on an affected individual, followed by bioinformatics analysis including alignment to the GRCh38 reference genome, variant calling, and annotation. Variants were filtered based on rarity, predicted functional impact, and autosomal recessive inheritance pattern. Candidate variants were validated and assessed by Sanger sequencing and segregation analysis. Protein modeling was performed to evaluate the structural impact of the identified variant. RESULTS: A novel homozygous missense variant NM_000817:c.1700G>A;p.Arg567Gln in GAD1 was identified. Segregation analysis confirmed co-segregation of the variant with the affected phenotype. Protein modeling suggested that the variant may disrupt GAD1 enzymatic function involved in gamma-aminobutyric acid synthesis. CONCLUSION: This study emphasizes the significance of genetic investigation in familial cases and the crucial role that GAD1 mutations play in neurodevelopmental disorders with intellectual disability. The results advance the knowledge of molecular causes of ARID and broaden the mutational range.

Pakistani

Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants.

Variant interpretation in the era of massively parallel sequencing is challenging. Although many resources and guidelines are available to assist with this task, few integrated end-to-end tools exist. Here, we present the Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE), a web- and cloud-based platform for annotation, identification, and classification of variations in known or putative disease genes. Starting from a set of variants in variant call format (VCF), variants are annotated, ranked by putative pathogenicity, and presented for formal classification using a decision-support interface based on published guidelines from the American College of Medical Genetics and Genomics (ACMG). The system can accept files containing millions of variants and handle single-nucleotide variants (SNVs), simple insertions/deletions (indels), multiple-nucleotide variants (MNVs), and complex substitutions. PeCanPIE has been applied to classify variant pathogenicity in cancer predisposition genes in two large-scale investigations involving >4000 pediatric cancer patients and serves as a repository for the expert-reviewed results. PeCanPIE was originally developed for pediatric cancer but can be easily extended for use for nonpediatric cancers and noncancer genetic diseases. Although PeCanPIE's web-based interface was designed to be accessible to non-bioinformaticians, its back-end pipelines may also be run independently on the cloud, facilitating direct integration and broader adoption. PeCanPIE is publicly available and free for research use.

Child

Clinical Variant Interpretation with the Integrative Genomics Viewer (IGV) for Molecular Pathologists.

The integrative genomics viewer (IGV) is a pivotal tool in clinical genomics, enabling the visualization and interpretation of complex sequencing data. Bringing clinical knowledge to bear with visual evaluation of sequencing results is the primary means by which molecular pathologists and other professionals assess and finalize cases. A variety of software tools can assist, but their relationship to the underlying data must be understood and applied systematically. This study includes essential background on next-generation sequencing (NGS) data file types (e.g., FASTQ, BAM, VCF) with a discussion of their format and purpose. We then describe features of IGV that derive nuances from these files. We utilize a series of curated practical cases based on clinical vignettes through which the reader will interact with clinical NGS sequencing data using the IGV software to review various types of clinically relevant variants relative to the human reference genome. These clinical vignettes have been curated to describe examples of some of the complexities of interpretation of genomic data, and how utilizing IGV as part of a routine workflow can provide additional interpretive information for variants beyond routine bioinformatic software algorithm variant calls. The visual inspection of genomic variants utilizing the tools within IGV can unmask subtle contextual cues (i.e., variant allele frequency, strand bias, tissue-specific context) that can influence the interpretation of genomic variants. Although this study focuses on using IGV for the detection and interpretation of somatic variants, the provided applications can be extrapolated for use in the germline setting, including analysis of complex variants and detection of mosaicism.

Humans