PubMed HealthSearch

PubMed · 41883150

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

Abstract

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Barthélémy Caron, Antonio Rausell. 2026-05-03. NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.. https://doi.org/10.1093/bioinformatics%2Fbtag138

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

European ash pangenome reveals widespread structural variation and genetic basis of low ash dieback susceptibility.

European Ash (Fraxinus excelsior) is a keystone tree species, whose populations are being decimated by ash dieback disease (ADB) - better characterisation of genetic variants associated with low susceptibility to the disease is needed. Here, we develop a F. excelsior pangenome to more fully capture sequence variability within this species compared with a linear reference genome, using a geographically diverse set of fifty F. excelsior samples. We identify 362,965 structural variants (SVs), including 174 Mb of sequence absent from the linear reference genome (22% of the linear reference size), and identify 3,412 high-confidence dispensable genes (those present only in some individuals). We use the pangenome to analyse existing genomic data from over 1,200 individuals, revealing 220 single nucleotide polymorphisms (SNPs) showing consistent allele frequency shifts between healthy individuals and those highly damaged by ADB, across UK seed sources, explicitly demonstrating the existence of a shared genetic component to low ADB susceptibility.

Polymorphism, Single Nucleotide

Identification of Candidate Genes Associated with Growth Traits in Procambarus clarkii Using Whole-Genome Resequencing.

Growth is a critical economic trait in all aquaculture industries. To address issues such as germplasm degradation, a comprehensive understanding of the growth and development mechanisms, along with genetic improvement strategies, for Procambarus clarkii (P. clarkii) is urgently required. In this study, we performed whole-genome resequencing on 89 individuals from five cultured stocks to investigate growth traits (body length) and identified a total of 46,919,297 high-quality single nucleotide polymorphisms (SNPs). Based on these SNPs, we conducted principal component analysis (PCA), phylogenetic analysis, and population genetic structure analysis. Furthermore, we performed selective sweep analysis (using FST, Pi, and XP-CLR) and a genome-wide association study (GWAS) to identify genetic variants associated with growth traits. The results revealed significant genetic differentiation among the five cultured stocks, with the Ma'anshan cultured stock exhibiting the fastest linkage disequilibrium (LD) decay. Additionally, long-term aquaculture in different geographical regions resulted in distinct genetic differences among cultured stocks. Through selective sweep analysis, the intersection of FST, Pi, and XP-CLR across the five populations yielded several growth-related candidate genes: Nephrin, Somatostatin, zinc finger protein 154, and yeti. Subsequent the GWAS identified two candidate genes associated with growth traits: Cullin-associated and neddylation-dissociated protein 1 (CAND1) and Baculoviral IAP repeat-containing protein 8 (BIRC8). These genes are presumed to play pivotal roles in the growth and development of P. clarkii. Overall, our findings provide new insights into the genetic mechanisms underlying growth and development in P. clarkii, and these identified genes serve as promising candidates for further functional studies and genetic improvement of this species.

Polymorphism, Single Nucleotide

IBAS: Interaction-bridged association studies discovering novel genes underlying complex traits.

Genetic contributions to complex traits are often mediated through coordinated gene-gene interaction networks, yet most existing association frameworks focus on marginal single-gene effects and overlook higher-order dependency structures. Direct modeling of interactions remains challenging due to combinatorial complexity and statistical instability. We introduce Interaction-Bridged Association Study (IBAS), a general framework that incorporates pathway-level interaction patterns into genotype-phenotype association analysis without explicitly enumerating interactions. IBAS leverages transcriptomic reference data to construct low-dimensional representations of pathway activity, which guide SNP-weighting and gene-level association testing within a kernel-based framework. In perturbation-based simulations, IBAS demonstrates improved stability and reproducibility compared to conventional TWAS and gene-based methods, while maintaining well-calibrated Type I error under phenotype permutation. Application to the WTCCC datasets identifies both known and novel genes across multiple complex diseases, including candidates with modest marginal effects missed by standard approaches. These findings are supported by replication in an independent cohort, and analyses across multiple reference tissues revealing both shared and tissue-specific signals. Overall, IBAS provides a statistically robust and computationally tractable framework for incorporating interaction effects into association mapping, extending beyond the single-gene paradigm and enabling more comprehensive characterization of complex trait. IBAS is available on GitHub at: https://github.com/QingrunZhangLab/IBAS.

Polymorphism, Single Nucleotide