PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “variant interpretation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The Subtle Crisis: Public Domain Genomes and the Ethics of Translational Infrastructure.

Public domain human genomic resources are infrastructure: tools researchers use to ask basic biological questions whose answers are then translated into products and care. Translational science now asks them to support an expanding set of tasks, including clinical variant interpretation for diverse populations, pharmacogenomic prescribing, polygenic risk prediction, and the training of clinical artificial intelligence. The corpus of public domain genomes, due in large part to upstream recruitment choices, is not fit for these purposes, and the gap between discovery and translation is widening. This essay argues that closing the gap requires treating public domain genomic infrastructure as a particular object of translational bioethics rather than a technical precondition for it. The limited number of genomes in the public domain relative to the broader genomic record, and the typology-friendliness of how that record represents human variation, are two faces of the same set of upstream choices. Reversing them is not a matter of more sampling under existing terms; it is a matter of building infrastructure of a particular kind; infrastructure made from people. That category, common in genomics but absent from the rest of science, demands an ethical apparatus the field has not yet built. Here we consider the commitments such an apparatus requires, and argue that where, how, and with whom we build genomic infrastructure is itself an ethics question the field has largely declined to ask.

Humans↗

Next-generation newborn screening: feasibility of combined genetic and biochemical testing for 95 treatable inherited metabolic disorders.

INTRODUCTION: Next-generation sequencing (NGS) is gaining attention in newborn screening (NBS) for its ability to detect treatable genetic disorders, especially those without a biochemical footprint. However, NGS-NBS requires interpreting variants without phenotype information or family trio analysis. Biochemical tests, preferably in dried blood spots (DBS), are therefore useful to confirm the pathogenicity of variants identified by NGS-NBS and increase its specificity and sensitivity. OBJECTIVES: We aimed to explore the potential of combined genetic-biochemical testing for 95 treatable Inherited Metabolic Disorders (IMD) considered eligible for NGS-NBS (100 genes) previously identified by our research group. METHODS: We reviewed the Collaborative Laboratory Integrated Reports (CLIR) and carried out systematic literature reviews in PubMed and Embase to identify biochemical tests for 95 IMD. Biochemical tests conducted on DBS were differentiated from tests that require referral. RESULTS: We identified DBS-biochemical tests for 72 of the 95 IMD (77/100 genes). DBS-based biochemical tests for 55 IMD (60 genes) are already implemented in NBS. For the other 23 IMD, biochemical tests in non-DBS specimens are reported, although some are less sensitive when measured at neonatal age in presymptomatic infants. CONCLUSION: We present a comprehensive overview of current biochemical tests for 95 IMD. These tests can be used to confirm inconclusive NGS-NBS results, and combined genetic-biochemical testing is expected to improve both the negative and positive predictive values of NBS programs.

Humans↗

Quantitative natural history modeling of HPDL-related disease based on cross-sectional data reveals genotype-phenotype correlations.

PURPOSE: Biallelic HPDL variants have been identified as the cause of a progressive childhood-onset movement disorder, with a broad clinical spectrum from severe neurodevelopmental disorder to juvenile-onset pure hereditary spastic paraplegia type 83. This study aims at delineating the geno- and phenotypic spectra of patients with HPDL-related disease, quantitatively modeling the natural history, and uncovering genotype-phenotype associations. METHODS: A cross-sectional analysis of 90 published and 1 novel case was performed, using a Human-Phenotype-Ontology-based approach. Unsupervised phenotypic clustering was used alongside in silico analyses to identify distinct patient subgroups. RESULTS: The study models the natural history of the HPDL-related disease in a global cohort, clarifying the molecular and phenotypic spectrum and identifying 3 distinct subgroups characterized by differences in onset, clinical trajectories, and survival. It establishes genotype-phenotype associations, showing that the presence of moderately pathogenic missense variants in 1 allele leads to a milder, spastic paraplegic phenotype with later disease onset, whereas biallelic, highly pathogenic missense or truncating variants are associated with a more severe phenotype and reduced life span. CONCLUSION: Quantitative and unbiased natural history modeling in HPDL-related disease reveals significant genotype-phenotype associations, providing a foundation for variant interpretation, anticipatory guidance, and choice of outcome measures in future prospective and functional studies.

Humans↗

Charting the phenotypic landscape of mitochondrial diseases through a systematic evaluation of pathogenic mitochondrial DNA and nuclear gene variants.

PURPOSE: Primary mitochondrial diseases (PMD) arise from variants in the mitochondrial or nuclear genomes. Phenotype-based recognition of specific PMD genotypes remains difficult, prolonging the diagnostic odyssey. We expanded the MitoPhen database to characterize phenotypic variation across PMD more systematically. METHODS: Individual-level data on mitochondrial DNA disorders, nuclear-encoded mitochondrial diseases, and single large-scale mitochondrial DNA deletions were manually curated with Human Phenotype Ontology (HPO) terms to produce MitoPhen v2. Principal-component analysis summarized system-level abnormalities; HPO-level enrichment and mean phenotype-similarity scores were then used to distinguish common PMD genotypes. RESULTS: MitoPhen v2 adds 3940 individuals to the original release, now encompassing 1597 publications, 10,626 individuals, and 117 genotypes. Among 7586 affected cases, 72,861 HPO terms were recorded. Principal-component analysis revealed 6 phenotype dimensions capturing most system-level variance. At the HPO level, we observed genotype-specific enrichments and identified 111 gene-phenotype links absent from the current HPO database. Using MT-TL1, single large-scale mitochondrial DNA deletions, and POLG as exemplars, phenotype-similarity scores reliably separated individuals with these genotypes from those without. CONCLUSION: MitoPhen v2 enabled systematic, genotype-aware analysis of heterogeneous PMD phenotypes and highlighted the diagnostic value of structured, individual-level data. Phenotype-similarity metrics from such data sets can refine variant interpretation in large rare-disease cohorts and provide a transferable framework for other phenotypically complex genetic disorders.

Humans↗

Integration of multi-omics data uncovers novel germline susceptibility candidates in early-onset colorectal cancer.

Colorectal cancer (CRC) is increasingly diagnosed in individuals under 50 years of age, yet the underlying genetic predisposition remains largely unexplained, particularly in mismatch repair (MMR)-proficient cases. This study aimed to identify novel hereditary CRC susceptibility genes by integrating germline and tumour whole-exome sequencing (WES) with transcriptomic profiling across a cohort of early-onset CRC (EOCRC) patients. Tumours were categorised using Consensus Molecular Subtypes (CMS) classification and analysed for mutational signature and burden. We used a novel 'All vs One' multi-omic integration approach to identify loss-of-function rare germline variants with concordant gene expression alterations in tumour tissue. Five candidate genes (ADCY4, NOXO1, CDHR2, ARHGAP10, EEF2K) were prioritised based on this approach and potential biological relevance in CRC. These findings highlight the molecular heterogeneity of EOCRC and demonstrate the utility of multi-omic approaches in refining germline variant interpretation. Integrating tumour transcriptomics enhances gene discovery efforts and supports a more comprehensive understanding of CRC heritability in younger individuals.

Humans↗

Genomic and computational analysis of variants in telomere regulatory genes in subjects with bone marrow failure.

Telomere Biology Disorders (TBDs) are a genetically heterogeneous and often under-recognized cause of Bone Marrow Failure Syndromes (BMFS), driven by defective telomere maintenance and progressive telomere attrition. We performed an integrated genomic, telomeric and computational analysis in 118 subjects presenting clinical features of BMFS to delineate the contribution of Telomere Regulatory Genes (TRGs) variants to disease pathogenesis. Whole exome sequencing (WES) identified pathogenic (18.18%), likely pathogenic (27.27%) and rare variants of uncertain significance (54.54%) in 27 subjects (22.9%) across five TRGs: RTEL1, TERT, TINF2, NOP10, and WRAP53. Telomere Length (TL) assessment revealed significant telomere shortening in TRG variant-positive subjects compared with age-matched controls, with the most profound attrition observed in individuals harboring de novo TINF2 gene variants. RTEL1 emerged as the most frequently affected gene, with recurrent clustering of variants within its C-terminal regulatory region. A familial NOP10 variant, Asp12His, segregated with cutaneous pigmentation and hematological abnormalities consistent with the established role of NOP10 in dyskeratosis congenita, further broadening the known mutational spectrum of the gene. Structure-guided in-silico analyses predicted that both novel and recurrent variants disrupt protein stability, telomerase assembly or trafficking and shelterin complex integrity. Reduced TERT expression and a significant inverse correlation between telomere length and clinical severity further underscored the functional impact of TRG defects. Collectively, this study provides the first comprehensive characterization of TRG variants in the Indian BMFS cohort and highlights the utility of integrating genomic sequencing, telomere length measurement and computational modeling to improve diagnostic precision, variant interpretation and clinical stratification in TBDs.

Journal Article↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Prediction of human missense variant effects from functional evidence.

Prediction of missense variant effects remains a critical bottleneck in both research and diagnostic genetics. Current predictors typically rely on clinical outcomes or population patterns rather than direct measures of functional impact, leading to limited generalizability and data circularity. Here we present FuncVEP, a family of variant effect predictors trained on diverse functional data to predict the functional impact of missense variants. FuncVEP generalizes across datasets and outperforms 48 existing predictors across a wide range of benchmarks, improving accuracy from 78.8% to 84.6% on functional benchmarks and from 90.1% to 92.4% on clinical benchmarks. From a discovery perspective, we identified 210 new gene-phenotype associations involving 494 genes linked to inborn errors of immunity in the UK Biobank and the Mount Sinai Million Health Discoveries Program. FuncVEP substantially improved the discovery rate relative to state-of-the-art predictors. Overall, FuncVEP provides a robust, scalable solution for variant interpretation, advancing both diagnostic precision and gene discovery.

Humans↗

Clinical proteomics in inborn errors of metabolism: from biomarker discovery to implementation.

INTRODUCTION: Inborn errors of metabolism (IEMs) are rare, heterogeneous disorders traditionally diagnosed through genetic testing, enzyme assays, and metabolite measurements. However, these tools often do not fully explain phenotypic variability, organ involvement, disease progression, or treatment response. Clinical proteomics provides a complementary functional layer by capturing changes in protein abundance, proteoforms, post-translational modifications (PTM), and biological pathways, offering insights beyond genotype- and metabolite-based approaches. AREAS COVERED: This review examines the role of high-resolution mass spectrometry and computational proteomics in biomarker discovery and clinical decision-making for IEMs. It focuses on their contribution to diagnosis, variant interpretation, patient stratification, and treatment monitoring. Disease-specific applications are discussed, with the strongest evidence in lysosomal storage disorders, mitochondrial diseases, congenital disorders of glycosylation, and selected neurodegenerative or renal metabolic conditions. The literature search was performed in PubMed, Scopus, Web of Science, and Google Scholar, covering peer-reviewed articles available up to 2026, with emphasis on methodological advances and translational applications in clinical proteomics for IEMs. EXPERT OPINION: Proteomics will not replace established diagnostic tools, but it can help address clinically actionable questions in selected contexts. Translation into clinical practice will require standardized workflows, multicenter validation, clinically anchored endpoints, and integration with other omics approaches.

Humans↗

Diagnostic testing for Rett syndrome by DHPLC and direct sequencing analysis of the MECP2 gene: identification of several novel mutations and polymorphisms.

Rett syndrome (RTT) is an X-linked dominant neurodevelopmental disorder affecting 1/10,000-15,000 girls. The disease-causing gene was identified as MECP2 on chromosome Xq28, and mutations have been found in approximately 80% of patients diagnosed with RTT. Numerous mutations have been identified in de novo and rare familial cases, and they occur primarily in the methyl-CpG-binding and transcriptional-repression domains of MeCP2. Our first diagnostic strategy used bidirectional sequencing of the entire MECP2 coding region. Subsequently, we implemented a two-tiered strategy that used denaturing high-performance liquid chromatography (DHPLC) for initial screening of nucleotide variants, followed by confirmatory sequencing analysis. If a definite mutation was not identified, then the entire MECP2 coding region was sequenced, to reduce the risk of false negatives. Collectively, we tested 228 unrelated female patients with a diagnosis of possible (209) or classic (19) RTT and found MECP2 mutations in 83 (40%) of 209 and 16 (84%) of 19 of the patients, respectively. Thirty-two different mutations were identified (8 missense, 9 nonsense, 1 splice site, and 14 frameshifts), of which 12 are novel and 9 recurrent in unrelated patients. Seven unclassified variants and eight polymorphisms were detected in 228 probands. Interestingly, we found that T203M, previously reported as a missense mutation in an autistic patient, is actually a benign polymorphism, according to parental analysis performed in a second case identified in this study. These findings highlight the complexities of missense variant interpretation and emphasize the importance of parental DNA analysis for establishing an etiologic relation between a possible mutation and disease. Overall, we found a 98.8% concordance rate between DHPLC and sequence analyses. One mutation initially missed by the DHPLC screening was identified by sequencing. Modified conditions subsequently enabled its detection, underscoring the need for multiple optimized conditions for DHPLC analysis. We conclude that this two-tiered approach provides a sensitive, robust, and efficient strategy for RTT molecular diagnosis.

Alleles↗

PanelAppRex aggregates disease gene panels and facilitates sophisticated search.

MOTIVATION: Gene panel data are essential for variant interpretation and genomic diagnostics, but existing resources are fragmented, inconsistently annotated, and not easily accessible for programmatic use. We developed PanelAppRex, a harmonised dataset and interactive search tool that integrates over 58 000 curated gene-disease panel associations. It supports natural language-style queries by gene, phenotype, disease group, and mode of inheritance, with results returned in machine-readable export formats. RESULTS: The resulting dataset includes standardised gene identifiers, disease annotations, mode of inheritance, and literature support, enabling seamless integration into bioinformatic pipelines. We benchmarked 15 case studies spanning immunology, neurology, and additional disease areas. Under the recommended usage, in which the union of returned panels is considered, the causal gene was recovered in every case. Across all returned panels, the causal gene was present in 85.6% of panels. For manual interface interpretation, the causal gene was present in the user-selected best-fit panel(s) in all 15 benchmarked cases. AVAILABILITY: The platform data is openly available at Zenodo https://doi.org/10.5281/zenodo.15736689, with source code at https://github.com/DylanLawless/PanelAppRex, and demonstration page at https://panelapprex.github.io/landing_page. The dataset is maintained for a minimum of two years following publication.

Journal Article↗

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans↗

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics↗

RNA splicing and cardiovascular disease: a guide for cardiologists.

Alternative splicing (AS) is a fundamental RNA processing mechanism, which generates different RNA transcripts and consequently different protein isoforms from a single gene. This increases the diversity of proteins within an organism and can fine-tune biological processes. This review examines how cardiac-enriched RNA-binding proteins establish heart-specific splicing programs governing aspects of cardiac development, function, and disease. Developmentally, coordinated sarcomeric isoform switches underpin the foetal-to-adult transition and further isoform rewiring in ion channel and kinase genes determine electrophysiology and excitation-contraction coupling. AS contributes to the pathogenesis of several cardiomyopathies and emerging datasets suggest that pathological hypertrophy engages distinct splicing signatures compared with physiological hypertrophy. This review summarizes diagnostic and prognostic opportunities arising from bulk, long-read, and single-cell/nucleus transcriptomics, which resolve cell type-specific isoforms and disease-associated switches. Circulating RNA biomarkers (including splice ratios and circularRNAs) may signify myocardial remodelling and arrhythmic risk. Integrative approaches that link AS with proteomics and genomics improve variant interpretation, reveal previously unannotated protein isoforms, and enable tracking of disease progression and therapy response. Finally, an outline of therapeutic strategies to modulate AS in cardiovascular disease (CVD), including antisense oligonucleotides, small molecules, and genome-editing modalities (CRISPR, base, and prime editing), is provided. The major challenges that remain before splice-targeting therapeutics can be targeted to treat cardiovascular disease are highlighted. Lessons from neuromuscular indications establish clinical feasibility of splicing correction and motivate translation to cardiology. Together, mechanistic insight, biomarker development, and therapeutic innovation position RNA splicing as a tractable axis for precision cardiovascular medicine.

Humans↗

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals↗

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Short-read sequencing cannot determine whether co-occurring variants within a cancer gene lie on the same allele (cis) or opposing alleles (trans), a distinction with direct therapeutic consequences: trans configurations confirm biallelic tumor suppressor inactivation, whereas cis configurations generate compound oncogenic alleles with enhanced activity. Among 768 patients with prostate, breast, or ovarian cancers, we used mutational signatures to nominate cryptic genomic instability cases lacking a causative biallelic event on short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%) through methylation detection, long insertion resolution, and structural variant characterization, confirming trans inactivation in every resolved tumor suppressor case. Analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, 78.7% of which exceeded the 500 bp short-read phasing limit, and long-read phasing revealed recurrent compound cis alleles in NOTCH1, PIK3CA, PDGFRB, and KIT. Haplotype phasing addresses an overlooked gap in cancer variant interpretation and warrants integration into precision oncology.

Journal Article↗

Proteome-scale prediction of molecular mechanisms underlying dominant genetic diseases.

Many dominant genetic disorders result from protein-altering mutations, acting primarily through dominant-negative (DN), gain-of-function (GOF), and loss-of-function (LOF) mechanisms. Deciphering the mechanisms by which dominant diseases exert their effects is often experimentally challenging and resource intensive, but is essential for developing appropriate therapeutic approaches. Diseases that arise via a LOF mechanism are more amenable to be treated by conventional gene therapy, whereas DN and GOF mechanisms may require gene editing or targeting by small molecules. Moreover, pathogenic missense mutations that act via DN and GOF mechanisms are more difficult to identify than those that act via LOF using nearly all currently available variant effect predictors. Here, we introduce a tripartite statistical model made up of support vector machine binary classifiers trained to predict whether human protein coding genes are likely to be associated with DN, GOF, or LOF molecular disease mechanisms. We test the utility of the predictions by examining biologically and clinically meaningful properties known to be associated with the mechanisms. Our results strongly support that the models are able to generalise on unseen data and offer insight into the functional attributes of proteins associated with different mechanisms. We hope that our predictions will serve as a springboard for researchers studying novel variants and those of uncertain clinical significance, guiding variant interpretation strategies and experimental characterisation. Predictions for the human UniProt reference proteome are available at https://osf.io/z4dcp/.

Humans↗

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins.

INTRODUCTION: The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. METHODS: We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. RESULTS: We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. DISCUSSION: The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

DNA-binding proteins↗