PubMed HealthSearch

SEARCH · PubMed Health

Results for “Variant interpretation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Genomic and computational analysis of variants in telomere regulatory genes in subjects with bone marrow failure.

Telomere Biology Disorders (TBDs) are a genetically heterogeneous and often under-recognized cause of Bone Marrow Failure Syndromes (BMFS), driven by defective telomere maintenance and progressive telomere attrition. We performed an integrated genomic, telomeric and computational analysis in 118 subjects presenting clinical features of BMFS to delineate the contribution of Telomere Regulatory Genes (TRGs) variants to disease pathogenesis. Whole exome sequencing (WES) identified pathogenic (18.18%), likely pathogenic (27.27%) and rare variants of uncertain significance (54.54%) in 27 subjects (22.9%) across five TRGs: RTEL1, TERT, TINF2, NOP10, and WRAP53. Telomere Length (TL) assessment revealed significant telomere shortening in TRG variant-positive subjects compared with age-matched controls, with the most profound attrition observed in individuals harboring de novo TINF2 gene variants. RTEL1 emerged as the most frequently affected gene, with recurrent clustering of variants within its C-terminal regulatory region. A familial NOP10 variant, Asp12His, segregated with cutaneous pigmentation and hematological abnormalities consistent with the established role of NOP10 in dyskeratosis congenita, further broadening the known mutational spectrum of the gene. Structure-guided in-silico analyses predicted that both novel and recurrent variants disrupt protein stability, telomerase assembly or trafficking and shelterin complex integrity. Reduced TERT expression and a significant inverse correlation between telomere length and clinical severity further underscored the functional impact of TRG defects. Collectively, this study provides the first comprehensive characterization of TRG variants in the Indian BMFS cohort and highlights the utility of integrating genomic sequencing, telomere length measurement and computational modeling to improve diagnostic precision, variant interpretation and clinical stratification in TBDs.

Journal Article

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models

Prediction of human missense variant effects from functional evidence.

Prediction of missense variant effects remains a critical bottleneck in both research and diagnostic genetics. Current predictors typically rely on clinical outcomes or population patterns rather than direct measures of functional impact, leading to limited generalizability and data circularity. Here we present FuncVEP, a family of variant effect predictors trained on diverse functional data to predict the functional impact of missense variants. FuncVEP generalizes across datasets and outperforms 48 existing predictors across a wide range of benchmarks, improving accuracy from 78.8% to 84.6% on functional benchmarks and from 90.1% to 92.4% on clinical benchmarks. From a discovery perspective, we identified 210 new gene-phenotype associations involving 494 genes linked to inborn errors of immunity in the UK Biobank and the Mount Sinai Million Health Discoveries Program. FuncVEP substantially improved the discovery rate relative to state-of-the-art predictors. Overall, FuncVEP provides a robust, scalable solution for variant interpretation, advancing both diagnostic precision and gene discovery.

Humans

Clinical proteomics in inborn errors of metabolism: from biomarker discovery to implementation.

INTRODUCTION: Inborn errors of metabolism (IEMs) are rare, heterogeneous disorders traditionally diagnosed through genetic testing, enzyme assays, and metabolite measurements. However, these tools often do not fully explain phenotypic variability, organ involvement, disease progression, or treatment response. Clinical proteomics provides a complementary functional layer by capturing changes in protein abundance, proteoforms, post-translational modifications (PTM), and biological pathways, offering insights beyond genotype- and metabolite-based approaches. AREAS COVERED: This review examines the role of high-resolution mass spectrometry and computational proteomics in biomarker discovery and clinical decision-making for IEMs. It focuses on their contribution to diagnosis, variant interpretation, patient stratification, and treatment monitoring. Disease-specific applications are discussed, with the strongest evidence in lysosomal storage disorders, mitochondrial diseases, congenital disorders of glycosylation, and selected neurodegenerative or renal metabolic conditions. The literature search was performed in PubMed, Scopus, Web of Science, and Google Scholar, covering peer-reviewed articles available up to 2026, with emphasis on methodological advances and translational applications in clinical proteomics for IEMs. EXPERT OPINION: Proteomics will not replace established diagnostic tools, but it can help address clinically actionable questions in selected contexts. Translation into clinical practice will require standardized workflows, multicenter validation, clinically anchored endpoints, and integration with other omics approaches.

Humans

PanelAppRex aggregates disease gene panels and facilitates sophisticated search.

MOTIVATION: Gene panel data are essential for variant interpretation and genomic diagnostics, but existing resources are fragmented, inconsistently annotated, and not easily accessible for programmatic use. We developed PanelAppRex, a harmonised dataset and interactive search tool that integrates over 58 000 curated gene-disease panel associations. It supports natural language-style queries by gene, phenotype, disease group, and mode of inheritance, with results returned in machine-readable export formats. RESULTS: The resulting dataset includes standardised gene identifiers, disease annotations, mode of inheritance, and literature support, enabling seamless integration into bioinformatic pipelines. We benchmarked 15 case studies spanning immunology, neurology, and additional disease areas. Under the recommended usage, in which the union of returned panels is considered, the causal gene was recovered in every case. Across all returned panels, the causal gene was present in 85.6% of panels. For manual interface interpretation, the causal gene was present in the user-selected best-fit panel(s) in all 15 benchmarked cases. AVAILABILITY: The platform data is openly available at Zenodo https://doi.org/10.5281/zenodo.15736689, with source code at https://github.com/DylanLawless/PanelAppRex, and demonstration page at https://panelapprex.github.io/landing_page. The dataset is maintained for a minimum of two years following publication.

Journal Article

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

RNA splicing and cardiovascular disease: a guide for cardiologists.

Alternative splicing (AS) is a fundamental RNA processing mechanism, which generates different RNA transcripts and consequently different protein isoforms from a single gene. This increases the diversity of proteins within an organism and can fine-tune biological processes. This review examines how cardiac-enriched RNA-binding proteins establish heart-specific splicing programs governing aspects of cardiac development, function, and disease. Developmentally, coordinated sarcomeric isoform switches underpin the foetal-to-adult transition and further isoform rewiring in ion channel and kinase genes determine electrophysiology and excitation-contraction coupling. AS contributes to the pathogenesis of several cardiomyopathies and emerging datasets suggest that pathological hypertrophy engages distinct splicing signatures compared with physiological hypertrophy. This review summarizes diagnostic and prognostic opportunities arising from bulk, long-read, and single-cell/nucleus transcriptomics, which resolve cell type-specific isoforms and disease-associated switches. Circulating RNA biomarkers (including splice ratios and circularRNAs) may signify myocardial remodelling and arrhythmic risk. Integrative approaches that link AS with proteomics and genomics improve variant interpretation, reveal previously unannotated protein isoforms, and enable tracking of disease progression and therapy response. Finally, an outline of therapeutic strategies to modulate AS in cardiovascular disease (CVD), including antisense oligonucleotides, small molecules, and genome-editing modalities (CRISPR, base, and prime editing), is provided. The major challenges that remain before splice-targeting therapeutics can be targeted to treat cardiovascular disease are highlighted. Lessons from neuromuscular indications establish clinical feasibility of splicing correction and motivate translation to cardiology. Together, mechanistic insight, biomarker development, and therapeutic innovation position RNA splicing as a tractable axis for precision cardiovascular medicine.

Humans

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Short-read sequencing cannot determine whether co-occurring variants within a cancer gene lie on the same allele (cis) or opposing alleles (trans), a distinction with direct therapeutic consequences: trans configurations confirm biallelic tumor suppressor inactivation, whereas cis configurations generate compound oncogenic alleles with enhanced activity. Among 768 patients with prostate, breast, or ovarian cancers, we used mutational signatures to nominate cryptic genomic instability cases lacking a causative biallelic event on short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%) through methylation detection, long insertion resolution, and structural variant characterization, confirming trans inactivation in every resolved tumor suppressor case. Analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, 78.7% of which exceeded the 500 bp short-read phasing limit, and long-read phasing revealed recurrent compound cis alleles in NOTCH1, PIK3CA, PDGFRB, and KIT. Haplotype phasing addresses an overlooked gap in cancer variant interpretation and warrants integration into precision oncology.

Journal Article

Proteome-scale prediction of molecular mechanisms underlying dominant genetic diseases.

Many dominant genetic disorders result from protein-altering mutations, acting primarily through dominant-negative (DN), gain-of-function (GOF), and loss-of-function (LOF) mechanisms. Deciphering the mechanisms by which dominant diseases exert their effects is often experimentally challenging and resource intensive, but is essential for developing appropriate therapeutic approaches. Diseases that arise via a LOF mechanism are more amenable to be treated by conventional gene therapy, whereas DN and GOF mechanisms may require gene editing or targeting by small molecules. Moreover, pathogenic missense mutations that act via DN and GOF mechanisms are more difficult to identify than those that act via LOF using nearly all currently available variant effect predictors. Here, we introduce a tripartite statistical model made up of support vector machine binary classifiers trained to predict whether human protein coding genes are likely to be associated with DN, GOF, or LOF molecular disease mechanisms. We test the utility of the predictions by examining biologically and clinically meaningful properties known to be associated with the mechanisms. Our results strongly support that the models are able to generalise on unseen data and offer insight into the functional attributes of proteins associated with different mechanisms. We hope that our predictions will serve as a springboard for researchers studying novel variants and those of uncertain clinical significance, guiding variant interpretation strategies and experimental characterisation. Predictions for the human UniProt reference proteome are available at https://osf.io/z4dcp/.

Humans

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins.

INTRODUCTION: The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. METHODS: We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. RESULTS: We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. DISCUSSION: The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

DNA-binding proteins

Novel compound heterozygous POR variants in a neonate with Antley-Bixler syndrome and 46,XY DSD: a case report and literature review.

BACKGROUND: Cytochrome P450 oxidoreductase deficiency (PORD) is an ultra-rare autosomal recessive disorder within the congenital adrenal hyperplasia (CAH) spectrum, characterized by a broad clinical spectrum involving steroidogenesis defects, genital anomalies, and skeletal abnormalities. CASE PRESENTATION: We report a phenotypically female neonate with a 46,XY karyotype whose postnatal diagnostic evaluation was initiated after newborn screening revealed elevated 17-hydroxyprogesterone (17-OHP) concentration. The patient presented with mild hypertelorism, mild nasal hypoplasia, and low-set bilateral ears, along with female external genitalia consistent with disorder of sex development (DSD) and anal atresia. Radiological evaluation revealed femoral bowing and subsequent fracture. The craniofacial and skeletal abnormalities were consistent with the features of Antley-Bixler syndrome (ABS). Endocrine evaluation revealed elevated progesterone, markedly reduced testosterone, and secondary hyperaldosteronism. Genetic analysis identified three novel variants in the POR gene (NM_001395413.1): the patient harbored a paternal c.1187_1195dup (p.Pro396_Glu398dup) variant and two maternally inherited variants in cis, c.1447G>A (p.Gly483Ser) and c.1806 + 4_1806 + 28del. Protein structural modeling predicted that the p.Pro396_Glu398dup and p.Gly483Ser may disrupt the flavin adenine dinucleotide (FAD)-binding domain. RNA sequencing (RNA-seq) confirmed that the intronic variant c.1806 + 4_1806 + 28del caused aberrant splicing, resulting in partial intron retention and predicted impairment of the nicotinamide adenine dinucleotide phosphate (NADPH)-binding domain. According to American College of Medical Genetics and Genomics (ACMG) guidelines and incorporating functional evidence, c.1187_1195dup and c.1806 + 4_1806 + 28del were reclassified as likely pathogenic (LP), whereas c.1447G>A remained a variant of uncertain significance (VUS). CONCLUSIONS: This study describes a neonate with PORD caused by three novel POR variants and expands the known clinical spectrum of PORD by identifying rare manifestations including anal atresia and hearing loss. RNA-seq provided valuable functional evidence for variant interpretation and facilitated accurate molecular diagnosis. These findings highlight the importance of integrating genetic phasing, transcript-level functional analysis, and comprehensive clinical evaluation for precise diagnosis and counseling in rare endocrine disorders.

Humans

Clinical implementation of next-generation sequencing in tertiary health system: the Rijeka retrospective study.

OBJECTIVE: This study aimed to assess the diagnostic yield, clinical indications, and utility of next-generation sequencing (NGS) testing since its implementation through collaboration between the University of Rijeka Faculty of Medicine and the Clinical Hospital Centre Rijeka. MATERIALS AND METHODS: This retrospective study included patients referred between 2018 and 2023 from the Clinical Hospital Centre Rijeka to the University of Rijeka Faculty of Medicine for genetic testing, primarily using exome sequencing. RESULTS: Between April 2018 and December 2023, 412 patients were referred for exome sequencing, of whom 353 (85.7%) underwent diagnostic genetic testing. A notable increase in tests ordered was observed over time. Patients were most frequently referred from Pediatrics (55.0%), Neurology (29.5%), Cardiology (7.4%), Ophthalmology (3.4%), and others (4.7%). A diagnosis was confirmed in 103/353 patients, corresponding to an overall diagnostic yield of 29.2%, and an adjusted diagnostic yield of 27.2% after collapsing related individuals into single family units. In these confirmed cases, 83 distinct disorders involving 71 unique genes were identified, with most patients showing heterozygous variants and several recurrent disorders and genes. Variants of uncertain significance were reported in 35/353 (9.9%) patients. CONCLUSION: The 27.2% diagnostic yield demonstrates effective integration of NGS into tertiary clinical practice. The recent introduction of medical genetics specialization is expected to further improve referral quality, variant interpretation, and overall diagnostic outcomes.

exome sequencing

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model

ATM Variants and Breast Cancer Risk in North Macedonia: Focus on the Regionally Enriched p.(Leu2492Arg) Variant.

BACKGROUND: Germline pathogenic variants (PVs) in the ataxia-telangiectasia mutated (ATM) gene are established moderate-risk factors for breast cancer (BC), however, population-specific variant spectra and the clinical significance of many missense variants remain incompletely characterized. AIMS: To evaluate the prevalence of ATM variants in a large cohort of patients with BC from North Macedonia and compare it with that in the general population, with a particular focus on the frequency of the p.(Leu2492Arg) variant and its distribution relative to global genomic datasets. STUDY DESIGN: Retrospective case–control study. METHODS: ATM variants were analyzed in 1,211 patients with BC from North Macedonia using a targeted hereditary cancer gene panel. These findings were compared with those from 1,303 population-based controls analyzed by clinical exome or whole-exome sequencing. RESULTS: Pathogenic ATM variants were identified in 1.9% of BC cases and 0.4% of controls, indicating a significantly increased risk of BC [odds ratio (OR) = 5.02, p = 0.0006]. Most PVs were protein-truncating, with six recurrent variants accounting for over 70% of detections, suggesting regional enrichment. Carriers showed a significantly higher prevalence of human epidermal growth factor receptor 2-positive tumors (OR = 2.92, p = 0.0189). Variants of uncertain significance were observed at comparable frequencies in cases and controls. The p.(Leu2492Arg) missense variant was more frequently detected in cases than in controls (1.9% vs. 1.1%; OR = 1.78, p = 0.086) and exhibited a markedly higher allele frequency in this population than in global databases. CONCLUSION: These findings confirm ATM as a clinically relevant BC susceptibility gene in North Macedonia and highlight the population-specific enrichment of both PVs and the p.(Leu2492Arg) missense variant. The results emphasize the importance of using population-matched controls and regional genomic data for accurate risk assessment and variant interpretation.

Humans

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Conventional short-read sequencing cannot determine whether co-occurring variants within a cancer gene reside on the same allele (cis) or on opposing alleles (trans), a distinction with direct biological and therapeutic consequences. Trans configurations confirm biallelic tumor suppressor inactivation and inform therapy selection, while cis configurations generate compound oncogenic alleles with enhanced activity. We analyzed 768 patients with prostate, breast, or ovarian cancers in the PROBLEM cohort, using mutational signatures to nominate cryptic genomic instability cases where the causative biallelic event was not apparent from short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%), leveraging its unique advantages in direct methylation detection, long insertion resolution, and complex structural variant characterization, confirming trans biallelic inactivation in all resolved tumor suppressor cases. Systematic analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, of which 78.7% exceeded the 500 bp short-read phasing limit. Long-read phasing further revealed recurrent compound cis oncogenic alleles in NOTCH1, PIK3CA, PDGFRB, and KIT with functionally synergistic activity. Haplotype phasing resolves a systematically overlooked gap in cancer variant interpretation and warrants broader integration into precision oncology workflows.

Journal Article

Unraveling the complex genetic landscape of OTOF-related hearing loss: a deep dive into cryptic variants and haplotype phasing.

BACKGROUND: Pathogenic variants in OTOF are a major cause of auditory synaptopathy. However, challenges remain in interpreting OTOF variants, including difficulties in confirming haplotype phasing using traditional short-read sequencing (SRS) due to the large gene size, the potential incomplete penetrance of certain variants, and difficulties in assessing variants at non-canonical splice sites. This study aims to revisit the genetic landscape of OTOF variants in a Taiwanese non-syndromic auditory neuropathy spectrum disorder (ANSD) cohort using a combination of sequencing technologies, predictive tools, and experimental validations. METHODS: We performed SRS to analyze OTOF variants in 65 unrelated Taiwanese patients diagnosed with non-syndromic ANSD, complemented by long-read sequencing (LRS) for haplotype phasing. A prediction-to-validation pipeline was implemented to assess the pathogenicity of cryptic variants using SpliceAI software and minigene assays. RESULTS: Biallelic pathogenic OTOF variants were identified in 33 patients (50.8%), while monoallelic variants were found in five patients. Three novel variants, c.3864G > A (p.Ala1288 =), c.4501G > A (p.Ala1501Thr), and c.5813 + 2T > C, were detected. The pathogenicity of two non-canonical mis-splicing variants, c.3894 + 5G > C and c.3864G > A (p.Ala1288 =), was confirmed by minigene assays. LRS-based haplotype phasing revealed that the common missense variant c.5098G > C (p.Glu1700Gln) and the novel variant c.5975A > G (p.Lys1992Arg) are in cis and form a founder pathogenic allele in the Taiwanese population. CONCLUSIONS: Our study highlights the genetic heterogeneity of DFNB9 and emphasizes the importance of population-specific variant interpretation. The integration of advanced sequencing technologies, predictive algorithms, and functional validation assays will improve the accuracy of molecular diagnosis and inform personalized treatment strategies for individuals with DFNB9.

Humans