PubMed HealthSearch

SEARCH · PubMed Health

Results for “Variant classification”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

WilsonGenAI a deep learning approach to classify pathogenic variants in Wilson Disease.

BACKGROUND: Advances in Next Generation Sequencing have made rapid variant discovery and detection widely accessible. To facilitate a better understanding of the nature of these variants, American College of Medical Genetics and Genomics and the Association of Molecular Pathologists (ACMG-AMP) have issued a set of guidelines for variant classification. However, given the vast number of variants associated with any disorder, it is impossible to manually apply these guidelines to all known variants. Machine learning methodologies offer a rapid way to classify large numbers of variants, as well as variants of uncertain significance as either pathogenic or benign. Here we classify ATP7B genetic variants by employing ML and AI algorithms trained on our well-annotated WilsonGen dataset. METHODS: We have trained and validated two algorithms: TabNet and XGBoost on a high-confidence dataset of manually annotated, ACMG & AMP classified variants of the ATP7B gene associated with Wilson's Disease. RESULTS: Using an independent validation dataset of ACMG & AMP classified variants, as well as a patient set of functionally validated variants, we showed how both algorithms perform and can be used to classify large numbers of variants in clinical as well as research settings. CONCLUSION: We have created a ready to deploy tool, that can classify variants linked with Wilson's disease as pathogenic or benign, which can be utilized by both clinicians and researchers to better understand the disease through the nature of genetic variants associated with it.

Hepatolenticular Degeneration

'Truthsets' for clinical validation of large-scale functional assays: Practice recommendations from Cancer Variant Interpretation Group UK (CanVIG-UK).

BACKGROUND: Large-scale functional assays, including multiplex assays of variant effect, have substantial potential to resolve variants of uncertain significance (VUS), particularly for rare missense variants where clinical and population evidence are limited. The ClinGen assay-level clinical validation framework described by Brnich et al provided baseline guidance for the use of functional data for variant classification. However, clear consensus regarding construction of variant 'truthsets' by which to clinically validate functional data remains lacking. METHODS: CanVIG-UK developed consensus recommendations for truthset construction through an iterative national consultation process involving the CanVIG Steering Advisory Group (CStAG), wider CanVIG-UK membership, and engagement with international functional genomics experts. Consultation was based on previous analyses of 2,120 truthset constructions examining the impact of truthset composition on evidence point allocation within the ClinGen assay-level clinical validation framework. RESULTS: Across several consultations, CanVIG-UK established nine guiding principles and seven best-practice recommendations for assay-level clinical validation, using the assumed context of an assay for a cancer susceptibility gene where loss-of-function is the mechanism of pathogenicity. The principal recommendation stipulates, where assays are intended for use in interpretation of largely missense variants, the truthset used to validate should comprise only missense variants. Rather than mixtures of different variant types which may serve to over-estimate assay performance. Additional recommendations support option for relaxation of truthset stringency to improve power, augmentation of benign missense truthsets with systematically derived 'proxy-clinical' benign variants, independent clinical validation separate from assayist-defined validation, and careful evaluation of missense score distributions against that of protein-truncating and synonymous variants. Guidance is also provided for scenarios with limited pathogenic truthset availability and for assays reporting multiple deleterious zones or readouts. CONCLUSIONS: The CanVIG-UK principles and recommendations for truthset construction upon the ClinGen assay-level clinical validation framework, while aiming to form a baseline for future discussion regarding other functional and disease contexts and helping to address the gap between publication of new data and routine clinical implementation.

Journal Article

Distinct rates of VUS reclassification are observed when subclassifying VUS by evidence level.

PURPOSE: Genetic testing commonly yields a plethora of variants of uncertain significance (VUS) that can lead to ongoing uncertainty for patients and their caregivers. Although all VUS hold uncertainty, some VUS have more evidence in support of pathogenicity, whereas others have more evidence of a benign role. Sharing these nuances can help guide the investment in follow-up clinical and research investigations and may, at times, influence medical decision making despite appreciated uncertainty. METHODS: Four clinical laboratories have been subclassifying VUS to help prioritize investigation and guide reporting decisions. Each laboratory developed a distinct approach for how these subclasses are used in their laboratories and, in some cases, displayed on reports. We examined the composition of each laboratory's VUS subclasses and the likelihood variants from each subclass were reclassified toward pathogenic or benign. RESULTS: We found that variants in the lowest subclass of VUS were never reclassified as likely pathogenic or pathogenic, whereas those in the highest subclass were much more likely to be reclassified as pathogenic or likely pathogenic. CONCLUSION: Given that forthcoming professional guidance in variant classification will advise the use of VUS subclasses, the experience of our laboratories in using VUS subclasses can inform future practices.

Humans

Opportunistic screening for broad range of medically relevant secondary findings: Laboratory benefits and burdens.

PURPOSE: Exome and genome sequencing enable opportunistic screening for secondary findings (SFs). We report on exome analysis for a broad range of medically relevant SFs in the setting of the Incidental Genomics randomized clinical trial (NCT03597165). METHODS: Participants had exome sequencing and were randomized to receive only primary cancer findings (control) or cancer findings and a choice of SFs (intervention). RESULTS: Across 279 participants, there were 4441 unique variants in SF genes: 5.0% (221) were reportable pathogenic/likely pathogenic variants, and 81.4% (3615) were nonreportable variants of uncertain significance (VUS). Intervention arm participants had on average 2.6 (SD 1.66, range 0-9) pathogenic/likely pathogenic variants and 29.5 VUS (SD 13.2, range 2-74). SFs for monogenic disease risk were reported in 35.3% (49/139) of participants (American College of Medical Genetics and Genomics non-cancer subset in 1.4%) and carrier status in 89.3% (117/131). In the intervention arm, variant filtration was 7.7 times longer per case (95% CI 5.3 to 11.3, P < .0001), variant classification was 13.3 times longer (95% CI 10.6 to 16.5, P < .0001), and report preparation was 3.3 times longer (95% CI 2.6 to 4.1, P < .0001). CONCLUSION: Although the yield of reportable SFs was high, this was accompanied by many nonreportable VUS and increased efforts for exome analysis.

Humans

Xeroderma pigmentosum variant with multisystem involvement.

BACKGROUND: Xeroderma pigmentosum (XP) is a hereditary disorder characterized by recessive inheritance and elevated rates of skin carcinogenesis. There are seven complementation groups (A through G) for which the genetic defect results in a failure to repair DNA damage from UV light and sunlight; one group, the variant, fails to replicate UV-damaged DNA correctly. Patients in XP groups A, B, D, and G have associated neurologic problems, the most severe being known as the DeSanctis-Cacchione syndrome. OBSERVATIONS: We describe a patient with XP from consanguineous parents who has severe multisystem involvement similar to that of the DeSanctis-Cacchione syndrome. Extensive laboratory investigation showed that cells from this patient exhibit DNA replication after irradiation with UV light that is characteristic of the XP variant. The cells also show normal sensitivity to UV light and normal excision repair, consistent with XP variant classification. The presence of the neurologic symptoms is quite unusual in an XP variant. CONCLUSION: Our patient clearly fits into the XP variant category based on normal survival, caffeine toxic reaction, photoproduct excision and repair, and the deficient replication of UV-damaged DNA. This patient seems to be rare, however, among XP variants in displaying severe neurologic symptoms. Because of the consanguineous parents, the possibility that some of this patient's findings are from non-XP-related abnormalities must also be entertained. However, other consanguineous patients with XP variant, eg, XPIOCA, have been described who do not show neurologic abnormalities. In view of the difficulty of defining an XP group from clinical symptoms alone, we urge the term xeroderma pigmentosum variant be used only in the context of the laboratory studies of patients with XP that contain normal repair but deficient semiconservative replication of UV-damaged DNA.

Cell Survival

Personalizing CA125 Levels Using Tumor Marker Variants: A Case-Control Analysis of Diagnostic Performance for Pancreatic Cancer.

BACKGROUND: Cancer antigen 125 (CA125) is widely recognized as a useful biomarker for the surveillance of patients with ovarian and other cancers. Prior genome-wide association studies have identified variants that influence CA125 levels. We evaluated the utility of stratifying CA125 levels by such variants and evaluated diagnostic performance in control subjects and patients with pancreatic ductal adenocarcinoma (PDAC). METHODS: We measured CA125 levels in 807 control subjects and 450 patients with PDAC and genotyped 10 variants involving four genes (GAL3ST2, MSLN, D2HGDH, and MUC16). We compared CA125 levels in controls by variant and generated variant-defined CA125 cutoffs and then classified cases and controls into functional groups based on their variant profile. We used this variant classification to evaluate the diagnostic performance of CA125 in patients with PDAC. RESULTS: Six variants associated with CA125 levels were used to group controls into one of four groups. Mean CA125 levels in the highest variant group were approximately fourfold higher than in the lowest group. African Americans were more likely to have a variant group associated with low CA125 levels. After setting diagnostic cutoffs by variant group, the diagnostic sensitivity of CA125 for PDAC was 20.2% at 98% specificity (areas under the ROC curve, 0.702), not significantly different from a uniform CA125 diagnostic cutoff (areas under the ROC curve, 0.700). CONCLUSIONS: Gene variants can be used to generate personalized CA125 reference ranges. This approach did not significantly improve CA125's diagnostic performance for pancreatic cancer, but it merits evaluation in other diagnostic settings, such as detecting ovarian cancer. IMPACT: Gene variants can be used to personalize CA125 levels.

Humans

Third Patient With Biallelic Variants in SMAD6 With an Overlapping Phenotype: Developmental Delays, Dysmorphic Features, and Cardiovascular Abnormalities.

SMAD6 encodes an inhibitory SMAD protein that modulates BMP and TGF-&#x3b2; signaling. Heterozygous pathogenic variants in SMAD6 have been primarily associated with aortic valve disease, radioulnar synostosis, and nonsyndromic sagittal and metopic synostosis. However, only two syndromic patients with biallelic variants have been reported in the literature. We report a 4-year-old girl with neurodevelopmental delays, dysmorphic features, complex congenital heart disease, renal asymmetry, and arterial tortuosity. Whole exome sequencing showed two homozygous SMAD6 variants of uncertain significance: c.161G>T (p.Gly54Val) and c.1A>G (p.Met1?). This is the third patient with biallelic SMAD6 variants associated with skeletal changes, more complex cardiovascular phenotype, facial dysmorphism, and novel arterial abnormalities. This suggests biallelic variants may cause a distinct and potentially more severe autosomal recessive syndrome. Functional investigation is needed to determine the molecular consequences of biallelic SMAD6 variants and to inform variant classification and mechanism. This report characterizes a potential unique genetic syndrome associated with biallelic SMAD6 variants, highlighting the importance of additional sequencing, vascular imaging, and multidisciplinary care coordination for these patients.

SMAD6

Deciphering the Role of LNX2 as a Potential Contributor to Neurodevelopmental Disorders.

BACKGROUND/OBJECTIVES: Attention-deficit/hyperactivity disorder (ADHD) is a common neurodevelopmental condition characterized by a complex and multifactorial genetic architecture. In this study, we report a male patient, born to non-consanguineous healthy parents, presenting with ADHD and oppositional defiant disorder (ODD). METHODS: Trio-based whole-exome sequencing (WES) was performed in the proband and both parents. Variant classification was performed according to American College of Medical Genetics and Genomics (ACMG) guidelines, and the potential pathogenicity of the identified variant was further assessed through multiple in silico prediction algorithms and protein structural analyses. RESULTS: WES identified a homozygous variant in the LNX2 gene (NM_153371.4: c.1165G>A, p.Ala389Thr), classified as a variant of uncertain significance (VUS) and supported by multiple in silico predictions. LNX2 is expressed during brain development and encodes an E3 ubiquitin ligase involved in neuronal differentiation and synaptic function. The identified variant is located within the PDZ2 domain, a functionally relevant region involved in protein-protein interactions. Although the variant is reported in population databases (gnomAD ID: rs148429804), it has not been associated with any clinical phenotype, and its presence in the homozygous state has been reported only once, remaining extremely rare and lacking clinical annotation. Structural modelling predicted localized rearrangement of the hydrogen-bonding network within the PDZ2 domain without major conformational changes. Integrative transcriptomic, and single-cell analyses further supported the biological relevance of LNX2 in neurodevelopment, highlighting its preferential association with neuronal projection-cell networks, synaptic vesicle trafficking pathways, and neuron-specific regulatory programs. CONCLUSION: Although the identified LNX2 variant cannot be considered causative for the patient's phenotype and a definitive disease-gene relationship cannot be established based on a single individual, the complementary genetic, structural, and transcriptomic findings support the biological plausibility of LNX2 as a candidate gene for neurodevelopmental disorders. Additional independent patients and functional studies will be required to clarify its contribution to human disease.

Child

Clinical and genetic variant re-analysis among pediatric probands undergoing genetic testing for arrhythmia syndromes.

BACKGROUND: Despite increases in genetic testing, longitudinal data regarding changes in diagnostic yield and variant reclassification for inherited arrhythmia syndromes are limited. OBJECTIVE: Determine longitudinal changes in diagnostic yield and variant classification. METHODS: Single-center retrospective study of probands <18 years undergoing genetic testing for suspected inherited cardiac conditions associated with arrhythmias, 2007 to 2018. Variants were classified as diagnostic (pathogenic/likely pathogenic), non-diagnostic (benign/likely benign [B/LB]), or variants of uncertain significance (VUS). Variant reclassification was performed in October 2023 using VarSome and American College of Medical Genetics criteria. We evaluated results by era (early 2007-2013 vs. later 2014-2018, coinciding with Sanger and next-generation sequencing, respectively) and by likelihood of disease based on clinical evaluation. RESULTS: Of 306 probands, initial testing was 23.2% diagnostic, 55.6% non-diagnostic (33.7% no variant, 21.9% B/LB), and 21.2% VUS. When comparing eras, diagnostic yield decreased (34.1%-15.3%), VUS increased (9.3%-29.9%), and non-diagnostic remained similar (55% to 57%). Variants for 22.7% (46/203) of probands with &#x2265;1 variant changed: 9.9% of diagnostic variants (7/71) downgraded to VUS or non-diagnostic, and 60.0% of VUS changed (23.1% upgraded, 36.9% downgraded). B/LB variants did not change. Probands with higher disease likelihood had 6-times the odds of diagnostic results compared to lower disease likelihood, regardless of era (odds ratio 6.3, 95% confidence interval 3.2-12.4, P < .0001). CONCLUSION: Variant reclassification led to changes in 23% of probands, both downgrading and upgrading status, even among probands initially thought to be pathogenic. When comparing later to earlier eras, VUS variants increased while diagnostic yield decreased. Findings support the need for variant re-interpretation and periodic reclassification over time.

Humans

What Should a Clinical Cardiologist Know About Cardiogenetics?

Inherited cardiovascular diseases are becoming increasingly prominent in clinical practice, significantly impacting diagnosis, risk assessment, and family screening strategies. Progress in genetic testing has broadened access to cardiogenetic evaluations, while also presenting new challenges in interpreting variants and incorporating findings into clinical care. This narrative review explores 20 essential questions that clinical cardiologists may face when dealing with suspected or confirmed inherited cardiac conditions. Organized as a practical, question-driven guide, it outlines when to consider a genetic cause, how to choose and interpret genetic tests, and how to manage patients regardless of their genetic test results. The review emphasizes variant classification based on American College of Medical Genetics and Genomics criteria, the importance of clinical context in interpreting uncertain results, and the principles behind family cascade screening. Particular attention is given to the management of relatives who carry a genetic variant but show no symptoms, and to the current limitations of genetic testing technologies (eg, performance). Ethical considerations, including the appropriate timing of testing in children minors, are also discussed. By connecting genetic insights with clinical cardiology, this review aims to support practical, informed decision making and promote effective collaboration with cardiogenetic specialists.

Humans

The impact of tokenizer selection in genomic language models.

MOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130.

Natural Language Processing

Shared inheritance reveals landscape of somatic and germline cancer risk in TP53.

Pathogenic variants in TP53, the key tumor suppressor gene underlying Li-Fraumeni syndrome (LFS), are among the best-established causes of inherited cancer predisposition. However, large-scale sequencing has revealed that many apparently pathogenic TP53 variants detected in blood are the result of somatic clonal expansions, complicating risk interpretation. Using blood-derived whole-exome data from 469,391 UK Biobank participants, we combined the variant allele fraction (VAF) with haplotype-sharing analysis to distinguish germline and somatic TP53 variants. Germline variants were concentrated at sites linked to partial loss of p53 function and lower disease penetrance, whereas classic LFS alleles appeared to be predominantly somatically acquired. Classic LFS alleles at high VAF conferred markedly increased risk of hematological malignancy but not solid tumors, indicating an important contribution from large TP53-mutant clonal expansions. The prevalence of somatic clonal expansion also correlated with missense variant pathogenicity, suggesting that somatic activity provides an informative in vivo proxy for functional impact. These results provide new insights into TP53-associated cancer risk at the population level, demonstrate that somatic rather than germline risk predominates in middle-aged healthy adults, and provide a scalable framework for variant classification in large-scale population genomics.

Humans

A systematic method for clinical description and classification of personality variants. A proposal.

A systematic method for clinical description and classification of both normal and abnormal personality variants is proposed based on a general biosocial theory of personality. Three dimensions of personality are defined in terms of the basic stimulus-response characteristics of novelty seeking, harm avoidance, and reward dependence. The possible underlying genetic and neuroanatomical bases of observed variation in these dimensions are reviewed and considered in relation to adaptive responses to environmental challenge. The functional interaction of these dimensions leads to integrated patterns of differential response to novelty, punishment, and reward. The possible tridimensional combinations of extreme (high or low) variants on these basic stimulus-response characteristics correspond closely to traditional descriptions of personality disorders. This reconciles dimensional and categorical approaches to personality description. It also implies that the underlying structure of normal adaptive traits is the same as that of maladaptive personality traits, except for schizotypal and paranoid disorders.

Animals

Pan-cancer analysis of biallelic inactivation in tumor suppressor genes identifies KEAP1 zygosity as a predictive biomarker in lung cancer.

The canonical model of tumor suppressor gene (TSG)-mediated oncogenesis posits that loss of both alleles is necessary for inactivation. Here, through allele-specific analysis of sequencing data from 48,179 cancer patients, we define the prevalence, selective pressure for, and functional consequences of biallelic inactivation across TSGs. TSGs largely assort into distinct classes associated with either pan-cancer (Class 1) or lineage-specific (Class 2) patterns of selection for biallelic loss, although some TSGs are predominantly monoallelically inactivated (Class 3/4). We demonstrate that selection for biallelic inactivation can be utilized to identify driver genes in non-canonical contexts, including among variants of unknown significance (VUSs) of several TSGs such as KEAP1. Genomic, functional, and clinical data collectively indicate that KEAP1 VUSs phenocopy established KEAP1 oncogenic alleles and that zygosity, rather than variant classification, is predictive of therapeutic response. TSG zygosity is therefore a fundamental determinant of disease etiology and therapeutic sensitivity.

Kelch-Like ECH-Associated Protein 1

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Increased yield of genetic diagnoses in inherited heart diseases using expanded genome and RNA-splicing analyses.

PURPOSE: The Australian Genomics Cardiovascular Disorders Flagship investigated genome sequencing as a first-line genetic test in 600 individuals with cardiomyopathy, primary arrhythmia syndromes, or congenital heart disease. Analysis of disease-specific virtual gene panels achieved a genetic diagnosis in 38% of participants. We sought to increase genetic diagnosis yields by analyzing lesser-evidenced disease genes, the mitochondrial genome, and by functional analysis of predicted splice-altering variants. METHODS: Genome sequences of 520 participants with cardiomyopathy or primary arrhythmia syndromes were reanalyzed in 572 cardiac genes and the mitochondrial genome. Participants with congenital heart disease were excluded. Variants predicted in silico to disrupt splicing were assessed with blood RNA and minigenes. RESULTS: A new genetic diagnosis was achieved in 4% (19/520) of participants, including deep intronic and mitochondrial genome variants. Ten participants had diagnostic variants in lesser evidenced disease genes; 9 had splicing variant pathogenicity functionally validated. Eleven participants had a newly identified variant of uncertain significance with high suspicion of pathogenicity, warranting clinical review. Our data supported the gene-disease association of 1 new cardiomyopathy gene, TBX20. CONCLUSION: Identifying new gene-disease relationships, maintaining contemporary gene panels, and integrating functional studies to refine splicing variant classifications increase genetic diagnoses for cardiomyopathies and primary arrhythmia syndromes.

Humans

Survey of diagnostic laboratories highlights need for improved standards in somatic genomic testing and reporting.

There is a growing international need to support somatic genomic testing, standardised variant curation and improved patient access to molecular profiling for somatic conditions, including cancer. We conducted a survey of scope, curation, reporting and sharing practices of diagnostic laboratories performing somatic testing in Australia and New Zealand. Laboratories with accreditation (n&#x2009;=&#x2009;41) were invited in 2023 to complete a semi-structured, 25-question interview. Responses were received for 27 laboratories (66% response rate) offering solid tumour, haematological malignancy and non-cancer services. Only 36% of laboratories offered tests capturing the full breadth of variants, from single-nucleotide variants to gene fusions. Knowledge sharing was rare, with only one laboratory submitting variant classifications to a public knowledge base. Most laboratories (96%) conducted somatic testing in oncology. Of cancer laboratories, 35% offered testing considered capable of comprehensive genomic profiling (CGP). Almost half of cancer laboratories had already adopted the 2022 ClinGen/CGC/VICC oncogenicity guidelines, and 84% were using AMP/ASCO/CAP 2017 clinical significance guidelines. Only 47% of mixed discipline&#xa0;cancer laboratories reported biomarkers such as tumour mutational burden, with wide variation in reporting of matched therapy options. Our study has generated a unique overview of somatic laboratory practices in the region, and areas for global standardisation in somatic molecular testing and reporting. We also provide a model for practice and guideline uptake assessment, for application by other country-wide networks. This is particularly relevant in anticipation of CGP mainstreaming, with the increasing complexity of sequencing interpretation for laboratories and clinicians.

Humans

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics