PubMed HealthSearch

SEARCH · PubMed Health

Results for “Structural variants (SVs)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A pangenome framework uncovers the role of deletions in repeated evolution of cave-derived traits.

Structural variants (SVs) are increasingly recognized as key contributors to adaptive evolution, yet they remain underexplored compared with single-nucleotide variation. To understand how large-scale genomic changes shape repeated evolution, we leveraged multiple levels of sequence data across the powerful evolutionary model system of the Mexican tetra fish (Astyanax mexicanus). We constructed one of the first pangenome graphs from a naturally evolving vertebrate, enabling comprehensive discovery of SVs among 120 fish from 11 populations. We discover substantial amounts of structural variation and explore the roles of genomic biases and selection in shaping the distribution of these variants. More than 2400 high-confidence cave-specific deletions are enriched in biological pathways involved in vision, metabolism, and behavior and cluster nonrandomly in quantitative trait loci linked to cavefish traits. Additionally, 67 genes harbor unique deletions between independent cavefish lineages. These reused genes show evidence of population-specific selection (99% contain selective sweeps compared with 8%-15% in genes lacking SVs), indicating that deletions likely rose in frequency through repeated positive selection rather than drift. Together, these results reveal that recurrent deletion events have repeatedly contributed to the evolution of cave-adapted phenotypes and highlight deletions as underexplored contributors of adaptive evolution in extreme environments.

Animals

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We develop a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implement it in the SuPreMo-Akita software. With this tool, we rank hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This reveals that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritize candidate variants that disrupt ASD CREints and validate our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study suggests that disrupted genome folding is a potential genetic mechanism in a subset of ASD cases and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Humans

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We developed a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implemented it in the SuPreMo-Akita software (Gjoni and Pollard 2024). With this tool, we ranked hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This revealed that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritized candidate variants that disrupt ASD CREints and validated our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study establishes disrupted genome folding as a potential genetic mechanism in ASD and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Journal Article

Lineage-associated small inversions disrupt dosT, dnaE2, and a promoter-adjacent region in some Mycobacterium tuberculosis isolates.

UNLABELLED: Large molecular inversions in the genome of Mycobacterium tuberculosis (Mtb) due to factors like the presence of insertion sequences and transposases are widely known. However, smaller inversions within coding sequences and non-coding control elements are rarely reported. The present study aims to identify inversions and their potential impact on Mtb biology in a lineage-specific manner. Structural variants (SVs) could only be detected by long reads. For this, we simulated long reads by de novo assembling the short-read sequencing data sets and subsequently aligned representative strains from each lineage using the Progressive Mauve algorithm. Independently, long-read sequencing from the Pacific Biosciences platform was acquired and analyzed using the structural variant identification method. Variants were merged, and Fisher's exact test was carried out to identify the inversion association with lineages. To visualize deoxyribonucleic acid (DNA) features, the DNA-features-viewer tool was used. Simulated reads from short-read sequencing gave indications of lineage (L)-specific inversions. The long-read sequencing approach led to the identification of seven unique inversions: two positively associated with L1, one positively associated with L3, two negatively associated with L4, and two positively associated with L3 but negatively associated with L4 (P < 0.05). The inversions encompassed primarily non-essential genes like sdaA, dosT, Rv2026c, dnaE2, Rv1341, Rv1342, and lprD. An interesting inversion was observed in the upstream control element of purB and Rv0776c. The study sheds light on small inversions that may be causing alterations in expression, formation of fusion genes, and nonsense mutations that may have a role in lineage-specific phenotypic changes. IMPORTANCE: The role of mutations like SNPs and INDELs and their association with drug resistance is well known in Mycobacterium tuberculosis (Mtb). However, structural variations, especially inversions, are largely overlooked and unreported. In this paper, publicly available whole-genome sequencing datasets from Illumina and Pacific Biosciences-Oxford Nanopore Technologies platform have been used to detect inversions and report seven unreported Mtb lineage-specific small inversions.

Mycobacterium tuberculosis

Pan-Genome Analysis Reveals Local Adaptation to Climate Driven by Introgression in Oak Species.

The genetic base of local adaptation has been extensively studied in natural populations. However, a comprehensive genome-wide perspective on the contribution of structural variants (SVs) and adaptive introgression to local adaptation remains limited. In this study, we performed de novo assembly and annotation of 22 representative accessions of Quercus variabilis, identifying a total of 543,372 SVs. These SVs play crucial roles in shaping genomic structure and influencing gene expression. By analyzing range-wide genomic data, we identified both SNPs and SVs associated with local adaptation in Q. variabilis and Quercus acutissima. Notably, SV-outliers exhibit selection signals that did not overlap with SNP-outliers, indicating that SNP-based analyses may not detect the same candidate genes associated with SV-outliers. Remarkably, 29%-37% of candidate SNPs were located in a 250&#x2005;kb region on chromosome 9, referred to as Chr9-ERF. This region contains 8 duplicated ethylene-responsive factor (ERF) genes, which may have contributed to local adaptation of Q. variabilis and Q. acutissima. We also found that a considerable number of candidate SNPs were shared between Q. variabilis and Q. acutissima in the Chr9-ERF region, suggesting a pattern of repeated selection. We further demonstrated that advantageous variants in this region were introgressed from western populations of Q. acutissima into Q. variabilis, providing compelling evidence that introgression facilitates local adaptation. This study offers a valuable genomic resource for future studies on oak species and highlights the importance of pan-genome analysis in understating mechanism driving adaptation and evolution.

Quercus

Signals of Natural Selection Across Regions of Low Recombination in Wild Populations of the Purple Sea Urchin, Strongylocentrotus purpuratus.

Structural variants (SVs) are increasingly recognized as important components of genetic architecture. Yet our understanding of the evolutionary forces maintaining SVs in natural populations is limited. Chromosomal inversions in particular can facilitate local adaptation in populations with high gene flow, including many marine species. The purple sea urchin (Strongylocentrotus purpuratus) is a powerful system to study these dynamics due to its high gene flow, lack of population structure, and broad latitudinal range. We analyzed whole genome sequence data from 137 individuals sampled across seven populations to identify regions of low recombination using scans for elevated linkage disequilibrium and genetic differentiation. Such regions may arise from structural variants, including chromosomal inversions. We identified nine regions showing signatures of reduced recombination, including three way genotype clustering, long range linkage, and hanging bridge patterns frequently associated with inversion polymorphisms. The regions were polymorphic within locations and along the species range with three loci showing concordant signatures of balancing and spatially heterogeneous selection based on enrichment of outliers and distinct patterns of allelic age. Additionally, these loci showed enrichment for genes associated with biomineralization and development. Our results provide the first evidence for regions of low recombination in the purple sea urchin genome, several of which display genomic signatures consistent with structural variants such as chromosomal inversions. These findings add to growing evidence that regions of reduced recombination constitute an important component of standing genetic variation in natural populations and may play a key role in adaptation to heterogeneous environments.

Strongylocentrotus purpuratus

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article

Third-generation whole-genome sequencing reveals the role of CNTNAP2 as a tumor suppressor gene in high-risk neuroblastomas.

BACKGROUND: Neuroblastoma is a common and aggressive pediatric sympathetic nervous system tumor. Genomic structural variants (SVs) contribute substantially to neuroblastoma, yet remain under-characterized in high-risk neuroblastomas. We aimed to elucidate neuroblastoma pathogenesis using third-generation whole-genome sequence high-risk cases to identify driver aberrations and explore potential therapeutic strategies. METHODS: We analyzed third-generation whole-genome sequencing data of 20 high-risk neuroblastoma samples and combined the findings with those obtained from the analysis of clinical samples, in vitro models, and public datasets. RESULTS: The contactin-associated protein-like 2 (CNTNAP2) gene was observed to be frequently aberrated because of structural variants in high-risk neuroblastoma samples. CNTNAP2 expression was significantly correlated with favorable histology and could be used to predict prognosis using clinical samples and neuroblastoma datasets. Overexpression and knockdown experiments and transcriptomic analysis revealed that CNTNAP2 was primarily involved in neuronal differentiation and axon guidance pathways; moreover, CNTNAP2 was required for neuroblastoma differentiation and affected cancer stemness. Immunoprecipitation and mass spectrometry revealed that CNTNAP2 interacted with cytoskeletal proteins like drebrin 1 (DBN1) and myosin-heavy chain 9 (MYH9). CNTNAP2 dynamically reorganises actin and microtubules for DBN1-mediated neuronal differentiation. CNTNAP2 also reduces CTNNB1 transcription and &#x3b2;-catenin pathway activation by inhibiting MYH9 nuclear translocation. CNTNAP2 overexpression in neuroblastoma cell lines resulted in cell cycle arrest, decreased cell proliferation and metastasis. CONCLUSIONS: The recurrent loss of CNTNAP2 in neuroblastoma contributes to an aggressive phenotype by impairing neuronal differentiation and increasing cancer stemness. These findings may serve as a foundation for developing therapeutic strategies to overcome barriers to differentiation.

Humans

Comparisons Between Large-Scale Genomic Variants and SNPs in Driving Population Divergence and Local Adaptation.

Genomic variations, such as indels (2-49 bp) and structural variants (SVs, &#x2265;50 bp), are larger-scale mutations than single nucleotide polymorphisms (SNPs) and can substantially impact evolutionary processes, including speciation, adaptation, and phenotypes. Despite their functional importance, integrative population genetic analyses that jointly consider genome-wide SNPs, indels, and SVs remain under-explored. The ground tit (Pseudopodoces humilis), an endemic species to the Qinghai-Tibet Plateau (QTP), exhibits divergence across distinct glacial refugia, accompanied by habitat and morphological divergence, making it an excellent example for investigating how different types of genomic variants contribute to population divergence and local adaptation. Here, by retrieving 81 whole-genome sequence data, over 13 million SNPs, 2 million indels, and 22,101 SVs were identified. Variants were unevenly distributed across the genome, characterized by distinct hotspot regions. Indels and SVs revealed four genetic clusters consistent with previous SNP-based results, thereby validating the reliability of our variant datasets. FST and genotype-environment association (GEA) analyses independently revealed numerous candidate indels and SVs; each showed minimal overlap with previously identified SNPs, and were enriched in similar functional pathways such as signal transduction, skeletal muscle development, water transport, DNA repair, reproduction, nervous system development, and immunity. Collectively, our results demonstrated that indels and SVs could capture additional signatures besides SNPs. Furthermore, similar but distinct gene functions among different types of genomic variants collectively and complementarily drive genomic divergence across environmental gradients in such a high-elevation endemic species, underscoring its evolutionary relevance in local adaptation.

indels

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28&#xd7;, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS&#xa0;&#x2248;&#xa0;1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (H&#x2092;&#xa0;&#x2248;&#xa0;0.71) and nucleotide diversity (&#x3c0;&#xa0;&#x2248;&#xa0;7&#xa0;&#xd7;&#xa0;10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation

Chromosome-scale genome remodeling in tumor evolution: Copy number alterations and structural variants as two sides of the same coin.

Chromosome-scale genomic rearrangements are a dominant force in tumor evolution. Copy-number alterations (CNAs) and structural variants (SVs) constitute two complementary axes of this process. Although detection technologies now deliver near-comprehensive catalogs, technical resolution has outpaced conceptual integration. In this review, we frame CNAs and SVs as inextricable facets of chromosomal aberrations. They reshape cancer genomes through altered gene dosage and three-dimensional regulatory rewiring. CNAs quantify the gene-dosage imbalance, yet arise through mechanistically distinct routes. Segmental CNAs typically require chromosomal breakage, and therefore often coincide with SV junctions. By contrast, whole-chromosome aneuploidy and whole-genome doubling (WGD) primarily reflect mitotic or cytokinetic failure and can occur without local breakpoints, while nevertheless reshaping the karyotypic landscape and seeding subsequent structural complexity. SVs, in turn, range from unbalanced events that alter copy number to ostensibly balanced exchanges that predominantly rewire regulatory architecture. Despite their diverse and sometimes catastrophic architectures, SVs are ultimately rooted in double-strand break formation and error-prone resolution. By integrating CNAs and SVs within a unified mechanistic and functional framework, we aim to convert catalogs into concepts and distill the organizing principles that govern tumor genome evolution.

Humans

Structural variants linked to Alzheimer's disease and other common age-related clinical and neuropathologic traits.

BACKGROUND: Alzheimer's disease (AD) is a complex neurodegenerative disorder with substantial genetic influence. While genome-wide association studies (GWAS) have identified numerous risk loci for late-onset AD (LOAD), the functional mechanisms underlying most of these associations remain unresolved. Large genomic rearrangements, known as structural variants (SVs), represent a promising avenue for elucidating such mechanisms within some of these loci. METHODS: By leveraging data from two ongoing cohort studies of aging and dementia, the Religious Orders Study and Rush Memory and Aging Project (ROS/MAP), we performed genome-wide association analysis testing 20,205 common SVs from 1088 participants with whole genome sequencing (WGS) data. A range of Alzheimer's disease and other common age-related clinical and neuropathologic traits were examined. RESULTS: First, we mapped SVs across 81 AD risk loci and discovered 22 SVs in linkage disequilibrium (LD) with GWAS lead variants and directly associated with the phenotypes tested. The strongest association was a deletion of an Alu element in the 3'UTR of the TMEM106B gene, in high LD with the respective AD GWAS locus and associated with multiple AD and AD-related disorders (ADRD) phenotypes, including tangles density, TDP-43, and cognitive resilience. The deletion of this element was also linked to lower TMEM106B protein abundance. We also found a 22-kb deletion associated with depression in ROS/MAP and bearing similar association patterns as GWAS SNPs at the IQCK locus. In addition, we leveraged our catalog of SV-GWAS to replicate and characterize independent findings in SV-based GWAS for AD and five other neurodegenerative diseases. Among these findings, we highlight the replication of genome-wide significant SVs for progressive supranuclear palsy (PSP), including markers for the 17q21.31 MAPT locus inversion and a 1483-bp deletion at the CYP2A13 locus, along with other suggestive associations, such as a 994-bp duplication in the LMNTD1 locus, suggestively linked to AD and a 3958-bp deletion at the DOCK5 locus linked to Lewy body disease (LBD) (P&#x2009;=&#x2009;3.36&#x2009;&#xd7;&#x2009;10-4). CONCLUSIONS: While still limited in sample size, this study highlights the utility of including analysis of SVs for elucidating mechanisms underlying GWAS loci and provides a valuable resource for the characterization of the effects of SVs in neurodegenerative disease pathogenesis.

Humans

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

European ash pangenome reveals widespread structural variation and genetic basis of low ash dieback susceptibility.

European Ash (Fraxinus excelsior) is a keystone tree species, whose populations are being decimated by ash dieback disease (ADB) - better characterisation of genetic variants associated with low susceptibility to the disease is needed. Here, we develop a F. excelsior pangenome to more fully capture sequence variability within this species compared with a linear reference genome, using a geographically diverse set of fifty F. excelsior samples. We identify 362,965 structural variants (SVs), including 174&#x2009;Mb of sequence absent from the linear reference genome (22% of the linear reference size), and identify 3,412 high-confidence dispensable genes (those present only in some individuals). We use the pangenome to analyse existing genomic data from over 1,200 individuals, revealing 220 single nucleotide polymorphisms (SNPs) showing consistent allele frequency shifts between healthy individuals and those highly damaged by ADB, across UK seed sources, explicitly demonstrating the existence of a shared genetic component to low ADB susceptibility.

Polymorphism, Single Nucleotide

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans