PubMed HealthSearch

SEARCH · PubMed Health

Results for “Structural variants (SVs)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

The value of structural variants to conservation genomics in the pangenome era.

Structural variants (SVs) comprise an axis of genetic diversity with strong consequences for phenotype and fitness, making them a potentially important target for conservation genomics. Here, we review how and why SVs can play a role in conservation genomics; the different types of SVs and how they can affect phenotype; and how pangenomes and long-read sequencing are illuminating their evolution in populations, including small populations and those of conservation concern. SVs comprise multinucleotide mutations including insertions, deletions, transpositions, inversions, and other multinucleotide mutations, often overlapping genes and other functional genome regions. As a result, SVs often play important roles in phenotypic evolution and local adaptation and can contribute substantially to genetic load in inbred populations. However, our understanding of the factors influencing SV diversity in populations is still in its infancy and is complicated by the vast range of sizes, effects, and mechanisms of formation of these mutations. We argue that SVs are an important axis of genetic diversity which should be characterized alongside more traditional metrics of genetic diversity in conservation contexts. There are a number of analytical challenges to detecting and studying SVs, but analyses aimed at understanding the role of SVs in inbreeding load and population health are rapidly becoming realizable goals, accelerated by new technologies and analytical approaches. New tools, including population-scale long-read sequencing and pangenome approaches, are beginning to make SVs accessible in ways which can be readily applied in conservation settings.

Genomic Structural Variation

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals

Sawfish: improving long-read structural variant discovery and genotyping with local haplotype modeling.

MOTIVATION: Structural variants (SVs) play an important role in evolutionary and functional genomics but are challenging to characterize. High-accuracy, long-read sequencing can substantially improve SV characterization when coupled with effective calling methods. While state-of-the-art long-read SV callers are highly accurate, further improvements are achievable by systematically modeling local haplotypes during SV discovery and genotyping. RESULTS: We describe sawfish, an SV caller for mapped high-quality long reads incorporating systematic SV haplotype modeling to improve accuracy and resolution. Assessment against the draft Genome in a Bottle (GIAB) SV benchmark from the T2T-HG002-Q100 diploid assembly shows that sawfish has the highest accuracy among state-of-the-art long-read SV callers across every tested SV size group. Additionally, sawfish maintains the highest accuracy at every tested depth level from 10- to 32-fold coverage, such that other callers required at least 30-fold coverage to match sawfish accuracy at 15-fold coverage. Sawfish also shows the highest accuracy in the GIAB challenging medically relevant genes benchmark, demonstrating improvements in both comprehensive and medically relevant contexts.When joint-genotyping seven samples from CEPH-1463, sawfish has over 9000 more pedigree-concordant calls than other state-of-the-art SV callers, with the highest proportion of concordant SVs (81%). Sawfish's quality model enables selection for an even higher proportion of concordant SVs (88%), while still calling nearly 5000 more pedigree-concordant SVs than other callers. These results demonstrate that sawfish improves on the state-of-the-art for long-read SV calling accuracy across both individual and joint-sample analyses. AVAILABILITY AND IMPLEMENTATION: Sawfish source code, pre-compiled Linux binaries, and documentation are released on GitHub: https://github.com/PacificBiosciences/sawfish.

Haplotypes

Optical genome mapping enhanced by refined variant interpretation in pediatric acute lymphoblastic leukemia.

Reliable detection of structural variants (SVs) and copy number variations (CNVs) is crucial in the contemporary diagnostics of pediatric B-cell acute lymphoblastic leukemia (B-ALL). However, limitations of commonly used conventional and molecular cytogenetic methods may hinder the accurate genetic characterization of patients. Optical genome mapping (OGM) offers a reliable alternative by enabling high-resolution, genome-wide detection of CNVs and SVs. Chromosomal aberrations were screened using OGM in 51 children with B-ALL. The results were compared with those of karyotyping, fluorescence in situ hybridization (FISH), digital multiplex ligation-dependent probe amplification (digitalMLPA), and targeted RNA sequencing (RNA-seq). OGM data showed high congruency with karyotyping and FISH findings, detecting clinically relevant variants beyond G-banding results and unraveling a complex KMT2A fusion undetected by FISH. Gene fusions involved in complex ETV6::RUNX1 translocations, but not detected by RNA-seq, were confirmed using FISH. Normalization of OGM copy number values with DNA-index-improved concordance with FISH-derived copy numbers in near-tri/tetraploid cases. In the peripheral regions of OGM variants (fringe-zones), a novel evaluation strategy called 'FriZone' was applied, which significantly improved the concordance between OGM and digitalMLPA. In addition, a co-segregation analysis revealed strong associations between ETV6::RUNX1 fusion and deletions of ETV6, RAG2, and NR3C2. OGM uncovered complex rearrangements undetected by widely used methods in 15% of cases, improving genetic classification and risk stratification in 10% of the patients. The FriZone analysis and normalization by DNA-index provide a refined, more accurate approach to OGM variant interpretation, facilitating the efficient application of OGM in clinical diagnostics. © 2026 The Author(s). The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

Humans

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Analysis of 14q12 microdeletions reveals novel regulatory loci for the neurodevelopmental disorder-related gene FOXG1.

Up to 17% of neurodevelopmental disorders (NDDs) can be explained by pathogenic structural variants (SVs) that disrupt coding regions and elicit gene dosage defects. However, noncoding SVs which can perturb cis-regulatory elements (CREs) and downstream gene expression are understudied. In this study, we describe multiple 14q12 deletions downstream of NDD-related gene FOXG1 in individuals with overlapping phenotypes of FOXG1 haploinsufficiency. We show that deletion of a minimum region of overlap (MRO) reduced FOXG1 expression, disrupted CREs and altered FOXG1's native genomic interactions. Deleting the MRO did not fully eliminate FOXG1 expression, indicating that multiple CREs likely cooperate to regulate FOXG1 and would need to be deleted to completely prevent expression. The transcriptomic profiles of MRO loss overlap in part with FOXG1 loss, including direct FOXG1 targets, indicating converging molecular pathways. These findings expand the scope of FOXG1's complex regulatory region, and more broadly, of regulatory SVs in NDD susceptibility.

Forkhead Transcription Factors

Analysis of deep-resequencing data of 984 soybean accessions reveals structural variations underlying agronomic traits.

Genomic structural variants (SVs) are major sources of genetic variation and have profound impacts on phenotypic traits. However, their functional effects remain largely unexplored in soybean. Here, we resequence 940 soybean accessions. Together with 44 publicly available datasets, we identify 602,281 SVs. Using a graph-based genome, we detect an additional 58,760 presence/absence variations (PAVs) that broadly affect gene expression. Population genomic analyses reveal that SVs serve as a core driving force for soybean domestication and improvement. Integrating SVs with QTLs for oil and protein content, and performing GWAS on 27 traits, we identify key functional SVs. These include transposable element insertions altering seed coat color, multiple insertions within a cytochrome P450 gene modifying flower and hypocotyl color, and a GmMATE1 deletion enhancing seed size. Together, our study establishes a comprehensive SV map of soybean, offering a valuable resource for dissecting the genetic basis of complex traits to accelerate molecular breeding.

Glycine max

Genomic and genetic dissection underlying seedling drought resilience in oats.

Drought threatens global crop yields, and common oat, a vital nutritional source for food and feed, is particularly constrained in the semi‑arid regions where it is widely cultivated. Here, we report two high-quality genome assemblies for drought-resilient (Borris37) and drought-sensitive (XymC06) oat accessions with distinct seedling survival rates and genome sizes of 10.92 Gb and 10.96 Gb, and construct comprehensive landscapes of insertion‑deletions (InDels) and structural variants (SVs). Integrating population-level genomic, transcriptomic and phenotypic (seedling survival rate), we demonstrate that InDels and SVs underpin divergent drought resilience and identify 52 candidate genes associated with drought resistance whose expression is significantly modulated by these variants. Borris37 accumulates 36 favorable alleles of these genes. An InDel in the AsNF-YB3 promoter enhances binding to AsARF1, upregulating AsNF‑YB3 under drought, and overexpression of AsNF‑YB3 reduces ROS accumulation. Our findings provide resources and targets for drought‑resistance breeding in oat, thereby supporting global food security.

Drought Resistance

Pan-genome based on chromosome sequences of wild and cultivated Agaricus bisporus.

Agaricus bisporus, one of the most widely cultivated mushrooms around the world, plays an important role in economy and agriculture. In this study, by employing long-reads generated by PacBio and Nanopore sequencing, we assembled six novel high-quality genomes (of which three are telomere-to-telomere assemblies) with sizes 29.6 ~ 30.8 Mb and N50 lengths of 2.5 ~ 2.6 Mb. Combined with public genome data of nine strains, we successfully established a pan-genome of A. bisporus, comprising a total of 14,626 clusters of protein coding genes, of which 50.70%, 7.45%, 24.74% and 17.01% are defined as core, soft core, dispensable, and private clusters, respectively. A total of 5,646 non- redundant structural variants (SVs) were identified among wild and cultivated strains and the genes associated with SV were mapped. This work provides valuable whole-genome sequences and genomic resources across wild and cultivated strains of the most widely cultivated mushroom species for functional analyses of genomes.

Agaricus

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

The human IG heavy chain constant gene locus is enriched for large structural variants and coding polymorphisms that vary among human populations.

The human immunoglobulin heavy chain constant (IGHC) domain of antibodies (Ab) is responsible for effector functions critical to immunity. This domain is encoded by genes in the IGHC locus, where descriptions of genomic diversity remain incomplete. We utilized long-read sequencing to build an IGHC haplotype/variant catalog from 105 individuals of diverse ancestry. We discovered uncharacterized single nucleotide variants (SNV) and large structural variants (SVs, n=7), representing new genes and alleles enriched for non-synonymous substitutions, highlighting potential functional effects. Of the 221 identified IGHC alleles, 192 were novel. SNV, SV, and gene allele/genotype frequencies revealed population differentiation, including (i) hundreds of SNVs in African and East Asian populations exceeding a fixation index (FST) of 0.3, and (ii) an IGHG4 haplotype carrying coding variants uniquely enriched in Asian populations. Our results illuminate missing signatures of IGHC diversity and establish a new foundation for investigating IGHC germline variation in Ab function and disease.

Journal Article

The landscape of structural variation in pediatric cancer.

Structural variants (SVs) account for over 60% of the driver variants in pediatric cancer, and in many cases act as the cancer initiating event. To study SVs from a pan-cancer perspective, we analyzed 1,616 pediatric cancer genomes in 16 major cancer types of hematological malignancies (n = 908), brain tumors (n = 183), and solid tumors (n = 525) and compared their profiles to those of 2,203 adult cancers. The SV burden varied ~100-fold across pediatric cancer types and demonstrated an 8- to 16-fold reduction compared to adult brain and solid tumors but was comparable in pediatric versus adult hematological malignancies. Recurrent SV hotspots occurred uniquely in pediatric acute lymphoblastic leukemias (ALLs) in proximity to RAG-mediated recombination signal sequences (RSS) and disrupted multiple immune-related loci as well as 69 genes, which often involved cryptic RSS sites. By contrast, such hotspots affected only immune-related loci but not driver genes in adult lymphoid cancers. Eight SV signatures extracted from the cohort had varying distributions across cancer types, with clustered translocations reflecting templated insertions in osteosarcoma, and medium-sized deletions (10 kb to 1 Mb) enriched in cancers with RAG-mediated deletions. Intra-patient evolutionary analysis in 13 patients with multiple spatiotemporally distinct samples revealed that RAG-mediated recombination in leukemia and complex rearrangements in solid tumors occurred both early in disease initiation and continuously during later diversification, contributing to clonal heterogeneity. Finally, we found that both driver genes and fragile sites were the two genomic regions most frequently disrupted by SVs. The unique and diverse SV landscapes that emerged from this comprehensive analysis expand the scope of RSS-mediated mutagenesis in pediatric ALL and will be a valuable resource for guiding future functional studies and the design of clinical genomic testing in pediatric cancer.

Journal Article

Long-read sequencing of single cell-derived melanoma subclones reveals divergent and parallel genomic and epigenomic evolutionary trajectories.

Tumor evolution is driven by various mutational processes, ranging from single-nucleotide variants (SNVs) to large structural variants (SVs) to dynamic shifts in DNA methylation. Current short-read sequencing methods struggle to accurately capture the full spectrum of these genomic and epigenomic alterations due to inherent technical limitations. To overcome that, here we introduce an approach for long-read sequencing of single-cell derived subclones, and use it to profile 23 subclones of a mouse melanoma cell line, characterized with distinct growth phenotypes and treatment responses. We develop a computational framework for harmonization and joint analysis of different variant types in the evolutionary context. Uniquely, our framework enables detection of recurrent amplifications of putative driver genes, generated by independent SVs across different lineages, suggesting parallel evolution. In addition, our approach revealed gradual and lineage-specific methylation changes associated with aggressive clonal phenotypes. We also show our set of phylogeny-constrained variant calls along with openly released sequencing data can be a valuable resource for the development of new computational methods.

Journal Article

Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions.

Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions ≥50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs ≥4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..

Journal Article

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article