PubMed HealthSearch

SEARCH · PubMed Health

Results for “Structural Variant (SV)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of ∼2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria × ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986 bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria

Longitudinal ctDNA tracking in early and recurrent breast cancer using an ultrasensitive structural variant-based assay: an extended analysis from the TRACER study.

BACKGROUND: Detection of circulating tumor DNA (ctDNA) following curative-intent therapy is prognostic of disease recurrence in early-stage breast cancer (EBC). An ultrasensitive structural variant (SV)-based ctDNA assay was evaluated previously in a 100-patient EBC cohort treated with neoadjuvant therapy, demonstrating high sensitivity, specificity, and a long lead-time to relapse. The stability of primary tumor-specific SVs at and after metastatic recurrence and their utility for longer-term ctDNA monitoring had not been established. PATIENTS AND METHODS: An updated retrospective analysis of ctDNA dynamics was conducted in an expanded cohort of 121 patients with EBC treated with neoadjuvant therapy. Plasma samples were collected at key clinical timepoints and serially in several patients who experienced metastatic recurrence. Clinical variables were abstracted from medical records. Associations between ctDNA detection, dynamics, and clinical outcomes were evaluated in the early-stage and metastatic settings. RESULTS: Thirty of 121 patients experienced clinical recurrence (28 distant, 2 local) over a median follow-up of 4.2 years (range 0.5-8.8; 25 ctDNA evaluable with adjuvant timepoints). All patients with detectable ctDNA in the adjuvant setting developed metastatic recurrence (22/22). Median lead time from ctDNA detection to metastatic recurrence was 346 days (range 0-1937). Among recurrent cases, 79% of primary tumor-specific SVs (n = 17 patients, tumor fraction ≥0.1%) remained detectable in plasma [range 7% (1/14 SV)-100% (15/15); median: 92%]. ctDNA dynamics in the recurrent metastatic setting demonstrated a strong relationship with radiographic outcomes in evaluable patients (n = 9). CONCLUSION: This SV-based digital PCR assay provided ultrasensitive ctDNA detection in an expanded EBC cohort, maintaining 100% positive predictive value for metastatic recurrence. In patients with recurrence, ctDNA dynamics were concordant with radiographic outcomes. Prospective studies evaluating the clinical utility of longitudinal ctDNA monitoring are warranted.

MRD

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

needLR: long-read structural variant annotation with population-scale frequency estimation.

SUMMARY: We present needLR, a structural variant (SV) annotation tool that can be used for filtering and prioritization of candidate pathogenic SVs from long-read sequencing data using population allele frequencies, annotations for genomic context, and gene-phenotype associations. When using population data from 500 presumably healthy individuals to evaluate nine test cases with known pathogenic SVs, needLR assigned allele frequencies to over 97.5% of all detected SVs and reduced the average number of novel genic SVs to 121 per case while retaining all known pathogenic variants. AVAILABILITY AND IMPLEMENTATION: needLR is implemented in bash with dependencies including Truvari v4.2.2, BEDTools v2.31.1, and BCFtools v1.19. Source code, documentation, and pre-computed population allele frequency data are freely available at https://github.com/jgust1/needLR under an MIT license and archived on Zenodo at https://zenodo.org/records/19463479.

Software

Refining the genetic diagnostic puzzle: A case report on a Chinese ARPKD patient with a reciprocal balanced translocation and c.2507 T > C (p.V836A) in PKHD1.

INTRODUCTION: Autosomal recessive polycystic kidney disease (ARPKD) ranks among the most severe chronic kidney diseases (CKD). Its primary cause is variants in the Polycystic Kidney and Hepatic Disease 1 gene (PKHD1). The clinical spectrum of ARPKD varies widely, ranging from mild late-onset symptoms to severe perinatal mortality. However, achieving an early genetic diagnosis in ARPKD patients before clinical symptoms appear proves challenging. CASE PRESENTATION: This case is a 4-year-old boy who experienced a convulsion characterized by a generalized tonic attack lasting approximately 3-5 minutes and later sought treatment to our hospital. However, routine abdominal ultrasound examination accidentally detected that he had diffuse liver lesions, splenomegaly, and bilateral renal enlargement with renal pelvis dilation. Given the uncertainty regarding the underlying cause of the patient's structural abnormalities and convulsions, karyotyping, whole exome sequencing (WES), structural variant analysis (SV analysis) of whole genome sequencing (WGS) were recommended. The result of SV analysis revealed that he has an RBT impacting PKHD1 and the precise location of breakpoints was confirmed through Long-Range Polymerase Chain Reaction (LR-PCR). However, WES did not screen out pathogenic variants initially, the WES data was reviewed subsequently based on SV analysis results. CONCLUSION: We identified an infrequent variant combination, c.2507T>C (p.V836A) in PKHD1 and an RBT with broken PKHD1, which extends the genetic spectrum of ARPKD, and provide a basis for further genetic counselling to the family.

Humans

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

Pan-genome based on chromosome sequences of wild and cultivated Agaricus bisporus.

Agaricus bisporus, one of the most widely cultivated mushrooms around the world, plays an important role in economy and agriculture. In this study, by employing long-reads generated by PacBio and Nanopore sequencing, we assembled six novel high-quality genomes (of which three are telomere-to-telomere assemblies) with sizes 29.6 ~ 30.8 Mb and N50 lengths of 2.5 ~ 2.6 Mb. Combined with public genome data of nine strains, we successfully established a pan-genome of A. bisporus, comprising a total of 14,626 clusters of protein coding genes, of which 50.70%, 7.45%, 24.74% and 17.01% are defined as core, soft core, dispensable, and private clusters, respectively. A total of 5,646 non- redundant structural variants (SVs) were identified among wild and cultivated strains and the genes associated with SV were mapped. This work provides valuable whole-genome sequences and genomic resources across wild and cultivated strains of the most widely cultivated mushroom species for functional analyses of genomes.

Agaricus

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Sawfish: improving long-read structural variant discovery and genotyping with local haplotype modeling.

MOTIVATION: Structural variants (SVs) play an important role in evolutionary and functional genomics but are challenging to characterize. High-accuracy, long-read sequencing can substantially improve SV characterization when coupled with effective calling methods. While state-of-the-art long-read SV callers are highly accurate, further improvements are achievable by systematically modeling local haplotypes during SV discovery and genotyping. RESULTS: We describe sawfish, an SV caller for mapped high-quality long reads incorporating systematic SV haplotype modeling to improve accuracy and resolution. Assessment against the draft Genome in a Bottle (GIAB) SV benchmark from the T2T-HG002-Q100 diploid assembly shows that sawfish has the highest accuracy among state-of-the-art long-read SV callers across every tested SV size group. Additionally, sawfish maintains the highest accuracy at every tested depth level from 10- to 32-fold coverage, such that other callers required at least 30-fold coverage to match sawfish accuracy at 15-fold coverage. Sawfish also shows the highest accuracy in the GIAB challenging medically relevant genes benchmark, demonstrating improvements in both comprehensive and medically relevant contexts.When joint-genotyping seven samples from CEPH-1463, sawfish has over 9000 more pedigree-concordant calls than other state-of-the-art SV callers, with the highest proportion of concordant SVs (81%). Sawfish's quality model enables selection for an even higher proportion of concordant SVs (88%), while still calling nearly 5000 more pedigree-concordant SVs than other callers. These results demonstrate that sawfish improves on the state-of-the-art for long-read SV calling accuracy across both individual and joint-sample analyses. AVAILABILITY AND IMPLEMENTATION: Sawfish source code, pre-compiled Linux binaries, and documentation are released on GitHub: https://github.com/PacificBiosciences/sawfish.

Haplotypes

The human IG heavy chain constant gene locus is enriched for large structural variants and coding polymorphisms that vary among human populations.

The human immunoglobulin heavy chain constant (IGHC) domain of antibodies (Ab) is responsible for effector functions critical to immunity. This domain is encoded by genes in the IGHC locus, where descriptions of genomic diversity remain incomplete. We utilized long-read sequencing to build an IGHC haplotype/variant catalog from 105 individuals of diverse ancestry. We discovered uncharacterized single nucleotide variants (SNV) and large structural variants (SVs, n=7), representing new genes and alleles enriched for non-synonymous substitutions, highlighting potential functional effects. Of the 221 identified IGHC alleles, 192 were novel. SNV, SV, and gene allele/genotype frequencies revealed population differentiation, including (i) hundreds of SNVs in African and East Asian populations exceeding a fixation index (FST) of 0.3, and (ii) an IGHG4 haplotype carrying coding variants uniquely enriched in Asian populations. Our results illuminate missing signatures of IGHC diversity and establish a new foundation for investigating IGHC germline variation in Ab function and disease.

Journal Article