PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

A haploid wild yeast resource for exploring the natural ecology of Saccharomyces cerevisiae.

Saccharomyces cerevisiae occurs predominantly in the diploid state in nature, limiting genetic analyses of wild populations. Here, we establish a haploid collection from 32 Taiwanese S. cerevisiae isolates through targeted HO disruption. This resource spans predomesticated Asian wild lineages and enables the investigation of reproductive isolation and ecological trait variation. Although all pairwise hybridizations formed zygotes, many yielded reduced spore viability, revealing strong postzygotic barriers. Genome analyses associated reduced hybrid fertility with lineage-specific structural variation, including elevated levels of intra-chromosomal inversions in H413-8/TW1 and inter-chromosomal rearrangements in PD35A/CHN-V, rather than sequence divergence alone. Phenotyping revealed ecological differentiation, with TW1 favoring cooler growth and a natural hybrid exhibiting heterosis with expanded thermotolerance. Most wild strains grew poorly on maltose, whereas anthropogenic strains displayed enhanced utilization linked to MAL + regulatory alleles and maltose-specific transporters. Together, this haploid collection links structural variation and metabolic divergence to ecological and reproductive differentiation in wild S. cerevisiae.

Saccharomyces cerevisiae

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

STK11 Mutations and Deletions Define an Aggressive Molecular Subgroup of Cervical Adenocarcinoma.

Cervical adenocarcinoma accounts for 15%-20% of cervical cancers and is associated with poorer survival and reduced response to screening and immunotherapy compared with squamous cell carcinoma (SCC). The genomic drivers underlying this molecular subgroup remain incompletely characterized. Whole-exome sequencing was performed on 302 invasive cervical cancers from Guatemala and Venezuela. Structural variation analysis was conducted using SNP-array and whole-genome sequencing data. Findings were replicated in more than 4600 additional cervical cancer samples from TCGA, AACR Project GENIE, MSKCC, and Caris datasets. TP53 mutations were more frequent in adenocarcinoma than SCC, particularly in HPV-negative tumors. STK11 alterations, including mutations and focal deletions, were significantly enriched in HPV-positive adenocarcinomas compared with SCC and affected 23% of adenocarcinomas overall. Whole-genome analyses identified recurrent focal deletions, inversions, chromosomal rearrangements, and breakage-fusion-bridge events involving chromosome 19p and STK11 that were not detected by exome sequencing alone. STK11 alterations were associated with younger age at diagnosis, poorer overall survival, and inferior outcomes following immune checkpoint inhibitor (ICI) therapy. STK11 alterations significantly co-occurred with YAP1 amplification but were largely mutually exclusive with PIK3CA mutation. Cervical adenocarcinomas also demonstrated significantly lower CD274 (PD-L1) expression than SCC. STK11 alterations define a distinct molecular subgroup of cervical adenocarcinoma characterized by structural disruption of chromosome 19p, younger age at onset, and poorer clinical outcomes. These findings have implications for molecular classification and future targeted therapeutic approaches in cervical cancer.

Humans

Panmixia in a Widespread Butterfly: High Dispersal and Ecological Generalism Buffer Against Landscape Fragmentation.

Habitat fragmentation is widely expected to reduce population connectivity and increase genetic differentiation, although the strength of these effects depends on species-specific traits such as dispersal ability. Here, we investigated the population genetic structure of the cosmopolitan butterfly, Pieris rapae L. (Lepidoptera: Pieridae), across western Germany using genome-wide single-nucleotide polymorphism (SNP) data. To analyze the effects of landscape structure on genetic connectivity, we applied a paired study design comprising four landscape pairs, each consisting of a highly intensified, modern agricultural landscape and a more heterogeneous, traditional landscape. Our results revealed no evidence of genetic differentiation. Pairwise FST values were close to zero; we detected no isolation by distance, and clustering analyses supported a single genetic population. No meaningful associations between genetic variation and environmental variables were detected, with landscape effects explaining less than 0.4% of genomic variation. Consequently, we found no evidence for stronger genetic structuring in modern compared to more connected traditional landscapes. Our results suggest that extensive habitat fragmentation does not necessarily translate into reduced genetic connectivity in highly mobile, generalist species. In P. rapae , high dispersal ability and ecological generalism appear to buffer against the genetic consequences of landscape modification, resulting in panmictic population structure even across strongly contrasting agricultural landscapes.

Pieris rapae

Pericentromeric heterochromatin and A-T contents during Robertsonian fusion in the house mouse.

The pericentromeric heterochromatin of meiotic trivalents formed by the Robertsonian (Rb) chromosomes and the two homologous acrocentrics in the house mouse was evaluated by static cytophotometry after selective staining. To reveal pericentromeric heterochromatin specifically, C-banding Giemsa and Hoechst 33258 stains were utilized. Five different Rb chromosomes were investigated and none of them possessed less pericentromeric heterochromatin than the sum of the two homologous acrocentrics. Moreover the total A-T (DAPI) and DNA (PI) content was quantitatively evaluated, by flow cytometry, in G0/G1 nuclei belonging to four different Rb mouse populations, karyotypically characterized by the presence of up to nine Rb chromosomes. Again there were no significant difference, of DAPI and PI content, in the Rb populations nor between any of them and the NMRI/HAN strain with forty acrocentric chromosomes. We conclude that the main consequence of Robertsonian processes (i.e. the rapid variation of the karyotype structure) does not imply detectable quantitative variation in the genome portion involved in the Rb process. We also discuss the possibility that the high rate of Rb exchange in the house mouse could be favoured by the simultaneous effects of undetectable losses of chromosomal material, high repetitiveness of the DNA involved, the presence of the same major type of satellite DNA over each chromosome and the all acrocentric constitution of the karyotype.

Adenine

Panorama of Chromosomal Instability in Lung Cancer.

Lung cancer is a highly heterogeneous disease primarily driven by tobacco smoking. About 20% of lung cancers occur among patients who have never smoked (LCINS) with differences in patient ancestry, sex, tumor histology, and clinical features. Our understanding of chromosomal instability in lung cancer, especially LCINS, is still limited. Here, we perform a comprehensive study of 182,429 somatic structural variations (SVs) detected in 1,209 whole-genome sequenced lung cancers, of which 864 LCINS. SVs are more abundant in tumors from patients who have smoked (LCSS); however, they are more complex and play more important roles in tumorigenesis in LCINS. EGFR mutations and KRAS mutations profoundly and independently shape the SV landscape. EGFR-mutant tumors have higher SV burden and more cancer-driving SVs. In contrast, KRAS mutations are associated with lower SV burden and less driver SVs. We decompose 16 SV signatures for both complex and simple SVs that likely represent divergent molecular mechanisms. The SV breakpoints have distinct distributions across the genome depending on the signatures due to mutagenic mechanisms and positive selection. Many established cancer-driving genes are recurrently rearranged by multiple SV signatures suggesting functional convergence of these genome instability mechanisms.

Journal Article

Dual-dimensional profiling of host genomic variations and HPV integration in PD-L1-stratified cervical cancer via Oxford Nanopore Technology.

BACKGROUND: The integration of human papillomavirus (HPV) DNA into the host genome is a key step in the development of HPV-associated cervical cancer (CC). However, the genomic characteristics of host genomic variations and HPV integration within the context of programmed death-ligand 1 (PD-L1) expression stratification have not been systematically investigated. METHODS: Whole-genome sequencing was performed using Oxford Nanopore Technology (ONT) on six samples (three from the high PD-L1 expression group and three from the low PD-L1 expression group). The characteristics of host genomic variations under different PD-L1 expression stratifications were explored, including structural variations (SV), copy number variations (CNV), single nucleotide polymorphisms (SNP), and insertion-deletions (Indel). Subsequently, the distribution features of HPV integration sites were analyzed, different integration types were identified, and pathway analysis was conducted. RESULTS: Whole-genome SV analysis revealed that the total number of SVs and the composition of mutation types were similar between the high and low PD-L1 expression groups, with insertions (INS) and deletions (DEL) predominating in both. These variations were primarily enriched in intergenic regions and introns. In the low PD-L1 expression group, integration events were observed at multiple chromosomal loci, with the most frequent integration occurring in the KLF5 gene region on chromosome 13. No frequently integrated loci were identified in the high PD-L1 expression group. Additionally, four distinct HPV integration breakpoint patterns were preliminarily identified and analyzed. CONCLUSION: PD-L1 expression stratification did not significantly alter the overall genomic instability of the host. However, differences were observed in the distribution patterns of HPV integration sites. These findings provide new insights into the genomic heterogeneity of CC under different PD-L1 expression backgrounds and may lay the groundwork for future research exploring stratified immunotherapy based on HPV integration features.

Humans

Optical mapping in Black genomes: Distinct LCR22 structures and 22q11.2 deletion syndrome mechanisms.

PURPOSE: The genomic architecture of 22q11.2 deletion syndrome (22q11.2DS) has primarily been studied in White populations, despite evidence suggesting a lower prevalence in Black individuals. This study aims to improve our understanding of the population-specific organization of 22q11.2 genomic structures. METHODS: Optical mapping data from 106 genomes, representing various Black and White individuals, were analyzed to assess the structure and variation of the 22q11.2 low copy repeats (LCR22s). RESULTS: Extensive variability in copy-number and orientation of LCR22 elements was observed between Black and White genomes. Several novel copy-number variants and haplotype configurations were identified, some being private or more prevalent within specific groups. Notably, copy-number variants diversity was particularly striking among Black genomes. Comparisons of Black and White families with de novo 22q11.2DS probands revealed unique nonallelic homologous recombination scenarios, with Black families exhibiting recombination patterns that are not previously observed. CONCLUSION: Perhaps the unique and highly variable LCR22 haplotype configurations in Black individuals contribute to the lower observed prevalence of 22q11.2DS by inhibiting the likelihood of nonallelic homologous recombination, the mechanism that leads to the syndrome.

Humans

Primulina pan-genome reveals differential gene retention following whole-genome duplications and provides insights into edaphic specialization.

Primulina, a genus of >200 species specialized to extreme soils, provides a model for edaphic adaptation. We assemble seven genomes and construct a pan-genome spanning nine species from karst, Danxia, and acidic soils. Comparative analyses reveal that karst-adapted species have smaller genomes. Two lineage-specific whole-genome duplications (WGDs) exhibit biased duplicate loss in large gene families but preferential retention of transcription factors, indicating combined adaptive and nonadaptive forces. Pan-genome analyses identify ion channel and transporter genes enriched in variant hotspots and under positive selection in karst lineages. Candidate genes for drought and salt stress tolerance include ABC transporters and ion channels. Notably, an ABC transporter shows positive selection in karst species and unique structural variation in non-karst species. Together, our findings show that genome downsizing, biased post-WGD retention, and evolution of ion-transport pathways shape adaptation to extreme soils. The Primulina pan-genome provides a resource for dissecting mechanisms underlying edaphic specialization.

Gene Duplication

Decoding missense variants pleiotropy in the immune GPCR P2RY8.

G protein-coupled receptors (GPCRs) form the largest family of cell surface receptors and remain a central focus in pharmacology and drug discovery. Despite extensive structural and pharmacological studies, the functional impact of missense variation across GPCRs remains poorly understood, particularly for receptors involved in immune regulation. In this issue of Cell Genomics, LaFlam et al.1 systematically map P2RY8 variant functions using deep mutational scanning (DMS) combined with structural biology approaches, revealing pleiotropy and mechanisms linking GPCR variation to B cell confinement and lymphoma.

Humans

Lineage-associated small inversions disrupt dosT, dnaE2, and a promoter-adjacent region in some Mycobacterium tuberculosis isolates.

UNLABELLED: Large molecular inversions in the genome of Mycobacterium tuberculosis (Mtb) due to factors like the presence of insertion sequences and transposases are widely known. However, smaller inversions within coding sequences and non-coding control elements are rarely reported. The present study aims to identify inversions and their potential impact on Mtb biology in a lineage-specific manner. Structural variants (SVs) could only be detected by long reads. For this, we simulated long reads by de novo assembling the short-read sequencing data sets and subsequently aligned representative strains from each lineage using the Progressive Mauve algorithm. Independently, long-read sequencing from the Pacific Biosciences platform was acquired and analyzed using the structural variant identification method. Variants were merged, and Fisher's exact test was carried out to identify the inversion association with lineages. To visualize deoxyribonucleic acid (DNA) features, the DNA-features-viewer tool was used. Simulated reads from short-read sequencing gave indications of lineage (L)-specific inversions. The long-read sequencing approach led to the identification of seven unique inversions: two positively associated with L1, one positively associated with L3, two negatively associated with L4, and two positively associated with L3 but negatively associated with L4 (P < 0.05). The inversions encompassed primarily non-essential genes like sdaA, dosT, Rv2026c, dnaE2, Rv1341, Rv1342, and lprD. An interesting inversion was observed in the upstream control element of purB and Rv0776c. The study sheds light on small inversions that may be causing alterations in expression, formation of fusion genes, and nonsense mutations that may have a role in lineage-specific phenotypic changes. IMPORTANCE: The role of mutations like SNPs and INDELs and their association with drug resistance is well known in Mycobacterium tuberculosis (Mtb). However, structural variations, especially inversions, are largely overlooked and unreported. In this paper, publicly available whole-genome sequencing datasets from Illumina and Pacific Biosciences-Oxford Nanopore Technologies platform have been used to detect inversions and report seven unreported Mtb lineage-specific small inversions.

Mycobacterium tuberculosis

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common &#x223c;4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

A chromosome-scale genome of Capsicum pubescens provides insights into candidate terpene-associated gene clusters and pan variation of terpene synthases.

A chromosome-scale genome of Capsicum pubescens and comparative pan-TPS analysis support structural characterization and gene-level prioritization of a chromosome-9 terpene-associated candidate locus in this accession. Capsicum pubescens is one of the five domesticated Capsicum species, mainly cultivated in mid- to high-elevation regions of the Americas. Despite its distinctive morphology and fruit traits, genomic resources for C. pubescens remain less developed than those for the widely cultivated C. annuum. Here, we assembled a chromosome-scale reference genome for accession HNUCP0001, spanning 3.70&#xa0;Gb with a scaffold N50 of 278.01&#xa0;Mb. Comparative genomics revealed 679 significantly expanded gene families enriched in sesquiterpenoid and triterpenoid biosynthesis. Genome-wide biosynthetic gene-cluster mining identified multiple terpene-associated candidate loci, which were subsequently prioritized using genome-derived structural criteria and Capsicum pubescens-specific expression evidence. Subsequently, we curated the terpene synthase (TPS) repertoire and, across 16 Capsicum genomes, resolved 36 TPS orthogroups with pronounced presence/absence variation, highlighting dynamic lineage-specific diversification. Together, these analyses establish HNUCP0001 as an accession-specific genomic resource and provide a comparative framework for prioritizing terpene-associated TPS genes and candidate BGCs in Capsicum. These candidate loci, together with accession-level transcriptomic and metabolomic evidence, offer testable hypotheses for future functional studies of specialized terpenoid metabolism in C. pubescens.

Alkyl and Aryl Transferases

Integrated multi-omics analyses provide new insights into genomic variation landscape and regulatory network candidate genes associated with walnut endocarp.

Persian walnut (Juglans regia) is an economically important nut oil tree; the fruit has a hard endocarp/shell to protect seeds, thus playing a key role in its evolution, and the shell thickness is an important trait for walnut breeding. However, the genomic landscape and the gene regulatory networks associated with walnut shell development remain to be systematically elucidated. Here, we report a high-quality genome assembly of the walnut cultivar 'Xiangling' and construct a graphic structure pan-genome of eight Juglans species to reveal the genetic variations at the genome level. We re-sequence 285 accessions to characterize the genomic variation landscape. Through genome-wide association studies (GWAS), we identified 19 loci associated with more than 268 loci that underwent selection during walnut domestication and improvement. Multi-omics analyses, including transcriptomics, metabolomics, DNA methylation, and spatial transcriptomics across eleven developmental stages, revealed several candidate genes related to secondary cell biosynthesis and lignin accumulation. This integrated multi-omics approach revealed several candidate genes associated with secondary cell biosynthesis and lignin accumulation, such as UGP, MYB308, MYB83, NAC043, NAC073, CCoAOMT1, CCoAOMT7, CHS2, CESA7, LAC7, COBL4, and IRX12. Overexpression of JrUGP and JrMYB308 in Arabidopsis thaliana confirmed their roles in lignin biosynthesis and cell wall thickening. Consequently, our comprehensive multi-omics findings offer novel insights into walnut genetic variation and network regulation of endocarp development and shell thickness, which enable further genome-informed breeding strategies for walnut cultivar improvement.

Juglans

Polygenic and monogenic adaptation drive evolutionary rescue at different magnitudes of environmental change.

Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.

Arabidopsis

Evolution and domestication-trait associations of ultra-long centromere haplotypes in pepper plants.

Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9&#x2009;Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.

Centromere

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics