PubMed HealthSearch

SEARCH · PubMed Health

Results for “DNA Copy Number Variations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Ribosomal DNA copy number variation associates with hematological profiles and renal function in the UK Biobank.

The phenotypic impact of genetic variation of repetitive features in the human genome is currently understudied. One such feature is the multi-copy 47S ribosomal DNA (rDNA) that codes for rRNA components of the ribosome. Here, we present an analysis of rDNA copy number (CN) variation in the UK Biobank (UKB). From the first release of UKB whole-genome sequencing (WGS) data, a discovery analysis in White British individuals reveals that rDNA CN associates with altered counts of specific blood cell subtypes, such as neutrophils, and with the estimated glomerular filtration rate, a marker of kidney function. Similar trends are observed in other ancestries. A range of analyses argue against reverse causality or common confounder effects, and all core results replicate in the second UKB WGS release. Our work demonstrates that rDNA CN is a genetic influence on trait variance in humans.

Humans

DNA Methylation Profiling of Pediatric Ectomesenchymoma Supports Embryonal Rhabdomyosarcoma-Like Epigenetic Identity.

Ectomesenchymoma is a rare, biphenotypic pediatric tumor combining rhabdomyoblastic and neuroectodermal differentiation. We characterize two novel cases through integrated genomics and the first report of genome-wide DNA methylation profiling. Both tumors harbored RAS-pathway mutations (HRAS p.Gly13Arg; NRAS p.Gln61His). Methylation analysis, including microdissected components, consistently aligned ectomesenchymoma with the embryonal rhabdomyosarcoma superfamily, revealing a shared myogenic epigenetic program despite neural differentiation. Shared copy-number profiles across distinct histological regions supported a monoclonal origin. Overall, our data support a close biological relationship between ectomesenchymoma and embryonal rhabdomyosarcoma and indicate that RAS-pathway testing and methylation profiling can significantly refine diagnostic precision.

Humans

The genetic control of rapid genome content divergence in Arabidopsis thaliana.

Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1043 resequenced Arabidopsis thaliana genomes using a novel K-mer-based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 candidate trans-acting loci associated with repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, and DNA methylation regulation. The results are consistent with purifying selection acting against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.

Arabidopsis

Exploring precision risk in pediatric vesicoureteral reflux: Innate immune gene variations and reflux outcomes in the RIVUR cohort.

INTRODUCTION: Children with vesicoureteral reflux (VUR) are at increased risk for morbidity from recurrent urinary tract infections (UTIs), yet the factors influencing spontaneous VUR resolution remain poorly defined. This study evaluates whether genetic variations in key urinary innate immune effectors (DEFA1A3, DMBT1, and RNASE7) influences VUR resolution and interacts with prophylaxis to alter clinical response. METHODS: We conducted a secondary analysis of 303 RIVUR participants with available DEFA1A3 and DMBT1 copy number variation (CNV) data and RNASE7 rs1263872 genotype. Primary outcomes were (1) VUR improvement (decrease in grade) and (2) VUR resolution at study exit. Multivariable logistic regression models included genotype, treatment, and their interactions, adjusting for age, sex, baseline grade (high vs low), laterality, bowel/bladder dysfunction, and any UTI. Internal validation used 2000-sample bootstrap with bias-corrected and accelerated confidence intervals and influence diagnostics. RESULTS: Clinical covariates did not significantly predict VUR improvement. Children with DEFA1A3 CNV >5 had higher odds of improvement (OR 2.36, 95% CI 1.12-4.96, p = 0.023), an effect that remained significant in bootstrap analyses. High-grade VUR was associated with lower odds of resolution (OR 0.34, 95% CI 0.12-0.94, p = 0.038). A significant interaction was observed between prophylaxis and high DMBT1 copy number for VUR resolution (interaction OR 2.99, 95% CI 1.11-8.04, p = 0.031); no interaction was seen for improvement. RNASE7 rs1263872 was not associated with either outcome. CONCLUSION: Innate immune gene variation may contribute to heterogeneity in VUR outcomes. High DEFA1A3 copy number was associated with reflux improvement and a DMBT1-prophylaxis interaction was associated with reflux resolution. The results of this study is hypothesis-generating and prompt further evaluation to assess whether a subset of children may experience structural benefit from prophylaxis or have a more favorable natural history based on their innate immune genotype.

Humans

Coalescing single-cell genomes and transcriptomes to decode breast cancer progression.

Understanding epithelial lineages of breast cancer and genotype-phenotype relationships requires direct measurements of the genome and transcriptome of the same single cells at scale. To achieve this, we developed wellDR-seq, a high-genomic-resolution, high-throughput method to simultaneously profile the genome and transcriptome of thousands of single cells. We profiled 33,646 single cells from 12 estrogen-receptor-positive breast cancers and identified ancestral subclones in multiple patients that showed a luminal hormone-responsive lineage, indicating a potential cell of origin. In contrast to bulk studies, wellDR-seq enabled the study of subclone-level gene-dosage relationships, which showed near-linear correlations in large chromosomal segments and extensive variation at the single-gene level. We identified dosage-sensitive and dosage-insensitive genes, including many breast cancer genes as well as sporadic copy-number aberrations in non-cancer cells. Overall, these data reveal complex relationships between copy number and gene expression in single cells, improving our understanding of breast cancer progression.

Breast Neoplasms

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

A stratified urine-based molecular diagnostic and prognostic model for non-muscle-invasive bladder cancer management.

BACKGROUND: Non-muscle-invasive bladder cancer (NMIBC) is characterized by a high recurrence rate requiring lifelong cystoscopic surveillance. Existing urine-based molecular assays mainly rely on mutations or methylation, which fail to capture large-scale genomic instability. Copy number variation (CNV) profiling offers complementary information on tumor evolution and aggressiveness, but its application in urinary diagnosis remains limited. We aimed to integrate CNV and DNA methylation signals from urinary DNA to establish a noninvasive and biologically informed stratified diagnostic model for NMIBC recurrence surveillance and risk stratification. METHODS: Urine samples were prospectively collected from 91 patients (75 evaluable) between June 2021 and August 2023. Shallow whole-genome sequencing (sWGS) was used to detect CNVs at chromosomal arm and focal gene levels, while ONECUT2 promoter methylation was quantified by qPCR. Diagnostic and prognostic performance was evaluated by ROC analysis, Kaplan-Meier survival, and stratified recurrence assessment. RESULTS: We evaluated a stratified diagnostic model combining CNV and ONECUT2 methylation testing in a cohort of 79 patients. CNV analysis alone showed high specificity (0.923) for NMIBC diagnosis. A combined model, using CNV as an initial screen followed by ONECUT2 methylation testing in CNV-positive cases, achieved a sensitivity of 0.783, specificity of 0.981, and a negative predictive value (NPV) of 0.911. This approach reduced the number of required ONECUT2 tests by 35% and identified a high proportion of true-negative patients (98.1%), which may help reduce unnecessary cystoscopy procedures. The model also demonstrated significant prognostic value, with the molecularly defined high-risk group showing significantly shorter recurrence-free survival (RFS) than the low-risk group (median RFS: 4.33 months vs. not reached; p&#x2009;<&#x2009;0.001). Additional, in patients with initially negative cystoscopy after urine sample collection, the model demonstrated a predictive accuracy of 0.922 for recurrence, with molecular positivity observed a median of 9.6 months prior to clinical diagnosis. CONCLUSIONS: Integrating CNV and DNA methylation profiling from urinary DNA provides a powerful and noninvasive molecular framework for NMIBC surveillance. By combining early epigenetic changes with genomic instability signals, this approach enhances recurrence risk assessment and enables earlier detection compared with conventional cystoscopy. It offers a practical route toward personalized and adaptive post-treatment monitoring of NMIBC. TRIAL REGISTRATION: NCT04994197.

Humans

The Fire Ant Social Chromosome Exerts a Major Influence on Genome Regulation.

Supergenes underlying complex trait polymorphisms ensure that sets of coadapted alleles remain genetically linked. Despite their prevalence in nature, the mechanisms of supergene effects on genome regulation are poorly understood. In the fire ant Solenopsis invicta, a supergene containing over 500 individual genes influences trait variation in multiple castes to collectively underpin a colony level social polymorphism. Here, we present results of an integrative investigation of supergene effects on gene regulation. We present analyses of ATAC-seq data to investigate variation in chromatin accessibility by supergene genotype and STARR-seq data to characterize enhancer activity by supergene haplotype. Integration with gene co-expression analyses, newly mapped intact transposable elements (TEs), and previously identified copy number variants (CNVs) collectively reveals widespread effects of the supergene on chromatin structure, gene transcription, and regulatory element activity, with a genome-wide bias for open chromatin and increased expression in the presence of the derived supergene haplotype, particularly in regions that harbor intact TEs. Integrated consideration of CNVs and regulatory element divergence suggests each evolved in concert to shape the expression of supergene encoded factors, including several transcription factors that may directly contribute to the trans-regulatory footprint of a heteromorphic social chromosome. Overall, we show how genome structure in the form of a supergene has wide-reaching effects on gene regulation and gene expression.

Animals

Genomic Analysis of Circulating Tumor Cells at the Single-Cell Level.

Circulating tumor cells (CTCs) have a great potential for noninvasive diagnosis and real-time monitoring of cancer. A comprehensive evaluation of four whole genome amplification (WGA)/next-generation sequencing workflows for genomic analysis of single CTCs, including PCR-based (GenomePlex and Ampli1), multiple displacement amplification (Repli-g), and hybrid PCR- and multiple displacement amplification-based [multiple annealing and loop-based amplification cycling (MALBAC)] is reported herein. To demonstrate clinical utilities, copy number variations (CNVs) in single CTCs isolated from four patients with squamous non-small-cell lung cancer were profiled. Results indicate that MALBAC and Repli-g WGA have significantly broader genomic coverage compared with GenomePlex and Ampli1. Furthermore, MALBAC coupled with low-pass whole genome sequencing has better coverage breadth, uniformity, and reproducibility and is superior to Repli-g for genome-wide CNV profiling and detecting focal oncogenic amplifications. For mutation analysis, none of the WGA methods were found to achieve sufficient sensitivity and specificity by whole exome sequencing. Finally, profiling of single CTCs from patients with non-small-cell lung cancer revealed potentially clinically relevant CNVs. In conclusion, MALBAC WGA coupled with low-pass whole genome sequencing is a robust workflow for genome-wide CNV profiling at single-cell level and has great potential to be applied in clinical investigations. Nevertheless, data suggest that none of the evaluated single-cell sequencing workflows can reach sufficient sensitivity or specificity for mutation detection required for clinical applications.

Carcinoma, Non-Small-Cell Lung

PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase framework.

MOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv.

DNA Copy Number Variations

Graph-KIR: graph-based KIR copy number estimation and allele calling using short-read sequencing data.

MOTIVATION: The Killer-cell Immunoglobulin-like Receptor (KIR) is a highly polymorphic region in the human genome, associated with autoimmune diseases and organ transplantation. The sequences of KIR genes are highly similar among star alleles as well as in between individual genes, with the copy number of each KIR gene typically ranging from 0 to 4. In this study, we introduce Graph-KIR, a tool designed to estimate gene copy numbers and predict full-resolution (7-digit, encompassing both coding and non-coding sequence variations) from a whole genome sequencing (WGS) sample. RESULTS: Graph-KIR is capable of independently typing KIR alleles per sample with no reliance on the distribution of any framework gene in a cohort. In a set of 100 simulated samples, Graph-KIR demonstrated 99.2% accuracy in copy number estimation and high F1-score of allele typing: 91.79% at 7-digit resolution, 97.37% at 5-digit resolution, and 97.11% at 3-digit resolution. Graph-KIR outperforms existing tools such as Geny (96.39% F1-score), PING's WGS version (92.77% F1-score), and T1K (90.44% F1-score) at 5-digit resolution. By analyzing the results on 44 HPRC samples, Graph-KIR achieves better F1-score than Geny and PING at 7-digit resolution. The release of Graph-KIR adds another valuable tool to assist users in accurately estimating copy numbers and calling alleles of KIR genes from WGS samples. AVAILABILITY AND IMPLEMENTATION: The Graph-KIR and paper-related pipeline codes are available at https://github.com/linnil1/KIR_graph.

Receptors, KIR

Effects of particulate air pollution on BPDE-DNA adducts, telomere length, and mitochondrial DNA copy number in human exhaled breath condensate and BEAS-2B cells.

Traffic-related particulate matter (PM) and polycyclic aromatic hydrocarbons (PAHs) have been linked to respiratory diseases and cancer risk in humans. Genomic damage, including benzo[a]pyrene diolepoxide (BPDE)-DNA adducts as well as alterations in telomere length (TL) and mitochondrial DNA copy number (mtDNA-CN) are associated with respiratory diseases. This study aimed to investigate the association between exposure to traffic-related particulate pollutants and genomic damage in exhaled breath condensate (EBC) in human subjects and a bronchial epithelial cell line (BEAS-2B). Among the 60 healthy recruited subjects, residents living in high-traffic-congested areas were exposed to higher concentrations of PM2.5 (1.66-fold, p&#xa0;<&#xa0;0.01), UFPs (1.79-fold, p&#xa0;<&#xa0;0.01), PM2.5-PAHs (1.50-fold, p&#xa0;<&#xa0;0.01), and UFPs-PAHs (1.35-fold, p&#xa0;<&#xa0;0.05), than those in low-traffic-congested areas. In line with increased exposure to particulate air pollution, the high-traffic-exposed group had significantly increased BPDE-DNA adducts (1.40-fold, p&#xa0;<&#xa0;0.05), TL shortening (1.24-fold, p&#xa0;<&#xa0;0.05), and lower mtDNA-CN (1.38-fold, p&#xa0;<&#xa0;0.05) in EBC. The observations in the human study linking exposure to PM2.5, UFPs, PM2.5-PAHs, and UFPs-PAHs with the aforementioned biological effects were confirmed by an in vitro cell-based study, in which BEAS-2B cells were treated with diesel exhaust particulate matter (DEP) containing fine and ultrafine PM and PAHs. Increased BPDE-DNA adducts levels, shortened TL, and decreased mtDNA-CN were also found in treated BEAS-2B cells. The shortened TL and decreased mtDNA-CN were in part mediated by decreased transcript levels of hTERT, and SIRT1, which are involved in telomerase activity and mitochondrial biogenesis, respectively. These results suggest that exposure to traffic-related particulate pollutants can cause genomic instability in respiratory cells, which may increase the health risk of respiratory diseases and the development of cancer.

Humans

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to detect CNVs, which can be used in genetic association tests. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (those present on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has led to excessive reduction in overall sensitivity since arrays can have an undesirably low probe overlap. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4906 individuals genotyped across three different arrays, we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g. use of consensus probes only). We further demonstrate that MarkerMatch matches the CNV detection from current practice in terms of F1 score and PPV for larger CNVs. We also optimize MarkerMatch parameters, DMAX and Method, and find an optimal DMAX setting at 10&#x2009;kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis. AVAILABILITY: The R package for MarkerMatch is available at: https://github.com/FranjoIM/MarkerMatch. The code used for analysis and implementation is available at: https://doi.org/10.5281/zenodo.18460979. The live notebook is available at https://fivankovic.notion.site/2026-markermatch.

DNA Copy Number Variations

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean&#x2009;=&#x2009;0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Simultaneous detection of glyphosate and glufosinate target-site resistance in Eleusine indica via multiplex TaqMan qPCR.

BACKGROUND: Continuous use of glyphosate followed by glufosinate-ammonium has selected for multiple resistance to both herbicides in Eleusine indica worldwide. Managing such resistant weeds requires fast, accurate molecular detection assay. To address this critical need, we developed a robust multiplex TaqMan quantitative (q)PCR assay that simultaneously detects five well-characterized target-site resistance markers in E.&#x2009;indica: EPSPS copy number variation; T102I in EPSPS; P106A and P106S in EPSPS; and S59G in GS1-1. RESULTS: The multiplex qPCR assay showed analytical specificity when tested on genomic DNA from nine reference accessions: three susceptible, three glyphosate-resistant (with EPSPS CNV) and three multiple-resistant. Subsequent analysis of 56 field-collected samples demonstrated 98.2% concordance (55 of 56) with Sanger sequencing across all five resistance-associated markers: EPSPS CNV, T102I, P106A, P106S and GS1-1 S59G, confirming the reliability and practical value of the multiplex qPCR assay. Only samples 7-8 showed discordance at EPSPS position 102, where Sanger chromatograms showed overlapping peaks at this position, which is likely to be a result of heterozygous mutation distribution among amplified EPSPS gene copies. This case further underscores the advantages of the multiplex qPCR assay over Sanger sequencing in detection sensitivity and accuracy. Moreover, a strong correlation (R2&#x2009;=&#x2009;0.8935) in gene copy number estimation between the two methods across all samples further supports the reliability of the qPCR assay. CONCLUSIONS: In summary, this study delivers a simple, robust and high-throughput diagnostic tool for the rapid, simultaneous identification of dual herbicide target-site resistance in goosegrass, offering superior sensitivity, quantitative resolution and throughput compared with Sanger sequencing. &#xa9; 2026 Society of Chemical Industry.

Herbicides

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome&#x2011;wide coverage of only 0.1-5&#xd7;, sWGS data display a pronounced zero&#x2011;inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several&#x2011;fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy&#x2011;number gains (false positives), and true deletions often become indistinguishable from pervasive zero&#x2011;coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

Binary vector copy number engineering improves Agrobacterium-mediated transformation.

The copy number of a plasmid is linked to its functionality, yet there have been few attempts to optimize higher-copy-number mutants for use across diverse origins of replication in different hosts. We use a high-throughput growth-coupled selection assay and a directed evolution approach to rapidly identify origin of replication mutations that influence copy number and screen for mutants that improve Agrobacterium-mediated transformation (AMT) efficiency. By introducing these mutations into binary vectors within the plasmid backbone used for AMT, we observe improved transient transformation of Nicotiana benthamiana in four diverse tested origins (pVS1, RK2, pSa and BBR1). For the best-performing origin, pVS1, we isolate higher-copy-number variants that increase stable transformation efficiencies by 60-100% in Arabidopsis thaliana and 390% in the oleaginous yeast Rhodosporidium toruloides. Our work provides an easily deployable framework to generate plasmid copy number variants that will enable greater precision in prokaryotic genetic engineering, in addition to improving AMT efficiency.

Genetic Vectors