PubMed HealthSearch

SEARCH · PubMed Health

Results for “Copy Number Variations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Determinants of functional burden pleiotropy and gene dosage responses across human traits.

Pleiotropic and monotonic effects of gene dosage are central to understanding comorbidities in developmental pediatric and psychiatric disorders, yet the underlying biological processes are not well characterized. Here we develop a functional burden analysis to investigate the association of all protein-coding copy-number variants, genome-wide, with 43 complex traits in approximately 500,000 UK Biobank participants. We test variant associations disrupting 172 tissue or cell-type gene sets, finding associations for all traits, which we replicate in the All of Us cohort. Functional burden pleiotropy, defined as the number of traits significantly associated with a gene set, correlates with genetic constraint and is higher for brain than non-brain functions, even after normalizing for genetic constraint. Levels of pleiotropy, measured by burden correlation, are similar in deletions and loss-of-function single-nucleotide variants, and higher than in common variants and duplications. Most gene dosage responses are non-monotonic, with deletions and duplications showing same-direction effects, and monotonic responses decrease with genetic constraint. We observe associations between functional gene sets and traits for either deletions or duplications, but rarely both, with negatively correlated effect sizes. Together, these results link genetic constraint and brain-specific mechanisms to the whole-body multimorbidity of neurodevelopmental and psychiatric conditions.

Humans

Aneuploidy selects for the acquisition of driver genes in breast cancer.

Chromosome instability is highly prevalent in cancer and drives large-scale chromosomal imbalances, known as aneuploidies1-4. How aneuploidy contributes to tumorigenesis remains difficult to study due to the vast numbers of genes affected. Here we established a CRISPR knockout- and activation-linked assay (CRISPR-KOALA), enabling high-throughput bidirectional genetic screens in immunocompetent mouse models of cancer. We developed a compendium of the ten most frequent human chromosome-arm-level alterations in basal-like breast cancer (BLBC), a disease type that is driven by large copy-number alterations (CNAs)5-8. Using CRISPR-KOALA, we screened the mouse orthologues of 3,752 genes on these arms and identified 90 cancer driver genes, the function of the vast majority of which is unknown. These genes drive distinct signalling pathways including MAPK, HIPPO and WNT, reflecting the high degree of BLBC heterogeneity. Manipulating the identified cancer driver genes overcomes the need for CNAs in Trp53-mutant BLBC mouse models. Mechanistically, we identify that PLGRKT is a potent oncogene that lies on chromosome 9p and show that its tumour-promoting activity is associated with highly stress-resistant mitochondria and an increased ability to detoxify reactive oxygen species. Together, our findings reveal that arm-level CNAs can function to select specific driver genes to promote heterogeneous biological processes.

Animals

Binary vector copy number engineering improves Agrobacterium-mediated transformation.

The copy number of a plasmid is linked to its functionality, yet there have been few attempts to optimize higher-copy-number mutants for use across diverse origins of replication in different hosts. We use a high-throughput growth-coupled selection assay and a directed evolution approach to rapidly identify origin of replication mutations that influence copy number and screen for mutants that improve Agrobacterium-mediated transformation (AMT) efficiency. By introducing these mutations into binary vectors within the plasmid backbone used for AMT, we observe improved transient transformation of Nicotiana benthamiana in four diverse tested origins (pVS1, RK2, pSa and BBR1). For the best-performing origin, pVS1, we isolate higher-copy-number variants that increase stable transformation efficiencies by 60-100% in Arabidopsis thaliana and 390% in the oleaginous yeast Rhodosporidium toruloides. Our work provides an easily deployable framework to generate plasmid copy number variants that will enable greater precision in prokaryotic genetic engineering, in addition to improving AMT efficiency.

Genetic Vectors

Comprehensive molecular profiling of multiple myeloma identifies refined copy number and expression subtypes.

Multiple myeloma is a treatable, but currently incurable, hematological malignancy of plasma cells characterized by diverse and complex tumor genetics for which precision medicine approaches to treatment are lacking. The Multiple Myeloma Research Foundation's Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile study ( NCT01454297 ) is a longitudinal, observational clinical study of newly diagnosed patients with multiple myeloma (n = 1,143) where tumor samples are characterized using whole-genome sequencing, whole-exome sequencing and RNA sequencing at diagnosis and progression, and clinical data are collected every 3 months. Analyses of the baseline cohort identified genes that are the target of recurrent gain-of-function and loss-of-function events. Consensus clustering identified 8 and 12 unique copy number and expression subtypes of myeloma, respectively, identifying high-risk genetic subtypes and elucidating many of the molecular underpinnings of these unique biological groups. Analysis of serial samples showed that 25.5% of patients transition to a high-risk expression subtype at progression. We observed robust expression of immunotherapy targets in this subtype, suggesting a potential therapeutic option.

Humans

MelanoDB: A dataset of clinical and molecular features of patients with advanced melanoma treated with MAPK inhibitors.

MAPK inhibitors (MAPKi) have revolutionized the treatment of patients with advanced melanoma. However, primary and acquired resistance mechanisms limit their efficacy. Predicting MAPKi response from the tumor baseline features remains challenging due to the limited size of patient cohorts. Therefore, we collected data from nine different patient cohorts (total n = 417 patients with advanced melanoma treated with MAPKi) to identify clinical and molecular features. Our curated dataset, named MelanoDB, includes whole or partial exome sequencing data for 191 patients, copy number alteration information for 66 patients, and gene expression data for 132 patients. We provide a web application to explore the integrated dataset and data distribution across the collected studies, and we share this dataset with the scientific community according to the Findable, Accessible, Interoperable, Reusable (FAIR) principles.

Humans

Transfer learning with multiomics integration and deep neural networks reveals drug resistance mechanisms in cancer.

Drug resistance remains one of the primary challenges in effective cancer therapy. In this study, we employed a deep neural network (DNN)-based transfer learning (TL) approach to predict drug response and uncover drug resistance mechanisms. We integrated gene expression, somatic mutation, and copy number aberration (CNA) data with drug response profiles using multi-omics integration (MI). We used the Genomics of Drug Sensitivity in Cancer (GDSC) data for training and incorporated drugs with same pathways into the training models. We then evaluated drug response predictions on independent in-vivo PDX Encyclopedia (PDX) and ex-vivo the Cancer Genome Atlas (TCGA) datasets. In addition, we conducted pathway enrichment analyses to elucidate the mechanisms underlying drug resistance for paclitaxel, 5-fluorouracil (5-FU), gemcitabine, and cetuximab. We also applied Fisher's exact test (FET) to assess potential associations between drug resistance and the presence of mutations or CNAs. Our pan-drug models outperformed other methods based on the area under the precision-recall curve (AUCPR). Our pathway enrichment analyses revealed LDHB-mediated pyruvate metabolism and FYN-mediated focal adhesion might have pivotal roles in paclitaxel resistance, while PINK1-mediated mitophagy might be critical in 5-FU resistance. In addition to transcriptional activation, FET suggested that CNAs in LDHB and PINK1 may also be associated with resistance to paclitaxel and 5-FU, respectively. Furthermore, enrichment results for paclitaxel and cetuximab indicated shared resistance mechanisms between the two drugs. Importantly, our findings are consistent with prior experimental studies, providing literature-based validation of our results. Overall, our DNN-based TL approach achieved strong predictive performance across PDX & TCGA datasets and enrichment analyses provided valuable biological insights into drug resistance mechanisms.

Humans

Single-copy transgenic mice with chosen-site integration.

We describe a general way of introducing transgenes into the mouse germ line for comparing different sequences without the complications of variation in copy number and insertion site. The method uses homologous recombination in embryonic stem (ES) cells to generate mice having a single copy of a transgene integrated into a chosen location in the genome. To test the method, a single copy murine bcl-2 cDNA driven by either a chicken beta-actin promoter or a human beta-actin promoter has been inserted immediately 5' to the X-linked hypoxanthine phosphoribosyltransferase locus by a directly selectable homologous recombination event. The level of expression of the targeted bcl-2 transgene in ES cells is identical in independently isolated homologous recombinants having the same promoter yet varies between the different promoters. In contrast, the expression of bcl-2 transgenes having the same (chicken beta-actin) promoter varies drastically when they are independently integrated at random insertion sites. Both promoters direct broad expression of the single-copy transgene in mice derived from the respective targeted ES cells. In vitro and in vivo, the human beta-actin promoter consistently directed a higher level of transgene expression than the chicken beta-actin promoter.

Animals

Single-cell copy number calling and event history reconstruction.

MOTIVATION: Copy number alterations are driving forces of tumour development and the emergence of intra-tumour heterogeneity. A comprehensive picture of these genomic aberrations is therefore essential for the development of personalised and precise cancer diagnostics and therapies. Single-cell sequencing offers the highest resolution for copy number profiling down to the level of individual cells. Recent high-throughput protocols allow for the processing of hundreds of cells through shallow whole-genome DNA sequencing. The resulting low read-depth data poses substantial statistical and computational challenges to the identification of copy number alterations. RESULTS: We developed SCICoNE, a statistical model and MCMC algorithm tailored to single-cell copy number profiling from shallow whole-genome DNA sequencing data. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells. We show the accuracy of this approach in evaluations on simulated data and demonstrate its practicability in applications to two breast cancer samples from different sequencing protocols. AVAILABILITY AND IMPLEMENTATION: SCICoNE is available at https://github.com/cbg-ethz/SCICoNE.

Single-Cell Analysis

CCRR: a user-friendly platform for analyzing complex chromosomal rearrangements in tumors.

SUMMARY: Complex chromosomal rearrangements in tumors involve intricate genomic alterations that significantly affect gene function and contribute to cancer development. Identifying these events is crucial for cancer research but is often challenging due to the complexity and limitations of existing tools. We developed the Complex Chromosomal Rearrangements Resolver (CCRR), a comprehensive, reproducible, and user-friendly platform for analyzing complex rearrangements in tumors. CCRR integrates multiple SV and CNV detection tools within a Docker container environment, simplifying installation and configuration. It can be easily deployed, automating the execution and merging of results, providing high-confidence consensus SV and CNV calls, allowing researchers to efficiently analyze complex chromosomal rearrangements in tumors without extensive bioinformatics expertise. CCRR also includes a web server for one-click analysis and customized visualization. AVAILABILITY AND IMPLEMENTATION: The CCRR platform is freely available at https://www.ccrr.life. Source code and executables can be accessed at https://github.com/laslk/CCRR. An archived version is available at Zenodo: https://doi.org/10.5281/zenodo.15386513.

Software

MPAC: a computational framework for inferring pathway activities from multi-omic data.

MOTIVATION: Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. RESULTS: We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g. associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell composition. Our MPAC R package enables similar multi-omic analyses on new datasets. AVAILABILITY AND IMPLEMENTATION: The MPAC package is available at Bioconductor https://bioconductor.org/packages/MPAC.

Humans

ELViS: an R package for estimating copy number levels of viral genomic segments at base-resolution.

MOTIVATION: Tumor viruses account for ∼10% of cancer diagnoses. Virally induced tumorigenesis is understood as direct signaling through oncogenes such as E6 and E7 genes in the case of human papillomavirus. Furthermore, pathogen characteristics such as viral oncogene dose may impact the disease course. To our knowledge, no tool has been proposed to assess the intra-viral copy number alterations that define the gene dose of viral oncogenes and associated suppressive pathways native to the pathogen's normal life cycle. RESULTS: We propose an R package, "ELViS," that analyzes viral copy number changes from DNA sequencing of whole viral genomes. The method adjusts for viral load with 2D transformation and segmentation to offer the relative viral gene doses. AVAILABILITY AND IMPLEMENTATION: The ELViS R package is available from https://bioconductor.org/packages/ELViS. This article used controlled access data from dbGaP (phs001713.v1.p1).

Software

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

Phylogenetic screening of the human genome: identification of differentially hybridizing repetitive sequence families.

The phi-screen, a method of phylogenetic screening, can be employed to detect repetitive sequence families that differentially hybridize between closely related species. Such differences may involve sequence divergence or variations in copy number, including total presence versus absence of a family of repeated DNA. We present the results of a phi-screen comparing the human genome to that of the prosimian, Galago crassicaudatus. Three human repetitive families that are divergent or not present in galago have been detected. One of these families is described in detail; it is similar among the anthropoids but is present in a lower copy number and/or divergent form in prosimians. The family is clearly related to the transposon-like human element (THE) described by Paulson et al. (1985). THEs have long terminal repeats reminiscent of retroviruses but are unique in that they have no sequence similarity to known mammalian retroviruses. The sequence of a solo long terminal repeat, found unassociated with THE internal sequence, is presented. This family member, THE p2, is bordered by a 5-bp target-site repeat and is interrupted by the insertion of an Alu element. A solo THE element sequenced by Wiginton et al. (1986) contains an insertion of Alu at precisely the same position as does THE p2.

Animals

Multiple features of cell-free mtDNA for predicting transarterial chemoembolization response in hepatocellular carcinoma.

BACKGROUND: Transarterial chemoembolization (TACE) is the primary treatment modality for advanced HCC, yet its efficacy assessment and prognosis prediction largely depend on imaging and serological markers that possess inherent limitations in terms of real-time capability, sensitivity, and specificity. Here, we explored whether multiple features of cell-free mitochondrial DNA (cf-mtDNA), including copy number, mutations, and fragmentomics, could be used to predict the response and prognosis of patients with HCC undergoing TACE treatment. METHODS: A total of 60 plasma cell-free DNA samples were collected from 30 patients with HCC before and after the first TACE treatment and then subjected to capture-based mtDNA sequencing and whole-genome sequencing. RESULTS: Comprehensive analyses revealed a clear association between cf-mtDNA multiple features and tumor characteristics. Based on cf-mtDNA multiple features, we also developed HCC death and progression risk prediction models. Kaplan-Meier curve analyses revealed that the high-death risk or high-progression-risk group had significantly shorter median overall survival (OS) and progression-free survival than the low-death risk or low-progression-risk group (all p<0.05). Moreover, the change in cf-mtDNA multiple features before and after TACE treatment exhibited an exceptional ability to predict the risk of death and progression in patients with HCC (log-rank test, all p<0.01; HRs: 0.36 and 0.33, respectively). Furthermore, we observed the consistency of change between the cf-mtDNA multiple features and copy number variant burden before and after TACE treatment in 40.00% (12/30) patients with HCC. CONCLUSIONS: Altogether, we developed a novel strategy based on profiling of cf-mtDNA multiple features for prognosis prediction and efficacy evaluation in patients with HCC undergoing TACE treatment.

Humans

Strong but diffuse genetic divergence underlies differentiation in an incipient species of marine stickleback.

Understanding how lineages proceed along the "speciation continuum" and how species boundaries are maintained over time remain central questions in evolutionary biology. Populations early in the speciation process can give us detailed insight into the reproductive barriers that first initiate speciation. In this study, we explore the nature of genomic divergence between two sympatric marine stickleback ecotypes from Atlantic Canada, "whites" and "commons". Males of each ecotype exhibit distinct nuptial colorations, nesting habits, and parental care strategies. Using population genomic analyses of SNPs and copy number variants (CNVs; deletions and duplications) we show that whites and commons consistently form distinct populations. We uncover genomic differentiation in the white ecotype characteristic of an incipient species, showing extremely low genome-wide differentiation (FST) and very recent divergence (~1 kya). Demographic analysis detected very low levels of ongoing gene flow between populations. Our results and prior genomic studies suggest that reproductive isolation is being maintained between ecotypes despite recent evidence that hybridization in nature does occur. Contrary to other systems, we found many small, but dispersed regions of high differentiation throughout the genome rather than explicitly within chromosomal inversions or the sex chromosomes. On chromosomes VII and XVI, we identified CNVs overlapping genes enriched for olfaction, which may play a role in differences in reproductive strategies between ecotypes. Ultimately, our results demonstrate that genome-wide rather than localized differences can underlie the early stages of divergence, and that this pattern is corroborated by both SNPs and CNVs.

Copy Number Variation

Reproducible autosomal gene expression changes with loss of typical X and Y complement across tumor types.

Although there are known sex differences in cancer incidence, severity, and treatment, the sex chromosomes are typically excluded from genomic analyses because of the unique technical challenges associated with assessing their copy number, sequence variation, and expression. Here we assess sex chromosome complement in three widely-used human genomics datasets from normal (non-cancerous) tissues, primary tumors, and cancer cell lines and study the effects on genome-wide gene expression. Expected sex chromosome complements based on reported patient sex were observed in non-cancerous tissues, but about half of tumors and cancer cell lines showed loss of typical sex chromosome gene expression across tissue types with three categories: loss of chromosome Y (LOY), loss of chromosome X (LOX) and reactivation of the inactive X chromosome (XaXa). Genes consistently differentially expressed in tumors with loss of chromosome X, loss of chromosome Y, or loss of X chromosome inactivation are associated with the hallmarks of cancer and include both sex-linked and autosomal genes from nearly all chromosomes, druggable genes, and genes with molecular functions relevant to cancer signaling, such as kinase activity. Strikingly, tumors that are X0, including tumors from female patients that have lost an X chromosome and tumors from male patients that have lost a Y chromosome, cluster together by gene expression profile. Patients with tumors that have LOX or LOY had poorer survival outcomes compared to those with tumors that had maintained their sex chromosome complement. Further, LOX and LOY eliminates nearly all of the differential gene expression between tumors from different patient sexes, affecting sex chromosomal and autosomal gene expression. Going forward, considering patient sex as well as the entire genome, including assessment of the sex chromosome complement, will provide additional insights into personalized tumor etiology, progression, treatment, and patient outcome.

Journal Article

Unique genetic basis of the distinct antibiotic potency of high acetic acid production in the probiotic yeast Saccharomyces cerevisiae var. boulardii.

The yeast Saccharomyces boulardii has been used worldwide as a popular, commercial probiotic, but the basis of its probiotic action remains obscure. It is considered conspecific with budding yeast Saccharomyces cerevisiae, which is generally used in classical food applications. They have an almost identical genome sequence, making the genetic basis of probiotic potency in S. boulardii puzzling. We now show that S. boulardii produces at 37&#xb0;C unusually high levels of acetic acid, which is strongly inhibitory to bacterial growth in agar-well diffusion assays and could be vital for its unique application as a probiotic among yeasts. Using pooled-segregant whole-genome sequence analysis with S. boulardii and S. cerevisiae parent strains, we succeeded in mapping the underlying QTLs and identified mutant alleles of SDH1 and WHI2 as the causative alleles. Both genes contain a SNP unique to S. boulardii (sdh1 F317Y and whi2 S287*) and are fully responsible for its high acetic acid production. S. boulardii strains show different levels of acetic acid production, depending on the copy number of the whi2 S287* allele. Our results offer the first molecular explanation as to why S. boulardii could exert probiotic action as opposed to S. cerevisiae They reveal for the first time the molecular-genetic basis of a probiotic action-related trait in S. boulardii and show that antibacterial potency of a probiotic microorganism can be due to strain-specific mutations within the same species. We suggest that acquisition of antibacterial activity through medium acidification offered a selective advantage to S. boulardii in its ecological niche and for its application as a probiotic.

Acetic Acid