PubMed HealthSearch

SEARCH · PubMed Health

Results for “Gene copy number”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Gene copy number effects in the mer operon of plasmid NR1.

The level of resistance to Hg2+ determined by the inducible mer operon of plasmid NR1 was essentially the same for three gene copy number variants in Escherichia coli, less in Proteus mirabilis, and intermediate in P. mirabilis "transitioned" to a high r-determinant gene copy number. Cell-free volatilization rates of radioactive mercury indicated increasing levels of intracellular mercuric reductase enzyme from low- to high-gene copy number forms in P. mirabilis and from low- to high-copy number forms in E. coli, but the additional enzyme in E. coli was effectively cryptic.

Enterobacteriaceae

The combined effect of the gene copy number and chaperone overexpression on the recombinant bovine chymosin production in Pichia pastoris, with mutant ADH2 promoter.

Chymosin is an enzyme used to coagulate milk, in the cheese industry. This study aimed to increase recombinant production of the chymosin in Pichia pastoris by determining the optimum copy number and overproduction of a Protein Disulfide Isomerase (PpPDI) chaperon protein. Bos taurus chymosin was expressed under the control of a mutant ADH2 promoter. The clones containing 1-4 gene copy numbers of the chymosin were constructed using the in vitro cloning method, and the effect of chaperone protein on chymosin secretion was investigated. The enzyme production levels are 4, 6.3, 4.5, and 3 IMCU/mL for 1, 2, 3, and 4-copy clones. The secreted chymosin levels increased up to two copies, and increasing the number of copies decreased the secretion level. Therefore, PpPDI was over-expressed in the clones regulated with the ADH2 promoter. The over-expression of PDI gene increased chymosin secretion in clones compared to the counterpart host. However, the highest chymosin level was obtained with C2 (2-copy chymosin containing clone; 6.3 IMCU/mL) and C2P2 (2-copy chymosin/2-copy PDI containing clone; 8.2 IMCU/mL). The maximum production was 39 IMCU/mL with the clone C2P2 in the fermenter scale production. The enzyme activity increased approximately 2-fold by adding two copies of the chaperone protein. The combined effect of gene copy number and chaperone overexpression on chymosin production was investigated. Two copies of the chymosin and PpPDI genes were the optimum among the tested clones.

Animals

Integrating mutation, copy number, and gene expression data to identify driver genes of recurrent chromosome-arm losses.

Aneuploidy is a hallmark of cancer, yet the genes driving recurrent chromosome-arm losses remain largely unknown. We present a systematic framework integrating mutation, copy number, and gene expression data to identify candidate driver genes of cancer type-specific recurrent chromosome-arm losses across 20 cancer types, using ∼7,500 tumors from The Cancer Genome Atlas. By analyzing focal deletions and point mutations that co-occur, or are mutually exclusive, with chromosome-arm losses, we pinpoint 322 candidate drivers associated with 159 recurring events. Our approach identifies known aneuploidy drivers such as TP53 and PTEN, while revealing multiple additional candidates, including tumor suppressors not previously linked to aneuploidy. We leverage expression changes associated with chromosome-arm losses to propose cancer-promoting pathway-level alterations. Integrating these findings highlights key candidate drivers that underlie the observed expression alterations, reinforcing their biological relevance. We provide a comprehensive catalog of candidate driver genes for recurrently lost chromosome-arms in human cancer.

Humans

Transfer of antibiotic resistance genes from soil to rice in paddy field.

The global spread and distribution of antibiotic resistance genes (ARGs) has received much attention whereas knowledge about the transmission of ARGs from one matrix to another is still insufficient. In this study, the paddy fields fertilized with chemical fertilizer, swine compost, and no fertilizer were investigated to assess the transfer of ARGs from soil to rice. Soil and plant samples were collected at day 0, 7, 30 and 79 representing various stages of paddy growth. High throughput qPCR was applied to quantify ARGs using a set of 144 primers. Gene copy number of ARGs measured in soil initially decreased and then increased in soil with no fertilizer and chemical fertilizer, indicating that crop planting and flooding conditions did influence the ARGs profiles in soil. Application of swine compost significantly enhanced the relative abundance and gene copy number of ARGs in paddy soil. Rice seedlings contained substantial amount of ARGs and their relative abundance continually decreased after transplant. Compared with initial stage, detection frequencies of ARGs increased in soil without swine compost at harvest time (day 79), indicating the transmission of ARGs from irrigation water to soil. Detection frequencies of ARGs increased in soil and rice root with swine compost at harvest time, indicating the transfer of ARGs from swine compost to soil and rice root. There was no significant difference in abundance and diversity of ARGs in rice grains with these three different fertilizations. The source of the ARGs in rice grain still needs further exploration.

Oryza

Genomic insights into karyotype evolution and adaptive mechanisms in Polygonaceae species.

Polygonaceae, with ecological versatility and global distribution, is an ideal system for investigating plant adaptation. However, the genomic mechanisms underlying its karyotype evolution and environmental resilience remain unclear. We herein present chromosome-level genomes of 11 species from 10 Polygonaceae genera. Our analyses reveal that Gypsy retrotransposons are key drivers of genome size variations in Polygonaceae. We reconstructed a Polygonaceae ancestral karyotype comprising 28 proto-chromosomes and elucidated evolutionary trajectories via extensive chromosomal rearrangements. Furthermore, we constructed a cross-genus super pan-genome for Polygonaceae, identifying 80,055 gene families, of which 9,845 (12.30%) are core gene families. Private genes are found to contribute significantly to interspecific differences in adaptability. Notably, gene copy number variations are identified as a critical factor influencing adaptations to diverse niches involving species-specific increases in metabolic pathways. This study provides a genomic framework for Polygonaceae karyotype plasticity and adaptive innovation, offering insights into plant evolution under environmental challenges.

Karyotype

Regulated expression by readthrough translation from a plasmid-encoded beta-galactosidase.

We have characterized expression of beta-galactosidase from a plasmid cloning vehicle, pBGP120, which carries most of the lacZ gene and contains a single EcoRI site near the end of lacZ. In addition, we have examined expression of heterologous DNA inserted at the position of the EcoRI site. The EcoRI site was shown to be within the sequence coding for beta-galactosidase and its precise location and phase were deduced. Insertion of heterologous EcoRI-generated DNA fragments altered the molecular weight of the plasmid-encoded beta-galactosidase polypeptide. Those insertions that were in the correct phase were expressed at a high level as a fused protein. The different forms of beta-galactosidase polypeptides produced by various hybrid plasmids were all stable proteins. The level of expression of the plasmid-encoded beta-galactosidase was several times higher than maximal expression of chromosome-encoded beta-galactosidase, suggesting that expression is proportional to gene copy number. The expression of the plasmid lacZ gene was controlled by cyclic AMP. When grown in a cya strain (DG74), expression was dependent on exogenous cyclic AMP. Although in normal strains there was insufficient lac repressor to inactivate all copies of the plasmid, repressor regulation was restored when the plasmid was grown in a strain (M96) that overproduces the lac repressor.

Cyclic AMP

Simultaneous detection of glyphosate and glufosinate target-site resistance in Eleusine indica via multiplex TaqMan qPCR.

BACKGROUND: Continuous use of glyphosate followed by glufosinate-ammonium has selected for multiple resistance to both herbicides in Eleusine indica worldwide. Managing such resistant weeds requires fast, accurate molecular detection assay. To address this critical need, we developed a robust multiplex TaqMan quantitative (q)PCR assay that simultaneously detects five well-characterized target-site resistance markers in E. indica: EPSPS copy number variation; T102I in EPSPS; P106A and P106S in EPSPS; and S59G in GS1-1. RESULTS: The multiplex qPCR assay showed analytical specificity when tested on genomic DNA from nine reference accessions: three susceptible, three glyphosate-resistant (with EPSPS CNV) and three multiple-resistant. Subsequent analysis of 56 field-collected samples demonstrated 98.2% concordance (55 of 56) with Sanger sequencing across all five resistance-associated markers: EPSPS CNV, T102I, P106A, P106S and GS1-1 S59G, confirming the reliability and practical value of the multiplex qPCR assay. Only samples 7-8 showed discordance at EPSPS position 102, where Sanger chromatograms showed overlapping peaks at this position, which is likely to be a result of heterozygous mutation distribution among amplified EPSPS gene copies. This case further underscores the advantages of the multiplex qPCR assay over Sanger sequencing in detection sensitivity and accuracy. Moreover, a strong correlation (R2 = 0.8935) in gene copy number estimation between the two methods across all samples further supports the reliability of the qPCR assay. CONCLUSIONS: In summary, this study delivers a simple, robust and high-throughput diagnostic tool for the rapid, simultaneous identification of dual herbicide target-site resistance in goosegrass, offering superior sensitivity, quantitative resolution and throughput compared with Sanger sequencing. © 2026 Society of Chemical Industry.

Herbicides

Evolutionary expansion of the NF-Y gene family in bivalves and divergent subunit responses to thermal and pathogenic stress in the noble scallop.

Nuclear factor Y (NF-Y) is a conserved eukaryotic transcription factor complex that specifically interacts with the CCAAT motif. Prior research has demonstrated that this gene family participates in various biological processes, encompassing growth, development, and stress responses, across a broad spectrum of organisms. However, research on the role of the NF-Y family in bivalves remains limited. In this study, we comprehensively identified the NF-Y family in 34 bivalve species, and further investigated its expression in the noble scallop Chlamys nobilis. A total of 296 NF-Y genes were identified and classified into three subfamilies, NF-YA, NF-YB, and NF-YC. Phylogenetic analysis revealed that NF-YA and NF-YC have remained relatively conserved, whereas NF-YB has undergone significant expansion. Additionally, while substantial disparities in gene copy numbers exist across species, the motif composition and exon-intron structures within each subfamily demonstrate notable conservation. Tissue expression profiling revealed distinct expression patterns among CnNF-Y genes, with several members exhibiting relatively high transcript abundance in gonadal tissues. Furthermore, qRT-PCR results demonstrated that CnNF-YA2, CnNF-YB6, and CnNF-YC were significantly and continuously upregulated under heat stress. Conversely, several genes, particularly CnNF-YA2, CnNF-YB3, and CnNF-YB4, exhibited dynamic transcriptional responses to Vibrio parahaemolyticus exposure. These findings enhance our understanding of the evolutionary trajectory and functional diversification of the NF-Y gene family in bivalves, laying a theoretical foundation for future research on thermal adaptation, immune regulation, and molecular breeding in scallops.

Animals

Lambda phage promoter used to enhance expression of a plasmid-cloned gene.

A 50-fold (or greater) increase in the production of phage 21 repressor was obtained by construction of a plasmid in which the 21cI (repressor) gene could be transcribed from lambdaPL. The enhancement due to increased 21cI gene copy number and transcription from lambdaPL were at least five-fold and ten-fold, respectively. The plasmid was constructed in vitro by recombination of EcoRI-generated DNA fragments. The use of the DNA fragment containing lambdaPL in obtaining expression of cloned genes is discussed.

Coliphages

Evolutionary Diversification and Functions of the Candidate Male Killing Gene wmk.

Symbiont-mediated male killing (MK) is a mechanism that selectively eliminates male offspring, often by disrupting sex-specific developmental processes. In Drosophila melanogaster, the WO-mediated killing gene wmk from Wolbachia prophage WO transgenically reproduces the MK phenotype, yet how the gene evolves and functions across diverse Wolbachia has not been systematically investigated. We analyzed 32 Wolbachia genomes available in the NCBI database to study wmk homologs across different arthropod hosts, reproductive parasitism functions, and Wolbachia supergroups. First, we report at least five distinct wmk phylogenetic clusters (Types I to V), often organized in multigenic dyads or triads. Second, among MK Wolbachia, there is a significantly higher number of wmk genes and diversity in Lepidoptera strains than in Drosophila strains, which exclusively harbor wmk Types I and III. Third, there are three patterns of wmk sequence and genomic organizational changes in Drosophila MK strains that associate with different evolutionary trajectories underpinning the MK phenotype. Fourth, single and combinatory transgenic expression of Types I and III in D. melanogaster uncovers male-biased lethality associated with Type I; however, dual expression of the Types together elicits a major reduction in offspring number. Fifth, wmk genes have low expression level across D. melanogaster developmental stages relative to the cifA and cifB genes, which could explain why cytoplasmic incompatibility is expressed in this system. These findings establish a complex and phylogenetically informed genetic basis of wmk-induced lethality, highlighting the role of gene copy number and expression, wmk Types, and host background in shaping the phenotype.

Animals

Comparative prevalence of the mercury resistance gene merA in human feces, food, and environmental water from Japan, Vietnam, and Ghana.

In this study, we investigated the prevalence and abundance of the mercury resistance gene merA in human feces, retail chicken meat, and environmental water samples collected from Japan, Vietnam, and Ghana. A real-time PCR assay developed in this study demonstrated high specificity toward merA sequences from more than 12 bacterial species. Using this assay, merA was detected in 6.8% of human fecal samples in Japan (n = 29), in contrast to significantly higher rates observed in Vietnam (70.2%, n = 47) and Ghana (97.4%, n = 39). Similar geographic trends were evident in the chicken meat samples: 18.5% in Japan (n = 27), 66% in Vietnam (n = 91), and 90% in Ghana (n = 10). Environmental water samples showed a consistently high merA detection rate across all countries (75-100%, n = 21), with substantially higher gene copy numbers in Vietnam and Ghana than in Japan. merA was detected in some water samples, even when total mercury concentrations were below the detection limit, indicating that molecular detection may offer greater sensitivity than traditional physicochemical methods. Mercury-resistant bacteria were successfully isolated and cultured, and Citrobacter freundii was identified as the representative strain. Genomic analysis revealed that merA was located on an IncFIB plasmid, flanked by insertion sequences, suggesting its potential for horizontal gene transfer. These findings highlight merA as a promising biomarker for environmental mercury exposure and support the utility of fecal merA analysis as a proxy for assessing mercury-related public health risks.

Humans

Metabolomics and genomics reveal high diversity and concentrations of cyanopeptides during a Microcystis bloom.

Cyanobacterial blooms are an immense global problem that release complex mixtures of poorly characterized biologically active cyanopeptides into freshwater. In this study, metabolomics and genomics were used to assess the diversity and concentrations of cyanopeptides during a dense Microcystis bloom during the late summer of 2023 in Lake Champlain, a large transboundary lake situated between Canada and the United States. Despite the relatively low genetic diversity of the bloom determined by 16S rRNA metabarcoding, 151 cyanopeptides were detected by non-targeted metabolomics. This represents the most recorded cyanopeptides from a single lake plankton bloom event to date. Fifty-two cyanopeptides were previously reported and 99 represent putative new structures. Standards from the microcystin, cyanopeptolin, microginin, and anabaenopeptin groups were used to either quantify or approximate respective cyanopeptide concentrations over the sampling period. Cyanopeptolins were the most diverse (n = 68) cyanopeptides and the second most abundant, reaching 12,892 μg/L. Microginins were the second most diverse (n = 24) and reached the highest concentrations (18,262 μg/L). Anabaenopeptins were the third most diverse (n = 17) cyanopeptides, reaching 4,818 μg/L. Only 8 microcystins were detected, reaching 4,935 μg/L, where MC-LR was the dominant congener. Target cyanopeptide biosynthesis genes for microcystins (mcyE), cyanopeptolins (mcnC), anabaenopeptins (apnD), microviridins (mdnC), and aeruginosins (aerA) were also quantified using digital droplet PCR (ddPCR). The gene copy numbers for mcyE, mcnC, and apnD were highly correlated with their corresponding cyanopeptide concentrations. Overall, the studied Microcystis bloom produced a very diverse cyanopeptide mixture with high cyanopeptide concentrations including non-microcystin groups.

Microcystis

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Pervasive positive selection on X-linked ampliconic genes in primates.

Mammalian sex chromosomes harbour ampliconic gene families, which are multi-copy genes with ≥97% sequence identity, predominantly expressed in testis tissue and essential for male fertility. The amplification of testis-specific genes is conserved across mammals, yet the specific gene families that expand show striking lineage-specific variation. Previous studies suggest a dynamic turnover with adaptive evolution for several of these families, but their analysis has been limited by the quality of reference genomes of repetitive regions. To characterise the molecular evolutionary processes of ampliconic gene families on both sex chromosomes, we analysed telomere-to-telomere genome assemblies from eight primate species spanning 25 million years of evolution. We identified 53 X-linked and 19 Y-linked ampliconic gene families with dynamic copy number variation. Gene conversion through palindromic pairing and tandem arrays maintained high sequence similarity despite accumulating mutations. X-linked families maintained conserved chromosomal positions despite copy number changes, whereas Y-linked families showed frequent positional turnover. Strikingly, multiple X-linked families (GAGE, SSX, CSAG, and VCX) showed pervasive positive selection across the primate phylogeny and multiple (MAGEB, CT45, HSFX) showed lineage specific positive selection. Y-linked families predominantly evolve under purifying selection. Examining intraspecific copy number variation of the X-linked ampliconic families in chimpanzees, humans, and gorillas, we found variation among individuals but clear differences between species, with the largest families varying the most. These patterns could suggest that sperm competition, meiotic drive, or dosage-dependent selection drive the rapid, lineage-specific evolution of testis-expressed ampliconic genes in primates.

Journal Article

Synthesis of ribosomal proteins in merodiploid strains and in minicells of Escherichia coli.

Merodiploid strains of Escherichia coli containing episomes which carry one or several of the ribosomal protein (r-protein) transcriptional units were analysed to see whether the increase in the number of gene copies leads to an increased synthesis of the respective r-proteins. It was found that the amount of ribosomal proteins was (with the only exception of ribosomal protein S20) independent of the number of gene copies present. The comparison of the in vivo stability of r-proteins in haploid and merodiploid strains did not, within the time resolution of the experiment, provide any evidence for an increased rate of degradation of those proteins coded by more than one gene copy. These results indicate a tight coupling between the amount of ribosomal proteins synthesized and the level required irrespective of the number of gene copies present. With the aid of minicells from a strain containing the episome F'101 which carries the thr-leu segment of the chromosome it was demonstrated that (i) in vivo synthesis of r-protein S20 could proceed in the absence of the synthesis of ribosomal RNA and of other r-proteins, and (ii) r-protein S20 was degraded under conditions where it was not assembled into ribosomes.

Chromosomes, Bacterial

Graph-KIR: graph-based KIR copy number estimation and allele calling using short-read sequencing data.

MOTIVATION: The Killer-cell Immunoglobulin-like Receptor (KIR) is a highly polymorphic region in the human genome, associated with autoimmune diseases and organ transplantation. The sequences of KIR genes are highly similar among star alleles as well as in between individual genes, with the copy number of each KIR gene typically ranging from 0 to 4. In this study, we introduce Graph-KIR, a tool designed to estimate gene copy numbers and predict full-resolution (7-digit, encompassing both coding and non-coding sequence variations) from a whole genome sequencing (WGS) sample. RESULTS: Graph-KIR is capable of independently typing KIR alleles per sample with no reliance on the distribution of any framework gene in a cohort. In a set of 100 simulated samples, Graph-KIR demonstrated 99.2% accuracy in copy number estimation and high F1-score of allele typing: 91.79% at 7-digit resolution, 97.37% at 5-digit resolution, and 97.11% at 3-digit resolution. Graph-KIR outperforms existing tools such as Geny (96.39% F1-score), PING's WGS version (92.77% F1-score), and T1K (90.44% F1-score) at 5-digit resolution. By analyzing the results on 44 HPRC samples, Graph-KIR achieves better F1-score than Geny and PING at 7-digit resolution. The release of Graph-KIR adds another valuable tool to assist users in accurately estimating copy numbers and calling alleles of KIR genes from WGS samples. AVAILABILITY AND IMPLEMENTATION: The Graph-KIR and paper-related pipeline codes are available at https://github.com/linnil1/KIR_graph.

Receptors, KIR