PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genes, Duplicate”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Closely linked H2B genes in the marine copepod, Tigriopus californicus indicate a recent gene duplication or gene conversion event.

Two nonallelic histone gene clusters were characterized in the marine copepod, Tigriopus californicus. The DNA sequence of one of the clusters reveals six genes in the contiguous arrangement of H2B, H1, H3, H4, H2B and H2A. The order of genes within the second cluster is H3, H4, H2B and H2A. There is no evidence for the presence of an H1 gene in this cluster. Comparison of the three copepod H2B genes reveals a high degree of similarity between the 5' upstream regions and between the amino terminal halves of the two H2B genes found within the same cluster. From these data we infer that gene duplication and/or gene conversion events occurred within this cluster in the recent past.

Alleles↗

Not born equal: increased rate asymmetry in relocated and retrotransposed rodent gene duplicates.

Duplicated genes frequently evolve at different rates. This asymmetry is evidence of natural selection's ability to discriminate between the 2 copies, subjecting them to different levels of purifying selection or even permitting adaptive evolution of one or both copies. However, if gene duplication creates pairs of protein-coding sequences that are initially identical, this raises the question of how selection tells the 2 copies apart. Here, we investigated asymmetric sequence divergence of recently duplicated genes in rodents and related this to 2 possible sources of such asymmetry: gene relocation as a consequence of duplication and retrotransposition as a mechanism of gene duplication. We found that most young rodent duplicates that have been relocated were created by retrotransposition. The degree of rate asymmetry in gene pairs where one copy has been relocated (either by retrotransposition or DNA-based duplication) is greater than in pairs formed by local DNA-based duplication events. Furthermore, by considering the direction of transposition for distant duplicates, we found a consistent tendency for retrogenes to undergo accelerated protein evolution relative to their static paralogs, whereas DNA-based transpositions showed no such tendency. Finally, we demonstrate that the faster sequence evolution of retrogenes correlates with the profound alteration of their expression pattern that is precipitated by retrotransposition.

Animals↗

The evolutionary fate and consequences of duplicate genes.

Gene duplication has generally been viewed as a necessary source of material for the origin of evolutionary novelties, but it is unclear how often gene duplicates arise and how frequently they evolve new functions. Observations from the genomic databases for several eukaryotic species suggest that duplicate genes arise at a very high rate, on average 0.01 per gene per million years. Most duplicated genes experience a brief period of relaxed selection early in their history, with a moderate fraction of them evolving in an effectively neutral manner during this period. However, the vast majority of gene duplicates are silenced within a few million years, with the few survivors subsequently experiencing strong purifying selection. Although duplicate genes may only rarely evolve new functions, the stochastic silencing of such genes may play a significant role in the passive origin of new species.

Amino Acid Substitution↗

Subfunctionalization of duplicated genes as a transition state to neofunctionalization.

BACKGROUND: Gene duplication has been suggested to be an important process in the generation of evolutionary novelty. Neofunctionalization, as an adaptive process where one copy mutates into a function that was not present in the pre-duplication gene, is one mechanism that can lead to the retention of both copies. More recently, subfunctionalization, as a neutral process where the two copies partition the ancestral function, has been proposed as an alternative mechanism driving duplicate gene retention in organisms with small effective population sizes. The relative importance of these two processes is unclear. RESULTS: A set of lattice model genes that fold and bind to two peptide ligands with overlapping binding pockets, but not a third ligand present in the cell was designed. Each gene was duplicated in a model haploid species with a small constant population size and no recombination. One set of models allowed subfunctionalization of binding events following duplication, while another set did not allow subfunctionalization. Modeling under such conditions suggests that subfunctionalization plays an important role, but as a transition state to neofunctionalization rather than as a terminal fate of duplicated genes. There is no apparent selective pressure to maintain redundancy. CONCLUSION: Subfunctionalization results in an increase in the preservation of duplicated gene copies, including those that are neofunctionalized, but never represents a substantial fraction of duplicate gene copies at any evolutionary time point and ultimately leads to neofunctionalization of those preserved copies. This conclusion also may reflect changes in gene function after duplication with time in real genomes.

Animals↗

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software↗

The early stages of duplicate gene evolution.

Gene duplications are one of the primary driving forces in the evolution of genomes and genetic systems. Gene duplicates account for 8-20% of the genes in eukaryotic genomes, and the rates of gene duplication are estimated at between 0.2% and 2% per gene per million years. Duplicate genes are believed to be a major mechanism for the establishment of new gene functions and the generation of evolutionary novelty, yet very little is known about the early stages of the evolution of duplicated gene pairs. It is unclear, for example, to what extent selection, rather than neutral genetic drift, drives the fixation and early evolution of duplicate loci. Analysis of recently duplicated genes in the Arabidopsis thaliana genome reveals significantly reduced species-wide levels of nucleotide polymorphisms in the progenitor and/or duplicate gene copies, suggesting that selective sweeps accompany the initial stages of the evolution of these duplicated gene pairs. Our results support recent theoretical work that indicates that fates of duplicate gene pairs may be determined in the initial phases of duplicate gene evolution and that positive selection plays a prominent role in the evolutionary dynamics of the very early histories of duplicate nuclear genes.

Arabidopsis↗

Rapid subfunctionalization accompanied by prolonged and substantial neofunctionalization in duplicate gene evolution.

Gene duplication is the primary source of new genes. Duplicate genes that are stably preserved in genomes usually have divergent functions. The general rules governing the functional divergence, however, are not well understood and are controversial. The neofunctionalization (NF) hypothesis asserts that after duplication one daughter gene retains the ancestral function while the other acquires new functions. In contrast, the subfunctionalization (SF) hypothesis argues that duplicate genes experience degenerate mutations that reduce their joint levels and patterns of activity to that of the single ancestral gene. We here show that neither NF nor SF alone adequately explains the genome-wide patterns of yeast protein interaction and human gene expression for duplicate genes. Instead, our analysis reveals rapid SF, accompanied by prolonged and substantial NF in a large proportion of duplicate genes, suggesting a new model termed subneofunctionalization (SNF). Our results demonstrate that enormous numbers of new functions have originated via gene duplication.

Amino Acid Sequence↗

Rapid evolution of expression and regulatory divergences after yeast gene duplication.

Although gene duplication is widely believed to be the major source of genetic novelty, how the expression or regulatory network of duplicate genes evolves remains poorly understood. In this article, we propose an additive expression distance between duplicate genes, so that the evolutionary rate of expression divergence after gene duplication can be estimated through phylogenomic analysis. We have analyzed yeast genome sequences, microarrays, and transcriptional regulatory networks, showing a >10-fold increase in the initial rate for both expression and regulatory network evolution after gene duplication but only an approximately 20% rate increase in the early stage for protein sequences. Based on the estimated age distribution of yeast duplicate genes, we roughly estimate that the initial rate of expression divergence shortly after gene duplication is 2.9 x 10(-9) per year, whereas the baseline rate for very ancient gene duplication is 0.14 x 10(-9) per year. Relative expression rate tests suggest that the expression of duplicate genes tends to evolve asymmetrically, that is, the expression of one copy evolves rapidly, whereas the other one largely maintains the ancestral expression profile. Our study highlights the crucial role of early rapid evolution after gene/genome duplication for continuously increasing the complexity of the yeast regulatory network.

Evolution, Molecular↗

The porcine taurochenodeoxycholic acid 6alpha-hydroxylase (CYP4A21) gene: evolution by gene duplication and gene conversion.

Porcine taurochenodeoxycholic acid 6alpha-hydroxylase, cytochrome P450 4A21 (CYP4A21), differs from other members of the CYP4A subfamily in terms of structural features and catalytic activity. CYP4A21 participates in the formation of hyocholic acid, a species-specific primary bile acid in the pig. The CYP4A21 gene was investigated and found to be approx. 13 kb in size and split into 12 exons. The intron-exon organization of the CYP4A21 gene corresponds to that of CYP4A fatty acid hydroxylase genes in other species. Comparison with a genomic segment of a pig CYP4A fatty acid hydroxylase gene ( CYP4A24 ) revealed a sequence identity with CYP4A21 that extends beyond the exons, indicating a common origin by gene duplication. A pronounced sequence identity was found also within the proximal 5'-flanking regions, whereas the patterns of mRNA expression of CYP4A21 and CYP4A fatty acid hydroxylases in pig liver differ. Sequence comparison aiming to elucidate the origin of the unique features of CYP4A21 revealed a region of decreased sequence identity from exon 6 to exon 8, strongly suggesting that gene conversion could have contributed to the evolution of CYP4A21.

5' Flanking Region↗

The evolutionary demography of duplicate genes.

Although gene duplication has generally been viewed as a necessary source of material for the origin of evolutionary novelties, the rates of origin, loss, and preservation of gene duplicates are not well understood. Applying steady-state demographic techniques to the age distributions of duplicate genes censused in seven completely sequenced genomes, we estimate the average rate of duplication of a eukaryotic gene to be on the order of 0.01/ gene/million years, which is of the same order of magnitude as the mutation rate per nucleotide site. However, the average half-life of duplicate genes is relatively small, on the order of 4.0 million years. Significant interspecific variation in these rates appears to be responsible for differences in species-specific genome sizes that arise as a consequence of a quasi-equilibrium birth-death process. Most duplicated genes experience a brief period of relaxed selection early in their history and a minority exhibit the signature of directional selection, but those that survive more than a few million years eventually experience strong purifying selection. Thus, although most theoretical work on the gene-duplication process has focused on issues related to adaptive evolution, the origin of a new function appears to be a very rare fate for a duplicate gene. A more significant role of the duplication process may be the generation of microchromosomal rearrangements through reciprocal silencing of alternative copies, which can lead to the passive origin of post-zygotic reproductive barriers in descendant lineages of incipient species.

Animals↗

The major shifts of human duplicated genes.

Since many gene duplications in the human genome are ancient duplications going back to the origin of vertebrates, the question may be asked about the fate of such duplicated genes at the compositional genome transitions that occurred between cold- and warm-blooded vertebrates. Indeed, at that transition, about half of the (GC-poor) genes of cold-blooded vertebrates (the genes of the gene-dense "ancestral genome core") underwent a GC enrichment to become the genes of the "genome core" of warm-blooded vertebrates. Since the compositional distribution of the human duplicated genes investigated (1111 pairs) mimics the general distribution of human genes (about 50% GC(3)-poor and 50% GC(3)-rich genes, the border being at 60% GC(3)), we considered two possibilities, namely that the compositional transition affected either (i) about half of the copies on a random basis, or (ii) preferentially only one copy of the duplicated genes. The two possibilities could be distinguished if each copy is put into one of two subsets according to its GC(3) level. Indeed, in the first case, the two distributions would be similar, whereas in the second case, the two distributions would be different, one copy having maintained the ancestral GC-poor composition, and one copy having undergone the compositional change. Using this approach, we could show that, by far and large, one copy of the duplicated genes preferentially underwent the GC enrichment. This result implies that this copy, which had possibly acquired a different function and/or regulation, was preferentially translocated into the gene-dense compartment of the genome, the "ancestral genome core", namely the "gene space" which underwent the compositional transition at the emergence of warm-blooded vertebrates.

Adaptation, Physiological↗

Different evolutionary patterns between young duplicate genes in the human genome.

BACKGROUND: Following gene duplication, two duplicate genes may experience relaxed functional constraints or acquire different mutations, and may also diverge in function. Whether the two copies will evolve in different patterns remains unclear, however, because previous studies have reached conflicting conclusions. In order to resolve this issue, by providing a general picture, we studied 250 independent pairs of young duplicate genes from the whole human genome. RESULTS: We showed that nearly 60% of the young duplicate gene pairs have evolved at the amino-acid level at significantly different rates from each other. More than 25% of these gene pairs also showed significantly different ratios of nonsynonymous to synonymous rates (Ka/Ks ratios). Moreover, duplicate pairs with different rates of amino-acid substitution also tend to differ in the Ka/Ks ratio, with the fast-evolving copy tending to have a slightly higher Ks than the slow-evolving one. Lastly, a substantial portion of fast-evolving copies have accumulated amino-acid substitutions evenly across the protein sequences, whereas most of the slow-evolving copies exhibit uneven substitution patterns. CONCLUSIONS: Our results suggest that duplicate genes tend to evolve in different patterns following the duplication event. One copy evolves faster than the other and accumulates amino-acid substitutions evenly across the sequence, whereas the other copy evolves more slowly and accumulates amino-acid substitutions unevenly across the sequence. Such different evolutionary patterns may be largely due to different functional constraints on the two copies.

Amino Acid Substitution↗

Rapid and asymmetric divergence of duplicate genes in the human gene coexpression network.

BACKGROUND: While gene duplication is known to be one of the most common mechanisms of genome evolution, the fates of genes after duplication are still being debated. In particular, it is presently unknown whether most duplicate genes preserve (or subdivide) the functions of the parental gene or acquire new functions. One aspect of gene function, that is the expression profile in gene coexpression network, has been largely unexplored for duplicate genes. RESULTS: Here we build a human gene coexpression network using human tissue-specific microarray data and investigate the divergence of duplicate genes in it. The topology of this network is scale-free. Interestingly, our analysis indicates that duplicate genes rapidly lose shared coexpressed partners: after approximately 50 million years since duplication, the two duplicate genes in a pair have only slightly higher number of shared partners as compared with two random singletons. We also show that duplicate gene pairs quickly acquire new coexpressed partners: the average number of partners for a duplicate gene pair is significantly greater than that for a singleton (the latter number can be used as a proxy of the number of partners for a parental singleton gene before duplication). The divergence in gene expression between two duplicates in a pair occurs asymmetrically: one gene usually has more partners than the other one. The network is resilient to both random and degree-based in silico removal of either singletons or duplicate genes. In contrast, the network is especially vulnerable to the removal of highly connected genes when duplicate genes and singletons are considered together. CONCLUSION: Duplicate genes rapidly diverge in their expression profiles in the network and play similar role in maintaining the network robustness as compared with singletons.

Chromosome Mapping↗

Gene duplications in 21-hydroxylase deficiency: the importance of accurate molecular diagnosis in carrier detection and prenatal diagnosis.

BACKGROUND: The detection of 21-OH deficiency (21OHD) carriers in the general population requires that misinterpretations of apparently severe mutations in alleles carrying duplicated genes be avoided. Prenatal treatment prevents virilization in female fetuses and genetic counseling may be offered to couples in which one partner is either a patient or a carrier. This paper proposes a semiquantitative PCR method involving primer extension that distinguishes the severe point mutation Q318X in single gene copy alleles from the normal/nondeficient variant in gene-duplicated alleles. SAMPLES AND METHODS: DNA from 65 individuals carrying Q318X variants, that of 85 partners of 21OHD carriers or patients, and one fetal sample (as well as the DNA of his family) were analyzed. 21OHD alleles were studied by gene-specific PCR/allele-specific oligonucleotides hybridization for common mutations, Southern analysis, complementary direct sequencing and microsatellite typing. Primer extension analysis of the Q318X variants using fluorescent dideoxynucleotides was performed on CYP21A2 gene-specific PCR-amplified DNA samples from controls, patients, potential carriers and prenatal samples. RESULTS: Different fluorescence patterns were seen for the severe mutation (single gene copy) and the nondeficient (gene-duplicated) alleles carrying Q318X. The normal/mutant fluorescence peak (N/M) ratio was < 1 in all heterozygous carriers (mean 0.83; min. 0.70; max. 0.95). In all normal individuals carrying the gene-duplicated Q318X normal variant, the N/M ratio was > 1 (mean 1.69; min. 1.44; max. 2.02). CONCLUSION: The proposed method discriminated between the severe Q318X mutation and the normal Q318X variant in gene duplication, and could be a useful complementary tool in prenatal diagnosis and carrier detection.

Adrenal Hyperplasia, Congenital↗

Sharing of transcription factors after gene duplication in the yeast Saccharomyces cerevisiae.

In a set of 190 duplicate gene pairs in yeast Saccharomyces cerevisiae, the sharing of transcription factors tended to decrease with increased divergence in coding sequence, at both synonymous and nonsynonymous sites. Our results showed a significantly higher sharing of transcription factors by duplicated gene pairs falling within duplicated genomic blocks than in other duplicated gene pairs; and genes in duplicated blocks also showed significantly greater conservation at the coding sequence level. In spite of the overall trends, there were certain gene pairs, both in duplicated blocks and in other genomic regions, which were highly divergent in coding sequence and yet had identical patterns of transcription factor binding. These results suggest that functional differentiation of genes after duplication is a multi-dimensional process, with different duplicate pairs differentiating in different ways.

Base Sequence↗

Duplicated genes evolve slower than singletons despite the initial rate increase.

BACKGROUND: Gene duplication is an important mechanism that can lead to the emergence of new functions during evolution. The impact of duplication on the mode of gene evolution has been the subject of several theoretical and empirical comparative-genomic studies. It has been shown that, shortly after the duplication, genes seem to experience a considerable relaxation of purifying selection. RESULTS: Here we demonstrate two opposite effects of gene duplication on evolutionary rates. Sequence comparisons between paralogs show that, in accord with previous observations, a substantial acceleration in the evolution of paralogs occurs after duplication, presumably due to relaxation of purifying selection. The effect of gene duplication on evolutionary rate was also assessed by sequence comparison between orthologs that have paralogs (duplicates) and those that do not (singletons). It is shown that, in eukaryotes, duplicates, on average, evolve significantly slower than singletons. Eukaryotic ortholog evolutionary rates for duplicates are also negatively correlated with the number of paralogs per gene and the strength of selection between paralogs. A tally of annotated gene functions shows that duplicates tend to be enriched for proteins with known functions, particularly those involved in signaling and related cellular processes; by contrast, singletons include an over-abundance of poorly characterized proteins. CONCLUSIONS: These results suggest that whether or not a gene duplicate is retained by selection depends critically on the pre-existing functional utility of the protein encoded by the ancestral singleton. Duplicates of genes of a higher biological import, which are subject to strong functional constraints on the sequence, are retained relatively more often. Thus, the evolutionary trajectory of duplicated genes appears to be determined by two opposing trends, namely, the post-duplication rate acceleration and the generally slow evolutionary rate owing to the high level of functional constraints.

Animals↗

The 2.15 A crystal structure of Mycobacterium tuberculosis chorismate mutase reveals an unexpected gene duplication and suggests a role in host-pathogen interactions.

Chorismate mutase catalyzes the first committed step toward the biosynthesis of the aromatic amino acids, phenylalanine and tyrosine. While this biosynthetic pathway exists exclusively in the cell cytoplasm, the Mycobacterium tuberculosis enzyme has been shown to be secreted into the extracellular medium. The secretory nature of the enzyme and its existence in M. tuberculosis as a duplicated gene are suggestive of its role in host-pathogen interactions. We report here the crystal structure of homodimeric chorismate mutase (Rv1885c) from M. tuberculosis determined at 2.15 A resolution. The structure suggests possible gene duplication within each subunit of the dimer (residues 35-119 and 130-199) and reveals an interesting proline-rich region on the protein surface (residues 119-130), which might act as a recognition site for protein-protein interactions. The structure also offers an explanation for its regulation by small ligands, such as tryptophan, a feature previously unknown in the prototypical Escherichia coli chorismate mutase. The tryptophan ligand is found to be sandwiched between the two monomers in a dimer contacting residues 66-68. The active site in the "gene-duplicated" monomer is occupied by a sulfate ion and is located in the first half of the polypeptide, unlike in the Saccharomyces cerevisiae (yeast) enzyme, where it is located in the later half. We hypothesize that the M. tuberculosis chorismate mutase might have a role to play in host-pathogen interactions, making it an important target for designing inhibitor molecules against the deadly pathogen.

Amino Acid Sequence↗

Duplicated genes evolve independently after polyploid formation in cotton.

Of the many processes that generate gene duplications, polyploidy is unique in that entire genomes are duplicated. This process has been important in the evolution of many eukaryotic groups, and it occurs with high frequency in plants. Recent evidence suggests that polyploidization may be accompanied by rapid genomic changes, but the evolutionary fate of discrete loci recently doubled by polyploidy (homoeologues) has not been studied. Here we use locus-specific isolation techniques with comparative mapping to characterize the evolution of homoeologous loci in allopolyploid cotton (Gossypium hirsutum) and in species representing its diploid progenitors. We isolated and sequenced 16 loci from both genomes of the allopolyploid, from both progenitor diploid genomes and appropriate outgroups. Phylogenetic analysis of the resulting 73.5 kb of sequence data demonstrated that for all 16 loci (14.7 kb/genome), the topology expected from organismal history was recovered. In contrast to observations involving repetitive DNAs in cotton, there was no evidence of interaction among duplicated genes in the allopolyploid. Polyploidy was not accompanied by an obvious increase in mutations indicative of pseudogene formation. Additionally, differences in rates of divergence among homoeologues in polyploids and orthologues in diploids were indistinguishable across loci, with significant rate deviation restricted to two putative pseudogenes. Our results indicate that most duplicated genes in allopolyploid cotton evolve independently of each other and at the same rate as those of their diploid progenitors. These indications of genic stasis accompanying polyploidization provide a sharp contrast to recent examples of rapid genomic evolution in allopolyploids.

Biological Evolution↗