PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genes, Duplicate”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Structural insight into gene duplication, gene fusion and domain swapping in the evolution of PLP-independent amino acid racemases.

The X-ray crystal structure has revealed two similar alpha/beta domains of aspartate racemase (AspR) from Pyrococcus horikoshii OT3, and identified a pseudo mirror-symmetric distribution of the residues around its active site [Liu et al. (2002) J. Mol. Biol. 319, 479-489]. Structural homology and functional similarity between the two domains suggested that this enzyme evolved from an ancestral domain by gene duplication and gene fusion. We have expressed solely the C-terminal domain of this AspR and determined its three-dimensional structure by X-ray crystallography. The high structural stability of this domain supports the existence of the ancestral domain. In comparison with other amino acid racemases (AARs), we suggest that gene duplication and gene fusion are conventional ways in the evolution of pyridoxal 5'-phosphate-independent AARs.

Amino Acid Isomerases↗

Structure and genetics of the partially duplicated gene RP located immediately upstream of the complement C4A and the C4B genes in the HLA class III region. Molecular cloning, exon-intron structure, composite retroposon, and breakpoint of gene duplication.

The correlation of many HLA-associated autoimmune and genetic diseases with the polymorphic complement C4 genes may be attributed to the presence of disease susceptibility genes in the close proximity of C4. We have cloned and characterized a pair of partially duplicated genes, RP1 and RP2, located 611 base pairs upstream of the human C4A and C4B genes, respectively. The putative RP protein, consisting of 364 amino acid residues, is basic and highly hydrophilic. There is a bipartite nuclear localization signal at residues 114-131 and therefore RP may be a nuclear protein. Northern blot analysis suggested that RP is ubiquitously expressed. The 5' region of the RP1 gene is CpG rich, which is a characteristic of housekeeping genes. The RP1 gene contains nine exons. Located in the fourth intron is a cluster of Alu elements, and a newly defined composite retroposon SVA with a SINE, multiple copies of GC-rich VNTRs and an Alu element altogether enclosed by direct terminal repeats. Members of SVA are also present in the complement C2 gene located about 20 kilobases upstream of RP1 in the HLA and in the cytochrome CYP1A1 gene. Determination of the DNA sequences for RP2 from two different HLA haplotypes revealed identical hybrid sequences which resulted from fusion of RP with the tenascin-like Gene X and truncation of the 5' regions of both genes. Cumulative data suggest that the four tandemly arranged genes RP, complement C4, steroid 21-hydroxylase (CYP21), and Gene X altogether form a modular structure, RCCX. The number of RCCX modules varies from one to three or more in the population. Absence of the truncated genes RP2 and Gene XA have been detected in genomes with single RCCX modules. Duplication of the RCCX modules probably occurred before the speciation of great apes and humans as they contain the same breakpoint region of RP and Gene X gene duplication.

Alleles↗

Molecular cloning and chromosomal localization of the mouse decay-accelerating factor genes. Duplicated genes encode glycosylphosphatidylinositol-anchored and transmembrane forms.

Regulation of complement activation is essential in the prevention of damage to autologous tissue. This activity is mediated by the presence of specific complement regulatory proteins on the surface of host cells. In humans, one molecule involved in this regulation is a 70-kDa glycoprotein that has been designated decay-accelerating factor (DAF). We present the full-length cDNA sequence and chromosomal localization of the mouse genetic homologue of the human DAF gene. Interestingly, two classes of cDNA clones were obtained that, rather than representing alternately spliced mRNAs, were derived from two separate but closely related linked genes. Both genes encoded proteins with an amino-terminal signal sequence, followed by four short consensus repeats and a domain rich in serine and threonine. Hydrophilicity plots and alignment with human DAF predicted that one gene encoded a glycosylphosphatidylinositol-anchored form of mouse Daf with 64% nucleotide and 47% amino acid identity to human DAF. The product encoded by the second gene was predicted to have an alternate amino-terminal signal sequence and carboxyl-terminal membrane-spanning and cytoplasmic domains. The two mouse Daf genes share 85% nucleotide and 78% amino acid identities, and have been designated Daf-glycosylphosphatidylinositol and Daf-transmembrane to reflect the two alternate mechanisms of membrane attachment. mRNA expression analysis indicated that the two mouse Daf genes were differentially expressed in the adult mouse. Chromosome localization studies mapped the mouse Daf genes to chromosome 1, where they segregated with the C4-binding protein (C4bp) gene.

Amino Acid Sequence↗

Epigenetic silencing may aid evolution by gene duplication.

Gene duplication is commonly regarded as the main evolutionary path toward the gain of a new function. However, even with gene duplication, there is a loss-versus-gain dilemma: most newly born duplicates degrade to pseudogenes, since degenerative mutations are much more frequent than advantageous ones. Thus, something additional seems to be needed to shift the loss versus gain equilibrium toward functional divergence. We suggest that epigenetic silencing of duplicates might play this role in evolution. This study began when we noticed in a previous publication (Lynch M, Conery JS [2000] Science 291:1151-1155) that the frequency of functional young gene duplicates is higher in organisms that have cytosine methylation (H. sapiens, M. musculus, and A. thaliana) than in organisms that do not have methylated genomes (S. cerevisiae, D. melanogaster, and C. elegans). We find that genome data analysis confirms the likelihood of much more efficient functional divergence of gene duplicates in mammals and plants than in yeast, nematode, and fly. We have also extended the classic model of gene duplication, in which newly duplicated genes have exactly the same expression pattern, to the case when they are epigenetically silenced in a tissue- and/or developmental stage-complementary manner. This exposes each of the duplicates to negative selection, thus protecting from "pseudogenization." Our analysis indicates that this kind of silencing (i) enhances evolution of duplicated genes to new functions, particularly in small populations, (ii) is quite consistent with the subfunctionalization model when degenerative but complementary mutations affect different subfunctions of the gene, and (iii) furthermore, may actually cooperate with the DDC (duplication-degeneration-complementation) process.

Animals↗

The evolution of gene duplicates.

Gene and genome duplications have given rise to enormous variability among species in the number of genes within their genomes. Gene copies have in turn played important roles in adaptation, having been implicated in the evolution of the immune response, insecticide resistance, efficient protein synthesis, and vertebrate body plans. In this chapter, we discuss the life history of gene duplications, from their first appearance within a population, through the period during which they rise in frequency or disappear, to their long-term fate. At each phase, we discuss the evolutionary processes that have influenced the dynamics of gene duplications and shaped their ultimate roles within a population. We argue that there is no evidence that organisms have evolved strategies to promote gene duplication in order to permit adaptive evolution. In contrast, many mechanisms exist to silence or eliminate duplicated genes, suggesting that selection has acted largely to reduce the rate of gene duplication. We also argue that natural selection has functioned as an effective sieve, increasing the representation of beneficial gene duplicates among those that establish within a population and that play a long-term role in evolution. To refine our understanding of how selection acts on new gene duplications, we provide a model incorporating a single-copy gene, its gene duplicate, and selection either favoring heterozygotes or eliminating deleterious mutations. Although both forms of selection can increase the initial rate of spread of a gene duplicate, the efficacy with which they do so differs dramatically. Heterozygote advantage always increases the rate of spread and can have a large impact. In contrast, masking deleterious mutations never has a large effect on the rate of spread of the duplicate, and this minor effect can be negative as well as positive. In both cases, the degree of linkage between the two gene copies affects the rate of spread of the duplication. Finally, we discuss evolutionary processes that occur over longer periods after a gene duplication has become established within a population. These long-term processes include maintenance, inactivation, and diversification in function. Consideration of each of the short-term and long-term processes affecting duplicated genes illustrates the subtle ways in which selection has acted to shape genomic structure.

Animals↗

Gene duplication and gene conversion shape the evolution of archaeal chaperonins.

Chaperonins are multi-subunit double-ring complexes that mediate the folding of nascent or denatured proteins. Gene duplication has been a potent force in the evolution of chaperonins in Archaea. Here we show that gene conversion has also been an important factor. We utilized a novel maximum likehood-based phylogenetic method for scanning DNA sequence alignments for regions of anomalous phylogenetic signal, such as those affected by gene conversion. Our results suggest that in crenarchaeotes, where an ancient gene duplication producing alpha and beta subunits took place in the common ancestor of the Pyrodictium, Aeropyrum, Pyrobaculum and Sulfolobus lineages, multiple independent gene conversions have occurred between the alpha and beta genes independently in each of these groups. Significantly, the conversions have repeatedly homogenized the region of the gene encoding the substrate-binding domain. This suggests that while the alpha and beta subunits in crenarchaeotes share only 50-60% overall amino acid sequence identity, they do not possess distinct roles in the binding of substrate. Cryptic gene conversion between distantly related paralogs may be more common than is currently appreciated, and could be a significant factor in slowing the functional differentiation of proteins encoded by duplicate genes long after their duplication.

Amino Acid Sequence↗

Gene duplication and gene conversion in the Caenorhabditis elegans genome.

A comprehensive analysis of duplication and gene conversion for 7394 Caenorhabditis elegans genes (about half the expected total for the genome) is presented. Of the genes examined, 40% are involved in duplicated gene pairs. Intrachromosomal or cis gene duplications occur approximately two times more often than expected. In general the closer the members of duplicated gene pairs are, the more likely it is that gene orientation is conserved. Gene conversion events are detectable between only 2% of the duplicated pairs. Even given the excesses of cis duplications, there is an excess of gene conversion events between cis duplicated pairs on every chromosome except the X chromosome. The relative rates of cis and trans gene conversion and the negative correlation between conversion frequency and DNA sequence divergence for unconverted regions of converted pairs are consistent with previous experimental studies in yeast. Three recent, regional duplications, each spanning three genes are described. All three have already undergone substantial deletions spanning hundreds of base pairs. The relative rates of duplication and deletion may contribute to the compactness of the C. elegans genome.

Animals↗

The probability of duplicate gene preservation by subfunctionalization.

It has often been argued that gene-duplication events are most commonly followed by a mutational event that silences one member of the pair, while on rare occasions both members of the pair are preserved as one acquires a mutation with a beneficial function and the other retains the original function. However, empirical evidence from genome duplication events suggests that gene duplicates are preserved in genomes far more commonly and for periods far in excess of the expectations under this model, and whereas some gene duplicates clearly evolve new functions, there is little evidence that this is the most common mechanism of duplicate-gene preservation. An alternative hypothesis is that gene duplicates are frequently preserved by subfunctionalization, whereby both members of a pair experience degenerative mutations that reduce their joint levels and patterns of activity to that of the single ancestral gene. We consider the ways in which the probability of duplicate-gene preservation by such complementary mutations is modified by aspects of gene structure, degree of linkage, mutation rates and effects, and population size. Even if most mutations cause complete loss-of-subfunction, the probability of duplicate-gene preservation can be appreciable if the long-term effective population size is on the order of 10(5) or smaller, especially if there are more than two independently mutable subfunctions per locus. Even a moderate incidence of partial loss-of-function mutations greatly elevates the probability of preservation. The model proposed herein leads to quantitative predictions that are consistent with observations on the frequency of long-term duplicate gene preservation and with observations that indicate that a common fate of the members of duplicate-gene pairs is the partitioning of tissue-specific patterns of expression of the ancestral gene.

Biological Evolution↗

Preservation of duplicate genes by complementary, degenerative mutations.

The origin of organismal complexity is generally thought to be tightly coupled to the evolution of new gene functions arising subsequent to gene duplication. Under the classical model for the evolution of duplicate genes, one member of the duplicated pair usually degenerates within a few million years by accumulating deleterious mutations, while the other duplicate retains the original function. This model further predicts that on rare occasions, one duplicate may acquire a new adaptive function, resulting in the preservation of both members of the pair, one with the new function and the other retaining the old. However, empirical data suggest that a much greater proportion of gene duplicates is preserved than predicted by the classical model. Here we present a new conceptual framework for understanding the evolution of duplicate genes that may help explain this conundrum. Focusing on the regulatory complexity of eukaryotic genes, we show how complementary degenerative mutations in different regulatory elements of duplicated genes can facilitate the preservation of both duplicates, thereby increasing long-term opportunities for the evolution of new gene functions. The duplication-degeneration-complementation (DDC) model predicts that (1) degenerative mutations in regulatory elements can increase rather than reduce the probability of duplicate gene preservation and (2) the usual mechanism of duplicate gene preservation is the partitioning of ancestral functions rather than the evolution of new functions. We present several examples (including analysis of a new engrailed gene in zebrafish) that appear to be consistent with the DDC model, and we suggest several analytical and experimental approaches for determining whether the complementary loss of gene subfunctions or the acquisition of novel functions are likely to be the primary mechanisms for the preservation of gene duplicates. For a newly duplicated paralog, survival depends on the outcome of the race between entropic decay and chance acquisition of an advantageous regulatory mutation. Sidow 1996(p. 717) On one hand, it may fix an advantageous allele giving it a slightly different, and selectable, function from its original copy. This initial fixation provides substantial protection against future fixation of null mutations, allowing additional mutations to accumulate that refine functional differentiation. Alternatively, a duplicate locus can instead first fix a null allele, becoming a pseudogene. Walsh 1995 (p. 426) Duplicated genes persist only if mutations create new and essential protein functions, an event that is predicted to occur rarely. Nadeau and Sankoff 1997 (p. 1259) Thus overall, with complex metazoans, the major mechanism for retention of ancient gene duplicates would appear to have been the acquisition of novel expression sites for developmental genes, with its accompanying opportunity for new gene roles underlying the progressive extension of development itself. Cooke et al. 1997 (p. 362)

Animals↗

Closely linked H2B genes in the marine copepod, Tigriopus californicus indicate a recent gene duplication or gene conversion event.

Two nonallelic histone gene clusters were characterized in the marine copepod, Tigriopus californicus. The DNA sequence of one of the clusters reveals six genes in the contiguous arrangement of H2B, H1, H3, H4, H2B and H2A. The order of genes within the second cluster is H3, H4, H2B and H2A. There is no evidence for the presence of an H1 gene in this cluster. Comparison of the three copepod H2B genes reveals a high degree of similarity between the 5' upstream regions and between the amino terminal halves of the two H2B genes found within the same cluster. From these data we infer that gene duplication and/or gene conversion events occurred within this cluster in the recent past.

Alleles↗

The evolutionary fate and consequences of duplicate genes.

Gene duplication has generally been viewed as a necessary source of material for the origin of evolutionary novelties, but it is unclear how often gene duplicates arise and how frequently they evolve new functions. Observations from the genomic databases for several eukaryotic species suggest that duplicate genes arise at a very high rate, on average 0.01 per gene per million years. Most duplicated genes experience a brief period of relaxed selection early in their history, with a moderate fraction of them evolving in an effectively neutral manner during this period. However, the vast majority of gene duplicates are silenced within a few million years, with the few survivors subsequently experiencing strong purifying selection. Although duplicate genes may only rarely evolve new functions, the stochastic silencing of such genes may play a significant role in the passive origin of new species.

Amino Acid Substitution↗

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software↗

The evolutionary demography of duplicate genes.

Although gene duplication has generally been viewed as a necessary source of material for the origin of evolutionary novelties, the rates of origin, loss, and preservation of gene duplicates are not well understood. Applying steady-state demographic techniques to the age distributions of duplicate genes censused in seven completely sequenced genomes, we estimate the average rate of duplication of a eukaryotic gene to be on the order of 0.01/ gene/million years, which is of the same order of magnitude as the mutation rate per nucleotide site. However, the average half-life of duplicate genes is relatively small, on the order of 4.0 million years. Significant interspecific variation in these rates appears to be responsible for differences in species-specific genome sizes that arise as a consequence of a quasi-equilibrium birth-death process. Most duplicated genes experience a brief period of relaxed selection early in their history and a minority exhibit the signature of directional selection, but those that survive more than a few million years eventually experience strong purifying selection. Thus, although most theoretical work on the gene-duplication process has focused on issues related to adaptive evolution, the origin of a new function appears to be a very rare fate for a duplicate gene. A more significant role of the duplication process may be the generation of microchromosomal rearrangements through reciprocal silencing of alternative copies, which can lead to the passive origin of post-zygotic reproductive barriers in descendant lineages of incipient species.

Animals↗

Different evolutionary patterns between young duplicate genes in the human genome.

BACKGROUND: Following gene duplication, two duplicate genes may experience relaxed functional constraints or acquire different mutations, and may also diverge in function. Whether the two copies will evolve in different patterns remains unclear, however, because previous studies have reached conflicting conclusions. In order to resolve this issue, by providing a general picture, we studied 250 independent pairs of young duplicate genes from the whole human genome. RESULTS: We showed that nearly 60% of the young duplicate gene pairs have evolved at the amino-acid level at significantly different rates from each other. More than 25% of these gene pairs also showed significantly different ratios of nonsynonymous to synonymous rates (Ka/Ks ratios). Moreover, duplicate pairs with different rates of amino-acid substitution also tend to differ in the Ka/Ks ratio, with the fast-evolving copy tending to have a slightly higher Ks than the slow-evolving one. Lastly, a substantial portion of fast-evolving copies have accumulated amino-acid substitutions evenly across the protein sequences, whereas most of the slow-evolving copies exhibit uneven substitution patterns. CONCLUSIONS: Our results suggest that duplicate genes tend to evolve in different patterns following the duplication event. One copy evolves faster than the other and accumulates amino-acid substitutions evenly across the sequence, whereas the other copy evolves more slowly and accumulates amino-acid substitutions unevenly across the sequence. Such different evolutionary patterns may be largely due to different functional constraints on the two copies.

Amino Acid Substitution↗

Duplicated genes evolve independently after polyploid formation in cotton.

Of the many processes that generate gene duplications, polyploidy is unique in that entire genomes are duplicated. This process has been important in the evolution of many eukaryotic groups, and it occurs with high frequency in plants. Recent evidence suggests that polyploidization may be accompanied by rapid genomic changes, but the evolutionary fate of discrete loci recently doubled by polyploidy (homoeologues) has not been studied. Here we use locus-specific isolation techniques with comparative mapping to characterize the evolution of homoeologous loci in allopolyploid cotton (Gossypium hirsutum) and in species representing its diploid progenitors. We isolated and sequenced 16 loci from both genomes of the allopolyploid, from both progenitor diploid genomes and appropriate outgroups. Phylogenetic analysis of the resulting 73.5 kb of sequence data demonstrated that for all 16 loci (14.7 kb/genome), the topology expected from organismal history was recovered. In contrast to observations involving repetitive DNAs in cotton, there was no evidence of interaction among duplicated genes in the allopolyploid. Polyploidy was not accompanied by an obvious increase in mutations indicative of pseudogene formation. Additionally, differences in rates of divergence among homoeologues in polyploids and orthologues in diploids were indistinguishable across loci, with significant rate deviation restricted to two putative pseudogenes. Our results indicate that most duplicated genes in allopolyploid cotton evolve independently of each other and at the same rate as those of their diploid progenitors. These indications of genic stasis accompanying polyploidization provide a sharp contrast to recent examples of rapid genomic evolution in allopolyploids.

Biological Evolution↗

Comprehensive identification and analysis of clusters of tandemly duplicated genes reveal their contributions to adaptive evolution of green plants.

Tandem gene duplication occurred more frequently compared with the episodic whole-genome duplication (WGD), providing a continuous supply of genetic material for evolutionary innovation and adaptation to changing environments. The rising roles of clusters of tandemly duplicated genes (CTDGs) in the evolution of phenotypic diversity have been unraveled in mammals. However, the content and biological roles of CTDGs remain largely unknown in plants. Here, we comprehensively identified CTDGs in 220 published plant genomes representing major lineages of green plants. The number of CTDGs showed great variation across taxa, ranging from 0 to 6028. The size of CTDGs varied from 2 to 47 genes, with small clusters containing two members predominating. Interestingly, significant expansion of CTDGs was found in early-diverging land plants and is closely associated with the evolution of key traits (e.g., ABA response, plant cuticle, UV-B resistance) required for plants to conquer terrestrial environments. Functional enrichment analysis revealed conserved and specialized functional profiles among different sizes of CTDGs in both Arabidopsis thaliana and the bryophyte Physcomitrium patens. Small CTDGs were enriched in fundamental stress responses, including protein modification, signal transduction, and responses to diverse stress stimuli, while large CTDGs were enriched in more sophisticated processes such as plant hormone biosynthesis and signaling, plant-microbe interactions, and reproductive processes. Expression pattern analyses of CTDGs under different stress conditions in A. thaliana and P. patens revealed that the highest number of CTDGs showed differential expression under drought stress, suggesting important roles of CTDGs in the evolution of desiccation tolerance in early land plants. The results of this study provide new additions to our knowledge about the abundance of CTDGs across green plants and reveal their important contributions to enable plants to overcome stressful environments on land.

Gene Duplication↗

Extent of gene duplication in the genomes of Drosophila, nematode, and yeast.

We conducted a detailed analysis of duplicate genes in three complete genomes: yeast, Drosophila, and Caenorhabditis elegans. For two proteins belonging to the same family we used the criteria: (1) their similarity is > or =I (I = 30% if L > or = 150 a.a. and I = 0.01n + 4.8L(-0.32(1 + exp(-L/1000))) if L < 150 a.a., where n = 6 and L is the length of the alignable region), and (2) the length of the alignable region between the two sequences is > or = 80% of the longer protein. We found it very important to delete isoforms (caused by alternative splicing), same genes with different names, and proteins derived from repetitive elements. We estimated that there were 530, 674, and 1,219 protein families in yeast, Drosophila, and C. elegans, respectively, so, as expected, yeast has the smallest number of duplicate genes. However, for the duplicate pairs with the number of substitutions per synonymous site (K(S)) < 0.01, Drosophila has only seven pairs, whereas yeast has 58 pairs and nematode has 153 pairs. After considering the possible effects of codon usage bias and gene conversion, these numbers became 6, 55, and 147, respectively. Thus, Drosophila appears to have much fewer young duplicate genes than do yeast and nematode. The larger numbers of duplicate pairs with K(S) < 0.01 in yeast and C. elegans were probably largely caused by block duplications. At any rate, it is clear that the genome of Drosophila melanogaster has undergone few gene duplications in the recent past and has much fewer gene families than C. elegans.

Animals↗

A sensitive method for detecting variation in copy numbers of duplicated genes.

Gene duplications are common in the vertebrate genome, and duplicated loci often show a variation in copy number that may have important phenotypic effects. Here we describe a powerful method for quantification of duplicated copies based on pyrosequencing. A reliable quantification was obtained by amplification of the duplication break-point and a corresponding nonduplicated sequence in a competitive PCR assay. A comparison with an independent method for quantification based on the Invader technology revealed an excellent correlation between the two methods. The pyrosequencing-based method was evaluated by analyzing variation in copy number at the duplicated KIT/Dominant white locus in pigs. We were able to distinguish haplotypes at this locus by combining the duplication breakpoint test with a diagnostic test for a functionally important splice mutation in the duplicated gene. An extensive allelic variation, including the presence of a new allele carrying a single KIT copy expected to encode a truncated KIT receptor, was revealed when analyzing white pigs from commercial lines.

Algorithms↗