PubMed HealthSearch

SEARCH · PubMed Health

Results for “Gene Duplication”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software

Comprehensive identification and analysis of clusters of tandemly duplicated genes reveal their contributions to adaptive evolution of green plants.

Tandem gene duplication occurred more frequently compared with the episodic whole-genome duplication (WGD), providing a continuous supply of genetic material for evolutionary innovation and adaptation to changing environments. The rising roles of clusters of tandemly duplicated genes (CTDGs) in the evolution of phenotypic diversity have been unraveled in mammals. However, the content and biological roles of CTDGs remain largely unknown in plants. Here, we comprehensively identified CTDGs in 220 published plant genomes representing major lineages of green plants. The number of CTDGs showed great variation across taxa, ranging from 0 to 6028. The size of CTDGs varied from 2 to 47 genes, with small clusters containing two members predominating. Interestingly, significant expansion of CTDGs was found in early-diverging land plants and is closely associated with the evolution of key traits (e.g., ABA response, plant cuticle, UV-B resistance) required for plants to conquer terrestrial environments. Functional enrichment analysis revealed conserved and specialized functional profiles among different sizes of CTDGs in both Arabidopsis thaliana and the bryophyte Physcomitrium patens. Small CTDGs were enriched in fundamental stress responses, including protein modification, signal transduction, and responses to diverse stress stimuli, while large CTDGs were enriched in more sophisticated processes such as plant hormone biosynthesis and signaling, plant-microbe interactions, and reproductive processes. Expression pattern analyses of CTDGs under different stress conditions in A. thaliana and P. patens revealed that the highest number of CTDGs showed differential expression under drought stress, suggesting important roles of CTDGs in the evolution of desiccation tolerance in early land plants. The results of this study provide new additions to our knowledge about the abundance of CTDGs across green plants and reveal their important contributions to enable plants to overcome stressful environments on land.

Gene Duplication

ERCnet: Phylogenomic Prediction of Interaction Networks in the Presence of Gene Duplication.

Assigning gene function from genome sequences is a rate-limiting step in molecular biology research. A protein's position within an interaction network can potentially provide insights into its molecular mechanisms. Phylogenetic analysis of evolutionary rate covariation (ERC) in protein sequence has been shown to be effective for large-scale prediction of functional relationships and interactions. However, gene duplication, gene loss, and other sources of phylogenetic incongruence are barriers for analyzing ERC on a genome-wide basis. Here, we developed ERCnet, a bioinformatic program designed to overcome these challenges, facilitating efficient all-versus-all ERC analyses for large protein sequence datasets. We simulated proteome datasets and found that ERCnet achieves combined false positive and negative error rates well below 10% and that our novel "branch-by-branch" length measurements outperforms "root-to-tip" approaches in most cases, offering a valuable new strategy for performing ERC. We also compiled a sample set of 35 angiosperm genomes to test the performance of ERCnet on empirical data, including its sensitivity to user-defined analysis parameters such as input dataset size and branch-length measurement strategy. We investigated the overlap between ERCnet runs with different species samples to understand how species number and composition affect predicted interactions and to identify the protein sets that consistently exhibit ERC across angiosperms. Our systematic exploration of the performance of ERCnet provides a roadmap for design of future ERC analyses to predict functional interactions in a wide array of genomic datasets. ERCnet code is freely available at https://github.com/EvanForsythe/ERCnet.

Gene Duplication

Recent gene duplication and structural remodeling drive rapid lineage-specific gene family evolution in plants.

Gene duplication promotes the generation of novel gene functions and trait diversity across species. Here, we present DupHIST, a computational pipeline that reconstructs the hierarchical timing of gene duplications by integrating maximum likelihood (ML)-based phylogeny with substitution-derived timing via statistical smoothing. Applied to over 4.5 million genes from 114 plant genomes, we successfully inferred duplication histories across nearly 130,000 orthogroups. This large-scale analysis showed that 53.0% of genes arose from recent, lineage-specific duplications, with high concentrations in particular multi-copy families. Among these, NLR, C48, and P450 families exemplified how recently duplicated genes undergo rapid stepwise structural remodeling. This process was primarily driven by small-scale mutations, including insertions, deletions, and frameshifts, that rapidly accumulated shortly after duplication. By resolving the precise duplication order, we reconstructed these architectural changes, thereby enabling both the inference of putative ancestral structures and the exploration of functional diversification arising from structural remodeling. Structure-based clustering further uncovered that recently duplicated, uncharacterized genes retain core domain structures resembling known functional proteins even across phylogenetically distant species lacking sequence homology. Our findings reveal that recent gene duplications and subsequent structural remodeling represent a widespread and lineage-specific force driving rapid diversification of gene families in plants.

Gene duplication history

Evolution of the differential regulation of duplicate genes after polyploidization.

In the 50 million years since the polyploidization event that gave rise to the catostomid family of fishes the duplicate genes encoding isozymes have undergone different fates. Ample opportunity has been available for regulatory evolution of these duplicate genes. Approximately half the duplicate genes have lost their expressions during this time. Of the duplicate genes remaining, the majority have diverged to different extents in their expression within and among adult tissues. The pattern of divergence of duplicate gene expression is consistent with the accumulation of mutations at regulatory genes. The absence of a correlation of extent of divergence of gene expression with the level of genetic variability for isozymes at these loci is consistent with the view that the rates of regulatory gene and structural gene evolution are uncoupled. The magnitude of divergence of duplicate gene expressions varies among tissues, enzymes, and species. Little correlation was found with the extent of divergence of duplicate gene expression within a species and its degree of morphological "conservatism", although species pairs which are increasingly taxonomically distant are less likely to share specific patterns of differential gene expression. Probable phylogenetic times of origin of several patterns of differential gene expression have been proposed. Some patterns of differential gene expression have evolved in recent evolutionary times and are specific to one or a few species, whereas at least one pattern of differential gene expression is present in nearly all species and probably arose soon after the polyploidization event. Multilocus isozymes, formed by polyploidization, provide a useful model system for studying the forces responsible for the maintenance of duplicate genes and the evolution of these once identical genes to new spatially and temporally specific patterns of regulation.

Animals

Diverse evolutionary rates and gene duplication patterns among families of functional olfactory receptor genes in humans.

In humans, odors are detected by ~400 functional olfactory receptor (OR) genes. The superfamily of functional OR genes can be further divided into tens of families. In large part, the OR genes have experienced extensive tandem duplications, which have led to gene gains and losses. However, whether different OR gene families have experienced distinct modes of gene duplication has yet to be reported. We conducted comparative genomic and evolutionary analyses for human functional OR genes. Based on analysis of human-mouse 1-1 orthologs, we found that human functional OR genes show higher-than-average evolutionary rates, and there are significant differences among families of functional OR genes. Via comparison with seven vertebrate outgroups, families of human functional OR genes show different extents of gene synteny conservation. Although the superfamily of human functional OR genes is enriched in tandem and proximal duplications, there are particular families which are enriched in segmental duplications. These findings suggest that human functional OR genes may be governed by different evolutionary mechanisms and that large-scale gene duplications have contributed to the early evolution of human functional OR genes.

Humans

Gene duplication in tetraploid fish: model for gene silencing at unlinked duplicated loci.

Several groups of fishes, including salmonids and catastomids, appear to have originated through genome duplication events. However, these two groups retain approximately 50% of the loci examined as functioning duplicates, despite the passage of 50 million years or more of mutation and selection. Although other effects are not excluded, this apparently slow rate of duplicate silencing can be explained in terms of the effects of selection against defective double homozygotes to unlinked duplicates. We have derived a computer simulation of genetic drift that affords direct evaluation of the effects of population size (N), mutation rate (micron), initial allele frequencies, back mutation, fitness, and time on the probability of fixation for null alleles at unlinked duplicate loci. The results show that this probability is approximately linearly related to population size for N greater than or equal to 10(3). Specifically, for naive populations, the time for 50% probability of gene silencing is approximately equal to 15N + micron-3/4 generations. The retention of 50% of the loci as functional duplicates may therefore result from the large effective size of salmonid and catastomid populations. The results also show that, under most conditions for populations of 2000--3000 or larger, unlinked duplicate loci will be sustained in the functional state longer than tandem (linked) duplicates and hence are available for evolution of new functions for a longer time.

Alleles

Tandem gene duplication facilitates intertidal adaptation in atypical mangrove plants.

Mangrove plants, originating from inland ancestors, have independently adapted to extreme intertidal zones characterized by salt and hypoxia stress. While typical mangroves exhibit specialized phenotypes, like viviparous seeds and salt secretion, atypical clades that have thrived without such traits are particularly suitable for exploring the molecular and physiological basis underlying plant adaptation to intertidal zones. We assembled a chromosome-level genome of an atypical mangrove, Scyphiphora hydrophylacea, the only mangrove species in Gentianales. Similar to other mangroves, S. hydrophylacea colonized intertidal zones during climatic optimum periods of sea-level rise. Despite lacking recent whole-genome duplications (WGDs), its genome acquired extensive tandem gene duplications (TDs), leading to the rapid expansion of key salt- and hypoxia-related genes. Transcriptome data further corroborated that TD-driven gene expansions contribute to stress tolerance. Specifically, the expansion of genes involved in cation transmembrane transport, osmotic regulation, and oxidative stress response may enhance salinity tolerance, and the expansion of signal transduction and energy metabolism genes in hypoxia-response pathways may confer waterlogging tolerance. Therefore, in the absence of large-scale gene duplication, the rapid expansion of core genes involved in salt and hypoxia tolerance through tandem duplication may represent a key force driving the adaptation of atypical mangroves. These findings also provide valuable insights for crop improvement strategies aimed at enhancing environmental resilience while maintaining phenotypic stability.

Gene Duplication

Coexpression of neighboring genes in Caenorhabditis elegans is mostly due to operons and duplicate genes.

In many eukaryotic species, gene order is not random. In humans, flies, and yeast, there is clustering of coexpressed genes that cannot be explained as a trivial consequence of tandem duplication. In the worm genome this is taken a step further with many genes being organized into operons. Here we analyze the relationship between gene location and expression in Caenorhabditis elegans and find evidence for at least three different processes resulting in local expression similarity. Not surprisingly, the strongest effect comes from genes organized in operons. However, coexpression within operons is not perfect, and is influenced by some distance-dependent regulation. Beyond operons, there is a relationship between physical distance, expression similarity, and sequence similarity, acting over several megabases. This is consistent with a model of tandem duplicate genes diverging over time in sequence and expression pattern, while moving apart owing to chromosomal rearrangements. However, at a very local level, nonduplicate genes on opposite strands (hence not in operons) show similar expression patterns. This suggests that such genes may share regulatory elements or be regulated at the level of chromatin structure. The central importance of tandem duplicate genes in these patterns renders the worm genome different from both yeast and human.

Animals

Production of a soluble form of fumarate reductase by multiple gene duplication in Escherichia coli K12.

1. Ampicillin-hyperresistant mutants of Escherichia coli K12 bearing multiple gene duplications in the ampC (beta-lactamase) gene region of the chromosome overproduced at least six proteins with molecular weights 97,000, 80,000, 72,000, 49,000, 33,000 and 26,500 during anaerobic growth. All but two of the proteins (80,000-Mr and 49,000-Mr) were also overproduced during aerobic growth. The distribution of the proteins in soluble and particulate cell fractions was investigated. 2. The 33,000-Mr and 72,000-Mr components were identified as beta-lactamase and the amp-linked frdA gene product, fumarate reductase, respectively. Co-sedimentation of the 26,500-Mr component with the fumarate reductase suggested that the smaller protein could be functionally related to the reductase. The lack of correspondence between the amplified proteins and the products of other amp-linked genes, aspA and mop(groE), indicated that these genes are not included in the repetitive sequence. 3. Fumarate reductase activities were amplified up to 32-fold by the multiple gene duplications. Two forms of fumarate reductase were produced: particulate (membrane-bound) and soluble (cytoplasmic). Production of the soluble form occurred when the binding capacity of the membrane was saturated. Both forms of fumarate reductase were enzymically active but the soluble form was readily inactivated under assay conditions.

Centrifugation, Density Gradient

Diversification of Cellulose Synthase (CESA) Genes in Mosses Suggests Both Ancient and Recent Gene duplications.

Cellulose is an important polysaccharide that constitutes all plant cell walls, giving them strength and stability. The plant cellulose synthase (CESA) gene family, which encodes the catalytic subunits of cellulose synthesis complexes (CSCs), has diversified independently in several plant lineages, providing an interesting model for understanding selection for gene duplication. Here we quantified the presence of CESA genes across mosses to understand how the process of gene family diversification occurred in this group and how it parallels diversification in other groups. We first examined the CESA gene family in eight species of mosses across seven families for which whole genome assemblies were available. We then identified CESA genes from additional species, for which only short-read sequence data was available, by using BLAST searches and targeted gene assemblies. We validated this approach by comparing the assembled paralogs from the short-read data to the genes identified from whole genome assemblies in the eight reference species. This approach allowed us to identify paralogs directly from short-read data and greatly expand our sample set. Results from the combined empirical data support the hypothesis that CESA genes diversified within the moss lineage at least as early as the mesozoic period, during or possibly even prior to the onset of moss diversification, but also continue to diversify within modern species. In addition, we found evidence for purifying selection as the dominant force shaping these genes and observed that different lineages experienced different levels of evolutionary constraint. Lastly, our approach to assemble paralogs has the potential to allow researchers to improve analyses of gene duplication events.

Physcomitrium patens

Partial gene duplication and posterior pituitary peptide.

We have compiled the dipeptide frequencies in 100 known protein sequences. We suggest that dipeptides which occur with low frequencies can be used to locate proteins where partial gene duplication may have taken place. The 48 residue sequence of posterior pituitary peptide contains two Cys Trp pairs. The adjacent portions of the sequence are compatible with a partial gene duplication in the evolutionary history of posterior pituitary peptide.

Amino Acid Sequence

The evolution of protein sequences by repetitious gene duplication: clostridial flavodoxin.

Internal regularities of amino acid sequences of flavodoxins, FMN-containing, low molecular weight flavoproteins, were statistically examined using the minimum mutation method. The sequence of Clostridium pasteurianum flavodoxin shows statistically significant evidence of repetitious internal gene duplications at different levels of structure. Peptide pairs with a low chance probabilitiy of occurrence were frequently observed at a shift of 5 residues. The pairs with the lowest chance probabilities are a pair of heptapeptides at positions39--45 vs. 44--50, a 5 residue shift (p = 9 x 10(-6)). Most of the related pairs are consistent and could best be explained by the repeating pentapeptide sequence: (Lys-Gly-Ala-Asp-Val-)n and appropriate gaps. Internal repetitions with longer shifts were also suggested for other flavodoxins. Repetitious gene duplication is proposed for the early stages of flavodoxin evolution.

Amino Acid Sequence

Horsetail (Equisetum arvense) ferredoxins I and II Amino acid sequences and gene duplication.

Amino acid sequences of two ferredoxins isolated from Equisetum arvense were determined by conventional procedures. Ferredoxins I and II of E. arvense had 95 and 93 residues, respectively, and nearly identical sequences each with only one amino acid difference from ferredoxins. I and II of E. telmateia (1). The overall structural characteristics of these two ferredoxins were therefore very similar to those of E. telmateia ferredoxins. Ferredoxins I and II from E. arvense differ in 31 sites and those from E. telmateia in 29 sites from each other. These facts suggested that duplication of the ferredoxin gene in one organism occurred at an early evolutionary stage long before the divergence of the two horsetail species. The number of differences in amino acids between horsetail ferredoxins and other chloroplast-type ferredoxins indicated that the duplication occurred after divergence of horsetails from other plants. Comparing green plant ferredoxins, it was estimated that this gene duplication occurred about 250 million years ago. Some comments on the unique amino acid substitutions in horsetail ferredoxins are also presented.

Amino Acid Sequence

Structural evidence for gene duplication in the evolution of the acid proteases.

X-ray studies of acid proteases indicate a bilobal structure with a well defined active site cleft. An intramolecular twofold symmetry axis relates two topologically similar domains and the active site residues. A possible mechanism for evolution by gene duplication, divergence and gene fusion is presented.

Amino Acid Sequence

A comprehensive examination of protein sequences for evidence of internal gene duplication.

We have implemented a routine procedure for screening protein sequences for evidence of intragenic duplications. We tested 163 protein sequences representing 116 superfamilies of unrelated proteins. Twenty superfamilies contain proteins with internal gene duplications. The intragenic duplications detected can be divided into two major types. (1) One or more duplications of all or part of a gene produce a protein with two or several detectable regions of sequence homology. Sequences from 18 superfamilies contained this type of duplication. (2) Repeated reduplication of a small DNA segment can produce a protein that is repetitive over most of its length. Three superfamilies contain such repetitive sequences. We also investigated the limits of detection of ancient duplications using sequences derived by random mutation of a model sequence consisting of ten 10-residue repeats. The original repetitive nature of the sequence was usually detected after 250 point mutations even though the ancestral segment could not be accurately reconstructed.

Amino Acid Sequence

Gene duplication at an isocitrate dehydrogenase locus in Scaphiopus.

Spadefoot toads of the subgenus Scaphiopus have two isocitrate dehydrogenase loci, with no intergenic interaction between them. Toads of the subgenus Spea have three Idh loci, with intergenic enzymes formed between two of them, providing strong evidence for their homology and the origin of one through a duplication process. The Idh phenotype of interspecific hybrids is consistent with the theory of a gene duplication.

Animals