PubMed HealthSearch

SEARCH · PubMed Health

Results for “Duplication events”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The phosphofructokinase genes of yeast evolved from two duplication events.

Yeast phosphofructokinase (PFK) is an octameric enzyme composed of four alpha-subunits and four beta-subunits, encoded by the genes PFK1 and PFK2, respectively. PFK1 was mapped 23 cM distal to ADE3 on chromosome VII, and PFK2 30 cM proximal to RNA1 on chromosome XIII. The entire nucleotide sequences for the two genes were obtained by sequencing both DNA strands. Only one major open reading frame was found for each gene. They encode 987 aa for PFK1 (Mr 107,984) and 959 aa for PFK2 (Mr 104,589). Both genes show a biased codon usage. The deduced amino acid sequences showed: (i) 20% homology between the N- and the C-terminal halves of each subunit, (ii) 55% homology between the two subunits, and (iii) significant homologies to the PFK sequences from human and rabbit muscle (42%), Escherichia coli (34%), and Bacillus (36%). These data support the view that two gene duplication events occurred in the evolution of the yeast PFK genes. The first duplication event took place soon after the separation of prokaryotic and eukaryotic lineage and the second in Saccharomyces later in the phylogeny. Functional domains in the yeast subunits were deduced by comparison to the rabbit muscle enzyme.

Amino Acid Sequence

Evolution of the sarafotoxin/endothelin superfamily of proteins.

Sixteen protein and nucleic acid sequences from the vasoconstrictor sarafotoxin/endothelin/endothelin-like superfamily of peptides were studied, and the evolutionary relationships between the sarafotoxin and endothelin gene families as well as the phylogenetic topology within each gene family and the three endothelin subfamilies was reconstructed. The endothelin gene family has diverged from an ancestral gene that has experienced an exon duplication event followed by two gene duplication events. The sarafotoxins' lineage diverged from the ancestral gene prior to the first endothelin gene duplication event. Analysis of the resulting phylogenetic trees revealed that in several lineages, the peptides have independently accumulated identical replacements in position 2, therefore supporting the hypothesis that residue 2 is crucial to their activity.

Amino Acid Sequence

Gene and Genome Duplication in Spiders.

Gene and genome duplications are widely observed across various organisms, including plants, yeasts, and animals. Numerous studies link gene duplications to the emergence of novel phenotypes, supporting the hypothesis that duplication events are advantageous for adaptive evolution. Whole-genome duplications (WGD) are especially prevalent in plants and have also occurred ancestrally in vertebrates. However, large-scale duplication events in other animal groups remain understudied, partly due to limited genomic resources. Arthropods, particularly insects, represent one of the most diverse animal clades in terms of both species and phenotypic diversity. With increasing availability of chromosome-level genomes, large-scale duplications appear to be rare in insects but are more frequent in chelicerates (e.g. spiders, scorpions, and horseshoe crabs). This makes chelicerates an intriguing group for comparing the mechanisms, fates, and evolutionary impacts of large-scale duplications with those seen in plants and vertebrates. In this review, we synthesize and discuss current research on WGD in spiders and discuss different scenarios for genes following gene duplication events (conservation, nonfunctionalization, subfunctionalization, specialization, drift, neofunctionalization) in the context of experimental studies. We hypothesize if there might be common trajectories after duplication and how these could be tested.

Animals

The evolutionary history of the sarafotoxin/endothelin/endothelin-like superfamily.

The evolutionary relationships among 17 protein and nucleic acid sequences from the sarafotoxin/endothelin/endothelin-like superfamily of peptides were studied. The endothelin/endothelin-like gene family has diverged from an ancestral gene that has experienced an exon duplication event followed by two complete gene duplications. The sarafotoxin lineage diverged from the ancestral gene prior to the first gene duplication event. In several lineages, the peptides have independently accumulated identical amino acid replacements in position 2. This finding supports the hypothesis that residue 2 is crucial to biological activity.

Amino Acid Sequence

Identification of a heat-shock pseudogene from Caenorhabditis elegans.

While characterizing the hsp70 gene family from Caenorhabditis elegans we encountered an unusual member of this family. Sequence data reveal that the hsp-2ps gene is a pseudogene of the constitutively expressed, heat-inducible hsp-1 gene. Two stop codons generated near the 5' end of the sequence as well as several frameshift mutations and a large internal deletion confirm the identification of hsp-2ps as a pseudogene. The nucleotide substitution rate of the third codon position was twice that of the first and second codon positions, suggesting that the hsp-2ps gene was nonfunctional since the time of the duplication event. The hsp-2ps gene duplicates a region of the hsp-1 gene that lies exclusively within the transcribed region and retains the introns. We feel that the hsp-2ps gene was produced by a transpositional duplication event, which occurred approximately 8.5 million years ago.

Amino Acid Sequence

Episode clustering in phylogenetic networks.

MOTIVATION: The classical duplication episode clustering (EC) model introduced by Guigó et al. in the 1990s provides a foundational approach for inferring genomic duplication events crucial to understanding genome evolution. This model clusters single gene duplications from a collection of gene trees at locations in the species tree to minimize the total number of such locations, called duplication episodes. However, it does not capture reticulate evolutionary histories. RESULTS: Here, we introduce NetEC, a novel extension of this problem to phylogenetic networks. To solve NetEC, we first develop a polynomial-time dynamic programming (DP) algorithm for testing whether a given set of network nodes can serve as episode locations. We then propose a main inference algorithm that utilizes this DP component to optimize the episode count; while the feasibility test runs in polynomial time, the full optimization has exponential worst-case complexity, and an optional heuristic mode is provided for larger instances. We also propose an extended episode analysis procedure that identifies additional genomic duplication candidates below reticulation nodes, complementing the main algorithm by resolving potential upward clustering of duplications induced by reticulation. We evaluate our method on simulated data and on an empirical Pandanales dataset comprising over 29 000 gene trees, demonstrating exact and accurate inference of genomic duplication events even in the presence of multiple reticulations. AVAILABILITY AND IMPLEMENTATION: All experiments were conducted using the NetEC tool (https://github.com/ppgorecki/netec), with all input data, scripts, and parameter settings for reproduction available in the same repository.

Phylogeny

Genome-wide cyclin gene evolution in Arabidopsis and Brassica reveals polyploidization-driven duplication and flowering-time associations.

Cyclin genes are plant cell cycle regulators that play essential roles in growth, development, and reproduction. However, the evolutionary dynamics and genomic organization of cyclin genes across the Brassicaceae family remain poorly understood, particularly in the context of allotetraploid genome evolution. Here, we investigated the diversity, expansion mechanisms, and potential functional diversification of cyclin genes across ten Brassicaceae genomes, including four Arabidopsis and six Brassica species. A total of 1087 cyclin genes representing 23 cyclin types were identified. Comparative genomic analyses revealed that cyclin gene expansion was strongly influenced by polyploidization in Brassica species, with 1845 duplication events involving 1063 genes. Whole-genome duplication was the predominant mechanism driving expansion, while both inter- and intra-genomic duplications contributed to gene retention in tetraploid Brassica species, with the highest duplication frequency observed in Brassica juncea. Across genomes, 120 physical gene clusters were identified, including homogeneous and heterogeneous types. Ortholog analysis between progenitor and allotetraploid species identified 852 orthologous pairs involving 366 genes, indicating extensive conservation following allotetraploid formation. Phylogenetic analysis resolved cyclins into three major clades, while expression-based clustering in Brassica napus grouped genes into four major clusters, suggesting functional diversification. Integration of pan-genomic and flowering-time QTL analyses further identified two cyclin genes, Bna21cycA2 and Bna113cycD4, which contain amino acid polymorphisms and represent putative candidate variations potentially associated with flowering-time variation across multiple genomes. These findings provide new insights into the evolutionary expansion, retention, and potential functional divergence of cyclin genes in Brassicaceae and highlight candidate loci for future functional studies and crop improvement.

Evolution, Molecular

Myogenin is in an evolutionarily conserved linkage group on human chromosome 1q31-q41 and unlinked to other mapped muscle regulatory factor genes.

Myogenin is a member of a family of muscle-specific regulatory factors which includes MyoD1, Myf-5, and Myf-6 (also called MRF4 and herculin). Extensive regions of sequence homology in genes for these three factors suggest duplication events associated with their evolution. In the present study, the chromosomal location of the myogenin gene in humans (MYOG), mice (Myog), and Chinese hamsters (MYOG) was determined using in situ hybridization to human metaphase chromosomes as well as segregation analysis among interspecific somatic cell hybrid panels and interspecific backcrossed mice. We localize the gene encoding myogenin to human chromosome 1q31-q41 within a linkage group homologous with a region on mouse chromosome 1 and Chinese hamster chromosome 5. The results verify the nonlinkage of MYOG to MYOD1, MYF5, and MYF6 genes and indicate that events associated with the duplication of MYOG with respect to MYOD1, MYF5, or MYF6 loci were not chromosome-wide.

Animals

Genome-wide identification and expression analysis of the UGT gene family in honeysuckle.

BACKGROUND: The UGT gene family plays critical roles in regulating plant growth, development, stress responses, and secondary metabolite synthesis. Although UGT proteins have been studied in numerous plant species, research on the UGT family in honeysuckle (Lonicera japonica Thunb.) remains limited. RESULTS: In this study, a comprehensive genome-wide analysis of the UGT gene family was performed in honeysuckle. A total of 224 unique LjUGT genes were identified and classified into 21 distinct subfamilies (T71-T92 without T77) based on the phylogenetic analysis. These genes were unevenly distributed on the 9 chromosomes. Eighteen segmental duplication events and 61 tandem duplications were identified, of which only 3 were positive selection. Integrated analysis of promoter cis-acting elements, transcription factors, targeted miRNAs, and interacting proteins suggested that the expression and function of the LjUGT genes may be regulated by transcription factors and proteins through binding to the various binding sites and cis-acting elements, thereby putatively participating in diverse biological processes, including hormone signaling, stress response, and metabolism. The expression pattern analysis of LjUGTs in different tissues and under stress conditions indicated that Lj2A1135G32, Lj5A236T61, Lj6A350T83, and Lj7A737T47 emerged as candidate genes potentially associated with development, 46 genes showed expression changes under all 6 abiotic stresses, suggesting broad stress responsiveness. Additionally, there 7 genes were identified as candidate hub genes that may correlate with the low temperature stress tolerance in honeysuckle according to the WGCNA results, and further verification by qRT-PCR confirmed that Lj4A99G61 and Lj9A591T82 can be regarded as key candidate genes for in-depth research. CONCLUSIONS: This study systematically identified 224 LjUGT genes in honeysuckle for the first time and characterized their physicochemical properties, phylogenetic relationship, and expression patterns. These findings provide a foundational resource for hypothesis-driven investigations into the functions and action mechanisms of LjUGTs.

Lonicera

A beta-galactosidase deletion mutant of Lactobacillus bulgaricus reverts to generate an active enzyme by internal DNA sequence duplication.

Several spontaneous Lac- deletion derivatives of the beta-galactosidase gene of Lactobacillus bulgaricus were analyzed for their phenotypic stability. We found that one of these mutants, lac139, carrying a deletion of 30 bp within the gene, was able to revert to a Lac+ phenotype. Genetical analysis of revertants indicated that an internal region of 72 bp was duplicated immediately next to the deletion site. The region involved in the duplication event is flanked by direct repeated sequences of 13 bp in length. Both events, the deletion and the duplication, were mediated by the presence of such short direct repeats. Enzymatic studies of the purified proteins indicated identical kinetic parameters, but showed considerable instability of the revertant protein.

Base Sequence

Molecular evolution of the members of the Snq2/Pdr18 subfamily of Pdr transporters in the Hemiascomycete yeasts.

The transporters of the ATP-Binding Cassette (ABC) Superfamily involved in the Multidrug Resistance (MDR) phenomena are also known as ABC-Pleiotropic Drug Resistance (PDR) proteins. The homologs of the Saccharomyces cerevisiae SNQ2 and PDR18 genes were identified in 171 yeast genomes, representing 68 different hemiascomycetous species. All early-divergent yeast species analyzed in this work lack Snq2/Pdr18 homologs, suggesting that the origin of these ABC-PDR genes in hemiascomycete yeasts resulted from a horizontal transfer event. The evolutionary pathway of the Snq2/Pdr18 protein subfamily in pathogenic Candida species was also reconstructed, revealing a main gene lineage leading to the Candida albicans SNQ2 gene. The results indicate that, after the gene duplication event at the origin of the SNQ2/PDR18 paralogs, the PDR18 ortholog has been under strong diversifying selection and suggest that a small portion of the sequence of the SNQ2 ancestral ortholog might have been under mild positive selection. The results also showed that strong positive selection was exerted over one of the two paralogs generated by the Whole Genome Duplication (WGD) event, corresponding to the duplicate at the origin of a "short-lived" WGD sublineage.

Evolution, Molecular

Comparative genomic analysis of Artemisia argyi reveals asymmetric expansion of terpene synthases and conservation of artemisinin biosynthesis.

Artemisia argyi, a perennial herb of the Asteraceae family, possesses significant therapeutic and economic value. We present a 7.88 Gb chromosome-level haplotype-resolved genome assembly, revealing its unique evolutionary trajectory. The karyotype (2n = 34) of A. argyi is that of an autotetraploid, which underwent gametic chromosome fusion prior to species-specific whole-genome duplication (WGD-3). The genome exhibits pronounced multivalent chromosome pairing and frequent recombination among homologous groups. Asymmetrical evolution following WGD-3 is a hallmark feature, evidenced by imbalanced allelic gene loss and widespread neofunctionalization. The terpene synthase (TPS) gene family exemplifies this pattern, having expanded through four duplication events in A. argyi. Recent tandem duplications and allelic functional differentiation have generated substantial gene functional diversity. Notably, we identified a tandem-duplicated six-copy ADS homolog (AarADS)-a key TPS gene in the artemisinin biosynthetic pathway of Artemisia annua (AanADS)-localized exclusively to a single chromosome in A. argyi. Unlike AanADS, which converts farnesyl pyrophosphate (FPP) to amorpha-4,11-diene, AarADS catalyzes FPP to α-bisabolol. Evolutionary analysis suggested that AanADS acquired its specialized function via a derived mutation in the A. annua lineage. This study elucidates the genomic evolution underpinning A. argyi's distinctive medicinal properties.

Alkyl and Aryl Transferases

Haplotype-resolved genome of Forsythia suspensa reveals the reticulate evolution in Oleaceae and a novel gene cluster regulating stamen development.

The olive family (Oleaceae) comprises numerous species of economic, horticultural, and medicinal importance. Despite its significance, the evolutionary history of this complex family remains enigmatic. Here, we generated a high-quality haplotype-resolved genome of Forsythia suspensa, a distylous species that occupies a key phylogenetic position in Oleaceae. The 2 haplotypes exhibit significant allelic divergence with potential allele-specific regulation. We reconstructed the polyploidization history of Oleaceae by confirming and precisely dating a shared whole-genome triplication and an independent whole-genome duplication event. We revealed a complex reticulate evolution that gave rise to the tribe Oleeae: an initial hybridization between Forsythieae (♂) and Jasmineae (♀), a subsequent backcrossing event, and a final whole-genome duplication. We identified a novel tandemly duplicated pectin methylesterase inhibitor gene cluster that regulates filament length and pollen size via restricting cell elongation in the long-styled morph. Dosage augmentation via stepwise cluster formation (0.99 to 3.83 Mya) may contribute to maintaining stamen traits of the long-styled morph. These FsPMEIs are co-expressed with many cell wall-related genes, suggesting a functional link in cell wall modification. Our study reveals the reticulate evolution in Oleaceae and a novel gene cluster controlling stamen development in F. suspensa and provides valuable haplotype-resolved genomic resources for heterostylous species, offering novel framework and molecular pathways to understand plant adaptive evolution.

Forsythia

Molecular history of gene conversions in the primate fetal gamma-globin genes. Nucleotide sequences from the common gibbon, Hylobates lar.

Comparative and phylogenetic analyses of homologous sequences from closely related species reveal genetic events which have happened in the past and thus provide considerable insight into molecular genetic processes. One such process which has been especially important in the evolution of multigene families is gene conversion. The fetal gamma 1 and gamma 2-globin genes of catarrhine primates (humans, apes, and Old World monkeys) underwent numerous gene conversion events after they arose from a gene duplication event 25-35 million years ago. By including the gamma 1- and gamma 2-globin gene sequences from the common gibbon, Hylobates lar, the present work expands the gamma-globin data set to represent all major groups of hominoid primates. A computer-assisted algorithm is introduced which reveals converted DNA segments and provides results very similar to those obtained by site-by-site evolutionary reconstruction. Both methods provide strong evidence for at least 14 different converted stretches in catarrhine primates as well as five conversions in ancestral lineages. Features of gene conversions generalized from this molecular history are 1) conversions are restricted to regions maintaining high degrees of sequence similarity, 2) one gene may dominate in converting another gene, 3) sequences involved in conversions may accumulate changes more rapidly than expected, and 4) certain elements, such as polypurine/polypyrimidine [Y)n) and (TG)n elements, appear to be hotspots for initiating or terminating conversion events.

Amino Acid Sequence

Physical mapping of the human carbonic anhydrase gene cluster on chromosome 8.

A cluster of genes encoding the three cytoplasmic carbonic anhydrase isozymes CAI, CAII, and CAIII lie on the long arm of chromosome 8 (8q22) in humans. These genes have been mapped using pulsed-field gel electrophoresis. The genes lie in the order CA2, CA3, CA1. CA2 and CA3 are separated by 20 kb and are transcribed in the same direction, away from CA1. CA1 is separated from CA3 by over 80 kb and is transcribed in the direction opposite to CA2 and CA3. The arrangement of the genes is consistent with proposals that the duplication event which gave rise to CA1 predated the duplication which gave rise to CA2 and CA3. The order of these three genes differs from that suggested for the mouse based on recombination frequency.

Blotting, Southern

Gene duplication and concerted evolution of the GPDH locus in natural populations of Drosophila melanogaster.

The sn-glycerol-3-phosphate dehydrogenase (GPDH, EC 1, 1, 1, 8) locus of Drosophila melanogaster is polymorphic with respect to the number of tandemly duplicated genes in natural populations. The duplicated genes were cloned and the nucleotide sequences were determined. The duplication deletes both the first and second exons and has a size of 4500 b.p. The fact that there is no sequence variation at the junction point of the duplicated units among strains suggests a single origin for the duplication event. Comparison of the nucleotide sequences among the duplicates indicates that the frequent transfer of genetic information occurs from one to the other of the duplicates on the same chromosome either by gene conversion or by unequal crossing over. Because the GPDH duplication is partial and therefore a kind of pseudogene, the observed polymorphism of the number of tandemly duplicated GPDH genes appears to have been driven mainly by random genetic drift.

Animals

Identification of members of the P-glycoprotein multigene family.

Overproduction of P-glycoprotein is intimately associated with multidrug resistance. This protein appears to be encoded by a multigene family. Thus, differential expression of different members of this family may contribute to the complexity of the multidrug resistance phenotype. Three lambda genomic clones isolated from a hamster genomic library represent different members of the hamster P-glycoprotein gene family. Using a highly conserved exon probe, we found that the hamster P-glycoprotein gene family consists of three genes. We also found that the P-glycoprotein gene family consists of three genes in mice but has only two genes in humans and rhesus monkeys. The hamster P-glycoprotein genes have similar exon-intron organizations within the 3' region encoding the cytoplasmic domains. We propose that the hamster P-glycoprotein gene family arose from gene duplication. The hamster pgp1 and pgp2 genes appear to be more closely related to each other than either gene is to the pgp3 gene. We speculate that the hamster pgp1 and pgp2 genes arose from a recent gene duplication event and that primates did not undergo this duplication and therefore contain only two P-glycoprotein genes.

ATP Binding Cassette Transporter, Subfamily B, Mem

Parallel origins of duplications and the formation of pseudogenes in mitochondrial DNA from parthenogenetic lizards (Heteronotia binoei; Gekkonidae).

Analysis of mitochondrial DNAs (mtDNAs) from parthenogenetic lizards of the Heteronotia binoei complex with restriction enzymes revealed an approximately 5-kb addition present in all 77 individuals. Cleavage site mapping suggested the presence of a direct tandem duplication spanning the 16S and 12S rRNA genes, the control region and most, if not all, of the gene for the subunit 1 of NADH dehydrogenase (ND1). The location of the duplication was confirmed by Southern hybridization. A restriction enzyme survey provided evidence for modifications to each copy of the duplicated sequence, including four large deletions. Each gene affected by a deletion was complemented by an intact version in the other copy of the sequence, although for one gene the functional copy was heteroplasmic for another deletion. Sequencing of a fragment from one copy of the duplication which encompassed the tRNA(leu)(UUR) and parts of the 16S rRNA and ND1 genes, revealed mutations expected to disrupt function. Thus, evolution subsequent to the duplication event has resulted in mitochondrial pseudogenes. The presence of duplications in all of these parthenogens, but not among representatives of their maternal sexual ancestors, suggests that the duplications arose in the parthenogenetic form. This provides the second instance in H. binoei of mtDNA duplication associated with the transition from sexual to parthenogenetic reproduction. The increased incidence of duplications in parthenogenetic lizards may be caused by errors in mtDNA replication due to either polyploidy or hybridity of their nuclear genomes.

Amino Acid Sequence