PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Purifying selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

The origin of the Jingwei gene and the complex modular structure of its parental gene, yellow emperor, in Drosophila melanogaster.

Jingwei (jgw) is the first gene found to be of sufficiently recent origin in Drosophila to offer insights into the origin of a gene. While its chimerical gene structure was partially resolved as including a retrosequence of alcohol dehydrogenase (ADH:), the structure of its non-ADH: parental gene, the donor of the N-terminal domain of jgw, is unclear. We characterized this non-ADH: parental locus, yellow emperor (ymp), by cloning it, mapping it onto the polytene chromosomes, sequencing the entire locus, and examining its expression patterns in Drosophila melanogaster. We show that ymp is located in the 96-E region; the N-terminal domain of ymp has donated the non-ADH: portion of jgw via a duplication. The similar 5' portions of the gene and its regulatory sequences give rise to similar testis-specific expression patterns in ymp and jgw in Drosophila teissieri. Furthermore, between-species comparison of ymp revealed purifying selection in the protein sequence, suggesting a functional constraint in ymp. While the structure of ymp provides clear information for the molecular origin of the new gene jgw, it unexpectedly casts a new light on the concept of genes. We found, for the first time, that the single locus of the ymp gene encompasses three major molecular mechanisms determining structure of eukaryotic genes: (1) the 5' exons of ymp are involved in an exon-shuffling event that has created the portion recruited by jgw; (2) using alternative cleavage sites and alternative splicing sites, the 3' exon groups of ymp produce two proteins with nonhomologous C-terminal domains, both exclusively in the testis; and (3) in the opposite strand of the third intron of ymp is an essential gene, musashi (msi), which encodes an RNA-binding protein. The composite gene structure of ymp manifests the complexity of the gene concept, which should be considered in genomic research, e.g., gene finding.

Alternative Splicing↗

Phylogeny, rates of evolution, and patterns of codon usage among sea urchin retroviral-like elements, with implications for the recognition of horizontal transfer.

Phylogenetic relationships, rates of evolution, and codon usage were investigated in a family of retrotransposons (SURL elements) found in echinoids. The phylogeny of SURL element reverse transcriptase sequences from 10 echinoid species clearly shows the phylogenetic signature of the host taxa as well as paralogous sequences that diverged prior to speciation events. Two subfamilies (1 and 5) of SURL element reverse transcriptase sequences are recognized that diverged prior to the radiation of the Echinometridae. Comparisons of synonymous versus nonsynonymous substitutions indicate that SURL elements have been active in echinoid genomes and have evolved under purifying selection for millions of years. Rates of synonymous substitution for reverse transcriptase are similar to rates of single-copy DNA evolution and to rates of synonymous substitution for the H3 and H4 histone genes, contradicting the assumption that rates of evolution are accelerated in retrotransposons. Finally, codon usage in SURL elements is biased for codons ending in A or U relative to 42 sea urchin nuclear genes. Biased codon usage is sometimes cited as evidence for horizontal transfer, but in the case of SURL elements this bias occurs in spite of a long history of vertical transmission rather than because of horizontal transfer.

Animals↗

Origin and evolution of HLA class I pseudogenes.

The HLA (human major histocompatibility complex, or MHC) region includes three types of class I MHC genes: (1) class Ia loci (HLA-A, -B, and -C), which are highly expressed and polymorphic; (2) class Ib loci, which have much reduced expression and polymorphism; and (3) unexpressed class I pseudogenes. Phylogenetic analysis suggests that both class Ib loci and class I pseudogenes arose by duplication of class Ia loci. 5' regulatory elements conserved in class Ia genes are not conserved in either class Ib or class I pseudogenes, but the former show evidence of purifying selection at nonsynonymous sites in coding regions that is lacking in the latter. In the HLA class I region, there is little correspondence between map distance and phylogenetic distance, suggesting a complex history of tandem gene duplications. Furthermore, separate phylogenetic analysis of different gene regions suggests that several class I genes are evolutionary chaemeras, which presumably have arisen as a result of a process of interlocus recombination. For example, the 5' flanking region of the HLA-92 pseudogene was donated by a gene related to HLA-B and -C; and the HLA-A gene arose when exons 1-3 (and intervening introns) were donated to a gene related to the HLA-70 pseudogene by a gene related to HLA-B and -C.

Animals↗

Rapid evolution of goat and sheep globin genes following gene duplication.

Statistical analyses of DNA sequences of globin genes (beta A, beta C, and gamma) from goat and sheep (including new sequence information for the second intron of sheep beta A and gamma, kindly provided by A. Davis and A. W. Nienhuis) indicate that the rates of nonsynonymous substitution in these genes have been greatly accelerated following the gene duplication separating gamma and the ancestor of beta A and beta C and the gene duplication separating beta A and beta C. In both cases the acceleration was apparently due to relaxation of purifying selection (functional constraints) rather than advantageous mutations because acceleration occurred only in less important parts of the beta globin chain. The rates of nonsynonymous substitution in these genes are estimated to be about 2.3 x 10(-9) per site per year, which is three times higher than that for the divergence between human beta and mouse beta major globin genes. Our analyses further suggest that the rate of synonymous substitution in functional genes and the rate of substitution in pseudogenes are approximately equal and are between 2.8 x 10(-9) and 5.0 x 10(-9) and that the rate of substitution in introns is about 3.0 x 10(-9). The divergence time between beta A and beta C and that between gamma and the beta A-beta C pair are about 12 and 30 million years, respectively. The proportion of transition mutations is estimated to be 64%, two times higher than expected under random mutation but considerably lower than the 96% estimated for animal mitochondrial DNA.

Animals↗

Concerted evolution of the immunoglobulin VH gene family.

With the aim of understanding the concerted evolution of the immunoglobulin VH multigene family, a phylogenetic tree for the DNA sequences of 16 mouse and five human germ line genes was constructed. This tree indicates that all genes in this family have undergone substantial evolutionary divergence. The most closely related genes so far identified in the mouse genome seem to have diverged about 6 million years (MY) ago, whereas the most distantly related genes diverged about 300 MY ago. This suggests that gene duplication caused by unequal crossing-over or gene conversion occurs very slowly in this gene family. The rate of occurrence of gene duplication in the VH gene family has been estimated to be 5 x 10(-7) per gene per year, which seems to be at least about 100 times lower than that for the rRNA gene family. This low rate of concerted evolution in the VH gene family helps retain intergenic genetic variability that in turn contributes to antibody diversity. Because of accumulation of destructive mutations, however, about one-third of the mouse and human VH genes seem to have become nonfunctional. Many of these pseudogenes have apparently originated recently, but some of them seem to have existed in the genome for more than 10 MY. The rate of nucleotide substitution for the complementarity-determining regions (CDRs) is as high as that of pseudogenes. This suggests that there is virtually no purifying selection operating in the CDRs and that germ line mutations are effectively used for generating antibody diversity.

Amino Acid Sequence↗

Nearly identical allelic distributions of xanthine dehydrogenase in two populations of Drosophila pseudoobscura.

In a previous study, Keith (1983) showed by sequential gel electrophoresis of the esterase-5 protein in Drosophila pseudoobscura that a highly polymorphic locus with many alleles can have very similar frequency distributions in populations separated by 500 km. The present work studies another highly polymorphic locus, xanthine dehydrogenase, in the same California population samples, using the same technique to distinguish allelic classes. Twelve electromorphs were found in one population and 15 in the other. Both populations shared a single very frequent (approximately 60%) allele, as well as five other alleles in low but similar frequencies. In addition, each population had an array of unique alleles present only once in one population sample but absent in the other. A statistical test against the stationary distribution for neutral alleles shows that, if the populations are at equilibrium, then purifying selection is operating on xanthine dehydrogenase. The extremely close similarity in frequency distributions of the alleles between populations for both the xanthine dehydrogenase and esterase-5 loci, despite differences in allele frequency distribution between loci, strongly emphasizes the importance of migration in influencing genic diversity in these populations.

Alleles↗

Polymorphism and evolution of influenza A virus genes.

The nucleotide sequences of four genes of the influenza A virus (nonstructural protein, matrix protein, and a few subtypes of hemagglutinin and neuraminidase) are compiled for a large number of strains isolated from various locations and years, and the evolutionary relationship of the sequences is investigated. It is shown that all of these genes or subtypes are highly polymorphic and that the polymorphic sequences (alleles) are subject to rapid turnover in the population, their average age being much less than that of higher organisms. Phylogenetic analysis suggests that most polymorphic sequences within a subtype or a gene appeared during the last 80 years and that the divergence among the subtypes of hemagglutinin genes might have occurred during the last 300 years. The high degree of polymorphism in this RNA virus is caused by an extremely high rate of mutation, estimated to be 0.01/nucleotide site/year. Despite the high rate of mutation, most influenza virus genes are apparently subject to purifying selection, and the rate of nucleotide substitution is substantially lower than the mutation rate. There is considerable variation in the substitution rate among different genes, and the rate seems to be lower in nonhuman viral strains than in human strains. The difference might be responsible for the so-called freezing effect in some viral strains.

Amino Acid Sequence↗

Nucleotide sequence comparison of the rp49 gene region between Drosophila subobscura and D. melanogaster.

A 1.6-kb fragment encompassing the rp49 gene, which codes for a ribosomal protein, has been cloned and sequenced in Drosophila subobscura. The rp49 coding region has accumulated 46 nucleotide differences out of 402 bp since D. subobscura diverged from D. melanogaster. Forty-three percent of the effectively silent sites have changed since both species diverged. Both silent and replacement differences are distributed at random between the two exons of the gene. The frequency of silent differences in exons does not differ from that observed in the 5' leader sequence and in the intron. The frequency of silent differences in exon and intron sites is much greater than the number of amino acid replacement differences. This observation indicates strong purifying selection against amino acid replacements.

Animals↗

Molecular population genetics of Escherichia coli: DNA sequence diversity at the celC, crr, and gutB loci of natural isolates.

The DNA sequences of three genes--celC, crr, and gutB--have been determined for each of 11 or 12 natural isolates of Escherichia coli from the ECOR collection. These genes encode the phosphoenolpyruvate-dependent phosphotransferase-system enzyme III proteins specific for beta-glucoside sugars (celC), glucose (crr), and glucitol (gutB), respectively. There is little evidence of recombination at or among these loci; among these strains, relationships inferred from each gene are largely consistent with each other and with the relationship inferred from multilocus enzyme electrophoresis. DNA sequence diversity is similar for all three genes, particularly when silent (synonymous) sites only are considered. This is surprising because there is much stronger codon usage bias at crr than at celC or gutB. The extent of divergence in the protein sequences encoded by these three genes varies considerably. The constitutively expressed glucose-specific enzyme is completely conserved. It is surprising that the inducible glucitol-specific enzyme, which is functional, is more variable than the cellobiose-specific enzyme, which is cryptic; the latter might be expected to be under less (if any) purifying selection.

Amino Acid Sequence↗

Improvement of oxidative and thermostability of N-carbamyl-d-amino Acid amidohydrolase by directed evolution.

N-Carbamyl-D-amino acid amidohydrolase (N-carbamoylase), which is currently employed in the industrial production of unnatural D-amino acid in conjunction with D-hydantoinase, has low oxidative and thermostability. We attempted the simultaneous improvement of the oxidative and thermostability of N-carbamoylase from Agrobacterium tumefaciens NRRL B11291 by directed evolution using DNA shuffling. In a second generation of evolution, the best mutant 2S3 with improved oxidative and thermostability was selected, purified and characterized. The temperature at which 50% of the initial activity remains after incubation for 30 min was 73 degrees C for 2S3, whereas it was 61 degrees C for wild-type enzyme. Treatment of wild-type enzyme with 0.2 mM hydrogen peroxide for 30 min at 25 degrees C resulted in a complete loss of activity, but 2S3 retained about 79% of the initial activity under the same conditions. The K(m) value of 2S3 was estimated to be similar to that of wild-type enzyme; however k(cat) was decreased, leading to a slightly reduced value of k(cat)/K(m), compared with wild-type enzyme. DNA sequence analysis revealed that six amino acid residues were changed in 2S3 and substitutions included Q23L, V40A, H58Y, G75S, M184L and T262A. The stabilizing effects of each amino acid residue were investigated by incorporating mutations individually into wild-type enzyme. Q23L, H58Y, M184L and T262A were found to enhance both oxidative and thermostability of the enzyme and of them, T262A showed the most significant effect. V40A and G75S gave rise to an increase only in oxidative stability. The positions of the mutated amino acid residues were identified in the structure of N-carbamoylase from Agrobacterium sp. KNK 712 and structural analysis of the stabilizing effects of each amino acid substitution was also carried out.

Agrobacterium tumefaciens↗

Free fatty acids trigger apoptosis and inhibit cell cycle progression in human vascular endothelial cells.

Plasma free fatty acid (FFA) concentrations are increased in states of insulin resistance and impair endothelial function. Because the underlying mechanisms are largely unknown, we examined selected, purified FFAs' (100-300 micromol/l, 24-48 h) action on apoptosis, cell cycle distribution, and associated gene/protein expression in human umbilical vein endothelial cells (HUVECs). Stearic acid, but not oleic acid, time and concentration dependently increased endothelial apoptosis by fivefold (n=6, P<0.01), whereas polyunsaturated FFAs (PUFAs; linoleic, gamma-linolenic, and arachidonic acid) exerted proapoptotic activity only at 300 micromol/l (P<0.05). Proapoptotic FFA action increased with FFAs' number of double bonds and with protein expression of the apoptosis promotor bak. The G0/G1 cell cycle arrest (n=6, P<0.05) induced by stearic acid (+14%) and PUFAs (+30%) is reflected by up-regulation of p21(WAF-1/Cip1). In addition, all FFAs concentration dependently reduced (P<0.05) gene/protein expression of clusterin (-54%), NF-kappaB's inhibitor, IkappaBalpha (-50%), endothelin-1 (-44%), and endothelial NO synthase (-44%). Plasma samples obtained from individuals with elevated plasma FFAs (372+/-22 micromol/l) increased endothelial apoptosis by 4.2-fold (P<0.001, n=10) compared with intra-individually matched low plasma FFA (56+/-21 micromol/l) conditions, underlining the results obtained by defined FFA stimulation. In conclusion, FFA structure differently affects endothelial cell proliferation and apoptosis, both representing key factors in the development of micro- and macrovascular dysfunction.

Apoptosis↗

Sequence variation and phylogenetic history of the mouse Ahr gene.

The Ahr locus encodes for the aryl hydrocarbon receptor (AHR), which plays an important toxicological and developmental role. Sequence variation in this gene was studied in 13 different mouse lines that included eight laboratory strains, two Mus musculus subspecies and three additional Mus species. The data presented represent the largest study of sequence variation across multiple mouse lines in a single gene (approximately equal to 15.9 kb/mouse line). Among all mice, the average frequency of all polymorphisms in the intronic regions was 20.3 variants/kb and the average exonic frequency was 14.1 variants/kb. For substitutions alone, the average frequencies in the intronic and exonic regions for all mice were 13.3 and 8.9 substitutions/kb, respectively. Between laboratory strains, the average intronic and exonic frequencies for all polymorphisms dropped to 5.4 and 2.9 variants/kb, respectively. There were 111 non-synonymous polymorphisms that resulted in 42 different amino acid changes, of which only 10 amino acid changes had been previously identified. Based on the nucleotide sequence, the phylogenetic history of the gene showed mice from the Ahr(b2) and Ahr(d) alleles in separate branches while mice from the Ahr(b1) and Ahr(b3) alleles exhibited a more complex history. Evolutionarily, the AHR protein as a whole appears to be under purifying selective pressure (K(a) : K(s) ratio = 0.237). Despite significant functional constraint in the basic helix-loop-helix and PAS domains, ligand binding is not constrained to the high-affinity allele, which supports further the role of the AHR in development and its importance beyond the adaptive response to environmental toxicants.

Amino Acid Sequence↗

The evolution of functionally novel proteins after gene duplication.

A widely cited model of the evolution of functionally novel proteins (here called the model of mutation during non-functionality (MDN model)) holds that, after gene duplication, one gene copy is redundant and free to accumulate substitutions at random. By chance, some of these substitutions may suit the protein encoded by such a non-functional gene to a new function, which it can subsequently assume. Several lines of evidence contradict this hypothesis: (i) comparison of expressed duplicate genes from the tetraploid frog Xenopus laevis suggests that such genes are subject to purifying selection and are thus not free to accumulate substitutions at random; (ii) in a number of multi-gene families, there is now evidence that functionally distinct proteins have arisen not as a result of chance fixation of neutral variants but rather as a result of positive Darwinian selection; and (iii) the phenomenon of gene sharing, in which a single gene encodes a protein having two distinct functions, shows that gene duplication is not a necessary prerequisite to the evolution of a new protein function. A model for the evolution of new protein is proposed under which a period of gene sharing ordinarily precedes the evolution of functionally distinct proteins. Gene duplication then allows each daughter gene to specialize for one of the functions of the ancestral gene.(ABSTRACT TRUNCATED AT 250 WORDS)

Biological Evolution↗

Excessive homoplasy in an evolutionarily constrained protein.

The evolution of monomorphic proteins among closely related species has not been examined in detail. To investigate this phenomenon, the glycerol-3-phosphate dehydrogenase (Gpdh) locus was sequence in a broad range of Drosophila species. Although purifying selection to remove amino acid variation is the dominant force in the evolution of Gpdh, some replacements have occurred. The sequences were compared in the context of the phylogeny of the genus, revealing a high proportion of amino acid parallelism and reversal (homoplasy) at four sites. The level of homoplasy is significantly greater than that seen in other proteins for which multiple sequences are available, showing that Gpdh is strongly constrained by both the number of amino acid differences and the types of changes allowed. These four sites evolve at a much higher rate than do the other variable positions in the protein, accounting for half of the interspecific amino acid replacements. However, unlike typical hypervariable sites, where multiple changes to several different amino acids are seen, evolutionary 'flip-flopping' between two amino acid states defines this new class of hypervariable site.

Amino Acid Sequence↗

Müllerian mimicry: an examination of Fisher's theory of gradual evolutionary change.

In 1927, Fisher suggested that Müllerian mimicry evolution could be gradual and driven by predator generalization. A competing possibility is the so-called two-step hypothesis, entailing that Müllerian mimicry evolves through major mutational leaps of a less-protected species towards a better-protected, which sets the stage for coevolutionary fine-tuning of mimicry. At present, this hypothesis seems to be more widely accepted than Fisher's suggestion. We conducted individual-based simulations of communities with predators and two prey types to assess the possibility of Fisher's process leading to a common prey appearance. We found that Fisher's process worked for initially relatively similar appearances. Moreover, by introducing a predator spectrum consisting of several predator types with different ranges of generalization, we found that gradual evolution towards mimicry occurred also for large initial differences in prey appearance. We suggest that Fisher's process together with a predator spectrum is a realistic alternative to the two-step hypothesis and, furthermore, it has fewer problems with purifying selection. We also examined the factors influencing gradual evolution towards mimicry and found that not only the relative benefits from mimicry but also the mutational schemes of the prey types matter.

Adaptation, Biological↗

Acetylcholinesterase genes within the Diptera: takeover and loss in true flies.

It has recently been reported that the synaptic acetylcholinesterase (AChE) in mosquitoes is encoded by the ace-1 gene, distinct and divergent from the ace-2 gene, which performs this function in Drosophila. This is an unprecedented situation within the Diptera order because both ace genes derive from an old duplication and are present in most insects and arthropods. Nevertheless, Drosophila possesses only the ace-2 gene. Thus, a secondary loss occurred during the evolution of Diptera, implying a vital function switch from one gene (ace-1) to the other (ace-2). We sampled 78 species, representing 50 families (27% of the Dipteran families) spread over all major subdivisions of the Diptera, and looked for ace-1 and ace-2 by systematic PCR screening to determine which taxonomic groups within the Diptera have this gene change. We show that this loss probably extends to all true flies (or Cyclorrhapha), a large monophyletic group of the Diptera. We also show that ace-2 plays a non-detectable role in the synaptic AChE in a lower Diptera species, suggesting that it has non-synaptic functions. A relative molecular evolution rate test showed that the intensity of purifying selection on ace-2 sequences is constant across the Diptera, irrespective of the presence or absence of ace-1, confirming the evolutionary importance of non-synaptic functions for this gene. We discuss the evolutionary scenarios for the takeover of ace-2 and the loss of ace-1, taking into account our limited knowledge of non-synaptic functions of ace genes and some specific adaptations of true flies.

Acetylcholinesterase↗

Parallel genetic adaptation amid a background of changing effective population sizes in divergent yellow perch (Perca flavescens) populations.

Aquatic ecosystems are highly dynamic environments vulnerable to natural and anthropogenic disturbances. High-economic-value fisheries are one of many ecosystem services affected by these disturbances, and it is critical to accurately characterize the genetic diversity and effective population sizes of valuable fish stocks through time. We used genome-wide data to reconstruct the demographic histories of economically important yellow perch (Perca flavescens) populations. In two isolated and genetically divergent populations, we provide independent evidence for simultaneous increases in effective population sizes over both historic and contemporary time scales including negative genome-wide estimates of Tajima's D, 3.1 times more single nucleotide polymorphisms than adjacent populations, and contemporary effective population sizes that have increased 10- and 47-fold from their minimum, respectively. The excess of segregating sites and negative Tajima's D values probably arose from mutations accompanying historic population expansions with insufficient time for purifying selection, whereas linkage disequilibrium-based estimates of Ne also suggest contemporary increases that may have been driven by reduced fishing pressure or environmental remediation. We also identified parallel, genetic adaptation to reduced visual clarity in the same two habitats. These results suggest that the synchrony of key ecological and evolutionary processes can drive parallel demographic and evolutionary trajectories across independent populations.

Animals↗

Genetic variation in porcine reproductive and respiratory syndrome virus isolates in the midwestern United States.

The nucleotide sequence of a 3266 bp region encompassing open reading frames (ORFs) 2 through 7 of the porcine reproductive and respiratory syndrome virus (PRRSV) was determined for 10 isolates recovered from the midwestern United States. Pairwise comparisons showed that genetic distances between isolates ranged from 2.5% to 7.9% (mean 5.8% +/- 0.2%) whereas the Lelystad strain from Europe was, on average, 34.8% divergent from US clones. Thus, US and European PRRSV isolates represent genetically distinct clusters of the same virus. ORF 5, which encodes the envelope glycoprotein, was the most polymorphic [total nucleotide diversity (pi) = 0.097 +/- 0.007] and ORF 6, encoding the viral M protein, was the most conserved (pi = 0.038 +/- 0.003). The substantial differences in nucleotide diversity among ORFs suggests that the virus is evolving by processes other than simple accumulation of random neutral mutations. In support of this hypothesis, statistical analyses of the nucleotide sequence provided strong evidence for intragenic recombination or gene conversion in ORFs 2, 3, 4, 5 and 7, but not in ORF 6. An excess of synonymous (silent) substitutions was observed in all six ORFs, indicating an evolutionary pressure to conserve amino acid sequences. Taken together, the data indicate that despite intragenic recombination among extant PRRSV isolates, purifying selection has acted to maintain the primary structure of individual ORFs.

Animals↗