PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Purifying selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Mechanisms and rates of birth and death of dispersed duplicated genes during the evolution of a multigene family in diploid and tetraploid wheats.

A family of 5 genes that evolved within the past 1.9 Myr in diploid wheat was characterized. The ancestral gene, ALP-A1, is on chromosome 1A and encodes an aci-reductone dioxygenase-like protein. The duplicated genes ALP-A2, ALP-A3, ALP-A4.1, and ALP-A4.2 acquired complete coding sequences but lost the original promoter. They are on chromosomes 4A, 2A, 6A and 6A, respectively, and evolved sequentially, the youngest duplicated gene always producing the next duplicate. It is shown that dispersed gene duplication rate consists of the primary rate (duplications of ancestral genes) and the secondary rate (duplications of genes that had been generated by recent duplications). The primary rate was 2.5 x 10(-3) gene(-1) Myr(-1) in diploid wheat. The secondary rate was 5.2 x 10(-2) gene(-1) Myr(-1) in the ALP family. The 20-fold acceleration of the secondary rate was caused by the insertion of the ALP-A2 gene into a novel type transposon. Only the ALP-A1 and ALP-A3 genes are transcribed. The transcription of ALP-A3 is directed by a promoter within a DNA fragment similar to a CACTA type of DNA transposons, making ALP-A3 a new gene. The ALP-A3 transcript is longer than that of the ALP-A1. The half-life of ALP duplicated genes was estimated to be 0.87 Myr. Strong purifying selection acting on the ancestral gene ALP-A1 was undiminished by the evolution of duplicated genes. The evolution of the ALP family shows that repeated elements facilitate both gene duplication and expression of duplicated genes and highlights their importance for the evolution of gene repertoire in large plant genomes.

Amino Acid Sequence↗

Phylogenetic evidence for deleterious mutation load in RNA viruses and its contribution to viral evolution.

Populations of RNA viruses are often characterized by abundant genetic variation. However, the relative fitness of these mutations is largely unknown, although this information is central to our understanding of viral emergence, immune evasion, and drug resistance. Here we develop a phylogenetic method, based on the distribution of nonsynonymous and synonymous changes, to assess the relative fitness of polymorphisms in the structural genes of 143 RNA viruses. This reveals that a substantial proportion of the amino acid variation observed in natural populations of RNA viruses comprises transient deleterious mutations that are later purged by purifying selection, potentially limiting virus adaptability. We also demonstrate, for the first time, the existence of a relationship between amino acid variability and the phylogenetic distribution of polymorphisms. From this relationship, we propose an empirical threshold for the maximum viable deleterious mutation load in RNA viruses.

Amino Acid Sequence↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗

Evolutionary conservation suggests a regulatory function of AUG triplets in 5'-UTRs of eukaryotic genes.

By comparing sequences of human, mouse and rat orthologous genes, we show that in 5'-untranslated regions (5'-UTRs) of mammalian cDNAs but not in 3'-UTRs or coding sequences, AUG is conserved to a significantly greater extent than any of the other 63 nt triplets. This effect is likely to reflect, primarily, bona fide evolutionary conservation, rather than cDNA annotation artifacts, because the excess of conserved upstream AUGs (uAUGs) is seen in 5'-UTRs containing stop codons in-frame with the start AUG and many of the conserved AUGs are found in different frames, consistent with the location in authentic non-coding sequences. Altogether, conserved uAUGs are present in at least 20-30% of mammalian genes. Qualitatively similar results were obtained by comparison of orthologous genes from different species of the yeast genus Saccharomyces. Together with the observation that mammalian and yeast 5'-UTRs are significantly depleted in overall AUG content, these findings suggest that AUG triplets in 5'-UTRs are subject to the pressure of purifying selection in two opposite directions: the uAUGs that have no specific function tend to be deleterious and get eliminated during evolution, whereas those uAUGs that do serve a function are conserved. Most probably, the principal role of the conserved uAUGs is attenuation of translation at the initiation stage, which is often additionally regulated by alternative splicing in the mammalian 5'-UTRs. Consistent with this hypothesis, we found that open reading frames starting from conserved uAUGs are significantly shorter than those starting from non-conserved uAUGs, possibly, owing to selection for optimization of the level of attenuation.

5' Untranslated Regions↗

Molecular cloning and sequencing of cDNAs encoding three heavy-chain precursors of the inter-alpha-trypsin inhibitor in Syrian hamster: implications for the evolution of the inter-alpha-trypsin inhibitor heavy chain family.

Complementary DNAs encoding precursors of the three heavy chains (HC1, HC2, HC3) of the inter-alpha-trypsin inhibitor in Syrian hamster liver were sequenced. The deduced amino acid sequence of the HC1 precursor was 87, 82, and 79% identical with those of the HC1 precursors from mouse, man and pig, respectively. The HC2 and HC3 precursors showed similar degrees of sequence identity with the corresponding human and mouse HC precursors. When the hamster HC1 precursor was compared with its own HC2 and HC3 precursors, however, even the most highly conserved segment consisting of 565 amino acid residues, i.e., about 2/3 of the whole molecule, showed only about 35 and 65% sequence identity, respectively. Essentially the same results were obtained on the intra-species comparisons of three subfamilies in man and mouse. Thus, the interspecies conservation of a given HC subfamily is much greater than the similarity between the three different HC subfamilies within a given species. These results suggest that (i) higher vertebrates possess three HC genes which have been evolving independently of each other under purifying selection; (ii) the diversification of the three HC subfamilies, for which the middle regions of the molecules were mainly responsible, occurred before eutherian radiation; and (iii) each HC subfamily may have unique function(s), although at present virtually nothing is known about the functional differences between the three HC subfamilies.

Alpha-Globulins↗

The evolution of Ty1-copia group retrotransposons in gymnosperms.

A diverse collection of Ty1-copia group retrotransposons has been characterized from the genome of Picea abies (Norway spruce) by degenerate PCR amplification of a region of the reverse transcriptase gene. The occurrence of these retrotransposable elements in the gymnosperms was investigated by Southern blot hybridization analysis. The distribution of the different retrotransposons across the gymnosperms varies greatly. All of the retrotransposon clones isolated are highly conserved within the Picea (spruce) genus, many are also present in Pinus (pine) and/or Abies (fir) genera, and some share strongly homologous sequences with one or more of cedar, larch, Sequoia, cypress, and Ginkgo. Further subclones of one of the most strongly conserved retrotransposon sequences, Tpa28, were obtained from Ginkgo and P. abies. Comparisons of individual sequence pairs between the two species show nucleotide cross-homologies of around 80%-85%, corresponding to nucleotide substitution rates similar to those of nuclear protein-coding genes. Analysis of Tpa28 consensus sequences reveals that strong purifying selection has acted on this retrotransposon in the lineages connecting Ginkgo and Picea. Collectively, these data suggest, first, that the evolution of the Ty1-copia retrotransposon group in the gymnosperms is dominated by germ line vertical transmission, with strong selection for reverse transcriptase sequence, and, second, that extinction of individual retrotransposon types has been comparatively rare in gymnosperm species lineages compared with angiosperms. If this very high level of sequence conservation is a general property of the retrotransposons, then their extreme sequence diversity implies that they are extremely ancient, and the major element lineages seen today may have arisen early in eukaryote evolution. The data are also consistent with horizontal transmission of particular retrotransposons between species, but such a mechanism is unnecessary to explain the results.

Amino Acid Sequence↗

Dating the origin of the African human T-cell lymphotropic virus type-i (HTLV-I) subtypes.

To investigate the origin of the African PTLV-I virus, we phylogenetically analyzed the available HTLV-I and STLV-I strains. We also attempted to date the presumed interspecies transmissions that resulted in the African HTLV-I subtypes. Molecular-clock analysis was performed using the Tamura-Nei substitution model and gamma distributed rate heterogeneity based on the maximum-likelihood topology of the combined long-terminal-repeat and env third-codon-position sequences. Since the molecular clock was not rejected and no evidence for saturation was found, a constant rate of evolution at these positions for all 33 HTLV-I and STLV-I strains was reasonably assumed. The spread of PTLV-I in Africa is estimated to have occurred at least 27,300 +/- 8,200 years ago. Using the available strains, the HTLV-If subtype appears to have emerged within the last 3,000 years, and the HTLV-Ia, HTLV-Ib, HTLV-Id, and HTLV-Ie subtypes appear to have diverged between 21,100 and 5,300 years ago. Interspecies transmissions, most probably simian to human, must have occurred around that time and probably continued later. When the synonymous and nonsynonymous substitution ratios were compared, it was clear that purifying selection was the driving force for PTLV-I evolution in the env gene, irrespective of the host species. Due to the small number of strains in some of the investigated groups, these data on selective pressure should be taken with caution.

Africa↗

Gene genealogies, cryptic species, and molecular evolution in the human pathogen Coccidioides immitis and relatives (Ascomycota, Onygenales).

Previous genealogical analyses of population structure in Coccidioides immitis revealed the presence of two cryptic and sexual species in this pathogenic fungus but did not clarify their origin and relationships with respect to other taxa. By combining the C. immitis data with those of two of its closest relatives, the free-living saprophytes Auxarthron zuffianum and Uncinocarpus reesii, we show that the C. immitis species complex is monophyletic, indicating a single origin of pathogenicity. Cryptic species also were found in both A. zuffianum and U. reesii, indicating that they can be found in both pathogenic and free-living fungi. Our study, together with a few others, indicates that the current list of known fungal species might be augmented by a factor of at least two. However, at least in the C. immitis, A. zuffianum, and U. reesii complexes, cryptic species represent subdivisions at the tips of deep monophyletic clades and thus well within the existing framework of generic classification. An analysis of silent and expressed divergence and polymorphism values between and within the taxa identified by genealogical concordance did not reveal faster evolution in C. immitis as a consequence of adaptation to the pathogenic habit, nor did it show positive Darwinian evolution in a region of a dioxygenase gene (tcrP gene coding for 4-HPPD) known to cause antigenic responses in humans. Instead, the data suggested relative stasis, indicative of purifying selection against mostly deleterious mutations. Two introns in the same gene fragment were considerably more divergent than exons and were unalignable between species complexes but had very low polymorphism within taxa.

4-Hydroxyphenylpyruvate Dioxygenase↗

Whole chloroplast genome comparison of rice, maize, and wheat: implications for chloroplast gene diversification and phylogeny of cereals.

The fully sequenced chloroplast genomes of maize (subfamily Panicoideae), rice (subfamily Bambusoideae), and wheat (subfamily Pooideae) provide the unique opportunity to investigate the evolution of chloroplast genes and genomes in the grass family (Poaceae) by whole-genome comparison. Analyses of nucleotide sequence variations in 106 cereal chloroplast genes with tobacco sequences as the outgroup suggested that (1) most of the genic regions of the chloroplast genomes of maize, rice, and wheat have evolved at similar rates; (2) RNA genes have highly conservative evolutionary rates relative to the other genes; (3) photosynthetic genes have been under strong purifying selection; (4) between the three cereals, 14 genes which account for about 28% of the genic region have evolved with heterogeneous nucleotide substitution rates; and (5) rice genes tend to have evolved more slowly than the others at loci where rate heterogeneity exists. Although the mechanism that underlies chloroplast gene diversification is complex, our analyses identified variation in nonsynonymous substitution rates as a genetic force that generates heterogeneity, which is evidence of selection in chloroplast gene diversification at the intrafamilial level. Phylogenetic trees constructed with the variable nucleotide sites of the chloroplast genes place maize basal to the rice-wheat clade, revealing a close relationship between the Bambusoideae and Pooideae.

Chloroplasts↗

Rapid divergence of gene duplicates on the Drosophila melanogaster X chromosome.

The recent sequencing of several eukaryotic genomes has generated considerable interest in the study of gene duplication events. The classical model of duplicate gene evolution is that recurrent mutation ultimately results in one copy becoming a pseudogene, and only rarely will a beneficial new function evolve. Here, we study divergence between coding sequence duplications in Drosophila melanogaster as a function of the linkage relationship between paralogs. The mean K(a)/K(s) between all duplicates in the D. melanogaster genome is 0.2803, indicating that purifying selection is maintaining the structure of duplicate coding sequences. However, the mean K(a)/K(s) between duplicates that are both on the X chromosome is 0.4701, significantly higher than the genome average. Further, the distribution of K(a)/K(s) for these X-linked duplicates is significantly shifted toward higher values when compared with the distributions for paralogs in other linkage relationships. Two models of molecular evolution provide qualitative explanations of these observations-relaxation of selective pressure on the duplicate copies and, more likely, positive selection on recessive adaptations. We also show that there is an excess of X-linked duplicates with low K(s), suggesting a larger proportion of relatively young duplicates on the D. melanogaster X chromosome relative to autosomes.

Animals↗

Color vision of ancestral organisms of higher primates.

The color vision of mammals is controlled by photosensitive proteins called opsins. Most mammals have dichromatic color vision, but hominoids and Old World (OW) monkeys enjoy trichromatic vision, having the blue-, green-, and red-sensitive opsin genes. Most New World (NW) monkeys are either dichromatic or trichromatic, depending on the sex and genotype. Trichromacy in higher primates is believed to have evolved to facilitate the detection of yellow and red fruits against dappled foliage, but the process of evolutionary change from dichromacy to trichromacy is not well understood. Using the parsimony and the newly developed Bayesian methods, we inferred the amino acid sequences of opsins of ancestral organisms of higher primates. The results suggest that the ancestors of OW and NW monkeys lacked the green gene and that the green gene later evolved from the red gene. The fact that the red/green opsin gene has survived the long nocturnal stage of mammalian evolution and that it is under strong purifying selection in organisms that live in dark environments suggests that this gene has another important function in addition to color vision, probably the control of circadian rhythms.

Amino Acid Sequence↗

The origin and evolution of Ebola and Marburg viruses.

Molecular evolutionary analyses for Ebola and Marburg viruses were conducted with the aim of elucidating evolutionary features of these viruses. In particular, the rate of nonsynonymous substitutions for the glycoprotein gene of Ebola virus was estimated to be, on the average, 3.6 x 10(-5) per site per year. Marburg virus was also suggested to be evolving at a similar rate. Those rates were a hundred times slower than those of retroviruses and human influenza A virus, but were of the same order of magnitude as that of the hepatitis B virus. When these rates were applied to the degree of sequence divergence, the divergence time between Ebola and Marburg viruses was estimated to be more than several thousand years ago. Moreover, most of the nucleotide substitutions were transitions and synonymous for Marburg virus. This suggests that purifying selection has operated on Marburg virus during evolution.

DNA, Viral↗

Mitochondrial DNA and bindin gene sequence evolution among allopatric species of the sea urchin genus Arbacia.

Sea urchins of the genus Arbacia (order Stirodonta) have discontinuous allopatric distributions ranging over thousands of kilometers. Mitochondrial DNA (mtDNA) sequences were used to reconstruct phylogenetic relationships of four Arbacia species and their geographic populations. There is little evidence of genetic structuring of populations within species, except in two cases at range extremes. The mtDNA sequence differentiation between species suggests that divergence occurred about 4-9 MYA. Gene sequences encoding the sperm protein bindin and its intron were obtained and compared with the mtDNA phylogeny. Sea urchins among the well-studied echinoid order Camarodonta, with degrees of mtDNA divergence similar to those of Arbacia species, are known to have remarkable variation in bindin. However, in Arbacia, little variation in deduced amino acid sequences of bindin was found, indicating that purifying selection acts on the protein. In contrast, bindin intron sequences showed much differentiation, including numerous insertion/deletions. Fertilization experiments performed between a divergent pair of Arbacia species from the Atlantic and Pacific Oceans revealed no evidence of blocks to gamete recognition. In Arbacia, fertilization specificities may have evolved relatively slowly as a result of extensive gene flow within species, greater functional constraint on the bindin polypeptide, or reduced selective pressure for species recognition in singly occurring species.

Amino Acid Sequence↗

The rate heterogeneity of nonsynonymous substitutions in mammalian mitochondrial genes.

Substitution rates at the three codon positions (r1, r2, and r3) of mammalian mitochondrial genes are in the order of r3 > r1 > r2, and the rate heterogeneity at the three positions, as measured by the shape parameter of the gamma distribution (alpha 1, alpha 2, and alpha 3), is in the order of alpha 3 > alpha 1 > alpha 2. The causes for the rate heterogeneity at the three codon positions remain unclear and, in particular, there has been no satisfactory explanation for the observation of alpha 1 > alpha 2. I attempted to dissect the causes of rate heterogeneity by studying the pattern of nonsynonymous substitutions with respect to codon positions in 10 mitochondrial genes from 19 mammalian species. Nonsynonymous substitutions involve more different amino acid replacements at the second than at the first codon position, which results in r1 > r2. The difference between r1 and r2 increases with the intensity of purifying selection, and so does the rate heterogeneity in nonsynonymous substitutions among sites at the same codon position. All mitochondrial genes appear to have functionally important and unimportant codons, with the latter having all three codon positions prone to nonsynonymous substitutions. Within the functionally important codons, the second codon position is much more conservative than the codon position. This explains why alpha 1 > alpha 2. The result suggests that overweighting of the second codon position in phylogenetic analysis may be a misguided practice.

Animals↗

The core domain of retrotransposon integrase in Hordeum: predicted structure and evolution.

Propagation of long terminal repeat (LTR)-bearing retrotransposons and retroviruses requires integrase (IN, EC 2.7.7.-), encoded by the retroelements themselves, which mediates the insertion of cDNA copies back into the genome. An active retrotransposon family, BARE-1, comprises approximately 7% of the barley (Hordeum vulgare subsp. vulgare) genome. We have generated models for the secondary and tertiary structure of BARE-1 IN and demonstrate their similarity to structures for human immunodeficiency virus 1 and avian sarcoma virus INs. The IN core domains were compared for 80 clones from 28 Hordeum accessions representative of the diversity of the genus. Based on the structural model, variations in the predicted, aligned translations from these clones would have minimal structural and functional effects on the encoded enzymes. This indicates that Hordeum retrotransposon IN has been under purifying selection to maintain a structure typical of retroviral INs. These represent the first such analyses for plant INs.

Amino Acid Sequence↗

Evolution of sea urchin retroviral-like (SURL) elements: evidence from 40 echinoid species.

We conducted a phylogenetic survey of sea urchin retroviral-like (SURL) retrotransposable elements in 33 species of the class Echinoidea (sea urchins, sand dollars, and heart urchins). A 263-bp fragment from the coding region of the reverse transcriptase (RT) gene was amplified, cloned, and sequenced. Phylogenetic relationships of the elements isolated from independent clones, along with those from seven additional echinoid species obtained earlier by Springer et al., were compared with host phylogeny. Vertical transmission and the presence of paralogous sequences that diverged prior to host speciation can explain most of the phylogenetic relationships among SURL elements. Rates of evolution were estimated from cases in which SURL and host phylogenies were concordant. In agreement with conclusions reached previously by Springer et al., average rates of synonymous substitution were comparable with those of single-copy sea urchin DNA. High ratios of synonymous to nonsynonymous substitution suggest that the RT of the elements is under strong purifying selection. However, a high proportion (approximately 15%) of elements with deleterious frameshifts and stop codons and an increase of the ratio of synonymous to nonsynonymous substitutions with divergence time show that in the short term this selection is relaxed. Despite the predominance of vertical transmission, sequence similarity of 83%-94% for SURL elements from hosts that have been separated for 200 Myr suggests four cases of apparent horizontal transfer between the ancestors of the extant echinoid species. In three additional cases, elements with identical RT sequences were found in sea urchin species separated for a minimum of 3 Myr. Thus, horizontal transfer plays a role in the evolution of this retrotransposon family.

Animals↗

Sequence evolution of the CCR5 chemokine receptor gene in primates.

The chemokine receptor CCR5 can serve as a coreceptor for M-tropic HIV-1 infection and both M-tropic and T-tropic SIV infection. We sequenced the entire CCR5 gene from 10 nonhuman primates: Pongo pygmaeus, Hylobates leucogenys, Trachypithecus francoisi, Trachypithecus phayrei, Pygathrix nemaeus, Rhinopithecus roxellanae, Rhinopithecus bieti, Rhinopithecus avunculus, Macaca assamensis, and Macaca arctoides. When compared with CCR5 sequences from humans and other primates, our results demonstrate that: (1) nucleotide and amino acid sequences of CCR5 among primates are highly homologous, with variations slightly concentrated on the amino and carboxyl termini; and (2) site Asp13, which is critical for CD4-independent binding of SIV gp120 to Macaca mulatta CCR5, was also present in all other nonhuman primates tested here, suggesting that those nonhuman primate CCR5s might also bind SIV gp120 without the presence of CD4. The topologies of CCR5 gene trees constructed here conflict with the putative opinion that the snub-nosed langurs compose a monophyletic group, suggesting that the CCR5 gene may not be a good genetic marker for low-level phylogenetic analysis. The evolutionary rate of CCR5 was calculated, and our results suggest a slowdown in primates after they diverged from rodents. The synonymous mutation rate of CCR5 in primates is constant, about 1.1 x 10(-9) synonymous mutations per site per year. Comparisons of Ka and Ks suggest that the CCR5 genes have undergone negative or purifying selection. Ka/Ks ratios from cercopithecines and colobines are significantly different, implying that selective pressures have played different roles in the two lineages.

Animals↗

Why mitochondrial genes are most often found in nuclei.

A very small fraction of the proteins required for the propagation and function of mitochondria are coded by their genomes, while nuclear genes code the vast majority. We studied the migration of genes between the two genomes when transfer mechanisms mediate this exchange. We could calculate the influence of differential mutation rates, as well as that of biased transfer rates, on the partitioning of genes between the two genomes. We observe no significant difference in partitioning for haploid and diploid cell populations, but the effective size of cell populations is important. For infinitely large effective populations, higher mutation rates in mitochondria than in nuclear genomes are required to drive mitochondrial genes to the nuclear genome. In the more realistic case of finite populations, gene transfer favoring the nucleus and/or higher mutation rates in the mitochondrion will drive mitochondrial genes to the nucleus. We summarize experimental data that identify a gene transfer process mediated by vacuoles that favors the accumulation of mitochondrial genes in the nuclei of modern cells. Finally, we compare the behavior of mitochondrial genes for which transfer to the nucleus is neutral or influenced by purifying selection.

Cell Nucleus↗