PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Microcystin biosynthesis in planktothrix: genes, evolution, and manipulation.

Microcystins represent an extraordinarily large family of cyclic heptapeptide toxins that are nonribosomally synthesized by various cyanobacteria. Microcystins specifically inhibit the eukaryotic protein phosphatases 1 and 2A. Their outstanding variability makes them particularly useful for studies on the evolution of structure-function relationships in peptide synthetases and their genes. Analyses of microcystin synthetase genes provide valuable clues for the potential and limits of combinatorial biosynthesis. We have sequenced and analyzed 55.6 kb of the potential microcystin synthetase gene (mcy) cluster from the filamentous cyanobacterium Planktothrix agardhii CYA 126. The cluster contains genes for peptide synthetases (mcyABC), polyketide synthases (PKSs; mcyD), chimeric enzymes composed of peptide synthetase and PKS modules (mcyEG), a putative thioesterase (mcyT), a putative ABC transporter (mcyH), and a putative peptide-modifying enzyme (mcyJ). The gene content and arrangement and the sequence of specific domains in the gene products differ from those of the mcy cluster in Microcystis, a unicellular cyanobacterium. The data suggest an evolution of mcy clusters from, rather than to, genes for nodularin (a related pentapeptide) biosynthesis. Our data do not support the idea of horizontal gene transfer of complete mcy gene clusters between the genera. We have established a protocol for stable genetic transformation of Planktothrix, a genus that is characterized by multicellular filaments exhibiting continuous motility. Targeted mutation of mcyJ revealed its function as a gene coding for a O-methyltransferase. The mutant cells produce a novel microcystin variant exhibiting reduced inhibitory activity toward protein phosphatases.

Amino Acid Sequence↗

The limits of selection during maize domestication.

The domestication of all major crop plants occurred during a brief period in human history about 10,000 years ago. During this time, ancient agriculturalists selected seed of preferred forms and culled out seed of undesirable types to produce each subsequent generation. Consequently, favoured alleles at genes controlling traits of interest increased in frequency, ultimately reaching fixation. When selection is strong, domestication has the potential to drastically reduce genetic diversity in a crop. To understand the impact of selection during maize domestication, we examined nucleotide polymorphism in teosinte branched1, a gene involved in maize evolution. Here we show that the effects of selection were limited to the gene's regulatory region and cannot be detected in the protein-coding region. Although selection was apparently strong, high rates of recombination and a prolonged domestication period probably limited its effects. Our results help to explain why maize is such a variable crop. They also suggest that maize domestication required hundreds of years, and confirm previous evidence that maize was domesticated from Balsas teosinte of southwestern Mexico.

Agriculture↗

Transposition and exon shuffling by group II intron RNA molecules in pieces.

In the realms of RNA, transposable elements created by self-inserting introns recombine novel combinations of exon sequences in the background of replicating molecules. Although intermolecular RNA recombination is a wide-spread phenomenon reported for a variety of RNA-containing viruses, direct evidence to support the theory that modern splicing systems, together with the exon-intron structure, have evolved from the ability of RNA to recombine, is lacking. Here, we used an in vitro deletion-complementation assay to demonstrate trans-activation of forward and reverse self-splicing of a fragmented derivative of the group II intron bI1 from yeast mitochondria. We provide direct evidence for the functional interchangeability of analogous but non-identical domain 1 RNA molecules of group II introns that result in trans-activation of intron transposition and RNA-based exon shuffling. The data extend theories on intron evolution and raise the intriguing possibility that naturally fragmented group III and spliceosomal introns themselves can create transposons, permitting rapid evolution of protein-coding sequences by splicing reactions.

Base Sequence↗

The monomeric and dimeric mannose-binding proteins from the Orchidaceae species Listera ovata and Epipactis helleborine: sequence homologies and differences in biological activities.

The Orchidaceae species Listera ovata and Epipactis helleborine contain two types of mannose-binding proteins. Using a combination of affinity chromatography on mannose-Sepharose-4B and ion exchange chromatography on a Mono-S column eight different mannose-binding proteins were isolated from the leaves of Listera ovata. Whereas seven of these mannose-binding proteins have agglutination activity and occur as dimers composed of lectin subunits of 11-13 kDa, the eighth mannose-binding protein is a monomer of 14 kDa devoid of agglutination activity. Moreover, the monomeric mannose-binding protein does not react with an antiserum raised against the dimeric lectin and, in contrast to the lectins, is completely inactive when tested for antiretroviral activity against human immunodeficiency virus type 1 and type 2. Mannose-binding proteins with similar properties were also found in the leaves of Epipactis helleborine. However, in contrast to Listera only one lectin was found in Epipactis. Despite the obvious differences in molecular structure and biological activities molecular cloning of different mannose-binding proteins from Listera and Epipactis has shown that these proteins are related and some parts of the sequences show a high degree of sequence homology indicating that they have been conserved through evolution.

Amino Acid Sequence↗

Gene tree parsimony vs uninode coding for phylogenetic reconstruction.

have suggested that there are important weaknesses of gene tree parsimony in reconstructing phylogeny in the face of gene duplication, weaknesses that are addressed by method of uninode coding. Here, we discuss Simmons and Freudenstein's criticisms and suggest a number of reasons why gene tree parsimony is preferable to uninode coding. During this discussion we introduce a number of recent developments of gene tree parsimony methods overlooked by Simmons and Freudenstein. Finally, we present a re-analysis of data from that produces a more reasonable phylogeny than that found by Simmons and Freudenstein, suggesting that gene tree parsimony outperforms uninode coding, at least on these data.

Algorithms↗

A gamma mixture model better accounts for among site rate heterogeneity.

MOTIVATION: Variation of substitution rates across nucleotide and amino acid sites has long been recognized as a characteristic of molecular sequence evolution. Evolutionary models that account for this rate heterogeneity usually use a gamma density function to model the rate distribution across sites. This density function, however, may not fit real datasets, especially when there is a multimodal distribution of rates. Here, we present a novel evolutionary model based on a mixture of gamma density functions. This model better describes the among-site rate variation characteristic of molecular sequence evolution. The use of this model may improve the accuracy of various phylogenetic methods, such as reconstructing phylogenetic trees, dating divergence events, inferring ancestral sequences and detecting conserved sites in proteins. RESULTS: Using diverse sets of protein sequences we show that the gamma mixture model better describes the stochastic process underlying protein evolution. We show that the proposed gamma mixture model fits protein datasets significantly better than the single-gamma model in 9 out of 10 datasets tested. We further show that using the gamma mixture model improves the accuracy of model-based prediction of conserved residues in proteins. AVAILABILITY: C++ source codes are available from the authors upon request.

Chromosome Mapping↗

A test of neutral molecular evolution based on nucleotide data.

The neutral theory of molecular evolution predicts that regions of the genome that evolve at high rates, as revealed by interspecific DNA sequence comparisons, will also exhibit high levels of polymorphism within species. We present here a conservative statistical test of this prediction based on a constant-rate neutral model. The test requires data from an interspecific comparison of at least two regions of the genome and data on levels of intraspecific polymorphism in the same regions from at least one species. The model is rejected for data from the region encompassing the Adh locus and the 5' flanking sequence of Drosophila melanogaster and Drosophila sechellia. The data depart from the model in a direction that is consistent with the presence of balanced polymorphism in the coding region.

Alcohol Dehydrogenase↗

Molecular evolution between Drosophila melanogaster and D. simulans: reduced codon bias, faster rates of amino acid substitution, and larger proteins in D. melanogaster.

Both natural selection and mutational biases contribute to variation in codon usage bias within Drosophila species. This study addresses the cause of codon bias differences between the sibling species, Drosophila melanogaster and D. simulans. Under a model of mutation-selection-drift, variation in mutational processes between species predicts greater base composition differences in neutrally evolving regions than in highly biased genes. Variation in selection intensity, however, predicts larger base composition differences in highly biased loci. Greater differences in the G+C content of 34 coding regions than 46 intron sequences between D. melanogaster and D. simulans suggest that D. melanogaster has undergone a reduction in selection intensity for codon bias. Computer simulations suggest at least a fivefold reduction in Nes at silent sites in this lineage. Other classes of molecular change show lineage effects between these species. Rates of amino acid substitution are higher in the D. melanogaster lineage than in D. simulans in 14 genes for which outgroup sequences are available. Surprisingly, protein sizes are larger in D. melanogaster than in D. simulans in the 34 genes compared between the two species. A substantial fraction of silent, replacement, and insertion/deletion mutations in coding regions may be weakly selected in Drosophila.

Amino Acids↗

An evolutionary model for protein-coding regions with conserved RNA structure.

Here we present a model of nucleotide substitution in protein-coding regions that also encode the formation of conserved RNA structures. In such regions, apparent evolutionary context dependencies exist, both between nucleotides occupying the same codon and between nucleotides forming a base pair in the RNA structure. The overlap of these fundamental dependencies is sufficient to cause "contagious" context dependencies which cascade across many nucleotide sites. Such large-scale dependencies challenge the use of traditional phylogenetic models in evolutionary inference because they explicitly assume evolutionary independence between short nucleotide tuples. In our model we address this by replacing context dependencies within codons by annotation-specific heterogeneity in the substitution process. Through a general procedure, we fragment the alignment into sets of short nucleotide tuples based on both the protein coding and the structural annotation. These individual tuples are assumed to evolve independently, and the different tuple sets are assigned different annotation-specific substitution models shared between their members. This allows us to build a composite model of the substitution process from components of traditional phylogenetic models. We applied this to a data set of full-genome sequences from the hepatitis C virus where five RNA structures are mapped within the coding region. This allowed us to partition the effects of selection on different structural elements and to test various hypotheses concerning the relation of these effects. Of particular interest, we found evidence of a functional role of loop and bulge regions, as these were shown to evolve according to a different and more constrained selective regime than the nonpairing regions outside the RNA structures. Other potential applications of the model include comparative RNA structure prediction in coding regions and RNA virus phylogenetics.

Base Sequence↗

Molecular basis and evolutionary origins of color diversity in great star coral Montastraea cavernosa (Scleractinia: Faviida).

Natural pigments are normally products of complex biosynthesis pathways where many different enzymes are involved. Corals and related organisms of class Anthozoa represent the only known exception: in these organisms, each of the host-tissue colors is essentially determined by a sequence of a single protein, homologous to the green fluorescent protein (GFP) from Aequorea victoria. This direct sequence-color linkage provides unique opportunity for color evolution studies. We previously reported the general phylogenetic analysis of GFP-like proteins, which suggested that the present-day diversity of reef colors originated relatively recently and independently within several lineages. The present work was done to get insight into the mechanisms that gave rise to this diversity. Three colonies of the great star coral Montastraea cavernosa (Scleractinia, Faviida) were studied, representing distinct color morphs. Unexpectedly, these specimens were found to express the same collection of GFP-like proteins, produced by at least four, and possibly up to seven, different genetic loci. These genes code for three basic colors-cyan, green, and red-and are expressed differently relative to one another in different morphs. Phylogenetic analysis of the new sequences indicated that the three major gene lineages diverged before separation of some coral families. Our results suggest that color variation in M. cavernosa is not a true polymorphism, but rather a manifestation of phenotypic plasticity (polyphenism). The family level depth of its evolutionary roots indicates that the color diversity is adaptively significant. Relative roles of gene duplication, gene conversion, and point mutations in its evolution are discussed.

Animals↗

Evolution of the albumin: alpha-fetoprotein ancestral gene from the amplification of a 27 nucleotide sequence.

The genes for alpha-fetoprotein and albumin arose by duplication of an ancestral gene that contained three genetic domains. These domains were generated by the triplication of a primordial genetic domain composed of five exons or subdomains. That the primordial domain itself arose by amplification of a simpler sequence is suggested by nucleotide sequence homologies among the subdomains of the mouse alpha-fetoprotein gene. A detailed analysis of these homologies reveals that each of the five subdomain families contains remnants of a 27-base-long repeat from which the entire alpha-fetoprotein coding sequence has been assembled. A consensus sequence for the 27 nucleotide repeat is derived, and the positions of the repeats within each subdomain are described. A model is proposed for the evolution of the primordial domain by the amplification and divergence of the 27 base-pair sequence, along with the condensation of the repeats into subdomains separated by intervening sequences. It is postulated that the role of intervening sequences may be to limit sequence amplification in genes such as alpha-fetoprotein and albumin whose protein products cannot tolerate size variation.

Albumins↗

The genome of herpes simplex virus: structure, replication and evolution.

The objectives of this paper are to discuss the structure and genetic content of the genome of herpes simplex virus type 1 (HSV-1), the nature of virus DNA replicative processes, and aspects of the evolution of the virus DNA, in particular those bearing on DNA replication. We are in the late stages of determining the complete sequence of the DNA of HSV-1, which contains about 155,000 base pairs, and thus the treatment is primarily from a viewpoint of DNA sequence and organization. The genome possesses around 75 genes, generally densely arranged and without long range ordering. Introns are present in only a few genes. Protein coding sequences have been predicted, and the functions of the proteins are being pursued by various means, including use of existing genetic and biochemical data, computer based analyses, expression of isolated genes and use of oligopeptide antisera. Many proteins are known to be virion structural components, or to have regulatory roles, or to function in synthesis of virus DNA. Many, however, still lack an assigned function. Two classes of genetic entities necessary for virus DNA replication have been characterized: cis-acting sequences, which include origins of replication and packaging signals, and genes encoding proteins involved in replication. Aside from enzymes of nucleotide metabolism, the latter include DNA polymerase, DNA binding proteins, and five species detected by genetic assays, but of presently unknown functions. Complete genome sequences are now known for the related alphaherpesvirus varicella-zoster virus and for the very distinct gammaherpesvirus Epstein-Barr virus. Comparisons between the three sequences show various homologies, and also several types of divergence and rearrangement, and so allow models to be proposed for possible events in the evolution of present day herpesvirus genomes. Another aspect of genome evolution is seen in the wide range of overall base compositions found in present day herpesvirus DNAs. Finally, certain herpesvirus genes are homologous to non-herpesvirus genes, giving a glimpse of more remote relationships.

Base Sequence↗

Serial SimCoal: a population genetics model for data from multiple populations and points in time.

UNLABELLED: We present Serial SimCoal, a program that models population genetic data from multiple time points, as with ancient DNA data. An extension of SIMCOAL, it also allows simultaneous modeling of complex demographic histories, and migration between multiple populations. Further, we incorporate a statistical package to calculate relevant summary statistics, which, for the first time allows users to investigate the statistical power provided by, conduct hypothesis-testing with, and explore sample size limitations of ancient DNA data. AVAILABILITY: Source code and Windows/Mac executables at http://www.stanford.edu/group/hadlylab/ssc.html CONTACT: senka@stanford.edu.

Biological Evolution↗

Localization and expression of the closely linked cyanelle genes for RNase P RNA and two transfer RNAs.

The genomic region encoding the RNA subunit of the cyanelle RNase P has been characterized. rnpB, which has no homologue in chloroplasts, is flanked by two tRNA genes on the complementary DNA strand. Transcriptional control elements of all three genes have been experimentally determined. Comparison of the sequenced region with the corresponding loci of chloroplast genomes from vascular plants suggests that major inversions may have led to a possible loss or severe truncation of the RNase P RNA coding region during the course of plastid evolution.

Base Sequence↗

Bio++: a set of C++ libraries for sequence analysis, phylogenetics, molecular evolution and population genetics.

BACKGROUND: A large number of bioinformatics applications in the fields of bio-sequence analysis, molecular evolution and population genetics typically share input/output methods, data storage requirements and data analysis algorithms. Such common features may be conveniently bundled into re-usable libraries, which enable the rapid development of new methods and robust applications. RESULTS: We present Bio++, a set of Object Oriented libraries written in C++. Available components include classes for data storage and handling (nucleotide/amino-acid/codon sequences, trees, distance matrices, population genetics datasets), various input/output formats, basic sequence manipulation (concatenation, transcription, translation, etc.), phylogenetic analysis (maximum parsimony, markov models, distance methods, likelihood computation and maximization), population genetics/genomics (diversity statistics, neutrality tests, various multi-locus analyses) and various algorithms for numerical calculus. CONCLUSION: Implementation of methods aims at being both efficient and user-friendly. A special concern was given to the library design to enable easy extension and new methods development. We defined a general hierarchy of classes that allow the developer to implement its own algorithms while remaining compatible with the rest of the libraries. Bio++ source code is distributed free of charge under the CeCILL general public licence from its website http://kimura.univ-montp2.fr/BioPP.

Algorithms↗