PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

On the origin of the genetic code.

The main theories which have been suggested to explain the origin of genetic code organization are discussed. The coevolution theory, which considers the genetic code as a map of the biosynthetic relationships between amino acids, seems to be based on a mechanism that links it closely to certain stages of the origin of metabolism, which makes it preferable to other theories proposed as explanations of genetic code origin. Relationships, incompatibilities and compromises between the various theories are highlighted and these seem to indicate a certain lack of clarity in this field of research.

Animals↗

An analysis of the metabolic theory of the origin of the genetic code.

A computer program was used to test Wong's coevolution theory of the genetic code. The codon correlations between the codons of biosynthetically related amino acids in the universal genetic code and in randomly generated genetic codes were compared. It was determined that many codon correlations are also present within random genetic codes and that among the random codes there are always several which have many more correlations than that found in the universal code. Although the number of correlations depends on the choice of biosynthetically related amino acids, the probability of choosing a random genetic code with the same or greater number of codon correlations as the universal genetic code was found to vary from 0.1% to 34% (with respect to a fairly complete listing of related amino acids). Thus, Wong's theory that the genetic code arose by coevolution with the biosynthetic pathways of amino acids, based on codon correlations between biosynthetically related amino acids, is statistical in nature.

Amino Acids↗

Error minimization and coding triplet/binding site associations are independent features of the canonical genetic code.

The canonical genetic code has been reported both to be error minimizing and to show stereochemical associations between coding triplets and binding sites. In order to test whether these two properties are unexpectedly overlapping, we generated 200,000 randomized genetic codes using each of five randomization schemes, with and without randomization of stop codons. Comparison of the code error (difference in polar requirement for single-nucleotide codon interchanges) with the coding triplet concentrations in RNA binding sites for eight amino acids shows that these properties are independent and uncorrelated. Thus, one is not the result of the other, and error minimization and triplet associations probably arose independently during the history of the genetic code. We explicitly show that prior fixation of a stereochemical core is consistent with an effective later minimization of error.

Base Sequence↗

The robust statistical bases of the coevolution theory of genetic code origin.

A paper (Amirnovin R, J Mol Evol 44:473-476, 1997) seems to undermine the validity of the coevolution theory of genetic code origin by shedding doubt on the connection between the biosynthetic relationships between amino acids and the organization of the genetic code, at a time when the literature on the topic takes this for granted. However, as a few papers cite this paper as evidence against the coevolution theory, and to cast aside all doubt on the subject, we have decided to reanalyze the statistical bases on which this theory is founded. We come to the following conclusions: (1) the methods used in the above referred paper contain certain mistakes, and (2) the statistical foundations on which the coevolution theory is based are extremely robust. We have done this by critically appraising Amirnovin's paper and suggesting an alternative method based on the generation of random codes which, along with the method reported in the literature, allows us to evaluate the significance, in the genetic code, of different sets of amino acid pairs in biosynthetic relationships. In particular, by using this method and after building up a certain set of amino acid pairs reflecting the expectations of the coevolution theory, we show that the presence of this set in the genetic code would be obtained, purely by chance, with a probability of 6x10(-5). This observation seems to provide particularly strong support to the coevolution theory.

Amino Acids↗

Coding coenzyme handles: a hypothesis for the origin of the genetic code.

The coding coenzyme handle hypothesis suggests that useful coding preceded translation. Early adapters, the ancestors of present-day anticodons, were charged with amino acids acting as coenzymes of ribozymes in a metabolically complex RNA world. The ancestral aminoacyl-adapter synthetases could have been similar to present-day self-splicing tRNA introns. A codon-anticodon-discriminator base complex embedded in these synthetases could have played an important role in amino acid recognition. Extension of the genetic code proceeded through the take-over of nonsense codons by novel amino acids, related to already coded ones either through precursor-product relationship or physicochemical similarity. The hypothesis is open for experimental tests.

Amino Acyl-tRNA Synthetases↗

Preliminary amino acid sequences of transplantation antigens: genetic and evolutionary implications.

Preliminary amino acid sequence data on the transplantation antigens of mouse and man have led to provocative hypotheses about the genetic organization and evolution of genes coded by the major histocompatibility complex of mammals. New microsequencing techniques should permit a detailed analysis of these gene products and an eventual choice among the alternative hypotheses now posed. These data have made it apparent that the H-2 complex is a fascinating and complicated chromosomal region which will continue for some time to intrigue immunologists, geneticists, biochemists, and cell biologists.

Alleles↗

On the relationships between the genetic code coevolution hypothesis and the physicochemical hypothesis.

This paper analyzes the relationships between the genetic code coevolution hypothesis and the physicochemical hypothesis by means of a comparative study of the precursor-product amino acid pairs on which the former hypothesis is based. Even if the coevolution between the biosynthetic relationships of amino acids and the organization of the genetic code is not questioned in this paper, the results and the arguments used lead us to believe that the selective pressures considered essential by the physicochemical postulates, played a more active role than that of the precursor-product relationships in defining the allocation of these amino acids in the genetic code. It is furthermore pointed out that the two evolutionary hypothesis might be aspects of the same selective pressure, and thus difficult to differentiate.

Amino Acids↗

Selection on codon usage for error minimization at the protein level.

Given the structure of the genetic code, synonymous codons differ in their capacity to minimize the effects of errors due to mutation or mistranslation. I suggest that this may lead, in protein-coding genes, to a preference for codons that minimize the impact of errors at the protein level. I develop a theoretical measure of error minimization for each codon, based on amino acid similarity. This measure is used to calculate the degree of error minimization for 82 genes of Drosophila melanogaster and 432 rodent genes and to study its relationship with CG content, the degree of codon usage bias, and the rate of nucleotide substitution. I show that (i) Drosophila and rodent genes tend to prefer codons that minimize errors; (ii) this cannot be merely the effect of mutation bias; (iii) the degree of error minimization is correlated with the degree of codon usage bias; (iv) the amino acids that contribute more to codon usage bias are the ones for which synonymous codons differ more in the capacity to minimize errors; and (v) the degree of error minimization is correlated with the rate of nonsynonymous substitution. These results suggest that natural selection for error minimization at the protein level plays a role in the evolution of coding sequences in Drosophila and rodents.

Amino Acids↗

Increased frequency of cysteine, tyrosine, and phenylalanine residues since the last universal ancestor.

Analysis of extant proteomes has the potential of revealing how amino acid frequencies within proteins have evolved over biological time. Evidence is presented here that cysteine, tyrosine, and phenylalanine residues have substantially increased in frequency since the three primary lineages diverged more than three billion years ago. This inference was derived from a comparison of amino acid frequencies within conserved and non-conserved residues of a set of proteins dating to the last universal ancestor in the face of empirical knowledge of the relative mutability of these amino acids. The under-representation of these amino acids within last universal ancestor proteins relative to their modern descendants suggests their late introduction into the genetic code. Thus, it appears that extant ancient proteins contain evidence pertaining to early events in the formation of biological systems.

Amino Acid Sequence↗

Complementary coding conforms to the primeval comma-less code.

The hypothesis that the universal genetic code is adapted to double-strand coding is supported by its remarkable compatibility with the RNY comma-less hypothesis. Coding by a triplet code on a polynucleotide double-strand allows for enciphering of five additional messages with reference to a chosen primary reading frame. Assuming the acceptance of coupled mutations on both strands, the best codon register for two overlapping messages can be inferred. The idea of evolutionarily compatible coding of two proteins by one nucleotide double-strand is extended to complementary coding for one protein in folded, single-stranded RNA.

Animals↗

Origin and properties of non-coding ORFs in the yeast genome.

In a recent paper we have estimated the total number of protein coding open reading frames (ORFs) in the Saccharomyces cerevisiae genome, based on their properties, at about 4800. This number is much smaller than the 5800-6000 which is widely accepted. In this paper we analyse differences between the set of ORFs with known phenotypes annotated in the Munich Information Centre for Protein Sequences (MIPS) database and ORFs for which the probability of coding, counted by us, is very low. We have found that many of the latter ORFs have properties of antisense sequences of coding ORFs, which suggests that they could have been generated by duplication of coding sequences. Since coding sequences generate ORFs inside themselves, with especially high frequency in the antisense sequences, we have looked for homology between known proteins and hypothetical polypeptides generated by ORFs under consideration in all the six phases. For many ORFs we have found paralogues and orthologues in phases different than the phase which had been assumed in the MIPS database as coding.

Algorithms↗

Genomic evolution drives the evolution of the translation system.

Our thesis is that the characteristics of the translational machinery and its organization are selected in part by evolutionary pressure on genomic traits have nothing to do with translation per se. These genomic traits include size, composition, and architecture. To illustrate this point, we draw parallels between the structure of different genomes that have adapted to intracellular niches independently of each other. Our starting point is the general observation that the evolutionary history of organellar and parasitic bacteria have favored bantam genomes. Furthermore, we suggest that the constraints of the reductive mode of genomic evolution account for the divergence of the genetic code in mitochondria and the genetic organization of the translational system observed in parasitic bacteria. In particular, we associate codon reassignments in animal mitochondria with greatly simplified tRNA populations. Likewise, we relate the organization of translational genes in the obligate intracellular parasite Rickettsia prowazekii to the processes supporting the reductive mode of genomic evolution. Such findings provide strong support for the hypothesis that genomes of organelles and of parasitic bacteria have arisen from the much larger genomes of ancestral bacteria that have been reduced by intrachromosomal recombination and deletion events. A consequence of the reductive mode of genomic evolution is that the resulting translation systems may deviate markedly from conventional systems.

Animals↗

Coevolution theory of the genetic code at age thirty.

The coevolution theory of the genetic code, which postulates that prebiotic synthesis was an inadequate source of all twenty protein amino acids, and therefore some of them had to be derived from the coevolving pathways of amino acid biosynthesis, has been assessed in the light of the discoveries of the past three decades. Its four fundamental tenets regarding the essentiality of amino acid biosynthesis, role of pretran synthesis, biosynthetic imprint on codon allocations and mutability of the encoded amino acids are proven by the new knowledge. Of the factors that guided the evolutionary selection of the universal code, the relative contributions of Amino Acid Biosynthesis: Error Minimization: Stereochemical Interaction are estimated to first approximation as 40,000,000:400:1, which suggests that amino acid biosynthesis represents the dominant factor shaping the code. The utility of the coevolution theory is demonstrated by its opening up experimental expansions of the code and providing a basis for locating the root of life.

Amino Acids↗

Can the genetic code be mathematically described?

From a mathematical point of view, the genetic code is a surjective mapping between the set of the 64 possible three-base codons and the set of 21 elements composed of the 20 amino acids plus the Stop signal. Redundancy and degeneracy therefore follow. In analogy with the genetic code, non-power integer-number representations are also surjective mappings between sets of different cardinality and, as such, also redundant. However, none of the non-power arithmetics studied so far nor other alternative redundant representations are able to match the actual degeneracy of the genetic code. In this paper we develop a slightly more general framework that leads to the following surprising results: i) the degeneracy of the genetic code is mathematically described, ii) a new symmetry is uncovered within this degeneracy, iii) by assigning a binary string to each of the codons, their classification into definite parity classes according to the corresponding sequence of bases is made possible. This last result is particularly appealing in connection with the fact that parity coding is the basis of the simplest strategies devised for error correction in man-made digital data transmission systems.

Algorithms↗

On the 28-gon symmetry inherent in the genetic code intertwined with aminoacyl-tRNA synthetases--the Lucas series.

Despite considerable efforts it has remained unclear what principle governs the selection of the 20 canonical amino acids in the genetic code. Based on a previous study of the 28-gonal and rotational symmetric arrangement of the 20 amino acids in the genetic code, new analyses of the organization of the genetic code system together with their intrinsic relation to the two classes of aminoacyl-tRNA synthetases are reported in this work. A close inspection revealed how the enzymes and the 20 gene-encoded amino acids are intertwined on the polyhedron model. Complementary and cooperative symmetries between class I and class II aminoacyl-tRNA synthetases displayed by a 28-gon organization are discussed, and we found that the two previously suggested evolutionary axes within the genetic code overlap the symmetry axes within the two classes of aminoacyl-tRNA synthetases. Moreover, it has been shown that the side-chain carbon-atom numbers (2, 1, 3, 4 and 7) in the overwhelming majority of the amino acids recognized by each of the two classes of aminoacyl-tRNA synthetases are determined by a mathematical relationship, the Lucas series. A stepwise co-evolutionary selection logic of the amino acids is manifested by the amino acid side-chain carbon-atom number balance at '17', when grouping the genetic code doublets in the 28-gon organization. The number '17' equals the sum of the initial five numbers in the Lucas series, which are 2, 1, 3, 4 and 7.

Amino Acids↗

The case for an error minimizing set of coding amino acids.

The fidelity of the translation machinery largely depends on the accuracy by which the tRNAs within the living cells are charged. Aminoacyl-tRNA synthetases (aaRSs) attach amino acids to their cognate tRNAs ensuring the fidelity of translation in coding sequences. Based on the sequence analysis and catalytic domain structure, these enzymes are classified into two major groups of 10 enzymes each. In this study, we have generally tackled the role of aaRSs in decreasing the effects of mistranslations and consequently the evolution of the translation machinery. To this end, a fitness function was introduced in order to measure the accuracy by which each tRNA is charged with its cognate amino acid. Our results suggest that the aaRSs are very well optimized in "load minimization" based on their classes and their mechanisms in distinguishing the correct amino acids. Besides, our results support the idea that from an evolutionary point, a selectional pressure on the translational fidelity seems to be responsible in the occurrence of the 20 coding amino acids.

Amino Acids↗

Rewiring the keyboard: evolvability of the genetic code.

The genetic code evolved in two distinct phases. First, the 'canonical' code emerged before the last universal ancestor; subsequently, this code diverged in numerous nuclear and organelle lineages. Here, we examine the distribution and causes of these secondary deviations from the canonical genetic code. The majority of non-standard codes arise from alterations in the tRNA, with most occurring by post-transcriptional modifications, such as base modification or RNA editing, rather than by substitutions within tRNA anticodons.

Animals↗

The puzzling origin of the genetic code.

Recent results add to the mystery of the origin of the genetic code. In spite of early doubts, RNA can discriminate between hydrophobic amino acids under certain contexts. Moreover, codon reassignment, which has taken place in several organisms and mitochondria, is not a random process. Finally, phylogenies of some aminoacyl-tRNA synthetases suggest that the entire code was not completely assigned at the time of the divergence of bacteria from nucleated cells.

Amino Acids↗