PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Amino acid composition of proteins as a product of molecular evolution.

The average amino acid composition of proteins is determined by the genetic code and by random base changes in evolution. Small but significant deviations from expected composition can be explained by selective constraint on amino acid substitutions. In particular, the deficiency of arginine in proteins has been caused by constraint, during evolution, on fixation of mutations substituting arginine for other amino acids.

Amino Acids

Primordial reading of genetic information.

From the consideration of general features of the anticodon loop and stem in tRNA and the properties of present-day translation, we put forward a plausible scenario to explain the evolution of the genetic code from a highly ambiguous triplet code to the present refined decoding system. Our model based on the reading of the code suggests that the anticodon of primordial tRNA could adopt either the 3' or the 5' stacked conformation permitting the formation of the "best two out of three" base pairs, either the first and second codon position or the second and third. Progressive acquisition of precise structural constraint and the modification of bases in the anticodon loop would give way eventually to the less ambiguous "two out of three" reading mechanism having only the 3' stacked conformation. Further adjustments of base composition and modification leads inevitably to the present generalized code. In this way the primordial code encoding 4-8 amino acids or related derivates evolves smoothly to the present code having 20 amino acids.

Biological Evolution

Physicochemical optimization in the genetic code origin as the number of codified amino acids increases.

We have assumed that the coevolution theory of genetic code origin (Wong JT, Proc Natl Acad Sci USA 72:1909-1912, 1975) is essentially correct. This theory makes it possible to identify at least 10 evolutionary stages through which genetic code organization might have passed prior to reaching its current form. The calculation of the minimization level of all these evolutionary stages leads to the following conclusions. (1) The minimization percentages increased linearly with the number of amino acids codified in the codes of the various evolutionary stages when only the sense changes are considered in the analysis. This seems to favor the physicochemical theory of genetic code origin even if, as discussed in the paper, this observation is also compatible with the coevolution theory. (2) For the first seven evolutionary stages of the genetic code, this trend is less clear and indeed is inverted when we consider the global optimisation of the codes due to both sense changes and synonymous changes. This inverse correlation between minimization percentages and the number of amino acids codified in the codes of the intermediate stages seems to favor neither the physicochemical nor the stereochemical theories of genetic code origin, as it is in the early and intermediate stages of code development that these theories would expect minimization to have played a crucial role, and this does not seem to be the case. However, these results are in agreement with the coevolution theory, which attributes a role to the physicochemical properties of amino acids that, while important, is nevertheless subordinate to the mechanism which concedes codons from the precursor amino acids to the product amino acids as the primary factor determining the evolutionary structuring of the genetic code. The results are therefore discussed in the context of the various theories proposed to explain genetic code origin.

Algorithms

The evolution of aminoacyl-tRNA synthetases, the biosynthetic pathways of amino acids and the genetic code.

In this paper the partition metric is used to compare binary trees deriving from (i) the study of the evolutionary relationships between aminoacyl-tRNA synthetases, (ii) the physicochemical properties of amino acids and (iii) the biosynthetic relationships between amino acids. If the tree defining the evolutionary relationships between aminoacyl-tRNA synthetases is assumed to be a manifestation of the mechanism that originated the organization of the genetic code, then the results appear to indicate the following: the hypothesis that regards the genetic code as a map of the biosynthetic relationships between amino acids seems to explain the organization of the genetic code, at least as plausibly as the hypotheses that consider the physicochemical properties of amino acids as the main adaptive theme that lead to the structuring of the code.

Amino Acids

Relationships among isoacceptor tRNAs seems to support the coevolution theory of the origin of the genetic code.

A new method for looking at relationships between nucleotide sequences has been used to analyze divergence both within and between the families of isoaccepting tRNA sets. A dendrogram of the relationships between 21 tRNA sets with different amino acid specificities is presented as the result of the analysis. Methionine initiator tRNAs are included as a separate set. The dendrogram has been interpreted with respect to the final stage of the evolutionary pathway with the development of highly specific tRNAs from ambiguous molecular adaptors. The location of the sets on the dendrogram was therefore analyzed in relation to hypotheses on the origin of the genetic code: the coevolution theory, the physicochemical hypothesis, and the hypothesis of ambiguity reduction of the genetic code. Pairs of 16 sets of isoacceptor tRNAs, whose amino acids are in biosynthetic relationships, occupied contiguous positions on the dendrogram, thus supporting the coevolution theory of the genetic code.

Codon

An unusual genetic code in nuclear genes of Tetrahymena.

We have cloned and partially sequenced two histone H3 genes of Tetrahymena thermophila. The DNA sequences strongly suggest that both genes are active in the vegetatively growing cell. Comparison of the derived amino acid sequences of these two genes with the actual sequence of Tetrahymena histone H3 results in the surprising conclusion that TAA codes for glutamine. This represents the first demonstration of a coding function for this termination codon of the "universal" code. This observation has important implications for the evolution of ciliates and of the genetic code.

Amino Acid Sequence

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

Partition of aminoacyl-tRNA synthetases in two different structural classes dating back to early metabolism: implications for the origin of the genetic code and the nature of protein sequences.

We describe, on the molecular level, a possible fuzzy and primordial translation apparatus capable of synthesizing polypeptides from nucleic acids in a world containing a mixture of coevolving molecules of RNA and proteins already arranged in metabolic cycles (including cofactors). Close attention is paid to template-free systems because they are believed to be the immediate ancestors of this primordial translation apparatus. The two classes of aminoacyl-tRNA synthetases (aaRSs), as seen today, are considered as the remnants of such a simple imprecise translation apparatus and are used as guidelines for the construction of the model. Earlier theoretical work by Bedian on a related system is invoked to show how specificity and stability could have been achieved automatically and rather quickly, starting from such an imprecise system, i.e., how the encoded synthesis of proteins could have appeared. Because of the binary nature of the underlying proto-code, the first genetically encoded proteins would then have been alternating copolymers with a high degree of degeneracy, but not random. Indeed, a clear signal for alternating hydrophobic and hydrophilic residues in present-day protein sequences can be detected. Later evolution of the genetic code would have proceeded along lines already discussed by Crick. However, in the initial stages, the translation apparatus proposed here is in fact very similar to the one postulated by Woese, only here it is given a molecular framework. This hypothesis departs from the paradigm of the RNA world in that it supposes that the origin of the genetic code occurred after the apparition of some functional (statistical) proteins first. Implications for protein design are also discussed.

Amino Acid Sequence

Progress toward the evolution of an organism with an expanded genetic code.

Several significant steps have been completed toward a general method for the site-specific incorporation of unnatural amino acids into proteins in vivo. An "orthogonal" suppressor tRNA was derived from Saccharomyces cerevisiae tRNA2Gln. This yeast orthogonal tRNA is not a substrate in vitro or in vivo for any Escherichia coli aminoacyl-tRNA synthetase, including E. coli glutaminyl-tRNA synthetase (GlnRS), yet functions with the E. coli translational machinery. Importantly, S. cerevisiae GlnRS aminoacylates the yeast orthogonal tRNA in vitro and in E. coli, but does not charge E. coli tRNAGln. This yeast-derived suppressor tRNA together with yeast GlnRS thus represents a completely orthogonal tRNA/synthetase pair in E. coli suitable for the delivery of unnatural amino acids into proteins in vivo. A general method was developed to select for mutant aminoacyl-tRNA synthetases capable of charging any ribosomally accepted molecule onto an orthogonal suppressor tRNA. Finally, a rapid nonradioactive screen for unnatural amino acid uptake was developed and applied to a collection of 138 amino acids. The majority of glutamine and glutamic acid analogs under examination were found to be uptaken by E. coli. Implications of these results are discussed.

Amino Acid Substitution

Codon usage in Homo sapiens: evidence for a coding pattern on the non-coding strand and evolutionary implications of dinucleotide discrimination.

This study reports the analysis of codon usage in 35 complete Homo sapiens genes. Both codon frequency and inter-codon interference exhibit patterns of evolutionary interest. There is a significant positive correlation between the frequency with which a given codon is used and the frequency with which its complement is used. Since the frequency of appearance of the complementary codon on the coding strand is equal to the frequency of appearance of the original codon on the non-coding strand, in the same phase, the non-coding strand is found to resemble the coding strand in triplet composition. The same effect has been observed in Escherichia coli. This preference for the use of certain complementary triplets as codons suggests that the evolution of the use of the genetic code depended to some extent upon the double-stranded nature of the coding material. In addition, the effect of discrimination against the use of two dinucleotides, CpG and UpA, is observed in codon usage and also in adjacent codon interference. Codons beginning with G, or A, are unlikely to be preceded by codons ending in C, or U, respectively. Consideration of codon assignment in the genetic code together with the observed CpG infrequency suggests that the evolution of the code may have been influenced by conditions in which the use of CpG dinucleotides was unfavorable. The infrequent use of UpA dinucleotides can be explained as the result of frameshift mutation during gene evolution.

Base Sequence

Selection, history and chemistry: the three faces of the genetic code.

The genetic code might be a historical accident that was fixed in the last common ancestor of modern organisms. 'Adaptive', 'historical' and 'chemical' arguments, however, challenge such a 'frozen accident' model. These arguments propose that the current code is somehow optimal, reflects the expansion of a more primitive code to include more amino acids, or is a consequence of direct chemical interactions between RNA and amino acids, respectively. Such models are not mutually exclusive, however. They can be reconciled by an evolutionary model whereby stereochemical interactions shaped the initial code, which subsequently expanded through biosynthetic modification of encoded amino acids and, finally, was optimized through codon reassignment. Alternatively, all three forces might have acted in concert to assign the 20 'natural' amino acids to their present positions in the genetic code.

Biological Evolution

Information theory and the genetic code.

The genetic code, which directs the protein biosynthesis, is an information system. Although all its details are not known at present, its essential characteristics are elucidated, as well for the replication or transcription as for the translation of the genetic message. A coherent picture now appears, which reveals the existence of an universal structure, the most fundamental features of which seem to obey some logic. A systematic approach has been devised, which aims to their integration in a theorectical scheme: many features of the code table can thus be interpreted as resulting from a unique principle of best resistance against the effects of mutations. Any group of triplets or amino-acids can be considered along this line. It is more difficult however, to analyse the coexistence of two (or more) different groups. In this work, we propose to extend our optimization principle into a more general one, which includes the notion of information as defined by Shannon. We explore some consequences of this new principle in the most simple models that one can build for the origin and evolution of the genetic code.

Genetic Code