On the origin and evolution of the genetic code. II. Origin of the genetic code as a primordial collector language. The pairing-release hypothesis.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The genetic code doublets can be divided into two octets of completely degenerate and ambiguous coding dinucleotides. These two octets have the algebraic property of lying on continuously connected planes on the group graph (a tesseract) of the Cartesian product of two Klein 4-groups of nucleotide exchange operators. The K X K group can also be broken into four cosets, one of which has completely degenerate coding elements, and another that has completely ambiguous coding elements. The two octets of coding doublets have the further algebraic property that the product of their internal exchange operators naturally divide into two exactly equivalent sets. These properties of the genetic code are relevant to unraveling error-detecting and error-correcting (proof-reading) aspects of the genetic code and may be helpful in understanding the context-sensitive grammar of genetic language.
The genetic code is characterized by hidden symmetry. Amino acids possessing common antiamino acids are located symmetrically in the graphic models of the code. There is only one exception--apolar amino acids V, M, I, L and F are asymmetrically arranged. Asymmetric disposition of these amino acids is apparently due to divergence in the course of structural evolution of amino acid families as a result of inclusion of new members into the coding system.
The genetic code determines not only the amino acid sequences of proteins but also mRNA stability. How is this hidden message read? Hia and colleagues have now identified human DHX29 as a reader of the mRNA stability code carried by codons, providing new mechanistic insights into translation-coupled gene regulation.
The genetic code, formerly thought to be frozen, is now known to be in a state of evolution. This was first shown in 1979 by Barrell et al. (G. Barrell, A. T. Bankier, and J. Drouin, Nature [London] 282:189-194, 1979), who found that the universal codons AUA (isoleucine) and UGA (stop) coded for methionine and tryptophan, respectively, in human mitochondria. Subsequent studies have shown that UGA codes for tryptophan in Mycoplasma spp. and in all nonplant mitochondria that have been examined. Universal stop codons UAA and UAG code for glutamine in ciliated protozoa (except Euplotes octacarinatus) and in a green alga, Acetabularia. E. octacarinatus uses UAA for stop and UGA for cysteine. Candida species, which are yeasts, use CUG (leucine) for serine. Other departures from the universal code, all in nonplant mitochondria, are CUN (leucine) for threonine (in yeasts), AAA (lysine) for asparagine (in platyhelminths and echinoderms), UAA (stop) for tyrosine (in planaria), and AGR (arginine) for serine (in several animal orders) and for stop (in vertebrates). We propose that the changes are typically preceded by loss of a codon from all coding sequences in an organism or organelle, often as a result of directional mutation pressure, accompanied by loss of the tRNA that translates the codon. The codon reappears later by conversion of another codon and emergence of a tRNA that translates the reappeared codon with a different assignment. Changes in release factors also contribute to these revised assignments. We also discuss the use of UGA (stop) as a selenocysteine codon and the early history of the code.
Francis Crick proposed that the almost universal genetic code could be nothing more than a 'frozen accident'. The Ising model, widely used in statistical mechanics, is used to explore patterns which could achieve a phase transition mimicking the frozen accident. Codons are considered as nodes and amino acids as spins. Monte Carlo simulations of the 64-node genetic code models are carried out. Anti-ferromagnetic interactions or a combination of ferro and anti-ferromagnetic interactions can lead to stable, regular patterns resembling the genetic code. It is indeed found that the 64-node Ising system exhibits critical slowing down dynamics, compatible with a freezing process. These models aim to simulate the formation of patterns similar to the genetic code through physical freezing processes, providing insights into potential mechanisms and processes that may have contributed to the formation and evolution of the genetic code.
The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.
Theories of the origin of the genetic code assign different weights to amino acid properties such as polarity and precursor-product relationship. Previous statistical work on the origin of the genetic code has produced controversial results. We analyze relationships between various amino acid and tRNA properties by one and the same statistical method. It is shown that polarities as well as precursor-product relationships are both likely to have been important in shaping the genetic code, together with codon swapping that left protein sequences intact.
INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.
The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.
The contemporary genetic code is reflective of a significant correlation between the properties of amino acids and their anticodons in a periodic manner. Almost all properties of amino acids showed a greater correlation to anticondonic than to codonic dinucleoside monophosphate properties. The polarity and bulkiness of amino acid side chains can be used to predict the anticodon with considerable confidence. The results are most consistent with predictions of the "direct interaction" and "ambiguity reduction" hypotheses for the origin of the genetic code.
The sequential fulfillment of the principle of succession necessarily guides the main steps of the genetic code evolution to be reflected in its structure. The general scheme of the code series formation is proposed basing on the idea of "group coding" (Woese, 1970). The genetic code supposedly evolved by means of successive divergence of pra-ARS's loci, accompanied by increasing specification of recognition capacity of amino acids and triplets. The sense of codons had not been changed on any step of stochastic code evolution. The formulated rules for code series formation produce a code version, similar to the contemporary one. Based on these rules the scheme of pra-ARS's divergence is proposed resulting in the grouping of amino acids by their polarity and size. Later steps in the evolution of the genetic code were probably based on more detailed features of the amino acids (for example, on their functional similarities like their interchangeabilities in isofunctional proteins).
The problem of the origin of life understandably counts as one of the most exciting questions in the natural sciences, but in spite of almost endless speculation on this subject, it is still far from its final solution. The complexity of the functional correlation between recent nucleic acids and proteins can e.g. give rise to the assumption that the genetic code (and life) could not originate on the Earth. It was Portelli (1975) who published the hypothesis that the genetic code could not originate during the history of the Earth. In his opinion the recent genetic code represents the informational message transmitted by living systems of the previous cycle of the Universe. Here however, we defend the existence of a certain strategy in the syntheses of the genetic code during the history of the Earth. The strategy of correlation between amino acid and nucleotide polymers made an increasing velocity of the chemical evolution possible, that is, it increased the velocity of formation of the genetic code. Thus, life with the recent genetic code could originate on the Earth within the present cycle of the Universe.
We report the relative stabilities, in the form of complex lifetimes, of complexes between the tRNAs complementary, or nearly so, in their anticodons. The results show striking parallels with the genetic coding rules, including the wobble interaction and the role of modified nucleotides S2U and V (a 5-oxyacetic acid derivative of U). One important difference between the genetic code and the pairing rules in the tRNA-tRNA interaction is the stability in the latter of the short wobble pairs, which the wobble hypothesis excludes. We stress the potential of U for translational errors, and suggest a simple stereochemical basis for ribosome-mediated discrimination against short wobble pairs. Surprisingly, the stability of anticodon-anticodon complexes does not vary systematically on base sequence. Because of the close similarity to the genetic coding rules, it is tempting to speculate that the interaction between two RNA loops may have been part of the physical basis for the evolutionary origin of the genetic code, and that this mechanism may still be utilized by folding the mRNA on the ribosome into a loop similar to the anticodon loop.
The contemporary genetic code and the process of protein biosynthesis most assuredly evolved from a simpler code and process. We believe that there was obligatory coevolution of the two and that the earlier code and process must have involved a more direct linkage between the amino acids and the information macromolecule. We propose that an early form of translating existed in which amino acids were attached directly to the 'messenger' RNA along the backbone as 2'OH aminoacyl esters. These esters then condensed with each other on the RNA backbone yielding a peptide covalently attached to the RNA, without the use of tRNA's and ribosomes. THis presentation is concerned with experimental data which indicate that such a simple translation system is possible and must have involved the following steps: (1) formation of the aminoacyl adenylate anhydride, (2) transfer of the amino acid from the adenylate to immidazole, (3) transfer of the amino acid from imidazole to 2'OH groups along the backbone of RNAs, (4) condensation of the amino acids to yield peptides. Steps (1)-(3) have been confirmed in chemical systems. Our preliminary evidence indicates step (4) is also possible. The aminoacylation of polyribonucleotides and the subsequent formation of peptides is a dynamic and experimentally accessible system for studying genetic coding specfities and our present studies are now concentrated on step (4), looking for such specifities.
In this paper the partition metric is used to compare binary trees deriving from (i) the study of the evolutionary relationships between aminoacyl-tRNA synthetases, (ii) the physicochemical properties of amino acids and (iii) the biosynthetic relationships between amino acids. If the tree defining the evolutionary relationships between aminoacyl-tRNA synthetases is assumed to be a manifestation of the mechanism that originated the organization of the genetic code, then the results appear to indicate the following: the hypothesis that regards the genetic code as a map of the biosynthetic relationships between amino acids seems to explain the organization of the genetic code, at least as plausibly as the hypotheses that consider the physicochemical properties of amino acids as the main adaptive theme that lead to the structuring of the code.
Distribution of amino acids in 68 representative proteins is compared with their distribution among 61 codons of the genetic code. Average amounts of lysine, aspartic acid, glutamic acid, and alanine are above the levels anticipated from the genetic code, and arginine, serine, leucine, cysteine, proline, and histidine are below such levels. Arginine plus lysine account for 11.0 percent of codons and aspartic acid plus glutamic acid account for 11.3 percent; thus the average charge is roughly neutral.
A correlation of various aspects of the protein structures and substrate and mechanistic specificities of the aminoacyl-tRNA synthetases has led to the identification of at least one family of enzymes probably derived from a common ancestral synthetase. While strong correlations exist only in one part of the array of 64 codons comprising the Genetic Code, this itself may be interpreted as a meaningful pattern, most consistent with a development of the present code from earlier codes containing fewer amino acids and fewer available codons. Specifically, strong correlations in the enzymes whose cognate tRNAs respond to codons containing a central pyrimidine, including the enzyme family of Ile-, Phe-, Val-, Met-, and Leu-tRNA synthetases, suggests that these enzymes evolved last, and that, therefore, an earlier version of the Genetic Code was comprised solely of codons containing a central purine. It is suggested that further study of the historical interrelationships of these enzymes could lead to a fairly detailed picture of how the Genetic Code developed.