On the origin and evolution of the genetic code. II. Origin of the genetic code as a primordial collector language. The pairing-release hypothesis.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The genetic code doublets can be divided into two octets of completely degenerate and ambiguous coding dinucleotides. These two octets have the algebraic property of lying on continuously connected planes on the group graph (a tesseract) of the Cartesian product of two Klein 4-groups of nucleotide exchange operators. The K X K group can also be broken into four cosets, one of which has completely degenerate coding elements, and another that has completely ambiguous coding elements. The two octets of coding doublets have the further algebraic property that the product of their internal exchange operators naturally divide into two exactly equivalent sets. These properties of the genetic code are relevant to unraveling error-detecting and error-correcting (proof-reading) aspects of the genetic code and may be helpful in understanding the context-sensitive grammar of genetic language.
The genetic code determines not only the amino acid sequences of proteins but also mRNA stability. How is this hidden message read? Hia and colleagues have now identified human DHX29 as a reader of the mRNA stability code carried by codons, providing new mechanistic insights into translation-coupled gene regulation.
Francis Crick proposed that the almost universal genetic code could be nothing more than a 'frozen accident'. The Ising model, widely used in statistical mechanics, is used to explore patterns which could achieve a phase transition mimicking the frozen accident. Codons are considered as nodes and amino acids as spins. Monte Carlo simulations of the 64-node genetic code models are carried out. Anti-ferromagnetic interactions or a combination of ferro and anti-ferromagnetic interactions can lead to stable, regular patterns resembling the genetic code. It is indeed found that the 64-node Ising system exhibits critical slowing down dynamics, compatible with a freezing process. These models aim to simulate the formation of patterns similar to the genetic code through physical freezing processes, providing insights into potential mechanisms and processes that may have contributed to the formation and evolution of the genetic code.
The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.
INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.
The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.
The contemporary genetic code is reflective of a significant correlation between the properties of amino acids and their anticodons in a periodic manner. Almost all properties of amino acids showed a greater correlation to anticondonic than to codonic dinucleoside monophosphate properties. The polarity and bulkiness of amino acid side chains can be used to predict the anticodon with considerable confidence. The results are most consistent with predictions of the "direct interaction" and "ambiguity reduction" hypotheses for the origin of the genetic code.
The sequential fulfillment of the principle of succession necessarily guides the main steps of the genetic code evolution to be reflected in its structure. The general scheme of the code series formation is proposed basing on the idea of "group coding" (Woese, 1970). The genetic code supposedly evolved by means of successive divergence of pra-ARS's loci, accompanied by increasing specification of recognition capacity of amino acids and triplets. The sense of codons had not been changed on any step of stochastic code evolution. The formulated rules for code series formation produce a code version, similar to the contemporary one. Based on these rules the scheme of pra-ARS's divergence is proposed resulting in the grouping of amino acids by their polarity and size. Later steps in the evolution of the genetic code were probably based on more detailed features of the amino acids (for example, on their functional similarities like their interchangeabilities in isofunctional proteins).
The problem of the origin of life understandably counts as one of the most exciting questions in the natural sciences, but in spite of almost endless speculation on this subject, it is still far from its final solution. The complexity of the functional correlation between recent nucleic acids and proteins can e.g. give rise to the assumption that the genetic code (and life) could not originate on the Earth. It was Portelli (1975) who published the hypothesis that the genetic code could not originate during the history of the Earth. In his opinion the recent genetic code represents the informational message transmitted by living systems of the previous cycle of the Universe. Here however, we defend the existence of a certain strategy in the syntheses of the genetic code during the history of the Earth. The strategy of correlation between amino acid and nucleotide polymers made an increasing velocity of the chemical evolution possible, that is, it increased the velocity of formation of the genetic code. Thus, life with the recent genetic code could originate on the Earth within the present cycle of the Universe.
We report the relative stabilities, in the form of complex lifetimes, of complexes between the tRNAs complementary, or nearly so, in their anticodons. The results show striking parallels with the genetic coding rules, including the wobble interaction and the role of modified nucleotides S2U and V (a 5-oxyacetic acid derivative of U). One important difference between the genetic code and the pairing rules in the tRNA-tRNA interaction is the stability in the latter of the short wobble pairs, which the wobble hypothesis excludes. We stress the potential of U for translational errors, and suggest a simple stereochemical basis for ribosome-mediated discrimination against short wobble pairs. Surprisingly, the stability of anticodon-anticodon complexes does not vary systematically on base sequence. Because of the close similarity to the genetic coding rules, it is tempting to speculate that the interaction between two RNA loops may have been part of the physical basis for the evolutionary origin of the genetic code, and that this mechanism may still be utilized by folding the mRNA on the ribosome into a loop similar to the anticodon loop.
The contemporary genetic code and the process of protein biosynthesis most assuredly evolved from a simpler code and process. We believe that there was obligatory coevolution of the two and that the earlier code and process must have involved a more direct linkage between the amino acids and the information macromolecule. We propose that an early form of translating existed in which amino acids were attached directly to the 'messenger' RNA along the backbone as 2'OH aminoacyl esters. These esters then condensed with each other on the RNA backbone yielding a peptide covalently attached to the RNA, without the use of tRNA's and ribosomes. THis presentation is concerned with experimental data which indicate that such a simple translation system is possible and must have involved the following steps: (1) formation of the aminoacyl adenylate anhydride, (2) transfer of the amino acid from the adenylate to immidazole, (3) transfer of the amino acid from imidazole to 2'OH groups along the backbone of RNAs, (4) condensation of the amino acids to yield peptides. Steps (1)-(3) have been confirmed in chemical systems. Our preliminary evidence indicates step (4) is also possible. The aminoacylation of polyribonucleotides and the subsequent formation of peptides is a dynamic and experimentally accessible system for studying genetic coding specfities and our present studies are now concentrated on step (4), looking for such specifities.
Distribution of amino acids in 68 representative proteins is compared with their distribution among 61 codons of the genetic code. Average amounts of lysine, aspartic acid, glutamic acid, and alanine are above the levels anticipated from the genetic code, and arginine, serine, leucine, cysteine, proline, and histidine are below such levels. Arginine plus lysine account for 11.0 percent of codons and aspartic acid plus glutamic acid account for 11.3 percent; thus the average charge is roughly neutral.
A correlation of various aspects of the protein structures and substrate and mechanistic specificities of the aminoacyl-tRNA synthetases has led to the identification of at least one family of enzymes probably derived from a common ancestral synthetase. While strong correlations exist only in one part of the array of 64 codons comprising the Genetic Code, this itself may be interpreted as a meaningful pattern, most consistent with a development of the present code from earlier codes containing fewer amino acids and fewer available codons. Specifically, strong correlations in the enzymes whose cognate tRNAs respond to codons containing a central pyrimidine, including the enzyme family of Ile-, Phe-, Val-, Met-, and Leu-tRNA synthetases, suggests that these enzymes evolved last, and that, therefore, an earlier version of the Genetic Code was comprised solely of codons containing a central purine. It is suggested that further study of the historical interrelationships of these enzymes could lead to a fairly detailed picture of how the Genetic Code developed.
A simple selforganizing model system of molecules is considered and it is demonstrated by a computer simulation, that a genetic code of 16 elements (aminoacids) can gradually be formed by such a system in the course of many generations. By a number of rare chance events, each suppressing other events of equal a priori probability, a single code results out of an immense number of possible codes of the same a priori probability. The result is discussed in relation to the uniqueness of the genetic code in living systems. The computer simulation emphasizes a particular step in a model pathway discussed elsewhere consisting of many assumed physicochemical steps leading to a genetic apparatus.
A new approach to the origin of the genetic code is proposed based on some regularities in the nucleotide distribution pattern of the code. The relative amounts of various amino acids in primitive proteins were possibly different from those in organisms living today. The primordial ratio was supposed to shift to the modern one guided by the action of primitive nucleotides. Each primitive tRNA had a discriminator site and, distinguished from it, an anticodon site. It also postulated that primordially each amino acid could correspond to a wide variety of codons. During the course of the evolutionary change, a selective mechanism worked among the protobionts so that less frequent nucleotides became associated with more abundant amino acids in the primordial conditions,thus finally leading to the present codon catalogue.
Explore the source record for details and available documents.
The theory is proposed that the structure of the genetic code was determined by the sequence of evolutionary emergence of new amino acids within the primordial biochemical system.