[Successful solution of the genetic code].
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The theory is proposed that the structure of the genetic code was determined by the sequence of evolutionary emergence of new amino acids within the primordial biochemical system.
It has been shown that Chou-Fasman conformational parameters of amino acids, which reflect their ability to adopt a definite conformation within the peptide chain, change very regularly within the genetic code, arranged in the manner discussed recently by Siemion and Stefanowicz (1992a) (BioSystems 27, 77-84). Two mutually perpendicular C2 axes of pseudosymmetry appear in the center of the diagrams (between ACY and ACR threonine codons) presenting the changes of P alpha and P beta parameters. The left and right parts of diagrams superimpose on each other quite well when the symmetry operation involving a proper axis is performed. This phenomenon is due, in our opinion, to the regular arrangement of equivalent codons in the 'one-step mutation' ring formed by 64 triplets of the genetic code.
The apparent dissociation constants of the complexes of AMP with the methyl esters of amino acids in aqueous solution exhibit good correlations with features of the genetic code and with the frequencies of occurrence of amino acid residues in proteins. Thus it is likely that chemically selective nucleotide-amino acid interactions were involved in the processes of chemical evolution that have led to the emergence of the genetic code. Based on these correlations a storage device for the information regarding nucleotide-amino acid interactions is proposed. It involves processes of simultaneous polymerization to polynucleotides and polypeptides.
A model is presented for the emergence of a primitive genetic code through the selection of a family of proteins capable of executing the code and catalyzing their own formation from polynucleotide templates. These proteins are assignment catalysts capable of modulating the rate of incorporation of different amino acids at the position of different codons. The starting point of the model is a polynucleotide based polypeptide construction process which maintains colinearity between template and product, but may not maintain a coded relationship between amino acids and codons. Among the primitive proteins made are assumed to be assignment catalysts characterized by structural and functional parameters which are used to formulate the production kinetics of these catalysts from available templates. Application of the model to the simple case of two letter codon and amino acid alphabets has been analyzed in detail. As the structural, functional, and kinetic parameters are varied, the dynamics undergoes many bifurcations, allowing an initially ambiguous system of catalysts to evolve to a coded, self-reproductive system. The proposed selective pressure of this evolution is the efficiency of utilization of monomers and energy. The model also simulates the qualitative features of suppression, in which a deleterious mutation is partly corrected by the introduction of translation error.
A pocket on the complex of four nucleotides (C4N), three anticodon bases and a discriminator base, has a lock and key relation to the corresponding amino acids. This relation can explain various general features of the universal and mitochondrian genetic codes, and therefore, could be the real molecular model of the genetic code. A beautiful matching among the amino acid- C4N complex, the hypermodified base next to the third anticodon base, and the ACC chain may be the good direct evidence for the existence of the C4N, as well as other various experimental evidences, which can easily be interpreted in terms of the C4N model.
The genetic code is comprised of a system concerning the distribution of doublets of the first two codon bases among amino acids. According to this system a definite order in the relative distribution of the first and the second codon bases coincides with a definite order among the common amino acids and their distribution for the number of hydrogen atoms per molecule (an unexpected parameter). The pattern of the relative distribution of the first and the second codon bases suggests it originated from a crystalline-like structure in which the set of bases AUGC served as an elementary structural unit and the base doublets played the role of structural analogs to the amino acids. These hypothetical crystalline-like aggregates are composed of the free molecules of amino acids and bases, and although different in their composition, should have an even number of hydrogen atoms per standard structural module.
The physical properties of amino acids were investigated in order to evaluate their possible relationship to the assignment of codons for amino acids in the genetic code. A comparison of the interconversion probability between amino acids and the distances between the amino acids for individual physical properties revealed a striking hierarchy among the physical properties. Surprisingly, it is the long-range/solvent interactions and not the short-range/stereochemical properties which are preferentially conserved in the genetic code.
The evolutionary relationships between transfer RNA (tRNA) molecules are analyzed by parsimony algorithms. The position of the topologies expected on the basis of the hypotheses made to explain the origin of the genetic code, on the frequency distribution of all the possible tree topologies of the evolutionary relationships between tRNAs seems to lead to the following conclusion: The hypothesis (Wong, J. T., Proc. Natl. Acad. Sci. USA, 1975, 72: 1909-1912) that sees the genetic code as a map of the biosynthetic relationships between amino acids seems to occupy a statistically significant position on these frequency distributions, thus reflecting a significant part of the tRNA phylogeny.
An evolutionary scheme is postulated in which the bases enter the genetic code in a definite temporal sequence and the correlated amino acids are assigned definite functions in the evolving system. The scheme requires a singlet code (guanine coding for glycine) evolving into a doublet code (guanine-cytosine doublet coding for gly (GG), ala (GC), arg (CG), pro (CC). The doublet code evolves into a triplet code. Polymerization of nucleotides is thought to have been by block polymerization rather than by a template mechanism. The proteins formed at first were simple structural peptides. No direct nucleotide-amino acid stereo-chemical interaction was required. Rather an adaptor-type indirect mechanism is thought to have been functioning since the origin.
The group I RNAs, of which the Tetrahymena ribosomal RNA intron is the most investigated example, catalyze their own splicing reactions. Splicing is initiated at a conserved site on the RNA that facilitates attack by exogenous guanosine (or its nucleotides) on the exon-intron junction. The guanosine site in the RNA's catalytic center also binds arginine, and is quite selective for the arginine side chain. This amino acid-RNA interaction is stereoselective, and L-arginine is preferred. Immediately at the site at which arginine binds there is one of only four RNA triplets in 92 group I RNA sequences: AGA/G and CGA/G. Thus the arginine contact site is within any of four different codons for arginine. Mutation of the conserved G in the middle of the triplet decreases affinity for the amino acid, showing that binding is sequence-specific. A pathway for the origin of the genetic code for arginine is suggested, based on the existence and properties of this sequence-specific, amino acid-specific RNA complex. The existence of a proto-ribosome related to the group I RNAs seems the most likely hypothesis. This notion is used to distinguish three periods in the development of the code. Restrained and exuberant hypotheses about the origin of the genetic code are distinguished, and some objections to these hypotheses are considered.
By means of an algorithm for finding rules in data, it is shown that the genetic code may be written as a codon-tree, independent of amino acid assignments. Considering this tree as a structural description, a low-complexity, context-free grammar of the code is built and its grammar complexity and grammar redundancy calculated. The relationship between the codon-tree and the hierarchy of amino acid categorizations previously introduced by the author is investigated. Interpreting the obtained code's structure as a record of its evolution, some inferences about the divergences of the code series are made.
An extensive analysis of the evolutionary relationships existing between transfer RNAs, performed using parsimony algorithms, is presented. After building up an estimate of the tRNA ancestral sequences, these sequences are then compared using certain methods. The results seem to suggest that the coevolution hypothesis (Wong, J.T., 1975, Proc. Natl. Acad. Sci. USA 72, 1909-1912) that sees the genetic code as a map of the biosynthetic relationships between amino acids is further supported by these results, as compared to the hypotheses that see the physicochemical properties of amino acids as the main adaptative theme that led to the structuring of the genetic code.
We demonstrate that serine instead of leucine is specified by the CUG codon in the yeast Candida maltosa. Evidence for this deviation from the universal genetic code was obtained by means of in vitro translation experiments. Depending on the cell-free system used, either serine, in the C. maltosa system, or leucine, in the control with the conventional wheat germ system, was found to be incorporated into the translation products of artificial CUG-containing mRNAs. Moreover, we were able to transfer the non-universal decoding of CUG to the wheat germ system by adding a tRNA fraction isolated from C. maltosa. This finding indicates the presence in C. maltosa of an unusual serine tRNA that recognizes CUG. As a consequence of the altered genetic code, expression in Saccharomyces cerevisiae of C. maltosa cytochrome P450 genes required an exchange of their CTG triplets by TCT encoding serine in order to produce the authentic proteins. In contrast, heterologous expression of the original C. maltosa genes resulted in the formation of still active but unstable enzymes probably subject to selective proteolysis in the host cells.
The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous ρ-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as ρ moves Q6→H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at α=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p≤10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.
Bovine-heart mitochondrial DNA from a single animal was isolated and fragments representative of the entire genome cloned into multicopy plasmid vectors to facilitate determination of its complete nucleotide sequence. We present here the sequence of the region covering the gene for cytochrome oxidase subunit II. Comparison of this sequence with the amino acid sequence of the homologous beef-heart protein has enabled the determination of most of the bovine mitochondrial genetic code. The code differs from the "universal" genetic code in that UGA codes for tryptophan and not termination, and AUA codes for methionine and not isoleucine. The only codon family not represented is the AGA/AGG pair normally used for arginine; evidence from other genes suggests that these code for termination in bovine mitochondria. The sequence presented also includes the adjacent tRNAAsp and tRNALys genes. The tRNAAsp gene is separated by one nucleotide from the 5' end of the COII gene and only three bases separate the 3' end of this gene and the adjacent tRNALys gene. This highly compact gene organisation is very similar to that found in the corresponding region of the human mitochondrial genome and the gene arrangement is identical. The structure of the respective bovine and human tRNAs vary primarily the "D-" and "T psi C-loops".
Disconnected recurrences of the stop signal, serine and arginine appear in the original representation of the genetic code, and of the stop signal, arginine, serine and leucine in the codon ring representation. To achieve connectedness along with structural continuity, a rook's tour representation is presented here. On the basis of structural similarities and disparities in their side groups, each of the 20 amino acids is associated with a domain comprised of from one to six contiguous squares on the chess board. As the rook moves on the chess board, it reaches all 64 squares in the ordering of the codon numbers, which prescribe the codons by a simple formula based on the position and size of the nucleotides in a triplet. Recurrences of the stop signal, arginine and serine occur naturally on the tour as the rook enters each of the latter domains for the second time. A mathematical equivalent of the rook's tour may enter as a programming device in the implementation of the code by the RNAs.
Only three tRNA genes are present within a sequenced 12.35 kbp region of the 15.8 kbp mtDNA of Chlamydomonas reinhardtii, a unicellular green alga. The corresponding tRNAs, whose anticodons are specific for TGG (Trp), CAA/G (Gln) and ATG (Met) codons, all display conventional secondary structures. The tRNA(Met) gene encodes an elongator rather than initiator species. The standard genetic code is used in C. reinhardtii mitochondria, but codon distribution is highly biased: in a collection of six identified protein coding genes, nine codons (including TGA) are not used at all, while four other sense codons occur very infrequently. In spite of the absence of certain codons, a minimum of 23 tRNAs (assuming separate initiator and elongator tRNAs(Met) are used) is needed to translate the C. reinhardtii mitochondrial genetic code. It appears unlikely that this minimal tRNA set is encoded by C. reinhardtii mtDNA.