Origin of the genetic code: a physical-chemical model of primitive codon assignments.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The genetic code has been influenced by directional mutation pressure affecting the base composition of DNA, sometimes in the direction of increased GC content and at other times, in the direction of AT. Such pressure led to changes in species-specific usages of codons and tRNA anticodons, and also in amino acid assignments of codons in mitochondria and in several intact organisms. These code changes are probably recent evolutionary events. The genetic code is not 'frozen', but instead it is still evolving.
Enthalpies (delta H++) and entropies (delta S++) of activation for the reaction of 18 N'-hydroxysuccinimide esters of N-protected proteinaceous amino acids with p-anisidine were measured and free enthalpies of activation (delta G++) at 25 degrees C were calculated on this basis. A regular correlation between delta G++s and the corresponding amino acid codons was found. To obtain this correlation all the codons had to be arranged in a closed ring in which the consecutive codons were connected by one-step mutational changes. One-step mutations appeared as a regular series: 2,3,3,3,1,3,3,3,1,3,3,3,1,3,3,3,2,3,3,3. (the numbers denote a codon position in which a change took place). There were three such 'one-step mutation periods' in the ring, each containing 20 codons (in each block of 16 codons with A, U and C, in the central position and 4 codons containing G in the central position). The end of the third period (UG) and the beginning of the first period were bridged by the four codons of glycine with G in the second position. The values of delta G++ change similarly in each period, increasing upon approaching Lys, Pro, and Ile. The periodical relation between the chemical reactivities of the coded amino acids (reflected by delta G++s) and the structure of their codons could be of importance for the origin of the genetic code i.e. for selection of proper codons for the definite amino acids.
There is a very close steric relationship between the codon-anticodon site which accounts for the genetic code dictionary and a polynucleotide replicase site. Protein biosynthesis must therefore have arisen out of a primaeval polynucleotide replicase system.
Explore the source record for details and available documents.
The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. The extent to which these biases are expressed in protein secondary structure is described through the analysis of the overall amino acid compositions of the alpha-helix, beta-sheet, beta-turn and random coil segments elucidated by X-ray crystallography. Given the concept of neutral mutation in proteins, the allocation of synonyms in the genetic code appears to protect secondary structures from amino acid changes and discourages the appearance of chemically complex residues. The level of protection is similar for each structural form, despite their clear preferences for certain amino acids. The organization of the code is therefore relevant to the preservation of conformation seen in the evolution of many protein families.
According to the earlier proposed hypothesis on the structural correspondence between amino acids and doublets from the first codon bases (Sukhodolets 1980), the UGA triplet corresponds to tryptophan and the AGX triplets - to the termination codons. It is notably this sense of the UGA and AGA, AGG, respectively, that was reported for mitochondrial codes. Thereby, a proposal is indirectly confirmed that meanings of the UGA (nonsense) and AGA, AGG (arginine) in the normal cytoplasmic code is the result of evolutionary changes.
Explore the source record for details and available documents.
Preliminary amino acid sequence data on the transplantation antigens of mouse and man have led to provocative hypotheses about the genetic organization and evolution of genes coded by the major histocompatibility complex of mammals. New microsequencing techniques should permit a detailed analysis of these gene products and an eventual choice among the alternative hypotheses now posed. These data have made it apparent that the H-2 complex is a fascinating and complicated chromosomal region which will continue for some time to intrigue immunologists, geneticists, biochemists, and cell biologists.
This paper analyzes the relationships between the genetic code coevolution hypothesis and the physicochemical hypothesis by means of a comparative study of the precursor-product amino acid pairs on which the former hypothesis is based. Even if the coevolution between the biosynthetic relationships of amino acids and the organization of the genetic code is not questioned in this paper, the results and the arguments used lead us to believe that the selective pressures considered essential by the physicochemical postulates, played a more active role than that of the precursor-product relationships in defining the allocation of these amino acids in the genetic code. It is furthermore pointed out that the two evolutionary hypothesis might be aspects of the same selective pressure, and thus difficult to differentiate.
For the first time it is shown that each of the three codon bases has a general correlation with a different, predictable amino acid property, depending on position within the codon. In addition to the previously recognized link between the mid-base and the hydrophobic-hydrophilic spectrum, we show that, with the exception of G, the first base is generally invariant within a synthetic pathway. G--coded amino acids show a different order, being found only at the head of the synthetic pathways. The redundancy of the nature of the third base has a previously unrecognised relationship with molecular weight. The bases U and A (transversions) are associated with the most sharply defined or opposite states in both the first and second position, C somewhat less so or intermediate, anf G neutral. The apparently systematic nature of these relationships has profound implications for the origin of the genetic code. It appears to be the remains of the first language of the cell, predating the tRNA/ribosome system, persisting with remarkably little change at a deeper level of organisation than the codon language.
The periodic variations obtained by correlating the relative positions of purines and pyrimidines (and of the four bases thymine, cytosine, adenine, and guanine) in a wide variety of genomes of wholly or partly known sequence suggest that there may be enough of an earlier comma-free coding system (i.e., only readable in one frame) still present to permit determination of the reading frame and approximate extent of the present protein coding stretches. The characteristics of these variations support the hypothesis that these primitive messages were formed of coding triplets having the form RNY (R = purine; Y = pyrimidine; and N = purine or pyrimidine). The base sequences and reading frames that have a minimal deviation from such a message are still good predictors of actual coding regions and reading frames in spite of the many mutations that have occurred since such a genetic code was last in use. In fact, the right frame for almost all the proteins in a number of viruses and various prokaryotes and eukaryotes is deduced purely from purine/pyrimidine information and not by using the normal start and stop signals.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Distribution of amino acids in 68 representative proteins is compared with their distribution among 61 codons of the genetic code. Average amounts of lysine, aspartic acid, glutamic acid, and alanine are above the levels anticipated from the genetic code, and arginine, serine, leucine, cysteine, proline, and histidine are below such levels. Arginine plus lysine account for 11.0 percent of codons and aspartic acid plus glutamic acid account for 11.3 percent; thus the average charge is roughly neutral.