PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The genetic code is one in a million.

Statistical and biochemical studies of the genetic code have found evidence of nonrandom patterns in the distribution of codon assignments. It has, for example, been shown that the code minimizes the effects of point mutation or mistranslation: erroneous codons are either synonymous or code for an amino acid with chemical properties very similar to those of the one that would have been present had the error not occurred. This work has suggested that the second base of codons is less efficient in this respect, by about three orders of magnitude, than the first and third bases. These results are based on the assumption that all forms of error at all bases are equally likely. We extend this work to investigate (1) the effect of weighting transition errors differently from transversion errors and (2) the effect of weighting each base differently, depending on reported mistranslation biases. We find that if the bias affects all codon positions equally, as might be expected were the code adapted to a mutational environment with transition/transversion bias, then any reasonable transition/transversion bias increases the relative efficiency of the second base by an order of magnitude. In addition, if we employ weightings to allow for biases in translation, then only 1 in every million random alternative codes generated is more efficient than the natural code. We thus conclude not only that the natural genetic code is extremely efficient at minimizing the effects of errors, but also that its structure reflects biases in these errors, as might be expected were the code the product of selection.

Amino Acids↗

An analytical model of gene evolution with six mutation parameters: an application to archaeal circular codes.

We develop here an analytical evolutionary model based on a trinucleotide mutation matrix 64 x 64 with six substitution parameters associated with the transitions and transversions in the three trinucleotide sites. It generalizes the previous models based on the nucleotide mutation matrices 4 x 4 and the trinucleotide mutation matrix 64 x 64 with three parameters. It determines at some time t the exact occurrence probabilities of trinucleotides mutating randomly according to six substitution parameters. An application of this model allows an evolutionary study of the common circular code COM and the 15 archaeal circular codes X which have been recently identified in several archaeal genomes. The main property of a circular code is the retrieval of the reading frames in genes, both locally, i.e. anywhere in genes and in particular without a start codon, and automatically with a window of a few nucleotides. In genes, the circular code is superimposed on the traditional genetic one. Very unexpectedly, the evolutionary model demonstrates that the archaeal circular codes can derive from the common circular code subjected to random substitutions with particular values for six substitutions parameters. It has a strong correlation with the statistical observations of three archaeal codes in actual genes. Furthermore, the properties of these substitution rates allow proposal of an evolutionary classification of the 15 archaeal codes into three main classes according to this model. In almost all the cases, they agree with the actual degeneracy of the genetic code with substitutions more frequent in the third trinucleotide site and with transitions more frequent that transversions in any trinucleotide site.

Archaea↗

Periodical changes of amino acid reactivity within the genetic code.

Enthalpies (delta H++) and entropies (delta S++) of activation for the reaction of 18 N'-hydroxysuccinimide esters of N-protected proteinaceous amino acids with p-anisidine were measured and free enthalpies of activation (delta G++) at 25 degrees C were calculated on this basis. A regular correlation between delta G++s and the corresponding amino acid codons was found. To obtain this correlation all the codons had to be arranged in a closed ring in which the consecutive codons were connected by one-step mutational changes. One-step mutations appeared as a regular series: 2,3,3,3,1,3,3,3,1,3,3,3,1,3,3,3,2,3,3,3. (the numbers denote a codon position in which a change took place). There were three such 'one-step mutation periods' in the ring, each containing 20 codons (in each block of 16 codons with A, U and C, in the central position and 4 codons containing G in the central position). The end of the third period (UG) and the beginning of the first period were bridged by the four codons of glycine with G in the second position. The values of delta G++ change similarly in each period, increasing upon approaching Lys, Pro, and Ile. The periodical relation between the chemical reactivities of the coded amino acids (reflected by delta G++s) and the structure of their codons could be of importance for the origin of the genetic code i.e. for selection of proper codons for the definite amino acids.

Amino Acid Sequence↗

An expanded genetic code with a functional quadruplet codon.

With few exceptions the genetic codes of all known organisms encode the same 20 amino acids, yet all that is required to add a new building block are a unique tRNA/aminoacyl-tRNA synthetase pair, a source of the amino acid, and a unique codon that specifies the amino acid. For example, the amber nonsense codon, TAG, together with orthogonal Methanococcus jannaschii or Escherichia coli tRNA/synthetase pairs have been used to genetically encode a variety of unnatural amino acids in E. coli and yeast, respectively. However, the availability of noncoding triplet codons ultimately limits the number of amino acids encoded by any organism. Here, we report the design and generation of an orthogonal synthetase/tRNA pair derived from archaeal tRNA(Lys) sequences that efficiently and selectively incorporates an unnatural amino acid into proteins in response to the quadruplet codon, AGGA. Frameshift suppression with L-homoglutamine (hGln) does not significantly affect protein yields or cell growth rates and is mutually orthogonal with amber suppression, permitting the simultaneous incorporation of two unnatural amino acids, hGln and O-methyl-L-tyrosine, at distinct positions within myoglobin. This work suggests that neither the number of available triplet codons nor the translational machinery itself represents a significant barrier to further expansion of the genetic code.

Amino Acids↗

Do universal codon-usage patterns minimize the effects of mutation and translation error?

BACKGROUND: Do species use codons that reduce the impact of errors in translation or replication? The genetic code is arranged in a way that minimizes errors, defined as the sum of the differences in amino-acid properties caused by single-base changes from each codon to each other codon. However, the extent to which organisms optimize the genetic messages written in this code has been far less studied. We tested whether codon and amino-acid usages from 457 bacteria, 264 eukaryotes, and 33 archaea minimize errors compared to random usages, and whether changes in genome G+C content influence these error values. RESULTS: We tested the hypotheses that organisms choose their codon usage to minimize errors, and that the large observed variation in G+C content in coding sequences, but the low variation in G+U or G+A content, is due to differences in the effects of variation along these axes on the error value. Surprisingly, the biological distribution of error values has far lower variance than randomized error values, but error values of actual codon and amino-acid usages are actually greater than would be expected by chance. CONCLUSION: These unexpected findings suggest that selection against translation error has not produced codon or amino-acid usages that minimize the effects of errors, and that even messages with very different nucleotide compositions somehow maintain a relatively constant error value. They raise the question: why do all known organisms use highly error-minimizing genetic codes, but fail to minimize the errors in the mRNA messages they encode?

Animals↗

On the origin of protein biosynthesis.

There is a very close steric relationship between the codon-anticodon site which accounts for the genetic code dictionary and a polynucleotide replicase site. Protein biosynthesis must therefore have arisen out of a primaeval polynucleotide replicase system.

Anticodon↗

Growth function of self-complementary circular codes.

In several papers Arquès and Michel studied the maximal circular codes consisting of words of length 3 (or trinucleotides) on the genetic alphabet {A, C, G, T}. We present here some additional information on these codes. In particular, we study the growth function of the self-complementary circular codes and we prove that among them exactly 528 are maximal.

Algorithms↗

The emergence of genetic coding in physical systems.

A simple model of molecular biological translation, based on the classification of polymers as either information carriers or functional catalysts, is used to analyse formal constraints on physical systems which utilise genetic coding. We investigate (i) how the structure-function relationship for coding assignment catalysts constrains the selection of genetic information which can sustain functional self-organisation and (ii) what general prerequisites must be satisfied for selection to give rise to an increase in functional complexity. This is done by considering two separate alphabets and defining the complete set of assignments from letters of one alphabet onto letters from the other. A code is defined as a set of assignments which maps each letter from the first alphabet onto a letter from the second alphabet. We enumerate all the embeddings of the assignment functions in the minimal sequence space of strings of letters from the second alphabet and demonstrate how the embeddings can be classified according to whether they allow different codes to be represented unambiguously in the minimal sequence space of strings of letters from the first alphabet. Non-minimal embeddings are also discussed. Finally, we consider how the mutual specification of letters of the two alphabets and assignment functions can be decomposed into more highly differentiated classes. Only a certain class of embeddings allows coding to be preserved under decomposition. We conclude that the evolution of increasing coding complexity can take place only when special conditions are satisfied regarding the structure-function relationship for the coding assignment catalysts.

Animals↗

Codon usage bias and mutation constraints reduce the level of error minimization of the genetic code.

Studies on the origin of the genetic code compare measures of the degree of error minimization of the standard code with measures produced by random variant codes but do not take into account codon usage, which was probably highly biased during the origin of the code. Codon usage bias could play an important role in the minimization of the chemical distances between amino acids because the importance of errors depends also on the frequency of the different codons. Here I show that when codon usage is taken into account, the degree of error minimization of the standard code may be dramatically reduced, and shifting to alternative codes often increases the degree of error minimization. This is especially true with a high CG content, which was probably the case during the origin of the code. I also show that the frequency of codes that perform better than the standard code, in terms of relative efficiency, is much higher in the neighborhood of the standard code itself, even when not considering codon usage bias; therefore alternative codes that differ only slightly from the standard code are more likely to evolve than some previous analyses suggested. My conclusions are that the standard genetic code is far from being an optimum with respect to error minimization and must have arisen for reasons other than error minimization.

Base Composition↗

Intramolecular interactions in aminoacyl cyclic-3',5'-nucleotides.

Polymerization of amino-acid acyl cyclic-3',5'-nucleotides is postulated to be the origin of RNA and associated protein in prebiotic molecular evolution. The enthalpy change in the intramolecular interaction between the nucleotide base and the amino-acid side chain determines the stability of the particular complex, resulting in a preferred association (or coding) of a base for a particular amino acid. The compounds studied were glycine acyl cyclic-3',5'-guanylate where the strong hydrogen bond between protonated glycine and guanine N7 gives an enthalpy change of -0.05 h. Similarly, hydrogen bonds in l-lysine acyl cyclic-3',5'-adenylate give an enthalpy change of -0.06 h. Hydrophobic interactions in l-phenylalanine acyl cyclic-3',5'-uridylate give an enthalpy change of -0.02 h and the corresponding value for l-proline acyl cyclic-3',5'-cytidylate is -0.01 h. These interactions were expected to be modified as the genetic code became a duplet and finally a triplet code. The interactions have been shown to be feasible from the overall enthalpy changes in the ZKE approximation at the MP2/6-31G* level.

Amino Acids↗

Genetic code redundancy and the evolutionary stability of protein secondary structure.

The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. The extent to which these biases are expressed in protein secondary structure is described through the analysis of the overall amino acid compositions of the alpha-helix, beta-sheet, beta-turn and random coil segments elucidated by X-ray crystallography. Given the concept of neutral mutation in proteins, the allocation of synonyms in the genetic code appears to protect secondary structures from amino acid changes and discourages the appearance of chemically complex residues. The level of protection is similar for each structural form, despite their clear preferences for certain amino acids. The organization of the code is therefore relevant to the preservation of conformation seen in the evolution of many protein families.

Amino Acids↗

Genetic code synonym quotas and amino acid complexity: cutting the cost of proteins?

The synonym quotas within the genetic code for the 20 common amino acids are examined in relation to the ways in which these amino acids can be marshalled into different sets on the basis of shared physico-chemical properties. This reveals which shared properties are encouraged or discouraged during the course of protein evolution by the arrangement of the code. A dominant theme is that the synonym quotas are allocated in favour of small and chemically uncomplicated residues, and to the disadvantage of large and chemically prominent ones. From amongst the various measurements that can be considered to quantitatively express aspects of amino acid residue "size and complexity" (e.g. side chain volume, bulkness and formula weight), formula weight has the highest correlation with the synonym quota for each amino acid. However, the correlation is weak. A specially derived "size/complexity" scale for the amino acids based on their relative atomic composition improved the correlation only marginally. The existence of another weak correlation between the synonym quotas and the general amino acid composition of proteins prompted an investigation of the correlations between this composition and the previously considered amino acid properties. Again, the highest correlations are with amino acid formula weight and "size/complexity", but in this instance the correlations are high enough to be truly significant. It is suggested that the biased synonym quotas in the genetic code are intended to ensure that proteins as a whole maintain a certain amino acid composition, even to the extent that the quotas include compensatory biases to counter opposing influences upon this composition caused by the processes of natural selection for protein function. It is the need for these compensatory biases that prevents a simple correlation between the quotas and measures of amino acid complexity. The final outcome, in which amino acids are deployed in functional proteins in approximate proportion to their chemical complexity may serve both as a means of minimising the negative consequences of random genetic mutation (by reducing the chance appearance of the more "disruptive" types of side chain in proteins) and as a means of ensuring the most economic use of biosynthetic resources. According to this reasoning, the code is not a "frozen accident"; it is universally appropriate because it provides the best compromise that can be achieved between biosynthetic cost and biological return in respect of the rate of protein evolution.

Amino Acids↗

Evolution of phage with chemically ambiguous proteomes.

BACKGROUND: The widespread introduction of amino acid substitutions into organismal proteomes has occurred during natural evolution, but has been difficult to achieve by directed evolution. The adaptation of the translation apparatus represents one barrier, but the multiple mutations that may be required throughout a proteome in order to accommodate an alternative amino acid or analogue is an even more daunting problem. The evolution of a small bacteriophage proteome to accommodate an unnatural amino acid analogue can provide insights into the number and type of substitutions that individual proteins will require to retain functionality. RESULTS: The bacteriophage Qbeta initially grows poorly in the presence of the amino acid analogue 6-fluorotryptophan. After 25 serial passages, the fitness of the phage on the analogue was substantially increased; there was no loss of fitness when the evolved phage were passaged in the presence of tryptophan. Seven mutations were fixed throughout the phage in two independent lines of descent. None of the mutations changed a tryptophan residue. CONCLUSIONS: A relatively small number of mutations allowed an unnatural amino acid to be functionally incorporated into a highly interdependent set of proteins. These results support the 'ambiguous intermediate' hypothesis for the emergence of divergent genetic codes, in which the adoption of a new genetic code is preceded by the evolution of proteins that can simultaneously accommodate more than one amino acid at a given codon. It may now be possible to direct the evolution of organisms with novel genetic codes using methods that promote ambiguous intermediates.

Allolevivirus↗

[Evolutionary changes in the genetic code, predictable on basis of the hypothesis of physical predetermination of the structure of codon bases].

According to the earlier proposed hypothesis on the structural correspondence between amino acids and doublets from the first codon bases (Sukhodolets 1980), the UGA triplet corresponds to tryptophan and the AGX triplets - to the termination codons. It is notably this sense of the UGA and AGA, AGG, respectively, that was reported for mitochondrial codes. Thereby, a proposal is indirectly confirmed that meanings of the UGA (nonsense) and AGA, AGG (arginine) in the normal cytoplasmic code is the result of evolutionary changes.

Amino Acids↗

Early fixation of an optimal genetic code.

The evolutionary forces that produced the canonical genetic code before the last universal ancestor remain obscure. One hypothesis is that the arrangement of amino acid/codon assignments results from selection to minimize the effects of errors (e.g., mistranslation and mutation) on resulting proteins. If amino acid similarity is measured as polarity, the canonical code does indeed outperform most theoretical alternatives. However, this finding does not hold for other amino acid properties, ignores plausible restrictions on possible code structure, and does not address the naturally occurring nonstandard genetic codes. Finally, other analyses have shown that significantly better code structures are possible. Here, we show that if theoretically possible code structures are limited to reflect plausible biological constraints, and amino acid similarity is quantified using empirical data of substitution frequencies, the canonical code is at or very close to a global optimum for error minimization across plausible parameter space. This result is robust to variation in the methods and assumptions of the analysis. Although significantly better codes do exist under some assumptions, they are extremely rare and thus consistent with reports of an adaptive code: previous analyses which suggest otherwise derive from a misleading metric. However, all extant, naturally occurring, secondarily derived, nonstandard genetic codes do appear less adaptive. The arrangement of amino acid assignments to the codons of the standard genetic code appears to be a direct product of natural selection for a system that minimizes the phenotypic impact of genetic error. Potential criticisms of previous analyses appear to be without substance. That known variants of the standard genetic code appear less adaptive suggests that different evolutionary factors predominated before and after fixation of the canonical code. While the evidence for an adaptive code is clear, the process by which the code achieved this optimization requires further attention.

Amino Acids↗