PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

The triplet genetic code had a doublet predecessor.

Information theoretic analysis of genetic languages indicates that the naturally occurring 20 amino acids and the triplet genetic code arose by duplication of 10 amino acids of class-II and a doublet genetic code having codons NNY and anticodons GNN. Evidence for this scenario is presented based on the properties of aminoacyl-tRNA synthetases, amino acids and nucleotide bases.

Amino Acid Motifs↗

Evolution of amino acid frequencies in proteins over deep time: inferred order of introduction of amino acids into the genetic code.

To understand more fully how amino acid composition of proteins has changed over the course of evolution, a method has been developed for estimating the composition of proteins in an ancestral genome. Estimates are based upon the composition of conserved residues in descendant sequences and empirical knowledge of the relative probability of conservation of various amino acids. Simulations are used to model and correct for errors in the estimates. The method was used to infer the amino acid composition of a large protein set in the Last Universal Ancestor (LUA) of all extant species. Relative to the modern protein set, LUA proteins were found to be generally richer in those amino acids that are believed to have been most abundant in the prebiotic environment and poorer in those amino acids that are believed to have been unavailable or scarce. It is proposed that the inferred amino acid composition of proteins in the LUA probably reflects historical events in the establishment of the genetic code.

Amino Acid Sequence↗

The code within the codons.

For the first time it is shown that each of the three codon bases has a general correlation with a different, predictable amino acid property, depending on position within the codon. In addition to the previously recognized link between the mid-base and the hydrophobic-hydrophilic spectrum, we show that, with the exception of G, the first base is generally invariant within a synthetic pathway. G--coded amino acids show a different order, being found only at the head of the synthetic pathways. The redundancy of the nature of the third base has a previously unrecognised relationship with molecular weight. The bases U and A (transversions) are associated with the most sharply defined or opposite states in both the first and second position, C somewhat less so or intermediate, anf G neutral. The apparently systematic nature of these relationships has profound implications for the origin of the genetic code. It appears to be the remains of the first language of the cell, predating the tRNA/ribosome system, persisting with remarkably little change at a deeper level of organisation than the codon language.

Amino Acid Sequence↗

Method to determine the reading frame of a protein from the purine/pyrimidine genome sequence and its possible evolutionary justification.

The periodic variations obtained by correlating the relative positions of purines and pyrimidines (and of the four bases thymine, cytosine, adenine, and guanine) in a wide variety of genomes of wholly or partly known sequence suggest that there may be enough of an earlier comma-free coding system (i.e., only readable in one frame) still present to permit determination of the reading frame and approximate extent of the present protein coding stretches. The characteristics of these variations support the hypothesis that these primitive messages were formed of coding triplets having the form RNY (R = purine; Y = pyrimidine; and N = purine or pyrimidine). The base sequences and reading frames that have a minimal deviation from such a message are still good predictors of actual coding regions and reading frames in spite of the many mutations that have occurred since such a genetic code was last in use. In fact, the right frame for almost all the proteins in a number of viruses and various prokaryotes and eukaryotes is deduced purely from purine/pyrimidine information and not by using the normal start and stop signals.

Biological Evolution↗

An expanding genetic code.

More than 30 novel amino acids have been genetically encoded in response to unique triplet and quadruplet codons including fluorescent, photoreactive and redox active amino acids, glycosylated and heavy atom derived amino acids in addition to those with keto, azido and acetylenic chains. In this article, we describe recent advances that make it possible to add new building blocks systematically to the genetic codes of bacteria, yeast and mammalian cells. Taken together these tools will enable the detailed investigation of protein structure and function, which is not possible with conventional mutagenesis. Moreover, by lifting the constraints of the existing 20-amino-acid code, it should be possible to generate proteins and perhaps entire organisms with new or enhanced properties.

Amino Acids↗

Recent emergence of the modern genetic code: a proposal.

This article proposes that the genetic code was not fully formed before the divergence of life into three kingdoms. Rather, at least arginine and tryptophan evolved after the diversification of archaea, bacteria and eukaryotes, and were spread by horizontal gene transfer. Evidence for this hypothesis is based on data suggesting that enzymes for biosynthesis of arginine and tryptophan, and for arginine tRNA ligase, have shorter divergence times than the underlying lineages. Also, many of these genes display "star" phylogenies. This proposal is an extension of the idea that the genetic code was unified because of the evolutionary pressure from horizontal gene transfer. These considerations further undermine the need to postulate the existence of a "last common ancestor"; a simpler model would be that multiple lineages gave rise to life today.

Archaea↗

The late stage of genetic code structuring took place at a high temperature.

The correlation between the optimal growth temperature of organisms and a thermophily index based on the propensity of amino acids to enter more frequently into (hyper)thermophile proteins is used to conduct an analysis aiming to establish whether genetic code structuring took place at a low or a high temperature. If the number of codons attributed to the various amino acids in the genetic code constitutes an estimate of the mean amino acid composition of proteins produced when the genetic code was definitively structured, then the thermophily index can also be associated to the genetic code. This value and the sampling of the variable thermophily index of different alignments of protein sequences from mesophile, thermophile and hyperthermophile species make it possible to establish, with an extremely high statistical confidence, that the late stage of genetic code structuring took place in a hyperthermophile (or thermophile) 'organism'. Moreover the 95% confidence interval of the temperature at which the genetic code was fixed turned out to be 91+/-24 degrees C. These observations seem to support the hypothesis that the origin of life might have taken place at a high temperature.

Algorithms↗

On error minimization in a sequential origin of the standard genetic code.

Distances between amino acids were derived from the polar requirement measure of amino acid polarity and Benner and co-workers' (1994) 74-100 PAM matrix. These distances were used to examine the average effects of amino acid substitutions due to single-base errors in the standard genetic code and equally degenerate randomized variants of the standard code. Second-position transitions conserved all distances on average, an order of magnitude more than did second-position transversions. In contrast, first-position transitions and transversions were about equally conservative. In comparison with randomized codes, second-position transitions in the standard code significantly conserved mean square differences in polar requirement and mean Benner matrix-based distances, but mean absolute value differences in polar requirement were not significantly conserved. The discrepancy suggests that these commonly used distance measures may be insufficient for strict hypothesis testing without more information. The translational consequences of single-base errors were then examined in different codon contexts, and similarities between these contexts explored with a hierarchical cluster analysis. In one cluster of codon contexts corresponding to the RNY and GNR codons, second-position transversions between C and G and transitions between C and U were most conservative of both polar requirement and the matrix-based distance. In another cluster of codon contexts, second-position transitions between A and G were most conservative. Despite the claims of previous authors to the contrary, it is shown theoretically that the standard code may have been shaped by position-invariant forces such as mutation and base content. These forces may have left heterogeneous signatures in the code because of differences in translational fidelity by codon position. A scenario for the origin of the code is presented wherein selection for error minimization could have occurred multiple times in disjoint parts of the code through a phyletic process of competition between lineages. This process permits error minimization without the disruption of previously useful messages, and does not predict that the code is optimally error-minimizing with respect to modern error. Instead, the code may be a record of genetic process and patterns of mutation before the radiation of modern organisms and organelles.

Amino Acids↗

Forces maintaining organellar genomes: is any as strong as genetic code disparity or hydrophobicity?

It remains controversial why mitochondria and chloroplasts retain the genes encoding a small subset of their constituent proteins, despite the transfer of so many other genes to the nucleus. Two candidate obstacles to gene transfer, suggested long ago, are that the genetic code of some mitochondrial genomes differs from the standard nuclear code, such that a transferred gene would encode an incorrect amino acid sequence, and that the proteins most frequently encoded in mitochondria are generally very hydrophobic, which may impede their import after synthesis in the cytosol. More recently it has been suggested that both these interpretations suffer from serious "false positives" and "false negatives": genes that they predict should be readily transferred but which have never (or seldom) been, and genes whose transfer has occurred often or early, even though this is predicted to be very difficult. Here I consider the full known range of ostensibly problematic such genes, with particular reference to the sequences of events that could have led to their present location. I show that this detailed analysis of these cases reveals that they are in fact wholly consistent with the hypothesis that code disparity and hydrophobicity are much more powerful barriers to functional gene transfer than any other. The popularity of the contrary view has led to the search for other barriers that might retain genes in organelles even more powerfully than code disparity or hydrophobicity; one proposal, concerning the role of proteins in redox processes, has received widespread support. I conclude that this abandonment of the original explanations for the retention of organellar genomes has been premature. Several other, relatively minor, obstacles to gene transfer certainly exist, contributing to the retention of relatively many organellar genes in most lineages compared to animal mtDNA, but there is no evidence for obstacles as severe as code disparity or hydrophobicity. One corollary of this conclusion is that there is currently no reason to suppose that engineering nuclear versions of the remaining mammalian mitochondrial genes, a feat that may have widespread biomedical relevance, should require anything other than sequence alterations obviating code disparity and causing modest reductions in hydrophobicity without loss of enzymatic function.

Animals↗

Designing a neural network for the constraint optimization of the fitness functions devised based on the load minimization of the genetic code.

Nonrandom patterns in codon assignments are supported by many statistical and biochemical studies in the last two decades. The canonical genetic code is known to be highly efficient in minimizing the effects of mistranslational errors and point mutations, an ability, which in term is designated "load minimization". Prior studies have included many attempts at quantitative estimation of the fraction of randomly generated codes, which in terms of load minimization, score higher than the canonical genetic code. In this study, a neural network, which estimates a highly optimized genetic code in a relatively short period of time has been devised. Several fitness functions were used throughout this text. Meanwhile, we have made use of two cost measure matrices, PAM74-100 and mutation matrix.

Algorithms↗

Four primordial modes of tRNA-synthetase recognition, determined by the (G,C) operational code.

In distinction to single-stranded anticodons built of G, C, A, and U bases, their presumable double-stranded precursors at the first three positions of the acceptor stem are composed almost invariably of G-C and C-G base pairs. Thus, the "second" operational RNA code responsible for correct aminoacylation seems to be a (G,C) code preceding the classic genetic code. Although historically rooted, the two codes were destined to diverge quite early. However, closer inspection revealed that two complementary catalytic domains of class I and class II aminoacyl-tRNA synthetases (aaRSs) multiplied by two, also complementary, G2-C71 and C2-G71 targets in tRNA acceptors, yield four (2 x 2) different modes of recognition. It appears therefore that the core four-column organization of the genetic code, associated with the most conservative central base of anticodons and codons, was in essence predetermined by these four recognition modes of the (G,C) operational code. The general conclusion follows that the genetic code per se looks like a "frozen accident" but only beyond the "2 x 2 = 4" scope. The four primordial modes of tRNA-aaRS recognition are amenable to direct experimental verification.

Amino Acyl-tRNA Synthetases↗

Amino acid composition of proteins: Selection against the genetic code.

Distribution of amino acids in 68 representative proteins is compared with their distribution among 61 codons of the genetic code. Average amounts of lysine, aspartic acid, glutamic acid, and alanine are above the levels anticipated from the genetic code, and arginine, serine, leucine, cysteine, proline, and histidine are below such levels. Arginine plus lysine account for 11.0 percent of codons and aspartic acid plus glutamic acid account for 11.3 percent; thus the average charge is roughly neutral.

Amino Acid Sequence↗

Exploring the energy landscape of the genetic code.

New insights into the arrangement of the genetic code table, based on the analysis of the physico-chemical properties of its molecular constituents, are reported in this paper. It will be demonstrated that the code has a twofold symmetry that is not apparent from the conventional code table, but becomes apparent when the codon-anticodon energies are listed for each triplet. The evolutionary development of the current code based on single base replacement mutations (transitions) from an 'iso-energetic' degenerated subset of 16 of the 64 codons is discussed. The energy landscape of all 64 codons is presented. A detailed analysis of the energy changes due to mutations in the 3rd, 1st or 2nd position of a codon reveals that the modern genetic code is highly robust. Changes come in small discrete steps that can be quantified in relation to the thermal noise of the system. The relation of the individual codon to its neighbours in the rearranged codon table can be completely understood based on thermodynamic considerations.

Biological Evolution↗

[Physico-chemical basis of the genetic code origin: stereochemical analysis of interactions of amino acids and nucleotides based on the progene hypothesis].

A progene hypothesis has been proposed earlier to explain the mechanism of origin of the self-reproducing genetic system. Progenes (precursors of the genetic system) are mixed anhydrides of an amino acid and deoxyribotrinucleotide at the 3'-gamma-terminal phosphate (NpNpNppp-AA); they are produced from dinucleotides (NpNp) and 3'-gamma-aminoacylnucleotidylates (Nppp-AA) as a result of specific interaction between amino acid and dinucleotide. The postulated mechanism of progene formation accounts for the selection of substances, including chirality, the origin of the genetic code as well as for the mechanisms of formation, self-reproduction and evolution of the simpliest genetic system ("gene--polypeptide"). A stereochemical analysis of the progene formation mechanism has allowed us to support the main statements of the hypothesis that relate to the origin of the genetic code and to selection of substances. Atomic groups that could be responsible for the specificity of interaction between dinucleotides and amino acids in progene formation have been revealed. Stereochemical evidence for the physicochemical basis of the origin of the existing genetic code have been produced: 1) a special role of the second nucleotide in the codon is demonstrated in amino acid coding by the progene hypothesis principle; 2) an advantage of T against U in such coding is demonstrated; 3) for 16 amino acids out of 20 an agreement has been obtained between the optimal dinucleotide as revealed by the stereochemical analysis and the codon dinucleotides; 4) an explanation for the third nucleotide selection mechanism is offered. A restoration of the prebiotic code, based on these results, has indicated that the code contains 32 codons, is statistical and group-wise. It encodes 7 groups of isofunctional amino acids: 3 overlapping groups of non-polar amino acids 1) medium-size hydrophobic amino acids (chiefly Val, n-Val and a-But), 2) small and medium-size non-polar amino acids (chiefly Ala Val, n-Val a-But and Gly), 3) small non-polar amino acids (Gly, Ala, a-But) and 4 groups of polar amino acids--1) hydroxy--+dicarbonic (Asp, Glu, Ser and Thr), 2) dicarbonic (Asp and Glu), 3) hydroxy (Ser and Thr) and 4) basic (Arg and Lys). The code includes about 20 amino acids among which are 15-17 canonical and a few common non-canonical. The prebiotic code explains many properties of the existing genetic code and is capable of evolving into the latter by way of a gradual replacement of the physicochemical coding mechanism by the enzymatic coding mechanism.

Amino Acids↗

Polymorphism, recombination and alternative unscrambling in the DNA polymerase alpha gene of the ciliate Stylonychia lemnae (Alveolata; class Spirotrichea).

DNA polymerase alpha is the most highly scrambled gene known in stichotrichous ciliates. In its hereditary micronuclear form, it is broken into >40 pieces on two loci at least 3 kb apart. Scrambled genes must be reassembled through developmental DNA rearrangements to yield functioning macronuclear genes, but the mechanism and accuracy of this process are unknown. We describe the first analysis of DNA polymorphism in the macronuclear version of any scrambled gene. Six functional haplotypes obtained from five Eurasian strains of Stylonychia lemnae were highly polymorphic compared to Drosophila genes. Another incompletely unscrambled haplotype was interrupted by frameshift and nonsense mutations but contained more silent mutations than expected by allelic inactivation. In our sample, nucleotide diversity and recombination signals were unexpectedly high within a region encompassing the boundary of the two micronuclear loci. From this and other evidence we infer that both members of a long repeat at the ends of the loci provide alternative substrates for unscrambling in this region. Incongruent genealogies and recombination patterns were also consistent with separation of the two loci by a large genetic distance. Our results suggest that ciliate developmental DNA rearrangements may be more probabilistic and error prone than previously appreciated and constitute a potential source of macronuclear variation. From this perspective we introduce the nonsense-suppression hypothesis for the evolution of ciliate altered genetic codes. We also introduce methods and software to calculate the likelihood of hemizygosity in ciliate haplotype samples and to correct for multiple comparisons in sliding-window analyses of Tajima's D.

Animals↗