PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A constructionist model predicting the emergence, complementarity and classification of the nucleotide bases.

We are proposing an analytical matrix that models a logical process simulating the emergence of the bases of vital nucleotides. The construction and properties of the matricial model are outlined. The Graph 1 matrix specifies a unique distribution pattern of eight terms coded in binary triplet configurations and obeys specific dynamic laws. The whole set of binary triplet configurations is in dynamic equilibrium. For that reason, it is possible to carry out an analysis by entering by any one of the eight terms on the condition that the operating mode obeys all the internal laws of orientation and symmetries inherent in the matrix. The four chemical elements at the origin of life are distributed following their atomic structures in the order hydrogen (H), carbon (C), nitrogen (N) and oxygen (O) and organized according to the model. The dynamic properties of the model necessitate the running of three successive circular periodic studies per analysis in order to show the emergence of the four bases--adenine, guanine, cytosine and thymine--precisely in that order. The fifth base of the nucleotides--uracil--shows up twice but always in an intermediate position, thus in transition, as it is the case for messenger ribonucleic acid (mRNA). We show also that the model provides for a logical explanation of the law of complementarity of the bases and their chemical classification. It is proposed that subsequent developments of the dynamic laws of this matrix may lead to the study of the logical operations for the formation of protein sequences and of their analysis and to genetic bioprogramming in general. Thus, a strictly logical and dynamic approach to molecular genetics is possible.

Base Sequence↗

Aminoacyl-tRNA synthetases, the genetic code, and the evolutionary process.

The aminoacyl-tRNA synthetases (AARSs) and their relationship to the genetic code are examined from the evolutionary perspective. Despite a loose correlation between codon assignments and AARS evolutionary relationships, the code is far too highly structured to have been ordered merely through the evolutionary wanderings of these enzymes. Nevertheless, the AARSs are very informative about the evolutionary process. Examination of the phylogenetic trees for each of the AARSs reveals the following. (i) Their evolutionary relationships mostly conform to established organismal phylogeny: a strong distinction exists between bacterial- and archaeal-type AARSs. (ii) Although the evolutionary profiles of the individual AARSs might be expected to be similar in general respects, they are not. It is argued that these differences in profiles reflect the stages in the evolutionary process when the taxonomic distributions of the individual AARSs became fixed, not the nature of the individual enzymes. (iii) Horizontal transfer of AARS genes between Bacteria and Archaea is asymmetric: transfer of archaeal AARSs to the Bacteria is more prevalent than the reverse, which is seen only for the "gemini group. " (iv) The most far-ranging transfers of AARS genes have tended to occur in the distant evolutionary past, before or during formation of the primary organismal domains. These findings are also used to refine the theory that at the evolutionary stage represented by the root of the universal phylogenetic tree, cells were far more primitive than their modern counterparts and thus exchanged genetic material in far less restricted ways, in effect evolving in a communal sense.

Amino Acids↗

Coding sequence divergence between two closely related plant species: Arabidopsis thaliana and Brassica rapa ssp. pekinensis.

To characterize the coding-sequence divergence of closely related genomes, we compared DNA sequence divergence between sequences from a Brassica rapa ssp. pekinensis EST library isolated from flower buds and genomic sequences from Arabidopsis thaliana. The specific objectives were (i) to determine the distribution of and relationship between K(a) and K(s), (ii) to identify genes with the lowest and highest K(a): K(s) values, and (iii) to evaluate how codon usage has diverged between two closely related species. We found that the distribution of K(a): K(s) was unimodal, and that substitution rates were more variable at nonsynonymous than synonymous sites, and detected no evidence that K(a) and K(s) were positively correlated. Several genes had K(a): K(s) values equal to or near zero, as expected for genes that have evolved under strong selective constraint. In contrast, there were no genes with K(a): K(s) >1 and thus we found no strong evidence that any of the 218 sequences we analyzed have evolved in response to positive selection. We detected a stronger codon bias but a lower frequency of GC at synonymous sites in A. thaliana than B. rapa. Moreover, there has been a shift in the profile of most commonly used synonymous codons since these two species diverged from one another. This shift in codon usage may have been caused by stronger selection acting on codon usage or by a shift in the direction of mutational bias in the B. rapa phylogenetic lineage.

Arabidopsis↗

An evolutionary analytical model of a complementary circular code simulating the protein coding genes, the 5' and 3' regions.

The self-complementary subset T0 = X0 [symbol: see text] ¿AAA, TTT¿ with X0 = ¿AAC, AAT, ACC, ATC, ATT, CAG, CTC, CTG, GAA, GAC, GAG, GAT, GCC, GGC, GGT, GTA, GTC, GTT, TAC, TTC¿ of 22 trinucleotides has a preferential occurrence in the frame 0 (reading frame established by the ATG start trinucleotide) of protein (coding) genes of both prokaryotes and eukaryotes. The subsets T1 = X1 [symbol: see text] ¿CCC¿ and T2 = X2 [symbol: see text] ¿GGG¿ of 21 trinucleotides have a preferential occurrence in the shifted frames 1 and 2 respectively (frame 0 shifted by one and two nucleotides respectively in the 5'-3' direction). T1 and T2 are complementary to each other. The subset T0 contains the subset X0 which has the rarity property (6 x 10(-8) to be a complementary maximal circular code with two permutated maximal circular codes X1 and X2 in the frames 1 and 2 respectively. X0 is called a C3 code. A quantitative study of these three subsets T0, T1, T2 in the three frames 0, 1, 2 of protein genes, and the 5' and 3' regions of eukaryotes, shows that their occurrence frequencies are constant functions of the trinucleotide positions in the sequences. The frequencies of T0, T1, T2 in the frame 0 of protein genes are 49, 28.5 and 22.5% respectively. In contrast, the frequencies of T0, T1, T2 in the 5' and 3' regions of eukaryotes, are independent of the frame. Indeed, the frequency of T0 in the three frames of 5' (respectively 3') regions is equal to 35.5% (respectively 38%) and is greater than the frequencies T1 and T2, both equal to 32.25% (respectively 31%) in the three frames. Several frequency asymmetries unexpectedly observed (e.g. the frequency difference between T1 and T2 in the frame 0), are related to a new property of the subset T0 involving substitutions. An evolutionary analytical model at three parameters (p, q, t) based on an independent mixing of the 22 codons (trinucleotides in frame 0) of T0 with equiprobability (1/22) followed by t approximately 4 substitutions per codon according to the proportions p approximately 0.1, q approximately 0.1 and r = 1 - p - q approximately 0.8 in the three codon sites respectively, retrieves the frequencies of T0, T1, T2 observed in the three frames of protein genes and explains these asymmetries. Furthermore, the same model (0.1, 0.1, t) after t approximately 22 substitutions per codon, retrieves the statistical properties observed in the three frames of the 5' and 3' regions. The complex behaviour of these analytical curves is totally unexpected and a priori difficult to imagine.

Animals↗

Distinct stages of protein evolution as suggested by protein sequence analysis.

Evolution of proteins encoded in nucleotide sequences began with the advent of the triplet code. The chronological order of the appearance of amino acids on the evolution scene and the steps in the evolution of the triplet code have been recently reconstructed (Trifonov, 2000b) on the basis of 40 different ranking criteria and hypotheses. According to the consensus chronology, the pair of complementary GGC and GCC codons for the amino acids alanine and glycine appeared first. Other codons appeared as complementary pairs as well, which divided their respective amino acids into two alphabets, encoded by triplets with either central purines or central pyrimidines: G, D, S, E, N, R, K, Q, C, H, Y, and W (Glycine alphabet G) and A, V, P, S, L, T, I, F, and M (Alanine alphabet A). It is speculated that the earliest polypeptide chains were very short, presumably of uniform length, belonging to two alphabet types encoded in the two complementary strands of the earliest mRNA duplexes. After the fusion of the minigenes, a mosaic of the alphabets would form. Traces of the predicted mosaic structure have been, indeed, detected in the protein sequences of complete prokaryotic genomes in the form of weak oscillations with the period 12 residues in the form of alteration of two types of 6 residue long units. The next stage of protein evolution corresponded to the closure of the chains in the loops of the size 25-30 residues (Berezovsky et al., 2000). Autocorrelation analysis of proteins of 23 complete archaebacterial and eubacterial genomes revealed that the preferred distances between valine, alanine, glycine, leucine, and isoleucine along the sequences are in the same range of 25-30 residues, indicating that the loops are primarily closed by hydrophobic interactions between the ends of the loops. The loop closure stage is followed by the formation of typical folds of 100-200 amino acids, via end-to-end fusion of the genes encoding the loop-size chains. This size was apparently dictated by the optimal ring closure for DNA. In both cases the closure into the ring (loop) rendered evolutionarily advantageous stability to the respective structures. Further gene fusions lead to the formation of modern multidomain proteins. Recombinational gene splicing is likely to have appeared after the DNA circularization stage.

Amino Acid Sequence↗

The Candida albicans CUG-decoding ser-tRNA has an atypical anticodon stem-loop structure.

In many Candida species, the leucine CUG codon is decoded by a tRNA with two unusual properties: it is a ser-tRNA and, uniquely, has guanosine at position 33 (G33). Using a combination of enzymatic (V1 RNase, RnI nuclease) and chemical (Pb(2+), imidazole) probing of the native Candida albicans ser-tRNACAG, we demonstrate that the overall tertiary structure of this tRNA resembles that of a ser-tRNA rather than a leu-tRNA, except within the anticodon arm where there is considerable disruption of the anticodon stem. Using non-modified in vitro transcripts of the C. albicans ser-tRNACAG carrying G, C, U or A at position 33, we demonstrate that it is specifically a G residue at this position that induces the atypical anticodon stem structure. Further quantitative evidence for an unusual structure in the anticodon arm of the G33-tRNA is provided by the observed change in kinetics of methylation of the G at position 37, by purified Escherichia coli m(1)G37 methyltransferase. We conclude that the anticodon arm distortion, induced by a guanosine base at position 33 in the anticodon loop of this novel tRNA, results in reduced decoding ability which has facilitated the evolution of this tRNA without extinction of the species encoding it.

Anticodon↗

The evolution of the protein synthesis system, II. From chemical evolution to biological evolution.

The sequence of events previously proposed for modern protein synthesis is reviewed. It begins with an abiological synthesis of a template, and evolves through two model autocatalytic systems to a primitive cell that has a rudimentary biological protein synthesis system. A possible scheme for the origin of tRNA's is described so as to fill the gap between the model and the modern system. Fragments of genes that existed in and around the primitive system are proposed to be precursors of tRNA's. Since these fragments must have been undesirable components for the system, the origin and evolution of tRNA's may be regarded as an excellent answer by the primitive system to adverse circumstances.

Animals↗

Specificity of protein-nucleic acid interaction and the biochemical evolution.

The water soluble carbodiimide mediated condensation of dipeptides of the general form Gly-X was carried out in the presence of mono- and poly-nucleotides. The observed yield of the tetrapeptide was found to be higher for peptide-nucleotide system of higher interaction specificity following mainly the anticodon-amino acid relationship (Basu, H.S. & Podder, S.K., 1981, Ind. J. Biochem. Biophys., 19, 251-253). The yield of the condensation product of L-peptide was more because of its higher interaction specificity. The extent of the racemization during the condensation of Gly-L-Phe, Gly-L-Tyr and Gly-D-Phe was found to be dependent on the specificity of the interaction--the higher the specificity, the lesser the racemization. The product formed was shown to have a catalytic effect on the condensation reaction. These data thus provide a mechanism showing how the specific interaction between amino acids/dipeptides and nucleic acids could lead to the formation of the 'primitive' translation machinery.

Biological Evolution↗

The evolution of the mammalian Y chromosome.

There is a predominant theory for the evolution of the mammalian Y chromosome. This theory hypothesizes that genes for sex determination and male-specific traits, as well as sequences for X-Y meiotic pairing, are conserved on the mammalian Y chromosome across all lineages and that all other Y chromosomal genes or sequences have been or will be lost in each mammalian lineage. There are effects of mouse Y chromosomal genes on behaviors and other traits that are not male specific. Under the predominant theory, these Y chromosomal genes could be the same as the conserved genes for sex determination or male-specific traits, or they could be genes that have been lost from the Y chromosomes of other mammalian lineages and that will eventually be lost from the Y chromosome of the rodent lineage. Recently, the evolution of the primate and rodent Y chromosomes has been studied at the DNA level. These studies are summarized and reviewed in this article. The findings of these studies are not fully consistent with the predominant theory for the evolution of the mammalian Y chromosome. Also, they imply that there are other possibilities for the phylogenetic history of Y chromosomal genes of mice with effects on behavior. These are that Y chromosomal genes with effects on mouse behaviors or other traits could be conserved genes other than those for sex determination or male-specific traits or that they could be novel genes on the Y chromosome of the rodent or Mus lineage.

Animals↗

Sequence similarities of protein kinase peptide substrates and inhibitors: comparison of their primary structures with immunoglobulin repeats.

Forty original sequences of peptide substrates and inhibitors of protein kinases and phosphatases were aligned in a chain matrix without artificial gaps. Fifteen protein kinase peptide substrates and inhibitors (PKSI peptides) contained a common dipeptide ArgArg and also additional important tetra-, tri- and dipeptide homologies. Three further peptide substrates were significantly similar to these peptides but lacked the ArgArg dipeptide. Sequence comparison of individual PKSI peptides revealed probabilistically restricted consensus sequence--PKSI motif--comprising 8 homologous and 13 non-randomly distributed amino acids without considering mutation analysis. This template motif was compared with the consensus sequences of 12 different immunoglobulin domains. In 11 of 12 these domains, the starts of homologous segments were found at nearly the same domain related sites, beginning with serine. A single-triplet mutation of any of the first two triplet bases that encode equally localized amino acids in each of the two sequence sets (PKSI and Ig) revealed additional homologies with the other set. A primary derived motif version composed of 9 homologous and seven non-randomly distributed amino acids was consequently established by its feedback projection into the original sequence sets. This procedure yielded a second preliminary motif version (revised motif) formed by a sequence of 9 homologous amino acids and two non-randomly distributed amino acids. In addition, three shorter oligopeptide motifs called important stereotypes were derived, based on repeated homology between Ig chains and the revised motif. The most extensive similarities in terms of these stereotypes occurred in the CH2 and CH4 domains of Ig peptides, and inhibitors of cAMP dependent protein kinase and protein kinase A. Further comparisons based on a reference sequence set arranged with the aid of feedback projection revealed a lower similarity between variable Ig chains reflected in a decreased number of homologous amino acids. Two final motif versions, FMC and FMV, were found in two different subsets of constant and variable Ig chains, respectively. FMC was composed of seven homologous and one non-randomly distributed amino acids forming the dispersed structure STLR(C)LVSD, whereas 6 homologous and one questionable amino acid constituted FMV. Only CH4 and CH1 domain segments contained all five high-incidence amino acids, which represented a higher level of similarity than homologous amino acids of all preliminary and final motifs. Four such amino acids were present also in three PKSI peptides. All similarities described here occur in domain segments positionally overlapping with the CDR1 region of variable chains. The results are discussed in terms of immunoglobulin evolution, the position of Fc receptor binding sites and degeneration or mutability of the triplets of motif-constituting amino acids.

Amino Acid Motifs↗