PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Nonorthologous replacement of lysyl-tRNA synthetase prevents addition of lysine analogues to the genetic code.

Insertion of lysine during protein synthesis depends on the enzyme lysyl-tRNA synthetase (LysRS), which exists in two unrelated forms, LysRS1 and LysRS2. LysRS1 has been found in most archaea and some bacteria, and LysRS2 has been found in eukarya, most bacteria, and a few archaea, but the two proteins are almost never found together in a single organism. Comparison of structures of LysRS1 and LysRS2 complexed with lysine suggested significant differences in their potential to bind lysine analogues with backbone replacements. One such naturally occurring compound, the metabolic intermediate S-(2-aminoethyl)-L-cysteine, is a bactericidal agent incorporated during protein synthesis via LysRS2. In vitro tests showed that S-(2-aminoethyl)-L-cysteine is a poor substrate for LysRS1, and that it inhibits LysRS1 200-fold less effectively than it inhibits LysRS2. In vivo inhibition by S-(2-aminoethyl)-L-cysteine was investigated by replacing the endogenous LysRS2 of Bacillus subtilis with LysRS1 from the Lyme disease pathogen Borrelia burgdorferi. B. subtilis strains producing LysRS1 alone were relatively insensitive to growth inhibition by S-(2-aminoethyl)-L-cysteine, whereas a WT strain or merodiploid strains producing both LysRS1 and LysRS2 showed significant growth inhibition under the same conditions. These growth effects arising from differences in amino acid recognition could contribute to the distribution of LysRS1 and LysRS2 in different organisms. More broadly, these data demonstrate how diversity of the aminoacyl-tRNA synthetases prevents infiltration of the genetic code by noncanonical amino acids, thereby providing a natural reservoir of potential antibiotic resistance.

Bacillus subtilis↗

[Comparative hierarchic structure of the genetic language].

The genetical texts and genetic language are built according to hierarchic principle and contain no less than 6 levels of coding sequences, separated by marks of punctuation, separation and indication: codons, cistrons, scriptons, replicons, linkage groups, genomes. Each level has all the attributes of the language. This hierarchic system expresses some general properties and regularities. The rules of genetic language being determined, the variability of genetical texts is generated by block-modular combinatorics on each level. Between levels there are some intermediate sublevels and module types capable of being combined. The genetic language is compared with two different independent linguistic systems: human natural languages and artificial programming languages. Genetic language is a natural one by its origin, but it is a typical technical language of the functioning genetic regulatory system--by its predestination. All three linguistic systems under comparison have evident similarity of the organization principles and hierarchical structures. This argues for similarity of their principles of appearance and evolution.

Genetic Code↗

How old is the genetic code? Statistical geometry of tRNA provides an answer.

The age of the molecular organization of life as expressed in the genetic code can be estimated from experimental data. Comparative sequence analysis of transfer RNA by the method of statistical geometry in sequence space suggests that about one-third of the present transfer RNA sequence divergence was present at the urkingdom level about the time when archaebacteria separated from eubacteria. It is concluded that the genetic code is not older than, but almost as old as our planet. While this result may not be unexpected, it was not clear until now that interpretable data exist that permit inferences about such early stages of life as the establishment of the genetic code.

Anticodon↗

On the crucial stages in the origin of animate matter.

Theories of the origin of life have proposed hypotheses to link inanimate to animate matter. The theory proposed here derived the crucial stages in the origin of animate matter directly from the basic properties of inanimate matter. It asked what were the general characteristics of the link, rather than what might have been its chemical details. Life and its origin are shown to be one continuous physicochemical process of replication, random variation, and natural selection. Since life exists here and now, animate properties must have been initiated in the past somewhere. According to the theory, life originated from an as yet unknown elementary autocatalyst which occurred spontaneously, then replicated autocatalytically. As it multiplied to macroscopic abundance, its replicas gradually exhausted their reactants. Random chemical drift initiated diversity among autocatalysts. Diversity led to competition. Competition and depletion of reactants slowed down the rates of net replication of the autocatalysts. Some reached negative rates and became extinct, while those which stayed positive "survived." Thus chemical natural selection appeared, the first step in the transition from inanimate to animate matter. It initiated the first animate property, fitness, i.e., the capacity to adapt to the environment and to survive. As the environment was depleted of reactants, it was enriched with sequels-namely, with decomposition products and all other products which accompany autocatalysis. The changing environment exerted a selective pressure on autocatalysts to replace dwindling reactants by accumulating sequels. Sequels that were incorporated into the autocatalytic process became internal components of complex autocatalytic systems. Primitive forms of metabolism and organization were thus initiated. They evolved further by the same mechanism to ever higher levels of complexity, such as homochirality (handedness) and membranal enclosure. Subsequent evolution by the same mechanism generated cellular metabolism, cell division, information carriers, and a genetic code. Theories of self-organization without natural selection are refuted.

Catalysis↗

Coding in the noncoding DNA strand: A novel mechanism of gene evolution?

The question whether the noncoding DNA strand had or still has the capability for encoding functional polypeptides has been addressed in several articles. The theoretical background of the views advocating this idea arose from two groups of findings. One of them was based on various observations implying that the genetic code was adapted for double-strand coding. The other group of theories arose from the observation of gene-length overlapping open reading frames (O-ORFs) on the antisense DNA strand in a number of genes. In fact, the above theories, which I term selectionist, conceive a novel conception of gene evolution, proposing that new genes can be created by the utilization of antisense DNA strand. In contrast, neutralist theory claims that the O-ORFs are mere by-products of evolutionary processes acting to create special codon usage and base distribution patterns in the coding sequences.

Codon↗

Changes in the amino acid code.

The genetic code is characterized by a pattern arising from "wobble-pairing" between codons and anticodons, so that one nucleotide in the first anticodon position can pair with more than one nucleotide in the third position of a codon. Earlier codes may have existed in which there were fewer anticodons than at present, so that these earlier codes contained fewer amino acids. The universal code was formerly thought to be the only currently existing code used by terrestrial species. It is now known that differences exist from the universal code in mitochondrial coding systems, and also that mitochondrial systems differ from each other. These findings lend support to the proposal that archetypal codes preceded the present universal code. Such archetypal codes may have had some resemblances to mitochondrial codes.

Amino Acids↗

Relationships between genomic base content and distribution of mass in coded proteins.

The aim of this research was to examine the possible significance of genome/protein relationships in terms of effects on distribution of mass, especially in proteins. Amino acid residues in proteins have side-chains and polypeptide segments. We use "SCM" (side-chain mass), "MCM" (main-chain mass), and "deltaM" (SCM-MCM) as the deviation from "mass balance." Total MCM of the 61 amino acids in the standard code, 3412, equals total SCM: they form a mass balanced set (mean deltaM = 0). Of 14 natural variants of the code, seven have slightly positive mean deltaM values and seven have slightly negative values. Codes with the standard amino acids assigned randomly to the 20 codon sets of the standard code have about one chance in 3,300 of producing a mass balanced set. In natural proteins, as %A + T increases, the proportion of the mass in the side-chains also increases, by about half the amount calculated for standard genes with various AT/GC ratios, partly due to selection of codons with greater variability in composition at synonymous sites. For 203 representative species (including organelles), the total protein mass is distributed approximately equally between SCM and MCM (overall mean deltaM/amino acid residue, -0.06). The attainment of some overall macromolecular mass balance may have been a criterion for selecting the codon/amino acid pairs. When both structural and dynamic requirements are considered, a genetic code based on hydrophobicity and mass balance as key properties seems likely.

AT Rich Sequence↗

Widespread and ancient distribution of a noncanonical genetic code in diplomonads.

Recently, a group of diplomonads has been found to use a genetic code in which TAA and TAG encode glutamine rather than termination. To survey the distribution of this characteristic in diplomonads, we sought to identify TAA and TAG codons at positions where glutamine is expected in genes for alpha-tubulin, elongation factor-1 alpha, and the gamma subunit of eukaryotic translation initiation factor-2. These sequences show that the variant genetic code is utilized by almost all diplomonads, with the genus Giardia alone using the universal genetic code. Comparative phylogenetic analysis reveals that the switch to this genetic code took place very early in the evolution of diplomonads and was likely a single event. Termination signals and downstream untranslated regions were also cloned from three Hexamita genes. In all three of these genes, the predicted TGA termination codon was found at the expected position. Interestingly, the untranslated regions of these genes are high in AT. This is incongruent with the coding regions, which are comparatively GC-rich.

Amino Acid Sequence↗

Statistical evidence for remnants of the primordial code in the acceptor stem of prokaryotic transfer RNA.

The specificity of interaction of amino acids with triplets in the acceptor helix stem of tRNA was investigated by means of a statistical analysis of 1400 tRNA sequences. The imprint of a prototypic genetic code at position 3-5 of the acceptor helix was detected, but only for those major amino acids, glycine, alanine, aspartic acid, and valine, that are formed by spark discharges of simple gases in the laboratory. Although remnants of the code at position 3-5 are typical for tRNAs of archaebacteria, eubacteria, and chloroplasts, eukaryotes do not seem to contain this code, and mitochondria take up an intermediary position. A duplication mechanism for the transposition of the original 3-5 code toward its present position in the anticodon stem of tRNA is proposed. From this viewpoint, the mode of evolution of mRNA and functional ribosomes becomes more understandable.

Amino Acyl-tRNA Synthetases↗

On the classes of aminoacyl-tRNA synthetases and the error minimization in the genetic code.

As a consequence of the existence of two classes of aminoacyl-tRNA synthetases (aaRSs), we defined two types of mutations: g (mutations that do not change the class of the involved amino acids) and u (those which change the class). We have found that the mean chemical distance resulting from g mutations is smaller than that corresponding to u mutations, indicating that g mutations are responsible for most of the known minimization of the genetic code. This supports models for the origin and evolution of the code, in which new amino acids were added after duplications or modification of existing aaRSs.

Amino Acids↗

A non-canonical genetic code in an early diverging eukaryotic lineage.

The nearly invariant nature of the 'Universal Genetic Code' attests to its early establishment in evolution and to the difficulty of altering it now, since so many molecules are required for, and depend upon, faithful translation. Nevertheless, variations on the universal code are known in a handful of genomes. We have found one such variant in diplomonads, an early-diverging eukaryotic lineage. Genes for alpha-tubulin, beta-tubulin and elongation factor 1 alpha (EF-1alpha) from two unclassified strains of Hexamitidae were found to contain TAA and TAG (TAR) triplets at positions suggesting a variant code in which TAR codes for glutamine. We found confirmation of this hypothesis by identifying genes encoding glutamine-tRNAs with CUA and UUA anticodons. The alpha-tubulin and EF-1alpha genes from two other diplomonads, Spironucleus muris and Hexamita inflata, were also sequenced and shown to contain no such non-canonical codons. However, tRNA genes with the anticodons UUA and CUA were found in H.inflata, suggesting that this diplomonad also uses these codons, albeit infrequently. The high GC content of these genomes and the presence of two isoaccepting tRNAs compound the difficulty of understanding how this variant code arose by strictly neutral means.

Amino Acid Sequence↗

Characterisation of a non-canonical genetic code in the oxymonad Streblomastix strix.

The genetic code is one of the most highly conserved characters in living organisms. Only a small number of genomes have evolved slight variations on the code, and these non-canonical codes are instrumental in understanding the selective pressures maintaining the code. Here, we describe a new case of a non-canonical genetic code from the oxymonad flagellate Streblomastix strix. We have sequenced four protein-coding genes from S.strix and found that the canonical stop codons TAA and TAG encode the amino acid glutamine. These codons are retained in S.strix mRNAs, and the legitimate termination codons of all genes examined were found to be TGA, supporting the prediction that this should be the only true stop codon in this genome. Only four other lineages of eukaryotes are known to have evolved non-canonical nuclear genetic codes, and our phylogenetic analyses of alpha-tubulin, beta-tubulin, elongation factor-1 alpha (EF-1 alpha), heat-shock protein 90 (HSP90), and small subunit rRNA all confirm that the variant code in S.strix evolved independently of any other known variant. The independent origin of each of these codes is particularly interesting because the code found in S.strix, where TAA and TAG encode glutamine, has evolved in three of the four other nuclear lineages with variant codes, but this code has never evolved in a prokaryote or a prokaryote-derived organelle. The distribution of non-canonical codes is probably the result of a combination of differences in translation termination, tRNAs, and tRNA synthetases, such that the eukaryotic machinery preferentially allows changes involving TAA and TAG.

Amino Acid Sequence↗

Dramatic events in ciliate evolution: alteration of UAA and UAG termination codons to glutamine codons due to anticodon mutations in two Tetrahymena tRNAs.

The three major glutamine tRNAs of Tetrahymena thermophila were isolated and their nucleotide sequences determined by post-labeling techniques. Two of these tRNAs show unusual codon recognition: a previously isolated tRNA(UmUA) and a second species with CUA in the anticodon (tRNA(CUA)). These two tRNAs recognize two of the three termination codons on natural mRNAs in a reticulocyte system. tRNA(UmUA) reads the UAA codon of alpha-globin mRNA and the UAG codon of tobacco mosaic virus (TMV) RNA, whereas tRNA(CUA) recognizes only UAG. This indicates that Tetrahymena uses UAA and UAG as glutamine codons and that UGA may be the only functional termination codon. A notable feature of these two tRNAs is their unusually strong readthrough efficiency, e.g. purified tRNA(CUA) achieves complete readthrough over the UAG stop codon of TMV RNA. The third major tRNA of Tetrahymena has a UmUG anticodon and presumably reads the two normal glutamine codons CAA and CAG. The sequence homology between tRNA(UmUG) and tRNA(UmUA) is 81%, whereas that between tRNA(CUA) and tRNA(UmUA) is 95%, indicating that the two unusual tRNAs evolved from the normal tRNA early in ciliate evolution. Possible events leading to an altered genetic code in ciliates are discussed.

Journal Article↗

Assembly of a class I tRNA synthetase from products of an artificially split gene.

The aminoacyl-tRNA synthetases arose early in evolution and established the rules of the genetic code through their specific interactions with amino acids and RNA molecules. About half of these tRNA charging enzymes are class I synthetases, which contain similar N-terminal nucleotide-fold-like structures that are joined to variable domains implicated in specific protein-tRNA contacts. Here, we show that a bacterial synthetase gene can be split into two nonoverlapping segments. We split the gene for Escherichia coli methionyl-tRNA synthetase (a class I synthetase) at several sites near the interdomain junction, such that one segment codes for the nucleotide-fold-containing domain and the other provides determinants for tRNA recognition. When the segments are folded together, they can recognize and charge tRNA, both in vivo and in vitro. We postulate that an early step in the assembly of systems to attach amino acids to specific RNA molecules may have involved specific interactions between discrete proteins that is reflected in the interdomain contacts of modern synthetases.

Amino Acid Sequence↗

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus↗

[The genetic language: grammar, semantics, evolution].

The genetic language is a collection of rules and regularities of genetic information coding for genetic texts. It is defined by alphabet, grammar, collection of punctuation marks and regulatory sites, semantics. There is a review of these general attributes of genetic language, including also the problems of synonymy and evolution. The main directions of theoretical investigations of genetic language and neighbouring questions are formulated: (1) cryptographic problems, (2) analysis of genetic texts, (3) theoretical-linguistic problems, (4) evolutionary linguistic questions. The problem of genetic language becomes one of the key ones of molecular genetics, molecular biology and gene engineering.

Codon↗