PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Multiplicity of species in some replicative systems.

In an attempt to explain the uniqueness of the coding mechanism of living cells as contrasted with the multispecies structure of ecosystems we examine two models of individuals with some replicative properties. In the first model the system generically remains in a multispecies state. Even though for some of these species the replicative probability is very high, they are unable to invade the system. In the second model, in which the death rate depends on the type of the species, the system relatively quickly reaches a single-species state and fluctuations might at most bring it to yet another single-species state.

Animals↗

On the syntactic structure and redundancy distribution of the genetic code.

By means of an algorithm for finding rules in data, it is shown that the genetic code may be written as a codon-tree, independent of amino acid assignments. Considering this tree as a structural description, a low-complexity, context-free grammar of the code is built and its grammar complexity and grammar redundancy calculated. The relationship between the codon-tree and the hierarchy of amino acid categorizations previously introduced by the author is investigated. Interpreting the obtained code's structure as a record of its evolution, some inferences about the divergences of the code series are made.

Algorithms↗

[Do repetitive DNA sequences have a biological function?].

By DNA reassociation kinetics it is known that the eucaryotic genome consists of non-repetitive DNA, middle-repetitive DNA and highly repetitive DNA. Whereas the majority of protein-coding genes is located on non-repetitive DNA, repetitive DNA forms a constitutive part of eucaryotic DNA and its amount in most cases equals or even substantially exceeds that of non-repetitive DNA. During the past years a large body of data on repetitive DNA has accumulated and these have prompted speculations ranging from specific roles in the regulation of gene expression to that of a selfish entity with inconsequential functions. The following article summarizes recent findings on structural, transcriptional and evolutionary aspects and, although by no means being proven, some possible biological functions are discussed.

Animals↗

Identification and simulation of shifted periodicities common to protein coding genes of eukaryotes, prokaryotes and viruses.

The distribution of nucleotides in protein coding genes is studied with autocorrelation functions. The autocorrelation function YRY(N)iYRY, analysing the occurrence probability of the i-motif YRY(N)iYRY (two motifs YRY separated by any i bases N, R = purine = Adenine or Guanine, Y = pyrimidine = Cytosine or Thymine, N = R or Y) in the protein coding genes of eukaryotes, prokaryotes and viruses, reveals the classical periodicity 0 modulo 3 associated with the normal frame 0 (maximal values of the function at i = 0, 3, 6, etc). The specification of YRY(N)iYRY on the alphabet [A, C, G, T] leads to 64 i-motifs: CAC(N)iCAC, CAC(N)iCAT, ..., TGF(N)iTGT. The 64 autocorrelation functions associated with these 64 i-motifs in protein coding genes have all the periodicity modulo 3, but, surprisingly, not always the expected periodicity 0 modulo 3. Two new types of periodicities are identified: a periodicity 1 modulo 3 associated with the shifted frame +1 (maximal values of the function at i = 1, 4, 7, etc) and a periodicity 2 modulo 3 associated with the shifted frame -1 (maximal values of the function at i = 2, 5, 8 etc). Furthermore, the classification of i-motifs according to the type of periodicity demonstrates a strong coherence relation between the 64 i-motifs, which is, in addition, common to the three gene populations, as the same i-motifs in the three gene populations have the same periodicities. The three periodicities 0, 1 and 2 modulo 3 can be simulated by an evolutionary model at two successive processes. The simulated genes are generated by a process of gene construction, with a stochastic automaton followed by a process of gene evolution with random insertions and deletions of trinucleotides simulating RNA editing. For almost all i-motifs, the autocorrelation functions in these simulated genes are strongly correlated with those in protein coding genes, for both the type and the probability level of periodicities. This paper describes the process of ribosomal frameshifting leading to the shifted periodicities, which may reveal overlapping genes or concatenated genes from different frames. It also presents the evolutionary aspects of the shifted periodicities. The shifted periodicities cannot be associated with the RNY model (Eigen & Schuster, 1978, Naturwissenschaften 65, 341-369) or the RRY model (Crick et al., 1976, Origins of Life 7, 389-397), but are compatible with the oligonucleotide mixing model (Arquès & Michel, 1990, Bull. math. Biol. 52, 741-772). Finally, a variant of the primitive translation model of Crick et al. (1976) is proposed to explain the shifted periodicities.

Animals↗

An evolutionary analytical model of a complementary circular code.

The subset X0=[AAC,AAT,ACC,ATC,ATT,CAG,CTC,CTG, GAA,GAC,GAG,GAT,GCC,GGC,GGT,GTA,GTC,GTT,TAC,TTC] of 20 trinucleotides has a preferential occurrence in the frame 0 (reading frame established by the ATG start trinucleotide) of protein (coding) genes of both prokaryotes and eukaryotes. This subset X0 is a complementary maximal circular code with two permutated maximal circular codes X1 and X2 in the frames 1 and 2 respectively (frame 0 shifted by one and two nucleotides respectively in the 5'-3' direction). X0 is called a C3 code (Arquès and Michel, 1997, J. Biosyst 44, 107-134). A quantitative study of these three subsets X0, X1 and X2 in the three frames 0, 1 and 2 of eukaryotic protein genes shows that their occurrence frequencies are constant functions of the trinucleotide positions in the sequences. The frequencies of X0, X1 and X2 in the frame 0 of eukaryotic protein genes are 48.5%, 29% and 22.5% respectively. These properties are not observed in the 5' and 3' regions of eukaryotes where X0, X1 and X2 occur with variable frequencies around the random value (1/3). Several frequency asymmetries unexpectedly observed, e.g. the frequency difference between X1 and X2 in the frame 0, are related to a new property of the C3 code X0 involving substitutions. An evolutionary analytical model at three parameters (p, q, t) based on an independent mixing of the 20 codons (trinucleotides in the frame 0) of X0 with equiprobability (1/20) followed by t approximately 4 substitutions per codon according to the proportions p approximately 0.1, q approximately 0.1 and r = 1 - p - q approximately 0.8 in the three codon sites respectively, retrieves the frequencies of X0, X1 and X2 observed in the three frames of protein genes and explains these asymmetries. The complex behaviour of these analytical curves is totally unexpected and a priori difficult to imagine. Finally, the evolutionary analytical method developed could be applied to the phylogenetic tree reconstruction and the DNA sequence alignment.

Animals↗

Multiple coding and the evolutionary properties of RNA secondary structure.

This article evaluates evolutionary properties of the transition from RNA primary sequence to RNA secondary structure. It focuses on the restrictions that the conservation of a protein code in an RNA sequence puts on its potential to evolve towards a specific secondary structure. Restricting the mutations to those that do not affect the coding for a protein restricts both the accessibility and the connectivity of the sequence space. The accessibility is restricted because only certain point mutations are allowed. The connectivity is restricted because no insertions and deletions are allowed. Simulating an evolutionary search process for a specific secondary structure shows that (i) the reduction of allowable point mutations allows for adaptation to some large-scale topology, but strongly reduces the possibility of small-scale adaptations, (ii) the abolition of insertions and deletions has very little effect on the results of the search process. During the evolutionary search process for a secondary structure with a specific topology and a high frequency of base-pairing the quasispecies moves into a subspace in which the similarity between secondary structures of neighboring sequences is relatively high. Increased similarity between second structures of neighboring sequences is also found in the Rev responsive element (RRE) in the lentiviruses Caprine arthritis-encephalitis virus and Visna virus. In these viruses a biased nucleotide frequency in the RRE region suggests that selection for the RRE RNA secondary structure affects the amino acid sequence of the env gene. Our results show a variation in the ruggedness of fitness landscapes which are based on a high degree of epistatic interactions. Fitness landscapes play an essential role, not only in biotic evolution, but also in all kinds of optimization processes (Hill Climbing, Simulated Annealing, Genetic Algorithms, etc). Variation in their ruggedness should therefore be taken into account in the analysis of these processes.

Amino Acid Sequence↗

Nucleotide-amino acid interactions and their relation to the genetic code.

The apparent dissociation constants of the complexes of AMP with the methyl esters of amino acids in aqueous solution exhibit good correlations with features of the genetic code and with the frequencies of occurrence of amino acid residues in proteins. Thus it is likely that chemically selective nucleotide-amino acid interactions were involved in the processes of chemical evolution that have led to the emergence of the genetic code. Based on these correlations a storage device for the information regarding nucleotide-amino acid interactions is proposed. It involves processes of simultaneous polymerization to polynucleotides and polypeptides.

Adenosine Monophosphate↗

Divergence of glutamate and glutamine aminoacylation pathways: providing the evolutionary rationale for mischarging.

Aminoacyl-tRNA for protein synthesis is produced through the action of a family of enzymes called aminoacyl-tRNA synthetases. A general rule is that there is one aminoacyl-tRNA synthetase for each of the standard 20 amino acids found in all cells. This is not universal, however, as a majority of prokaryotic organisms and eukaryotic organelles lack the enzyme glutaminyl-tRNA synthetase, which is responsible for forming Gln-tRNAGln in eukaryotes and in Gram-negative eubacteria. Instead, in organisms lacking glutaminyl-tRNA synthetase, Gln-tRNAGln is provided by misacylation of tRNAGln with glutamate by glutamyl-tRNA synthetase, followed by the conversion of tRNA-bound glutamate to glutamine by the enzyme Glu-tRNAGln amidotransferase. The fact that two different pathways exist for charging glutamine tRNA indicates that ancestral prokaryotic and eukaryotic organisms evolved different cellular mechanisms for incorporating glutamine into proteins. Here, we explore the basis for diverging pathways for aminoacylation of glutamine tRNA. We propose that stable retention of glutaminyl-tRNA synthetase in prokaryotic organisms following a horizontal gene transfer event from eukaryotic organisms (Lamour et al. 1994) was dependent on the evolving pool of glutamate and glutamine tRNAs in the organisms that acquired glutaminyl-tRNA synthetase by this mechanism. This model also addresses several unusual aspects of aminoacylation by glutamyl- and glutaminyl-tRNA synthetases that have been observed.

Acylation↗

Searching tRNA sequences for relatedness to aminoacyl-tRNA synthetase families.

tRNA sequences were analyzed for sequence features correlated with known classes of aminoacyl-tRNA synthetase enzymes. The tRNAs were searched for distinguishing nucleotides anywhere in their sequences. The analyses did not find nucleotides predictive of synthetase class membership. We conclude that such nucleotides never existed in tRNA sequences or that they existed and were lost from many of the tRNA sequences during evolution.

Amino Acid Sequence↗

Transfer-RNA, an early gene?

The theory of self-reproductive molecular systems involves the consequence that translation must have started from a selected distribution of RNA molecules, that comprised GC-rich sequences of a length less than 100 nucleotides. This implies a joint function of messenger and adaptor, which both had to be recruited from the same mutant distribution. The reconstruction of tRNA precursors yields such a molecule showing some reverberation of a codon pattern GNC. These findings suggest that tRNA has been the earliest component of the translation machinery.

Base Sequence↗

Simulation of protein evolution by random fixation of allowed codons.

Computer simulation of protein evolution is based on a simple model consisting of random fixation of allowed codons (RFAC). Random replacement of single nucleotides occurs in a DNA sequence. If this results in any of the synonomous codons for allowed amino acids the mutation is fixed, if not, there is no change in the DNA and the cycle is repeated. Multiple fixations at the same nucleotide site, back mutations, degenerate fixations and coincidental identity of amino acids all occur. RFAC simulation begins with a single DNA sequence and follows a phylogeny based on the fossil record. The rate of fixation at the level of DNA is constant. The model upon which RFAC simulation is based is the same as the neutral theory of molecular evolution. The simulation is therefore a test of this theory. The results of simulated and real evolution are compared for fibrinopeptides A in mammals and cytochromes C and hemoglobin alpha and beta chains in vertebrates. In each case the allowed variation at each site has been set equal to that observed, twice that observed and all protein amino acids. Rates of fixation vary from 2.4 X 10(-10) to 10(-8) accepted nucleotide fixations per codon per year. There is some, although never excellent, agreement between real and simulated evolution, the better fits are obtained in the cases of fibrinopeptides A and cytochromes C. The major source of discrepancy between real evolution and simulation is irregularities in the rates of real evolution. RFAC simulation is compared with the random evolutionary hit (REH) model, augmented maximum parsimony and the accepted point mutations (PAM) approach.

Amino Acids↗

The rates of evolution in some ribosomal components.

The rate of nucleotide substitution (k(nuc)) of 5s RNA was estimated to be (1.8 +/- 0.5) x 10(-10) per site per year by comparing the nucleotide sequences of human and Xenopus 5s RNA and using the geological time elapsed since the separation of mammals and amphibians. Similarly, k(nuc) of 5.8s rRNA was calculated to be 0.93 10(-1u) per site per year from the sequences of rat hepatoma cells and Saccbaromyces cerevisiae. For the comparison of these data with the amino acid substitution rate of known proteins, the k(nuc) values of 5s rRNA and 5.8s rRNA were converted to the rate of amino acid substitution (k(aa')). The k(aa') values in pauling units were 0.4 and 2 0.3, respectively. The average k(aa) of ribosomal proteins was also estimated to be 0.2 0.3 pauling from the N-terminal amino acid sequences of seventeen 30s ribosomal proteins of Bacillus stearothermopbilus and Eschericbia coli. Thus, the evolutionary rates of these ribosomal components studied here are similar to each other; they considerably slower than that of the known cellular proteins. Most, if not all, of the replacements in ribosomal proteins occurred between amino acids of a chemically similar nature.

Animals↗