PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

From amino acid landscape to protein landscape: analysis of genetic codes in terms of fitness landscape.

Assigning the values of a certain physicochemical property for individual amino acids to the corresponding codons, we can make an amino acid property "landscape" on a four valued three dimensional sequence space from a genetic code table. Eleven property landscapes made from the standard genetic code (SGC) were analyzed. The evaluation of correlation for each landscape is done by theta value, which represents the ratio of the mean slope (as an additive term) to the degree of roughness (as a nonadditive term). The theta-values for hydropathy indices, polarity, specific heat, and beta-sheet propensity were considerably large with respect to SGC. This implies that the additivity of the contribution from each letter holds for these properties. To clarify the meaning of the so-called mutational robustness of SGC, we next examined correlations between the amino acid property and the actual "site fitnesses" of a protein. The site fitnesses were derived from a set of binding preference scores of amino acid residues at every site in MHC class I molecule binding peptides (Udaka et al. in press). We found that the SGC's theta value for an amino acid property is correlated with the significance of the property in the protein function. Adaptive walk simulation on fitness (= affinity) landscapes in a base sequence space for these model peptides confirmed better evolvability due to the introduction of SGC.

Amino Acids↗

An extended RNA code and its relationship to the standard genetic code: an algebraic and geometrical approach.

An algebraic and geometrical approach is used to describe the primaeval RNA code and a proposed Extended RNA code. The former consists of all codons of the type RNY, where R means purines, Y pyrimidines, and N any of them. The latter comprises the 16 codons of the type RNY plus codons obtained by considering the RNA code but in the second (NYR type), and the third, (YRN type) reading frames. In each of these reading frames, there are 16 triplets that altogether complete a set of 48 triplets, which specify 17 out of the 20 amino acids, including AUG, the start codon, and the three known stop codons. The other 16 codons, do not pertain to the Extended RNA code and, constitute the union of the triplets YYY and RRR that we define as the RNA-less code. The codons in each of the three subsets of the Extended RNA code are represented by a four-dimensional hypercube and the set of codons of the RNA-less code is portrayed as a four-dimensional hyperprism. Remarkably, the union of these four symmetrical pairwise disjoint sets comprises precisely the already known six-dimensional hypercube of the Standard Genetic Code (SGC) of 64 triplets. These results suggest a plausible evolutionary path from which the primaeval RNA code could have originated the SGC, via the Extended RNA code plus the RNA-less code. We argue that the life forms that probably obeyed the Extended RNA code were intermediate between the ribo-organisms of the RNA World and the last common ancestor (LCA) of the Prokaryotes, Archaea, and Eucarya, that is, the cenancestor. A general encoding function, E, which maps each codon to its corresponding amino acid or the stop signal is also derived. In 45 out of the 64 cases, this function takes the form of a linear transformation F, which projects the whole six-dimensional hypercube onto a four-dimensional hyperface conformed by all triplets that end in cytosine. In the remaining 19 cases the function E adopts the form of an affine transformation, i.e., the composition of F with a particular translation. Graphical representations of the four local encoding functions and E, are illustrated and discussed. For every amino acid and for the stop signal, a single triplet, among those that specify it, is selected as a canonical representative. From this mapping a graphical representation of the 20 amino acids and the stop signal is also derived. We conclude that the general encoding function E represents the SGC itself.

Amino Acids↗

Complete DNA sequence of the mitochondrial genome of Cepaea nemoralis (Gastropoda: Pulmonata).

The nucleotide sequence of a mitochondrial genome of the pulmonate gastropod mollusc Cepaea nemoralis has been determined. Contained within the 14,100 basepairs (bp) are the two ribosomal RNA genes and 13 protein coding genes typical of metazoan mitochondrial genomes. The Cepaea mtDNA does contain a gene for ATPase subunit 8, like the clausiliid pulmonate, Albinaria, and the chiton, Katharina, but unlike the bivalve mollusc, Mytilus. The mitochondrial genetic code of Cepaea is proposed to be the same as that of Mytilus, Katharina, and Drosophila. Only 14 putative tRNA genes are presented, although there is sufficient unassigned sequence to encode the remainder of the expected total of 22 tRNA genes. These 14 tRNA genes are a mixture of standard cloverleaf structures and nonstandard structures containing TV replacement loops as seen in nematode and mosquito mitochondrial genomes. If the eight unidentified tRNA genes are indeed present, very little unassigned sequence would remain to serve as a control region. Genes are transcribed from both strands of the molecule. Base composition is the least biased for any reported animal mitochondrial genome and is also very little skewed between strands using measures independent of base composition. The Cepaea mitochondrial gene order is quite unlike that of any other reported metazoan mtDNA, with the exception of the recently reported partial sequences of Albinaria. No gene boundaries are shared among all the reported molluscan taxa, demonstrating a complete lack of conservation of mitochondrial gene order across the phylum Mollusca.

Animals↗

Designing better phages.

We propose a method to engineer the genome of bacteriophages to increase their effectiveness as antibacterial agents. Specifically, we exploit the redundancy of the triplet code to design genomes that avoid restriction sites while producing the same proteins as wild-type phages. We give an efficient algorithm to minimize the number of restriction sites against sets of cutter sequences, and demonstrate that that phage genomes can be significantly protected against surprisingly large sets of enzymes with no loss of function. Finally, we develop a model to explain why evolution has failed to eliminate many possible restriction sites despite selective pressure, thus motivating the need for genome-level sequence engineering.

Algorithms↗

The natural selection of the chemical elements.

Evolution is treated here in a novel way. DNA or any other code is considered to be conservative and therefore, once life began, it would prevent change. Change was imposed upon the DNA code as a stress resulting in vulnerability to "advantageous" DNA damage and mutation. In this respect it is the stress, the changing environment, that opened up a possibility of evolution once an early life form had optimised itself in primitive circumstances. Here I examine the initial slow-coming-to-terms with the environment of primitive life, and then its evolution as the environment forced the DNA into novel development by introducing chemical elements in new forms. The situation today is no different. Environmental change is hostile to present day life and will lead to further evolution.

Biological Evolution↗

Partial sequence of the mitochondrial genome of Littorina saxatilis: relevance to gastropod phylogenetics.

A 8022 base pair fragment from the mitochondrial DNA of the prosobranch gastropod Littorina saxatilis has been sequenced and shown to contain the complete genes for 12 transfer RNAs and five protein genes (CoII, ATPase 6, ATPase 8, ND1, ND6), two partial protein genes (CoI and cyt b), and two ribosomal RNAs (small and large subunits). The order of these constituent genes differs from those of other molluscan mitochondrial gene arrangements. Only a single rearrangement involving a block of protein coding genes and three tRNA translocations are necessary to produce identical gene orders between L. saxatilis and K. tunicata. However, only one gene boundary is shared between the L. saxatilis gene order and that of the pulmonate gastropod Cepaea nemoralis. This extends the observation that there is little conservation of mitochrondrial gene order amongst the Mollusca and suggests that radical mitochondrial DNA gene rearrangement has occurred on the branch leading to the pulmonates.

Amino Acid Sequence↗

Elucidating sequence codes: three codes for evolution.

The sequences are related to evolution in several ways. First, they carry traces of a distant past. Two sequence features point to the earliest sequence organization. The universal hidden GCU-periodical pattern in mRNA suggests the earliest codons: GCU and its nine-point-change derivatives. They code for seven amino acids that by several criteria are also the oldest. Together it makes the earliest form of the triplet code, still recognizable in the extant sequences. Another feature present in the sequences, apparently, since separation of prokaryotes and eukaryotes, is hidden genome segmentation. Both protein-coding and noncoding sequences appear to have been formed by fusion of standard size units, about 360 bp (120 aa) in eukaryotes and 450 bp (150 aa) in prokaryotes. Presumably, the units have been functioning at some stage of evolution as autonomous single-gene size elements. There are sequence designs that promote evolution. One such design suitable for fast adaptation is the tandem repetition of identical sequences, so that their copy numbers in the repeat arrays would modulate (tune) the expression of nearby genes. The tandem repeat expansion diseases illustrate this mechanism in a dramatic way: overtuning of the respective gene expression leads to the disease.

Adaptation, Biological↗

The biological origin of antibody diversity.

Antibody diversity has a compelling fascination for many scientists and over the years speculations have sometimes seemed more numerous than facts. Now the structural basis of antibody specificity is well defined. Amino acid sequences and recently three-dimensional structures of various immunoglobulins provide the most solid basis for discussing the origin of diversity. The novel pattern of variable (V) and Constant (C) regions of amino acid sequence has been resolved further to show the functional pattern of variability. Inheritance of separate V and C genes is accepted, but attempts to define more than one gene coding for each V region are considered here to be unnecessary. The pattern of variability is still best understood in terms of mutation and the presence or absence of various selective pressures. The major area of debate still hinges around the extent to which mutation and selection operate during evolution or somatically. Sequence data have now been generally interpreted to require multiple V genes carried in the germ line. A few individual VH genes have been mapped in close linkage to CH genes in the mouse. The apparent existence of three VH alleles in rabbits was a strong argument against multiple V genes. Now the three phenotypes have been shown to be due to alleles controlling the expression of three sets of VH genes all present on the same chromosome. That V-gene expression requires rejoining of V and C genes at the DNA level is now almost certain. Models for the joining process can draw on the precedents of transposable genetic elements, which are widespread in Nature. The total extent of antibody diversity remains a philosophical point. Estimates of the number of antibody molecules required for observed diversity are reduced by two recently documented proposals. Each antibody combining site apparently has many (estimated at 100) different specificities and most combinations of VH and VL regions probably form a viable site. A given combining site can be defined by its pattern of shared specificities. Several specific antibody repertoires have been measured and the size in each case is consistent with the stringency with which the specificity is selected. Repertoire size appears to be under genetic control, but there are problems in viewing the genotype through the veil of clonal selection. Molecular hybridization has been used recently in an attempt to count V and C genes directly. C genes are seen in DNA having nonreiterated sequences, as formal genetics predicts. Each V-region probe hybridizes at a similar rate to C-region probes. Interpretation of this result depends on the extent to which one V-region probe will reveal nonhomologous V genes. Previous estimates that many cross-hybridizing genes should have been seen if present are possibly exaggerated. It is argued here that the data are compatible with a germ-line gene for each probe studied. Maximum estimates for the number of germ-line genes are sufficient to account for antibody diversity...

Amino Acid Sequence↗

An evolutionary model of a complementary circular code.

The subset X0 = [sequence: see text] of 20 trinucleotides has a preferential occurrence in frame 0 (a reading frame established by the ATG start trinucleotide) of protein (coding) genes of both prokaryotes and eukaryotes. This subset X0++ has the rarity property (6 x 10(-8)) to be a complementary maximal circular code with two permutated maximal circular codes X1 and X2 in frames 1 and 2 respectively (frame 0 shifted by one and two nucleotides respectively in the 5'-3' direction). X0 is called a C3 code. A quantitative study of these three subsets X0, X1 and X2 in the three frames 0, 1 and 2 of eukaryotic protein genes shows that their occurrence frequencies are constant functions of the trinucleotide positions in the sequences. The frequencies of X0, X1 and X2 in frame 0 of the eukaryotic protein genes are 48.5%, 29% and 22.5% respectively. These properties are not observed in the 5' and 3' regions of eukaryotes where X0, X1 and X2 occur with variable frequencies around the random value (1/3). Several frequency asymmetries unexpectedly observed, e.g. the frequency difference between X1 and X2 in the frame 0, are related to a new property of the C3 code X0 involving substitutions. An evolutionary model at three parameters (p, q, k) based on an independent mixing of the 20 codons (trinucleotides in frame 0) of X0 with equiprobability (1/20) followed by k approximately 5 substitutions per codon in the three codon sites in proportions p approximately 0.1, q approximately 0.1 and r = 1-p-q approximately 0.8 respectively, retrieves the frequencies of X0, X1 and X2 observed in the three frames of protein genes and explains these asymmetries.

Animals↗

Sense in antisense?

A correspondence between open reading frames in sense and antisense strands is expected from the hypothesis that the prototypic triplet code was of general form RNY, where R is a purine base, N is any base, and Y is a pyrimidine. A deficit of stop codons in the antisense strand (and thus long open reading frames) is predicted for organisms with high G + C percentages; however, two bacteria (Azotobacter vinelandii, Rhodobacter capsulatum) have larger average antisense strand open reading frames than predicted from (G + C)%. The similar codon frequencies found in sense and antisense strands can be attributed to the wide distribution of inverted repeats (stem-loop potential) in natural DNA sequences.

Anticodon↗

Forbidden synonymous substitutions in coding regions.

In the evolution of highly conserved genes, a few "synonymous" substitutions at third bases that would not alter the protein sequence are forbidden or very rare, presumably as a result of functional requirements of the gene or the messenger RNA. Another 10% or 20% of codons are significantly less variable by synonymous substitution than are the majority of codons. The changes that occur at the majority of third bases are subject to codon usage restrictions. These usage restrictions control sequence similarities between very distant genes. For example, 70% of third bases are identical in calmodulin genes of man and trypanosome. Third-base similarities of distant genes for conserved proteins are mathematically predicted, on the basis of the G+C composition of third bases. These observations indicate the need for reexamination of methods used to calculate synonymous substitutions.

Actins↗

A theory for the origin of a self-replicating chemical system. I: Natural selection of the autogen from short, random oligomers.

A theory is described for the origin of a simple chemical system named an autogen, consisting of two short oligonucleotide sequences coding for two simple catalytic peptides. If the theory is valid, under appropriate conditions the autogen would be capable of self-reproduction in a truly genetic process involving both replication and translation. Limited catalytic ability, short oligomer sequences, and low selectivities leading to sloppy information transfer processes are shown to be adequate for the origin of the autogen from random background oligomers. A series of discrete steps, each highly probable if certain minimum requirements and boundary conditions are satisfied, lead to exponential increase in population of all components in the system due to autocatalysis and hypercyclic organization. Nucleation of the components and exponential increase to macroscopic amounts could occur in times on the order of weeks. The feasibility of the theory depends on a number of factors, including the capability of simple protoenzymes to provide moderate enhancements of the accuracies of replication and translation and the likelihood of finding an environment where all of the required processes can occur simultaneously. Regardless of whether or not the specific form proposed for the autogen proves to be feasible, the theory suggests that the first self-replicating chemical systems may have been extremely simple, and that the period of time required for chemical evolution prior to Darwinian natural selection may have been far shorter than generally assumed. Due to the short time required, this theory, unlike others on the origin of genetic processes, is potentially capable of direct experimental verification. A number of prerequisites leading up to such an experiment are suggested, and some have been fulfilled. If successful, such an experiment would be the first laboratory demonstration of the spontaneous emergence by natural selection of a genetic, self-replicating, and evolving molecular system, and might represent the first step in the prebiotic environment of true Darwinian evolution toward a living cell.

Base Sequence↗

Complete sequences of the highly rearranged molluscan mitochondrial genomes of the Scaphopod Graptacme eborea and the bivalve Mytilus edulis.

We have determined the complete sequence of the mitochondrial genome of the scaphopod mollusk Graptacme eborea (14,492 nts) and completed the sequence of the mitochondrial genome of the bivalve mollusk Mytilus edulis (16,740 nts). (The name Graptacme eborea is a revision of the species formerly known as Dentalium eboreum.) G. eborea mtDNA contains the 37 genes that are typically found and has the genes divided about evenly between the two strands, but M. edulis contains an extra trnM and is missing atp8, and it has all genes on the same strand. Each has a highly rearranged gene order relative to each other and to all other studied mtDNAs. G. eborea mtDNA has almost no strand skew, but the coding strand of M. edulis mtDNA is very rich in G and T. This is reflected in differential codon usage patterns and even in amino acid compositions. G. eborea mtDNA has fewer noncoding nucleotides than any other mtDNA studied to date, with the largest noncoding region only 24 nt long. Phylogenetic analysis using 2,420 aligned amino acid positions of concatenated proteins weakly supports an association of the scaphopod with gastropods to the exclusion of Bivalvia, Cephalopoda, and Polyplacophora, but it is generally unable to convincingly resolve the relationships among major groups of the Lophotrochozoa, in contrast to the good resolution seen for several other major metazoan groups.

Animals↗

"Word" preference in the genomic text and genome evolution: different modes of n-tuplet usage in coding and noncoding sequences.

Extensive work on n-tuplet occurrence in genomic sequences has revealed the correlation of their usage with sequence origin. Parallel to that, there exist different restrictions in the nucleotide composition of coding and noncoding sequences that may result in distinct modes of usage of n-tuplets. The relatively simple approaches described herein focus on such differences. They are based on simple summation measures of n-tuplet frequencies, computed after filtering the background nucleotide composition. Among the main targets of this work is to draw some conclusions on the qualitative differences in the composition of genomic sequences depending on their functionality. Moreover, an evolutionary model is formulated, including simple forms of ubiquitous events of genome dynamics: genomic fusions, genome shuffling due to transpositions, replication slippage, and point mutations. This model is shown to be able to reproduce all the statistical features of genomic sequences discussed herein.

Base Sequence↗

Avian IFN-gamma genes: sequence analysis suggests probable cross-species reactivity among galliforms.

Little is known about the evolution of cytokines in non-mammalian systems. To address this problem, we attempted to clone the gene for interferon-gamma (IFN-gamma) from a variety of avian species using oligonucleotide primers based on the sequence of the chicken IFN-gamma gene. The coding sequence and partial intron sequences were determined for four species, namely guinea fowl, ring-necked pheasant, Japanese quail, and turkey. To obtain sequence information on the gene extremities, a modified 5' and 3' RACE protocol was used. The sequence information showed that the coding regions of the IFN-gamma gene are highly conserved among the species studied (93.5%-96.7% and 87.8%-97.6% at the nucleotide and peptide levels, respectively) and are more conserved at the amino-terminal region (exons 1 and 2) than the carboxyl-terminal (exons 3 and 4). This high degree of overall identity at the predicted primary amino acid sequence level of the protein, including the deduced IFN-gamma receptor binding motifs, suggests that IFN-gamma may be cross-reactive among these species. Phylogenetic analysis shows that the similarity of the avian IFN-gamma sequences parallels the presumed evolutionary relationships between the species.

Amino Acid Sequence↗

Chance and necessity do not explain the origin of life.

Where and how did the complex genetic instruction set programmed into DNA come into existence? The genetic set may have arisen elsewhere and was transported to the Earth. If not, it arose on the Earth, and became the genetic code in a previous lifeless, physical-chemical world. Even if RNA or DNA were inserted into a lifeless world, they would not contain any genetic instructions unless each nucleotide selection in the sequence was programmed for function. Even then, a predetermined communication system would have had to be in place for any message to be understood at the destination. Transcription and translation would not necessarily have been needed in an RNA world. Ribozymes could have accomplished some of the simpler functions of current protein enzymes. Templating of single RNA strands followed by retemplating back to a sense strand could have occurred. But this process does not explain the derivation of "sense" in any strand. "Sense" means algorithmic function achieved through sequences of certain decision-node switch-settings. These particular primary structures determine secondary and tertiary structures. Each sequence determines minimum-free-energy folding propensities, binding site specificity, and function. Minimal metabolism would be needed for cells to be capable of growth and division. All known metabolism is cybernetic--that is, it is programmatically and algorithmically organized and controlled.

Animals↗

Isochores and the evolutionary genomics of vertebrates.

The nuclear genomes of vertebrates are mosaics of isochores, very long stretches (>>300kb) of DNA that are homogeneous in base composition and are compositionally correlated with the coding sequences that they embed. Isochores can be partitioned in a small number of families that cover a range of GC levels (GC is the molar ratio of guanine+cytosine in DNA), which is narrow in cold-blooded vertebrates, but broad in warm-blooded vertebrates. This difference is essentially due to the fact that the GC-richest 10-15% of the genomes of the ancestors of mammals and birds underwent two independent compositional transitions characterized by strong increases in GC levels. The similarity of isochore patterns across mammalian orders, on the one hand, and across avian orders, on the other, indicates that these higher GC levels were then maintained, at least since the appearance of ancestors of warm-blooded vertebrates. After a brief review of our current knowledge on the organization of the vertebrate genome, evidence will be presented here in favor of the idea that the generation and maintenance of the GC-richest isochores in the genomes of warm-blooded vertebrates were due to natural selection.

Animals↗

The complete genome sequence of the murine respiratory pathogen Mycoplasma pulmonis.

Mycoplasma pulmonis is a wall-less eubacterium belonging to the Mollicutes (trivial name, mycoplasmas) and responsible for murine respiratory diseases. The genome of strain UAB CTIP is composed of a single circular 963 879 bp chromosome with a G + C content of 26.6 mol%, i.e. the lowest reported among bacteria, Ureaplasma urealyticum apart. This genome contains 782 putative coding sequences (CDSs) covering 91.4% of its length and a function could be assigned to 486 CDSs whilst 92 matched the gene sequences of hypothetical proteins, leaving 204 CDSs without significant database match. The genome contains a single set of rRNA genes and only 29 tRNAs genes. The replication origin oriC was localized by sequence analysis and by using the G + C skew method. Sequence polymorphisms within stretches of repeated nucleotides generate phase-variable protein antigens whilst a recombinase gene is likely to catalyse the site-specific DNA inversions in major M.pulmonis surface antigens. Furthermore, a hemolysin, secreted nucleases and a glyco-protease are predicted virulence factors. Surprisingly, several of the genes previously reported to be essential for a self-replicating minimal cell are missing in the M.pulmonis genome although this one is larger than the other mycoplasma genomes fully sequenced until now.

Animals↗