Indications of an evolutionary pathway in the amino acid code.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
BACKGROUND: Evolutionary analysis may serve as a useful approach to identify and characterize host defense and viral proteins involved in genetic conflicts. We analyzed patterns of coding sequence evolution of genes with known (TRIM5alpha and APOBEC3G) or suspected (TRIM19/PML) roles in virus restriction, or in viral pathogenesis (PPIA, encoding Cyclophilin A), in the same set of human and non-human primate species. RESULTS AND CONCLUSION: This analysis revealed previously unidentified clusters of positively selected sites in APOBEC3G and TRIM5alpha that may delineate new virus-interaction domains. In contrast, our evolutionary analyses suggest that PPIA is not under diversifying selection in primates, consistent with the interaction of Cyclophilin A being limited to the HIV-1M/SIVcpz lineage. The strong sequence conservation of the TRIM19/PML sequences among primates suggests that this gene does not play a role in antiretroviral defense.
Apolipoprotein(a) is coded by one of the most polymorphic genes known in humans. In white and Asian populations variation in this gene is the major determinant of the plasma concentrations of the atherogenic lipoprotein(a) which varies enormously between individuals and considerably across populations. Recent studies have shown that the genetic architecture of the quantitative Lp(a) trait differs among major human groups. In Africans there is evidence for a transacting factor. Three types of variation have been identified in the apo(a) gene: a size polymorphism in the coding region (K IV type 2 repeats), a pentanucleotide repeat polymorphism in the promoter (5'PNRP) and sequence variation in coding and non-coding regions of the gene including a C/T polymorphism at +93 which creates an additional ATG start codon but also affects transcription. The causal +93 C/T effect is masked by linkage disequilibrium in white populations. Analysis of apo(a) K IV 6-10 exons revealed the existence of population-specific spectra of polymorphism in this domain. However further sequence variation which may provide clues for the understanding of the regulation of apo(a) concentrations still needs to be identified. DNA sequencing and phylogenetic analysis have demonstrated that two types of apo(a) exist, in phylogenetically distant mammalian lineages a K IV derived primate form and a K III-derived hedgehog form which are products of convergent evolution.
It is known that globin genes contain three exons with the middle exon coding for a four-helical supersecondary structure responsible for heme binding. Since this portion of the globin peptide chain can be structurally superimposed onto the cytochrome c and cytochrome b5 chains (Argos and Rossmann 1979), it can be inferred that the cytochrome c gene will contain only one coding sequence while the cytochrome b5 gene will be composed of three exons as found in the globin gene.
The principle of complementary hydropathy predicts that peptides coded for by opposing DNA strands will bind one another because highly hydrophilic amino acids will be complemented by hydrophobic ones and vice versa. This paper provides the chemical plausibility for such interactions. It is suggested that exons coding for interacting peptides were juxtaposed and co-evolved together. Present day genes are no longer thus arranged because of duplications and exon shuffling.
The underlying basis of the genetic code is specific aminoacylation of tRNAs by aminoacyl-tRNA synthetases. Although the code is conserved, bases in tRNA that establish aminoacylation are not necessarily conserved. Even when the bases are conserved, positions of backbone groups that contribute to aminoacylation may vary. We show here that, although the Escherichia coli and human cysteinyl-tRNA synthetases both recognize the same bases (U73 and the GCA anticodon) of tRNA for aminoacylation, they have different emphasis on the tRNA backbone. The E. coli enzyme recognizes two clusters of phosphate groups. One is at A36 in the anticodon and the other is in the core of the tRNA structure and includes phosphate groups at positions 9, 12, 14, and 60. Metal-ion rescue experiments show that those at positions 9, 12, and 60 are involved with binding divalent metal ions that are important for aminoacylation. The E. coli enzyme also recognizes 2'-hydroxyl groups within the same two clusters: at positions 33, 35, and 36 in the anticodon loop, and at positions 49, 55, and 61 in the core. The human enzyme, by contrast, recognizes few phosphate or 2'-hydroxy groups for aminoacylation. The evolution from the backbone-dependent recognition by the E. coli enzyme to the backbone-independent recognition by the human enzyme demonstrates a previously unrecognized shift that nonetheless has preserved the specificity for aminoacylation with cysteine.
Sequence analysis of genes in four species of ciliated protozoa and analysis of tRNAs in Tetrahymena has demonstrated that TAG and TAA encode glutamine or glutamic acid in these organisms and TGA is the only stop codon. Thus, it has generally been assumed that all ciliates use a nonuniversal genetic code in which TGA acts as the sole termination codon. We have sequenced the linear DNA molecules that carry an actin gene and a beta-tubulin gene from the ciliate Euplotes crassus. These genes are shown to use TAA as a termination codon based on homology to known actin and beta-tubulin gene sequences. In addition, we have sequenced a portion of the 3' terminus of the E. crassus H4 histone gene and show that it also uses TAA as a termination codon. These data indicate that the timing of genetic code changes in the ciliates must be reconsidered.
The genetic code that produces human teeth began to develop in primitive vertebrates around 500 million years ago. Some parts of the information appear to have been very stable, particularly the mineralized matrices, while others have evolved. The development of tooth shape and tooth number are very rigidly controlled by genes in each species and are responsive to relatively rapid genetic selection by the environment, as are bone shape and associated soft tissues. Their developmental independence is reflected in the ability of the embryonic tooth bud to develop in vitro. Part of the genetic control of tooth size and shape is correlated with the genetic control of size and shape of the jaws, but the jaws are more responsive to environmental variables modifying individual development than are the teeth. There are recognized mutations producing changes in teeth which act at all levels of control, the development of the embryonic bud, the morphogenesis of the bell stage, the production of enamel and dentin and the formation of the roots. The mechanisms of this genetic control are at the molecular and submolecular levels which are just beginning to be examined. Tooth germs are a good system for study of these processes, and changes in our knowledge will lead to increased understanding of the variation of teeth and its relation to the structure of other tissues both normal and abnormal.
Much animal communication takes place via symbolic codes, where each symbol's meaning is fixed by convention only and not by intrinsic meaning. It is unclear how understanding can arise among individuals utilizing such arbitrary codes, and specifically, whether evolution unaided by individual learning is sufficient to produce such understanding. Using a genetic algorithm implemented on a computer, I demonstrate that a significant though imperfect level of understanding can be achieved by organisms through evolution alone. The population as a whole settles on one particular scheme of coding/decoding information (there are no separate dialects). Several features of such evolving systems are explored and it is shown that the system as a whole is stable against perturbation along many different kinds of ecological parameters.
BACKGROUND: The periodic pattern of DNA in exons is a known phenomenon. It was suggested that one of the initial causes of periodicity could be the universal (RNY)npattern (R = A or G, Y = C or U, N = any base) of ancient RNA. Two major questions were addressed in this paper. Firstly, the cause of DNA periodicity, which was investigated by comparisons between real and simulated coding sequences. Secondly, quantification of DNA periodicity was made using an evolutionary algorithm, which was not previously used for such purposes. RESULTS: We have shown that simulated coding sequences, which were composed using codon usage frequencies only, demonstrate DNA periodicity very similar to the observed in real exons. It was also found that DNA periodicity disappears in the simulated sequences, when the frequencies of codons become equal. Frequencies of the nucleotides (and the dinucleotide AG) at each location along phase 0 exons were calculated for C. elegans, D. melanogaster and H. sapiens. Two models were used to fit these data, with the key objective of describing periodicity. Both of the models showed that the best-fit curves closely matched the actual data points. The first dynamic period determination model consistently generated a value, which was very close to the period equal to 3 nucleotides. The second fixed period model, as expected, kept the period exactly equal to 3 and did not detract from its goodness of fit. CONCLUSIONS: Conclusion can be drawn that DNA periodicity in exons is determined by codon usage frequencies. It is essential to differentiate between DNA periodicity itself, and the length of the period equal to 3. Periodicity itself is a result of certain combinations of codons with different frequencies typical for a species. The length of period equal to 3, instead, is caused by the triplet nature of genetic code. The models and evolutionary algorithm used for characterising DNA periodicity are proven to be an effective tool for describing the periodicity pattern in a species, when a number of exons in the same phase are analysed.
Analysis of the human and mouse genomes identified an abundance of conserved non-genic sequences (CNGs). The significance and evolutionary depth of their conservation remain unanswered. We have quantified levels and patterns of conservation of 191 CNGs of human chromosome 21 in 14 mammalian species. We found that CNGs are significantly more conserved than protein-coding genes and noncoding RNAS (ncRNAs) within the mammalian class from primates to monotremes to marsupials. The pattern of substitutions in CNGs differed from that seen in protein-coding and ncRNA genes and resembled that of protein-binding regions. About 0.3% to 1% of the human genome corresponds to a previously unknown class of extremely constrained CNGs shared among mammals.
Six allotypic specificities of the a series are found on rabbit immunoglobulins: a1, a2 and a3 are found both in domestic and wild rabbits Oryctolagus cuniculus; a100, a101 and a102 seem to be present only in wild rabbits. Each of these specificities is a family of variants always present together in a given serum. These variants can be studied through the cross-reactivities detected between the patterns of the a series. The results of studies of cross-reactivities between a1, a3 and the two specificities a100, and a102 and also the cross-reactivity between a2 and a minor variant of the a1 specificity suggested a hypothetical scheme. This hypothesis attempts to take into account the evolution of the specificities of the a series and their variants. This hypothesis also postulates the existence of a set of closely linked genes which control the synthesis of the variants of a given specificity. One could suppose that primordial allelic genes might have appeared from an ancestor gene. By duplication each allele would have led to the appearance of a set of genes coding for a given specificity. These genes might have evolved through mutations and recombinations. In wild rabbits, the observation of an allotype which seems to result from a recombination between the group of genes coding for the a2 variants and the group of genes coding for a3 variants argues in favors of the genetic recombination mechanism.
The mitochondrial DNA (mtDNA) size of the terrestrial gastropod Albinaria turrita was determined by restriction enzyme mapping and found to be approximately 14.5 kb. Its partial gene content and organization were examined by sequencing three cloned segments representing about one-fourth of the mtDNA molecule. Complete sequences of cytochrome c oxidase subunit II (COII), and ATPase subunit 8 (ATPase8), as well as partial sequences of cytochrome c oxidase subunit I (COI), NADH dehydrogenase subunit 6 (ND6), and the large ribosomal RNA (lrRNA) genes were determined. Nine putative tRNA genes were also identified by their ability to conform to typical mitochondrial tRNA secondary structures. An 82-nt sequence resembles a noncoding region of the bivalve Mytilus edulis, even though it might contain a tenth tRNA gene with an unusual 5-nt overlap with another tRNA gene. The genetic code of Albinaria turrita appears to be the same as that of Drosophila and Mytilus edulis. The structures of COI and COII are conservative, but those of ATPase8 and ND6 are diversified. The sequenced portion of the lrRNA gene (1,079 nt) is characterized by conspicuous deletions in the 5' and 3' ends; this gene represents the smallest coelomate lrRNA gene so far known. Sequence comparisons of the identified genes indicate that there is greater difference between Albinaria and Mytilus than between Albinaria and Drosophila. An evolutionary analysis, based on COII sequences, suggests a possible nonmonophyletic origin of molluskan mtDNA. This is supported also by the absence of the ATPase8 gene in the mtDNA of Mytilus and nematodes, while this gene is present in the mtDNA of Albinaria and Cepaea nemoralis and in all other known coelomate metazoan mtDNAs.
The paper recommends a broadening of Howard Pattee's seminal distinction between a dynamic and a linguistic mode of living systems. It is observed that even the dynamic mode is always a semiotic mode although indexical and analogically coded rather than symbolic and digitally coded. The analogically coded messages corresponds to a kind of tacit knowledge hidden in macromolecular structure and shape (e.g. molecular complementarity) and in organismic architecture and communication, i.e. in the semiotic interactions of the body. It is claimed that the origin of referential processes is tied to the flow of historical singularities. The function of analog and digital codes in evolutionary systems is discussed.
In spite of variations in the sequences of tRNAs, the genetic code (anticodon trinucleotides) is conserved in evolution. However, non-anticodon nucleotides which are species specific are known to prevent a given tRNA from functioning in all organisms. Conversely, species-specific tRNA contact residues in synthetases should also prevent cross-species acylation in a predictable way. To address this question, we investigated the relatively small tyrosine tRNA synthetase where contacts of Escherichia coli tRNA(Tyr) with the alpha2 dimeric protein have been localized by others to four specific sequence clusters on the three-dimensional structure of the Bacillus stearothermophilus enzyme. We used specific functional tests with a previously not-sequenced and not-characterized Mycobacterium tuberculosis enzyme and showed that it demonstrates species-specific aminoacylation in vivo and in vitro. The specificity observed fits exactly with the presence of the clusters characteristic of those established as important for recognition of E. coli tRNA. Conversely, we noted that a recent analysis of the tyrosine enzyme from the eukaryote pathogen Pneumocystis carinii showed just the opposite species specificity of tRNA recognition. According to our alignments, the sequences of the clusters diverge substantially from those seen with the M. tuberculosis, B. stearothermophilus and other enzymes. Thus, the presence or absence of species-specific residues in tRNA synthetases correlates in both directions with cross-species aminoacylation phenotypes, without reference to the associated tRNA sequences. We suggest that this kind of analysis can identify those synthetase-tRNA covariations which are needed to preserve the genetic code. These co-variations might be exploited to develop novel antibiotics against pathogens such as M. tuberculosis and P. carinii.
In aligning homologous protein sequences, it is generally assumed that amino acid substitutions subsequent in time occur independently of amino acid substitutions previous in time, i.e. that patterns of mutation are similar at low and high sequence divergence. This assumption is examined here and shown to be incorrect in an interesting way. Separate mutation matrices were constructed for aligned protein sequence pairs at divergences ranging from 5 to 100 PAM units (point accepted mutations per 100 aligned positions). From these, the corresponding log-odds (Dayhoff) matrices, normalized to 250 PAM units, were constructed. The matrices show that the genetic code influences accepted point mutations strongly at early stages of divergence, while the chemical properties of the side chains dominate at more advanced stages.
This is the first report of a complete mitochondrial genome sequence from a photosynthetic member of the stramenopiles, the chrysophyte alga Chrysodidymus synuroideus. The circular-mapping mitochondrial DNA (mtDNA) of 34 119 bp contains 58 densely packed genes (all without introns) and five unique open reading frames (ORFs). Protein genes code for components of respiratory chain complexes, ATP synthase and the mitoribosome, as well as one product of unknown function, encoded in many other protist mtDNAs (YMF16). In addition to small and large subunit ribosomal RNAs, 23 tRNAs are mtDNA-encoded, permitting translation of all codons present in protein-coding genes except ACN (Thr) and CGN (Arg). The missing tRNAs are assumed to be imported from the cytosol. Comparison of the C.SYNUROIDEUS: mtDNA with that of other stramenopiles allowed us to draw conclusions about mitochondrial genome organization, expression and evolution. First, we provide evidence that mitochondrial ORFs code for highly derived, unrecognizable versions of ribosomal or respiratory genes otherwise 'missing' in a particular mtDNA. Secondly, the observed constraints in mitochondrial genome rearrangements suggest operon-based, co-ordinated expression of genes functioning in common biological processes. Finally, stramenopile mtDNAs reveal an unexpectedly low variability in genome size and gene complement, testifying to substantial differences in the tempo of mtDNA evolution between major eukaryotic lineages.
In an attempt to explain the uniqueness of the coding mechanism of living cells as contrasted with the multispecies structure of ecosystems we examine two models of individuals with some replicative properties. In the first model the system generically remains in a multispecies state. Even though for some of these species the replicative probability is very high, they are unable to invade the system. In the second model, in which the death rate depends on the type of the species, the system relatively quickly reaches a single-species state and fluctuations might at most bring it to yet another single-species state.