PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

The possible role of assignment catalysts in the origin of the genetic code.

A model is presented for the emergence of a primitive genetic code through the selection of a family of proteins capable of executing the code and catalyzing their own formation from polynucleotide templates. These proteins are assignment catalysts capable of modulating the rate of incorporation of different amino acids at the position of different codons. The starting point of the model is a polynucleotide based polypeptide construction process which maintains colinearity between template and product, but may not maintain a coded relationship between amino acids and codons. Among the primitive proteins made are assumed to be assignment catalysts characterized by structural and functional parameters which are used to formulate the production kinetics of these catalysts from available templates. Application of the model to the simple case of two letter codon and amino acid alphabets has been analyzed in detail. As the structural, functional, and kinetic parameters are varied, the dynamics undergoes many bifurcations, allowing an initially ambiguous system of catalysts to evolve to a coded, self-reproductive system. The proposed selective pressure of this evolution is the efficiency of utilization of monomers and energy. The model also simulates the qualitative features of suppression, in which a deleterious mutation is partly corrected by the introduction of translation error.

Amino Acid Sequence↗

Worldwide polymorphism at the MC1R locus and normal pigmentation variation in humans.

While there have been many advances in our understanding of the genetics of pathological skin pigmentation in humans, our knowledge about what determines variation in normal skin color is still incomplete. Variation in one gene, melanocortin 1 receptor (MC1R), has been associated with red hair and fair skin in Europeans. However, this gene might also play an important role in shaping pigmentation of other human populations, where it experiences different selective pressures. Below we review what is currently known about polymorphism and selection at the MC1R coding and promoter regions in human populations, the pattern of MC1R evolution in nonhuman primates, and the interaction of MC1R with other genes.

Africa↗

Molecular determinants and guided evolution of species-specific RNA editing.

Most RNA editing systems are mechanistically diverse, informationally restorative, and scattershot in eukaryotic lineages. In contrast, genetic recoding by adenosine-to-inosine RNA editing seems common in animals; usually, altering highly conserved or invariant coding positions in proteins. Here I report striking variation between species in the recoding of synaptotagmin I (sytI). Fruitflies, mosquitoes and butterflies possess shared and species-specific sytI editing sites, all within a single exon. Honeybees, beetles and roaches do not edit sytI. The editing machinery is usually directed to modify particular adenosines by information stored in intron-mediated RNA structures. Combining comparative genomics of 34 species with mutational analysis reveals that complex, multi-domain, pre-mRNA structures solely determine species-appropriate RNA editing. One of these is a previously unreported long-range pseudoknot. I show that small changes to intronic sequences, far removed from an editing site, can transfer the species specificity of editing between RNA substrates. Taken together, these data support a phylogeny of sytI gene editing spanning more than 250 million years of hexapod evolution. The results also provide models for the genesis of RNA editing sites through the stepwise addition of structural domains, or by short walks through sequence space from ancestral structures.

Adenosine↗

Gene network polymorphism is the raw material of natural selection: the selfish gene network hypothesis.

Population genetics, the mathematical theory of modern evolutionary biology, defines evolution as the alteration of the frequency of distinct gene variants (alleles) differing in fitness over the time. The major problem with this view is that in gene and protein sequences we can find little evidence concerning the molecular basis of phenotypic variance, especially those that would confer adaptive benefit to the bearers. Some novel data, however, suggest that a large amount of genetic variation exists in the regulatory region of genes within populations. In addition, comparison of homologous DNA sequences of various species shows that evolution appears to depend more strongly on gene expression than on the genes themselves. Furthermore, it has been demonstrated in several systems that genes form functional networks, whose products exhibit interrelated expression profiles. Finally, it has been found that regulatory circuits of development behave as evolutionary units. These data demonstrate that our view of evolution calls for a new synthesis. In this article I propose a novel concept, termed the selfish gene network hypothesis, which is based on an overall consideration of the above findings. The major statements of this hypothesis are as follows. (1) Instead of individual genes, gene networks (GNs) are responsible for the determination of traits and behaviors. (2) The primary source of microevolution is the intraspecific polymorphism in GNs and not the allelic variation in either the coding or the regulatory sequences of individual genes. (3) GN polymorphism is generated by the variation in the regulatory regions of the component genes and not by the variance in their coding sequences. (4) Evolution proceeds through continuous restructuring of the composition of GNs rather than fixing of specific alleles or GN variants.

Alleles↗

Selection on rapidly evolving proteins in the Arabidopsis genome.

Genes that have undergone positive or diversifying selection are likely to be associated with adaptive divergence between species. One indicator of adaptive selection at the molecular level is an excess of amino acid replacement fixed differences per replacement site relative to the number of synonymous fixed differences per synonymous site (omega = K(a)/K(s)). We used an evolutionary expressed sequence tag (EST) approach to estimate the distribution of omega among 304 orthologous loci between Arabidopsis thaliana and A. lyrata to identify genes potentially involved in the adaptive divergence between these two Brassicaceae species. We find that 14 of 304 genes (approximately 5%) have an estimated omega > 1 and are candidates for genes with increased selection intensities. Molecular population genetic analyses of 6 of these rapidly evolving protein loci indicate that, despite their high levels of between-species nonsynonymous divergence, these genes do not have elevated levels of intraspecific replacement polymorphisms compared to previously studied genes. A hierarchical Bayesian analysis of protein-coding region evolution within and between species also indicates that the selection intensities of these genes are elevated compared to previously studied A. thaliana nuclear loci.

Arabidopsis↗

Molecular genetics and evolution of stomach and nonstomach lysozymes in the hoatzin.

Multiple genes of the hoatzin encoding stomach lysozyme c and closely related members of this calcium-binding lysozyme c group were cloned from a genomic DNA library and sequenced. There are a minimum of five genes represented among these sequences that encode two distinct groups of protein sequences. One group of three genes corresponds to the stomach lysozyme amino acid sequences, and the remaining genes encode predicted proteins that are more basic in character and share several sequence identities with the pigeon egg-white lysozyme rather than with the hoatzin stomach lysozymes. Despite these structural similarities between some of the hoatzin gene products and the pigeon lysozyme, phylogenetic analyses indicate that all of the hoatzin sequences are closely related to one another. This is borne out by the relatively small genetic distances even in the intronic regions, which are not subject to the selective pressures operating on the coding regions of the stomach lysozymes. These results suggest that multiple gene duplication events have occurred during the evolution of hoatzin lysozymes.

Amino Acid Sequence↗

Pathogenicity mechanisms of prokaryotic cells: an evolutionary view.

The success of pathogenic microbes depends on their ability to colonize host tissues and to counter host defense mechanisms. Microorganisms can produce overwhelming infection because of their relatively short generation times, and because they have evolved powerful mechanisms for generating phenotypic diversity as an efficient strategy for adapting to rapidly responding immune system defenses and the broad range of polymorphisms characteristic of different host tissues. Bacterial evolution may not be a continuous process, but more of a succession of temporally spaced major events. These events cause a non-gradual sequence of adaptations to a given environment. The pathogenicity islands are genetically unstable elements, and many of the genes coding for the adhesins, toxins and other virulence factors are present in pathogenicity islands, which almost certainly had former lives as accessory elements or as parts thereof, or were borne on functional accessory elements. Novel genes are also acquired by transduction (mediated by bacteriophages, plasmids or transposons), by conjugation (DNA transfer between cells) or by transformation (natural DNA uptake). Horizontal gene transfer from other species is a major source of variation and is fundamental to the genetic theory of adaptive evolution in prokaryotes.

Bacteria↗

r8s: inferring absolute rates of molecular evolution and divergence times in the absence of a molecular clock.

SUMMARY: Estimating divergence times and rates of substitution from sequence data is plagued by the problem of rate variation between lineages. R8s version 1.5 is a program which uses parametric, nonparametric and semiparametric methods to relax the assumption of constant rates of evolution to obtain better estimates of rates and times. Unlike most programs for rate inference or phylogenetics, r8s permits users to convert results to absolute rates and ages by constraining one or more node times to be fixed, minimum or maximum ages (using fossil or other evidence). Version 1.5 uses truncated Newton nonlinear optimization code with bound constraints, offering superior performance over previous versions. AVAILABILITY: The linux executable, C source code, sample data sets and user manual are available free at http://ginger.ucdavis.edu/r8s.

Algorithms↗

Drosophila melanogaster acetylcholinesterase gene. Structure, evolution and mutations.

Acetylcholinesterase is a key component of cholinergic neurotransmission. In Drosophila melanogaster, acetylcholinesterase is encoded by the Ace locus. We have determined the complete organization of the locus. The transcription unit is 34 kb (1 kb = 10(3) bases) long and encompasses ten exons. We have mapped the 5' end of the transcript, sequenced all the intron/exon boundaries, as well as the 3' end of the transcript. The deduced mature transcript is 4291 nucleotides long without poly(A). Sequencing of the promoter region reveals a potential TATA box and (GA)n motives. The Drosophila coding sequence is more split than its vertebrate counterparts, but the splicing sites of the two last exons are precisely conserved among Drosophila and vertebrate cholinesterases, and intriguingly also with the bovine thyroglobulin gene. Finally, a number of the mutations isolated in earlier genetic work are precisely placed on our molecular map in introns, exons and promoter regions. Among them, for example, a short deletion known to affect acetylcholinesterase level and tissue distribution removes promoter regions and the first non-coding exon.

Acetylcholinesterase↗

Small fitness effect of mutations in highly conserved non-coding regions.

Comparison of human and mouse genomes has revealed that many non-coding regions have levels of sequence conservation similar to protein-coding genes. These regions have attracted a lot of attention as potentially functional genomic sequences. However, little is known about the effect mutations in these conserved non-coding regions have on fitness and how many of them are present in the human genome as deleterious polymorphisms. To gain insight into the selective constraints imposed on conserved non-coding and protein-coding regions, we compared substitution rates in primate and rodent lineages and analyzed the density and allele frequencies of human polymorphism. Genomic regions conserved between primate and rodent groups show higher relative conservation within rodents than within primates. Thus, our analysis indicates a genome-wide relaxation of selective constraint in the primate lineage, which most likely resulted from a smaller effective population size. We found that this relaxation is much more profound in conserved non-coding regions than in protein-coding regions, and that mutations at a large proportion of sites in conserved non-coding regions are associated with very small fitness effect. Data on human polymorphism are also consistent with very weak selection in conserved non-coding regions. This staggering enrichment in sites at the borderline of neutrality can be explained by assuming an important role for synergistic epistasis in the evolution of non-coding regions. Our results suggest that most individual mutations in conserved non-coding regions are only slightly deleterious but are numerous and may have a significant cumulative impact on fitness.

Algorithms↗

In vivo evidence for non-universal usage of the codon CUG in Candida maltosa.

An alkane-assimilating yeast Candida maltosa had been studied in order to establish systems suitable for biotransformation of hydrophobic compounds. However, functional expression of heterologous genes tested for this purpose had not been successful in several cases. On the other hand, it had been reported that the codon CUG, a universal leucine codon, is read as serine in C. cylindracea. The same altered codon usage had also been suggested by in vitro experiments in some Candida yeasts which are phylogenetically closely related to C. maltosa. In this study we have shown that the failure in functional expression of a heterologous gene is due to the fact that the codon CUG is read as serine in C. maltosa. This conclusion was drawn from the following experimental results: (1) when a cytochrome P450 gene of C. maltosa containing a CTG codon was expressed in C. maltosa, the corresponding amino acid was found to be serine, and not leucine; (2) a tRNA gene with an almost identical structure to that of the tRNASerCAG gene of C. albicans could be isolated from the genome of C. maltosa; (3) the Saccharomyces cerevisiae URA3 gene, which has one CTG codon, could not complement the ura3 mutation of C. maltosa as itself, but when the CTG codon was changed to another leucine codon, CTC, the mutated gene could complement the ura3 mutation. The last result is the first example of succeeding in functional expression of a heterologous gene in Candida species having an altered codon usage by changing the CTG codon in the gene to another codon.

Amino Acid Sequence↗

Structure of rDNA in the mosquito Anopheles gambiae and rDNA sequence variation within and between species of the A. gambiae complex.

The structure of the rDNA repeating unit of Anopheles gambiae (Diptera: Culicidae) was determined by restriction endonuclease mapping and hybridization analyses on four independent clones obtained from a genomic library of a colony (G3) from the Gambia (West Africa). rDNA gene coding sequences are conserved, but much intragenomic and intraspecific (geographic) variation occurs in the intergenic spacer. Hybridization of subclones from spacer and coding sequences to genomic DNA that was isolated from single mosquitoes from laboratory colonies of four other A. gambiae complex species reveals conservation of coding sequences but concerted evolution in the intergenic spacers.

Africa, Western↗

Complete amino acid sequence of the Fc region of a human delta chain.

The complete amino acid sequence of an Fc-like fragment designated Fc delta (t) and obtained by limited proteolysis with trypsin of an intact myeloma IgD protein (NIG-65) has been determined. The fragment contains 226 amino acid residues and has a molecular weight of 32,000 per monomeric unit. It has three glucosamine oligosaccharides at asparagine residues 68, 159, and 210. Of these, glucosamine-159 is characteristic of the delta chain and has no counterpart position in any of the other classes. On the other hand, glucosamine-68 is shared by gamma, mu, and epsilon, and glucosamine-210 is shared by alpha and mu. Although the Fc delta (t) has the common framework structure of immunoglobulins, its sequence has many individual characteristics when its two domains are compared separately with the counterpart domain of other heavy chains. Such comparison has shown that the two Fc domains of the delta chain should be placed in an independent branch in topology; for all the other classes, the Fc domains are paired well with their counterparts. The comparison has also shown that there are three prominent gaps by which each domain can be divided into two homologous halves. For each class of immunoglobulin, a moderate degree of internal homology exists between the first half and the second half of each domain of the Fc, suggesting that the primordial gene may have coded for a unit about the size of a half domain. Based on this observation together with sequence comparisons, a possible genetic mechanism is proposed for the origin and evolution of the genes for immunoglobulin domains.

Amino Acid Sequence↗

Codon-level analysis of histone primary sequence: evidence of a repeat tetrapeptide origin and later inclusion of transcribed sequence.

This work is directed to the question of protein sequence conservation. By reference to the genetic code the aminoacyl sequence of histones H2A, H4, H3, H2B and H1 (fragment) were rewritten as the codon sequences. The N-terminal regions were set aside on the grounds of different composition and sequence. The remainder of the molecule could be referred to simple repeat-tetrapeptide proteins by codon composition (high Gxy, low xGy content) and by sequence. Random segments of three to six residues occur characterized by composition and sequence as originating from the complimentary DNA strand, i.e. as codon "transcript". Ancestral features are probably best seen in H3, point mutations appear to be more extensive in H2B and H1. Segments in reverse order in H2A and in "transcript" in H4 distinguish these two from the other three histones. There is a tenuous possibility the N-terminals also originated as repeat-tetrapeptide now intensively modified. At codon-level the 50S ribosomal protein (L7/L12) of E. coli has features in common with histones (including a palindrome-containing N-terminal). It has the composition and sequence of a well-conserved tetrapeptide-repeat strand (statistical support). If interpretations made here are substantially correct, the 50S r-protein illustrates a significant stage in evolution of histone codon strands.

Amino Acid Sequence↗

The human genes for S-adenosylhomocysteine hydrolase and adenosine deaminase are syntenic on chromosome 20.

Human-Chinese hamster cell hybrids and a monoclonal antibody to human S-adenosylhomocysteine hydrolase were used to identify chromosome 20 as the location of the human gene for this enzyme. The gene for adenosine deaminase had previously been mapped to this chromosome. The activity of S-adenosylhomocysteine hydrolase is dependent in vivo on that of adenosine deaminase, since the substrates for the deaminase, adenosine and deoxyadenosine, respectively, inhibit and inactivate S-adenosylhomocysteine hydrolase in genetic or drug-induced adenosine deaminase deficiency. This functional dependence and the likelihood that S-adenosylhomocysteine hydrolase, a eukaryotic enzyme, arose later than adenosine deaminase, which occurs in prokaryotes as well as eukaryotes, suggest that the occurrence of their genes on the same chromosome may have evolutionary significance. In addition, the unusual capacity of S-adenosylhomocysteine hydrolase to form stable complexes with adenosine and its cofactor, nicotinamide adenine dinucleotide, suggest that evolution of its gene may have involved recombination of a portion of the adenosine deaminase gene with an adenine nucleotide domain-coding sequence of another preexisting gene.

Adenosine Deaminase↗

Continued colonization of the human genome by mitochondrial DNA.

Integration of mitochondrial DNA fragments into nuclear chromosomes (giving rise to nuclear DNA sequences of mitochondrial origin, or NUMTs) is an ongoing process that shapes nuclear genomes. In yeast this process depends on double-strand-break repair. Since NUMTs lack amplification and specific integration mechanisms, they represent the prototype of exogenous insertions in the nucleus. From sequence analysis of the genome of Homo sapiens, followed by sampling humans from different ethnic backgrounds, and chimpanzees, we have identified 27 NUMTs that are specific to humans and must have colonized human chromosomes in the last 4-6 million years. Thus, we measured the fixation rate of NUMTs in the human genome. Six such NUMTs show insertion polymorphism and provide a useful set of DNA markers for human population genetics. We also found that during recent human evolution, Chromosomes 18 and Y have been more susceptible to colonization by NUMTs. Surprisingly, 23 out of 27 human-specific NUMTs are inserted in known or predicted genes, mainly in introns. Some individuals carry a NUMT insertion in a tumor-suppressor gene and in a putative angiogenesis inhibitor. Therefore in humans, but not in yeast, NUMT integrations preferentially target coding or regulatory sequences. This is indeed the case for novel insertions associated with human diseases and those driven by environmental insults. We thus propose a mutagenic phenomenon that may be responsible for a variety of genetic diseases in humans and suggest that genetic or environmental factors that increase the frequency of chromosome breaks provide the impetus for the continued colonization of the human genome by mitochondrial DNA.

Algorithms↗

RNA editing in plant mitochondria.

Comparative sequence analysis of genomic and complementary DNA clones from several mitochondrial genes in the higher plant Oenothera revealed nucleotide sequence divergences between the genomic and the messenger RNA-derived sequences. These sequence alterations could be most easily explained by specific post-transcriptional nucleotide modifications. Most of the nucleotide exchanges in coding regions lead to altered codons in the mRNA that specify amino acids better conserved in evolution than those encoded by the genomic DNA. Several instances show that the genomic arginine codon CGG is edited in the mRNA to the tryptophan codon TGG in amino acid positions that are highly conserved as tryptophan in the homologous proteins of other species. This editing suggests that the standard genetic code is used in plant mitochondria and resolves the frequent coincidence of CGG codons and tryptophan in different plant species. The apparently frequent and non-species-specific equivalency of CGG and TGG codons in particular suggests that RNA editing is a common feature of all higher plant mitochondria.

Amino Acid Sequence↗