PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Phylogeny, rates of evolution, and patterns of codon usage among sea urchin retroviral-like elements, with implications for the recognition of horizontal transfer.

Phylogenetic relationships, rates of evolution, and codon usage were investigated in a family of retrotransposons (SURL elements) found in echinoids. The phylogeny of SURL element reverse transcriptase sequences from 10 echinoid species clearly shows the phylogenetic signature of the host taxa as well as paralogous sequences that diverged prior to speciation events. Two subfamilies (1 and 5) of SURL element reverse transcriptase sequences are recognized that diverged prior to the radiation of the Echinometridae. Comparisons of synonymous versus nonsynonymous substitutions indicate that SURL elements have been active in echinoid genomes and have evolved under purifying selection for millions of years. Rates of synonymous substitution for reverse transcriptase are similar to rates of single-copy DNA evolution and to rates of synonymous substitution for the H3 and H4 histone genes, contradicting the assumption that rates of evolution are accelerated in retrotransposons. Finally, codon usage in SURL elements is biased for codons ending in A or U relative to 42 sea urchin nuclear genes. Biased codon usage is sometimes cited as evidence for horizontal transfer, but in the case of SURL elements this bias occurs in spite of a long history of vertical transmission rather than because of horizontal transfer.

Animals↗

Proteome composition and codon usage in spirochaetes: species-specific and DNA strand-specific mutational biases.

The genomes of the spirochaetes Borrelia burgdorferi and Treponema pallidum show strong strand-specific skews in nucleotide composition, with the leading strand in replication being richer in G and T than the lagging strand in both species. This mutation bias results in codon usage and amino acid composition patterns that are significantly different between genes encoded on the two strands, in both species. There are also substantial differences between the species, with T.pallidum having a much higher G+C content than B. burgdorferi. These changes in amino acid and codon compositions represent neutral sequence change that has been caused by strong strand- and species-specific mutation pressures. Genes that have been relocated between the leading and lagging strands since B. burgdorferi and T.pallidum diverged from a common ancestor now show codon and amino acid compositions typical of their current locations. There is no evidence that translational selection operates on codon usage in highly expressed genes in these species, and the primary influence on codon usage is whether a gene is transcribed in the same direction as replication, or opposite to it. The dnaA gene in both species has codon usage patterns distinctive of a lagging strand gene, indicating that the origin of replication lies downstream of this gene, possibly within dnaN. Our findings strongly suggest that gene-finding algorithms that ignore variability within the genome may be flawed.

Amino Acids↗

Codon usage bias and base composition of nuclear genes in Drosophila.

The nuclear genes of Drosophila evolve at various rates. This variation seems to correlate with codon-usage bias. In order to elucidate the determining factors of the various evolutionary rates and codon-usage bias in the Drosophila nuclear genome, we compared patterns of codon-usage bias with base compositions of exons and introns. Our results clearly show the existence of selective constraints at the translational level for synonymous (silent) sites and, on the other hand, the neutrality or near neutrality of long stretches of nucleotide sequence within noncoding regions. These features were found for comparisons among nuclear genes in a particular species (Drosophila melanogaster, Drosophila pseudoobscura and Drosophila virilis) as well as in a particular gene (alcohol dehydrogenase) among different species in the genus Drosophila. The patterns of evolution of synonymous sites in Drosophila are more similar to those in the prokaryotes than they are to those in mammals. If a difference in the level of expression of each gene is a main reason for the difference in the degree of selective constraint, the evolution of synonymous sites of Drosophila genes would be sensitive to the level of expression among genes and would change as the level of expression becomes altered in different species. Our analysis verifies these predictions and also identifies additional selective constraints at the translational level in Drosophila.

Animals↗

Codon usage in Plasmodium vivax nuclear genes.

Codon usage in Plasmodium vivax nuclear genes was analysed and compared with that in Plasmodium falciparum nuclear genes. Preferred codons were determined for P. vivax. Unlike P. falciparum, P. vivax genes are about 15% less A+T rich in the coding regions, with no obvious A+T bias at the third position of the codons. The amino-acid composition of P. vivax gene products is also different from that of P. falciparum. These results provide valuable information to facilitate gene cloning as well as expression and transfection studies for P. vivax.

Amino Acids↗

The rate of synonymous substitution in enterobacterial genes is inversely related to codon usage bias.

Genes sequences from Escherichia coli, Salmonella typhimurium, and other members of the Enterobacteriaceae show a negative correlation between the degree of synonymous-codon usage bias and the rate of nucleotide substitution at synonymous sites. In particular, very highly expressed genes have very biased codon usage and accumulate synonymous substitutions very slowly. In contrast, there is little correlation between the degree of codon bias and the rate of protein evolution. It is concluded that both the rate of synonymous substitution and the degree of codon usage bias largely reflect the intensity of selection at the translational level. Because of the high variability among genes in rates of synonymous substitution, separate molecular clocks of synonymous substitution might be required for different genes.

Biological Evolution↗

Codon usage adaptation in the ferredoxin-NADP+ oxidoreductase of Cyanophora paradoxa upon translocation from cyanoplast to nucleus.

Previous investigations of the petH gene of the biflagellated autotrophic protist Cyanophora paradoxa (Cp; Glaucocystophyta), descendant of an original endocyanome (symbiotic consortium of a eukaryote with an endocytobiotic cyanobacterium), established that: (i) the gene coding for a cyanoplast protein (FNR) is located on the nuclear genome; (ii) the sequence of the mature protein shows a high degree of amino-acid conservation to cyanobacterial homologs; (iii) the sequence of the transit peptide of the pre-protein displays poor, if any, homology to counterparts in higher plants. Here, we show that the G+C content and codon usage of this gene are most similar to a genuine nuclear gene. By contrast, the G+C content and codon usage display substantial differences to a collection of 30 cyanoplast encoded genes mainly attributable to alterations in the third codon position. Correspondence analysis on codon preference parameters corroborates the claim of codon usage adaptation of the translocated gene to the nuclear pattern. As a consequence, codon usage distances of genes of Cp encoded either by the nucleus or the cyanoplasts vs. homologous genes of the cyanobacterium, Anabaena, are notably different; this result has important phylogenetic implications.

Adaptation, Physiological↗

An analysis of codon usage in mammals: selection or mutation bias?

A new statistical test has been developed to detect selection on silent sites. This test compares the codon usage within a gene and thus does not require knowledge of which genes are under the greatest selection, that there exist common trends in codon usage across genes, or that genes have the same mutation pattern. It also controls for mutational biases that might be introduced by the adjacent bases. The test was applied to 62 mammalian sequences, and significant codon usage biases were detected in all three species examined (humans, rats, and mice). However, these biases appear not to be the consequence of selection, but of the first base pair in the codon influencing the mutation pattern at the third position.

Animals↗

Codon usage tabulated from international DNA sequence databases: status for the year 2000.

The frequencies of each of the 257 468 complete protein coding sequences (CDSs) have been compiled from the taxonomical divisions of the GenBank DNA sequence database. The sum of the codons used by 8792 organisms has also been calculated. The data files can be obtained from the anonymous ftp sites of DDBJ, Kazusa and EBI. A list of the codon usage of genes and the sum of the codons used by each organism can be obtained through the web site http://www.kazusa.or.jp/codon/. The present study also reports recent developments on the WWW site. The new web interface provides data in the CodonFrequency-compatible format as well as in the traditional table format. The use of the database is facilitated by keyword based search analysis and the availability of codon usage tables for selected genes from each species. These new tools will provide users with the ability to further analyze for variations in codon usage among different genomes.

Codon↗

Codon usage in Chlamydia trachomatis is the result of strand-specific mutational biases and a complex pattern of selective forces.

The patterns of synonymous codon choices of the completely sequenced genome of the bacterium Chlamydia trachomatis were analysed. We found that the most important source of variation among the genes results from whether the sequence is located on the leading or lagging strand of replication, resulting in an over representation of G or C, respectively. This can be explained by different mutational biases associated to the different enzymes that replicate each strand. Next we found that most highly expressed sequences are located on the leading strand of replication. From this result, replicational-transcriptional selection can be invoked. Then, when the genes located on the leading strand are studied separately, the correspondence analysis detects a principal trend which discriminates between lowly and highly expressed sequences, the latter displaying a different codon usage pattern than the former, suggesting selection for translation, which is reinforced by the fact that Ks values between orthologous sequences from C. trachomatis and Chlamydia pneumoniae are much smaller in highly expressed genes. Finally, synonymous codon choices appear to be influenced by the hydropathy of each encoded protein and by the degree of amino acid conservation. Therefore, synonymous codon usage in C.trachomatis seems to be the result of a very complex balance among different factors, which rises the problem of whether the forces driving codon usage patterns among microorganisms are rather more complex than generally accepted.

Amino Acids↗

Compositional pressure and translational selection determine codon usage in the extremely GC-poor unicellular eukaryote Entamoeba histolytica.

It is widely accepted that the compositional pressure is the only factor shaping codon usage in unicellular species displaying extremely biased genomic compositions. This seems to be the case in the prokaryotes Mycoplasma capricolum, Rickettsia prowasekii and Borrelia burgdorferi (GC-poor), and in Micrococcus luteus (GC-rich). However, in the GC-poor unicellular eukaryotes Dictyostelium discoideum and Plasmodium falciparum, there is evidence that selection, acting at the level of translation, influences codon choices. This is a twofold intriguing finding, since (1) the genomic GC levels of the above mentioned eukaryotes are lower than the GC% of any studied bacteria, and (2) bacteria usually have larger effective population sizes than eukaryotes, and hence natural selection is expected to overcome more efficiently the randomizing effects of genetic drift among prokaryotes than among eukaryotes. In order to gain a new insight about this problem, we analysed the patterns of codon preferences of the nuclear genes of Entamoeba histolytica, a unicellular eukaryote characterised by an extremely AT-rich genome (GC = 25%). The overall codon usage is strongly biased towards A and T in the third codon positions, and among the presumed highly expressed sequences, there is an increased relative usage of a subset of codons, many of which are C-ending. Since an increase in C in third codon positions is 'against' the compositional bias, we conclude that codon usage in E. histolytica, as happens in D. discoideum and P. falciparum, is the result of an equilibrium between compositional pressure and selection. These findings raise the question of why strongly compositionally biased eukaryotic cells may be more sensitive to the (presumed) slight differences among synonymous codons than compositionally biased bacteria.

Animals↗

Influence of intercodon and base frequencies on codon usage in filarial parasites.

Base frequency, codon usage, and intercodon identity were analyzed in five filarial parasite species representing five Onchocercidae genera. Wucheria bancrofti, Brugia malayi, Onchocerca volvulus, Acanthocheilonema viteae, and Dirofilaria immitis gene sequences were downloaded from NCBI, and analysis was performed using locally designed computer programs and other freely available applications. A clear sequence bias was observed among the nematode species examined. At the nucleotide level, AT basepairs were present in gene sequences at higher frequencies than GC. In addition, codons ending in A or T were used proportionately more than those with G or C in the third-codon position. In addition, the amino acids used most often corresponded to codons ending in AT basepairs. Intercodon base proportion was biased in that A was found most often at N4, second only to T in certain specific cases. Since all of these sequence biases were observed in a relatively consistent fashion among all of the organisms studied, we conclude that sequence bias is a genetic characteristic, which is associated with multiple filarial genera.

Animals↗

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence↗

Intron length and codon usage.

The correlation was shown between the length of introns and the codon usage of the coding sequences of the corresponding genes, which in some cases can be related to the level of gene expression. The link is positive in the unicellular organisms, i.e., genes with the longer introns show the higher bias of codon usage. It is most pronounced in baker's yeast, where it is definitely related to the level of gene expression--genes with the higher level of expression have the longer introns. The correlation is inverted in multicellular organisms as compared to unicellular ones. Some organisms, however, do not show the link. The presence or absence of the link does not seem to be related to the GC percent of the coding sequences.

Animals↗

The positive relationship between codon usage bias and translation initiation AUG context in Saccharomyces cerevisiae.

The relationship between the codon usage bias and the sequence context surrounding the AUG translation initiation codon was examined in 211 Saccharomyces cerevisiae mRNA sequences. The codon usage bias and the number of matches to optimal AUG context, (A/U)A(A/C)AA(A/C)AUGUC(U/C), for translation initiation showed a positive relationship, indicating that these two factors are evolutionally under the similar natural selection constraint at the translation level. A new index (AUGCAI = AUG Context Adaptation Index) for the measure of optimal AUG context was devised, and the importance of each position of AUG context was also examined.

Base Sequence↗

Nucleic acid composition, codon usage, and the rate of synonymous substitution in protein-coding genes.

Based on the rates of synonymous substitution in 42 protein-coding gene pairs from rat and human, a correlation is shown to exist between the frequency of the nucleotides in all positions of the codon and the synonymous substitution rate. The correlation coefficients were positive for A and T and negative for C and G. This means that AT-rich genes accumulate more synonymous substitutions than GC-rich genes. Biased patterns of mutation could not account for this phenomenon. Thus, the variation in synonymous substitution rates and the resulting unequal codon usage must be the consequence of selection against A and T in synonymous positions. Most of the variation in rates of synonymous substitution can be explained by the nucleotide composition in synonymous positions. Codon-anticodon interactions, dinucleotide frequencies, and contextual factors influence neither the rates of synonymous substitution nor codon usage. Interestingly, the nucleotide in the second position of codons (always a nonsynonymous position) was found to affect the rate of synonymous substitution. This finding links the rate of nonsynonymous substitution with the synonymous rate. Consequently, highly conservative proteins are expected to be encoded by genes that evolve slowly in terms of synonymous substitutions, and are consequently highly biased in their codon usage.

Animals↗

Translational selection shapes codon usage in the GC-rich genome of Chlamydomonas reinhardtii.

In unicellular species codon usage is determined by mutational biases and natural selection. Among prokaryotes, the influence of these factors is different if the genome is skewed towards AT or GC, since in AT-rich organisms translational selection is absent. On the other hand, in AT-rich unicellular eukaryotes the two factors are present. In order to understand if GC-rich genomes display a similar behavior, the case of Chlamydomonas reinhardtii was studied. Since we found that translational selection strongly influences codon usage in this species, we conclude that there is not a common pattern among unicellular organisms.

AT Rich Sequence↗

Codon usage as a tool to predict the cellular location of eukaryotic ribosomal proteins and aminoacyl-tRNA synthetases.

In spite of many efforts, the prediction of the location of proteins in eukaryotic cells (cytoplasm, mitochondrion or chloroplast) is still far from straightforward. In some cases (e.g. ribosomal proteins and aminoacyl-tRNA synthetases) both the cytoplasmic proteins and their organellar counterparts are encoded by the nuclear genome. A factorial correspondence analysis of the codon usage in yeast and Caenorhabditis elegans shows that the codon usage of those nuclear genes encoding ribosomal proteins or aminoacyl-tRNA synthetases is markedly different, depending on the final location of the proteins (cytoplasmic or mitochondrial). As a consequence, the location of such proteins-whose sequences are now frequently determined by systematic genomic sequencing-can be easily and quickly predicted. A WWW interface has been developed, aimed at providing a user-friendly tool for codon usage pattern analysis. It is available from http://www.genetique.uvsq.fr/afc.html

Amino Acyl-tRNA Synthetases↗

High-level periplasmic expression in Escherichia coli using a eukaryotic signal peptide: importance of codon usage at the 5' end of the coding sequence.

We investigated the ability of signal peptides of eukaryotic origin (human, mouse, and yeast) to efficiently direct model proteins to the Escherichia coli periplasm. These were compared against a well-characterized prokaryotic signal peptide-OmpA. Surprisingly, eukaryotic signal peptides can work very efficiently in E. coli, but require optimization of codon usage by codon-based mutagenesis of the signal peptide coding region. Analysis of the 5' of periplasmic and cytoplasmic E. coli genes shows some codon usage differences.

Amino Acid Sequence↗