PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

The rate of synonymous substitution in enterobacterial genes is inversely related to codon usage bias.

Genes sequences from Escherichia coli, Salmonella typhimurium, and other members of the Enterobacteriaceae show a negative correlation between the degree of synonymous-codon usage bias and the rate of nucleotide substitution at synonymous sites. In particular, very highly expressed genes have very biased codon usage and accumulate synonymous substitutions very slowly. In contrast, there is little correlation between the degree of codon bias and the rate of protein evolution. It is concluded that both the rate of synonymous substitution and the degree of codon usage bias largely reflect the intensity of selection at the translational level. Because of the high variability among genes in rates of synonymous substitution, separate molecular clocks of synonymous substitution might be required for different genes.

Biological Evolution↗

Accounting for background nucleotide composition when measuring codon usage bias: brilliant idea, difficult in practice.

The effective number of codons used in a gene is a commonly used measure of codon usage. It varies between 20 and 61 (standard genetic code) and indicates to which degree the entire genetic code is used. It is a drawback of this method that it does not take background composition into account. This led Novembre to introduce a variant called Nc' (Novembre JA. 2002. Accounting for background nucleotide composition when measuring codon usage bias. Mol Biol Evol 19:1390-4). In this letter, its properties are under the loupe, with special emphasis on phenomena relating to codon homozygosity. A theoretical misunderstanding regarding this estimator is explained in detail, notably Nc varies between 0 and 61 instead of 20 and 61 (with the standard genetic code). Practical examples from the genome of Pseudomonas aeruginosa are given which demonstrate that the problem is not just theoretical.

Base Composition↗

Codon usage adaptation in the ferredoxin-NADP+ oxidoreductase of Cyanophora paradoxa upon translocation from cyanoplast to nucleus.

Previous investigations of the petH gene of the biflagellated autotrophic protist Cyanophora paradoxa (Cp; Glaucocystophyta), descendant of an original endocyanome (symbiotic consortium of a eukaryote with an endocytobiotic cyanobacterium), established that: (i) the gene coding for a cyanoplast protein (FNR) is located on the nuclear genome; (ii) the sequence of the mature protein shows a high degree of amino-acid conservation to cyanobacterial homologs; (iii) the sequence of the transit peptide of the pre-protein displays poor, if any, homology to counterparts in higher plants. Here, we show that the G+C content and codon usage of this gene are most similar to a genuine nuclear gene. By contrast, the G+C content and codon usage display substantial differences to a collection of 30 cyanoplast encoded genes mainly attributable to alterations in the third codon position. Correspondence analysis on codon preference parameters corroborates the claim of codon usage adaptation of the translocated gene to the nuclear pattern. As a consequence, codon usage distances of genes of Cp encoded either by the nucleus or the cyanoplasts vs. homologous genes of the cyanobacterium, Anabaena, are notably different; this result has important phylogenetic implications.

Adaptation, Physiological↗

Significance of codon usage and irregularities of rare codon distribution in genes for expression of BspLU11III methyltransferases.

Genes of adenine-specific DNA-methyltransferase M.BspLU11IIIa and cytosine-specific DNA-methyltransferase M.BspLU11IIIb of the type IIG BspLU11III restriction-modification system from the thermophilic strain Bacillus sp. LU11 were expressed in E. coli. They contain a large number of codons that are rare in E. coli and are characterized by equal values of codon adaptation index (CAI) and expression level measure (E(g)). Rare codons are either diffused (M.BspLU11IIIa) or located in clusters (M.BspLU11IIIb). The expression level of the cytosine-specific DNA-methyltransferase was increased by a factor of 7.3 and that of adenine-specific DNA only by a factor of 1.25 after introduction of the plasmid pRARE supplying tRNA genes for six rare codons in E. coli. It can be assumed that the plasmid supplying minor tRNAs can strongly increase the expression level of only genes with cluster distribution of rare codons. Using heparin-Sepharose and phosphocellulose chromatography and gel filtration on Sephadex G-75 both DNA-methyltransferases were isolated as electrophoretically homogeneous proteins (according to the results of SDS-PAGE).

Amino Acid Sequence↗

An analysis of codon usage in mammals: selection or mutation bias?

A new statistical test has been developed to detect selection on silent sites. This test compares the codon usage within a gene and thus does not require knowledge of which genes are under the greatest selection, that there exist common trends in codon usage across genes, or that genes have the same mutation pattern. It also controls for mutational biases that might be introduced by the adjacent bases. The test was applied to 62 mammalian sequences, and significant codon usage biases were detected in all three species examined (humans, rats, and mice). However, these biases appear not to be the consequence of selection, but of the first base pair in the codon influencing the mutation pattern at the third position.

Animals↗

The strength of translational selection for codon usage varies in the three replicons of Sinorhizobium meliloti.

The genome of the nitrogen-fixing bacterium Sinorhizobium meliloti is composed of three replicons of 3.65 (chromosome), 1.35 (pSymA) and 1.68 Mb (pSymB), respectively. While the chromosome encodes for most of the housekeeping functions, the three elements may contribute to symbiosis, though pSymA is absolutely necessary for nodulation and nitrogen fixation, since it harbours all the characterized nodulation and symbiotic fixation genes. On the other hand, the majority of the sequences located in this megaplasmid are probably not expressed during the free-living stage of the organism. Since most of the sequences located in pSymA are transcribed only at the stage of bacteroids when most probably the fate of the bacterium is to die, the mutations occurring at this stage will not be fixed in the population. Therefore, if natural selection contributes to the codon usage pattern in this species, its effect will be much weaker for the genes placed in pSymA. A codon usage analysis of the genes comprising the three replicons is consistent with the conclusion that selection for translational speed shapes the codon usage of the two replicons which are important for competitive cell growth while the codon usage of the third replicon reflects primarily the mutational bias.

Base Sequence↗

Codon usage tabulated from international DNA sequence databases: status for the year 2000.

The frequencies of each of the 257 468 complete protein coding sequences (CDSs) have been compiled from the taxonomical divisions of the GenBank DNA sequence database. The sum of the codons used by 8792 organisms has also been calculated. The data files can be obtained from the anonymous ftp sites of DDBJ, Kazusa and EBI. A list of the codon usage of genes and the sum of the codons used by each organism can be obtained through the web site http://www.kazusa.or.jp/codon/. The present study also reports recent developments on the WWW site. The new web interface provides data in the CodonFrequency-compatible format as well as in the traditional table format. The use of the database is facilitated by keyword based search analysis and the availability of codon usage tables for selected genes from each species. These new tools will provide users with the ability to further analyze for variations in codon usage among different genomes.

Codon↗

Optimization of codon usage is required for effective genetic immunization against Art v 1, the major allergen of mugwort pollen.

BACKGROUND: As the major allergen of mugwort pollen, Art v 1 is an important target for specific immunotherapy. However, both recombinant protein as well as a gene vaccine for Art v 1 failed to be immunogenic in mice. In order to improve immunogenicity we focused on genetic immunization because interspecific differences of codon usage have been shown as an obstacle for effective induction of immune responses with gene vaccines encoding infectious pathogens. OBJECTIVE: In order to find out, whether codon usage might also be used to improve genetic immunization with allergen genes, the response against a gene vaccine expressing the wild-type gene of Art v 1 (pCMV-wtArt) was compared with a synthetic codon-optimized vector with human codon usage (pCMV-humArt). METHODS: Balb/c mice were injected intradermally with pCMV-wtArt or pCMV-humArt. In vitro expression levels of both constructs were compared in transfection experiments. Total immunoglobulin G (IgG), IgG1, IgG2a and IgE antibodies were analyzed by enzyme-linked immunosorbent assay and the anaphylactic activity of the sera was determined by allergen-specific degranulation of rat basophil leukemia-2H3 cells. RESULTS: No immune response was detectable with the gene vaccine expressing the wildtype Art v 1, but immunization with pCMV-humArt revealed a strong and allergen-specific induction of antibody responses. The antibodies recognized both the recombinant as well as the purified natural (glycosylated) Art v 1 molecule. The response type was Th1-biased, as indicated by high levels of IgG2a antibodies. Expression analysis with B16 mouse melanoma cells transfected with pCMV-humArt or pCMV-wtArt revealed an impaired expression of the wild-type vector but normal translation after recoding. CONCLUSION: The results demonstrate that optimization of codon usage offers a simple way to improve immunogenicity and therefore should be routinely considered in the development of gene vaccines for the treatment of allergy.

Allergens↗

Codon usage in Chlamydia trachomatis is the result of strand-specific mutational biases and a complex pattern of selective forces.

The patterns of synonymous codon choices of the completely sequenced genome of the bacterium Chlamydia trachomatis were analysed. We found that the most important source of variation among the genes results from whether the sequence is located on the leading or lagging strand of replication, resulting in an over representation of G or C, respectively. This can be explained by different mutational biases associated to the different enzymes that replicate each strand. Next we found that most highly expressed sequences are located on the leading strand of replication. From this result, replicational-transcriptional selection can be invoked. Then, when the genes located on the leading strand are studied separately, the correspondence analysis detects a principal trend which discriminates between lowly and highly expressed sequences, the latter displaying a different codon usage pattern than the former, suggesting selection for translation, which is reinforced by the fact that Ks values between orthologous sequences from C. trachomatis and Chlamydia pneumoniae are much smaller in highly expressed genes. Finally, synonymous codon choices appear to be influenced by the hydropathy of each encoded protein and by the degree of amino acid conservation. Therefore, synonymous codon usage in C.trachomatis seems to be the result of a very complex balance among different factors, which rises the problem of whether the forces driving codon usage patterns among microorganisms are rather more complex than generally accepted.

Amino Acids↗

Conservation of translation initiation sites based on dinucleotide frequency and codon usage in Escherichia coli K-12 (W3110): non-random distribution of A/T-rich sequences immediately upstream of the translation initiation codon.

Dinucleotide frequencies are useful for characterizing consensus elements as a minimum unit of nucleotide sequence because the neighborhood relations of nucleotide sequences are reflected in dinucleotides. Using a consensus score based on dinucleotide frequencies and intra-species codon usage heterogeneity, denoted by the Z1 parameter, we report the relationship between nucleotide conservation at the translation initiation sites of genes in the Escherichia coli K-12 genome (W3110) and codon usage in its downstream genes. Significant positive correlations were obtained in three regions centered at -13, -4, and +7, which correspond to the Shine-Dalgarno element, the A + T element immediately upstream of the translation initiation site, and the downstream box, respectively.

Base Sequence↗

Compositional pressure and translational selection determine codon usage in the extremely GC-poor unicellular eukaryote Entamoeba histolytica.

It is widely accepted that the compositional pressure is the only factor shaping codon usage in unicellular species displaying extremely biased genomic compositions. This seems to be the case in the prokaryotes Mycoplasma capricolum, Rickettsia prowasekii and Borrelia burgdorferi (GC-poor), and in Micrococcus luteus (GC-rich). However, in the GC-poor unicellular eukaryotes Dictyostelium discoideum and Plasmodium falciparum, there is evidence that selection, acting at the level of translation, influences codon choices. This is a twofold intriguing finding, since (1) the genomic GC levels of the above mentioned eukaryotes are lower than the GC% of any studied bacteria, and (2) bacteria usually have larger effective population sizes than eukaryotes, and hence natural selection is expected to overcome more efficiently the randomizing effects of genetic drift among prokaryotes than among eukaryotes. In order to gain a new insight about this problem, we analysed the patterns of codon preferences of the nuclear genes of Entamoeba histolytica, a unicellular eukaryote characterised by an extremely AT-rich genome (GC = 25%). The overall codon usage is strongly biased towards A and T in the third codon positions, and among the presumed highly expressed sequences, there is an increased relative usage of a subset of codons, many of which are C-ending. Since an increase in C in third codon positions is 'against' the compositional bias, we conclude that codon usage in E. histolytica, as happens in D. discoideum and P. falciparum, is the result of an equilibrium between compositional pressure and selection. These findings raise the question of why strongly compositionally biased eukaryotic cells may be more sensitive to the (presumed) slight differences among synonymous codons than compositionally biased bacteria.

Animals↗

Influence of intercodon and base frequencies on codon usage in filarial parasites.

Base frequency, codon usage, and intercodon identity were analyzed in five filarial parasite species representing five Onchocercidae genera. Wucheria bancrofti, Brugia malayi, Onchocerca volvulus, Acanthocheilonema viteae, and Dirofilaria immitis gene sequences were downloaded from NCBI, and analysis was performed using locally designed computer programs and other freely available applications. A clear sequence bias was observed among the nematode species examined. At the nucleotide level, AT basepairs were present in gene sequences at higher frequencies than GC. In addition, codons ending in A or T were used proportionately more than those with G or C in the third-codon position. In addition, the amino acids used most often corresponded to codons ending in AT basepairs. Intercodon base proportion was biased in that A was found most often at N4, second only to T in certain specific cases. Since all of these sequence biases were observed in a relatively consistent fashion among all of the organisms studied, we conclude that sequence bias is a genetic characteristic, which is associated with multiple filarial genera.

Animals↗

Whole genome analysis of non-optimal codon usage in secretory signal sequences of Streptomyces coelicolor.

Non-optimal (rare) codons have been suggested to reduce translation rate and facilitate secretion in Escherichia coli. In this study, the complete genome analysis of non-optimal codon usage in secretory signal sequences and non-secretory sequences of Streptomyces coelicolor was performed. The result showed that there was a higher proportion of non-optimal codons in secretory signal sequences than in non-secretory sequences. The increased tendency was more obvious when tested with the experimental data of secretory proteins from proteomics analysis. Some non-optimal codons for Arg (AGA, CGU and CGA), Ile (AUA) and Lys (AAA) were significantly over presented in the secretary signal sequences. It may reveal that a balanced non-optimal codon usage was necessary for protein secretion and expression in Streptomyces.

Base Sequence↗

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence↗

Intron length and codon usage.

The correlation was shown between the length of introns and the codon usage of the coding sequences of the corresponding genes, which in some cases can be related to the level of gene expression. The link is positive in the unicellular organisms, i.e., genes with the longer introns show the higher bias of codon usage. It is most pronounced in baker's yeast, where it is definitely related to the level of gene expression--genes with the higher level of expression have the longer introns. The correlation is inverted in multicellular organisms as compared to unicellular ones. Some organisms, however, do not show the link. The presence or absence of the link does not seem to be related to the GC percent of the coding sequences.

Animals↗

The positive relationship between codon usage bias and translation initiation AUG context in Saccharomyces cerevisiae.

The relationship between the codon usage bias and the sequence context surrounding the AUG translation initiation codon was examined in 211 Saccharomyces cerevisiae mRNA sequences. The codon usage bias and the number of matches to optimal AUG context, (A/U)A(A/C)AA(A/C)AUGUC(U/C), for translation initiation showed a positive relationship, indicating that these two factors are evolutionally under the similar natural selection constraint at the translation level. A new index (AUGCAI = AUG Context Adaptation Index) for the measure of optimal AUG context was devised, and the importance of each position of AUG context was also examined.

Base Sequence↗

Unexpected correlations between gene expression and codon usage bias from microarray data for the whole Escherichia coli K-12 genome.

Escherichia coli has long been regarded as a model organism in the study of codon usage bias (CUB). However, most studies in this organism regarding this topic have been computational or, when experimental, restricted to small datasets; particularly poor attention has been given to genes with low CUB. In this work, correspondence analysis on codon usage is used to classify E.coli genes into three groups, and the relationship between them and expression levels from microarray experiments is studied. These groups are: group 1, highly biased genes; group 2, moderately biased genes; and group 3, AT-rich genes with low CUB. It is shown that, surprisingly, there is a negative correlation between codon bias and expression levels for group 3 genes, i.e. genes with extremely low codon adaptation index (CAI) values are highly expressed, while group 2 show the lowest average expression levels and group 1 show the usual expected positive correlation between CAI and expression. This trend is maintained over all functional gene groups, seeming to contradict the E.coli-yeast paradigm on CUB. It is argued that these findings are still compatible with the mutation-selection balance hypothesis of codon usage and that E.coli genes form a dynamic system shaped by these factors.

Bias↗

Synonymous codon usage and gene function are strongly related in Oryza sativa.

The relationship between codon usage and gene function was investigated while considering a dataset of 2106 nuclear genes of Oryza sativa. The results of standard chi(2) test and F-statistic showed that for every 59 synonymous codons, a strongly significant association with gene functional categories existed in rice, indicating that codon usage was generally coordinated with gene function whether it was at the level of individual amino acids or at the level of nucleotides. However, it could not be directly said that the use of every codons differed significantly between any two functional categories. Notably, there existed large difference both in selection for biased codons or selection intensity among functional categories. Therefore, we identified at least two classes of genes: one group of genes, mainly belonging to the "METABOLISM" category, was tended to use G- and/or C-ending codons while the other was more biased to choose codons ending with A and/or U. The latter group contained genes of various functions, especially those genes classified into the "Nuclear Structure" category. These observations will be more important for molecular genetic engineering and genome functional annotation.

Chromosome Mapping↗