PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Translational selection shapes codon usage in the GC-rich genome of Chlamydomonas reinhardtii.

In unicellular species codon usage is determined by mutational biases and natural selection. Among prokaryotes, the influence of these factors is different if the genome is skewed towards AT or GC, since in AT-rich organisms translational selection is absent. On the other hand, in AT-rich unicellular eukaryotes the two factors are present. In order to understand if GC-rich genomes display a similar behavior, the case of Chlamydomonas reinhardtii was studied. Since we found that translational selection strongly influences codon usage in this species, we conclude that there is not a common pattern among unicellular organisms.

AT Rich Sequence↗

The role of context-dependent mutations in generating compositional and codon usage bias in grass chloroplast DNA.

The influence of local base composition on mutations in chloroplast DNA (cpDNA) is studied in detail and the resulting, empirically derived, mutation dynamics are used to analyze both base composition and codon usage bias. A 4 x 4 substitution matrix is generated for each of the 16 possible flanking base combinations (contexts) using 17,253 noncoding sites, 1309 of which are variable, from an alignment of three complete grass chloroplast genome sequences. It is shown that substitution bias at these sites is correlated with flanking base composition and that the A+T content of these flanking sites as well as the number of flanking pyrimidines on the same strand appears to have general influences on substitution properties. The context-dependent equilibrium base frequencies predicted from these matrices are then applied to two analyses. The first examines whether or not context dependency of mutations is sufficient to generate average compositional differences between noncoding cpDNA and silent sites of coding sequences. It is found that these two classes of sites exist, on average, in very different contexts and that the observed mutation dynamics are expected to generate significant differences in overall composition bias that are similar to the differences observed in cpDNA. Context dependency, however, cannot account for all of the observed differences: although silent sites in coding regions appear to be at the equilibrium predicted, noncoding cpDNA has a significantly lower A+T content than expected from its own substitution dynamics, possibly due to the influence of indels. The second study examines the codon usage of low-expression chloroplast genes. When context is accounted for, codon usage is very similar to what is predicted by the substitution dynamics of noncoding cpDNA. However, certain codon groups show significant deviation when followed by a purine in a manner suggesting some form of weak selection other than translation efficiency. Overall, the findings indicate that a full understanding of mutational dynamics is critical to understanding the role selection plays in generating composition bias and sequence structure.

Base Composition↗

Codon usage as a tool to predict the cellular location of eukaryotic ribosomal proteins and aminoacyl-tRNA synthetases.

In spite of many efforts, the prediction of the location of proteins in eukaryotic cells (cytoplasm, mitochondrion or chloroplast) is still far from straightforward. In some cases (e.g. ribosomal proteins and aminoacyl-tRNA synthetases) both the cytoplasmic proteins and their organellar counterparts are encoded by the nuclear genome. A factorial correspondence analysis of the codon usage in yeast and Caenorhabditis elegans shows that the codon usage of those nuclear genes encoding ribosomal proteins or aminoacyl-tRNA synthetases is markedly different, depending on the final location of the proteins (cytoplasmic or mitochondrial). As a consequence, the location of such proteins-whose sequences are now frequently determined by systematic genomic sequencing-can be easily and quickly predicted. A WWW interface has been developed, aimed at providing a user-friendly tool for codon usage pattern analysis. It is available from http://www.genetique.uvsq.fr/afc.html

Amino Acyl-tRNA Synthetases↗

[Synonymous codon usage bias in the rice cultivar 93-11 (Oryza sativa L. ssp. indica)].

By using the whole genome sequences and EST data from the indica rice cultivar 93-11, a detailed relative analysis is made of the effect of some impact factors on synonymous codon usage. The results showed that the gene expression level assessed by mRNA abundance is positive relative to the "codon adaptation index" (CAI, 0.227**), and "codon preference parameter" (CPP, 0.145**), but negative relative to "effective number of codons" (ENC, -0.147**), indicating that genes with higher expression showed more significant variation in codon usage. There are significant negative correlations between gene length and CAI, CPP (r = -0.413** and -0.480** respectively), but a positive correlation between gene length and ENC(r = 0.210**), which suggested a tendency of shorter genes to higher expression of the transcriptional activity in 93-11. From the results that a higher negative correlation between GC content and ENC(r = -0.740**), but higher positive correlations between GC content and CAI, CPP (r = 0.877** and 0.832**, respectively), we can concluded that the GC content in coding region gave far more contribution to codon usage bias than that mRNA abundance and gene length. Four kinds of bases showed a three-period distribution in the translation initiation region, the bias at the first codon sites, which located +4, and +6, in the downstream of ATG being the largest. That suggested that there was a strong action of natural selection on these specific positions in the 93-11 genome. In this paper twenty-five codons defined firstly as "optimal codons" in 93-11 may provide some more useful information for rice gene-transformation.

Base Composition↗

Codon usage and selection on proteins.

Selection pressures on proteins are usually measured by comparing homologous nucleotide sequences (Zuckerkandl and Pauling 1965). Recently we introduced a novel method, termed volatility, to estimate selection pressures on proteins on the basis of their synonymous codon usage (Plotkin and Dushoff 2003; Plotkin et al. 2004). Here we provide a theoretical foundation for this approach. Under the Fisher-Wright model, we derive the expected frequencies of synonymous codons as a function of the strength of selection on amino acids, the mutation rate, and the effective population size. We analyze the conditions under which we can expect to draw inferences from biased codon usage, and we estimate the time scales required to establish and maintain such a signal. We find that synonymous codon usage can reliably distinguish between negative selection and neutrality only for organisms, such as some microbes, that experience large effective population sizes or periods of elevated mutation rates. The power of volatility to detect positive selection is also modest--requiring approximately 100 selected sites--but it depends less strongly on population size. We show that phenomena such as transient hyper-mutators can improve the power of volatility to detect selection, even when the neutral site heterozygosity is low. We also discuss several confounding factors, neglected by the Fisher-Wright model, that may limit the applicability of volatility in practice.

Algorithms↗

High-level periplasmic expression in Escherichia coli using a eukaryotic signal peptide: importance of codon usage at the 5' end of the coding sequence.

We investigated the ability of signal peptides of eukaryotic origin (human, mouse, and yeast) to efficiently direct model proteins to the Escherichia coli periplasm. These were compared against a well-characterized prokaryotic signal peptide-OmpA. Surprisingly, eukaryotic signal peptides can work very efficiently in E. coli, but require optimization of codon usage by codon-based mutagenesis of the signal peptide coding region. Analysis of the 5' of periplasmic and cytoplasmic E. coli genes shows some codon usage differences.

Amino Acid Sequence↗

Synthesis of a new Cre recombinase gene based on optimal codon usage for mammalian systems.

The origin of the Cre recombinase gene is bacteriophage P1, and thus the codon usages are different from in mammals. In order to adapt this codon usage for mammals, we synthesized a "mammalian Cre recombinase gene" and examined its expression in Chinese hamster ovarian tumor (CHO) cells. Significant increases in protein production as well as mRNA levels were observed. When the recombination efficiency was compared using CHO cell transfectants having a cDNA containing loxP sites, the "mammalian Cre recombinase gene" recombined the loxP sites much more efficiently than the wild-type Cre recombinase gene.

Amino Acid Sequence↗

Codon usage in Homo sapiens: evidence for a coding pattern on the non-coding strand and evolutionary implications of dinucleotide discrimination.

This study reports the analysis of codon usage in 35 complete Homo sapiens genes. Both codon frequency and inter-codon interference exhibit patterns of evolutionary interest. There is a significant positive correlation between the frequency with which a given codon is used and the frequency with which its complement is used. Since the frequency of appearance of the complementary codon on the coding strand is equal to the frequency of appearance of the original codon on the non-coding strand, in the same phase, the non-coding strand is found to resemble the coding strand in triplet composition. The same effect has been observed in Escherichia coli. This preference for the use of certain complementary triplets as codons suggests that the evolution of the use of the genetic code depended to some extent upon the double-stranded nature of the coding material. In addition, the effect of discrimination against the use of two dinucleotides, CpG and UpA, is observed in codon usage and also in adjacent codon interference. Codons beginning with G, or A, are unlikely to be preceded by codons ending in C, or U, respectively. Consideration of codon assignment in the genetic code together with the observed CpG infrequency suggests that the evolution of the code may have been influenced by conditions in which the use of CpG dinucleotides was unfavorable. The infrequent use of UpA dinucleotides can be explained as the result of frameshift mutation during gene evolution.

Base Sequence↗

Recent selection on synonymous codon usage in Drosophila.

Evidence from a variety of sources indicates that selection has influenced synonymous codon usage in Drosophila. It has generally been difficult, however, to distinguish selection that acted in the distant past from ongoing selection. However, under a neutral model, polymorphisms usually reflect more recent mutations than fixed differences between species and may, therefore, be useful for inferring recent selection. If the ancestral state is preferred, selection should shift the frequency distribution of derived states/site toward lower values; if the ancestral is unpreferred, selection should increase the number of derived states/site. Polymorphisms were classified as ancestrally preferred or unpreferred for several genes of D. simulans and D. melanogaster. A computer simulation of coalescence was employed to derive the expected frequency distributions of derived states/site under various modifications of the Wright-Fisher neutral model, and distributions of test statistics (t and Mann-Whitney U) were derived by appropriate sampling. One-tailed tests were applied to transformed frequency data to assess whether the two frequency distributions deviated from neutral expectations in the direction predicted by selection on codon usage. Several genes from D. simulans appear to be subject to recent selection on synonymous codons, including one gene with low codon bias, esterase-6. Selection may also be acting in D. melanogaster.

Animals↗

Replicational and transcriptional selection on codon usage in Borrelia burgdorferi.

With more than 10 fully sequenced, publicly available prokaryotic genomes, it is now becoming possible to gain useful insights into genome evolution. Before the genome era, many evolutionary processes were evaluated from limited data sets and evolutionary models were constructed on the basis of small amounts of evidence. In this paper, I show that genes on the Borrelia burgdorferi genome have two separate, distinct, and significantly different codon usages, depending on whether the gene is transcribed on the leading or lagging strand of replication. Asymmetrical replication is the major source of codon usage variation. Replicational selection is responsible for the higher number of genes on the leading strands, and transcriptional selection appears to be responsible for the enrichment of highly expressed genes on these strands. Replicational-transcriptional selection, therefore, has an influence on the codon usage of a gene. This is a new paradigm of codon selection in prokaryotes.

Borrelia burgdorferi Group↗

Growth rate-optimised tRNA abundance and codon usage.

The abundance of different tRNAs in Escherichia coli at different growth rates correlates strongly with the usage of the corresponding cognate codons. On the assumption that the investment in the translation system is optimised to provide a maximal growth rate, the relationship between tRNA levels and codon usage can be predicted. When the complications due to different degeneracies and different association rate constants for the different tRNA-codon combinations are accounted for, recent data from the literature indicate that the predicted relations hold up very well: the tRNA levels correlate with codon frequencies in a way that would support a maximal growth rate. The relations can also be used to predict the association rate constant between an A-site codon and the cognate ternary complex. In the cases where they can be compared, the results agree reasonably well with experimental results from the literature.

Cell Division↗

Structural organization and unusual codon usage in the DNA polymerase gene from herpes simplex virus type 1.

We have analyzed the protein and nucleic acid sequences of the DNA polymerase from herpes simplex virus type 1 (HSV-1) to provide insight into the expression and possible structure of this enzyme. Extensive similarity between the amino acid sequence and that of the Epstein Barr virus DNA polymerase is reported. We describe probable structural similarities between these proteins and the use of these similarities to define structural and functional domains within the polymerase. Analysis of base composition and codon usage reveals that several genes from HSV-1, including DNA polymerase, exhibit a strong preference for guanine or cytosine at the third codon position. This preference may result from the high guanine + cytosine content of the virus and produces a highly restricted codon usage, different from that of the host cell. Consequences of the unusual codon usage for viral expression include the potential for extensive mRNA secondary structure.

Amino Acid Sequence↗

Comparison of codon usage measures and their applicability in prediction of microbial gene expressivity.

BACKGROUND: There are a number of methods (also called: measures) currently in use that quantify codon usage in genes. These measures are often influenced by other sequence properties, such as length. This can introduce strong methodological bias into measurements; therefore we attempted to develop a method free from such dependencies. One of the common applications of codon usage analyses is to quantitatively predict gene expressivity. RESULTS: We compared the performance of several commonly used measures and a novel method we introduce in this paper--Measure Independent of Length and Composition (MILC). Large, randomly generated sequence sets were used to test for dependence on (i) sequence length, (ii) overall amount of codon bias and (iii) codon bias discrepancy in the sequences. A derivative of the method, named MELP (MILC-based Expression Level Predictor) can be used to quantitatively predict gene expression levels from genomic data. It was compared to other similar predictors by examining their correlation with actual, experimentally obtained mRNA or protein abundances. CONCLUSION: We have established that MILC is a generally applicable measure, being resistant to changes in gene length and overall nucleotide composition, and introducing little noise into measurements. Other methods, however, may also be appropriate in certain applications. Our efforts to quantitatively predict gene expression levels in several prokaryotes and unicellular eukaryotes met with varying levels of success, depending on the experimental dataset and predictor used. Out of all methods, MELP and Rainer Merkl's GCB method had the most consistent behaviour. A 'reference set' containing known ribosomal protein genes appears to be a valid starting point for a codon usage-based expressivity prediction.

Chi-Square Distribution↗

The repertoire of transfer RNA genes is tuned to codon usage bias in the genomes of Phytophthora sojae and Phytophthora ramorum.

In all, 238 and 155 transfer (t)RNA genes were predicted from the genomes of Phytophthora sojae and P. ramorum, respectively. After omitting pseudogenes and undetermined types of tRNA genes, there remained 208 P. sojae tRNA genes and 140 P. ramorum tRNA genes. There were 45 types of tRNA genes, with distinct anticodons, in each species. Fourteen common anticodon types of tRNAs are missing altogether from the genome in the two species; however, these appear to be compensated by wobbling of other tRNA anticodons in a manner which is tied to the codon bias in Phytophthora genes. The most abundant tRNA class was arginine in both P. sojae and P. ramorum. A codon usage table was generated for these two organisms from a total of 9,803,525 codons in P. sojae and 7,496,598 codons in P. ramorum. The most abundant codon type detected from the codon usage tables was GAG (encoding glutamic acid), whereas the most numerous tRNA gene had a methionine anticodon (CAT). The correlation between the frequencies of tRNA genes and the codon frequencies in protein-coding genes was very low (0.12 in P. sojae and 0.19 in P. ramorum); however, the correlation between amino acid tRNA gene frequency and the corresponding amino acid codon frequency in P. sojae and P. ramorum was substantially higher (0.53 in P. sojae and 0.77 in P. ramorum). The codon usage frequencies of P. sojae and P ramorum were very strongly correlated (0.99), as were tRNA gene frequencies (0.77). Approximately 60% of orthologous tRNA gene pairs in P sojae and P. ramorum are located in regions that have conserved synteny in the two species.

Anticodon↗

Correlation between sequence conservation of the 5' untranslated region and codon usage bias in Mus musculus genes.

The codon adaptation index (CAI) values of all protein-coding sequences of the full-length cDNA libraries of Mus musculus were computed based on the RIKEN mouse full-length cDNA library. We have also computed the extent of consensus in flanking sequences of the initiator ATG codon based on the 'relative entropy' values of respective nucleotide positions (from -20 to +12 bp relative to the initiator ATG codon) for each group of genes classified by CAI values. With regard to the two nucleotides positions (-3 and +4) known to be highly conserved in Kozak's consensus sequence, a clear correlation between CAI values and relative entropy values was observed at position -3 but this was not significant at position +4, although a significant correlation was found at position -1 of the consensus sequence. Further, although no correlation was observed at any additional positions, relative entropy values were very high at positions -4, -6, and -8 in genes with high CAI values. These findings suggest that the extent of conservation in the flanking sequence of the initiator ATG codon including Kozak's consensus sequence was an important factor in modulation of the translation efficiency as well as synonymous codon usage bias particularly in highly expressed genes.

5' Untranslated Regions↗

Codon usage limitation in the expression of HIV-1 envelope glycoprotein.

BACKGROUND: The expression of both the env and gag gene products of human immunodeficiency virus type 1 (HIV-1) is known to be limited by cis elements in the viral RNA that impede egress from the nucleus and reduce the efficiency of translation. Identifying these elements has proven difficult, as they appear to be disseminated throughout the viral genome. RESULTS: Here, we report that selective codon usage appears to account for a substantial fraction of the inefficiency of viral protein synthesis, independent of any effect on improved nuclear export. The codon usage effect is not specific to transcripts of HIV-1 origin. Re-engineering the coding sequence of a model protein (Thy-1) with the most prevalent HIV-1 codons significantly impairs Thy-1 expression, whereas altering the coding sequence of the jellyfish green fluorescent protein gene to conform to the favored codons of highly expressed human proteins results in a substantial increase in expression efficiency. CONCLUSIONS: Codon-usage effects are a major impediment to the efficient expression of HIV-1 genes. Although mammalian genes do not show as profound a bias as do Escherichia coli genes, other proteins that are poorly expressed in mammalian cells can benefit from codon re-engineering.

Animals↗

Nucleotide sequence of the gene for the immunity protein to colicin A. Analysis of codon usage of immunity proteins as compared to colicins.

The nucleotide sequence of the structural gene for the immunity protein to colicin A (cai) has been established. This sequence consists of 534 base pairs. According to the predicted amino acid sequence, the polypeptide chain of this immunity protein comprises 178 amino acids and has a relative molecular mass of 20462. As expected from its localization in the inner membrane, large hydrophobic fragments are found along the polypeptide chain that also contains clusters of mostly positively charged residues. The cai like the ceiA genes encode proteins that are weakly expressed as compared to the corresponding colicins (A and E1). Codon usage reflects this difference. In contrast, the four genes for immunity to cloacin DF13 and to colicin E3 and for these bacteriocins, all of which are highly expressed and are organized in operon, display similar codon usage. These results are discussed with regards to the possible relationship between expressivity and codon usage.

Bacterial Proteins↗

Codon-usage bias versus gene conversion in the evolution of yeast duplicate genes.

Many Saccharomyces cerevisiae duplicate genes that were derived from an ancient whole-genome duplication (WGD) unexpectedly show a small synonymous divergence (K(S)), a higher sequence similarity to each other than to orthologues in Saccharomyces bayanus, or slow evolution compared with the orthologue in Kluyveromyces waltii, a non-WGD species. This decelerated evolution was attributed to gene conversion between duplicates. Using approximately 300 WGD gene pairs in four species and their orthologues in non-WGD species, we show that codon-usage bias and protein-sequence conservation are two important causes for decelerated evolution of duplicate genes, whereas gene conversion is effective only in the presence of strong codon-usage bias or protein-sequence conservation. Furthermore, we find that change in mutation pattern or in tDNA copy number changed codon-usage bias and increased the K(S) distance between K. waltii and S. cerevisiae. Intriguingly, some proteins showed fast evolution before the radiation of WGD species but little or no sequence divergence between orthologues and paralogues thereafter, indicating that functional conservation after the radiation may also be responsible for decelerated evolution in duplicates.

Codon↗