PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Growth rate-optimised tRNA abundance and codon usage.

The abundance of different tRNAs in Escherichia coli at different growth rates correlates strongly with the usage of the corresponding cognate codons. On the assumption that the investment in the translation system is optimised to provide a maximal growth rate, the relationship between tRNA levels and codon usage can be predicted. When the complications due to different degeneracies and different association rate constants for the different tRNA-codon combinations are accounted for, recent data from the literature indicate that the predicted relations hold up very well: the tRNA levels correlate with codon frequencies in a way that would support a maximal growth rate. The relations can also be used to predict the association rate constant between an A-site codon and the cognate ternary complex. In the cases where they can be compared, the results agree reasonably well with experimental results from the literature.

Cell Division↗

Structural organization and unusual codon usage in the DNA polymerase gene from herpes simplex virus type 1.

We have analyzed the protein and nucleic acid sequences of the DNA polymerase from herpes simplex virus type 1 (HSV-1) to provide insight into the expression and possible structure of this enzyme. Extensive similarity between the amino acid sequence and that of the Epstein Barr virus DNA polymerase is reported. We describe probable structural similarities between these proteins and the use of these similarities to define structural and functional domains within the polymerase. Analysis of base composition and codon usage reveals that several genes from HSV-1, including DNA polymerase, exhibit a strong preference for guanine or cytosine at the third codon position. This preference may result from the high guanine + cytosine content of the virus and produces a highly restricted codon usage, different from that of the host cell. Consequences of the unusual codon usage for viral expression include the potential for extensive mRNA secondary structure.

Amino Acid Sequence↗

Comparison of codon usage measures and their applicability in prediction of microbial gene expressivity.

BACKGROUND: There are a number of methods (also called: measures) currently in use that quantify codon usage in genes. These measures are often influenced by other sequence properties, such as length. This can introduce strong methodological bias into measurements; therefore we attempted to develop a method free from such dependencies. One of the common applications of codon usage analyses is to quantitatively predict gene expressivity. RESULTS: We compared the performance of several commonly used measures and a novel method we introduce in this paper--Measure Independent of Length and Composition (MILC). Large, randomly generated sequence sets were used to test for dependence on (i) sequence length, (ii) overall amount of codon bias and (iii) codon bias discrepancy in the sequences. A derivative of the method, named MELP (MILC-based Expression Level Predictor) can be used to quantitatively predict gene expression levels from genomic data. It was compared to other similar predictors by examining their correlation with actual, experimentally obtained mRNA or protein abundances. CONCLUSION: We have established that MILC is a generally applicable measure, being resistant to changes in gene length and overall nucleotide composition, and introducing little noise into measurements. Other methods, however, may also be appropriate in certain applications. Our efforts to quantitatively predict gene expression levels in several prokaryotes and unicellular eukaryotes met with varying levels of success, depending on the experimental dataset and predictor used. Out of all methods, MELP and Rainer Merkl's GCB method had the most consistent behaviour. A 'reference set' containing known ribosomal protein genes appears to be a valid starting point for a codon usage-based expressivity prediction.

Chi-Square Distribution↗

Correlation between sequence conservation of the 5' untranslated region and codon usage bias in Mus musculus genes.

The codon adaptation index (CAI) values of all protein-coding sequences of the full-length cDNA libraries of Mus musculus were computed based on the RIKEN mouse full-length cDNA library. We have also computed the extent of consensus in flanking sequences of the initiator ATG codon based on the 'relative entropy' values of respective nucleotide positions (from -20 to +12 bp relative to the initiator ATG codon) for each group of genes classified by CAI values. With regard to the two nucleotides positions (-3 and +4) known to be highly conserved in Kozak's consensus sequence, a clear correlation between CAI values and relative entropy values was observed at position -3 but this was not significant at position +4, although a significant correlation was found at position -1 of the consensus sequence. Further, although no correlation was observed at any additional positions, relative entropy values were very high at positions -4, -6, and -8 in genes with high CAI values. These findings suggest that the extent of conservation in the flanking sequence of the initiator ATG codon including Kozak's consensus sequence was an important factor in modulation of the translation efficiency as well as synonymous codon usage bias particularly in highly expressed genes.

5' Untranslated Regions↗

Codon usage limitation in the expression of HIV-1 envelope glycoprotein.

BACKGROUND: The expression of both the env and gag gene products of human immunodeficiency virus type 1 (HIV-1) is known to be limited by cis elements in the viral RNA that impede egress from the nucleus and reduce the efficiency of translation. Identifying these elements has proven difficult, as they appear to be disseminated throughout the viral genome. RESULTS: Here, we report that selective codon usage appears to account for a substantial fraction of the inefficiency of viral protein synthesis, independent of any effect on improved nuclear export. The codon usage effect is not specific to transcripts of HIV-1 origin. Re-engineering the coding sequence of a model protein (Thy-1) with the most prevalent HIV-1 codons significantly impairs Thy-1 expression, whereas altering the coding sequence of the jellyfish green fluorescent protein gene to conform to the favored codons of highly expressed human proteins results in a substantial increase in expression efficiency. CONCLUSIONS: Codon-usage effects are a major impediment to the efficient expression of HIV-1 genes. Although mammalian genes do not show as profound a bias as do Escherichia coli genes, other proteins that are poorly expressed in mammalian cells can benefit from codon re-engineering.

Animals↗

Nucleotide sequence of the gene for the immunity protein to colicin A. Analysis of codon usage of immunity proteins as compared to colicins.

The nucleotide sequence of the structural gene for the immunity protein to colicin A (cai) has been established. This sequence consists of 534 base pairs. According to the predicted amino acid sequence, the polypeptide chain of this immunity protein comprises 178 amino acids and has a relative molecular mass of 20462. As expected from its localization in the inner membrane, large hydrophobic fragments are found along the polypeptide chain that also contains clusters of mostly positively charged residues. The cai like the ceiA genes encode proteins that are weakly expressed as compared to the corresponding colicins (A and E1). Codon usage reflects this difference. In contrast, the four genes for immunity to cloacin DF13 and to colicin E3 and for these bacteriocins, all of which are highly expressed and are organized in operon, display similar codon usage. These results are discussed with regards to the possible relationship between expressivity and codon usage.

Bacterial Proteins↗

Codon-usage bias versus gene conversion in the evolution of yeast duplicate genes.

Many Saccharomyces cerevisiae duplicate genes that were derived from an ancient whole-genome duplication (WGD) unexpectedly show a small synonymous divergence (K(S)), a higher sequence similarity to each other than to orthologues in Saccharomyces bayanus, or slow evolution compared with the orthologue in Kluyveromyces waltii, a non-WGD species. This decelerated evolution was attributed to gene conversion between duplicates. Using approximately 300 WGD gene pairs in four species and their orthologues in non-WGD species, we show that codon-usage bias and protein-sequence conservation are two important causes for decelerated evolution of duplicate genes, whereas gene conversion is effective only in the presence of strong codon-usage bias or protein-sequence conservation. Furthermore, we find that change in mutation pattern or in tDNA copy number changed codon-usage bias and increased the K(S) distance between K. waltii and S. cerevisiae. Intriguingly, some proteins showed fast evolution before the radiation of WGD species but little or no sequence divergence between orthologues and paralogues thereafter, indicating that functional conservation after the radiation may also be responsible for decelerated evolution in duplicates.

Codon↗

Correlations between the compositional properties of human genes, codon usage, and amino acid composition of proteins.

We have analyzed the correlation that exists between the GC levels of third and first or second codon position for about 1400 human coding sequences. The linear relationship that was found indicates that the large differences in GC level of third codon positions of human genes are paralleled by smaller differences in GC levels of first and second codon positions. Whereas third codon position differences correspond to very large differences in codon usage within the human genome, the first and second codon position differences correspond to smaller, yet very remarkable, differences in the amino acid composition of encoded proteins. Because GC levels of codon positions are linearly correlated with the GC levels of the isochores harboring the corresponding genes, both codon usage and amino acid composition are different for proteins encoded by genes located in isochores of different GC levels. Furthermore, we have also shown that a linear relationship with a unit slope and a correlation coefficient of 0.77 exists between GC levels of introns and exons from the 238 human genes currently available for this analysis. Introns are, however, about 5% lower in GC, on average, than exons from the same genes.

Amino Acid Sequence↗

Seven GC-rich microbial genomes adopt similar codon usage patterns regardless of their phylogenetic lineages.

Seven GC-rich (group I) and three AT-rich (group II) microbial genomes are analyzed in this paper. The seven microbes in group I belong to different phylogenetic lineages, even different domains of life. The common feature is that they are highly GC-rich organisms, with more than 60% genomic GC content. Group II includes three bacteria, which belong to the same subdivision as Pseudomonas aeruginosa in group I. The genomic GC content of the three bacteria is in the range of 26-50%. It is shown that although the phylogenetic lineages of the organisms in group I are remote, the common feature of highly genomic GC content forces them to adopt similar codon usage patterns, which constitutes the basis of an algorithm using a set of universal parameters to recognize known genes in the seven genomes. The common codon usage pattern of function known genes in the seven genomes is GGS type, where G, G, and S are the bases of G, non-G, and G/C, respectively. On the contrary, although the phylogenetic lineages of the three bacteria in group II are quite close, the codon usage patterns of function known genes in these genomes are obviously distinct. There are no universal parameters to identify known genes in the three genomes in group II. It can be deduced that the genomic GC content is more important than phylogenetic lineage in gene recognition programs. We hope that the work might be useful for understanding the common characteristics in the organization of microbial genomes.

Algorithms↗

Graphic analysis of codon usage strategy in 1490 human proteins.

The frequencies of bases A (adenine), C (cytosine), G (guanine), and T (thymine) occurring in codon position i, denoted by ai, ci, gi, and ti, respectively (i = 1,2,3), have been calculated and diagrammatized for the 1490 human proteins in the codon usage table for primate genes compiled recently. Based on the characteristic graphs thus obtained, an overall picture of codon base distribution has been provided, and the relevant biological implication discussed. For the first codon position, it is shown in most cases that G is the most dominant base, and that the relationship g1 > a1 > c1 > t1 generally holds true. For the second codon position, A is generally the most dominant base and G is the one with the least occurrence frequently, with the relationship of a2 > t2 > c2 > g2. As to the third codon position, the values of g3 + c3 vary from 0.27 to 1, roughly keeping the relationship of c3 > g3 > a3 = t3 for the majority of cases. Interestingly, if the average frequencies for bases A, C, G, and T are defined as a = (a1 + a2 + a3)/3, c = (c1 + c2 + c3)/3, g = (g1 + g2 + g3)/3, and t = (t1 + t2 + t3)/3, respectively, we find that a2 + c2 + g2 + t2 < 1/3 is valid almost without exception. Such a characteristic inequality might reflect some inherent rule of codon usage, although its biological implications is unclear.(ABSTRACT TRUNCATED AT 250 WORDS)

Adenine↗

Different stop codon usage in two pseudohypotrich ciliates.

Based on rRNA phylogeny, morphologic and morphogenetic characters, two major groups of hypotrich ciliates can be distinguished: euhypotrichs and pseudohypotrichs. Through the sequencing of actin genes, we show here that, interestingly, the pseudohypotrichs Dyophrys sp. and Euplotes vannus have a different stop codon usage. In fact, the stop codon usage of the former species resembles that of euhypotrichs. This unexpected result is used to discuss the origin and acquisition of genetic code deviations in ciliates.

Actins↗

Synonymous codon usage in Pseudomonas aeruginosa PA01.

Pseudomonas aeruginosa PA01 has a large (6.7 Mbp) genome with a high (67%) G+C content. Codon usage in this species is dominated by this compositional bias, with the average G+C content at synonymously variable third positions of codons being 83%. Nevertheless, there is some variation of synonymous codon usage among genes. The nature and causes of this variation were investigated using multivariate statistical analyses. Three trends were identified. The major source of variation was attributable to genes with unusually low G+C content that are probably due to horizontal transfer. A lesser trend among genes was associated with the preferential use of putatively translationally optimal codons in genes expressed at high levels. In addition, genes on the leading strand of replication were on average more G+T-rich. Our findings contradict the results of two previous analyses, and the reasons for the discrepancies are discussed.

Amino Acids↗

The effect of codon usage on the oligonucleotide composition of the E. coli genome and identification of over- and underrepresented sequences by Markov chain analysis.

As shown in the accompanying paper (5), the oligonucleotide composition of the E. coli genome is highly asymmetric for sequences up to 6 bp in length when ranked from highest to lowest abundance. We show here that this largely reflects codon usage because heavily used codons were found in the highly abundant oligomers whereas rarely used codons, with some exceptions, occurred in sequences in low abundance. Furthermore, linear regression analysis revealed a strong correlation between the frequencies of each trinucleotide and its usage as a codon. Dinucleotides are also not randomly distributed across each codon position and the dinucleotide composition of genes that are transcribed but not translated (rRNA and tRNA genes) was highly related to that seen in genes encoding polypeptides. However, 45 tetra-, 8 penta-, and 6 hexanucleotides were significantly over- or underabundant by Markov chain analysis and could not be accounted for by codon usage. Of these underrepresented sequences, many were palindromes, including the Dam methylation site.

Base Sequence↗

Effects of codon usage versus putative 5'-mRNA structure on the expression of Fusarium solani cutinase in the Escherichia coli cytoplasm.

Matching the codon usage of recombinant genes to that of the expression host is a common strategy for increasing the expression of heterologous proteins in bacteria. However, while developing a cytoplasmic expression system for Fusarium solani cutinase in Escherichia coli, we found that altering codons to those preferred by E. coli led to significantly lower expression compared to the wild-type fungal gene, despite the presence of several rare E. coli codons in the fungal sequence. On the other hand, expression in the E. coli periplasm using a bacterial PhoA leader sequence resulted in high levels of expression for both the E. coli optimized and wild-type constructs. Sequence swapping experiments as well as calculations of predicted mRNA secondary structure provided support for the hypothesis that differential cytoplasmic expression of the E. coli optimized versus wild-type cutinase genes is due to differences in 5(') mRNA secondary structures. In particular, our results indicate that increased stability of 5(') mRNA secondary structures in the E. coli optimized transcript prevents efficient translation initiation in the absence of the phoA leader sequence. These results underscore the idea that potential 5(') mRNA secondary structures should be considered along with codon usage when designing a synthetic gene for high level expression in E. coli.

Amino Acid Sequence↗

[Bias of base composition and codon usage in pseudorabies virus genes].

The complete sequence of the Pseudorabies Virus (PRV) genomic DNA has not yet been determined, primarily because of the high content of G + C nucleotides of about 74%. We examined the base composition and codon usage of the 68 known PRV genes. As a result, we found a strong bias towards GC-rich codons especially NNC or NNG (N represents any one of four nucleotides) in PRV genes. This demonstrated that the usage bias of synonymous codon and amino acid is the main cause of the high G + C content of PRV. The results showed that the genome regions adjacent UL48, UL40, UL14, IE180 genes where the G + C content occurs as pronounced waves are corresponding to the replication origins. It was also found that the codon usage patterns of regulatory genes are apparently different from other PRV genes. A corresponding analysis of amino acid compositions indicated that the bias of codon usage could be related to the differences of gene function.

Amino Acids↗

Optimization of codon usage of poxvirus genes allows for improved transient expression in mammalian cells.

Transient expression of viral genes from certain poxviruses in uninfected mammalian cells can sometimes be unexpectedly inefficient. The reasons for poor expression levels can be due to a number of features of the gene cassette, such as cryptic splice sites, polymerase II termination sequences or motifs that lead to mRNA instability. Here we suggest that in some cases the problem of low protein expression in transfected mammalian cells may be due to inefficient codon usage. We have observed that for many poxvirus genes from the yatapoxvirus genus this deficiency can be overcome by synthesis of the gene with codon sequences optimized for expression in primate cells. This led us to examine colon usage across 2-dozen sequenced members of the Poxviridae. We conclude that codon usage is surprisingly divergent across the different Poxviridae genera but is much more conserved within a single genus. Thus, Poxviridae genera can be divided into distinct groups based on their observed codon bias. When viewed in this context, successful transient expression of transfected poxvirus genes in uninfected mammalian cells can be more accurately predicted based on codon bias. As a corollary, for specific poxvirus genes with less favorable codon usage, codon optimization can result in profoundly increased transient expression levels following transfection of uninfected mammalian cell lines.

Animals↗

The relationship among gene expression, folding free energy and codon usage bias in Escherichia coli.

Taking advantage of microarray data in Escherichia coli genome, the relationship among mRNA expression levels, folding free energy and codon usage bias are investigated. Our results indicate that mRNA expression is correlated to the stability of mRNA secondary structure and the codon usage bias. The decrease of the stability of mRNA structure contributes to the increase of mRNA expression. There is a negative correlation between codon adaptation index (CAI) and mRNA expression in genes with less stable structure. The relationship between the stability of mRNA structure and mRNA half-life indicates the stability of mRNA structure is different from mRNA half-life.

Codon↗