PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Correlations between the compositional properties of human genes, codon usage, and amino acid composition of proteins.

We have analyzed the correlation that exists between the GC levels of third and first or second codon position for about 1400 human coding sequences. The linear relationship that was found indicates that the large differences in GC level of third codon positions of human genes are paralleled by smaller differences in GC levels of first and second codon positions. Whereas third codon position differences correspond to very large differences in codon usage within the human genome, the first and second codon position differences correspond to smaller, yet very remarkable, differences in the amino acid composition of encoded proteins. Because GC levels of codon positions are linearly correlated with the GC levels of the isochores harboring the corresponding genes, both codon usage and amino acid composition are different for proteins encoded by genes located in isochores of different GC levels. Furthermore, we have also shown that a linear relationship with a unit slope and a correlation coefficient of 0.77 exists between GC levels of introns and exons from the 238 human genes currently available for this analysis. Introns are, however, about 5% lower in GC, on average, than exons from the same genes.

Amino Acid Sequence↗

Seven GC-rich microbial genomes adopt similar codon usage patterns regardless of their phylogenetic lineages.

Seven GC-rich (group I) and three AT-rich (group II) microbial genomes are analyzed in this paper. The seven microbes in group I belong to different phylogenetic lineages, even different domains of life. The common feature is that they are highly GC-rich organisms, with more than 60% genomic GC content. Group II includes three bacteria, which belong to the same subdivision as Pseudomonas aeruginosa in group I. The genomic GC content of the three bacteria is in the range of 26-50%. It is shown that although the phylogenetic lineages of the organisms in group I are remote, the common feature of highly genomic GC content forces them to adopt similar codon usage patterns, which constitutes the basis of an algorithm using a set of universal parameters to recognize known genes in the seven genomes. The common codon usage pattern of function known genes in the seven genomes is GGS type, where G, G, and S are the bases of G, non-G, and G/C, respectively. On the contrary, although the phylogenetic lineages of the three bacteria in group II are quite close, the codon usage patterns of function known genes in these genomes are obviously distinct. There are no universal parameters to identify known genes in the three genomes in group II. It can be deduced that the genomic GC content is more important than phylogenetic lineage in gene recognition programs. We hope that the work might be useful for understanding the common characteristics in the organization of microbial genomes.

Algorithms↗

Graphic analysis of codon usage strategy in 1490 human proteins.

The frequencies of bases A (adenine), C (cytosine), G (guanine), and T (thymine) occurring in codon position i, denoted by ai, ci, gi, and ti, respectively (i = 1,2,3), have been calculated and diagrammatized for the 1490 human proteins in the codon usage table for primate genes compiled recently. Based on the characteristic graphs thus obtained, an overall picture of codon base distribution has been provided, and the relevant biological implication discussed. For the first codon position, it is shown in most cases that G is the most dominant base, and that the relationship g1 > a1 > c1 > t1 generally holds true. For the second codon position, A is generally the most dominant base and G is the one with the least occurrence frequently, with the relationship of a2 > t2 > c2 > g2. As to the third codon position, the values of g3 + c3 vary from 0.27 to 1, roughly keeping the relationship of c3 > g3 > a3 = t3 for the majority of cases. Interestingly, if the average frequencies for bases A, C, G, and T are defined as a = (a1 + a2 + a3)/3, c = (c1 + c2 + c3)/3, g = (g1 + g2 + g3)/3, and t = (t1 + t2 + t3)/3, respectively, we find that a2 + c2 + g2 + t2 < 1/3 is valid almost without exception. Such a characteristic inequality might reflect some inherent rule of codon usage, although its biological implications is unclear.(ABSTRACT TRUNCATED AT 250 WORDS)

Adenine↗

Different stop codon usage in two pseudohypotrich ciliates.

Based on rRNA phylogeny, morphologic and morphogenetic characters, two major groups of hypotrich ciliates can be distinguished: euhypotrichs and pseudohypotrichs. Through the sequencing of actin genes, we show here that, interestingly, the pseudohypotrichs Dyophrys sp. and Euplotes vannus have a different stop codon usage. In fact, the stop codon usage of the former species resembles that of euhypotrichs. This unexpected result is used to discuss the origin and acquisition of genetic code deviations in ciliates.

Actins↗

Synonymous codon usage in Pseudomonas aeruginosa PA01.

Pseudomonas aeruginosa PA01 has a large (6.7 Mbp) genome with a high (67%) G+C content. Codon usage in this species is dominated by this compositional bias, with the average G+C content at synonymously variable third positions of codons being 83%. Nevertheless, there is some variation of synonymous codon usage among genes. The nature and causes of this variation were investigated using multivariate statistical analyses. Three trends were identified. The major source of variation was attributable to genes with unusually low G+C content that are probably due to horizontal transfer. A lesser trend among genes was associated with the preferential use of putatively translationally optimal codons in genes expressed at high levels. In addition, genes on the leading strand of replication were on average more G+T-rich. Our findings contradict the results of two previous analyses, and the reasons for the discrepancies are discussed.

Amino Acids↗

The effect of codon usage on the oligonucleotide composition of the E. coli genome and identification of over- and underrepresented sequences by Markov chain analysis.

As shown in the accompanying paper (5), the oligonucleotide composition of the E. coli genome is highly asymmetric for sequences up to 6 bp in length when ranked from highest to lowest abundance. We show here that this largely reflects codon usage because heavily used codons were found in the highly abundant oligomers whereas rarely used codons, with some exceptions, occurred in sequences in low abundance. Furthermore, linear regression analysis revealed a strong correlation between the frequencies of each trinucleotide and its usage as a codon. Dinucleotides are also not randomly distributed across each codon position and the dinucleotide composition of genes that are transcribed but not translated (rRNA and tRNA genes) was highly related to that seen in genes encoding polypeptides. However, 45 tetra-, 8 penta-, and 6 hexanucleotides were significantly over- or underabundant by Markov chain analysis and could not be accounted for by codon usage. Of these underrepresented sequences, many were palindromes, including the Dam methylation site.

Base Sequence↗

Effects of codon usage versus putative 5'-mRNA structure on the expression of Fusarium solani cutinase in the Escherichia coli cytoplasm.

Matching the codon usage of recombinant genes to that of the expression host is a common strategy for increasing the expression of heterologous proteins in bacteria. However, while developing a cytoplasmic expression system for Fusarium solani cutinase in Escherichia coli, we found that altering codons to those preferred by E. coli led to significantly lower expression compared to the wild-type fungal gene, despite the presence of several rare E. coli codons in the fungal sequence. On the other hand, expression in the E. coli periplasm using a bacterial PhoA leader sequence resulted in high levels of expression for both the E. coli optimized and wild-type constructs. Sequence swapping experiments as well as calculations of predicted mRNA secondary structure provided support for the hypothesis that differential cytoplasmic expression of the E. coli optimized versus wild-type cutinase genes is due to differences in 5(') mRNA secondary structures. In particular, our results indicate that increased stability of 5(') mRNA secondary structures in the E. coli optimized transcript prevents efficient translation initiation in the absence of the phoA leader sequence. These results underscore the idea that potential 5(') mRNA secondary structures should be considered along with codon usage when designing a synthetic gene for high level expression in E. coli.

Amino Acid Sequence↗

[Bias of base composition and codon usage in pseudorabies virus genes].

The complete sequence of the Pseudorabies Virus (PRV) genomic DNA has not yet been determined, primarily because of the high content of G + C nucleotides of about 74%. We examined the base composition and codon usage of the 68 known PRV genes. As a result, we found a strong bias towards GC-rich codons especially NNC or NNG (N represents any one of four nucleotides) in PRV genes. This demonstrated that the usage bias of synonymous codon and amino acid is the main cause of the high G + C content of PRV. The results showed that the genome regions adjacent UL48, UL40, UL14, IE180 genes where the G + C content occurs as pronounced waves are corresponding to the replication origins. It was also found that the codon usage patterns of regulatory genes are apparently different from other PRV genes. A corresponding analysis of amino acid compositions indicated that the bias of codon usage could be related to the differences of gene function.

Amino Acids↗

Optimization of codon usage of poxvirus genes allows for improved transient expression in mammalian cells.

Transient expression of viral genes from certain poxviruses in uninfected mammalian cells can sometimes be unexpectedly inefficient. The reasons for poor expression levels can be due to a number of features of the gene cassette, such as cryptic splice sites, polymerase II termination sequences or motifs that lead to mRNA instability. Here we suggest that in some cases the problem of low protein expression in transfected mammalian cells may be due to inefficient codon usage. We have observed that for many poxvirus genes from the yatapoxvirus genus this deficiency can be overcome by synthesis of the gene with codon sequences optimized for expression in primate cells. This led us to examine colon usage across 2-dozen sequenced members of the Poxviridae. We conclude that codon usage is surprisingly divergent across the different Poxviridae genera but is much more conserved within a single genus. Thus, Poxviridae genera can be divided into distinct groups based on their observed codon bias. When viewed in this context, successful transient expression of transfected poxvirus genes in uninfected mammalian cells can be more accurately predicted based on codon bias. As a corollary, for specific poxvirus genes with less favorable codon usage, codon optimization can result in profoundly increased transient expression levels following transfection of uninfected mammalian cell lines.

Animals↗

The relationship among gene expression, folding free energy and codon usage bias in Escherichia coli.

Taking advantage of microarray data in Escherichia coli genome, the relationship among mRNA expression levels, folding free energy and codon usage bias are investigated. Our results indicate that mRNA expression is correlated to the stability of mRNA secondary structure and the codon usage bias. The decrease of the stability of mRNA structure contributes to the increase of mRNA expression. There is a negative correlation between codon adaptation index (CAI) and mRNA expression in genes with less stable structure. The relationship between the stability of mRNA structure and mRNA half-life indicates the stability of mRNA structure is different from mRNA half-life.

Codon↗

[Optimized codon usage enhances the expression and immunogenicity of DNA vaccine encoding the HPV 6b E7 gene].

OBJECTIVE: To analyze the influence of optimal codon usage on the expression levels and immunogenicity of DNA vaccines, encoding the human papillomavirus type 6b (HPV 6b) E7 gene. METHODS: The full length E7 gene of HPV 6b was modified to substitute human preferred codon for rarely used codon, and three mutations were introduced into the pRB binding site of HPV 6b E7 to eliminate its transformation potential. The codon optimized and mutated E7 gene (hu-mE7) were cloned into the Kpn I and EcoR I site of the pcDNA3 mammalian expression vector, the in vitro expression of the hu-mE7 gene and the immunogenicity of hu-mE7 DNA vaccine were compared with the wt-E7gene. RESULTS: The in vitro expression of pcDNA3-hu-mE7 was much higher than the classical wt-E7 plasmid in monkey COS-1 cell line. Mice immunized intramuscularly with the pcDNA3-hu-mE7 showed that the codon modified E7 gene induced a stronger IFN-gamma ratios than the wt-E7 gene. CONCLUSIONS: These results suggest that the optimized codon usage contributes to the enhancement of gene expression and immunogenicity of HPV 6b E7 gene.

Animals↗

Using codon usage to predict genes origin: is the Escherichia coli outer membrane a patchwork of products from different genomes?

Analysis of the codon usage of genes coding for the structural components of the outer membrane in Escherichia coli, is consistent with the requirement for high expression of these genes. Because porins (which constitute the major protein component of the outer membrane), and LPS (which constitute the major outermost constituent of the outer membrane), are synthesized from genes displaying widely different codon usage, it is possible to investigate the origin of the outer membrane. The analysis predicts that the outer membrane might originate from a genome other than the genome coding for the major part of the cell. Such a special origin would explain in structural terms, the likely lethality of porins if they were inadvertently inserted within the inner membrane, giving rise to the Gram-negative bacterial type, having an envelope comprising two membranes, instead of a single cytoplasmic membrane and a murein sacculus.

Bacterial Outer Membrane Proteins↗

Codon replacement in the PGK1 gene of Saccharomyces cerevisiae: experimental approach to study the role of biased codon usage in gene expression.

The coding sequences of genes in the yeast Saccharomyces cerevisiae show a preference for 25 of the 61 possible coding triplets. The degree of this biased codon usage in each gene is positively correlated to its expression level. Highly expressed genes use these 25 major codons almost exclusively. As an experimental approach to studying biased codon usage and its possible role in modulating gene expression, systematic codon replacements were carried out in the highly expressed PGK1 gene. The expression of phosphoglycerate kinase (PGK) was studied both on a high-copy-number plasmid and as a single copy gene integrated into the chromosome. Replacing an increasing number (up to 39% of all codons) of major codons with synonymous minor ones at the 5' end of the coding sequence caused a dramatic decline of the expression level. The PGK protein levels dropped 10-fold. The steady-state mRNA levels also declined, but to a lesser extent (threefold). Our data indicate that this reduction in mRNA levels was due to destabilization caused by impaired translation elongation at the minor codons. By preventing translation of the PGK mRNAs by the introduction of a stop codon 3' and adjacent to the start codon, the steady-state mRNA levels decreased dramatically. We conclude that efficient mRNA translation is required for maintaining mRNA stability in S. cerevisiae. These findings have important implications for the study of the expression of heterologous genes in yeast cells.

Amino Acid Sequence↗

Preferential codon usage and two types of repetitive motifs in the fibroin gene of the Chinese oak silkworm, Antheraea pernyi.

In this paper we describe the peculiar structures and preferential codon usage found in wild silkworm fibroin genes. We determined a 1350 bp nucleotide sequence from the Chinese oak silkworm, Antheraea pernyi. The deduced amino acid sequence was partitioned into thirteen polyalanine-containing repetitive motifs, which was one of the characteristics of Antheraea fibroins. Eleven of these arrays can be classified into two types of motifs depending on difference in amino acid sequences following polyalanine. Repetitive motifs structurally similar to those of A. pernyi were detected in a homologue of the Japanese oak silkworm, Antheraea yamamai. The most remarkable feature of this study was preferential codon usage, especially seen in alanine synonymous codons within both homologues of Antheraea: isocodon GCA most frequently occurred in alanine isocodons. In contrast, GCU isocodon was the most abundant in Bombyx mori fibroin heavy chain that lacks polyalanine arrays. This result strongly suggests different modes of selective constraint between the two types of fibroin gene. The similar finding that GCA isocodon was most frequent in two dragline silk sequences of the spider, Nephila clavipes, is consistent with our results because of the repetitive polyalanine-containing arrays seen in spider dragline silk.

Amino Acid Sequence↗

Codon usage tabulated from the international DNA sequence databases.

Codon usage in 87 602 genes has been calculated using the nucleotide sequence data obtained from the GenBank Genetic Sequence Data Bank (Release 90.0; September 1995). The database is called the CUTG Database; the complete form of the database can be obtained by anonymous ftp from DDBJ and a part of the database, which lists the frequency of codon use in each organism, is made searchable through our World Wide Web server.

Base Sequence↗

Codon usage optimization of HIV type 1 subtype C gag, pol, env, and nef genes: in vitro expression and immune responses in DNA-vaccinated mice.

Codon usage optimization of human immunodeficiency virus type 1 (HIV-1) structural genes has been shown to increase protein expression in vitro as well as in the context of DNA vaccines in vivo; however, all optimized genes reported thus far are derived from HIV-1 (group M) subtype B viruses. Here, we report the generation and biological characterization of codon usage-optimized gag, pol, env (gp160, gp140, gp120), and nef genes from a primary (nonrecombinant) HIV-1 subtype C isolate. After transfection into 293T cells, optimized subtype C genes expressed one to two orders of magnitude more protein (as determined by immunoblot densitometry) than the corresponding wild-type constructs. This effect was most pronounced for gp160, gp140, Gag, and Pol (>250-fold), but was also observed for gp120 and Nef (45- and 20-fold, respectively). Optimized gp160- and gp140-derived glycoproteins were processed, incorporated into virus particles, and mediated virus entry when expressed in trans to complement an env-minus HIV-1 provirus. Mice immunized with optimized gp140 DNA developed antibody as well as CD4+ and CD8+ T cell immune responses that were orders of magnitude greater than those of mice immunized with wild-type gp140 DNA. These data confirm and extend previous studies of codon usage optimization of HIV-1 genes to the most prevalent group M subtype. Our panel of matched optimized and wild-type subtype C genes should prove valuable for studies of protein expression and function, the generation of subtype-specific immunological reagents, and the production of DNA-based sub-unit vaccines directed against a broader spectrum of viruses.

AIDS Vaccines↗

Codon usage domains over bacterial chromosomes.

The geography of codon bias distributions over prokaryotic genomes and its impact upon chromosomal organization are analyzed. To this aim, we introduce a clustering method based on information theory, specifically designed to cluster genes according to their codon usage and apply it to the coding sequences of Escherichia coli and Bacillus subtilis. One of the clusters identified in each of the organisms is found to be related to expression levels, as expected, but other groups feature an over-representation of genes belonging to different functional groups, namely horizontally transferred genes, motility, and intermediary metabolism. Furthermore, we show that genes with a similar bias tend to be close to each other on the chromosome and organized in coherent domains, more extended than operons, demonstrating a role of translation in structuring bacterial chromosomes. It is argued that a sizeable contribution to this effect comes from the dynamical compartimentalization induced by the recycling of tRNAs, leading to gene expression rates dependent on their genomic and expression context.

Amino Acids↗