PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Effects of an opal termination codon preceding the nsP4 gene sequence in the O'Nyong-Nyong virus genome on Anopheles gambiae infectivity.

The genomic RNA of an alphavirus encodes four different nonstructural proteins, nsP1, nsP2, nsP3, and nsP4. The polyprotein P123 is produced when translation terminates at an opal termination codon between nsP3 and nsP4. The polyprotein P1234 is produced when translational readthrough occurs or when the opal termination codon has been replaced by a sense codon in the alphavirus genome. Evolutionary pressures appear to have maintained genomic sequences encoding both a stop codon (opal) and an open reading frame (arginine) as a general feature of the O'nyong-nyong virus (ONNV) genome, indicating that both are required at some point. Alternate replication of ONNVs in both vertebrate and invertebrate hosts may determine predominance of a particular codon at this locus in the viral quasispecies. However, no systematic study has previously tested this hypothesis in whole animals. We report here the results of the first study to investigate in a natural mosquito host the functional significance of the opal stop codon in an alphavirus genome. We used a full-length cDNA clone of ONNV to construct a series of mutants in which the arginine between nsP3 and nsP4 was replaced with an opal, ochre, or amber stop codon. The presence of an opal stop codon upstream of nsP4 nearly doubled (75.5%) the infectivity of ONNV over that of virus possessing a codon for the amino acid arginine at the corresponding position (39.8%). Although the frequency with which the opal virus disseminated from the mosquito midgut did not differ significantly from that of the arginine virus on days 8 and 10, dissemination did began earlier in mosquitoes infected with the opal virus. Although a clear fitness advantage is provided to ONNV by the presence of an opal codon between nsP3 and nsP4 in Anopheles gambiae, sequence analysis of ONNV RNA extracted from mosquito bodies and heads indicated codon usage at this position corresponded with that of the virus administered in the blood meal. These results suggest that while selection of ONNV variants is occurring, de novo mutation at the position between nsP3 and nsP4 does not readily occur in the mosquito. Taken together, these results suggest that the primary fitness advantage provided to ONNV by the presence of an opal codon between nsP3 and nsP4 is related to mosquito infectivity.

Alphavirus↗

Conservation of tandem stop codons in yeasts.

BACKGROUND: It has been long thought that the stop codon in a gene is followed by another stop codon that acts as a backup if the real one is read through by a near-cognate tRNA. The existence of such 'tandem stop codons', however, remains elusive. RESULTS: Here we show that a statistical excess of stop codons has evolved at the third codon downstream of the real stop codon UAA in yeasts. Comparative analysis indicates that stop codons at this location are considerably more conserved than sense codons, suggesting that these tandem stop codons are maintained by selection. We evaluated the influence of expression levels of genes and other biological factors on the distribution of tandem stop codons. Our results suggest that expression level is an important factor influencing the presence of tandem stop codons. CONCLUSIONS: Our study demonstrates the existence of tandem stop codons, which represent one of many meaningful genomic features that are driven by relatively weak selective forces.

Base Sequence↗

Codon usage in plant peroxidase genes.

Codon preference and asymmetry in usage in the DNA sequences encoding the mature enzyme protein of 24 plant peroxidases from 12 different species were examined. Codon usage in highly conserved/non-conserved areas of the sequences was analysed, as well as possible deficiency/excess in CpG dinucleotides in the pairs of codon positions. Sequence relationships displayed by overall codon usage, dinucleotide frequencies within codons, and amino acid sequences were also studied. The main findings were: (1) Monocots clustered separately from dicots for overall codon usage and dinucleotide frequencies in codon positions 2 and 3, with six and seven clusters respectively discernible among these 24 peroxidase sequences. The monocot/dicot distinction disappeared in the four clusters among the mature protein amino acid sequences. Overall codon usage in sequences from monocotyledon and dicotyledon species differed, the monocots favouring codons with C or G in the third position. (2) Codon usage was biassed in many sequences, asymmetry was particularly noticeable in the monocots. (3) For repeated amino acids within conserved areas, codon preference appeared dependent on the order in which the repeated amino acid occurred, so that its usage of synonymous codons frequently balanced out.

Amino Acid Sequence↗

Codon selection in yeast.

Extreme codon bias is seen for the Saccharomyces cerevisiae genes for the fermentative alcohol dehydrogenase isozyme I (ADH-I) and glyceraldehyde-3-phosphate dehydrogenase. Over 98% of the 1004 amino acid residues analyzed by DNA sequencing are coded for by a select 25 of the 61 possible coding triplets. These preferred codons tend to be highly homologous to the anticodons of the major yeast isoacceptor tRNA species. Codons which necessitate site by side GC base pairs between the codons and the tRNA anticodons are always avoided whenever possible. Codons containing 100% G, C, A, U, GC, or AU are also avoided. This provides for approximately equivalent codon-anticodon binding energies for all preferred triplets. All sequenced yeast genes show a distinct preference for these same 25 codons. The degree of preference varies from greater than 90% for glyceraldehyde-3-phosphate dehydrogenase and ADH-I to less than 20% for iso-2 cytochrome c. The degree of bias for these 25 preferred triplets in each gene is correlated with the level of its mRNA in the cytoplasm. Genes which are strongly expressed are more biased than genes with a lower level of expression. A similar phenomenon is observed in the codon preferences of highly expressed genes in Escherichia coli. High levels of gene expression are well correlated with high levels of codon bias toward 22 of the 61 coding triplets. As in yeast, these preferred codons are highly complementary to the major cellular isoacceptor tRNA species. In at least four cases (Ala, Arg, Leu, and Val), these preferred E. coli codons are incompatible with the preferred yeast codons.

Alcohol Oxidoreductases↗

Replacement of the Escherichia coli trp operon attenuation control codons alters operon expression.

To test features of the current model of transcription attenuation in amino acid biosynthetic operons, alterations were introduced into the trp operon leader region and expression of the mutated operons was examined in miaA and miaA+ Escherichia coli strains that lacked the trp repressor. The miaA mutation prevents modification of the adenosine residue immediately 3' of the anticodon of tRNAs that interact with codons beginning with uridine. The undermodified tRNA(Trp) in miaA strains is thought to increase readthrough at the trp attenuator by slowing ribosome movement over two tandem Trp codons in the 14-codon leader peptide coding region. The rate of translation of these two "control codons" is thought to be the key step in determining the extent of transcription attenuation in the trp leader region. Sequential deletion of trpL DNA specifying the leader peptide initiation region, RNA segment 1, RNA segment 2 and RNA segment 3 alternately decreased and increased trp operon expression, a result consistent with previous findings in another bacterium and the generally accepted model for transcription attenuation. Replacement of the tandem Trp control codons by AGG-UGC (Arg-Cys) codons eliminated the miaA-dependent increase in transcription readthrough. Replacement of the Trp control codons by AGG-UGA (Arg-stop) codons caused complete readthrough at the trp attenuator as well as abolishing the miaA effect. Presumably, the ribosome terminating translation at the new UGA codon mimics the effect of a stalled ribosome at the Trp control codons. This finding suggests that ribosome dissociation at some stop codons is slow relative to the time required for transcription of the trp leader region. Thus, most ribosomes translating the trp leader peptide coding region may remain attached to the natural UGA stop codon until after the attenuation decision is made. The interpretation supports models for trp operon attenuation in which the elevated basal level readthrough is determined by occasional ribosome release prior to synthesis of the 3:4 terminator hairpin.

Amino Acid Sequence↗

In vivo evidence for non-universal usage of the codon CUG in Candida maltosa.

An alkane-assimilating yeast Candida maltosa had been studied in order to establish systems suitable for biotransformation of hydrophobic compounds. However, functional expression of heterologous genes tested for this purpose had not been successful in several cases. On the other hand, it had been reported that the codon CUG, a universal leucine codon, is read as serine in C. cylindracea. The same altered codon usage had also been suggested by in vitro experiments in some Candida yeasts which are phylogenetically closely related to C. maltosa. In this study we have shown that the failure in functional expression of a heterologous gene is due to the fact that the codon CUG is read as serine in C. maltosa. This conclusion was drawn from the following experimental results: (1) when a cytochrome P450 gene of C. maltosa containing a CTG codon was expressed in C. maltosa, the corresponding amino acid was found to be serine, and not leucine; (2) a tRNA gene with an almost identical structure to that of the tRNASerCAG gene of C. albicans could be isolated from the genome of C. maltosa; (3) the Saccharomyces cerevisiae URA3 gene, which has one CTG codon, could not complement the ura3 mutation of C. maltosa as itself, but when the CTG codon was changed to another leucine codon, CTC, the mutated gene could complement the ura3 mutation. The last result is the first example of succeeding in functional expression of a heterologous gene in Candida species having an altered codon usage by changing the CTG codon in the gene to another codon.

Amino Acid Sequence↗

Compositional correlation studies among the three different codon positions in 12 bacterial genomes.

Compositional distributions in the three codon positions of the coding sequences of 12 fully sequenced prokaryotic genomes, which are publicly available, were investigated. A universal compositional correlation was observed in most of the genomes under investigation irrespective of their overall genomic GC contents. In all the genomes, the GC contents at the first codon positions are always greater than the overall GC contents of the genomes whereas the reverse is true in the case of second codon positions. GC contents at the third codon positions are higher than the overall genomic GC contents in high GC containing genomes, and the opposite situation was found in case of low GC genomes except for Helicobacter pylori. In high-GC rich genomes, the GC contents at the first + second codon positions are less than the GC contents at the third codon positions, and they are low in low-GC genomes except for Helicobacter pylori. The distributions of four bases at the three different positions were also investigated for all 12 organisms. It was observed that in high-GC genomes G is the most dominant base and in low-GC genomes A is the most dominant base in the first codon positions. But purine bases, i.e., (A + G), predominantly occur in the first codon position. In the second codon position, A is the most dominant base in most of the organisms and G is the least dominant base in all the organisms. There is no unique regular pattern of individual bases at the third codon positions; however, there are significant differences in the occurrences of (G + C) contents in the third codon positions among the different organisms. Calculations of dinucleotide frequencies in 12 different organisms indicate that in GC-rich genomes GG, GC, CC, and CG dinucleotides are the most dominant whereas the reverse is true in case of low-GC genomes. Biological implications of these results are discussed in this paper.

Bacteria↗

Clustering of low usage codons and ribosome movement.

A model is presented in which the distribution of low-usage codons in a message is a major factor in determining the impact that they will have on the translation rate and distribution of ribosomes on that message. This model is based on the assumption that low-usage codons are translated more slowly than normal codons, an assumption supported by various lines of published experimental evidence. Although the parameters used to develop this model are somewhat arbitrary, the main conclusions of this paper are consistent with a wide variation in the values of those parameters. In the model, low-usage codons arranged in clusters are much more effective in blocking ribosome movement on the message than ones that are dispersed. The effective size of the cluster is limited to the dimensions of the ribosome. It has been estimated that ribosomes on a message are spaced at least 27 nucleotides or nine codons apart. A ribosome translating a cluster of nine codons in which some or all of the codons are low-usage will move more slowly than over a comparable stretch of message containing no low-usage codons. Owing to ribosome size, the ribosome immediately behind the stalled ribosome will move as slowly; it must wait for the stalled ribosome to move on before it can even begin to translate the difficult region containing the low-usage codons. When the low-usage codon cluster is at the 3' end, the message will eventually be occupied by a ribosome jam that will transmit back to the 5' end of the message. In the steady state, the slowing effect imposed by a cluster of nine low-usage codons at the 3' end of a message would be just as great as if the entire message was composed of them. If the cluster is situated in the middle of a message, the ribosomes will form a jam upstream of the cluster. The ribosome density downstream of the cluster will be considerably reduced from what it would be for the same message with no cluster. If the cluster is at the 5' end of the message, the density of ribosomes will be reduced over the entire length of the message but the overall translation rate per ribosome will be only slightly reduced. However, owing to the reduced number of ribosomes initiating, the efficiency of the message in protein synthesis will be considerably reduced.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Comparative studies on codon usage pattern of chloroplasts and their host nuclear genes in four plant species.

A detailed comparison was made of codon usage of chloroplast genes with their host (nuclear) genes in the four angiosperm species Oryza sativa, Zea mays, Triticum aestivum and Arabidopsis thaliana. The average GC content of the entire genes, and at the three codon positions individually, was higher in nuclear than in chloroplast genes, suggesting different genomic organization and mutation pressures in nuclear and chloroplast genes. The results of Nc-plots and neutrality plots suggested that nucleotide compositional constraint had a large contribution to codon usage bias of nuclear genes in O. sativa, Z. mays, and T. aestivum, whereas natural selection was likely to be playing a large role in codon usage bias in chloroplast genomes. Correspondence analysis and chi-test showed that regardless of the genomic environment (species) of the host, the codon usage pattern of chloroplast genes differed from nuclear genes of their host species by their AU-richness. All the chloroplast genomes have predominantly A- and/or U-ending codons, whereas nuclear genomes have G-, C- or U-ending codons as their optimal codons. These findings suggest that the chloroplast genome might display particular characteristics of codon usage that are different from its host nuclear genome. However, one feature common to both chloroplast and nuclear genomes in this study was that pyrimidines were found more frequently than purines at the synonymous codon position of optimal codons.

Arabidopsis↗

Codon usage patterns in cytochrome oxidase I across multiple insect orders.

Synonymous codon usage bias is determined by a combination of mutational biases, selection at the level of translation, and genetic drift. In a study of mtDNA in insects, we analyzed patterns of codon usage across a phylogeny of 88 insect species spanning 12 orders. We employed a likelihood-based method for estimating levels of codon bias and determining major codon preference that removes the possible effects of genome nucleotide composition bias. Three questions are addressed: (1) How variable are codon bias levels across the phylogeny? (2) How variable are major codon preferences? and (3) Are there phylogenetic constraints on codon bias or preference? There is high variation in the level of codon bias values among the 88 taxa, but few readily apparent phylogenetic patterns. Bias level shifts within the lepidopteran genus Papilio are most likely a result of population size effects. Shifts in major codon preference occur across the tree in all of the amino acids in which there was bias of some level. The vast majority of changes involves double-preference models, however, and shifts between single preferred codons within orders occur only 11 times. These shifts among codons in double-preference models are phylogenetically conservative.

Animals↗

Nonrandom intragenic variations in patterns of codon bias implicate a sequential interplay between transitional genetic drift and functional amino acid selection.

Although most codon third bases appear to be functionless, the synonymous codons so defined exhibit a strikingly nonrandom distribution (codon bias) within human and other genes. To examine this phenomenon further, we generated a database of DNA sequences encoding human transmembrane cell-surface receptor proteins. Using this database we show here that the guanine and cytosine content of codon third bases (GC3) varies intragenically with the nature of the specified receptor domains (transmembrane > extracellular > intracellular domains; p < 0.001), the phenotype of the encoded amino acids (hydrophobic > hydrophilic > neutral amino acids; p < 0.001), and the receptor affiliation of the transmembrane (G-protein-coupled receptors > receptor tyrosine kinases; p < 0.001). Within gene regions specifying transmembrane domains, GC3 declines as domain functionality becomes redundant with increasing hydrophobicity (p < 0.001). Codons containing the second-base cytosine (XCZ, which encodes neutral amino acids) are selectively depleted of third-base adenine content (A3: XCA codons) when encoding transmembrane domain residues, consistent with positive selection for transitional mutation of XCG to XTG (which encodes hydrophobic amino acids) rather than to the synonymous XCA. Supporting this XCG --> XTG mechanism of codon bias, the G3:A3 ratio of codons specifying the transmembrane amino acid glycine (GGZ) is intermediate between that of its functional homolog alanine (GCZ) and that of hydrophobic valine (GTZ), even though the C3:T3 ratios are similar. Conversely, nearest-neighbor analysis of third bases 5' to codons specifying valine and leucine (CTZ) confirms a significant difference in C3:T3 but not G3:A3 ratios (i.e., C3/G1 --> T3/G1 > C3/A1; p < 0.001), consistent with the functionally advantageous retention of hydrophobic residues. These data raise the possibility that patterns of intragenic codon bias reflect a balance between negative and positive selection, suggesting in turn that analysis of codon third-base usage may help to predict the functional significance of encoded products.

Amino Acids↗

Synonymous codon usage, accuracy of translation, and gene length in Caenorhabditis elegans.

In many unicellular organisms, invertebrates, and plants, synonymous codon usage biases result from a coadaptation between codon usage and tRNAs abundance to optimize the efficiency of protein synthesis. However, it remains unclear whether natural selection acts at the level of the speed or the accuracy of mRNAs translation. Here we show that codon usage can improve the fidelity of protein synthesis in multicellular species. As predicted by the model of selection for translational accuracy, we find that the frequency of codons optimal for translation is significantly higher at codons encoding for conserved amino acids than at codons encoding for nonconserved amino acids in 548 genes compared between Caenorhabditis elegans and Homo sapiens. Although this model predicts that codon bias correlates positively with gene length, a negative correlation between codon bias and gene length has been observed in eukaryotes. This suggests that selection for fidelity of protein synthesis is not the main factor responsible for codon biases. The relationship between codon bias and gene length remains unexplained. Exploring the differences in gene expression process in eukaryotes and prokaryotes should provide new insights to understand this key question of codon usage.

Animals↗

Comparative analysis of expressed sequences reveals a conserved pattern of optimal codon usage in plants.

Codon usage bias is a ubiquitous phenomenon, which may be caused by mutational bias, selection, or both. The patterns of codon usage in plants are not well understood. Datasets of expressed sequence tags (ESTs) available for many plant species provide the resources for large-scale comparative analysis of codon usage patterns. We developed a computational approach to translate EST or assembled contig sequences, and then used the coding information for comparative analysis of codon usage in 12 plant species, including 6 eudicots, 5 monocots and the green alga Chlamydomonas reinhardtii. While codon nucleotide composition is highly conserved within eudicots or monocots, there is a significant difference between these two major taxonomic groups of higher plants. The third nucleotide position of codons is AU-rich in the eudicot genomes (35-42% of G+C content), but GC-rich in the monocot genomes (59-61% of G+C content). To identify optimal codons in these species, we used EST counts to estimate gene transcript levels. It was demonstrated that codon usage bias is correlated positively with gene transcript levels. Interestingly, the use of optimal codons appears to be well conserved between eudicots and monocots, and to a lesser degree between the higher plants and C. reinhardtii. Most of the optimal codons end with a C or G base, regardless of the different nucleotide composition in these genomes. The results suggest that plant codon usage is affected by translational selection, and the selective pressure appears to be conserved in the plant kingdom.

Animals↗

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Codon usage in Kluyveromyces lactis and in yeast cytochrome c-encoding genes.

Codon usage (CU) in Kluyveromyces lactis has been studied. Comparison of CU in highly and lowly expressed genes reveals the existence of 21 optimal codons; 18 of them are also optimal in other yeasts like Saccharomyces cerevisiae or Candida albicans. Codon bias index (CBI) values have been recalculated with reference to the assignment of optimal codons in K. lactis and compared to those previously reported in the literature taking as reference the optimal codons from S. cerevisiae. A new index, the intrinsic codon deviation index (ICDI), is proposed to estimate codon bias of genes from species in which optimal codons are not known; its correlation with other index values, like CBI or effective number of codons (Nc), is high. A comparative analysis of CU in six cytochrome-c-encoding genes (CYC) from five yeasts is also presented and the differences found in the codon bias of these genes are discussed in relation to the metabolic type to which the corresponding yeasts belong. Codon bias in the CYC from K. lactis and S. cerevisiae is correlated to mRNA levels.

Amino Acids↗

Lactic acid bacteria as prime candidates for codon optimization.

In species having a strong correlation of expressivity and codon bias it has been shown that heterologous expression can be optimized by changing codons of the introduced gene towards the set of codons that the host organism naturally uses in its highly expressed genes. Even though two lactic acid bacteria are fully sequenced, there are no reports on attempts of codon optimization in the literature. In this report it is demonstrated that codons used in highly expressed genes tend to differ from the codons in lowly expressed genes, and that there is a strong correlation of codon bias and empirical expressivity (codon adaptation index) in Lactococcus lactis and Lactobacillus plantarum. This strongly suggests that codon optimization strategies could be applied to expression systems with lactic acid bacteria as producer strains. A good example of a candidate for codon optimization is the mouse interleukin-2 gene, which in its natural form has an extremely low codon adaptation index for expression in Lc. lactis.

Algorithms↗

Correlation of codon bias measures with mRNA levels: analysis of transcriptome data from Escherichia coli.

Although codon usage is often represented by a 61-dimensional vector, the ability of determining the codon bias in a gene relies on a uni-dimensional vector which measures the total bias in usage of synonymous codons. Codon usage is receiving more and more focus because codon biases might be valuable tools to predict and optimize gene/protein expression. How good any of these measures is for correlating codon usage with gene and protein expression has yet to be investigated. In this study, we correlated gene transcript levels in Escherichia coli with codon usage, using a number of different codon bias measures. We found that there is a significant correlation between transcript levels and codon bias measures, suggesting that these measures can be used to assess or predict gene expression. The codon bias measure performing best in this context was the codon adaptation index.

Codon↗