PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Codon usage patterns in Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster and Homo sapiens; a review of the considerable within-species diversity.

The genetic code is degenerate, but alternative synonymous codons are generally not used with equal frequency. Since the pioneering work of Grantham's group it has been apparent that genes from one species often share similarities in codon frequency; under the "genome hypothesis" there is a species-specific pattern to codon usage. However, it has become clear that in most species there are also considerable differences among genes. Multivariate analyses have revealed that in each species so far examined there is a single major trend in codon usage among genes, usually from highly biased to more nearly even usage of synonymous codons. Thus, to represent the codon usage pattern of an organism it is not sufficient to sum over all genes as this conceals the underlying heterogeneity. Rather, it is necessary to describe the trend among genes seen in that species. We illustrate these trends for six species where codon usage has been examined in detail, by presenting the pooled codon usage for the 10% of genes at either end of the major trend. Closely-related organisms have similar patterns of codon usage, and so the six species in Table 1 are representative of wider groups. For example, with respect to codon usage, Salmonella typhimurium closely resembles E. coli, while all mammalian species so far examined (principally mouse, rat and cow) largely resemble humans.

Amino Acids↗

Role of GC-biased mutation pressure on synonymous codon choice in Micrococcus luteus, a bacterium with a high genomic GC-content.

The GC (G + C, or G or C)-contents of codon silent positions in all two-codon sets and three codons AUY/A (IIe), and in most of the family boxes of Micrococcus luteus (genomic GC-content: 74%) are 95% to 100% in both the highly and weakly expressed genes. In some family boxes, there is a decrease in NNC codons and an increase in NNG codons from the highly expressed to weakly expressed genes without apparent involvement of NNU and NNA codons. From these observations, we conclude that the selective use of synonymous codons in M. luteus may be largely determined by GC-biased mutation pressure and that in the highly expressed genes tRNAs would act as a weak selection pressure in some family boxes. Available data suggest that the effect of selection pressure by tRNAs on the synonymous codon choice becomes more apparent in the highly expressed genes in eubacteria with intermediate GC-contents such as Escherichia coli and Bacillus subtilis, and that the U/C ratio of the codon third positions in NNU/C-type two-codon sets in the weakly expressed genes would represent the approximate magnitude of directional mutation pressure throughout eubacteria.

Base Sequence↗

Translation of the F protein of hepatitis C virus is initiated at a non-AUG codon in a +1 reading frame relative to the polyprotein.

The hepatitis C virus (HCV) genome contains an internal ribosome entry site (IRES) followed by a large open reading frame coding for a polyprotein that is cleaved into 10 proteins. An additional HCV protein, the F protein, was recently suggested to result from a +1 frameshift by a minority of ribosomes that initiated translation at the HCV AUG initiator codon of the polyprotein. In the present study, we reassessed the mechanism accounting for the synthesis of the F protein by measuring the expression in cultured cells of a luciferase reporter gene with an insertion encompassing the IRES plus the beginning of the HCV-coding region preceding the luciferase-coding sequence. The insertion was such that luciferase expression was either in the +1 reading frame relative to the HCV AUG initiator codon, mimicking the expression of the F protein, or in-frame with this AUG, mimicking the expression of the polyprotein. Introduction of a stop codon at various positions in-frame with the AUG initiator codon and substitution of this AUG with UAC inhibited luciferase expression in the 0 reading frame but not in the +1 reading frame, ruling out that the synthesis of the F protein results from a +1 frameshift. Introduction of a stop codon at various positions in the +1 reading frame identified the codon overlapping codon 26 of the polyprotein in the +1 reading frame as the translation start site for the F protein. This codon 26(+1) is either GUG or GCG in the viral variants. Expression of the F protein strongly increased when codon 26(+1) was replaced with AUG, or when its context was mutated into an optimal Kozak context, but was severely decreased in the presence of low concentrations of edeine. These observations are consistent with a Met-tRNA(i)-dependent initiation of translation at a non-AUG codon for the synthesis of the F protein.

Base Sequence↗

Synonymous codon usage variation among Giardia lamblia genes and isolates.

The pattern of codon usage in the amitochondriate diplomonad Giardia lamblia has been investigated. Very extensive heterogeneity was evident among a sample of 65 genes. A discrete group of genes featured unusual codon usage due to the amino acid composition of their products: these variant surface proteins (VSPs) are unusually rich in Cys and, to a lesser extent, Gly and Thr. Among the remaining 50 genes, correspondence analysis revealed a single major source of variation in synonymous codon usage. This trend was related to the extent of use of a particular subset of 21 codons which are inferred to be those which are optimal for translation; at one end of this trend were genes expected to be expressed at low levels with near random codon usage, while at the other extreme were genes expressed at high levels in which these optimal codons are used almost exclusively. These optimal codons all end in C or G so G + C content at silent sites varies enormously among genes, from values around 40%, expected to reflect the background level of the genome, up to nearly 100%. Although VSP genes are occasionally extremely highly expressed, they do not, in general, have high frequencies of optimal codons, presumably because their high expression is only intermittent. These results indicate that natural selection has been very effective in shaping codon usage in G. lamblia. These analyses focused on sequences from strains placed within G. lamblia "assemblage A"; a few sequences from other strains revealed extensive divergence at silent sites, including some divergence in the pattern of codon usage.

Amino Acids↗

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗

Genetic robustness and selection at the protein level for synonymous codons.

Synonymous codons are neutral at the protein level, therefore natural selection at the protein level should have no effect on their frequencies. Synonymous codons, however, differ in their capacity to reduce the effects of errors: after mutation, certain codons keep on coding for the same amino acid or for amino acids with similar properties, while other synonymous codons produce very different amino acids. Therefore, the impact of errors on a coding sequence (genetic robustness) can be measured by analysing its codon usage. I analyse the codon usage of sequenced nuclear and cytoplasmic genomes and I show that there is an extensive variation in genetic robustness at the DNA sequence level, both among genomes and among genes of the same genome. I also show theoretically that robustness can be adaptive, that is natural selection may lead to a preference for codons that reduce the impact of errors. If selection occurs only among the mutants of a codon (e.g. among the progeny before the adult phase), however, the codons that are more sensitive to the effects of mutations may increase in frequency because they manage to get rid more easily of deleterious mutations. I also suggest other possible explanations for the evolution of genetic robustness at the codon level.

Amino Acid Sequence↗

Distinct paths to stop codon reassignment by the variant-code organisms Tetrahymena and Euplotes.

The reassignment of stop codons is common among many ciliate species. For example, Tetrahymena species recognize only UGA as a stop codon, while Euplotes species recognize only UAA and UAG as stop codons. Recent studies have shown that domain 1 of the translation termination factor eRF1 mediates stop codon recognition. While it is commonly assumed that changes in domain 1 of ciliate eRF1s are responsible for altered stop codon recognition, this has never been demonstrated in vivo. To carry out such an analysis, we made hybrid proteins that contained eRF1 domain 1 from either Tetrahymena thermophila or Euplotes octocarinatus fused to eRF1 domains 2 and 3 from Saccharomyces cerevisiae. We found that the Tetrahymena hybrid eRF1 efficiently terminated at all three stop codons when expressed in yeast cells, indicating that domain 1 is not the sole determinant of stop codon recognition in Tetrahymena species. In contrast, the Euplotes hybrid facilitated efficient translation termination at UAA and UAG codons but not at the UGA codon. Together, these results indicate that while domain 1 facilitates stop codon recognition, other factors can influence this process. Our findings also indicate that these two ciliate species used distinct approaches to diverge from the universal genetic code.

Animals↗

Protein tagging at rare codons is caused by tmRNA action at the 3' end of nonstop mRNA generated in response to ribosome stalling.

It has been believed that protein tagging caused by consecutive rare codons involves tmRNA action at the internal mRNA site. We demonstrated previously that ribosome stalling either at sense or stop codons caused by certain arrest sequences could induce mRNA cleavage near the arrest site, resulting in nonstop mRNAs that are recognized by tmRNA. These findings prompted us to re-examine the mechanism of tmRNA tagging at a run of rare codons. We report here that either AGG or CGA but not AGA arginine rare-codon clusters inserted into a model crp mRNA encoding cAMP receptor protein (CRP) could cause an efficient protein tagging. We demonstrate that more than three consecutive AGG codons are needed to induce an efficient ribosome stalling therefore tmRNA tagging in our system. The tmRNA tagging was eliminated by overproduction of tRNAs corresponding to rare codons, indicating that a scarcity of the corresponding tRNA caused by the rare-codon cluster is an important factor for tmRNA tagging. Mass spectrometry analyses of proteins generated in cells lacking or possessing tmRNA encoding a protease-resistant tag sequence indicated that the truncation and tmRNA tagging occur within the cluster of rare codons. Northern and S1 analyses demonstrated that nonstop mRNAs truncated within the rare-codon clusters are detected in cells lacking tmRNA but not in cells expressing tmRNA. We conclude that a ribosome stalled by the rare codon induces mRNA cleavage, resulting in nonstop mRNAs that are recognized by tmRNA.

Amino Acid Sequence↗

Four-base codon-mediated incorporation of non-natural amino acids into proteins in a eukaryotic cell-free translation system.

Various four-base codons have been shown to work for the introduction of non-natural amino acids into proteins in an Escherichia coli cell-free translation system. Here, a four-base codon-mediated non-natural mutagenesis was applied to a eukaryotic rabbit reticulocyte cell-free translation system. Mutated streptavidin mRNAs containing four-base codons were prepared and added to a rabbit reticulocyte lysate in the presence of tRNAs that were aminoacylated with a non-natural amino acid and had the corresponding four-base anticodons. A Western blot analysis of translation products indicated that the four-base codons CGGU, CGCU, CCCU, CUCU, CUAU, and GGGU were efficiently decoded by the aminoacyl-tRNAs having the corresponding four-base anticodons. In contrast, the four-base codons AGGU, AGAU, CGAU, UUGU, UCGU, and ACGU were not decoded. The stop codon-derived four-base codons UAGU, UAAU, and UGAU were found to be inefficient, whereas the amber codon UAG and opal codon UGA were efficient for the incorporation of non-natural amino acids. The application of the expanded genetic code in a eukaryotic cell-free system opens the possibility of a four-base codon-mediated incorporation of non-natural amino acids into proteins in living eukaryotic cells.

Amino Acids↗

Intragenic spatial patterns of codon usage bias in prokaryotic and eukaryotic genomes.

To study the roles of translational accuracy, translational efficiency, and the Hill-Robertson effect in codon usage bias, we studied the intragenic spatial distribution of synonymous codon usage bias in four prokaryotic (Escherichia coli, Bacillus subtilis, Sulfolobus tokodaii, and Thermotoga maritima) and two eukaryotic (Saccharomyces cerevisiae and Drosophila melanogaster) genomes. We generated supersequences at each codon position across genes in a genome and computed the overall bias at each codon position. By quantitatively evaluating the trend of spatial patterns using isotonic regression, we show that in yeast and prokaryotic genomes, codon usage bias increases along translational direction, which is consistent with purifying selection against nonsense errors. Fruit fly genes show a nearly symmetric M-shaped spatial pattern of codon usage bias, with less bias in the middle and both ends. The low codon usage bias in the middle region is best explained by interference (the Hill-Robertson effect) between selections at different codon positions. In both yeast and fruit fly, spatial patterns of codon usage bias are characteristically different from patterns of GC-content variations. Effect of expression level on the strength of codon usage bias is more conspicuous than its effect on the shape of the spatial distribution.

Animals↗

Factors affecting codon usage in Yersinia pestis.

The complete genome of Yersinia pestis which was the causative agent of the systemic invasive infectious disease classically referred as plague, had been recently sequenced. In order to have a further insight into the synonymous codon usage evolution, factors shaping synonymous codon usage pattern of Yersinia pestis were analyzed in this paper. The coding sequences larger than or equal to 300 bp were used in codon usage analysis. Though "G"+"C" content in Y. pestis genome was slightly lower (47.64%), the highly expressed genes tended to use "C" or "G" at synonymous sites compared with lowly expressed genes. Conversely, lowly expressed genes tended to prefer "A" or "T" at synonymous positions. Gene expression level was strongly correlated with the first axis of the correspondence analysis (COA) (R=0.63, P<0.0001). By the analyses of the codon usage pattern of highly and lowly expressed genes, it was confirmed that gene expression level was partially responsible for the codon usage bias. GC-skew analysis showed that codon usage suffered replication-transcriptional selection. Codon adaptation index (CAI), frequency of "C"+"G" at the synonymous third position of codon (GC3s) and the effective number of codons (Nc) values showed some differences among different gene length groups. "G"+"C" content of genes was strongly correlated with the first axis of the COA (R=0.72, P<0.0001). It could be concluded that gene expressivity, replication-transcriptional selection, gene length and gene composition constraints were the main affecting factors of codon usage variation in Y. pestis.

Amino Acids↗

K-ras codon 12 and 13 mutations are correlated with differential patterns of tumor cell dissemination in colorectal cancer patients.

The aim of this prospective study was to relate the incidences of cytokeratin 20 (CK20) and guanylylcyclase C (GCC) in lymph node, liver, and bone marrow specimens of 245 colorectal cancer (CRC) patients with the K-ras oncogene status of the corresponding primary tumor. Qualitative RT-PCR detection of CK20 and GCC mRNA was used as marker of circulating epithelial cells (CEC). Samples were considered positive for CEC only when both markers were detected concomitantly. For the detection of K-ras mutations, a PCR-RFLP assay was used. In the group with K-ras mutated primary carcinomas (n=92), CEC were detected in 62% of lymph node-, 43% of liver-, and 2% of bone marrow samples. No statistical significance was found when comparing these results with those from patients with K-ras wild-type carcinoma (59%, 46%, and 0%, respectively). In contrast to this combined evaluation, separate analysis of K-ras codons 12 (n=75, 82%) and 13 (n=17, 18%) revealed significantly differing CEC incidences. Lymph node specimens from corresponding K-ras codon 13 mutated carcinomas showed a significantly higher CEC incidence (82%) than the groups with codon 12 mutation (57%, p<0.05) or K-ras wild-type sequence (59%, p<0.05). Unlike these findings in lymph nodes, liver biopsies from corresponding carcinomas with K-ras codon 12 mutation or wild-type sequence were significantly more often positive for CEC (31% and 29%) than specimens from K-ras codon 13 mutated primary CRC (12%, p<0.04, respectively). In conclusion, colorectal carcinomas with K-ras codon 12 mutation showed the same pattern of tumor cell dissemination as their K-ras wild-type counterparts. Since K-ras codon 12 mutations prevailed 4-fold over codon 13 mutations, combined analysis of the two codons showed the same result. However, sub-analysis of patients with K-ras codon 13 mutation revealed that the respective CEC incidence was significantly increased in lymph nodes, but decreased in liver biopsies.

Aged↗

Codon usage of human DNA viruses and its similarity to certain host genes.

Codon usages of DNA viruses had previously been shown to associate with their genome size. Codon usage of various human DNA viruses was compared to those of human genes to further understand viral codon usage and its roles in viral-host interaction. Codon usage bias in both large and small genome human DNA viruses was dominantly driven by translation selection. Non-optimal codon usage in small DNA viruses showed similarity to cell cycle-related genes, whereas codon usage of large DNA viruses was more diverse, herpesviruses showed more heterogeneity than human adenoviruses, while poxviruses showed a clear bimodal pattern. Some of the large DNA viruses such as herpes simplex and molluscum contagiosum viruses showed more optimal codon usage. Enrichment analysis identified some groups of human genes with similar codon usage to each group of these viruses. These host genes with similarity in codon usages to those of viruses may be efficiently expressed in infected cells and involved in their life cycle, pathogenesis and/or immune evasion.

Humans↗

Codon-specific and general inhibition of protein synthesis by the tRNA-sequestering minigenes.

The expression of minigenes in bacteria inhibits protein synthesis and cell growth. Presumably, the translating ribosomes, harboring the peptides as peptidyl-tRNAs, pause at the last sense codon of the minigene directed mRNAs. Eventually, the peptidyl-tRNAs drop off and, under limiting activity of peptidyl-tRNA hydrolase, accumulate in the cells reducing the concentration of specific aminoacylable tRNA. Therefore, the extent of inhibition is associated with the rate of starvation for a specific tRNA. Here, we used minigenes harboring various last sense codons that sequester specific tRNAs with different efficiency, to inhibit the translation of reporter genes containing, or not, these codons. A prompt inhibition of the protein synthesis directed by genes containing the codons starved for their cognate tRNA (hungry codons) was observed. However, a non-specific in vitro inhibition of protein synthesis, irrespective of the codon composition of the gene, was also evident. The degree of inhibition correlated directly with the number of hungry codons in the gene. Furthermore, a tRNA(Arg4)-sequestering minigene promoted the production of an incomplete beta-galactosidase polypeptide interrupted, during bacterial polypeptide chain elongation at sites where AGA codons were inserted in the lacZ gene suggesting ribosome pausing at the hungry codons.

Base Sequence↗

Site-specific codon bias in bacteria.

Sequences of the gapA and ompA genes from 10 genera of enterobacteria have been analyzed. There is strong bias in codon usage, but different synonymous codons are preferred at different sites in the same gene. Site-specific preference for unfavored codons is not confined to the first 100 codons and is usually manifest between two codons utilizing the same tRNA. Statistical analyses, based on conclusions reached in an accompanying paper, show that the use of an unfavored codon at a given site in different genera is not due to common descent and must therefore be caused either by sequence-specific mutation or sequence-specific selection. Reasons are given for thinking that sequence-specific mutation cannot be responsible. We are unable to explain the preference between synonymous codons ending in C or T, but synonymous choice between A and G at third sites is largely explained by avoidance of AG-G (where the hyphen indicates the boundary between codons). We also observed that the preferred codon for proline in Enterobacter cloacea has changed from CCG to CCA.

Bacterial Outer Membrane Proteins↗

Rapid determination of COL2A1 mutations in individuals with Stickler syndrome: analysis of potential premature termination codons.

Stickler syndrome is one of the milder phenotypes resulting from mutations in the gene that encodes type-II collagen, COL2A1. All COL2A1 mutations known to cause Stickler syndrome result in the formation of a premature termination codon within the type-II collagen gene. COL2A1 has 10 in-frame CGA codons, which can mutate to TGA STOP codons via a methylation-deamination mechanism. We have analyzed these sites in genomic DNA from a panel of 40 Stickler syndrome patients to test the hypothesis that mutations that cause Stickler syndrome preferentially occur at these bases. Polymerase chain reaction (PCR) amplification of genomic DNA containing each of the in-frame CGA codons was done by one of two methods: either using primers that amplify DNA that includes the CGA codon, or using allele-specific primers that either amplify normal sequence containing a CGA codon or amplify a mutant sequence containing a TGA codon. Analysis of PCR products by restriction endonuclease digestion or sequencing demonstrated the presence of a normal or mutated codon. TGA mutations were identified in eight patients, at five of the 10 in-frame CGA codons. The identification of these mutations in eight of 40 patients demonstrates that these sites are common sites for mutations in individuals with Stickler syndrome and, we propose, should be analyzed as a first step in the search for mutations that result in this disorder.

Alleles↗

Stop codon decoding in Candida albicans: from non-standard back to standard.

The human pathogen Candida albicans translates the standard leucine-CUG codon as serine. This genetic code change is mediated by a novel ser-tRNA(CAG), which induces aberrant mRNA decoding in vitro, resulting in retardation of the electrophoretic mobility of the polypeptides synthesized in its presence. These non-standard decoding events have been attributed to readthrough of the UAG and UGA stop codons encoded by the Brome Mosaic Virus RNA 4, which codes for the virion coat protein, and the rabbit globin mRNAs, respectively. In order to fully elucidate the behaviour of the C. albicans ser-tRNA(CAG) towards stop codons, we have used other cell-free translation systems and reporter genes. However, the reporter systems used encode several CUG codons, making it impossible to distinguish whether the slow migration of the polypeptides is caused by the replacement of leucines by serines at the CUG codons, readthrough, or a combination of both. Therefore, we have constructed new reporter systems lacking CUG codons and have used them to demonstrate that aberrant mRNA decoding in vitro is not a result from stop codon readthrough or any other non-standard translational event. Our data show that a single leucine to serine replacement at only one of the four CUG codons encoded by the BMV RNA-4 gene is responsible for the aberrant migration of the BMV coat protein on SDS-PAGE, suggesting that this amino acid substitution (ser for leu) significantly alters the structure of the virion coat protein. The data therefore show that the only aberrant event mediated by the ser-tRNA(CAG) is decoding of the leu-CUG codon as serine.

Amino Acid Substitution↗

Consecutive low-usage leucine codons block translation only when near the 5' end of a message in Escherichia coli.

Insertion of nine consecutive low-usage CUA leucine codons after codon 13 of a 313-codon test mRNA strongly inhibited its translation without apparent effect on translation of other mRNAs containing CUA codons. In contrast, nine consecutive high-usage CUG leucine codons at the same position had no apparent effect, and neither low- nor high-usage codons affected translation when inserted after codon 223 or 307. Additional experiments indicated that the strong positional effect of the low-usage codons could not be accounted for by differences in stability of the mRNAs or in stringency of selection of the correct tRNA. The positional effect could be explained if translation complexes are less stable near the beginning of a message: slow translation through low-usage codons early in the message may allow most translation complexes to dissociate before they read through.

Blotting, Northern↗