PubMed HealthSearch

SEARCH · PubMed Health

Results for “Codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

CGG: an unassigned or nonsense codon in Mycoplasma capricolum.

CGG is an arginine codon in the universal genetic code. We previously reported that in Mycoplasma capricolum, a relative of Gram-positive eubacteria, codon CGG did not appear in coding frames, including termination sites, and tRNA(ArgCCG) pairing with codon CGG, was not detected. These facts suggest that CGG is a nonsense (unassigned and untranslatable) codon--i.e., not assigned to arginine or to any other amino acid. We have investigated whether CGG is really an unassigned codon by using a cell-free translation system prepared from M. capricolum. Translation of synthetic mRNA containing in-frame CGG codons does not result in "read-through" to codons beyond the CGG codons--i.e., translation ceases just before CGG. Sucrose-gradient centrifugation profiles of the reaction mixture have shown that the bulk of peptide that has been synthesized is attached to 70S ribosomes and is released upon further incubation with puromycin. The result suggests that the peptide is in the P site of ribosome in the form of peptidyl-tRNA, leaving the A site empty. When in-frame CGG codons are replaced by UAA codons in mRNA, no read-through occurs beyond UAA, just as in the case of CGG. However, the synthesized peptide is released from 70S ribosomes, presumably by release factor 1. These data suggest strongly that CGG is an unassigned codon and differs from UAA in that CGG is not used for termination.

Amino Acid Sequence

Role of codon choice in the leader region of the ilvGMEDA operon of Serratia marcescens.

Leucine participates in multivalent repression of the Serratia marcescens ilvGMEDA operon by attenuation (J.-H. Hsu, E. Harms, and H.E. Umbarger, J. Bacteriol. 164:217-222, 1985), although there is only one single leucine codon that could be involved in this type of control. This leucine codon is the rarely used CUA. The contribution of this leucine codon to the control of transcription by attenuation was examined by replacing it with the commonly used leucine codon CUG and with a nonregulatory proline codon, CCG. These changes left intact the proposed secondary structure of the leader. The effects of the codon changes were assessed by placing the mutant leader regions upstream of the ilvGME structural genes or the cat gene and measuring acetohydroxy acid synthase II, transaminase B, or chloramphenicol acetyltransferase activities in cells grown under limiting and repressing conditions. The presence of the common leucine codon in place of the rare leucine codon reduced derepression by about 70%. Eliminating the leucine codon by converting it to proline abolished leucine control. Furthermore, a possible context effect of the adjacent upstream serine codon on leucine control was examined by changing it into a glycine codon.

Base Sequence

Context effects and inefficient initiation at non-AUG codons in eucaryotic cell-free translation systems.

The context requirements for recognition of an initiator codon were evaluated in vitro by monitoring the relative use of two AUG codons that were strategically positioned to produce long (pre-chloramphenicol acetyl transferase [CAT]) and short versions of CAT protein. The yield of pre-CAT initiated from the 5'-proximal AUG codon increased, and synthesis of CAT from the second AUG codon decreased, as sequences flanking the first AUG codon increasingly resembled the eucaryotic consensus sequence. Thus, under prescribed conditions, the fidelity of initiation in extracts from animal as well as plant cells closely mimics what has been observed in vivo. Unexpectedly, recognition of an AUG codon in a suboptimal context was higher when the adjacent downstream sequence was capable of assuming a hairpin structure than when the downstream region was unstructured. This finding adds a new, positive dimension to regulation by mRNA secondary structure, which has been recognized previously as a negative regulator of initiation. Translation of pre-CAT from an AUG codon in a weak context was not preferentially inhibited under conditions of mRNA competition. That result is consistent with the scanning model, which predicts that recognition of the AUG codon is a late event that occurs after the competition-sensitive binding of a 40S ribosome-factor complex to the 5' end of mRNA. Initiation at non-AUG codons was evaluated in vitro and in vivo by introducing appropriate mutations in the CAT and preproinsulin genes. GUG was the most efficient of the six alternative initiator codons tested, but GUG in the optimal context for initiation functioned only 3 to 5% as efficiently as AUG. Initiation at non-AUG codons was artifactually enhanced in vitro at supraoptimal concentrations of magnesium.

Animals

Evolution of the mitochondrial genetic code. I. Origin of AGR serine and stop codons in metazoan mitochondria.

AGA and AGG (AGR) are arginine codons in the universal genetic code. These codons are read as serine or are used as stop codons in metazoan mitochondria. The arginine residues coded by AGR in yeast or Trypanosoma are coded by arginine CGN throughout metazoan mitochondria. AGR serine sites in metazoan mitochondria are occupied mainly in corresponding sites in yeast or Trypanosoma mitochondria by UCN serine, AGY serine, or codons for amino acids other than serine or arginine. Based on these observations, we propose the following evolutionary events. AGR codons became unassigned because of deletion of tRNA Arg (UCU) and elimination of AGR codons by conversion to CGN arginine codons. Upon acquisition by serine tRNA of pairing ability with AGR codons, some codons for amino acids other than arginine mutated to AGR, and were captured by anticodon GCU in serine tRNA. During vertebrate mitochondrial evolution, AGR stop codons presumably were created from UAG stop by deletion of the first nucleotide U and by use of R as the third nucleotide that had existed next to the ancestral UAG stop.

Animals

Analysis of the stop codon context in plant nuclear genes.

A region of 18 nucleotides surrounding the stop codon (the stop codon context) in 748 plant nuclear genes was analyzed. Non-randomness was found both upstream and downstream from the stop codon, suggesting that these sequences may help in ensuring efficient termination of translation. The UAG amber codon is the least-used stop codon and the bias in the nucleotide distribution 5' and 3' to the stop codon was more pronounced for the amber codon than for the other stop codons. This might indicate that the codon context affects termination more at UAG than at UGA or UAA stop codons.

Base Composition

A study of the purine/pyrimidine codon occurrence with a reduced centered variable and an evaluation compared to the frequency statistic.

With the three-letter alphabet [R,Y,N] (R = purine, Y = pyrimidine, N = R or Y), there are 26 codons (NNN being excluded): RNN,...,NNY (six codons at two unspecified bases N), RRN,...,NYY (12 codons at one unspecified base N), RRR,...,YYY (eight specified codons). A statistical methodology that uses the codon frequency and a reduced centered variable leads to similar results for a codon occurrence study, regardless of gene function and regardless of a particular protein coding gene taxonomic population. Therefore, this variable can be considered a new codon usage index, whose use removes certain nonsignificant results found with the frequency statistic. This methodology identifies the common and rare codons (i.e., the codons having the highest and lowest occurrence) and leads to a model of codon evolution at three successive states: RNN, then RNY, and finally RYY. Some biological relations between this model and the YRY(N)6YRY preferential occurrence are also presented.

Base Sequence

Translational efficiency of the Escherichia coli adenylate cyclase gene: mutating the UUG initiation codon to GUG or AUG results in increased gene expression.

Roy et al. [Roy, A., Haziza, C. & Danchin, A. (1983) EMBO J. 2, 791-797] established that translation of Escherichia coli adenylate cyclase initiates at a UUG codon, and they suggested this might decrease the efficiency of translation. We investigated the effect of varying the initiation codon on the expression of the adenylate cyclase (cya) gene. Using oligonucleotide-directed mutagenesis, we changed the UUG initiation codon to GUG and the more common initiator AUG and assayed for cya gene expression in a number of ways. First, the GUG initiation codon, in place of UUG, doubled cya expression when cya was expressed from the dual cya P1/P2 promoters. The corresponding AUG codon construct was nonviable. Second, when the cya gene was placed under the transcriptional control of the thermoinducible phage lambda PL promoter, the relative amounts of cya gene product were 1:2:6 for the UUG, GUG, and AUG initiation codons, respectively. Finally, the cya P2 promoter, Shine-Dalgarno sequence, and the DNA corresponding to the first 86 codons of cya were fused to DNA encoding the E. coli galactokinase gene beginning at the second codon. The relative amounts of the fusion polypeptides, which had galactokinase activity, were 1:2:3 for the UUG, GUG, and AUG initiation codons, respectively. These results demonstrate that the cya UUG initiation codon limits cya expression at the level of translation.

Adenylyl Cyclases

Codon usage patterns in Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster and Homo sapiens; a review of the considerable within-species diversity.

The genetic code is degenerate, but alternative synonymous codons are generally not used with equal frequency. Since the pioneering work of Grantham's group it has been apparent that genes from one species often share similarities in codon frequency; under the "genome hypothesis" there is a species-specific pattern to codon usage. However, it has become clear that in most species there are also considerable differences among genes. Multivariate analyses have revealed that in each species so far examined there is a single major trend in codon usage among genes, usually from highly biased to more nearly even usage of synonymous codons. Thus, to represent the codon usage pattern of an organism it is not sufficient to sum over all genes as this conceals the underlying heterogeneity. Rather, it is necessary to describe the trend among genes seen in that species. We illustrate these trends for six species where codon usage has been examined in detail, by presenting the pooled codon usage for the 10% of genes at either end of the major trend. Closely-related organisms have similar patterns of codon usage, and so the six species in Table 1 are representative of wider groups. For example, with respect to codon usage, Salmonella typhimurium closely resembles E. coli, while all mammalian species so far examined (principally mouse, rat and cow) largely resemble humans.

Amino Acids

Role of GC-biased mutation pressure on synonymous codon choice in Micrococcus luteus, a bacterium with a high genomic GC-content.

The GC (G + C, or G or C)-contents of codon silent positions in all two-codon sets and three codons AUY/A (IIe), and in most of the family boxes of Micrococcus luteus (genomic GC-content: 74%) are 95% to 100% in both the highly and weakly expressed genes. In some family boxes, there is a decrease in NNC codons and an increase in NNG codons from the highly expressed to weakly expressed genes without apparent involvement of NNU and NNA codons. From these observations, we conclude that the selective use of synonymous codons in M. luteus may be largely determined by GC-biased mutation pressure and that in the highly expressed genes tRNAs would act as a weak selection pressure in some family boxes. Available data suggest that the effect of selection pressure by tRNAs on the synonymous codon choice becomes more apparent in the highly expressed genes in eubacteria with intermediate GC-contents such as Escherichia coli and Bacillus subtilis, and that the U/C ratio of the codon third positions in NNU/C-type two-codon sets in the weakly expressed genes would represent the approximate magnitude of directional mutation pressure throughout eubacteria.

Base Sequence

Codon usage of human DNA viruses and its similarity to certain host genes.

Codon usages of DNA viruses had previously been shown to associate with their genome size. Codon usage of various human DNA viruses was compared to those of human genes to further understand viral codon usage and its roles in viral-host interaction. Codon usage bias in both large and small genome human DNA viruses was dominantly driven by translation selection. Non-optimal codon usage in small DNA viruses showed similarity to cell cycle-related genes, whereas codon usage of large DNA viruses was more diverse, herpesviruses showed more heterogeneity than human adenoviruses, while poxviruses showed a clear bimodal pattern. Some of the large DNA viruses such as herpes simplex and molluscum contagiosum viruses showed more optimal codon usage. Enrichment analysis identified some groups of human genes with similar codon usage to each group of these viruses. These host genes with similarity in codon usages to those of viruses may be efficiently expressed in infected cells and involved in their life cycle, pathogenesis and/or immune evasion.

Humans

Essential factors determining codon usage in ubiquitin genes.

Ubiquitin is ubiquitous in all eukaryotes and its amino acid sequence shows extreme conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences coding for 13 ubiquitin genes from 11 species reported so far have been compiled and analyzed. The G + C content of codon third base reveals a positive linear correlation with the genome G + C content of the corresponding species. The slope strongly suggests that the overall G + C content of codons of polyubiquitin genes clearly reflects the genome G + C content by AT/GC substitutions at the codon third position. The G + C content of ubiquitin codon third base also shows a positive linear correlation with the overall G + C content of coding regions of compiled genes, indicating the codon choices among synonymous codons reflect the average codon usage pattern of corresponding species. On the other hand, the monoubiquitin gene, which is different from the polyubiquitin gene in gene organization, gene expression, and function of the encoding protein, shows a different codon usage pattern compared with that of the polyubiquitin gene. From comparisons of the levels of synonymous substitutions among ubiquitin repeats and the homology of the amino acid sequence of the tail of monomeric ubiquitin genes, we propose that the molecular evolution of ubiquitin genes occurred as follows: Plural primitive ubiquitin sequences were dispersed on genome in ancestral eukaryotes. Some of them situated in a particular environment fused with the tail sequence to produce monomeric ubiquitin genes that were maintained across species. After divergence of species, polyubiquitin genes were formed by duplication of the other primitive ubiquitin sequences on different chromosomes. Differences in the environments in which ubiquitin genes are embedded reflect the differences in codon choice and in gene expression pattern between poly- and monomeric ubiquitin genes.

Amino Acid Sequence

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals

Frequencies of codons in histones, tubulins and fibrinogen: bias due to interference between transcription signals and protein function.

The distribution of codons was studied in 65 proteins: 48 histones, 14 tubulins, and three fibrinogens, With the methodology used, (1) we confirmed that the preterminator state of a codon has no detectable effect on codon bias. (2) The well-known effect of CG suppression was visible. We also found that (3) some codons which are very rare, are equal to parts of known transcription signals. Thus, we advanced that to avoid signal interference, the use of these codons is suppressed when a synonymous codon is available. In addition we found that in the whole series of codons, transcription signals are less frequent than in a random sequence of equal composition. Finally we observed (4) that tryptophan is absent in histones. This absence was related not to the TGG codon itself, but to characteristics of the amino acid. We conclude that the functional constraints of a protein can influence, at least for synonymous codon usage, the evolution of its own coding sequence.

Animals

Human apolipoprotein B (apoB) mRNA: identification of two distinct apoB mRNAs, an mRNA with the apoB-100 sequence and an apoB mRNA containing a premature in-frame translational stop codon, in both liver and intestine.

Human apolipoprotein B (apoB) is present in plasma as two separate isoproteins, designated apoB-100 (512 kDa) and apoB-48 (250 kDa). ApoB is encoded by a single gene on chromosome 2, and a single nuclear mRNA is edited and processed into two separate apoB mRNAs. A 14.1-kilobase apoB mRNA codes for apoB-100, and the second mRNA, which codes for apoB-48, contains a premature stop codon generated by a single base substitution of cytosine to uracil at nucleotide 6538, which converts the translated CAA codon coding for the amino acid glutamine at residue 2153 in apoB-100 to a premature in-frame stop codon (UAA). Two 30-base synthetic oligonucleotides (nucleotides 6523-6552 of apoB mRNA), designated apoB-Stop and apoB-Gln, were synthesized containing the complementary sequence to the stop codon (UAA) and glutamine codon (CAA), respectively. Analysis of intestinal apoB mRNA by hybridization with apoB-Stop and apoB-Gln probes and sequence analysis of apoB clones in two independent human small intestinal cDNA libraries established that intestinal apoB mRNA contained both the apoB mRNA that codes for apoB-100 and the apoB mRNA containing the premature in-frame stop codon, which codes for apoB-48. Investigation of hepatic apoB mRNA and two hepatic cDNA libraries by hybridization with the apoB-Stop and apoB-Gln synthetic probes as well as by cDNA sequencing revealed that liver apoB mRNA also contains both the apoB-100 mRNA and the apoB-48 mRNA containing the stop codon. The combined results from these studies establish that both human intestine and liver contain the two distinct apoB mRNAs, an mRNA that codes for apoB-100 and an apoB mRNA that contains the premature stop codon, which codes for apoB-48. The premature in-frame stop codon is not tissue specific and is present in both human liver and intestine.

Apolipoprotein B-100

TTA codons in some genes prevent their expression in a class of developmental, antibiotic-negative, Streptomyces mutants.

In Streptomyces coelicolor A3(2) and the related species Streptomyces lividans 66, aerial mycelium formation and antibiotic production are blocked by mutations in bldA, which specifies a tRNA(Leu)-like gene product which would recognize the UUA codon. Here we show that phenotypic expression of three disparate genes (carB, lacZ, and ampC) containing TTA codons depends strongly on bldA. Site-directed mutagenesis of carB, changing its two TTA codons to CTC (leucine) codons, resulted in bldA-independent expression; hence the bldA product is the principal tRNA for the UUA codon. Two other genes (hyg and aad) containing TTA codons show a medium-dependent reduction in phenotypic expression (hygromycin resistance and spectinomycin resistance, respectively) in bldA mutants. For hyg, evidence is presented that the UUA codon is probably being translated by a tRNA with an imperfectly matched anticodon, giving very low levels of gene product but relatively high resistance to hygromycin. It is proposed that TTA codons may be generally absent from genes expressed during vegetative growth and from the structural genes for differentiation and antibiotic production but present in some regulatory and resistance genes associated with the latter processes. The codon may therefore play a role in developmental regulation.

Anti-Bacterial Agents

Diagrammatization of codon usage in 339 human immunodeficiency virus proteins and its biological implication.

The occurrence frequencies of bases A (adenine), C (cytosine, G (guanine), and T (thymine) occurring in the 1st, 2nd, and 3rd codon positions in the codon usage table of viral genes for the 339 human immunodeficiency virus (HIV) proteins compiled recently have been calculated and diagrammatized. For comparison, the corresponding diagrammatic representations for the 2681 human proteins from the codon usage table for primate genes are also presented. The analyzed results based on these characteristic diagrams indicate that considerably similar features have been found between HIV and human proteins for the 1st and 2nd codon positions; i.e., they are all occupied predominantly by purine, especially base A. However, a significant difference in the 3rd codon position between HIV and human proteins has been observed; i.e., human proteins are of high C + G content and low A + G content in the 3rd codon position, whereas the case is just the opposite for HIV proteins. The biological implication of such a duality on the codon bias of HIV against human proteins is discussed. It is suggested that the 1st and 2nd codon positions can be termed as the structure-determining position, and the 3rd codon position termed as the species-determining position. The diagrammatic representation and analysis method described here possess a great potential for the study of molecular evolution from the viewpoint of the genetic code for which data have been accumulated rapidly and will continue to grow at a much faster pace.

Base Composition

Novel in-frame two codon translational hop during synthesis of bovine placental lactogen in a recombinant strain of Escherichia coli.

A recombinant Escherichia coli strain was constructed for the overexpression of bovine placental lactogen (bPL), using a bPL structural gene containing 9 of the rare arginine codons AGA and AGG. When high level bPL synthesis was induced in this strain, cell growth was inhibited and bPL accumulated to less than 10% of total cell protein. In addition, about 2% of the recombinant bPL produced from this strain exhibited an altered trypsin digestion pattern. Amino acid residues 74 through 109 normally produce 2 tryptic peptides, but the altered form of bPL lacked these two peptides and instead had a new peptide which was missing arginine residue 86 and one of the two flanking leucine residues. The codon for arginine residue 86 was AGG and the codons for the flanking leucine residues 85 and 87 were TTG. When 5 of the 9 AGA and AGG codons in the bPL structural gene were changed to more preferred arginine codons, cell growth was not inhibited and bPL accumulated to about 30% of total cell protein. When bPL was purified from this modified strain, which included changing the arginine codon at position 86 from AGG to CGT, none of the altered form of bPL was produced. These observations are consistent with a model in which translational pausing occurs at the arginine residue 86 AGG codon because the corresponding arginyl-tRNA species is reduced by the high level of bPL synthesis, and a translational hop occurs from the leucine residue 85 TTG codon to the leucine residue 87 TTG codon. This observation represents the first report of an error in protein synthesis due to an in-frame translational hop within an open reading frame.

Amino Acid Sequence