PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Evolution of the genetic triplet code via two types of doublet codons.

Explaining the apparent non-random codon distribution and the nature and number of amino acids in the 'standard' genetic code remains a challenge, despite the various hypotheses so far proposed. In this paper we propose a simple new hypothesis for code evolution involving a progression from singlet to doublet to triplet codons with a reading mechanism that moves three bases each step. We suggest that triplet codons gradually evolved from two types of ambiguous doublet codons, those in which the first two bases of each three-base window were read ('prefix' codons) and those in which the last two bases of each window were read ('suffix' codons). This hypothesis explains multiple features of the genetic code such as the origin of the pattern of four-fold degenerate and two-fold degenerate triplet codons, the origin of its error minimising properties, and why there are only 20 amino acids.

Amino Acids↗

Global mRNA stability is not associated with levels of gene expression in Drosophila melanogaster but shows a negative correlation with codon bias.

A multitude of factors contribute to the regulation of gene expression in living cells. The relationship between codon usage bias and gene expression has been extensively studied, and it has been shown that codon bias may have adaptive significance in many unicellular and multicellular organisms. Given the central role of mRNA in post-transcriptional regulation, we hypothesize that mRNA stability is another important factor associated either with positive or negative regulation of gene expression. We have conducted genome-wide studies of the association between gene expression (measured as transcript abundance in public EST databases), mRNA stability, codon bias, GC content, and gene length in Drosophila melanogaster. To remove potential bias of gene length inherently present in EST libraries, gene expression is measured as normalized transcript abundance. It is demonstrated that codon bias and GC content in second codon position are positively associated with transcript abundance. Gene length is negatively associated with transcript abundance. The stability of thermodynamically predicted mRNA secondary structures is not associated with transcript abundance, but there is a negative correlation between mRNA stability and codon bias. This finding does not support the hypothesis that codon bias has evolved as an indirect consequence of selection favoring thermodynamically stable mRNA molecules.

Animals↗

Why are translationally sub-optimal synonymous codons used in Escherichia coli?

Natural selection favors certain synonymous codons which aid translation in Escherichia coli, yet codons not favored by translational selection persist. We use the frequency distributions of synonymous polymorphisms to test three hypotheses for the existence of translationally sub-optimal codons: (1) selection is a relatively weak force, so there is a balance between mutation, selection, and drift; (2) at some sites there is no selection on codon usage, so some synonymous sites are unaffected by translational selection; and (3) translationally sub-optimal codons are favored by alternative selection pressures at certain synonymous sites. We find that when all the data is considered, model 1 is supported and both models 2 and 3 are rejected as sole explanations for the existence of translationally sub-optimal codons. However, we find evidence in favor of both models 2 and 3 when the data is partitioned between groups of amino acids and between regions of the genes. Thus, all three mechanisms appear to contribute to the existence of translationally sub-optimal codons in E. coli.

Codon↗

Codon usage bias amongst plant viruses.

An internet database (DPVweb) was established containing details of all sequences of viruses, viroids and satellites of plants that are complete or that contain at least one complete gene (n>4600). The start and end positions of each feature (genes, non-translated regions etc) were recorded and checked for accuracy. Client software was written to enable easy selection of sequences and features of a chosen virus and to analyse codon usage bias. Codon usage was analysed for each gene of one example of each fully-sequenced plant virus. There were large differences in codon preferences, related to the nucleotide composition of the genome, particularly the GC content of the third codon position. There was no effect of gene size on codon bias. Genes from the same genome usually had similar coding strategies except where constrained by the overlap of reading frames. Although some synonymous codons were consistently used with low frequency by both plants and viruses, viruses were not generally adapted to use (or avoid) those codons most frequently used by their host plants and there was no obvious association with the type of transmission. Mutational bias, rather than translational selection appears to account for the majority of the variation detected. The software is available at http://www.dpvweb.net/analysis/codons.php.

Codon↗

Association of sporadic Creutzfeldt-Jakob disease with homozygous genotypes at PRNP codons 129 and 219 in the Korean population.

Human prion protein gene (PRNP) is considered an important gene in determining the incidence of human transmissible spongiform encephalopathies or prion diseases. Polymorphisms of PRNP at codon 129 in Europeans and codon 219 in Japanese may play an important role in the susceptibility to sporadic Creutzfeldt-Jakob disease (CJD); data regarding codon 129 in the Japanese population have led to divergent interpretations. In order to determine which, if any, of the PRNP genotypes in Korean people are associated with sporadic CJD, we examined the genotype and allelic distributions of human PRNP polymorphisms in 150 patients with sporadic CJD. All Korean sporadic CJD patients were Met/Met at codon 129, Glu/Glu at codon 219 and undeleted at the octarepeat region of PRNP. Our study showed significant differences in genotype frequency of PRNP at codon 129 (chi 2=8.8998, P=0.0117) or 219 (chi 2=12.6945, P=0.0004) between sporadic CJD and normal controls. Furthermore, the genotype frequency of the heterozygotes for codons 129 and/or 219 showed a significant difference between the normal population and sporadic CJD patients (chi 2=21.0780, P<0.0001).

Alleles↗

Optimization of codon usage of poxvirus genes allows for improved transient expression in mammalian cells.

Transient expression of viral genes from certain poxviruses in uninfected mammalian cells can sometimes be unexpectedly inefficient. The reasons for poor expression levels can be due to a number of features of the gene cassette, such as cryptic splice sites, polymerase II termination sequences or motifs that lead to mRNA instability. Here we suggest that in some cases the problem of low protein expression in transfected mammalian cells may be due to inefficient codon usage. We have observed that for many poxvirus genes from the yatapoxvirus genus this deficiency can be overcome by synthesis of the gene with codon sequences optimized for expression in primate cells. This led us to examine colon usage across 2-dozen sequenced members of the Poxviridae. We conclude that codon usage is surprisingly divergent across the different Poxviridae genera but is much more conserved within a single genus. Thus, Poxviridae genera can be divided into distinct groups based on their observed codon bias. When viewed in this context, successful transient expression of transfected poxvirus genes in uninfected mammalian cells can be more accurately predicted based on codon bias. As a corollary, for specific poxvirus genes with less favorable codon usage, codon optimization can result in profoundly increased transient expression levels following transfection of uninfected mammalian cell lines.

Animals↗

Identification of a novel in-frame translational stop codon in human intestine apoB mRNA.

Human apolipoprotein (apo) B exists in plasma as two isoproteins designated apoB-100 and apoB-48. ApoB-100 (512 kDa) and apoB-48 (250 kDa) are synthesized by the liver and intestine respectively. Analysis of apoB cDNA clones isolated from a human intestinal cDNA library revealed that the intestinal apoB mRNA contains a new in-frame translational stop codon. This premature stop codon is generated by a single base substitution of a 'C' to 'T' at nucleotide 6538 which converts the codon 'CAA' coding for the amino acid glutamine residue 2153 to an in-frame stop codon 'TAA'. The generation of a stop codon in the intestinal apoB mRNA appears to be tissue specific since it has not been reported in cDNA clones isolated from human liver cDNA libraries which code for the 4536 amino acid apoB-100. A potential polyadenylation signal sequence 'AATAAA' was also identified 390 bases downstream from the new stop codon. The new stop codon in the human intestinal apoB mRNA provides a potential mechanism for the biosynthesis of intestinal apoB-48.

Amino Acid Sequence↗

Effects of surrounding sequence on the suppression of nonsense codons.

Using a lacI-Z fusion system, we have determined the efficiency of suppression of nonsense codons in the I gene of Escherichia coli by assaying beta-galactosidase activity. We examined the efficiency of four amber suppressors acting on 42 different amber (UAG) codons at known positions in the I gene, and the efficiency of a UAG suppressor at 14 different UGA codons. The largest effects were found with the amber suppressor supE (Su2), which displayed efficiencies that varied over a 35-fold range, and with the UGA suppressor, which displayed a 170-fold variation in efficiency. Certain UGA sites were so poorly suppressed (less than 0.2%) by the UGA suppressor that they were not originally detected as nonsense mutations. Suppression efficiency can be correlated with the sequence on the 3' side of the codon being suppressed, and in many cases with the first base on the 3' side. In general, codons followed by A or G are well suppressed, and codons followed by U or C are poorly suppressed. There are exceptions, however, since codons followed by CUG or CUC are well suppressed. Models explaining the effect of the surrounding sequence on suppression efficiency are considered in the Discussion and in the accompanying paper.

Base Sequence↗

Sense codons are found in specific contexts.

The sequence environment of codons in structural genes has been investigated statistically, using computer methods. A set of Escherichia coli genes with abundant products was compared with a set having low gene product levels, in order to detect potential differences associated with expression. The results show striking non-randomness in the nucleotides occurring near codons. These effects are, unexpectedly, very much larger and more homogeneous among the genes with rare products. The intensity of effects in weakly expressed genes suggests that such non-random sequence environments decrease expression. In the weakly expressed set of genes, the 5' neighbor of a codon, and all positions of the 3' neighbor codon are biased. In the highly expressed genes, the first nucleotide of the next codon is a uniquely affected site. The distribution of non-randomness in weakly expressed genes suggests that sequence bias is primarily due to a constraint acting directly on the secondary or tertiary structure of the codon/anticodon. In highly expressed genes, the observed bias suggests an interaction between the codon/anticodon and a site outside the codon/anticodon. Much of the tendency to non-random near-neighbor sequences in weakly expressed genes can be ascribed to a correlation between nearby nucleotides and the wobble nucleotide of the codon, despite the fact that selection of such correlations will alter the amino acid sequence. The favored pattern, in genes expressed at low level, is R YYR or Y RRY. R indicates purine, Y indicates pyrimidine; the space is the boundary between codons. It seems likely that this preference for nearby sequences is the physical basis of the genetic context effect. Under this assumption such sequence biases will affect expression. On this basis, we predict new sites for contextual mutations which decrease expression, and suggest strategy for the design of messages having optimal translational activity.

Amino Acids↗

Codon contexts from weakly expressed genes reduce expression in vivo.

Nucleotides that neighbor codons in Escherichia coli genes are highly non-random. Furthermore, these context biases are stronger and extend farther from the codon in weakly expressed than in highly expressed genes. We therefore suggested that codon contexts are selected to reduce gene expression levels. We now compare the expression levels of lacZ genes containing two specific coding sequences (context inserts). One context insert represents contexts seen in weakly expressed genes (low variant); the other represents contexts seen in highly expressed genes (high variant). The two variants have identical nucleotide and codon compositions, and encode the same protein. A permutation of four nucleotides, which changes eight codon:codon interfaces of 1043, comprises the only difference between the high and low context variant genes. In three different lacZ mRNAs, the low variant was expressed at a level significantly below that of the high variant. This context effect depends entirely on translation of the contexts in the correct frame; its magnitude depends in part on the placement of other features (e.g. transcriptional pauses and terminators, or perhaps other slow codons or contexts) in the mRNAs. Changing the ribosome density on the message by changing the ribosome binding site distinguishes between dropoff, interference and polarity, three fundamentally different types of models for the context effect. The expression difference between context variants is eliminated by both increases and decreases in the ribosome initiation frequency, as uniquely predicted by the polarity model. In fact, data from all constructions are accommodated by a model in which slow translation of the low context insert increases rho-dependent transcriptional termination within the test gene. The data suggest that the rates of translational initiation and elongation are poised with respect to the rate of transcriptional elongation so that all are influential in setting the expression level of wild-type lacZ. We conclude that context-induced polarity will exist in genes wherever low and reproducible gene product levels have been selected.

Amino Acid Sequence↗

Codon recognition patterns as deduced from sequences of the complete set of transfer RNA species in Mycoplasma capricolum. Resemblance to mitochondria.

The nucleotide sequences of the complete set of tRNA species in Mycoplasma capricolum, a derivative of Gram-positive eubacteria, have been determined. This bacterium represents the first genetic system in which the sequences of all the tRNA species have been determined at the RNA level. There are 29 tRNA species: three for Leu, two each for Arg, Ile, Lys, Met, Ser, Thr and Trp, and one each for the other 12 amino acids as judged from aminoacylation and the anticodon nucleotide sequences. The number of tRNA species is the smallest among all known genetic systems except for mitochondria. The tRNA anticodon sequences have revealed several features characteristic of M. capricolum. (1) There is only one tRNA species each for Ala, Gly, Leu, Pro, Ser and Val family boxes (4-codon boxes), and these tRNAs all have an unmodified U residue at the first position of the anticodon. (2) There are two tRNAThr species having anticodons UGU and AGU; the first positions of these anticodons are unmodified. (3) There is only one tRNA with anticodon ICG in the Arg family box (CGN); this tRNA can translate codons CGU, CGC and CGA. No tRNA capable of translating codon CGG has been detected, suggesting that CGG is an unassigned codon in this bacterium. (4) A tRNATrp with anticodon UCA is present, and reads codon UGA as Trp. On the basis of these and other observations, novel codon recognition patterns in M. capricolum are proposed. A comparatively small total, 13, of modified nucleosides is contained in all M. capricolum tRNAs. The 5' end nucleoside of the T psi C-loop (position 54) of all tRNAs is uridine, not modified to ribothymidine. The anticodon composition, and hence codon recognition patterns, of M. capricolum tRNAs resemble those of mitochondrial tRNAs.

Amino Acids↗

tRNA anticodons with the modified nucleoside 2-methylthio-N6-(4-hydroxyisopentenyl)adenosine distinguish between bases 3' of the codon.

The modified nucleoside 2-methylthio-N6-(4-hydroxyisopentenyl)adenosine (ms2io6A) is present immediately to the 3' side of the anticodon (position 37) in tRNAs that read codons starting with uridine and hence include amber (UAG) suppressor tRNAs. We have used strains of Salmonella typhimurium that differ only in their ability to synthesize ms2io6A in order to determine specifically how this modified nucleoside influences the efficiency of amber suppression in two codon contexts differing by only which base is 3' of the codon. The results show that the presence of the modified nucleoside ms2io6A not only improves the efficiency of the suppressor tRNAs but also allows them to distinguish between at least two bases 3' of the codon. Thus, the presence of ms2io6A reduces the intrinsic codon context sensitivity of the tRNA and specifically counteracts an unfavourable nucleotide on the 3' side of the codon. The possible codon-anticodon interactions responsible for this effect are discussed.

Anticodon↗

Both forms of translational initiation factor IF2 (alpha and beta) are required for maximal growth of Escherichia coli. Evidence for two translational initiation codons for IF2 beta.

The gene infB codes for two forms of translational initiation factor IF2; IF2 alpha (97,300 Da) and IF2 beta (79,700 Da). IF2 beta arises from an independent translational event on a GUG codon located 471 bases downstream from IF2 alpha start codon. By site-directed mutagenesis we constructed six different mutations of this GUG codon. In all cases, IF2 beta synthesis was variably affected by the mutations but not abolished. We show that the residual expression of IF2 beta results from translational initiation on an AUG codon located 21 bases downstream from the mutated GUG. Furthermore, two forms of IF2 beta have been separated by fast protein liquid chromatography and the determination of their N-terminal sequences indicated that they resulted from two internal initiation events, one occurring on the previously identified GUG start codon, the other on the AUG codon immediately downstream. We conclude that two forms of IF2 beta exist in the cell, which differ by seven aminoacid residues at their N terminus. Only by mutating both IF2 beta start codons could we construct plasmids that express only IF2 alpha. A plasmid expressing only IF2 beta was obtained by deletion of the proximal region of the infB gene. Using a strain that carries a null mutation in the chromosomal copy of infB and a functional copy of the same gene on a thermosensitive lysogenic lambda phage, we could cure the lambda phage when the plasmids expressing only one form of IF2 were supplied in trans. We found that each one of the two forms of IF2, at near physiological levels, can support growth of Escherichia coli, but that growth is retarded at 37 degrees C. This result shows that both forms of IF2 are required for maximal growth of the cell and suggests that they have acquired some specialized but not essential function.

Amino Acid Sequence↗

Replication capacities of natural and artificial precore stop codon mutants of hepatitis B virus: relevance of pregenome encapsidation signal.

The emergence of hepatitis B virus variants unable to express HBe protein during late stage of viral infection may represent an important mechanism of viral persistence. The molecular mechanisms responsible for the elimination of HBe expression are nonsense or frameshift mutations or initiation codon mutations in part of its coding sequence, the precore region. So far only 2 of the 29 precore amino acid codons have been found mutated to stop codons in nature, although a total of 10 codons are convertible to stop codons by single nucleotide changes. Since the HBe-coding sequence is largely overlapped by the pregenome encapsidation signal (epsilon signal), a recently found cis-acting element required for the packaging of pregenomic RNA, the absence of other potential nonsense mutants could result from their impairment of the epsilon signal. Seven such potential stop codon mutants were constructed and tested for replication capacities by transfection into a hepatoma cell line. Five mutants were replication competent, but at levels lower than that of a prevalent natural stop codon mutant. The remaining two mutants were completely defective in DNA replication, which clearly explained why these two mutants are not found in nature. Northern blot analysis revealed wild-type levels of RNA transcription by these two mutants but complete lack of packaged pregenomic RNA. Additional studies lent further support to the importance of the epsilon signal in pregenome encapsidation and suggested relaxed sequence requirements for the computer-predicted hexanucleotide bulge region as compared to the hexanucleotide loop of the signal.

Base Sequence↗

Cell-type specific codon usage and differentiation.

This paper presents evidence derived from the selective use of codons in ca. 40 eukaryotic genes (or messages derived from them) that codon usage is one of the most conserved features of messages for specific cell products. The theory has been developed by several investigations that the kinds of products of many if not most kinds of differentiated cells is determined by the pattern of translation abilities each cell possesses as it differentiates. A correllary of this thesis is that the groups of code words used for products of specific cell types, INDEPENDENTLY OF THE SPECIES INVOLVED, should exclude specific kinds of code words in one cell type and not another. To test this thesis, the specific frequency of codon usage and non-usage has been collated from the recently published literature and subjected to appropriate computer analysis. We find that: (1) certain codons are not used at all in any message for globins; (2) the pattern of codon usage is characteristic of specific products from specific embryonic derivatives, e.g. erythrocytes; (3) that certain code words are discriminated against generally in nearly all vertebrate cell messages evaluated; and (4) that the cell type from which a message is derived can be identified, at least in the case of 9 globin messages derived from species separated by millions of generations, purely on the basis of codon usage. From these studies it can be inferred that some evolutionary factor prevents the use of "forbidden" code words in specific kinds of cells. We propose that this factor derives from the fact that a "silent" mutation to a code word which is untranslatable by a differentiated cell will be lethal in the homozygous condition and that the thesis that so-called "codon restriction" is an important determinative factor in limiting what differentiating cells can synthesize in many kinds of developing cells explains the available evidence more adequately than alternative theories.

Animals↗

Multiple upstream AUG codons mediate translational control of GCN4.

GCN4 encodes a transcriptional activator of amino acid biosynthetic genes in yeast that is regulated at the translational level. The 5' leader of GCN4 mRNA contains four small open-reading-frames. By constructing point mutations in the initiation codons of these sequences, we show that they are essential for translational repression of GCN4. Each upstream AUG codon can repress translation; however, the two 3' proximal AUG codons are much more inhibitory than the 5' proximal AUG codons. Unexpectedly, the first AUG codon is required for efficient GCN4 expression under starvation conditions. This positive function appears to involve antagonism of the inhibitory effect of the 3' proximal AUG codons since it is dispensable in the absence of these sequences. The interaction between the upstream AUG codons is modulated by the trans-acting factors GCN2 and GCD1 in response to amino acid availability.

Amino Acids↗

The ability of bovine mitochondrial transfer RNAMet to decode AUG and AUA codons.

The ability of bovine mitochondrial tRNA(Met) with the anticodon f5CAU (where f5C is 5-formylcytidine) to decode AUG and AUA codons was examined in a codon-dependent ribosomal binding assay. The AUG codon stimulated the binding of Met-tRNA(Met) to mitochondrial ribosomes in the presence of EF-Tu/TSmt. In contrast, the AUA codon did not promote the binding to mitochondrial Met-tRNA to the ribosome. To investigate the translation of the AUG and AUA codons more fully, an in vitro translation system from bovine liver mitochondria was developed. The activity of this system was greatly enhanced by the addition of 1 mM spermine and reached about half the activity observed with a comparable translational system from E coli. Two types of mRNA containing either AUG or AUA codons were synthesized using T7 RNA polymerase to transcribe their chemically synthesized genes. In the E coli system, the AUG-containing mRNA was translated as Met and the AUA-containing mRNA was translated as Ile. The AUG-containing mRNA but not the AUA-containing mRNA was translated as Met by the mitochondrial translational system. The process by which the AUA codon is translated as Met in the mitochondrial system remains to be clarified.

Animals↗

A universal compositional correlation among codon positions.

We have investigated the compositional distributions of third codon positions of genes from the 16 prokaryotes and seven eukaryotes for which the largest numbers of coding sequences are available in data banks. In prokaryotes, both narrow and broad distributions were found. In eukaryotes, distributions were very broad (except for Saccharomyces cerevisiae) and remarkably different for different genomes. In low-GC genomes, third codon positions were lower in GC than first + second codon positions and trailed towards high GC; the opposite situation was found for high-GC genomes. In all genomes, first codon positions were higher in GC than second codon positions. We then investigated the compositional correlations between third and first + second codon positions in prokaryotic genomes (the 16 mentioned above plus 87 additional ones) and in genome compartments of eukaryotes. A general, common relationship was found, which also holds within the same (heterogeneous) genomes. This universal correlation is due to the fact that the relative effects of compositional constraints on different codon positions are the same, on the average, whatever the genome under consideration.

Animals↗