PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Codon usage in higher plants, green algae, and cyanobacteria.

Codon usage is the selective and nonrandom use of synonymous codons by an organism to encode the amino acids in the genes for its proteins. During the last few years, a large number of plant genes have been cloned and sequenced, which now permits a meaningful comparison of codon usage in higher plants, algae, and cyanobacteria. For the nuclear and organellar genes of these organisms, a small set of preferred codons are used for encoding proteins. Codon usage is different for each genome type with the variation mainly occurring in choices between codons ending in cytidine (C) or guanosine (G) versus those ending in adenosine (A) or uridine (U). For organellar genomes, chloroplastic and mitochrondrial proteins are encoded mainly with codons ending in A or U. In most cyanobacteria and the nuclei of green algae, proteins are encoded preferentially with codons ending in C or G. Although only a few nuclear genes of higher plants have been sequenced, a clear distinction between Magnoliopsida (dicot) and Liliopsida (monocot) codon usage is evident. Dicot genes use a set of 44 preferred codons with a slight preference for codons ending in A or U. Monocot codon usage is more restricted with an average of 38 codons preferred, which are predominantly those ending in C or G. But two classes of genes can be recognized in monocots. One set of monocot genes uses codons similar to those in dicots, while the other genes are highly biased toward codons ending in C or G with a pattern similar to nuclear genes of green algae. Codon usage is discussed in relation to evolution of plants and prospects for intergenic transfer of particular genes.

Journal Article↗

Intragenic codon bias in a set of mouse and human genes.

To better conceptualize the mechanism underlying the evolution of synonymous codons, we have analysed intragenic codon usage in chosen "regions" of some mouse and human genes. We divided a given gene into two regions: one consisting of a trinucleotide repeat (TNR) and the other consisting of the "rest of the coding region" (RCR). Usually, a TNR is composed of a repetitive single codon, which may reflect its frequency in a gene. In contrast, a non-random frequency of a codon in the RCR versus TNR (or vice versa) of a gene should indicate a bias for that codon within the TNR. We examined this scenario by comparing codon frequency between the RCR and the cognate TNR(s) for a set of human and mouse genes. A TNR length of six amino acids or more was used to identify genes from the Genbank database. Twenty nine human and twenty one mouse genes containing TNRs coding for nine different amino acid runs were identified. The ratio of codon frequency in a TNR versus the corresponding RCR was expressed as "fold change" which was also regarded as a measure of codon bias (defined as preferential use either in TNR or in RCR). Chi-square values were then determined from the distribution of codon frequency in a TNR vs. the cognate RCR. At p<0.001, 22% and 27%, respectively, of human and mouse TNRs showed codon bias. Greater than 40% of the TNRs (29 out of 69 in human, and 18 of 42 in mouse) showed codon bias at p<0.05. In addition, we identify eight single-codon TNRs in mouse and ten in human genes. Thus, our results show intragenic codon bias in both mouse and human genes expressed in diverse tissue types. Since our results are independent of the Codon Adaptation Index (CAI) and starvation CAI, and since the tRNA repertoire in a cell or in a tissue is constant, our data suggest that other constraints besides tRNA abundance played a role in creating intragenic codon bias in these genes.

Amino Acids↗

5' contexts of Escherichia coli and human termination codons are similar.

The nearest 5' context of 2559 human stop codons was analysed in comparison with the same context of stop-like codons (UGG, UGC, UGU, CGA for UGA; CAA, UAU, UAC for UAA; and UGG, UAU, UAC, CAG for UAG). The non-random distribution of some nucleotides upstream of the stop codons was observed. For instance, uridine is over-represented in position -3 upstream of UAG. Several codons were shown to be over-represented immediately upstream of the stop codons: UUU(Phe), AGC(Ser), and the Lys and Ala codon families before UGA; AAG(Lys), GCG(Ala), and the Ser and Leu codon families before UAA; and UCA(Ser), AUG(Met), and the Phe codon family before UAG. In contrast, the Thr and Gly codon families were under-represented before UGA, while ACC(Thr) and the Gly codon family were under-represented before UAG and UAA respectively. In an earlier study, uridine was shown to be over-represented in position -3 before UGA in Escherichia coli [Arkov,A.L., Korolev,S.V. and Kisselev,L.L. (1993) Nucleic Acids Res., 21,2891-2897]. In that study, the codons for Lys, Phe and Ser were shown to be over-represented immediately upstream of E. coli stop codons. Consequently, E. coli and human termination codons have similar 5' contexts. The present study suggests that the 5' context of stop codons may modulate the efficiency of peptide chain termination and (or) stop codon readthrough in higher eukaryotes, and that the mechanisms of such a modulation in prokaryotes and higher eukaryotes may be very similar.

Amino Acids↗

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

Codon contexts in enterobacterial and coliphage genes.

This investigation of the codon context of enterobacteria, plasmid, and phage protein genes was based on a search for correlations between the presence of one base type at codon position III and the presence of another base type at some other position in adjacent codons. Enterobacterial genes were compared with eukaryotic sequences for codon context effects. In enterobacterial genes, base usage at codon position III is correlated with the third position of the upstream adjacent codon and with all three positions of the downstream codon. Plasmid genes are free of context biases. Phage genes are heterogeneous: MS2 codons have no biased context, whereas lambda genes partly follow the trends of the host bacterium, and T7 genes have biased codon contexts that differ from those of the host. It has been reported that two successive third-codon positions tend to be occupied by two purines or two pyrimidines in Escherichia coli genes of low expression level. Here, the extent to which highly expressed protein genes can modulate base usage at two successive codon positions III, given the constraints on codon usage and protein sequence that act on them, was quantified. This demonstrates that the above-mentioned favored patterns are not a characteristic of weakly expressed genes but occur in all genes in which codon context can vary appreciably. The correlation between successive third-codon positions is a distinct feature of enterobacteria and of some phages, one that may result from adaptation of gene structure to translational efficiency. Conversely, codon context in yeast and human genes is biased--but for reasons unrelated to translation.

Animals↗

Mutability of p53 hotspot codons to benzo(a)pyrene diol epoxide (BPDE) and the frequency of p53 mutations in nontumorous human lung.

p53 mutations are common in lung cancer. In smoking-associated lung cancer,the occurrence of G:C to T:A transversions at hotspot codons, e.g., 157, 248, 249,and 273, has been linked to the presence of carcinogenic chemicalsin tobacco smoke including polycyclic aromatic hydrocarbons suchas benzo(a)pyrene (BP). In the present study, we have used a highly sensitive mutation assay to determine the p53 mutation load in nontumorous human lung and to study the mutability of p53 codons 157, 248, 249, and 250 to benzo(a)pyrene-diol-epoxide (BPDE), an active metabolite of BP in human bronchial epithelial BEAS-2B cells. We determined the p53 mutational load at codons 157, 248, 249, and 250 in nontumorous peripheral lung tissue either from lung cancer cases among smokers or noncancer controls among smokers and nonsmokers. A 5-25-fold higher frequency of GTC(val) to TTC(phe) transversions at codon 157 was found in nontumorous samples (57%) from cancer cases (n = 14) when compared with noncancer controls (n = 8; P < 0.01). Fifty percent (7/14) of the nontumorous samples from lung cancer cases showed a high frequency of codon 249 AGG(arg) to AGT(ser) mutations (P < 0.02). Four of these seven samples with AGT(ser) mutations also showed a high frequency of codon 249 AGG(arg) to ATG(met) mutations, whereas only one sample showed a codon 250 CCC to ACC transversion. Tumor tissue from these lung cancer cases (38%) contained p53 mutations but were different from the above mutations found in the nontumorous pair. Noncancer control samples from smokers or nonsmokers did not contain any detectable mutations at codons 248, 249, or 250. BEAS-2B bronchial epithelial cells exposed to doses of 0.125, 0.5, and 1.0 microM BPDE, showed G:C to T:A transversions at codon 157 at a frequency of 3.5 x 10(-7), 4.4 x 10(-7), and 8.9 x 10(-7), respectively. No mutations at codon 157 were found in the DMSO-treated controls. These doses of BPDE induced higher frequencies, ranging from 4-12-fold, of G:C to T:A transversions at codon 248, G:C to T:A transversions and G:C to A:T transitions at codon 249, and C:G to T:A transitions at codon 250 when compared with the DMSO-treated controls. These data are consistent with the hypothesis that chemical carcinogens such as BP in cigarette smoke cause G:C to T:A transversions at p53 codons 157, 248, and 249 and that nontumorous lung tissues from smokers with lung cancer carry a high p53 mutational load at these codons.

7,8-Dihydro-7,8-dihydroxybenzo(a)pyrene 9,10-oxide↗

Effect of the relative position of the UGA codon to the unique secondary structure in the fdhF mRNA on its decoding by selenocysteinyl tRNA in Escherichia coli.

The fdhF mRNA for formate dehydrogenase H of Escherichia coli contains a UGA codon at position 140. This termination codon is decoded by selenocysteinyl tRNA (the selC product) with the aid of its own specific elongation factor, SelB. For this decoding, a unique secondary structure immediately downstream of the UGA codon has been shown to be essential (Zinoni, F., Heider, J., and Böck, A. (1990) Proc. Natl. Acad. Sci. U. S. A. 87, 4660-4664). We examined the positional effect of the UGA codon relative to the secondary structure on its decoding using a fdhF-lacZ fusion gene. When the UGA codon was separated by one codon (position -1) from the secondary structure, the UGA decoding, as measured by the beta-galactosidase activity, dropped to approximately 76% of the normal level but was still almost as fully dependent upon selC and selenium in the culture medium as in the case of the UGA codon in the normal position (position 0). However, when the UGA codon was separated by two codons (position -2), the decoding level further dropped to 20% of the normal level, and in addition, became dependent only on selC but independent of selenium. When the UGA codon was further separated by three codons (position -3), the decoding level of UGA (-3) became higher than the decoding of UGA (-2) and was completely independent from selC and selenium, indicating that the UGA codon was nonspecifically suppressed. A similar nonspecific suppression was observed for the UGA codon at position -4, but at a lower level. When two UGA codons were tandemly placed at positions 0 and -1, they were still able to be decoded at 17% of the normal level in a selC- and selenium-dependent manner. In the absence of the SelB function, the decoding level of UGA(0) dropped to 1.6% of the normal level, whereas the UGA(-1) decoding dropped to 7.5%. These results indicate that the UGA codon at position 0 is not only most effectively decoded by selenocysteinyl tRNA but also tightly blocked from its nonspecific suppression in the absence of any components required for the decoding.

Base Sequence↗

Codon usage bias and base composition in MHC genes in humans and common chimpanzees.

Codon bias and base composition in major histocompatibility complex (MHC) sequences have been studied for both class I and II loci in Homo sapiens and Pan troglodytes. There is low to moderate codon bias for the MHC of humans and chimpanzees. In the class I loci, the same level of moderate codon bias is seen for HLA-B, HLA-C, Patr-A, Patr-B, and Patr-C, while at HLA-A the level of codon bias is lower. There is a correlation between codon usage bias and G+C content in the A and B loci in humans and chimps, but not at the C locus. To examine the effect of diversifying selection on codon bias, we subdivided class I alleles into antigen recognition site (ARS) and non-ARS codons. ARS codons had lower bias than non-ARS codons. This may indicate that the constraint of codon bias on nucleotide substitution may be selected against in ARS codons. At the class II loci, there are distinct differences between alpha and beta chain genes with respect to codon usage, with the beta chain genes being much more biased. Species-specific differences in base composition were seen in exon 2 at the DRB1 locus, with lower GC content in chimpanzees. Considering the complex evolutionary history of MHC genes, the study of codon usage patterns provides us with a better understanding of both the evolutionary history of these genes and the evolution of synonymous codon usage in genes under natural selection.

Animals↗

Comparative analysis of the base composition and codon usages in fourteen mycobacteriophage genomes.

To study the possible codon usage and base composition variation in the bacteriophages, fourteen mycobacteriophages were used as a model system here and both the parameters in all these phages and their plating bacteria, M. smegmatis had been determined and compared. As all the organisms are GC-rich, the GC contents at third codon positions were found in fact higher than the second codon positions as well as the first + second codon positions in all the organisms indicating that directional mutational pressure is strongly operative at the synonymous third codon positions. Nc plot indicates that codon usage variation in all these organisms are governed by the forces other than compositional constraints. Correspondence analysis suggests that: (i) there are codon usage variation among the genes and genomes of the fourteen mycobacteriophages and M. smegmatis, i.e., codon usage patterns in the mycobacteriophages is phage-specific but not the M. smegmatis-specific; (ii) synonymous codon usage patterns of Barnyard, Che8, Che9d, and Omega are more similar than the rest mycobacteriophages and M. smegmatis; (iii) codon usage bias in the mycobacteriophages are mainly determined by mutational pressure; and (iv) the genes of comparatively GC rich genomes are more biased than the GC poor genomes. Translational selection in determining the codon usage variation in highly expressed genes can be invoked from the predominant occurrences of C ending codons in the highly expressed genes. Cluster analysis based on codon usage data also shows that there are two distinct branches for the fourteen mycobacteriophages and there is codon usage variation even among the phages of each branch.

Bacteriophages↗

Impact of premature stop codons on mRNA levels in infantile Sandhoff disease.

Sandhoff disease is an autosomal recessive lysosomal storage disease resulting from mutations of the HEXB gene encoding the beta subunit of beta-hexosaminidase A. Fibroblast lines from four patients with the infantile form of the disease were investigated for mutations by single strand conformation polymorphism analysis and direct sequencing of PCR products. Two of the cell lines were homozygous for a common, 16 kb deletion of the 5' end of HEXB gene. The two other cell lines contained the 16 kb deletion along with a second mutant allele generating a stop codon: in one case a nonsense mutation, C850-->T, which generated a stop codon at codon 284; and in the other, a single base deletion, delta T1344, which generated a stop codon at codon 451. One additional cell line investigated was a compound heterozygote for two frameshift mutations, delta G774 in exon 7 and delta AG1305-1306 in exon 11 (McInnes et al. 1992, Biochim. Biophys. Acta 1138: 315-317). Stop codons were generated in this cell line at codons 274 and 454, respectively. We took advantage of these genotypes to investigate the steady-state level of mRNA produced by cells containing stop codons using a competitive polymerase chain reaction technique. The mRNA levels were, as percent of normal per single gene dose: for the stop codon at codon 451, 30%; for those at codons 274 and 454, combined percentage of 1.7%; and at codon 284, 0.8%. These studies demonstrate a dramatic difference in the steady-state level of Hex beta mRNA in the cell lines with stop codons in close proximity to each other (codons 451 vs 454).(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

The comparative method rules! Codon volatility cannot detect positive Darwinian selection using a single genome sequence.

All established methods for detecting positive selection at the molecular level rely on comparisons between nucleotide sequences. An exceptional method that purports to detect selection on the basis of a single genomic sequence has recently been proposed. This method uses a measure called "codon volatility," defined for each codon as the ratio between the number of nonsynonymous codons that differ from the codon under study at a single nucleotide position and the number of sense codons that differ from the codon under study at a single nucleotide position. Here, we examine various properties of codon volatility and its derivatives and use simulation of evolutionary processes to determine whether they can be used to detect selective pressures. Codons for only four amino acids (glycine, leucine, arginine, and serine) show any variation in codon volatility. Thus, codon volatility is mainly a proxy for amino acid usage, rather than for codon usage, with 65% of all synonymous changes and 27% of all nonsynonymous changes being undetectable by this measure. Genes identified by the volatility method as being subject to positive selection tend to have idiosyncratic amino acid compositions (e.g., they are glycine rich or arginine poor). An additional property of codon volatility is the near zero variance of its mean expectation, which translates into overestimated statistical significance estimates, especially in the absence of corrections for multiple comparisons. A comparison with measures of selection inferred through comparative methodology reveals no relationship between the results of the two methods. Finally, we show that codon volatility can increase in the absence of positive Darwinian selection; that is, increased codon volatility is not indicative of positive selection.

Animals↗

Saccharomyces cerevisiae ribosomes recognize non-AUG initiation codons.

A series of Saccharomyces cerevisiae plasmids and mutant derivatives containing fusions of the Escherichia coli galactokinase gene, galK, to the yeast iso-1-cytochrome c CYC1 transcription unit were used to study the sequences affecting the initiation of translation in S. cerevisiae. When the CYC1 AUG initiation codon preceded the galK AUG codon and coding sequence and either the two AUGs were out of frame with each other or a nonsense codon was located between them, the expression of the galK gene was extremely low. Deletion of the CYC1 AUG and its surrounding sequences resulted in a 100-fold increase in galK expression. This dependence of galK expression on the elimination of the CYC1 AUG codon was used to select mutations in that codon. Then the ability of these altered initiation codons to serve in translational initiation was determined by reconstruction of the CYC1 gene 3' to and in frame with them. Initiation was found to occur at the codons UUG and AUA, but not at the codons AAA and AUC. Furthermore the codon UUG, when preceded by an A three nucleotides upstream, served as a better initiation codon than when a U was substituted for the A. The efficiency of translation from these non-AUG codons was quantitated by using a CYC1/galK protein-coding fusion and measuring cellular galactokinase levels. Initiation at the UUG codon was 6.9% as efficient as initiation at the wild-type AUG codon when preceded by an A three nucleotides upstream, but was over 10-fold less efficient when a U was substituted for that A. Initiation at AUA was 0.5% as efficient as at AUG. The effects of the sequences preceding the initiation codon are discussed in light of these results.

Base Sequence↗

Differences in codon bias cannot explain differences in translational power among microbes.

BACKGROUND: Translational power is the cellular rate of protein synthesis normalized to the biomass invested in translational machinery. Published data suggest a previously unrecognized pattern: translational power is higher among rapidly growing microbes, and lower among slowly growing microbes. One factor known to affect translational power is biased use of synonymous codons. The correlation within an organism between expression level and degree of codon bias among genes of Escherichia coli and other bacteria capable of rapid growth is commonly attributed to selection for high translational power. Conversely, the absence of such a correlation in some slowly growing microbes has been interpreted as the absence of selection for translational power. Because codon bias caused by translational selection varies between rapidly growing and slowly growing microbes, we investigated whether observed differences in translational power among microbes could be explained entirely by differences in the degree of codon bias. Although the data are not available to estimate the effect of codon bias in other species, we developed an empirically-based mathematical model to compare the translation rate of E. coli to the translation rate of a hypothetical strain which differs from E. coli only by lacking codon bias. RESULTS: Our reanalysis of data from the scientific literature suggests that translational power can differ by a factor of 5 or more between E. coli and slowly growing microbial species. Using empirical codon-specific in vivo translation rates for 29 codons, and several scenarios for extrapolating from these data to estimates over all codons, we find that codon bias cannot account for more than a doubling of the translation rate in E. coli, even with unrealistic simplifying assumptions that exaggerate the effect of codon bias. With more realistic assumptions, our best estimate is that codon bias accelerates translation in E. coli by no more than 60% in comparison to microbes with very little codon bias. CONCLUSIONS: While codon bias confers a substantial benefit of faster translation and hence greater translational power, the magnitude of this effect is insufficient to explain observed differences in translational power among bacterial and archaeal species, particularly the differences between slowly growing and rapidly growing species. Hence, large differences in translational power suggest that the translational apparatus itself differs among microbes in ways that influence translational performance.

Bacterial Physiological Phenomena↗

[Analysis of factors shaping S. pneumoniae codon usage].

Streptococcus pneumoniae is a Gram-positive bacteria causing community acquired pneumonia, bacteremia, meningitis and otitis media. As a human pathogen, S. pneumoniae is the most common bacterial cause of acute respiratory infection and otitis media and is estimated to result in over 3 million deaths in children every year worldwide. S. pneumoniae has played a pivotal role in the fields of genetics and microbiology. The complete genome of S. pneumoniae was sequenced and published recently. In order to have a further insight into the synonymous codon usage evolution and to study S. pneumoniae gene codon usage pattern in highly and lowly expressed genes, factors shaping synonymous codon usage pattern of S. pneumoniae were analyzed in this paper. Genes larger than of equal to 300bp of the complete genome of S. pneumoniae (1709 genes in total) were analyzed. The gene expression level (CAI, codon adaption index), RSCU (relative synonymous codon usage), Nc (effective codon numbers), A3s, T3s, G3s, C3s (the frequencies of the adenine, thymine, guanine and cytosine at the synonymous third position of codons, respectively), GC (frequency of guanine + cytosine in gene sequence), GC3s (frequency of guanine + cytosine at the synonymous third position of codons) values and multivariate statistics were calculated. The results show that there is a significant increment of cytosine (C) usage at the synonymous positions in highly expressed genes than lowly expressed genes, while lowly expressed genes tend to use guanine (G) at synonymous sites. Gene expression has a significant correlation with the first axis of correspondence analysis (COA; R = 0.86) and significant effects on codon usage by comparing the codon usage patterns of highly expressed genes and lowly expressed genes. The G + C content of genes has a moderately correlation with gene expression (R = 0.44) and the first axis of the COA (R = 0.51), and therefore shapes gene expression and codon usage in S. pneumoniae. The dataset is divided into 6 groups by gene length. Then, gene expression level, GC3s and Nc values are compared among 6 different gene length groups (> = 300 bp, 2000-2999 bp, 1500-1999 bp, 1000-1499 bp, 500-999 bp, < 500 bp). CAI, GC3s and Nc values show some differences among different gene length groups. Protein hydrophobicities do not show significant influence on codon usage pattern. In summary, the natural selection on gene expression level and the base composition of genes are the major factors affecting codon usage of S. pneumoniae. Gene length shapes codon usage of S. pneumoniae in a minor way.

Amino Acids↗

Codon usage in the A/T-rich bacterium Campylobacter jejuni.

Campylobacter jejuni is a Gram negative, microaerophilic pathogen that causes gastroenteritis in humans. The genome of C. jejuni is AT-rich, with a mol% G + C of 30.4. This high AT content was hypothesized to result in unique codon usage. In the present study, we analyzed the codon usage of sixty-seven C. jejuni genes and generated a codon frequency table. As predicted, the codon usage of C. jejuni revealed a strong bias towards codons ending in A or U. In addition to determining codon usage frequencies, the relative synonymous codon usage values were calculated to identify rare and optimal codons. Seventeen codons were identified as optimal and twelve codons as rare. Thirty-two codons exhibited little or no bias. A plot of the effective number of codons versus the third position %G + C values for the sixty-seven genes revealed that C. jejuni uses an average of 39 of the 61 codons to encode proteins. These data will be useful for various molecular analyses including selection of degenerate primers to screen C. jejuni-genomic DNA libraries.

Adenine↗

Codon bias as a factor in regulating expression via translation rate in the human genome.

We study the interrelations between tRNA gene copy numbers, gene expression levels and measures of codon bias in the human genome. First, we show that isoaccepting tRNA gene copy numbers correlate positively with expression-weighted frequencies of amino acids and codons. Using expression data of more than 14,000 human genes, we show a weak positive correlation between gene expression level and frequency of optimal codons (codons with highest tRNA gene copy number). Interestingly, contrary to non-mammalian eukaryotes, codon bias tends to be high in both highly expressed genes and lowly expressed genes. We suggest that selection may act on codon bias, not only to increase elongation rate by favoring optimal codons in highly expressed genes, but also to reduce elongation rate by favoring non-optimal codons in lowly expressed genes. We also show that the frequency of optimal codons is in positive correlation with estimates of protein biosynthetic cost, and suggest another possible action of selection on codon bias: preference of optimal codons as production cost rises, to reduce the rate of amino acid misincorporation. In the analyses of this work, we introduce a new measure of frequency of optimal codons (FOP'), which is unaffected by amino acid composition and is corrected for background nucleotide content; we also introduce a new method for computing expected codon frequencies, based on the dinucleotide composition of the introns and the non-coding regions surrounding a gene.

Algorithms↗

alpha-L-iduronidase premature stop codons and potential read-through in mucopolysaccharidosis type I patients.

alpha-L-Iduronidase is a glycosyl hydrolase involved in the sequential degradation of the glycosaminoglycans heparan sulphate and dermatan sulphate. A deficiency in alpha-L-iduronidase results in the lysosomal accumulation and urinary secretion of partially degraded glycosaminoglycans and is the cause of the lysosomal storage disorder mucopolysaccharidosis type I (MPS I; Hurler and Scheie syndromes; McKusick 25280). The premature stop codons Q70X and W402X are two of the most common alpha-l-iduronidase gene (IDUA) mutations accounting for up to 70% of MPS I disease alleles in some populations. Here, we have reported a new mutation, making a total of 15 different mutations that can cause premature IDUA stop codons and have investigated the biochemistry of these mutations. Natural stop codon read-through was dependent on the fidelity of the codon when evaluated at Q70X and W402X in CHO-K1 cells, but the three possible stop codons TAA, TAG and TGA, had different effects on mRNA stability and this effect was context dependent. In CHO-K1 cells expressing the Q70X and W402X mutations, the level of gentamicin-enhanced stop codon read-through was slightly less than the increment in activity caused by a lower fidelity stop codon. In this system, gentamicin had more effect on read-through for the TAA and TGA stop codons when compared to the TAG stop codon. In an MPS I patient study, premature TGA stop codons were associated with a slightly attenuated clinical phenotype, when compared to classical Hurler syndrome (e.g. W402X/W402X and Q70X/Q70X genotypes with TAG stop codons). Natural read-through of premature stop codons is a potential explanation for variable clinical phenotype in MPS I patients. Enhanced stop codon read-through is a potential treatment strategy for a large sub-group of MPS I patients.

Animals↗

Constraints on codon context in Escherichia coli genes. Their possible role in modulating the efficiency of translation.

The constraints on nucleotide sequences of highly and weakly expressed genes from Escherichia coli have been analysed and compared. Differences in synonymous codon spectra in highly and weakly expressed genes lead to different frequencies of nucleotides (in the first and third codon positions) and dinucleotides in the two groups of genes. It has been found that the choice of synonymous codons in highly expressed genes depends on the nucleotides adjacent to the codon. For example, lysine is preferably encoded by the AAA codon if guanosine is 3' to the lysine codon (AAA-G, P less than 10(-9)). And, on the contrary, AAG is used more often than AAA (P less than 0.001) if cytidine is 3' adjacent to lysine. Guanosine occurs more frequently than adenosine 5' to all the lysine codons (AAR, P less than 10(-5), i.e. NNG codons are preferred over the synonymous NNA codons 5' to the positions of lysine in the genes. The context effect was observed in nonsense and missense suppression experiments. Therefore, a hypothesis has been suggested that the efficiency of translation of some codons (for which the constraints on the adjacent nucleotides were found) can be modulated by the codon context. The rules for preferable synonymous codon choice in highly expressed genes depending on the nucleotides surrounding the codon are presented. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Amino Acid Sequence↗