PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Improving the efficiency of the genetic code by varying the codon length--the perfect genetic code.

The function of DNA is to specify protein sequences. The four-base "alphabet" used in nucleic acids is translated to the 20 base alphabet of proteins (plus a stop signal) via the genetic code. The code is neither overlapping nor punctuated, but has mRNA sequences read in successive triplet codons until reaching a stop codon. The true genetic code uses three bases for every amino acid. The efficiency of the genetic code can be significantly increased if the requirement for a fixed codon length is dropped so that the more common amino acids have shorter codon lengths and rare amino acids have longer codon lengths. More efficient codes can be derived using the Shannon-Fano and Huffman coding algorithms. The compression achieved using a Huffman code cannot be improved upon. I have used these algorithms to derive efficient codes for representing protein sequences using both two and four bases. The length of DNA required to specify the complete set of protein sequences could be significantly shorter if transcription used a variable codon length. The restriction to a fixed codon length of three bases means that it takes 42% more DNA than the minimum necessary, and the genetic code is 70% efficient. One can think of many reasons why this maximally efficient code has not evolved: there is very little redundancy so almost any mutation causes an amino acid change. Many mutations will be potentially lethal frame-shift mutations, if the mutation leads to a change in codon length. It would be more difficult for the machinery of transcription to cope with a variable codon length. Nevertheless, in the strict and narrow sense of coding for protein sequences using the minimum length of DNA possible, the Huffman code derived here is perfect.

Algorithms↗

Codon bias in actin multigene families and effects on the reconstruction of phylogenetic relationships.

Codon usage patterns and phylogenetic relationships in the actin multigene family have been analyzed for three dipteran species--Drosophila melanogaster, Bactrocera dorsalis, and Ceratitis capitata. In certain phylogenetic tree reconstructions, using synonymous distances, some gene relationships are altered due to a homogenization phenomenon. We present evidence to show that this homogenization phenomenon is due to codon usage bias. A survey of the pattern of synonymous codon preferences for 11 actin genes from these three species reveals that five out of the six Drosophila actin genes show high degrees of codon bias as indicated by scaled chi 2 values. In contrast to this, four out of the five actin genes from the other species have low codon bias values. A Monte Carlo contingency test indicates that for those Drosophila actin genes which exhibit codon bias, the patterns of codon usage are different compared to actin genes from the other species. In addition, the genes exhibiting codon bias also appear to have reduced rates of synonymous substitution. The homogenization phenomenon seen in terms of synonymous substitutions is not observed for nonsynonymous changes. Because of this homogenization phenomenon, "trees" constructed based on synonymous substitutions will be affected. These effects can be overt in the case of multigene families, but similar distortions may underlie reconstructions based on single-copy genes which exhibit codon usage bias.

Actins↗

Essential factors determining codon usage in ubiquitin genes.

Ubiquitin is ubiquitous in all eukaryotes and its amino acid sequence shows extreme conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences coding for 13 ubiquitin genes from 11 species reported so far have been compiled and analyzed. The G + C content of codon third base reveals a positive linear correlation with the genome G + C content of the corresponding species. The slope strongly suggests that the overall G + C content of codons of polyubiquitin genes clearly reflects the genome G + C content by AT/GC substitutions at the codon third position. The G + C content of ubiquitin codon third base also shows a positive linear correlation with the overall G + C content of coding regions of compiled genes, indicating the codon choices among synonymous codons reflect the average codon usage pattern of corresponding species. On the other hand, the monoubiquitin gene, which is different from the polyubiquitin gene in gene organization, gene expression, and function of the encoding protein, shows a different codon usage pattern compared with that of the polyubiquitin gene. From comparisons of the levels of synonymous substitutions among ubiquitin repeats and the homology of the amino acid sequence of the tail of monomeric ubiquitin genes, we propose that the molecular evolution of ubiquitin genes occurred as follows: Plural primitive ubiquitin sequences were dispersed on genome in ancestral eukaryotes. Some of them situated in a particular environment fused with the tail sequence to produce monomeric ubiquitin genes that were maintained across species. After divergence of species, polyubiquitin genes were formed by duplication of the other primitive ubiquitin sequences on different chromosomes. Differences in the environments in which ubiquitin genes are embedded reflect the differences in codon choice and in gene expression pattern between poly- and monomeric ubiquitin genes.

Amino Acid Sequence↗

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids↗

The evolution of codon preferences in Drosophila: a maximum-likelihood approach to parameter estimation and hypothesis testing.

Synonymous codon usage in related species may differ as a result of variation in mutation biases, differences in the overall strength and efficiency of selection, and shifts in codon preference-the selective hierarchy of codons within and between amino acids. We have developed a maximum-likelihood method to employ explicit population genetic models to analyze the evolution of parameters determining codon usage. The method is applied to twofold degenerate amino acids in 50 orthologous genes from D. melanogaster and D. virilis. We find that D. virilis has significantly reduced selection on codon usage for all amino acids, but the data are incompatible with a simple model in which there is a single difference in the long-term Ne, or overall strength of selection, between the two species, indicating shifts in codon preference. The strength of selection acting on codon usage in D. melanogaster is estimated to be |Nes| approximately 0.4 for most CT-ending twofold degenerate amino acids, but 1.7 times greater for cysteine and 1.4 times greater for AG-ending codons. In D. virilis, the strength of selection acting on codon usage for most amino acids is only half that acting in D. melanogaster but is considerably greater than half for cysteine, perhaps indicating the dual selection pressures of translational efficiency and accuracy. Selection coefficients in orthologues are highly correlated (rho = 0.46), but a number of genes deviate significantly from this relationship.

Amino Acids↗

Codon context effects in missense suppression.

After our first observation of codon context effects in missense suppression ( Murgola & Pagel , 1983), we measured the suppression of missense mutations at two positions in trpA in Escherichia coli. The suppressible codons in the trpA messenger RNA were the lysine codons, AAA and AAG, and the glutamic acid codons, GAA and GAG. The mRNA sites of the codons correspond to amino acids 211 and 234 of the trpA polypeptide, positions at which glycine is the wild-type amino acid. Our data demonstrated codon context effects with both pairs of codons. The results indicate that suppression of AAA and AAG by mutant lysine transfer RNAs was more efficient at 211 than at 234, whereas suppression of GAA and GAG by two different mutant glycine tRNAs was more efficient at 234 than at 211. In general, the context effects were more pronounced with GAG and AAG than with GAA and AAA. (In some instances it appeared that suppression of GAA or AAA at a given position was more effective than suppression of GAG or AAG.) By contrast, no context effects were observed with a glyT suppressor of AAA and AAG, a glyT GAA/G-suppressor, and a glyU suppressor of GAG. Our observation of this phenomenon in missense suppression demonstrates that codon context can affect polypeptide elongation and that the effects can be different depending on the codons and tRNAs examined. It is suggested that tRNA-tRNA interaction on the ribosome is involved in the observed context effects.

Codon↗

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals↗

The relationship between palindrome avoidance and intragenic codon usage variations: a Monte Carlo study.

Several studies have shown that codon usage within genes varies, as it seems dependent on both codon context and codon position within the gene. Given that palindromes in addition often are avoided in genomes, this study aimed at finding out if intragenic variations in codon usage may be a way to control the amount and location of palindromes. A Monte Carlo algorithm was written which resampled the codons in genes while keeping the amino acid sequence of the translation product constant. On the resampled sequences, palindromes were counted and their intragenic positions mapped. Escherichia coli K12 uses type II restriction-modification systems and displays pronounced codon usage phenomena. Using this as a reference organism it was clearly shown that the number of palindromes in genes is generally lower than the amount of palindromes in resampled genes; thus, the succession of codons seems to be a way to decrease the number of palindromes. The intragenic position of palindromes in resampled sequences, however, was largely equal to the position in the native genes, so codon usage phenomena are unlikely to be a way to control the intragenic position of palindromes. The analysis was repeated on two bacteriophages and gave similar same results, even though the virus genomes are much smaller. Studies on the endosymbionts Buchnera sp. APS and Wigglesworthia sp., which seemingly have no type II restriction-modification systems, showed that in these species there is only weak evidence for codon usage acting to control the number of palindromes.

Algorithms↗

The 'effective number of codons' revisited.

Frank Wright [Gene 87 (1990) 23] derived a formula for calculation of a quantity termed the 'effective number of codons' (Nc) based on codon homozygosities. This quantity is a number between 20 and 61 and tells to what degree the codon usage in a gene is biased, i.e., it approaches 20 codons for the extremely biased genes, and approaches 61 for the genes where all possible codons are used with no preference. Among the different measures of codon bias Nc is considered the most useful and has found widespread use in papers dealing with codon usage phenomena. In this paper, the mathematical behaviours of codon homozygosities and Nc are evaluated, using Escherichia coli as the model organism. The results indicate that the classical formula for calculation of Nc could appropriately be substituted under circumstances, where there is bias discrepancy, i.e., when one amino acid (or more) within a degeneracy group is associated with strong codon bias while at the same time others in the same degeneracy group have little bias. An alternative estimator, termed Nc, is proposed and tested against Nc, and performs better when there is such bias discrepancy.

Codon↗

Exploring synonymous codon usage preferences of disulfide-bonded and non-disulfide bonded cysteines in the E. coli genome.

High-quality data about protein structures and their gene sequences are essential to the understanding of the relationship between protein folding and protein coding sequences. Firstly we constructed the EcoPDB database, which is a high-quality database of Escherichia coli genes and their corresponding PDB structures. Based on EcoPDB, we presented a novel approach based on information theory to investigate the correlation between cysteine synonymous codon usages and local amino acids flanking cysteines, the correlation between cysteine synonymous codon usages and synonymous codon usages of local amino acids flanking cysteines, as well as the correlation between cysteine synonymous codon usages and the disulfide bonding states of cysteines in the E. coli genome. The results indicate that the nearest neighboring residues and their synonymous codons of the C-terminus have the greatest influence on the usages of the synonymous codons of cysteines and the usage of the synonymous codons has a specific correlation with the disulfide bond formation of cysteines in proteins. The correlations may result from the regulation mechanism of protein structures at gene sequence level and reflect the biological function restriction that cysteines pair to form disulfide bonds. The results may also be helpful in identifying residues that are important for synonymous codon selection of cysteines to introduce disulfide bridges in protein engineering and molecular biology. The approach presented in this paper can also be utilized as a complementary computational method and be applicable to analyse the synonymous codon usages in other model organisms.

Amino Acids↗

Rare codon clusters at 5'-end influence heterologous expression of archaeal gene in Escherichia coli.

Proteins from hyperthermophilic microorganisms are attractive candidates for novel biocatalysts because of their high resistance to temperature extremes. However, archaeal genes are usually poorly expressed in Escherichia coli because of differences in codon usage. Genes from the thermoacidophilic archaea Sulfolobus solfataricus and Thermoplasma acidophilum contain high proportions of rare codons for arginine, isoleucine, and leucine, which are recognized by the tRNAs encoded by the argU, ileY, and leuW genes, respectively, and which are rarely used in E. coli. To examine the effects of these rare codons on heterologous expression, we expressed the Sso_gnaD and Tac_gnaD genes from S. solfataricus and T. acidophilum, respectively, in E. coli. The Sso_gnaD product was expressed at very low levels when the open reading frame (ORF) was cloned in pRSET and expressed in E. coli BL21(DE3), and was expressed at much higher levels in the E. coli BL21(DE3)-CodonPlus RIL strain, which contains extra copies of the argU, ileY, and leuW tRNA genes. In contrast, Tac_gnaD was expressed at similar levels in both E. coli strains. Comparison of the Sso_gnaD and Tac_gnaD gene sequences revealed that the 5'-end of the Sso_gnaD sequence was rich in AGA(arg) and ATA(Ile) codons. These codons were replaced with the codons commonly used in E. coli by polymerase chain reaction-mediated site-directed mutagenesis. The results of expression studies showed that a non-tandem repeat of rare codons is critical in the observed interference in heterologous expression of this gene. We concluded that the level of heterologous expression of Sso_gnaD in E. coli was limited by the clustering of the rare codons in the ORF, rather than on the rare codon frequency.

Codon↗

Molecular mechanism of stop codon recognition by eRF1: a wobble hypothesis for peptide anticodons.

We propose that the amino acid residues 57/58 and 60/61 of eukaryotic release factors (eRF1s) (counted from the N-terminal Met of human eRF1) are responsible for stop codon recognition in protein synthesis. The proposal is based on amino acid exchanges in these positions in the eRF1s of two ciliates that reassigned one or two stop codons to sense codons in evolution and on the crystal structure of human eRF1. The proposed mechanism of stop codon recognition assumes that the amino acid residues 57/58 interact with the second and the residues 60/61 with the third position of a stop codon. The fact that conventional eRF1s recognize all three stop codons but not the codon for tryptophan is attributed to the flexibility of the helix containing these residues. We suggest that the helix is able to assume a partly relaxed or tight conformation depending on the stop codon recognized. The restricted codon recognition observed in organisms with unconventional eRF1s is attributed mainly to the loss of flexibility of the helix due to exchanged amino acids.

Amino Acid Sequence↗

Synonymous codon usage in Cryptosporidium parvum: identification of two distinct trends among genes.

The usage of alternative synonymous codons in the apicomplexan Cryptosporidium parvum has been investigated. A data set of 54 genes was analysed. Overall, A- and U-ending codons predominate, as expected in an A+T-rich genome. Two trends of codon usage variation among genes were identified using correspondence analysis. The primary trend is in the extent of usage of a subset of presumably translationally optimal codons, that are used at significantly higher frequencies in genes expected to be expressed at high levels. Fifteen of the 18 codons identified as optimal are more G+C-rich than the otherwise common codons, so that codon selection associated with translation opposes the general mutation bias. Among 40 genes with lower frequencies of these optimal codons, a secondary trend in G+C content was identified. In these genes, G+C content at synonymously variable third positions of codons is correlated with that in 5' and 3' flanking sequences, indicative of regional variation in G+C content, perhaps reflecting regional variation in mutational biases.

Animals↗

Frequencies of codons in histones, tubulins and fibrinogen: bias due to interference between transcription signals and protein function.

The distribution of codons was studied in 65 proteins: 48 histones, 14 tubulins, and three fibrinogens, With the methodology used, (1) we confirmed that the preterminator state of a codon has no detectable effect on codon bias. (2) The well-known effect of CG suppression was visible. We also found that (3) some codons which are very rare, are equal to parts of known transcription signals. Thus, we advanced that to avoid signal interference, the use of these codons is suppressed when a synonymous codon is available. In addition we found that in the whole series of codons, transcription signals are less frequent than in a random sequence of equal composition. Finally we observed (4) that tryptophan is absent in histones. This absence was related not to the TGG codon itself, but to characteristics of the amino acid. We conclude that the functional constraints of a protein can influence, at least for synonymous codon usage, the evolution of its own coding sequence.

Animals↗

Thermophilic prokaryotes have characteristic patterns of codon usage, amino acid composition and nucleotide content.

A number of recent studies have shown that thermophilic prokaryotes have distinguishable patterns of both synonymous codon usage and amino acid composition, indicating the action of natural selection related to thermophily. On the other hand, several other studies of whole genomes have illustrated that nucleotide bias can have dramatic effects on synonymous codon usage and also on the amino acid composition of the encoded proteins. This raises the possibility that the thermophile-specific patterns observed at both the codon and protein levels are merely reflections of a single underlying effect at the level of nucleotide composition. Moreover, such an effect at the nucleotide level might be due entirely to mutational bias. In this study, we have compared the genomes of thermophiles and mesophiles at three levels: nucleotide content, codon usage and amino acid composition. Our results indicate that the genomes of thermophiles are distinguishable from mesophiles at all three levels and that the codon and amino acid frequency differences cannot be explained simply by the patterns of nucleotide composition. At the nucleotide level, we see a consistent tendency for the frequency of adenine to increase at all codon positions within the thermophiles. Thermophiles are also distinguished by their pattern of synonymous codon usage for several amino acids, particularly arginine and isoleucine. At the protein level, the most dramatic effect is a two-fold decrease in the frequency of glutamine residues among thermophiles. These results indicate that adaptation to growth at high temperature requires a coordinated set of evolutionary changes affecting (i) mRNA thermostability, (ii) stability of codon-anticodon interactions and (iii) increased thermostability of the protein products. We conclude that elevated growth temperature imposes selective constraints at all three molecular levels: nucleotide content, codon usage and amino acid composition. In addition to these multiple selective effects, however, the genomes of both thermophiles and mesophiles are often subject to superimposed large changes in composition due to mutational bias.

Amino Acids↗

Compositional pressure and translational selection determine codon usage in the extremely GC-poor unicellular eukaryote Entamoeba histolytica.

It is widely accepted that the compositional pressure is the only factor shaping codon usage in unicellular species displaying extremely biased genomic compositions. This seems to be the case in the prokaryotes Mycoplasma capricolum, Rickettsia prowasekii and Borrelia burgdorferi (GC-poor), and in Micrococcus luteus (GC-rich). However, in the GC-poor unicellular eukaryotes Dictyostelium discoideum and Plasmodium falciparum, there is evidence that selection, acting at the level of translation, influences codon choices. This is a twofold intriguing finding, since (1) the genomic GC levels of the above mentioned eukaryotes are lower than the GC% of any studied bacteria, and (2) bacteria usually have larger effective population sizes than eukaryotes, and hence natural selection is expected to overcome more efficiently the randomizing effects of genetic drift among prokaryotes than among eukaryotes. In order to gain a new insight about this problem, we analysed the patterns of codon preferences of the nuclear genes of Entamoeba histolytica, a unicellular eukaryote characterised by an extremely AT-rich genome (GC = 25%). The overall codon usage is strongly biased towards A and T in the third codon positions, and among the presumed highly expressed sequences, there is an increased relative usage of a subset of codons, many of which are C-ending. Since an increase in C in third codon positions is 'against' the compositional bias, we conclude that codon usage in E. histolytica, as happens in D. discoideum and P. falciparum, is the result of an equilibrium between compositional pressure and selection. These findings raise the question of why strongly compositionally biased eukaryotic cells may be more sensitive to the (presumed) slight differences among synonymous codons than compositionally biased bacteria.

Animals↗

Codon adaptation and synonymous substitution rate in diatom plastid genes.

Diatom plastid genes are examined with respect to codon adaptation and rates of silent substitution (Ks). It is shown that diatom genes follow the same pattern of codon usage as other plastid genes studied previously. Highly expressed diatom genes display codon adaptation, or a bias toward specific major codons, and these major codons are the same as those in red algae, green algae, and land plants. It is also found that there is a strong correlation between Ks and variation in codon adaptation across diatom genes, providing the first evidence for such a relationship in the algae. It is argued that this finding supports the notion that the correlation arises from selective constraints, not from variation in mutation rate among genes. Finally, the diatom genes are examined with respect to variation in Ks among different synonymous groups. Diatom genes with strong codon adaptation do not show the same variation in synonymous substitution rate among codon groups as the flowering plant psbA gene which, previous studies have shown, has strong codon adaptation but unusually high rates of silent change in certain synonymous groups. The lack of a similar finding in diatoms supports the suggestion that the feature is unique to the flowering plant psbA due to recent relaxations in selective pressure in that lineage.

Adaptation, Physiological↗

Aminoglycoside antibiotics mediate context-dependent suppression of termination codons in a mammalian translation system.

The translation machinery recognizes codons that enter the ribosomal A site with remarkable accuracy to ensure that polypeptide synthesis proceeds with a minimum of errors. When a termination codon enters the A site of a eukaryotic ribosome, it is recognized by the release factor eRF1. It has been suggested that the recognition of translation termination signals in these organisms is not limited to a simple trinucleotide codon, but is instead recognized by an extended tetranucleotide termination signal comprised of the stop codon and the first nucleotide that follows. Interestingly, pharmacological agents such as aminoglycoside antibiotics can reduce the efficiency of translation termination by a mechanism that alters this ribosomal proofreading process. This leads to the misincorporation of an amino acid through the pairing of a near-cognate aminoacyl tRNA with the stop codon. To determine whether the sequence context surrounding a stop codon can influence aminoglycoside-mediated suppression of translation termination signals, we developed a series of readthrough constructs that contained different tetranucleotide termination signals, as well as differences in the three bases upstream and downstream of the stop codon. Our results demonstrate that the sequences surrounding a stop codon can play an important role in determining its susceptibility to suppression by aminoglycosides. Furthermore, these distal sequences were found to influence the level of suppression in remarkably distinct ways. These results suggest that the mRNA context influences the suppression of stop codons in response to subtle differences in the conformation of the ribosomal decoding site that result from aminoglycoside binding.

Animals↗