PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Chloroplast genes transferred to the nuclear plant genome have adjusted to nuclear base composition and codon usage.

During plant evolution, some plastid genes have been moved to the nuclear genome. These transferred genes are now correctly expressed in the nucleus, their products being transported into the chloroplast. We compared the base compositions, the distributions of some dinucleotides and codon usages of transferred, nuclear and chloroplast genes in two dicots and two monocots plant species. Our results indicate that transferred genes have adjusted to nuclear base composition and codon usage, being now more similar to the nuclear genes than to the chloroplast ones in every species analyzed.

Base Composition↗

A backtranslation method based on codon usage strategy.

This study describes a method for the backtranslation of an aminoacidic sequence, an extremely useful tool for various experimental approaches. It involves two computer programs CLUSTER and BACKTR written in Fortran 77 running on a VAX/VMS computer. CLUSTER generates a reliable codon usage table through a cluster analysis, based on a chi 2-like distance between the sequences. BACKTR produces backtranslated sequences according to different options when use is made of the codon usage table obtained in addition to selecting the least ambiguous potential oligonucleotide probes within an aminoacidic sequence. The method was tested by applying it to 158 yeast genes.

Amino Acid Sequence↗

Subrepeats within the BR1 beta repeat unit in Chironomus pallidivittatus can be classified into different types depending on codon usage.

A new type of repeat unit was isolated from Balbiani ring 1 of Chironomus pallidivittatus and designated BR1 beta repeat. It consists of a constant and a subrepeated part, like previously described units belonging to the core blocks of the BR genes. The subrepeated part contains 10-codon subrepeats with an arrangement similar to the subrepeats of the previously described BR2 beta gene. The present unit differs from earlier reported core units firstly in a much lower number of copies (about 15) per genome, which are tandemly arranged. Secondly, the number of subrepeats per BR1 beta repeat unit can show great variations. On the basis of the pattern of codon usage, three types of subrepeats can be distinguished. One type lies 5'-proximal in the subrepeat array and consists of variable numbers of subrepeats almost identical at the nucleotide level. The last complete subrepeat represents another type, with consistent differences in codon usage as compared to subrepeats of the proximal type. Finally, there is an intermediate type represented by the subrepeat preceding the distal one. Here, codon characteristics from proximal and distal subrepeats are mixed in a patchy and irregular way. The evolution of the arrays can be understood either as being the result of subrepeat formation in two steps (occurring before and after amplification of whole repeat units) or as the result of a continuous process in which there is evidence for participation of gene conversion.

Amino Acid Sequence↗

Novel anticodon composition of transfer RNAs in Micrococcus luteus, a bacterium with a high genomic G + C content. Correlation with codon usage.

The number and relative amount of isoacceptor tRNAs for each amino acid in Micrococcus luteus, a Gram-positive bacterium with high genomic G + C content, have been determined by sequencing their anticodon loop and its adjacent regions and by selective labelling of tRNAs. Thirty-one tRNA species with 29 different anticodon sequences have been detected. All the tRNAs have G or C at the anticodon first position except for tRNA(ICGArg) and tRNA(NGASer), in response to the abundant usage of NNC and NNG codons. No tRNA with the anticodon UNN capable of translating codon NNA has been detected, in accordance with a very low or zero usage of NNA codons. The relative amount of isoacceptor tRNAs for an amino acid determined by selective labelling strongly correlates with usage of the corresponding codons. On the basis of these and other observations in this and other eubacterial species, we conclude that the relative amount and anticodon composition of isoacceptor tRNA species are flexible, and their changes are mainly adaptive phenomena that have been primarily affected by codon usage, which in turn is affected by directional mutation pressure.

Anticodon↗

Codon usage in the vertebrate hemoglobins and its implications.

A study of codon usage in vertebrate hemoglobins revealed an evolutionary trend toward elevated numbers of CpG codon boundary pairs in mammalian hemoglobin alpha genes. Selection for CpG codon boundaries countering the generally observed CpG suppression is strongly suggested by these data. These observations parallel recently published experimental results that indicate that constitutive expression of the human alpha-globin gene appears to be determined by regulatory information encoded within the structural gene. The possibility is raised that, in the absence of selection, CpG decay can be used to date the evolutionary origin of a mammalian alpha pseudogene from its active alpha gene.

Animals↗

Nucleotide sequence of the Caulobacter crescentus flaF and flbT genes and an analysis of codon usage in organisms with G + C-rich genomes.

The Caulobacter crescentus flaFG region encodes trans-acting, regulatory factors that modulate flagellin synthesis during flagellum biogenesis. In this study, sequence analysis and experiments utilizing a promoterless cat gene demonstrated that the flaF and flbT genes have overlapping transcripts with the same orientation. In addition, the 5' ends of the flgL and flbA genes were located. A sequence resembling an Rho-factor-independent terminator was found in the 3' region of the flaF gene. This region was uniquely A + T-rich and the encoded mRNA contained an inverted repeat sequence which could form a stable stem-loop structure followed by nine U-residues. The codon usage of C. crescentus genes was examined and indicated a preference for specific codons from each of the synonymous codon groups. Furthermore, comparison to the codon usage of other organisms with G + C-rich genomes indicated a strong preference for the same codons preferred by C. crescentus.

Amino Acid Sequence↗

High-level expression in mammalian cells of recombinant house dust mite allergen ProDer p 1 with optimized codon usage.

BACKGROUND: The major house dust mite allergen Der p 1 is associated with allergic diseases such as asthma. Production of recombinant Der p 1 was previously attempted, but with limited success. The present study describes the expression of recombinant (rec) ProDer p 1, a recombinant precursor form of Der p 1, in CHO cells. METHODS: As optimization of the codon usage may allow successful overexpression of protein in mammalian cells, a synthetic gene encoding ProDer p 1 was designed on the basis of the codon usage frequently found in highly expressed human genes. Gene synthesis was accomplished from a set of 14 mutually priming overlapping oligonucleotides and after two runs of polymerase chain reaction. RESULTS: COS cells transiently transfected with the synthetic ProDer p 1 gene produced up to 5--10 times as much ProDer p 1 compared with the expression level obtained after transfection with the authentic gene. To stably express the recombinant allergen, CHO-K1 cells were transfected with the ProDer p 1 synthetic gene, and one amplified recombinant clone produced up to 30 mg of recProDer p 1 per liter in the culture medium before purification. recProDer p 1 was secreted as an enzymatically inactive single-chain molecule presenting three glycosylated immunoreactive forms of 41, 38 and 36 kD. When examined with respect to direct binding, recProDer p 1 and natural Der p 1 displayed very similar IgE reactivities. However, IgE inhibition and histamine release assays showed a much higher reactivity to natural Der p 1 compared to recProDer p 1. CONCLUSIONS: These data indicated that codon optimization represents an attractive strategy for high-level production of allergen in mammalian cells.

Amino Acid Sequence↗

Yeast regulatory gene PPR1. I. Nucleotide sequence, restriction map and codon usage.

The PPR1 gene of Saccharomyces cerevisiae controls the transcription of two unlinked structural genes URA1 and URA3. The primary structure of this eukaryotic regulatory gene and its flanking regions has been established by the dideoxynucleotide chain termination method. Our data show an open reading frame of 2712 nucleotides, corresponding to 904 amino acid residues. The 3' untranslated messenger RNA region presents consensus yeast termination and polyadenylation sequences. The pattern of codon usage in the gene is clearly random. This result is discussed in relation to protein abundance and is compared with the codon usage in 20 yeast structural and regulatory genes and with that found for Escherichia coli genes.

Amino Acid Sequence↗

Essential factors determining codon usage in ubiquitin genes.

Ubiquitin is ubiquitous in all eukaryotes and its amino acid sequence shows extreme conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences coding for 13 ubiquitin genes from 11 species reported so far have been compiled and analyzed. The G + C content of codon third base reveals a positive linear correlation with the genome G + C content of the corresponding species. The slope strongly suggests that the overall G + C content of codons of polyubiquitin genes clearly reflects the genome G + C content by AT/GC substitutions at the codon third position. The G + C content of ubiquitin codon third base also shows a positive linear correlation with the overall G + C content of coding regions of compiled genes, indicating the codon choices among synonymous codons reflect the average codon usage pattern of corresponding species. On the other hand, the monoubiquitin gene, which is different from the polyubiquitin gene in gene organization, gene expression, and function of the encoding protein, shows a different codon usage pattern compared with that of the polyubiquitin gene. From comparisons of the levels of synonymous substitutions among ubiquitin repeats and the homology of the amino acid sequence of the tail of monomeric ubiquitin genes, we propose that the molecular evolution of ubiquitin genes occurred as follows: Plural primitive ubiquitin sequences were dispersed on genome in ancestral eukaryotes. Some of them situated in a particular environment fused with the tail sequence to produce monomeric ubiquitin genes that were maintained across species. After divergence of species, polyubiquitin genes were formed by duplication of the other primitive ubiquitin sequences on different chromosomes. Differences in the environments in which ubiquitin genes are embedded reflect the differences in codon choice and in gene expression pattern between poly- and monomeric ubiquitin genes.

Amino Acid Sequence↗

Codon usage patterns among genes for lepidopteran hemolymph proteins.

Patterns in codon usage were examined for the coding regions of the 23 known lepidopteran hemolymph proteins. Coding triplets are GC rich at the third position and a significant linear relationship between GC content of silent and nonsilent (replacement) sites was demonstrated. Intron GC content was significantly lower than in coding regions and no relationship between intron GC content and the same at silent and nonsilent sites was found. Though hemolymph proteins are all produced by the same tissue--fat body--significantly less bias was observed when all moth sequences were pooled than when sequences of the two major species were analyzed separately, as predicted by the genome hypothesis. In cases where no statistically significant bias was observed, polar or acidic/basic amino acids were almost exclusively involved. Calculation of codon adaptation indices (CAI) was of limited value in quantifying the degree of codon bias and probably reflects the complexity of multicellular-organism life cycles and the changing patterns of gene expression over different developmental stages.

Animals↗

Error minimization explains the codon usage of highly expressed genes in Escherichia coli.

Different organisms use synonymous codons with different preferences. Several measures have been introduced to compute the extent of codon usage bias within a gene or genome, among which the codon adaptation index (CAI) has been shown to be well correlated with mRNA levels of Escherichia coli. In this work an error adaptation index (eAI) is introduced, which estimates the level at which a gene can tolerate the effects of mistranslations. It is shown that the eAI has a strong correlation with CAI, as well as with mRNA levels, which suggests that the codons of highly expressed genes are selected so that mistranslation would have the minimum possible effect on the structure and function of the related proteins.

Base Composition↗

Codon usage in histone gene families of higher eukaryotes reflects functional rather than phylogenetic relationships.

The nucleic acid sequences coding for 23 H3 histone genes from a variety of species have been analyzed using a computer assisted alignment and analysis program. Although these histones are highly conserved within and between highly divergent species, they represent various classes of histones whose patterns of expression are distinctively regulated. Surprisingly, in dendrograms derived from these comparisons, H3 sequences cluster according to their modes of regulation rather than phylogenetically. These clusters are generated from highly distinctive patterns of codon usage within the functional gene classes. We suggest that one factor involved in specifying the differing codon usage patterns between functional classes is a difference in requirements for rapid translation of mRNA. In addition, the data presented here, together with structural and sequence information, suggest a heterodox evolutionary model in which genes related to the intron-bearing, basally expressed H3.3 vertebrate genes are the ancestors of the intronless H3.1 class of genes of higher eukaryotes. The H3.1 class must have arisen, therefore, following duplication of a primitive H3.3 gene, but prior to the plant-animal divergence. Implications of the data presented are discussed with regard to functional and evolutionary relationships.

Amino Acid Sequence↗

Synonymous codon usage bias and the expression of human glucocerebrosidase in the methylotrophic yeast, Pichia pastoris.

The lysosomal hydrolase glucocerebrosidase catalyzes the penultimate step in the breakdown of membrane glycosphingolipids. An inherited deficiency in this enzyme leads to the onset of Gaucher disease, the most common lysosomal storage disorder. Exogenous sources of this protein are required for biochemical and biophysical investigations and enzyme replacement therapy of Gaucher disease. Heterologous expression of glucocerebrosidase has been successful in mammalian and insect cell lines and although its use in enzyme replacement therapy of Gaucher disease has proven efficacious, current production levels limit the availability of the enzyme. Initial attempts to express human glucocerebrosidase using the methylotrophic yeast Pichia pastoris had limited success, despite significant levels of transcription. Using fragments of the glucocerebrosidase cDNA fused to the luciferase cDNA as a translational read-through reporter, the impact of synonymous codon usage bias on protein expression in P. pastoris was examined. A table of preferred codons was determined for P. pastoris and the codon usage of a 186-bp fragment of the glucocerebrosidase gene was optimized to that of the P. pastoris preferred set. A second construct with altered G+C content but no codon optimization was created for comparison. While the native glucocerebrosidase coding region limited luciferase activity to baseline levels, the codon optimized and G+C altered constructs increased luciferase activity 10.6- and 7.5-fold, respectively. Optimized G+C content, regardless of corresponding codon optimization, appears to be the major contributor to increased translational efficiency in this heterologous expression host.

Amino Acid Sequence↗

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence↗

Codon usage in plastid genes is correlated with context, position within the gene, and amino acid content.

Highly expressed plastid genes display codon adaptation, which is defined as a bias toward a set of codons which are complementary to abundant tRNAs. This type of adaptation is similar to what is observed in highly expressed Escherichia coli genes and is probably the result of selection to increase translation efficiency. In the current work, the codon adaptation of plastid genes is studied with regard to three specific features that have been observed in E. coli and which may influence translation efficiency. These features are (1) a relatively low codon adaptation at the 5' end of highly expressed genes, (2) an influence of neighboring codons on codon usage at a particular site (codon context), and (3) a correlation between the level of codon adaptation of a gene and its amino acid content. All three features are found in plastid genes. First, highly expressed plastid genes have a noticeable decrease in codon adaptation over the first 10-20 codons. Second, for the twofold degenerate NNY codon groups, highly expressed genes have an overall bias toward the NNC codon, but this is not observed when the 3' neighboring base is a G. At these sites highly expressed genes are biased toward NNT instead of NNC. Third, plastid genes that have higher codon adaptations also tend to have an increased usage of amino acids with a high G + C content at the first two codon positions and GNN codons in particular. The correlation between codon adaptation and amino acid content exists separately for both cytosolic and membrane proteins and is not related to any obvious functional property. It is suggested that at certain sites selection discriminates between nonsynonymous codons based on translational, not functional, differences, with the result that the amino acid sequence of highly expressed proteins is partially influenced by selection for increased translation efficiency.

Amino Acids↗

Evolution of relative synonymous codon usage in Human Immunodeficiency Virus type-1.

Mutation in Human Immunodeficiency Virus type-1 (HIV-1) is extremely rapid, a consequence of a low-fidelity viral reverse transcription process. The envelope gene has been shown to accumulate substitutions at a rate of approximately 1% per year and can frequently spend a long time in the host (approximately 10 years). The relative synonymous codon usage (RSCU) in HIV-1 is known to be different from that of the human host. However, by reengineering the protein coding sequences of HIV-1 to reflect the RSCU patterns observed in humans, a large increase in protein expression is observed. It is reasonable to suggest that within a host there may be a selective drive for change in the RSCU of HIV-1 towards human RSCU. To test this hypothesis we analyzed HIV-1 partial envelope sequences from eight patients sampled serially in time. For each sequence, an RSCU table was constructed. Sequences were labelled as "early" or "late" depending on whether they were sampled before or after the mid-point of the study. Using the RSCU values as descriptor variables, a Principal Components Analysis (PCA) was performed. The first three components clearly discriminated between early and late sequences. We also constructed pooled groupwise RSCU tables for early and late sequences. The viral RSCU values of each of the groups were correlated with human RSCU. If there is selection for host-adaptation in RSCU, we expect that "late" viral RSCUs would tend to be more highly correlated with human RSCU than "early" viral RSCUs. In fact, tests of significance suggest that this is the case. However, closer examination of the data revealed that the apparent trend towards human RSCU can be attributed to the homogenization of the codon usage by mutation pressure rather than host adaptation.

Algorithms↗

Organization and codon usage of the streptomycin operon in Micrococcus luteus, a bacterium with a high genomic G + C content.

The DNA sequence of the Micrococcus luteus str operon, which includes genes for ribosomal proteins S12 (str or rpsL) and S7 (rpsG) and elongation factors (EF) G (fus) and Tu (tuf), has been determined and compared with the corresponding sequence of Escherichia coli to estimate the effect of high genomic G + C content (74%) of M. luteus on the codon usage pattern. The gene organization in this operon and the deduced amino acid sequence of each corresponding protein are well conserved between the two species. The mean G + C content of the M. luteus str operon is 67%, which is much higher than that of E. coli (51%). The codon usage pattern of M. luteus is very different from that of E. coli and extremely biased to the use of G and C in silent positions. About 95% (1,309 of 1,382) of codons have G or C at the third position. Codon GUG is used for initiation of S12, EF-G, and EF-Tu, and AUG is used only in S7, whereas GUG initiates only one of the EF-Tu's in E. coli. UGA is the predominant termination codon in M. luteus, in contrast to UAA in E. coli.

Amino Acid Sequence↗

Codon usage decreases the error minimization within the genetic code.

The genetic code is not random but instead is organized in such a way that single nucleotide substitutions are more likely to result in changes between similar amino acids. This fidelity, or error minimization, has been proposed to be an adaptation within the genetic code. Many models have been proposed to measure this adaptation within the genetic code. However, we find that none of these consider codon usage differences between species. Furthermore, use of different indices of amino acid physicochemical characteristics leads to different estimations of this adaptation within the code. In this study, we try to establish a more accurate model to address this problem. In our model, a weighting scheme is established for mistranslation biases of the three different codon positions, transition/transversion biases, and codon usage. Different indices of amino acids' physicochemical characteristics are also considered. In contrast to pervious work, our results show that the natural genetic code is not fully optimized for error minimization. The genetic code, therefore, is not the most optimized one for error minimization, but one that balances between flexibility and fidelity for different species.

Amino Acid Substitution↗