PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Comparison and evolutionary analysis of the glycosomal glyceraldehyde-3-phosphate dehydrogenase from different Kinetoplastida.

In this work, we present the sequences and a comparison of the glycosomal GAPDHs from a number of Kinetoplastida. The complete gene sequences have been determined for some species (Crithidia fasciculata, Herpetomonas samuelpessoai, Leptomonas seymouri, and Phytomonas sp), whereas for other species (Trypanosoma brucei gambiense, Trypanosoma congolense, Trypanosoma vivax, and Leishmania major), only partial sequences have been obtained by PCR amplification. The structure of all available glycosomal GAPDH genes was analyzed in detail. Considerable variations were observed in both their nucleotide composition and their codon usage. The GC content varies between 64.4% in L. seymouri and 49.5% in the previously sequenced GAPDH gene from Trypanoplasma borreli. A highly biased codon usage was found in C. fasciculata, with only 34 triplets used, whereas in T. borreli 57 codons were employed. No obvious correlation could be observed between the codon usage and either the nucleotide composition or the level of gene expression. The glycosomal GAPDH is a very well-conserved enzyme. The maximal overall difference observed in the amino acid sequences is only 25%. Specific insertions and extensions are retained in all sequences. The residues involved in catalysis, substrate, and inorganic phosphate binding are fully conserved, whereas some variability is observed in the cofactor-binding pocket. The implications of these data for the design of new trypanocidal drugs targeted against GAPDH are discussed. All available gene and amino acid sequences of glycosomal GAPDHs were used for a phylogenetic analysis. The division of the Kinetoplastida into two suborders, Bodonina and Trypanosomatina, was well supported. Within the letter group, the Trypanosoma species appeared to be monophyletic, whereas the other trypanosomatids form a second clade.

Amino Acid Sequence↗

Over-representation of Chi sequences caused by di-codon increase in Escherichia coli K-12.

Chi sequences (5'-GCTGGTGG-3') are cis-acting 8 bp sequence elements that enhance homologous recombination promoted by the RecBCD pathway in Escherichia coli. The genome of E. coli K-12 MG1655 contains 1009 Chi sequences and this frequency far exceeds the expected value for occurrence of an 8 bp sequence in a genome of this size. It is generally thought that the over-representation of Chi sequences indicates that they have been selected for during evolution because of their function in recombination. The genes from three E. coli strains (K-12, O157 and CFT) were classified into three categories (island, match to other E. coli, and backbone). Island genes have a different base composition and codon usage in comparison with those in the backbone genes, therefore they were relatively new and not yet adapted to the base composition patterns and codon usage typical of the recipient genome. The over-representation of Chi sequences was examined by comparing Chi frequencies and codon frequencies between island and backbone genes. The difference in the CTGGTG di-codon frequency between the backbone and island genes was correlated with the frequency of Chi sequences which were translated in the Leu-Val (-G/CTG/GTG/G-) reading frame in the K-12 strain. These results suggest that the main reading frame of Chi sequences increased as a result of the di-codon CTG-GTG increasing under a genome-wide pressure for adapting to the codon usage and base composition of the E. coli K-12 strain, and that the RecBCD recombinase might adjust its recognition sequence to a frequently occurring oligomer such as G-CTG-GTG-G.

Base Composition↗

Expression of the green fluorescent protein in Paramecium tetraurelia.

In this paper we describe the expression of green fluorescent protein (GFP) as a reporter in vivo to monitor transformation in Paramecium cells. This is not trivial because of the limited number of strong promoters available for heterologous expression and the very high AT content of the genomic DNA, the consequence of which is a very aberrant codon usage. Taking into account differences in codon usage we selected and modified the original GFP open reading frame (ORF) from Aequorea victoria and placed the altered ORF into the Paramecium expression vector pPXV. Injection of the linearized plasmid into the macronucleus resulted in a cytoplasmic fluorescence signal in the clonal descendants, which was proportional to the number of copies injected. Southern hybridization indicated the establishment and replication of the plasmid during vegetative growth. Expression was also monitored by Northern and Western analysis. The results indicate that the modified GFP can be used in Paramecium as a reporter for transformation as an alternative to selection with antibiotics and that it may also be used to construct and localize fusion proteins.

Animals↗

Genetic analysis of Carboxydothermus hydrogenoformans carbon monoxide dehydrogenase genes cooF and cooS.

Carboxydothermus hydrogenoformans is an extremely thermophilic, Gram-positive bacterium growing on carbon monoxide (CO) as single carbon and energy source and producing only H(2) and CO(2). Carbon monoxide dehydrogenase is a key enzyme for CO metabolism. The carbon monoxide dehydrogenase genes cooF and cooS from C. hydrogenoformans were cloned and sequenced. These genes showed the highest similarity to the cooF genes from the archaeon Archaeoglobus fulgidus and the cooS gene from the bacterium Rhodospirillum rubrum, respectively. The cooS gene was identified immediately downstream of cooF, however, the cooF and cooS genes from C. hydrogenoformans have substantially different codon usage, and the cooF gene Arg codon usage pattern, dominated by AGA and AGG, resembles the archaeal pattern. The data therefore suggest lateral transfer of these genes, possibly from different donor species.

Aldehyde Oxidoreductases↗

A novel method of analyzing proline synonymous codons in E. coli.

Proline is a special imino acid in protein and the isomerization of the prolyl peptide bond has notable biological significance and influences the final structure of protein greatly, so the correlation between proline synonymous codon usage and local amino acid, the correlation between proline synonymous codon usage and the isomerization of the prolyl peptide bond were both investigated in the Escherichia coli genome by using a novel method based on information theory. The results show that in peptide chain, the residue at the first position C-terminal influences the usage of proline synonymous codon greatly and proline synonymous codons contain some factors influencing the isomerization of the prolyl peptide bond.

Codon↗

Correlations between the coding and non-coding regions in DNA.

In this paper various aspects of codon usage and k-tuple correlations in the DNA are compared. It is shown that the correlation structures of the coding and the non-coding regions are very similar and that codon usage is reasonably specific for large groups of organisms. These results suggest that the origin of codon usage is related to the origin and structure of the DNA.

Amino Acid Sequence↗

Disparate sequence characteristics of the Erysiphe graminis f.sp. hordei glyceraldehyde-3-phosphate dehydrogenase gene.

The Erysiphe graminis f.sp. hordei (Egh) glyceraldehyde-3-phosphate dehydrogenase (gpd) gene was isolated and characterized. It contains typical promoter elements and has three introns, one of which is positioned in the 5' untranslated region of the gene. The deduced amino-acid sequence has 87% similarity to gpd genes from other Ascomycete fungi. This is at the same level as previously estimated among these fungi. Comparison at the DNA level reveal similarities of only around 70%, which is 10% lower than previously reported. In an evolutionary tree based on the sequences from 18 fungal gpd genes, Egh falls into the group of Ascomycetes located at a basal position. The regulatory region of the Egh gpd gene has no homology to corresponding sequences in other filamentous Ascomycetes. Codon usage was determined for the four characterized Egh genes (tub2, Egh7, Egh16 and gpd) and found to be similar for all four genes. The results of the codon-usage analysis suggest that Egh is more flexible than other fungi in the choice of nucleotides at the wobble position. Codon-usage preferences in Egh and barley genes indicate a level of difference which may be exploited to discriminate between fungal and plant genes in sequence mixtures. The Egh gpd promoter appears to be superior to that of the Egh beta-tubulin gene (tub2) for driving the E. coli beta-glucuronidase (GUS) gene in transformation experiments.

Ascomycota↗

Substitution rates in Drosophila nuclear genes: implications for translational selection.

The relationships between synonymous and nonsynonymous substitution rates and between synonymous rate and codon usage bias are important to our understanding of the roles of mutation and selection in the evolution of Drosophila genes. Previous studies used approximate estimation methods that ignore codon bias. In this study we reexamine those relationships using maximum-likelihood methods to estimate substitution rates, which accommodate the transition/transversion rate bias and codon usage bias. We compiled a sample of homologous DNA sequences at 83 nuclear loci from Drosophila melanogaster and at least one other species of Drosophila. Our analysis was consistent with previous studies in finding that synonymous rates were positively correlated with nonsynonymous rates. Our analysis differed from previous studies, however, in that synonymous rates were unrelated to codon bias. We therefore conducted a simulation study to investigate the differences between approaches. The results suggested that failure to properly account for multiple substitutions at the same site and for biased codon usage by approximate methods can lead to an artifactual correlation between synonymous rate and codon bias. Implications of the results for translational selection are discussed.

Animals↗

The action of selection on codon bias in the human genome is related to frequency, complexity, and chronology of amino acids.

BACKGROUND: The question of whether synonymous codon choice is affected by cellular tRNA abundance has been positively answered in many organisms. In some recent works, concerning the human genome, this relation has been studied, but no conclusive answers have been found. In the human genome, the variation in base composition and the absence of cellular tRNA count data makes the study of the question more complicated. In this work we study the relation between codon choice and tRNA abundance in the human genome by correcting relative codon usage for background base composition and using a measure based on tRNA-gene copy numbers as a rough estimate of tRNA abundance. RESULTS: We term major codons to be those codons with a relatively large tRNA-gene copy number for their corresponding amino acid. We use two measures of expression: breadth of expression (the number of tissues in which a gene was expressed) and maximum expression level among tissues (the highest value of expression of a gene among tissues). We show that for half the amino acids in the study (8 of 16) the relative major codon usage rises with breadth of expression. We show that these amino acids are significantly more frequent, are smaller and simpler, and are more ancient than the rest of the amino acids. Similar, although weaker, results were obtained for maximum expression level. CONCLUSION: There is evidence that codon bias in the human genome is related to selection, although the selection forces acting on codon bias may not be straightforward and may be different for different amino acids. We suggest that, in the first group of amino acids, selection acts to enhance translation efficiency in highly expressed genes by preferring major codons, and acts to reduce translation rate in lowly expressed genes by preferring non-major ones. In the second group of amino acids other selection forces, such as reducing misincorporation rate of expensive amino acids, in terms of their size/complexity, may be in action. The fact that codon usage is more strongly related to breadth of expression than to maximum expression level supports the notion, presented in a recent study, that codon choice may be related to the tRNA abundance in the tissue in which a gene is expressed.

Amino Acids↗

Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages.

In the plant chloroplast genome the codon usage of the highly expressed psbA gene is unique and is adapted to the tRNA population, probably due to selection for translation efficiency. In this study the role of selection on codon usage in each of the fully sequenced chloroplast genomes, in addition to Chlamydomonas reinhardtii, is investigated by measuring adaptation to this pattern of codon usage. A method is developed which tests selection on each gene individually by constructing sequences with the same amino acid composition as the gene and randomly assigning codons based on the nucleotide composition of noncoding regions of that genome. The codon bias of the actual gene is then compared to a distribution of random sequences. The data indicate that within the algae selection is strong in Cyanophora paradoxa, affecting a majority of genes, of intermediate intensity in Odontella sinensis, and weaker in Porphyra purpurea and Euglena gracilis. In the plants, selection is found to be quite weak in Pinus thunbergii and the angiosperms but there is evidence that an intermediate level of selection exists in the liverwort Marchantia polymorpha. The role of selection is then further investigated in two comparative studies. It is shown that average relative codon bias is correlated with expression level and that, despite saturation levels of substitution, there is a strong correlation among the algae genomes in the degree of codon bias of homologous genes. All of these data indicate that selection for translation efficiency plays a significant role in determining the codon bias of chloroplast genes but that it acts with different intensities in different lineages. In general it is stronger in the algae than the higher plants, but within the algae Euglena is found to have several unusual features which are noted. The factors that might be responsible for this variation in intensity among the various genomes are discussed.

Chloroplasts↗

Correlation of codon bias measures with mRNA levels: analysis of transcriptome data from Escherichia coli.

Although codon usage is often represented by a 61-dimensional vector, the ability of determining the codon bias in a gene relies on a uni-dimensional vector which measures the total bias in usage of synonymous codons. Codon usage is receiving more and more focus because codon biases might be valuable tools to predict and optimize gene/protein expression. How good any of these measures is for correlating codon usage with gene and protein expression has yet to be investigated. In this study, we correlated gene transcript levels in Escherichia coli with codon usage, using a number of different codon bias measures. We found that there is a significant correlation between transcript levels and codon bias measures, suggesting that these measures can be used to assess or predict gene expression. The codon bias measure performing best in this context was the codon adaptation index.

Codon↗

How optimized is the translational machinery in Escherichia coli, Salmonella typhimurium and Saccharomyces cerevisiae?

The optimization of the translational machinery in cells requires the mutual adaptation of codon usage and tRNA concentration, and the adaptation of tRNA concentration to amino acid usage. Two predictions were derived based on a simple deterministic model of translation which assumes that elongation of the peptide chain is rate-limiting. The highest translational efficiency is achieved when the codon recognized by the most abundant tRNA reaches the maximum frequency. For each codon family, the tRNA concentration is optimally adapted to codon usage when the concentration of different tRNA species matches the square-root of the frequency of their corresponding synonymous codons. When tRNA concentration and codon usage are well adapted to each other, the optimal content of all tRNA species carrying the same amino acid should match the square-root of the frequency of the amino acid. These predictions are examined against empirical data from Escherichia coli, Salmonella typhimurium, and Saccharomyces cerevisiae.

Codon↗

Gene synthesis, bacterial expression and purification of the Rickettsia prowazekii ATP/ADP translocase.

The Rickettsia prowazekii ATP/ADP translocase (Tlc) is the first member of a new family of ATP/ADP exchangers that includes both prokaryotic and eukaryotic proteins. We optimized the codon usage for expression of tlc in Escherichia coli by means of gene synthesis, expressed the synthetic gene in E. coli, and purified a modified Tlc that contained a C-terminal tag of 10 consecutive histidine residues by immobilized metal affinity chromatography. Although codon usage in R. prowazekii is very different from E. coli, the optimization of the codon usage by itself was insufficient to improve expression. However, the change of the cloning vector from pET11a to pT7-5 led to a 3-10-fold increase in the specific ATP transport rate by cells expressing the synthetic construct. The authenticity of the purified protein was confirmed by N-terminal amino acid sequencing and a matrix assisted laser desorption/ionization mass spectrometry.

Amino Acid Sequence↗

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence↗

Synonymous and nonsynonymous substitution rates in diatoms: a comparison between chloroplast and nuclear genes.

Rates of synonymous and nonsynonymous nucleotide substitutions and codon usage bias (ENC) were estimated for a number of nuclear and chloroplast genes in a sample of centric and pennate diatoms. The results suggest that DNA evolution has taken place, on an average, at a slower rate in the chloroplast genes than in the nuclear genes: a rate variation pattern similar to that observed in land plants. Synonymous substitution rates in the chloroplast genes show a negative association with the degree of codon usage bias, suggesting that genes with a higher degree of codon usage bias have evolved at a slower rate. While this relationship has been shown in both prokaryotes and multicellular eukaryotes, it has not been demonstrated before in diatoms.

Cell Nucleus↗

Coding sequence divergence between two closely related plant species: Arabidopsis thaliana and Brassica rapa ssp. pekinensis.

To characterize the coding-sequence divergence of closely related genomes, we compared DNA sequence divergence between sequences from a Brassica rapa ssp. pekinensis EST library isolated from flower buds and genomic sequences from Arabidopsis thaliana. The specific objectives were (i) to determine the distribution of and relationship between K(a) and K(s), (ii) to identify genes with the lowest and highest K(a): K(s) values, and (iii) to evaluate how codon usage has diverged between two closely related species. We found that the distribution of K(a): K(s) was unimodal, and that substitution rates were more variable at nonsynonymous than synonymous sites, and detected no evidence that K(a) and K(s) were positively correlated. Several genes had K(a): K(s) values equal to or near zero, as expected for genes that have evolved under strong selective constraint. In contrast, there were no genes with K(a): K(s) >1 and thus we found no strong evidence that any of the 218 sequences we analyzed have evolved in response to positive selection. We detected a stronger codon bias but a lower frequency of GC at synonymous sites in A. thaliana than B. rapa. Moreover, there has been a shift in the profile of most commonly used synonymous codons since these two species diverged from one another. This shift in codon usage may have been caused by stronger selection acting on codon usage or by a shift in the direction of mutational bias in the B. rapa phylogenetic lineage.

Arabidopsis↗

Microevolutionary divergence pattern of the segmentation gene hunchback in Drosophila.

To study the microevolutionary processes shaping the evolution of the segmentation gene hunchback (hb) from Drosophila melanogaster, we cloned and sequenced the gene from 12 isofemale lines representing wild-type populations of D. melanogaster, as well as from the closely related species Drosophila sechellia, Drosophila orena, and Drosophila yakuba. We find a relatively low degree of sequence variation in D. melanogaster (theta = 0.0017), which is, however, consistent with its chromosomal location in a region of low recombination. Tests of neutrality do not reject a neutral-evolution model for the whole region. However, pairwise tests with different subregions indicate that there is a relative excess of polymorphic sites in the leader and the intron. Codon usage pattern analysis shows a particularly biased codon usage in the highly conserved regions, which is in line with the hypothesis that selection on translational accuracy is the driving force behind such a bias. A comparison of the expression pattern of hb in different sibling species of D. melanogaster reveals some regulatory changes in D. yakuba, which could be interpreted as changes in the timing of secondary expression domains.

Animals↗

BIGPROBE: a computer program that predicts the sequence of long oligonucleotide probes with high reliability.

We have written a computer program, BIGPROBE, which facilitates the design of long nucleic acid probes from the partial or complete amino acid sequence of a protein. BIGPROBE relies upon information on codon usage, intercodon dinucleotide frequency, and potential probe self-complementarity. We have examined the accuracy with which the program predicts coding sequences using sample human and rat genes and probe lengths of 30-60 nucleotides. Rat probe sequences selected by BIGPROBE using either codon usage or dinucleotide frequency data alone averaged 86-92% homology with the known exons of the corresponding gene sequences. Predictive accuracy with rat gene probes could be improved to 89-94%, depending upon probe length, by applying codon usage and dinucleotide frequency data in combination. Similar accuracy was achieved for human genes.

Amino Acid Sequence↗