PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

DNA G+C content of the third codon position and codon usage biases of human genes.

The human genome, as in other eukaryotes, has a wide heterogeneity in the DNA base composition. The evolutionary basis for this heterogeneity has been unknown. A previous study of the human genome (846 genes analyzed) has shown that, in the major range of the G+C content in the third codon position (0.25-0.75), biases from the Parity Rule 2 (PR2) among the synonymous codons of the four-codon amino acids are similar except in the highest G+C range (Sueoka, N., 1999. Translation-coupled violation of Parity Rule 2 in human genes is not the cause of heterogeneity of the DNA G+C content of third codon position. Gene 238, 53-58.). PR2 is an intra-strand rule where A=T and G=C are expected when there are no biases between the two complementary strands of DNA in mutation and selection rates (substitution rates). In this study, 14,026 human genes were analyzed. In addition, the third codon positions of two-codon amino acids were analyzed. New results show the following: (a) The G+C contents of the third codon position of human genes are scattered in the G+C range of 0.22-0.96 in the third codon position. (b) The PR2 biases are similar in the range of 0.25-0.75, whereas, in the high G+C range (0.75-0.96; 13% of the genes), the PR2-bias fingerprints are different from those of the major range. (c) Unlike the PR2 biases, the G+C contents of the third codon position for both four-codon and two-codon amino acids are all correlated almost perfectly with the G+C content of the third codon position over the total G+C ranges. These results support the notion that the directional mutation pressure, rather than the directional selection pressure, is mainly responsible for the heterogeneity of the G+C content of the third codon position.

Amino Acids↗

Alternative CUG codon usage (Ser for Leu) in Pichia farinosa and the effect of a mutated killer gene in Saccharomyces cerevisiae.

The halotolerant yeast Pichia farinosa KK1 strain produces a killer toxin termed SMKT (salt-mediated killer toxin). Mass spectrometry and Edman sequencing of peptides from the mature SMKT and secreted protoxin demonstrate that positions specified by the CUG codon contain unmodified serine (Ser) in P.farinosa. In order to express the authentic SMK1 product in Saccharomyces cerevisiae, which uses the universal genetic code, the three CUG codons corresponding to Ser87, Ser137 and Ser206 in the SMK1 gene were changed to universal Ser codons by site-directed mutagenesis. The expression of the modified SMK1 gene with universal Ser codons was lethal in S.cerevisiae, as well as that of the unmodified SMK1 gene with the CUG codons. The secretion of protoxin with the authentic amino acid sequence from the modified SMK1 was significantly increased, whereas the transcription level of SMK1 was not affected in the presence or absence of CUG codon. Our results provide the first in vivo evidence that non-universal decoding of CUG is used in a hemiascomycetous yeast, P.farinosa.

Amino Acid Substitution↗

Nucleotide substitution pattern in rice paralogues: implication for negative correlation between the synonymous substitution rate and codon usage bias.

Understanding the correlation between synonymous substitution rate and GC content is essential to decipher the gene evolution. However, it has been controversial on their relationship. We analyzed the GC content and synonymous substitution rate in 1092 paralogues produced by two large-scale duplication events in the rice genome. According to the GC content at the third codon sites (GC3), the paralogues were classified into GC3-rich and GC3-poor genes. By referring to their outgroup sequences, we inferred the last common ancestor of sister paralogues and, consequently, calculated the average synonymous substitution rate for two gene classes. The results suggest that average synonymous substitution rate is lower in GC3-rich genes than that in GC3-poor genes, indicating that the synonymous substitution rate is negatively correlated with GC content in the rice genome. Through characterizing the synonymous nucleotide substitution pattern, we found a strong synonymous nucleotide substitution frequency bias from AT to GC in GC3-rich genes. This indicates possible limitations of commonly used methods developed to estimate the synonymous substitution rate. Their estimates might produce misleading results on correlation between the synonymous substitution rate and GC content.

Base Composition↗

Selective charging of tRNA isoacceptors explains patterns of codon usage.

We modeled how the charged levels of different transfer RNAs (tRNAs) that carry the same amino acid (isoacceptors) respond when this amino acid becomes growth-limiting. The charged levels will approach zero for some isoacceptors (such as tRNA2Leu) and remain high for others (such as tRNA4Leu), as determined by the concentrations of isoacceptors and how often their codons occur in protein synthesis. The theory accounts for (synonymous) codons for the same amino acid that are used in ribosome-mediated transcriptional attenuation, the choices of synonymous codons in trans-translating transfermessenger RNA, and the overrepresentation of rare codons in messenger RNAs for amino acid biosynthetic enzymes.

Amino Acids↗

Preliminary indication of unusual codon usage in the DNA coding sequence of the attachment protein of Mycoplasma pneumoniae.

From a Mycoplasma pneumoniae genomic library, three recombinant clones encoding approximately one-third of the attachment (P1) gene were identified. P1 fusion proteins expressed by these clones in Escherichia coli were found to be much smaller than expected from the sizes of the cloned DNA fragments. Nucleotide sequence analysis revealed the presence of UGA codons in the open reading frames of two of the clones, explaining the incomplete translation of the inserts. Sequencing data further revealed that two of the recombinant clones did have similar but not identical carboxyl-end sequences. This finding suggests the existence of more than one genomic DNA sequence coding for the 3'-end of the P1 gene. Potential transcriptional regulatory sequences, a possible termination signal at the 3'-end of the P1 gene and possible promoter-like structures, have been recognized.

Amino Acid Sequence↗

Strong homology between the small subunit of ribulose-1,5-bisphosphate carboxylase/oxygenase of two species of Acetabularia and the occurrence of unusual codon usage.

Amino acid sequences of the small subunit of ribulose-1,5-bisphosphate carboxylase (SSU) of Acetabularia cliftonii and A. mediterranea were derived from five cDNA sequences of each of the two species of algae and by direct amino acid sequence determination of the isolated protein. An homology of more than 96% between the proteins indicates the close relationship between the two algae. All ten cDNAs in the reading frame display the termination codons TAA and/or TAG at various positions, which seem to code for the amino acid glutamine when compared with the amino acid sequence from the mature protein. This is reminiscent of proteins from ciliates where TAA and TAG also code for glutamine.

Acetabularia↗

Biased codon usage near intron-exon junctions: selection on splicing enhancers, splice-site recognition or something else?

Two groups recently argued that, in human genes, synonymous sites near intron-exon junctions undergo selection for correct splicing. However, neither study controlled for the possibility of an underlying nucleotide bias at the ends of exons. In this article, we show that generalized A and T enrichment exists, which could be independent of splicing regulation. Evidence for selection between synonymous codons that are associated with splicing enhancers remains after controlling for this bias, whereas support for cryptic splice-site avoidance is diminished.

Animals↗

A site-directed integration system for the nonuniversal CUG(Ser) codon usage species Pichia farinosa by electroporation.

Halotolerant yeast, Pichia farinosa, is a valuable yeast strain in fermentation industry because it produces high yield of glycerol and xylitol, and can tolerate both contamination and high-density growth during fermentation. However, the lack of genetic manipulation tools makes it less popular as a gene engineering strain. Expression systems commonly used in other yeast systems, such as Saccharomyces cerevisiae and Pichia pastoris cannot be used in P. farinosa because it translates universal Leu codon CUG as Ser. Here we reported a modified expression vector and a transformation system with enhanced efficiency in P. farinosa. The results showed that cells of OD(600 )0.8-1.0 with DTT treatment can obtain high transformation efficiency. The optimized electroporation condition was 900 V, 25 microF, and 200 Omega. The DNA concentration did not influence the transformation. Our system provides the potential not only for applying P. farinosa as an industrial strain of gene engineering, but also for studying gene function in its native host.

Amino Acid Sequence↗

A comparative study of mutations in Escherichia coli and Salmonella typhimurium shows that codon conservation is strongly correlated with codon usage.

Escherichia coli and Salmonella typhimurium are closely related species of enteric bacteria, having diverged from 120 to 160 million years ago, according to the estimate of Ochman & Wilson (1987. J. Mol. Evol.26, 74-86). In order to study base substitution mutations in the genomes of these bacteria, we have compared pairs of genes for the same product in the two species, and have selected a sample in which the protein length is the same in both E. coli and S. typhimurium. From the alignment of these gene pairs, we observe that frequently used codons are more conserved than infrequently used codons, i.e., the apparent mutation rate is higher for rare codons than for popular codons.

Codon↗

Co-variation of tRNA abundance and codon usage in Escherichia coli at different growth rates.

We have used two-dimensional polyacrylamide gel electrophoresis to fractionate tRNAs from Escherichia coli. A sufficiently high degree of resolution was obtained for 44 out of 46 tRNA species in E. coli to be resolved into individual electrophoretic components. These isolated components were identified by hybridization to tRNA-specific oligonucleotide probes. Systematic measurements of the abundance of each individual tRNA isoacceptor in E. coli, grown at rates varying from 0.4 to 2.5 doublings per hour, were made with the aid of this electrophoretic protocol. We find that there is a biased distribution of the tRNA abundance at all growth rates, and that this can be roughly correlated with the values of codon frequencies in the mRNA pools calculated for bacteria growing at different rates. The tRNA species cognate to abundant codons increase in concentration as the growth rate increases but not as dramatically as might be anticipated. The levels of most of the tRNA isoacceptors cognate to less abundant codons remain unchanged with increasing growth rates. The result of these changes in tRNA abundance is that the relative increase in the amounts of major tRNA species in the bacteria growing at the fastest growth rates is more modest than previous estimates from this laboratory suggested. Furthermore, a systematic error in previous estimates of ribosomal RNA content of the bacteria has been detected. This will account for the quantitative discrepancies between the previous and the present data for tRNA abundance.

Base Sequence↗

Optimized codon usage and chromophore mutations provide enhanced sensitivity with the green fluorescent protein.

The green fluorescent protein (GFP) from Aequorea victoria is a versatile reporter protein for monitoring gene expression and protein localization in a variety of cells and organisms. Despite many early successes using this reporter, wild type GFP is suboptimal for most applications due to low fluorescence intensity when excited by blue light (488 nm), a significant lag in the development of fluorescence after protein synthesis, complex photoisomerization of the GFP chromophore and poor expression in many higher eukaryotes. To improve upon these qualities, we have combined a mutant of GFP with a significantly larger extinction coefficient for excitation at 488 nm with a re-engineered GFP gene sequence containing codons preferentially found in highly expressed human proteins. The combination of improved fluorescence intensity and higher expression levels yield an enhanced GFP which provides greater sensitivity in most systems.

Animals↗

Efficient synthesis of secreted murine interleukin-2 by Saccharomyces cerevisiae: influence of 3'-untranslated regions and codon usage.

Several expression vectors were compared which directed the synthesis of secreted murine interleukin-2 (mIL2) in the culture medium of Saccharomyces cerevisiae. We used the prepro-sequence of the alpha 1 mating-factor precursor as a secretion signal in S. cerevisiae in combination with different promoters. The yield of mature mIL2 was significantly improved by deleting the major part of the 3'-untranslated region (UTR). In Northern-blotting experiments we showed that a destabilizing sequence present in the 3' UTR might be responsible for rapid degradation of the mIL2 mRNA. The highest expression (about 10 micrograms/ml) was obtained under control of the GAL1 promoter in an S. cerevisiae strain where the regulatory GAL4 gene was overexpressed. No difference in expression level was observed in a construct wherein twelve consecutive codons were replaced by optimal codons for S. cerevisiae.

Animals↗

Codon usage and mistranslation. In vivo basal level misreading of the MS2 coat protein message.

The coat protein of the small RNA virus MS2 shows charge heterogeneity in vivo. In most strains there is a basic satellite of the native protein. We have shown that this basic satellite is greatly diminished or absent in strains with the streptomycin-resistant allele, rpsL, a mutation which leads to increased translational accuracy. Further, the satellite is present in cells where the coat protein is encoded by duplex DNA. Tryptic digests of the satellite show that it contains new lysine-containing peptides which appear to be the same as those found in derivatives of coat protein which have a lysine for asparagine substitution. Sequencing of the NH2-terminal 19 amino acids of the satellite protein shows that the asparagine codon AAU at amino acid 12 is misread approximately 8 times more frequently than the AAC at amino acid 3. We conclude that the satellite species is the result of basal level lysine for asparagine substitution. These substitutions are most likely caused by preferential misreading of AAU codons at a frequency of approximately 5 X 10(-3), 10-fold higher than the average error frequency.

Amino Acid Sequence↗

Heterogeneity in regional GC content and differential usage of codons and amino acids in GC-poor and GC-rich regions of the genome of Apis mellifera.

The honeybee (Apis mellifera) has a genome with a wide variation in GC content showing 2 clear modal GC values, in some ways reminiscent of an isochore-like structure. To gain insight into causes and consequences of this pattern, we used a comparative approach to study the genome-wide alignment of primarily coding sequence of A. mellifera with Drosophila melanogaster and Anopheles gambiae. The latter 2 species show a higher average GC content than A. mellifera and no indications of bimodality, suggesting that the GC-poor mode is a derived condition in honeybee. In A. mellifera, synonymous sites of genes generally adopt the GC content of the region in which they reside. A large proportion of genes in GC-poor regions have not been assigned to the honeybee assembly because of the low sequence complexity of their genome neighborhood. The synonymous substitution rate between A. mellifera and the other species is very close to saturation, but analyses of nonsynonymous substitutions as well as amino acid substitutions indicate that the GC-poor regions are not evolving faster than the GC-rich regions. We describe the codon usage and amino acid usage and show that they are remarkably heterogeneous within the honeybee genome between the 2 different GC regions. Specifically, the genes located in GC-poor regions show a much larger deviation in both codon usage bias and amino acid usage from the Dipterans than the genes located in the GC-rich regions.

Amino Acids↗

Co-evolution of base composition and codon usage in Xenopus laevis and human globin genes with long-range DNA organization of their genome.

Eucaryotic DNA is punctuated by many A+T-rich segments that we named A+T-rich linkers. Two types of these A+T-rich linkers can be distinguished: (i) isolated A+T-rich linkers, and (ii) A+T-rich linkers crowded in clusters. We have analysed the distribution of A+T-rich linker across the alpha- and beta-globin gene domain in Xenopus laevis and human genomes using isodenaturation and electron microscopy. Comparison of our data with those previously obtained for the avian globin genes leads us to conclude that genes can be harboured indifferently in either domain. A correlation is established between the presence of A+T-rich linker inside introns and flanking regions and the A+T content of the coding sequence. For the coding sequence, a high A+T content is strongly correlated with high A+T content in the codon's third position and weakly in the first position.

Animals↗

Specific amino acid content and codon usage account for the existence of overlapping ORFS.

Here we present a novel hypothesis for the origin of overlapping open reading frames (O-ORFs) observed in the 'non-coding frames' of several genes of yeast chromosome II. By computer analysis it was found that the specific amino acid content and base distribution pattern at certain genomic locations and the presence of O-ORFs were related. This observation prompt us to conclude that these O-ORFs are mere statistical curiosities without any biological function, which is in contrast to the hypotheses proposed by other authors.

Amino Acids↗

Protein evolution and codon usage bias on the neo-sex chromosomes of Drosophila miranda.

The neo-sex chromosomes of Drosophila miranda constitute an ideal system to study the effects of recombination on patterns of genome evolution. Due to a fusion of an autosome with the Y chromosome, one homolog is transmitted clonally. Here, I compare patterns of molecular evolution of 18 protein-coding genes located on the recombining neo-X and their homologs on the nonrecombining neo-Y chromosome. The rate of protein evolution has significantly increased on the neo-Y lineage since its formation. Amino acid substitutions are accumulating uniformly among neo-Y-linked genes, as expected if all loci on the neo-Y chromosome suffer from a reduced effectiveness of natural selection. In contrast, there is significant heterogeneity in the rate of protein evolution among neo-X-linked genes, with most loci being under strong purifying selection and two genes showing evidence for adaptive evolution. This observation agrees with theory predicting that linkage limits adaptive protein evolution. Both the neo-X and the neo-Y chromosome show an excess of unpreferred codon substitutions over preferred ones and no difference in this pattern was observed between the chromosomes. This suggests that there has been little or no selection maintaining codon bias in the D. miranda lineage. A change in mutational bias toward AT substitutions also contributes to the decline in codon bias. The contrast in patterns of molecular evolution between amino acid mutations and synonymous mutations on the neo-sex-linked genes can be understood in terms of chromosome-specific differences in effective population size and the distribution of selective effects of mutations.

Animals↗

Codon usage and secondary structure of MS2 phage RNA.

MS2 is an RNA bacteriophage (3569 bases). The secondary structure of the RNA has been determined, and is known to play an important role in regulating translation. Paired regions of the genome have a higher G+C content than unpaired regions. It has been suggested that this reflects selection for high G+C content to encourage pairing, but a re-analysis of the data together with computer simulation suggest that it is an automatic consequence in any RNA sequence of the way it folds up to minimise its free energy. It has also been suggested that the three registers in which pairing can occur in a coding region are used differentially to optimise the use of the redundancy of the genetic code, but re-analysis of the data shows only weak statistical support for this hypothesis.

Base Composition↗