PubMed HealthSearch

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Codon usage determines translation rate in Escherichia coli.

We wish to determine whether differences in translation rate are correlated with differences in codon usage or with differences in mRNA secondary structure. We therefore inserted a small DNA fragment in the lacZ gene either directly or flanked by a few frame-shifting bases, leaving the reading frame of the lacZ gene unchanged. The fragment was chosen to have "infrequent" codons in one reading frame and "common" codons in the other. The insert in these constructs does not seem to give mRNAs that are able to form extensive secondary structures. The translation time for these modified lacZ mRNAs was measured with a reproducibility better than plus or minus one second. We found that the mRNA with infrequent codons inserted has an approximately three-seconds longer translation time than the one with common codons. In another set of experiments we constructed two almost identical lacZ genes in which the lacZ mRNAs have the potential to generate stem structures with stabilities of about -75 kcal/mol. In this way we could investigate the influence of mRNA structure on translation rate. This type of modified gene was generated in two reading frames with either common or infrequent codons similar to our first experiments. We find that the yield of protein from these mRNAs is reduced, probably due to the action in vivo of an RNase. Nevertheless, the data do not indicate that there is any effect of mRNA secondary structure on translation rate. In contrast, our data persuade us that there is a difference in translation rate between infrequent codons and common codons that is of the order of sixfold.

Bacterial Proteins

Codon usage and evolutionary rates of proteins.

The 61 codons and the three terminators were counted in the coding sequences of 31 families of proteins of higher vertebrates. The protein families were ordered according to their evolutionary rate. In each family, the ratio between the Observed and Expected frequency of each codon was obtained (O/E ratio). A strong and significant positive correlation was observed between the O/E ratio of the eight codons AAC, TAT, ATA, GAA, ACA, AAT, ATG and CGA and the evolutionary rate of the protein. A negative and significant correlation was observed for codons AAG and GAG. It was advanced that the functional constraints of proteins can influence the usage of codons, particularly for those trimers which are components of signal sequences. It was also observed that the O/E ratios of the terminators are negatively correlated with the evolutionary rate of the protein they terminate, and the correlation is significant for TAA and TGA, which in vertebrates might be older than TAG.

Animals

Codon usage can affect efficiency of translation of genes in Escherichia coli.

By inserting synthetic oligonucleotides into a highly expressed gene in E. coli it has been shown that unfavourable codon usage can reduce the maximum translation rate of a protein. However, in the case of the codon used (AGG), a significant effect on translation was only seen at very high transcription rates from a gene containing multiple copies of the unfavourable codon.

Bacterial Proteins

The relationship between base composition and codon usage in bacterial genes and its use for the simple and reliable identification of protein-coding sequences.

Bacterial genes that code for proteins appear to possess a codon usage characteristic of their overall base composition. This results in different but predictable non-random distributions of nucleotides within codons, permitting the recognition of protein-coding sequences in a wide range of bacterial species. The nature of this distribution depends on the base composition of the coding sequence. The position-specific differences are especially conspicuous in genes of extreme G + C content, allowing the particularly reliable prediction of the reading frame and coding strand of experimentally determined DNA sequences. This finding has been exploited to identify the coding sequence of the viomycin phosphotransferase (vph) gene of Streptomyces vinaceus. An easily applied computer program ("Frame") has been written to carry out and display such analyses.

Bacterial Proteins

Nucleotide sequence and codon usage of the elongation factor Tu(EF-Tu) gene from Mycoplasma pneumoniae.

The Mycoplasma pneumoniae tuf gene, encoding the elongation factor protein Tu, was cloned and sequenced. The nucleotide sequence of the mycoplasmal gene showed about 60% homology to the sequences of tuf genes of other prokaryotes, yeast mitochondria and Euglena gracilis chloroplasts, and about 75% similarity was found when comparing the deduced amino acid sequences of the various Tu proteins. The relatively low G + C content (40%) of the M. pneumoniae DNA was reflected in a low G + C content (44.6%) of the tuf gene, and in a preferential use of adenine and uracil at the third position of codons, yet codon usage analysis revealed the presence of almost all of the codons of the genetic code in the mycoplasmal gene. Southern blot hybridization of digested DNAs of 11 Mollicutes species with the entire M. pneumoniae tuf gene and with its 5' part suggested the presence of one copy only of this gene in the representative species of the Mollicutes. In this respect, the Mollicutes resemble Gram-positive bacteria and differ from the Gram-negative bacteria, which carry two copies of the tuf gene.

Amino Acid Sequence

Codon usage in Plasmodium falciparum.

The codon frequencies used in 7874 codons from 17 sequences of Plasmodium falciparum have been examined. The frequency distribution is markedly biased. A and C occur with similar frequency in all positions but G is predominantly in the first base and T is predominantly in the last position. This information can be used to predict the coding strand and reading frame of P. falciparum genes.

Animals

Comparison of dinucleotide frequency and codon usage in Toxoplasma and Plasmodium: evolutionary implications.

The weight-averaged observed/expected dinucleotide frequencies for the sum total of the coding regions of five Toxoplasma genes were compared with the same parameters previously determined for the coding regions of 21 Plasmodium genes. In addition, codon usage in the five Toxoplasma genes was compared with that in the 21 Plasmodium genes, and the percent distribution of amino acids in the Toxoplasma protein pool and the Plasmodium protein pool were compared with that in a general protein pool of 314 proteins. The results are consistent with the hypothesis that, contrary to currently held opinion, the genera Toxoplasma and Plasmodium are not especially closely related.

Animals

Evident diversity of codon usage patterns of human genes with respect to chromosome banding patterns and chromosome numbers; relation between nucleotide sequence data and cytogenetic data.

The sequences of the human genome compiled in DNA databases are now about 10 megabase pairs (Mb), and thus the size of the sequences is several times the average size of chromosome bands at high resolution. By surveying this large quantity of data, it may be possible to clarify the global characteristics of the human genome, that is, correlation of gene sequence data (kb-level) to cytogenetic data (Mb-level). By extensively searching the GenBank database, we calculated codon usages in about 2000 human sequences. The highest G + C percentage at the third codon position was 97%, and that of about 250 sequences was 80% or more. The lowest G + C% was 27%, and that in about 150 sequences was 40% or less. A major portion of the GC-rich genes was found to be on special subsets of R-bands (T-bands and/or terminal R-bands). AT-rich genes, however, were mainly on G-bands or non-T-type internal R-bands. Average G + C% at the third position for individual chromosomes differed among chromosomes, and were related to T-band density, quinacrine dullness, and mitotic chiasmata density in the respective chromosomes.

Base Composition

Correlation between molecular clock ticking, codon usage fidelity of DNA repair, chromosome banding and chromatin compactness in germline cells.

The vertebrate genome is built of long DNA regions, relatively homogeneous in GC content, which likely correspond to bands on stained chromosomes. Large differences in composition have been found among DNA regions belonging to the same genome. They are paralleled by differences in codon usage in genes differently localized. The hypothesis presented here asserts that these differences in composition are caused by different mutational bias of alpha and beta DNA polymerases, these polymerases being involved to different extents in the repair of DNA lesions in compact and relaxed chromatin, respectively, in germline cells.

Animals

Nonrandom patterns of codon usage and of nucleotide substitutions in human alpha- and beta-globin genes: an evolutionary strategy reducing the rate of mutations with drastic effects?

Nucleotide substitutions within a structural gene can cause two principal "drastic" phenotypic effects at the protein level: translatable leads to untranslatable and nonpolar hydrophobic in equilibrium hydrophilic amino acid substitutions. The sequence of nucleotides in the structural human alpha- and beta-globin genes and their variants were examined to determine whether codon usage, patterns of nucleotide substitutions, or both, reduced the relative and absolute rates of these unfavorable mutations. Based on translation of abnormal hemoglobins, it is likely that all 61 nontermination codons are potentially translatable, though only 47 are normally used. Moreover, codons that can mutate to a termination codon are never used whenever the corresponding amino acid is specified also by triplets that cannot mutate to termination by a single-step mutation. Thus, the number of opportunities to mutate to an untranslatable codon is reduced to the minimum compatible with the amino acid composition of these chains. The relative rates of U in equilibrium non-U substitutions were much lower than those of other substitutions. Because U residues must be involved in most termination mutations and in all nonpolar hydrophobic in equilibrium hydrophilic amino acid substitutions, there is a considerable reduction of mutational events, causing drastic phenotypic effects. These findings are likely to be the end result of evolutionary selection by yet unknown mechanisms.

Base Sequence

Lactococcus lactis glyceraldehyde-3-phosphate dehydrogenase gene, gap: further evidence for strongly biased codon usage in glycolytic pathway genes.

The gene gap, encoding glyceraldehyde-3-phosphate dehydrogenase (EC 1.2.1.12), was isolated from a genomic library of Lactococcus lactis LM0230 DNA. Plasmids containing the L. lactis gene were able to complement a gap mutant of Escherichia coli. The nucleotide sequence of gap predicted a polypeptide chain of 337 amino acids for the enzyme and a subunit molecular mass of 36,043. The codon usage in gap and four other glycolytic genes from L. lactis showed a high degree of bias, when compared with 84 other chromosomal genes. Northern blot analysis of total L. lactis RNA showed that gap hybridized strongly with a 1.3 kb transcript. The 5' end of the transcript was determined by primer extension analysis to be a C located 35 bp upstream from the gap start codon. These transcript analyses, and the orientation of the open reading frames in the DNA flanking gap, indicated that in L. lactis gap is expressed on a monocistronic transcript. Nucleotide sequencing indicated that the DNA adjacent to gap did not encode other glycolytic pathway enzymes. The DNA sequence flanking gap contained two open reading frames (ORF156 and ORF211) of unknown function. The 3' end of a clpA homologue was identified in the sequence upstream of ORF156. The location of gap on the L. lactis DL11 chromosome map was determined to be between map coordinates 0.530 and 0.660.

Amino Acid Sequence

Spectinomycin operon of Micrococcus luteus: evolutionary implications of organization and novel codon usage.

The complete DNA sequence of the Micrococcus luteus spectinomycin (spc) operon and its adjacent regions has been determined. The sequence has revealed the presence of genes that are homologous to those of the Escherichia coli ribosomal and related proteins, L14, L24, L5, S8, L6, L18, S5, L30, L15, and secretion protein Y (sec Y), and the gene for adenylate kinase (adk). The gene arrangement in the spc operon is essentially the same as that of E. coli except for the absence in the M. luteus spc operon of the genes for S14 and X protein that exist in the E. coli spc operon. SecY and adk seem to be composed of another operon (adk operon) with at least an open reading frame. The deduced amino acid sequences for these ribosomal proteins are well conserved among the two species (40-65% identity). Reflecting the high genomic guanine and cytosine (GC) content of M. luteus (74%), the codon usage of the genes is extremely biased toward use of G and C, about 94% of the codon third positions being G or C. Seven codons, AUA, AAA, AGA, UUA, GUA, CUA, and CAA, all of which have A at the codon third positions, are completely absent in the M. luteus genes examined. Out of 11 genes in the M. luteus spc and adk operons, 5 (10) use GUG (UGA) and 6 (1) use AUG (UAA) as an initiation (termination) codon.

Amino Acid Sequence

Structural features of multiple nifH-like sequences and very biased codon usage in nitrogenase genes of Clostridium pasteurianum.

The structural gene (nifH1) encoding the nitrogenase iron protein of Clostridium pasteurianum has been cloned and sequenced. It is located on a 4-kilobase EcoRI fragment (cloned into pBR325) that also contains a portion of nifD and another nifH-like sequence (nifH2). C. pasteurianum nifH1 encodes a polypeptide (273 amino acids) identical to that of the isolated iron protein, indicating that the smaller size of the C. pasteurianum iron protein does not result from posttranslational processing. The 5' flanking region of nifH1 or nifH2 does not contain the nif promoter sequences found in several gram-negative bacteria. Instead, a sequence resembling the Escherichia coli consensus promoter (TTGACA-N17-TATAAT) is present before C. pasteurianum nifH2, and a TATAAT sequence is present before C pasteurianum nifH1. Codon usage in nifH1, nifH2, and nifD (partial) is very biased. A preference for A or U in the third position of the codons is seen. nifH2 could encode a protein of 272 amino acid residues, which differs from the iron protein (nifH1 product) in 23 amino acid residues (8%). Another nifH-like sequence (nifH3) is located on a nonadjacent EcoRI fragment and has been partially sequenced. C. pasteurianum nifH2 and nifH3 may encode proteins having several amino acids that are conserved in other proteins but not in C. pasteurianum iron protein, suggesting a possible role for the multiple nifH-like sequences of C. pasteurianum in the evolution of nifH. Among the nine sequenced iron proteins, only the C. pasteurianum protein lacks a conserved lysine residue which is near the extended C terminus of the other iron proteins. The absence of this positive charge in the C. pasteurianum iron protein might affect the cross-reactivity of the protein in heterologous systems.

Amino Acid Sequence

Human hemoglobin expression in Escherichia coli: importance of optimal codon usage.

The overexpression of a nonfusion product of human beta-globin in Escherichia coli from its cDNA sequence has been accomplished for the first time. Expression of beta-globin from its native cDNA required the use of the strong bacteriophage T7 promoter. In this system, beta-globin accumulated to approximately 10% of total E. coli proteins. alpha-Globin was not expressed in the T7 system using the native cDNA. For the expression of alpha-globin, synthetic genes containing optimal E. coli codons were constructed. Neither synthetic alpha- nor beta-globin gene alone was expressed from the lac or tac promoter. Globin expression was achieved when the two synthetic alpha- and beta-globin genes were combined as an operon downstream of the lac promoter. The two proteins combined intracellularly with endogenous heme, which was concomitantly overproduced to yield tetrameric hemoglobin as roughly 5-10% of total E. coli protein. Cloning the alpha- and beta-globin cDNAs in a construct identical with the lac promoter did not yield globin production, establishing the requirement for optimal codon usage. The recombinant beta-globin from the T7 expression system was purified and reconstituted in vitro with heme and native alpha chains. N-terminal analyses showed that the beta-globin produced in the T7 system and the tetrameric hemoglobin produced from the synthetic genes contained an additional beta 1 methionine residue. Two additional mutants, beta 1 Val----Met and beta 1 Val----Ala were produced using the T7 system. Functional and structural properties of the purified hemoglobins will be discussed in the following papers.

Amino Acid Sequence

Nucleotide sequence of simian virus 40 DNA: structure of the middle segment of the HindII + III restriction fragment B (sixth part of the T antigen gene) and codon usage.

We report here the nucleotide sequence of the simian virus 40 DNA region that lies between the EcoRII restriction endonuclease cleavage sites at map positions 0.214 and 0.281. The sequence was determined by partial chemical degradation of terminally labeled DNA fragments according to the procedure of Maxam and Gilbert. This region represents 6.7% of the SV40 genome and is located in the middle of HindII + III restriction fragment B. It is expressed as part of the early 19-S messenger RNA, which codes for the large-T antigen protein. Only one open reading frame for translation can be deduced from the message strand of the DNA and this reading frame connects in phase with the one of both neighboring fragments. This publication is the last in a series of papers about the T-antigen gene, and several properties of this gene and its product are discussed. The non-randomness of codon usage is similar to that previously discussed for the late part of the genome. Moreover, it appears that the choice of a third letter can be determined by the nature of the following codon; some codons which start with a pyrimidine are almost never preceded by an adenosine and some ANN-type codons are almost never preceded by a guanosine.

Amino Acid Sequence

Sequence, codon usage and cysteine periodicity of the SerH1 gene and in the encoded surface protein of Tetrahymena thermophila.

The temperature-regulated SerH1 gene coding for an immunodominant surface glycoprotein (i-Ag H1) of Tetrahymena thermophila has been sequenced. The gene is reproducibly rearranged during macronuclear development and steady state mRNA levels are present at < 36 degrees C. The deduced i-Ag H1 amino acid (aa) sequence is rich in Ser, Thr and Cys, and contains three periods each consisting of 85 aa punctuated by eight Cys with the general formula, CX6CX17CX2CX18CX2CX11CX2CX19 (where X = any aa). Such Cys periodicity is common to ciliate i-Ag. Codon usage in Tt, Paramecium primaurelia and P. tetraurelia i-Ag encoding genes is similar, with approx. 80% A+T in the 3' position which is in marked contrast to the approx. 54% 3' A+T in other ciliate genes.

Amino Acid Sequence

The codon usage of the nisZ operon in Lactococcus lactis N8 suggests a non-lactococcal origin of the conjugative nisin-sucrose transposon.

An 11.6 kb area downstream from the structural gene of nisin Z in the conjugative nisin-sucrose transposon of Lactococcus lactis subsp. lactis N8 was cloned and sequenced. Analysis of the sequence revealed eight open reading frames, nisZBTClPRK, followed by a putative rho-independent terminator (delta G degrees = -4.7 kcal/mol). The C-terminal hydrophilic domain of the NisK protein is homologous to the C-termini of several histidine kinases of bacterial two-component regulator systems, such as SpaK from Bacillus subtilis and KdpD and RcsC of Escherichia coli. The nisin Z biosynthetic genes were highly similar with the genes of the nisin A operons having, however, a 0-3% difference in the amino acid sequences of the individual proteins. The codon usage of eleven genes within the same conjugative transposon was calculated and found to be strikingly different from that of other lactococcal genes. This, together with the low GC-content (32%) compared to the 38% (G+C) of the lactococcal chromosome in general strongly suggests a non-lactococcal origin of this transposon.

Amino Acid Sequence