PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Evolution of the Adh locus in the Drosophila willistoni group: the loss of an intron, and shift in codon usage.

We report here the DNA sequence of the alcohol dehydrogenase gene (Adh) cloned from Drosophila willistoni. The three major findings are as follows: (1) Relative to all other Adh genes known from Drosophila, D. willistoni Adh has the last intron precisely deleted; PCR directly from total genomic DNA indicates that the deletion exists in all members of the willistoni group but not in any other group, including the closely related saltans group. Otherwise the structure and predicted protein are very similar to those of other species. (2) There is a significant shift in codon usage, especially compared with that in D. melanogaster Adh. The most striking shift is from C to U in the wobble position (both third and first position). Unlike the codon-usage-bias pattern typical of highly biased genes in D. melanogaster, including Adh, D. willistoni has nearly 50% G + C in the third position. (3) The phylogenetic information provided by this new sequence is in agreement with almost all other molecular and morphological data, in placing the obscura group closer to the melanogaster group, with the willistoni group farther distant but still clearly within the subgenus Sophophora.

Alcohol Dehydrogenase↗

Sequence analysis of the cDNA encoding human liver glycogen phosphorylase reveals tissue-specific codon usage.

We have cloned the cDNA encoding glycogen phosphorylase (1,4-alpha-D-glucan:orthophosphate alpha-D-glucosyl-transferase, EC 2.4.1.1) from human liver. Blot-hybridization analysis using a large fragment of the cDNA to probe mRNA from rabbit brain, muscle, and liver tissues shows preferential hybridization to liver RNA. Determination of the entire nucleotide sequence of the liver message has allowed a comparison with the previously determined rabbit muscle phosphorylase sequence. Despite an amino acid identity of 80%, the two cDNAs exhibit a remarkable divergence in G+C content. In the muscle phosphorylase sequence, 86% of the nucleotides at the third codon position are either deoxyguanosine or deoxycytidine residues, while in the liver homolog the figure is only 60%, resulting in a strikingly different pattern of codon usage throughout most of the sequence. The liver phosphorylase cDNA appears to represent an evolutionary mosaic; the segment encoding the N-terminal 80 amino acids contains greater than 90% G+C at the third codon position. A survey of other published mammalian cDNA sequences reveals that the data for liver and muscle phosphorylases reflects a bias in codon usage patterns in liver and muscle coding sequences in general.

Animals↗

Analysis of codon usage diversity of bacterial genes with a self-organizing map (SOM): characterization of horizontally transferred genes with emphasis on the E. coli O157 genome.

With increases in the amounts of available DNA sequence data, it has become increasingly important to develop tools for comprehensive systematic analysis and comparison of species-specific characteristics of protein-coding sequences for a wide variety of genomes. In the present study, we used a novel neural-network algorithm, a self-organizing map (SOM), to efficiently and comprehensively analyze codon usage in approximately 60,000 genes from 29 bacterial species simultaneously. This SOM makes it possible to cluster and visualize genes of individual species separately at a much higher resolution than can be obtained with principal component analysis. The organization of the SOM can be explained by the genome G+C% and tRNA compositions of the individual species. We used SOM to examine codon usage heterogeneity in the E. coli O157 genome, which contains 'O157-unique segments' (O-islands), and showed that SOM is a powerful tool for characterization of horizontally transferred genes.

Algorithms↗

Development of a reporter system for the yeast Schwanniomyces occidentalis: influence of DNA composition and codon usage.

In this paper we report on searching for suitable reporters to monitor gene expression and protein secretion in the amylolytic yeast Schwanniomyces occidentalis. Several potential reporter and marker genes, formerly shown to be functional in other yeasts, were cloned downstream from the homologous invertase gene (INV) promoter and their activity was followed in conditions of repression and derepression of the INV promoter. However, neither beta-glucuronidase nor beta-lactamase nor phleomycin resistance-conferring gene, all originating from E. coli, were expressed in S. occidentalis cells to such a level to allow for monitoring of their activity. All the reporter genes tested have a higher percentage of GC (47-62%) in their DNA compared to the DNA composition of S. occidentalis genes that are more AT-rich (36% GC). The codon usage of all the reporter genes also varies from that of 16 so far sequenced S. occidentalis genes. This suggests that an appropriate composition of DNA and a codon usage similar to S. occidentalis genes might be very important parameters for an efficient expression of a heterologous gene in Schwanniomyces occidentalis. Indeed, two genes originating from Staphylococcus aureus, with an AT-content in their DNA similar to that of S. occidentalis, were functionally expressed in S. occidentalis cells. Both a phleomycin resistance-conferring gene and a chloramphenicol acetyltransferase-encoding gene thus represent suitable reporters of gene expression and protein secretion in S. occidentalis. Additionally, we show in this work that the transcription-regulating region and the signal peptide sequence of the S. occidentalis invertase gene were efficient to direct gene expression and subsequent protein secretion in Saccharomyces cerevisiae.

Bacterial Proteins↗

Codon usage in Kluyveromyces lactis and in yeast cytochrome c-encoding genes.

Codon usage (CU) in Kluyveromyces lactis has been studied. Comparison of CU in highly and lowly expressed genes reveals the existence of 21 optimal codons; 18 of them are also optimal in other yeasts like Saccharomyces cerevisiae or Candida albicans. Codon bias index (CBI) values have been recalculated with reference to the assignment of optimal codons in K. lactis and compared to those previously reported in the literature taking as reference the optimal codons from S. cerevisiae. A new index, the intrinsic codon deviation index (ICDI), is proposed to estimate codon bias of genes from species in which optimal codons are not known; its correlation with other index values, like CBI or effective number of codons (Nc), is high. A comparative analysis of CU in six cytochrome-c-encoding genes (CYC) from five yeasts is also presented and the differences found in the codon bias of these genes are discussed in relation to the metabolic type to which the corresponding yeasts belong. Codon bias in the CYC from K. lactis and S. cerevisiae is correlated to mRNA levels.

Amino Acids↗

Codon usage pattern in alpha 2(I) chain domain of chicken type I collagen and its implications for the secondary structure of the mRNA and the synthesis pauses of the collagen.

A stability map of local secondary structure of the mRNA of the triple-helical alpha 2(I) chain domain of chicken type I collagen was obtained by plotting the free energy of the optimal secondary structure of a local segment in mRNA against the segment position along a base sequence of the mRNA. It was found that the positions of the minima of free energy in the plot coincide with the positions where synthesis pauses of the alpha-chain polypeptides of the corresponding sizes translated from the mRNA have been reported to occur (1). The codon usage pattern of each of the three major amino acids of the alpha-chain domain of the collagen, Gly, Pro and Ala, fluctuates considerably along the base sequence segments of the mRNA and a deviation of the pattern from that of the average of the whole alpha 2(I) chain domain mRNA, particularly for Gly codons, leads to a loss of the stability of the local secondary structure of the mRNA. The results suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA of the chicken collagen alpha 2(I) chain domain which leads to a nonuniform polypeptide elongation pattern.

Animals↗

Evidence for selective evolution in codon usage in conserved amino acid segments of human alphaherpesvirus proteins.

The genomes of human viruses herpes simplex 1 (HSV1) and varicella zoster (VZV), although similar in biology, largely concordant in gene order, and identical in many amino acid segments, differ widely in their genomic G + C (abbreviated S) content, which is high in HSV1 (68%) and low in VZV (46%). This paper analyzes several striking codon usage contrasts. The S difference in coding regions is dramatically large in codon site 3, S3, about 42%. The large difference in S3 is maintained at the same level in a subset of closely similar genes and even in corresponding identical amino acid blocks. A similar difference in S levels in silent site 1 (S1) is found in leucine and arginine. The difference in S3 levels occurs in every gene and in every multicodon amino acid form. The S difference also exists in amino acid usage, with HSV1 using significantly more codon types SSN, while VZV uses more codon types WWN (where W stands for A or T). The nonoverlapping and narrow histograms of S3 gene frequencies in both viruses suggest that the difference has arisen and been maintained by a process of selective rather than nonselective effects. This is in sharp contrast to the relatively large variance seen for highly similar genes in the human versus yeast analysis. Interpretations and hypotheses to explain the HSV1 vs VZV codon usage disparity relate to virus-host interactions, to the role of viral genes in DNA metabolism, to availability of molecular resources (molecular Gause exclusion principle), and to differences in genomic structure.

Amino Acids↗

Can codon usage bias explain intron phase distributions and exon symmetry?

More introns exist between codons (phase 0) than between the first and the second bases (phase 1) or between the second and the third base (phase 2) within the codon. Many explanations have been suggested for this excess of phase 0. It has, for example, been argued to reflect an ancient utility for introns in separating exons that code for separate protein modules. There may, however, be a simple, alternative explanation. Introns typically require, for correct splicing, particular nucleotides immediately 5' in exons (typically a G) and immediately 3' in the following exon (also often a G). Introns therefore tend to be found between particular nucleotide pairs (e.g., G|G pairs) in the coding sequence. If, owing to bias in usage of different codons, these pairs are especially common at phase 0, then intron phase biases may have a trivial explanation. Here we take codon usage frequencies for a variety of eukaryotes and use these to generate random sequences. We then ask about the phase of putative intron insertion sites. Importantly, in all simulated data sets intron phase distribution is biased in favor of phase 0. In many cases the bias is of the magnitude observed in real data and can be attributed to codon usage bias. It is also known that exons may carry either the same phase (symmetric) or different phases (asymmetric) at the opposite ends. We simulated a distribution of different types of exons using frequencies of introns observed in real genes assuming random combination of intron phases at the opposite sides of exons. Surprisingly the simulated pattern was quite similar to that observed. In the simulants we typically observe a prevalence of symmetric exons carrying phase 0 at both ends, which is common for eukaryotic genes. However, at least in some species, the extent of the bias in favor of symmetric (0,0) exons is not as great in simulants as in real genes. These results emphasize the need to construct a biologically relevant null model of successful intron insertion.

Animals↗

A method for measuring the non-random bias of a codon usage table.

We describe a new statistical method for measuring bias in the codon usage table of a gene. The test is based on the multinomial and Poisson distributions. The method is used to scan DNA sequences and measure the strength of codon preference. For E. Coli we show that the strength of codon preference is related to levels of gene expression. The method can also be used to compare base triplet frequencies with those expected from the base composition. This second type of codon bias test is useful for distinguishing coding from non-coding regions.

Base Sequence↗

The 'weighted sum of relative entropy': a new index for synonymous codon usage bias.

Shannon entropy from information theory has been applied to estimate the degree of deviation from equal usage of synonymous codons; however, previous attempts have failed to take into account all three aspects of amino acid usage, i.e. (i) the number of distinct amino acids, (ii) their relative frequencies, and (iii) their degree of codon degeneracy. A new index taking into account all of these aspects is proposed. The index, designated as the 'weighted sum of relative entropy' (E(w)), is defined as the sum of the relative entropy of each amino acid weighted by its relative frequency in the sequence. In this paper, we demonstrate that E(w) allows us to avoid some amino acid usage biases and can yield results contradictory to those obtained by previous methods.

Algorithms↗

Mononucleotide and dinucleotide frequencies, and codon usage in poliovirion RNA.

The polio type 1 (Mahoney) RNA sequence (1) has been analyzed in terms of the distribution of its mononucleotides, dinucleotides and trinucleotides (codons). The distribution of adenosine in the sequence is nonuniform, being lower at the 5' end and higher at the 3' end. The dinucleotide CG is relatively rare and the dinucleotides UG and CA are relatively more common than expected. Codon usage is decidedly nonrandom. Codons containing CG are avoided and those ending in adenosine are favored. The asymmetric use of mononucleotides, dinucleotides and codons in polio RNA is unexplained at the present time although the lowered CG frequency may be the result of a DNA origin for polio RNA.

Codon↗

Contrasts in codon usage of latent versus productive genes of Epstein-Barr virus: data and hypotheses.

Epstein-Barr virus (EBV) has two different modes of existence: latent and productive. There are eight known genes expressed during latency (and hardly at all during the productive phase) and about 70 other ("productive") genes. It is shown that the EBV genes known to be expressed during latency display codon usage strikingly different from that of genes that are expressed during lytic growth. In particular, the percentage of S3 (G or C in codon site 3) is persistently lower (about 20%) in all latent genes than in nonlatent genes. Moreover, S3 is lower in each multicodon amino acid form. Also, the percentage of S in silent codon sites 1 of leucine and arginine is lower in latent than in nonlatent genes. The largest absolute differences in amino acid usage between latent and nonlatent genes emphasize codon types SSN and WWN (W means nucleotide A or T and N is any nucleotide). Two principal explanations to account for the EBV latent versus productive gene codon disparity are proposed. Latent genes have codon usage substantially different from that of host cell genes to minimize the deleterious consequences to the host of viral gene expression during latency. (Productive genes are not so constrained.) It is also proposed that the latency genes of EBV were acquired recently by the viral genome. Evidence and arguments for these proposals are presented.

Amino Acid Sequence↗

Enhanced expression in tobacco of the gene encoding green fluorescent protein by modification of its codon usage.

The gene encoding green fluorescent protein (GFP) from Aequorea victoria was resynthesized to adapt its codon usage for expression in plants by increasing the frequency of codons with a C or a G in the third position from 32 to 60%. The strategy for constructing the synthetic gfp gene was based on the overlap extension PCR method using 12 long oligonucleotides as the starting material and as primers. The new gene contains 101 silent nucleotide changes compared to its wild-type counterpart used in this study. Several transgenic tobacco lines containing the wild-type gfp gene contained minute amounts of a smaller protein cross-reacting with GFP antiserum, whereas only one protein of the expected size was found in transgenics with the synthetic gfp gene. The smaller protein was probably encoded by a truncated gfp mRNA created by splicing of a 84 bp cryptic intron as detected by a reverse transcription-PCR technique. A comparison of GFP production in transgenics with the wild-type and the synthetic gfp gene under the control of the enhanced CaMV 35S promoter showed that the large-scale alterations in the gfp gene increased the frequency of high expressors in the transgenic population but hardly changed the maximum GFP concentrations. The latter phenomenon may be attributed to a reduced regeneration capacity of transformed cells with higher GFP concentrations.

Amino Acid Sequence↗

Translation of the downstream ORF from bicistronic mRNAs by human cells: Impact of codon usage and splicing in the upstream ORF.

Biochemistry textbooks describe eukaryotic mRNAs as monocistronic. However, increasing evidence reveals the widespread presence and translation of upstream open reading frames preceding the "main" ORF. DNA and RNA viruses infecting eukaryotes often produce polycistronic mRNAs and viruses have evolved multiple ways of manipulating the host's translation machinery. Here, we introduce an experimental model to study gene expression regulation from virus-like bicistronic mRNAs in human cells. The model consists of a short upstream ORF and a reporter downstream ORF encoding a fluorescent protein. We have engineered synonymous variants of the upstream ORF to explore large parameter space, including codon usage preferences, mRNA folding features, and splicing propensity. We show that human translation machinery can translate the downstream ORF from bicistronic mRNAs, albeit reporter protein levels are thousand times lower than those from the upstream ORF. Furthermore, synonymous recoding of the upstream ORF exclusively during elongation significantly influences its own translation efficiency, reveals cryptic splice signals, and modulates the probability of downstream ORF translation. Our results are consistent with a leaky scanning mechanism facilitating downstream ORF translation from bicistronic mRNAs in human cells, offering new insights into the role of upstream ORFs in translation regulation.

Humans↗

Nucleotide sequence of a macronuclear DNA molecule coding for alpha-tubulin from the ciliate Stylonychia lemnae. Special codon usage: TAA is not a translation termination codon.

The gene-sized macronuclear DNA of the hypotrichous ciliate Stylonychia lemnae contains two size classes of DNA molecules (1.85 and 1.73 kbp) coding for alpha-tubulin. Each macronucleus contains about 55000 copies of the 1.85 kbp molecules and about 17000 copies of the 1.73 kbp DNA molecules. Five macronuclear molecules of these sequences were cloned and sequenced, one, from the 1.85 kbp size class in its entirety. The 5 sequences fell into two classes suggesting that Stylonychia lemnae contains at least two different alpha-tubulin genes. All 5 clones show the codon TAA in the same nucleotide positions of the coding region. In this position the TAA codon cannot function as a translational stop codon and we suggest that this codon codes for the amino acid glutamine. The nucleotide sequence of the coding region as well as the encoded amino acid sequence is highly conserved compared to alpha-tubulin genes from vertebrates. The noncoding regions show several putative transcription-regulatory sequences as well as sequences presumably functioning as replication origins.

Base Sequence↗

Patterns of synonymous codon usage in Drosophila melanogaster genes with sex-biased expression.

The nonrandom use of synonymous codons (codon bias) is a well-established phenomenon in Drosophila. Recent reports suggest that levels of codon bias differ among genes that are differentially expressed between the sexes, with male-expressed genes showing less codon bias than female-expressed genes. To examine the relationship between sex-biased gene expression and level of codon bias on a genomic scale, we surveyed synonymous codon usage in 7276 D. melanogaster genes that were classified as male-, female-, or non-sex-biased in their expression in microarray experiments. We found that male-biased genes have significantly less codon bias than both female- and non-sex-biased genes. This pattern holds for both germline and somatically expressed genes. Furthermore, we find a significantly negative correlation between level of codon bias and degree of sex-biased expression for male-biased genes. In contrast, female-biased genes do not differ from non-sex-biased genes in their level of codon bias and show a significantly positive correlation between codon bias and degree of sex-biased expression. These observations cannot be explained by differences in chromosomal distribution, mutational processes, recombinational environment, gene length, or absolute expression level among genes of the different expression classes. We propose that the observed codon bias differences result from differences in selection at synonymous and/or linked nonsynonymous sites between genes with male- and female-biased expression.

Animals↗

Minor structural consequences of alternative CUG codon usage (Ser for Leu) in Candida albicans exoglucanase.

In some species of Candida the CUG codon is encoded as serine and not leucine. In the case of the exo-beta-1,3-glucanase from the pathogenic fungus C. albicans there are two such translational events, one in the prepro-leader sequence and the other at residue 64. Overexpression of active mature enzyme in a yeast host indicated that these two positions are tolerant to substitution. By comparing the crystal structure of the recombinant protein with that of the native (presented here), it is seen how either serine or leucine can be accommodated at position 64. Examination of the relatively few solved protein structures from C. albicans indicates that other CUG encoded serines are also found at non-essential surface sites. However such codon usage is rare in C. albicans, in contrast to C. rugosa, with direct implications for respective recombinant protein production.

Amino Acid Substitution↗

Successful lateral transfer requires codon usage compatibility between foreign genes and recipient genomes.

We present evidence supporting the notion that codon usage (CU) compatibility between foreign genes and recipient genomes is an important prerequisite to assess the selective advantage of imported functions, and therefore to increase the fixation probability of horizontal gene transfer (HGT) events. This contrasts with the current tendency in research to predict recent HGTs in prokaryotes by assuming that acquired genes generally display poor CU. By looking at the CU level (poor, typical, or rich) exhibited by putative xenologs still resembling their original CU, we found that most alien genes predominantly present typical CU immediately upon introgression, thereby suggesting that the role of CU amelioration in HGT has been overemphasized. In our strategy, we first scanned a representative set of 103 complete prokaryotic genomes for all pairs of candidate xenologs (exported/imported genes) displaying similar CU. We applied additional filtering criteria, including phylogenetic validations, to enhance the reliability of our predictions. Our approach makes no assumptions about the CU of foreign genes being typical or atypical within the recipient genome, thus providing a novel unbiased framework to study the evolutionary dynamics of HGT.

Archaea↗