PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Genomic analysis of influenza A viruses, including avian flu (H5N1) strains.

This study was designed to conduct genomic analysis in two steps, such as the overall relative synonymous codon usage (RSCU) analysis of the five virus species in the orthomyxoviridae family, and more intensive pattern analysis of the four subtypes of influenza A virus (H1N1, H2N2, H3N2, and H5N1) which were isolated from human population. All the subtypes were categorized by their isolated regions, including Asia, Europe, and Africa, and most of the synonymous codon usage patterns were analyzed by correspondence analysis (CA). As a result, influenza A virus showed the lowest synonymous codon usage bias among the virus species of the orthomyxoviridae family, and influenza B and influenza C virus were followed, while suggesting that influenza A virus might have an advantage in transmitting across the species barrier due to their low codon usage bias. The ENC values of the host-specific HA and NA genes represented their different HA and NA types very well, and this reveals that each influenza A virus subtype uses different codon usage patterns as well as the amino acid compositions. In NP, PA and PB2 genes, most of the virus subtypes showed similar RSCU patterns except for H5N1 and H3N2 (A/HK/1774/1999) subtypes which were suspected to be transmitted across the species barrier, from avian and porcine species to human beings, respectively. This distinguishable synonymous codon usage patterns in non-human origin viruses might be useful in determining the origin of influenza A viruses in genomic levels as well as the serological tests. In this study, all the process, including extracting sequences from GenBank flat file and calculating codon usage values, was conducted by Java codes, and these bioinformatics-related methods may be useful in predicting the evolutionary patterns of pandemic viruses.

Africa↗

Usage of the three termination codons: compilation and analysis of the known eukaryotic and prokaryotic translation termination sequences.

The published translation termination sequences have been compiled and analysed to aid the interpretation of experiments on termination codon usage in the Xenopus oocyte (Bienz et al. 1981). There are significant differences between prokaryotes and eukaryotes concerning the usage of the three termination codons and of tandem stops. In addition viruses show termination strategies that differ from those of their hosts. Preferred context sequences flanking termination codons are described. Contexts vary within the last codon according to the nature of the termination codon, but are uniform within the first triplet following the terminators.

Animals↗

Evidence for codon bias selection at the pre-mRNA level in eukaryotes.

We investigated codon usage patterns across eukaryotic exons. We have shown that in humans codon preference varies with distance from the splice sites. This is consistent with the distribution of RNA elements involved in splicing regulation. Our results provide the first evidence that selection at the pre-mRNA level influences codon usage in humans. We also show that systematic trends in codon usage are found in other eukaryotes, suggesting that pre-mRNA level selection for codon usage could be a widespread phenomenon in organisms that undergo RNA splicing.

Animals↗

Lineage-specific variations of congruent evolution among DNA sequences from three genomes, and relaxed selective constraints on rbcL in Cryptomonas (Cryptophyceae).

BACKGROUND: Plastid-bearing cryptophytes like Cryptomonas contain four genomes in a cell, the nucleus, the nucleomorph, the plastid genome and the mitochondrial genome. Comparative phylogenetic analyses encompassing DNA sequences from three different genomes were performed on nineteen photosynthetic and four colorless Cryptomonas strains. Twenty-three rbcL genes and fourteen nuclear SSU rDNA sequences were newly sequenced to examine the impact of photosynthesis loss on codon usage in the rbcL genes, and to compare the rbcL gene phylogeny in terms of tree topology and evolutionary rates with phylogenies inferred from nuclear ribosomal DNA (concatenated SSU rDNA, ITS2 and partial LSU rDNA), and nucleomorph SSU rDNA. RESULTS: Largely congruent branching patterns and accelerated evolutionary rates were found in nucleomorph SSU rDNA and rbcL genes in a clade that consisted of photosynthetic and colorless species suggesting a coevolution of the two genomes. The extremely accelerated rates in the rbcL phylogeny correlated with a shift from selection to mutation drift in codon usage of two-fold degenerate NNY codons comprising the amino acids asparagine, aspartate, histidine, phenylalanine, and tyrosine. Cysteine was the sole exception. The shift in codon usage seemed to follow a gradient from early diverging photosynthetic to late diverging photosynthetic or heterotrophic taxa along the branches. In the early branching taxa, codon preferences were changed in one to two amino acids, whereas in the late diverging taxa, including the colorless strains, between four and five amino acids showed changes in codon usage. CONCLUSION: Nucleomorph and plastid gene phylogenies indicate that loss of photosynthesis in the colorless Cryptomonas strains examined in this study possibly was the result of accelerated evolutionary rates that started already in photosynthetic ancestors. Shifts in codon usage are usually considered to be caused by changes in functional constraints and in gene expression levels. Thus, the increasing influence of mutation drift on codon usage along the clade may indicate gradually relaxed constraints and reduced expression levels on the rbcL gene, finally correlating with a loss of photosynthesis in the colorless Cryptomonas paramaecium strains.

Asparagine↗

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals↗

On the rate of DNA sequence evolution in Drosophila.

Analysis of the rate of nucleotide substitution at silent sites in Drosophila genes reveals three main points. First, the silent rate varies (by a factor of two) among nuclear genes; it is inversely related to the degree of codon usage bias, and so selection among synonymous codons appears to constrain the rate of silent substitution in some genes. Second, mitochondrial genes may have evolved only as fast as nuclear genes with weak codon usage bias (and two times faster than nuclear genes with high codon usage bias); this is quite different from the situation in mammals where mitochondrial genes evolve approximately 5-10 times faster than nuclear genes. Third, the absolute rate of substitution at silent sites in nuclear genes in Drosophila is about three times higher than the average silent rate in mammals.

Animals↗

An environmental signature for 323 microbial genomes based on codon adaptation indices.

BACKGROUND: Codon adaptation indices (CAIs) represent an evolutionary strategy to modulate gene expression and have widely been used to predict potentially highly expressed genes within microbial genomes. Here, we evaluate and compare two very different methods for estimating CAI values, one corresponding to translational codon usage bias and the second obtained mathematically by searching for the most dominant codon bias. RESULTS: The level of correlation between these two CAI methods is a simple and intuitive measure of the degree of translational bias in an organism, and from this we confirm that fast replicating bacteria are more likely to have a dominant translational codon usage bias than are slow replicating bacteria, and that this translational codon usage bias may be used for prediction of highly expressed genes. By analyzing more than 300 bacterial genomes, as well as five fungal genomes, we show that codon usage preference provides an environmental signature by which it is possible to group bacteria according to their lifestyle, for instance soil bacteria and soil symbionts, spore formers, enteric bacteria, aquatic bacteria, and intercellular and extracellular pathogens. CONCLUSION: The results and the approach described here may be used to acquire new knowledge regarding species lifestyle and to elucidate relationships between organisms that are far apart evolutionarily.

Bacteria↗

Predicted highly expressed genes of diverse prokaryotic genomes.

Our approach in predicting gene expression levels relates to codon usage differences among gene classes. In prokaryotic genomes, genes that deviate strongly in codon usage from the average gene but are sufficiently similar in codon usage to ribosomal protein genes, to translation and transcription processing factors, and to chaperone-degradation proteins are predicted highly expressed (PHX). By these criteria, PHX genes in most prokaryotic genomes include those encoding ribosomal proteins, translation and transcription processing factors, and chaperone proteins and genes of principal energy metabolism. In particular, for the fast-growing species Escherichia coli, Vibrio cholerae, Bacillus subtilis, and Haemophilus influenzae, major glycolysis and tricarboxylic acid cycle genes are PHX. In Synechocystis, prime genes of photosynthesis are PHX, and in methanogens, PHX genes include those essential for methanogenesis. Overall, the three protein families-ribosomal proteins, protein synthesis factors, and chaperone complexes-are needed at many stages of the life cycle, and apparently bacteria have evolved codon usage to maintain appropriate growth, stability, and plasticity. New interpretations of the capacity of Deinococcus radiodurans for resistance to high doses of ionizing radiation is based on an excess of PHX chaperone-degradation genes and detoxification genes. Expression levels of selected classes of genes, including those for flagella, electron transport, detoxification, histidine kinases, and others, are analyzed. Flagellar PHX genes are conspicuous among spirochete genomes. PHX genes are positively correlated with strong Shine-Dalgarno signal sequences. Specific regulatory proteins, e.g., two-component sensor proteins, are rarely PHX. Genes involved in pathways for the synthesis of vitamins record low predicted expression levels. Several distinctive PHX genes of the available complete prokaryotic genomes are highlighted. Relationships of PHX genes with stoichiometry, multifunctionality, and operon structures are discussed. Our methodology may be used complementary to experimental expression analysis.

Archaea↗

Bioinformatic analysis of the link between gene composition and expressivity in Saccharomyces cerevisiae and Schizosaccharomyces pombe.

The compositional non-randomness was studied in genes of Saccharomyces cerevisiae and Schizosaccharomyces pombe. In both species, codon usage is well correlated with expressivity (measured as the codon adaptation index). Both species generally display higher nucleotide non-randomness in the group of highly expressed genes than in the lowly expressed genes. The highly expressed genes in both species are furthermore characterized by marked peaks in non-randomness at N=3 upstream of start codons, N=2 downstream of start codons and at N=1 and N=7 downstream of stop codons, indicating that these nucleotides may be key elements in translational regulation. Intragenic variation in codon usage was also observed to be linked to expressivity. It is suggested that the firm link between expressivity and codon usage calls for codon optimization. Based on bioinformatic calculations, examples of proteins are given for which codon optimizations might be relevant.

Base Sequence↗

The nonrandom location of synonymous codons suggests that reading frame-independent forces have patterned codon preferences.

Biased codon usage is common in eukaryotic and prokaryotic genes. Evidence from Escherichia, Saccharomyces, and Drosophila indicates that it favors translational efficiency and accuracy. However, to date no functional advantages have been identified in the codon-anticodon interactions involving the most frequently used (preferred) codons. Here we present evidence that forces not related to the individual codon-anticodon interaction may be involved in determining which synonymous codons are preferred or avoided. We show that the "off-frame" trinucleotide motif preferences inferrable from Drosophila coding regions are often in the same direction as Drosophila's "in-frame" codon preferences, i.e., its codon usage. The off-frame preferences were inferred from the nonrandomness of the location of confamilial synonymous codons along coding regions-a pattern often described as a context dependence of nucleotide choice at synonymous positions or as codon-pair bias. We relied on randomizations of the location of confamilial codons that do not alter, and cannot be influenced by, the encoded amino acid sequences, codon usage, or base composition of the genes examined. The statistically significant congruency of in-frame and off-frame trinucleotide preferences suggests that the same kind of reading-frame-independent force(s) may also influence synonymous codon choice. These forces may have produced biases in codon usage that then led to the evolution of the translational advantages of these motifs as preferred codons. Under this scenario, tRNA pool size differences between preferred and nonpreferred codons initially were evolved to track the default overrepresentation of codons with preferred motifs. The motif preference hypothesis can explain the structuring of codon preferences and the similarities in the codon usages of distantly related organisms.

Amino Acid Sequence↗

The evolution of codon preferences in Drosophila: a maximum-likelihood approach to parameter estimation and hypothesis testing.

Synonymous codon usage in related species may differ as a result of variation in mutation biases, differences in the overall strength and efficiency of selection, and shifts in codon preference-the selective hierarchy of codons within and between amino acids. We have developed a maximum-likelihood method to employ explicit population genetic models to analyze the evolution of parameters determining codon usage. The method is applied to twofold degenerate amino acids in 50 orthologous genes from D. melanogaster and D. virilis. We find that D. virilis has significantly reduced selection on codon usage for all amino acids, but the data are incompatible with a simple model in which there is a single difference in the long-term Ne, or overall strength of selection, between the two species, indicating shifts in codon preference. The strength of selection acting on codon usage in D. melanogaster is estimated to be |Nes| approximately 0.4 for most CT-ending twofold degenerate amino acids, but 1.7 times greater for cysteine and 1.4 times greater for AG-ending codons. In D. virilis, the strength of selection acting on codon usage for most amino acids is only half that acting in D. melanogaster but is considerably greater than half for cysteine, perhaps indicating the dual selection pressures of translational efficiency and accuracy. Selection coefficients in orthologues are highly correlated (rho = 0.46), but a number of genes deviate significantly from this relationship.

Amino Acids↗

A theoretical analysis of codon adaptation index of the Boophilus microplus bm86 gene directed to the optimization of a DNA vaccine.

DNA vaccines utilize host cell molecules for gene transcription and translation to proteins, and the interspecific difference of codon usage is one of the major obstacles for effective induction of specific and strong immune response. In an attempt to improve codon usage effects of DNA vaccine on protein expression, a quantitative study was conducted to clarify the relationship of codon usage in the tick gene bm86 and its potential expression in bovine cells. The calculated relative synonymous codon usage (RSCU) and codon adaptation index (CAI) values of bm86 from Boophilus microplus and a set of 14 highly expressed genes from Bos taurus indicated that some codons utilized frequently in bm86 are rarely used in B. taurus genes and vice versa. The different translational efficiencies obtained suggested that after DNA vaccination using the wild bm86 gene, the protein Bm86 would be expressed in bovines, but it would not be the optimum sequence. However, using the codon-optimized bm86 gene to bovines, whose sequence was theoretically designed, would probably improve the level of the immune response generated against ticks.

Amino Acid Sequence↗

Revisiting the codon adaptation index from a whole-genome perspective: analyzing the relationship between gene expression and codon occurrence in yeast using a variety of models.

Highly expressed genes in many bacteria and small eukaryotes often have a strong compositional bias, in terms of codon usage. Two widely used numerical indices, the codon adaptation index (CAI) and the codon usage, use this bias to predict the expression level of genes. When these indices were first introduced, they were based on fairly simple assumptions about which genes are most highly expressed: the CAI was originally based on the codon composition of a set of only 24 highly expressed genes, and the codon usage on assumptions about which functional classes of genes are highly expressed in fast-growing bacteria. Given the recent advent of genome-wide expression data, we should be able to improve on these assumptions. Here, we measure, in yeast, the degree to which consideration of the current genome-wide expression data sets improves the performance of both numerical indices. Indeed, we find that by changing the parameterization of each model its correlation with actual expression levels can be somewhat improved, although both indices are fairly insensitive to the exact way they are parameterized. This insensitivity indicates a consistent codon bias amongst highly expressed genes. We also attempt direct linear regression of codon composition against genome-wide expression levels (and protein abundance data). This has some similarity with the CAI formalism and yields an alternative model for the prediction of expression levels based on the coding sequences of genes. More information is available at http://bioinfo.mbb.yale.edu/expression/codons.

Codon↗

Genetic robustness and selection at the protein level for synonymous codons.

Synonymous codons are neutral at the protein level, therefore natural selection at the protein level should have no effect on their frequencies. Synonymous codons, however, differ in their capacity to reduce the effects of errors: after mutation, certain codons keep on coding for the same amino acid or for amino acids with similar properties, while other synonymous codons produce very different amino acids. Therefore, the impact of errors on a coding sequence (genetic robustness) can be measured by analysing its codon usage. I analyse the codon usage of sequenced nuclear and cytoplasmic genomes and I show that there is an extensive variation in genetic robustness at the DNA sequence level, both among genomes and among genes of the same genome. I also show theoretically that robustness can be adaptive, that is natural selection may lead to a preference for codons that reduce the impact of errors. If selection occurs only among the mutants of a codon (e.g. among the progeny before the adult phase), however, the codons that are more sensitive to the effects of mutations may increase in frequency because they manage to get rid more easily of deleterious mutations. I also suggest other possible explanations for the evolution of genetic robustness at the codon level.

Amino Acid Sequence↗

Gene expression levels influence amino acid usage and evolutionary rates in endosymbiotic bacteria.

Most endosymbiotic bacteria have extremely reduced genomes, accelerated evolutionary rates, and strong AT base compositional bias thought to reflect reduced efficacy of selection and increased mutational pressure. Here, we present a comparative study of evolutionary forces shaping five fully sequenced bacterial endosymbionts of insects. The results of this study were three-fold: (i) Stronger conservation of high expression genes at not just nonsynonymous, but also synonymous, sites. (ii) Variation in amino acid usage strongly correlates with GC content and expression level of genes. This pattern is largely explained by greater conservation of high expression genes, leading to their higher GC content. However, we also found indication of selection favoring GC-rich amino acids that contrasts with former studies. (iii) Although the specific nutritional requirements of the insect host are known to affect gene content of endosymbionts, we found no detectable influence on substitution rates, amino acid usage, or codon usage of bacterial genes involved in host nutrition.

AT Rich Sequence↗

Translational features of human alpha 2b interferon production in Escherichia coli.

The yield of human alpha 2b interferon in Escherichia coli was optimized by replacement of low-usage arginine codons located in the mRNA 5' end. The differences observed among the various gene variants suggest that codon usage, Shine-Dalgarno-like sequences, and mRNA secondary structure contribute to the performance of E. coli translation machinery.

Amino Acid Sequence↗

Beta tubulin gene of the parasitic protozoan Leishmania mexicana.

A genomic DNA library was generated with Sau3A cut DNA derived from promastigotes of Leishmania mexicana amazonensis and the lambda vector EMBL3. The library was screened for beta tubulin clones using 32P-labeled heterologous probe of chicken beta tubulin cDNA. From the various genomic clones the one designated 23.1, which gave the simplest hybridization banding pattern, was further characterized. The leishmanial insert DNA was subcloned into plasmid vectors and the resulting clones were designated as T11, T28 and T50. Using these clones leishmanial beta tubulin coding region was sequenced by the dideoxy method. The result shows that the beta tubulin has 445 amino acids, a carboxyl terminal tyrosine, and no intron. Leishmanial beta tubulin has 93% amino acid sequence similarity with that of trypanosome and 82% with that of man: and there is a strong bias in codon usage for codons possessing guanine or cytosine in the third base.

Amino Acid Sequence↗

Production of (10E,12Z)-conjugated linoleic acid in yeast and tobacco seeds.

The polyenoic fatty acid isomerase from Propioniumbacterium acnes (PAI) was expressed in E. coli and biochemically characterized. PAI catalyzes the isomerization of a methylene-interrupted double bond system to a conjugated double bond system, creating (10E,12Z)-conjugated linoleic acid (CLA). PAI accepted a wide range of free polyunsaturated fatty acids as substrates ranging from 18:2 fatty acids to 22:6, converting them to fatty acids with two or three conjugated double bonds. For expression of PAI in yeast the PAI-sequence encoding 20 N-terminal amino acid residues was altered for optimal codon usage, yielding codon optimized PAI (coPAI). The percentage of 10,12-CLA of total esterified fatty acids was 8 times higher in yeast transformed with coPAI than in cells transformed with PAI. CLA was detected in amounts up to 5.7% of total free fatty acids in yeast transformed with coPAI but none was detected in yeast transformed with PAI. PAI or coPAI under the control of the constitutive CaMV 35S promoter or the seed-specific USP promoter was transformed into tobacco plants. CLA was only detected in seeds in coPAI-transgenic plants. The amount of CLA detected in esterified fatty acids was up to 0.3%, in free fatty acids up to 15%.

Actinomycetales↗