PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

On the coevolution of genes and genetic code.

The canonical genetic code acts efficiently in minimizing the effects of mistranslations and point mutations. In the work presented we have also considered the effects of single nucleotide insertions and deletions on the optimality of the genetic code. Our results suggest that the canonical genetic code compensates for the ins/del mutations as well as mistranslations and point mutations. On the other hand, we highlighted the point that ins/del mutations have a lesser impact on the selected genes of Saccharomyces cerevisiae compared to randomly generated ones. We hypothesized that the codon usage preferences in S. cerevisiae genes are responsible for the higher efficiency of translation machinery in this organism. Our results support the conjecture that codon usage preferences render the genetic code more effective in minimizing the effects of ins/del mutations.

Animals↗

Gene expression and molecular evolution.

The combination of complete genome sequence information and estimates of mRNA abundances have begun to reveal causes of both silent and protein sequence evolution. Translational selection appears to explain patterns of synonymous codon usage in many prokaryotes as well as a number of eukaryotic model organisms (with the notable exception of vertebrates). Relationships between gene length and codon usage bias, however, remain unexplained. Intriguing correlations between expression patterns and protein divergence suggest some general mechanisms underlying protein evolution.

Animals↗

Characterization of In0 of Pseudomonas aeruginosa plasmid pVS1, an ancestor of integrons of multiresistance plasmids and transposons of gram-negative bacteria.

Many multiresistance plasmids and transposons of gram-negative bacteria carry related DNA elements that appear to have evolved from a common ancestor by site-specific integration of discrete cassettes containing antibiotic resistance genes or sequences of unknown function. The site of integration is flanked by conserved segments coding for an integraselike protein and for sulfonamide resistance, respectively. These segments, together with the antibiotic resistance genes between them, have been termed integrons (H. W. Stokes and R. M. Hall, Mol. Microbiol. 3:1669-1683, 1989). We report here the characterization of an integron, In0, from Pseudomonas aeruginosa plasmid pVS1, which has an unoccupied integration site and hence may be an ancestor of more complex integrons. Codon usage of the integrase (int) and sulfonamide resistance (sul1) genes carried by this integron suggests a common origin. This contrasts with the codon usage of other antibiotic resistance genes that were presumably integrated later as cassettes during the evolution and spread of these DNA elements. We propose evolutionary schemes for (i) the genesis of the integrons by the site-specific integration of antibiotic resistance genes and (ii) the evolution of the integrons of multiresistance plasmids and transposons, in relation to the evolution of transposons related to Tn21.

Amino Acid Sequence↗

Correlations between mRNA expression levels and GC contents of coding and untranslated regions of genes in rodents.

Gene expression is regulated by a highly coordinated network of events whose efficiency may constrain the level of expression. Among other factors, natural selection for increased translational efficiency and/or fidelity may shape nucleotide composition and, hence, codon usage during evolution. Previous studies have shown that highly expressed genes in Saccharomyces cerevisiae, Caenorhabditis elegans, and Drosophila melanogaster have relatively higher codon usage biases. However, in the case of mammals, results have been equivocal. In this study, we assessed the correlation between nucleotide composition and mRNA expression levels of rodent genes measured by cDNA microarray and serial analysis of gene expression (SAGE) techniques. We found that mRNA expression levels were correlated with the third nucleotide position GC (GC3) content for both Rattus norvegicus (r = 0.246, p = 0.01; N = 110) and Mus musculus (r = 0.21, p = 0.0026; N = 203) genes. However, no significant correlation was evident between mRNA expression level and GC contents of 5'- and 3'-untranslated regions (UTRs) for either species. This suggests that, in rodents, nucleotide composition of coding sequences and UTRs might evolve differentially when considered along an expression gradient. Accordingly, it is possible that higher GC levels may present the rodent genes with a selective advantage for translational efficiency. However, the increase in GC3 content seems to level off above an expressional threshold (e.g., >or=threefold the median expression for R. norvegicus), suggesting that conflicting demands posed by different aspects of transcriptional and translational machineries (e.g., efficiency versus fidelity) may set an upper limit for GC3.

Animals↗

Overproduction from a cellulase gene with a high guanosine-plus-cytosine content in Escherichia coli.

A recombinant exoglucanase was expressed in Escherichia coli to a level that exceeded 20% of total cellular protein. To obtain this level of overproduction, the exoglucanase gene coding sequence was fused to a synthetic ribosome-binding site, an initiating ATG, and placed under the control of the leftward promoter of bacteriophage lambda contained on the runaway replication plasmid vector pCP3 (E. Remaut, H. Tsao, and W. Fiers, Gene 22:103-113, 1983). With the exception of an inserted asparagine adjacent to the initiating ATG, the highly expressed exoglucanase is identical to the native exoglucanase. The overproduced exoglucanase can be isolated easily in an enriched form as insoluble aggregates, and exoglucanase activity can be recovered by solubilization of the aggregates in 6 M urea or 5 M guanidine hydrochloride. Since the codon usage of the exoglucanase gene is so markedly different from that of E. coli genes, the overproduction of the exoglucanase in E. coli indicates that codon usage may not be a major barrier to heterospecific gene expression in this organism.

Actinomycetales↗

Gene gun DNA vaccination with Rev-independent synthetic HIV-1 gp160 envelope gene using mammalian codons.

DNA immunization with HIV envelope plasmids induce only moderate levels of specific antibodies which may in part be due to limitations in expression influenced by a species-specific and biased HIV codon usage. We compared antibody levels, Th1/Th2 type and CTL responses induced by synthetic genes encoding membrane bound gp160 versus secreted gp120 using optimized codons and the efficient gene gun immunization method. The in vitro expression of syn.gp160 as gp120 + gp41 was Rev independent and much higher than a classical wt.gp160 plasmid. Mice immunized with syn.gp160 and wt.gp160 generated low and inconsistent ELISA antibody titres whereas the secreted gp120 consistently induced faster seroconversion and higher antibody titres. Due to a higher C + G content the numbers of putative CpG immune (Th1) stimulatory motifs were highest in the synthetic gp160 gene. However, both synthetic genes induced an equally strong and more pronounced Th2 response with higher IgG1/IgG2a and IFNgamma/IL-4 ratios than the wt.gp160 gene. As for induction of CTL, synthetic genes induced a somewhat earlier response but did not offer any advantage over wild type genes at a later time point. Thus, optimizing codon usage has the advantage of rendering the structural HIV genes Rev independent. For induction of antibodies the level of expression, while important, seems less critical than optimal contact with antigen presenting cells at locations reached by the secreted gp120 protein. A proposed Th1 adjuvant effect of the higher numbers of CpG motifs in the synthetic genes was not seen using gene gun immunization which may be due to the low amount of DNA used.

AIDS Vaccines↗

Comparison of synonymous codon distribution patterns of bacteriophage and host genomes.

Synonymous codon usage patterns of bacteriophage and host genomes were compared. Two indexes, G + C base composition of a gene (fgc) and fraction of translationally optimal codons of the gene (fop), were used in the comparison. Synonymous codon usage data of all the coding sequences on a genome are represented as a cloud of points in the plane of fop vs. fgc. The Escherichia coli coding sequences appear to exhibit two phases, "rising" and "flat" phases. Genes that are essential for survival and are thought to be native are located in the flat phase, while foreign-type genes from prophages and transposons are found in the rising phase with a slope of nearly unity in the fgc vs. fop plot. Synonymous codon distribution patterns of genes from temperate phages P4, P2, N15 and lambda are similar to the pattern of E. coli rising phase genes. In contrast, genes from the virulent phage T7 or T4, for which a phage-encoded DNA polymerase is identified, fall in a linear curve with a slope of nearly zero in the fop vs. fgc plane. These results may suggest that the G + C contents for T7, T4 and E. coli flat phase genes are subject to the directional mutation pressure and are determined by the DNA polymerase used in the replication. There is significant variation in the fop values of the phage genes, suggesting an adjustment to gene expression level. Similar analyses of codon distribution patterns were carried out for Haemophilus influenzae, Bacillus subtilis, Mycobacterium tuberculosis and their phages with complete genomic sequences available.

Bacillus subtilis↗

Intragenic Hill-Robertson interference influences selection intensity on synonymous mutations in Drosophila.

Natural selection influences synonymous mutations and synonymous codon usage in many eukaryotes to improve the efficiency of translation in highly expressed genes. Recent studies of gene composition in eukaryotes have shown that codon usage also varies independently of expression levels, both among genes and at the intragenic level. Here, we investigate rates of evolution (Ks) and intensity of selection (gamma(s)) on synonymous mutations in two groups of genes that differ greatly in the length of their exons, but with equivalent levels of gene expression and rates of crossing-over in Drosophila melanogaster. We estimate gamma(s) using patterns of divergence and polymorphism in 50 Drosophila genes (100 kb of coding sequence) to take into account possible variation in mutation trends across the genome, among genes or among codons. We show that genes with long exons exhibit higher Ks and reduced gamma(s) compared to genes with short exons. We also show that Ks and gamma(s) vary significantly across long exons, with higher Ks and reduced gamma(s) in the central region compared to flanking regions of the same exons, hence indicating that the difference between genes with short and long exons can be mostly attributed to the central region of these long exons. Although amino acid composition can also play a significant role when estimating Ks and gamma(s), our analyses show that the differences in Ks and gamma(s) between genes with short and long exons and across long exons cannot be explained by differences in protein composition. All these results are consistent with the Interference Selection (IS) model that proposes that the Hill-Robertson (HR) effect caused by many weakly selected mutations has detectable evolutionary consequences at the intragenic level in genomes with recombination. Under the IS model, exon size and exon-intron structure influence the effectiveness of selection, with long exons showing reduced effectiveness of selection when compared to small exons and the central region of long exons showing reduced intensity of selection compared to flanking coding regions. Finally, our results further stress the need to consider selection on synonymous mutations and its variation--among and across genes and exons--in studies of protein evolution.

Animals↗

Highly expressed and alien genes of the Synechocystis genome.

Comparisons of codon frequencies of genes to several gene classes are used to characterize highly expressed and alien genes on the SYNECHOCYSTIS: PCC6803 genome. The primary gene classes include the ensemble of all genes (average gene), ribosomal protein (RP) genes, translation processing factors (TF) and genes encoding chaperone/degradation proteins (CH). A gene is predicted highly expressed (PHX) if its codon usage is close to that of the RP/TF/CH standards but strongly deviant from the average gene. Putative alien (PA) genes are those for which codon usage is significantly different from all four classes of gene standards. In SYNECHOCYSTIS:, 380 genes were identified as PHX. The genes with the highest predicted expression levels include many that encode proteins vital for photosynthesis. Nearly all of the genes of the RP/TF/CH gene classes are PHX. The principal glycolysis enzymes, which may also function in CO(2) fixation, are PHX, while none of the genes encoding TCA cycle enzymes are PHX. The PA genes are mostly of unknown function or encode transposases. Several PA genes encode polypeptides that function in lipopolysaccharide biosynthesis. Both PHX and PA genes often form significant clusters (operons). The proteins encoded by PHX and PA genes are described with respect to functional classifications, their organization in the genome and their stoichiometry in multi-subunit complexes.

Codon↗

Analysis of messages expressed by Echinostoma paraensei miracidia and sporocysts, obtained by random EST sequencing.

A lambdaZAP Express cDNA library was constructed with mRNA obtained from immature miracidia within eggs, hatched miracidia, and sporocysts of Echinostoma paraensei. This cDNA library was amplified and 213 expressed sequence tag (EST) sequences (averaging 466 nucleotides in length) were obtained. The mean percentage of unresolved bases within the EST sequences was 0.4%, ranging from 0 to 4.6%. The 213 ESTs represent 151 unique messages. BLAST (version 2.0.8) analysis disclosed that 64 unique E. paraensei messages (42.4%) had significant similarities (BLAST score < or =e-5), at deduced amino acid or nucleotide levels, with known sequences in the nonredundant GenBank databases or the dbEST database (NCBI). The remainder, 57.6% of the unique EST-encoded messages, scored nonsignificant hits. Most of the E. paraensei messages that could be assigned a cellular role based on sequence similarities were involved in gene/protein expression. Several ESTs scored highest similarities with sequences obtained from trematode species. A total of 22,560 nucleotides present in open reading frames from ESTs that aligned with known sequences was used to determine codon usage for E. paraensei. Analysis of a subset of eight ESTs that contained full-length open reading frames did not reveal a bias in codon usage. Also, EST sequences were found to contain 3' untranslated regions with an average length of 69.9 +/- 88.4 nucleotides (n = 46). The EST sequences were submitted to GenBank/dbEST, adding to the 51 available Echinostoma-derived sequences, to provide reference information for both phylogenetic analysis and study of general trematode biology.

Animals↗

Periodicity of DNA in exons.

BACKGROUND: The periodic pattern of DNA in exons is a known phenomenon. It was suggested that one of the initial causes of periodicity could be the universal (RNY)npattern (R = A or G, Y = C or U, N = any base) of ancient RNA. Two major questions were addressed in this paper. Firstly, the cause of DNA periodicity, which was investigated by comparisons between real and simulated coding sequences. Secondly, quantification of DNA periodicity was made using an evolutionary algorithm, which was not previously used for such purposes. RESULTS: We have shown that simulated coding sequences, which were composed using codon usage frequencies only, demonstrate DNA periodicity very similar to the observed in real exons. It was also found that DNA periodicity disappears in the simulated sequences, when the frequencies of codons become equal. Frequencies of the nucleotides (and the dinucleotide AG) at each location along phase 0 exons were calculated for C. elegans, D. melanogaster and H. sapiens. Two models were used to fit these data, with the key objective of describing periodicity. Both of the models showed that the best-fit curves closely matched the actual data points. The first dynamic period determination model consistently generated a value, which was very close to the period equal to 3 nucleotides. The second fixed period model, as expected, kept the period exactly equal to 3 and did not detract from its goodness of fit. CONCLUSIONS: Conclusion can be drawn that DNA periodicity in exons is determined by codon usage frequencies. It is essential to differentiate between DNA periodicity itself, and the length of the period equal to 3. Periodicity itself is a result of certain combinations of codons with different frequencies typical for a species. The length of period equal to 3, instead, is caused by the triplet nature of genetic code. The models and evolutionary algorithm used for characterising DNA periodicity are proven to be an effective tool for describing the periodicity pattern in a species, when a number of exons in the same phase are analysed.

Algorithms↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

Fine structural features of the chloroplast genome: comparison of the sequenced chloroplast genomes.

The entire nucleotide sequences of the rice, tobacco and liverwort chloroplast genomes have been determined. We compared all the chloroplast genes, open reading frames and spacer regions in the plastid genomes of these three species in order to elucidate general structural features of the chloroplast genome. Analyses of homology, GC content and codon usage of the genes enabled us to classify them into two groups: photosynthesis genes and genetic system genes. Based on comparisons of homology, GC content and codon usage, unidentified ORFs can also be assigned to each of these groups such that it is possible to speculate about the functions of products which may be produced by these ORFs. The spacer regions and intron sequences were compared and found to have no obvious homology between rice and liverwort or between tobacco and liverwort.

Base Composition↗

Variation in G + C-content and codon choice: differences among synonymous codon groups in vertebrate genes.

The relationship between G + C-content and codon usage in genes of human, mus, rat, bovine and chicken nuclear genomes was investigated. Correlation and lineal regression analyses were carried out on plots that related the frequency of each codon within each synonymous codon group to the G + C-content of the coding sequence as a whole. Under GC pressure, in most of the quartet codon groups there is a preferential choice of the C-ending codon, except in leucine and valine codon groups where the choice of the G-ending codon is preferred. Among ducts, the choice of codons specifying phenylalanine and glutamate shows the strongest dependence on G + C-content. The relationship found between G + C-content and codon usage in these genomes correlate with taxonomic distance.

Animals↗

Modeling sequencing errors by combining Hidden Markov models.

Among the largest resources for biological sequence data is the large amount of expressed sequence tags (ESTs) available in public and proprietary databases. ESTs provide information on transcripts but for technical reasons they often contain sequencing errors. Therefore, when analyzing EST sequences computationally, such errors must be taken into account. Earlier attempts to model error prone coding regions have shown good performance in detecting and predicting these while correcting sequencing errors using codon usage frequencies. In the research presented here, we improve the detection of translation start and stop sites by integrating a more complex mRNA model with codon usage bias based error correction into one hidden Markov model (HMM), thus generalizing this error correction approach to more complex HMMs. We show that our method maintains the performance in detecting coding sequences.

Algorithms↗

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids↗

Genotypic variation in the pol gene of HIV type 1 in an antiretroviral treatment-naive population in rural southwestern Uganda.

The majority of studies of HIV-1 drug resistance have involved subtype B viruses. Here we have characterized subtype distribution and determined the levels of polymorphism at protease (PR) and reverse transcriptase (RT) drug resistance positions, in antiretroviral treatment-naive HIV-positive Ugandan patients. We have also investigated codon usage variability at these positions and assessed intersubtype recombination within the pol gene. The study population consisted of 187 patients, from a cohort established by the UK Medical Research Council Programme on AIDS in Uganda in 1990. Results indicate that 28.3% of patients were infected with subtype A (n = 53), 64.2% subtype D (n = 120), 6.4% A/D recombinant (n = 12), and 1.1% subtype C (n = 2). Variation in amino acid usage at drug resistance-associated positions was minimal between the two main subtypes (A and D) in RT, but there was appreciable variation in PR. Codon usage, however, was considerably more variable between subtypes A and D in both PR and RT. Thus, while no natural high-level resistance to antiretroviral therapy was detected in this cohort, subtypes A and D may possess different genetic barriers to be overcome in order to achieve resistance. With the increasing introduction of antiretroviral therapy into Africa, such information will be vital in our understanding and evaluation of the development of drug resistance as it occurs, and how to interpret resistance data the type of which has rarely previously been seen. This analysis also significantly increases the number of Ugandan PR and RT sequences characterized to date.

Amino Acid Substitution↗

Exploring the potential of the bacterial carotene desaturase CrtI to increase the beta-carotene content in Golden Rice.

To increase the beta-carotene (provitamin A) content and thus the nutritional value of Golden Rice, the optimization of the enzymes employed, phytoene synthase (PSY) and the Erwinia uredovora carotene desaturase (CrtI), must be considered. CrtI was chosen for this study because this bacterial enzyme, unlike phytoene synthase, was expressed at barely detectable levels in the endosperm of the Golden Rice events investigated. The low protein amounts observed may be caused by either weak cauliflower mosaic virus 35S promoter activity in the endosperm or by inappropriate codon usage. The protein level of CrtI was increased to explore its potential for enhancing the flux of metabolites through the pathway. For this purpose, a synthetic CrtI gene with a codon usage matching that of rice storage proteins was generated. Rice plants were transformed to express the synthetic gene under the control of the endosperm-specific glutelin B1 promoter. In addition, transgenic plants expressing the original bacterial gene were generated, but the endosperm-specific glutelin B1 promoter was employed instead of the cauliflower mosaic virus 35S promoter. Independent of codon optimization, the use of the endosperm-specific promoter resulted in a large increase in bacterial desaturase production in the T(1) rice grains. However, this did not lead to a significant increase in the carotenoid content, suggesting that the bacterial enzyme is sufficiently active in rice endosperm even at very low levels and is not rate-limiting. The endosperm-specific expression of CrtI did not affect the carotenoid pattern in the leaves, which was observed upon its constitutive expression. Therefore, tissue-specific expression of CrtI represents the better option.

Bacterial Proteins↗