PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Characterization of a highly expressed lignin peroxidase-encoding gene from the basidiomycete Phanerochaete chrysosporium.

The genomic clone, LG2, encoding LiP2, the major lignin peroxidase (LiP) isozyme from Phanerochaete chrysosporium strain OGC101, was isolated and characterized. The 5'-untranslated region of LG2 contains sequences similar to CRE and XRE promoter elements. Comparison with its transcript indicates that eight introns, each less than 59 bp, interrupt the coding sequence. Comparison with genes encoding other LiP isozymes shows five related patterns of intron location, whose incidence coincides with described LiP structural subfamilies. Codon bias indices calculated for all known P. chrysosporium genes, including trpC and genes encoding LiP, MnP, and exo-cellobiohydrolase I, demonstrate that LG2 has the most biased codon usage. We conclude that subdivisions of the LiP family may be based on intron location in the encoding genes, and that ranking of isozyme production levels can be estimated by the extent of bias in codon usage in the cognate gene.

Amino Acid Sequence↗

Increasing expression of P450 and P450-reductase proteins from monocots in heterologous systems.

Monocotyledonous crop plants are usually more resistant to herbicides than grass weeds and most dicots. Their resistance to herbicides is mediated in many cases by P450 oxygenases. Monocots thus constitute an appealing source of P450 enzymes for manipulating herbicide resistance and recombinant forms of the major xenobiotic metabolizing mooxygenases are potential tools for the optimization of new active molecules. We report here the isolation and functional characterization of the first P450 and P450 reductase coding sequences from wheat. The first attempts at expressing these cDNAs in yeast and tobacco led to levels of protein, which were extremely low, often not even detectable. The wheat P450 cDNAs were efficiently transcribed, but no protein or activity was found. Wheat coding sequences, like those of other monocots, are characterized by a high GC content and by a related strong bias of codon usage, different from that observed in yeast or dicots. Complete recoding of genes being costly, the reengineering their 5'-end using a single PCR megaprimer designed to comply with codon usage of the host was attempted. It was sufficient to relieve translation inhibition and to obtain good levels of protein expression. The same strategy also resulted in a dramatic increase in protein expression in tobacco. A basis for the success of such a partial recoding strategy, much easier and cheaper than complete recoding of the cDNA, is proposed.

Amino Acid Sequence↗

The nucleotide sequence of an Escherichia coli operon containing genes for the tRNA(m1G)methyltransferase, the ribosomal proteins S16 and L19 and a 21-K polypeptide.

The nucleotide sequence of a 4.6-kb SalI-EcoRI DNA fragment including the trmD operon, located at min 56 on the Escherichia coli K-12 chromosome, has been determined. The trmD operon encodes four polypeptides: ribosomal protein S16 (rpsP), 21-K polypeptide (unknown function), tRNA-(m1G)methyltransferase (trmD) and ribosomal protein L19 (rplS), in that order. In addition, the 4.6-kb DNA fragment encodes a 48-K and a 16-K polypeptide of unknown functions which are not part of the trmD operon. The mol. wt. of tRNA(m1G)methyltransferase determined from the DNA sequence is 28 424. The probable locations of promoter and terminator of the trmD operon are suggested. The translational start of the trmD gene was deduced from the known NH2-terminal amino acid sequence of the purified enzyme. The intercistronic regions in the operon vary from 9 to 40 nucleotides, supporting the earlier conclusion that the four genes are co-transcribed, starting at the major promoter in front of the rpsP gene. Since it is known that ribosomal proteins are present at 8000 molecules/genome and the tRNA-(m1G)methyltransferase at only approximately 80 molecules/genome in a glucose minimal culture, some powerful regulatory device must exist in this operon to maintain this non-coordinate expression. The codon usage of the two ribosomal protein genes is similar to that of other ribosomal protein genes, i.e., high preference for the most abundant tRNA isoaccepting species. The trmD gene has a codon usage typical for a protein made in low amount in accordance with the low number of tRNA-(m1G)methyltransferase molecules found in the cell.

Bacterial Proteins↗

The cytochrome b region in the mitochondrial DNA of the ant Tetraponera rufoniger: sequence divergence in Hymenoptera may be associated with nucleotide content.

Polymerase chain reaction (PCR) followed by sequencing of single-stranded DNA yielded sequence information from the cytochrome b (cyt b) region in mitochondrial DNA from the ant Tetraponera rufoniger. Compared with the cyt b genes from Apis mellifera, Drosophila melanogaster, and D. yakuba, the overall A+T content (A+T%) of that of T. rufoniger is lower (69.9% vs 80.7%, 74.2%, and 73.9%, respectively) than those of the other three. The codon usage in the cyt b gene of T. rufoniger is biased although not as much as in A. mellifera, D. melanogaster, and D. yakuba; T. rufoniger has eight unused codons whereas D. melanogaster, D. yakuba, and A. mellifera have 21, 20, and 23, respectively. The inferred cyt b polypeptide chain (PPC) of T. rufoniger has diverged at least as much from a common ancestor with D. yakuba as has that of A. mellifera (approximately 3.5 vs approximately 2.9). Despite the lower A+T%, the relative frequencies of amino acids in the cyt b PPC of T. rufoniger are significantly (P < 0.05) associated with the content of adenine and thymine (A+T%) and size of codon families. The mitochondrially located cytochrome oxidase subunit II genes (CO-II) of endopterygote insects have significantly higher average A+T% (approximately 75%) than those of exopterygous (approximately 69%) and paleopterous (approximately 69%) insects. The increase in A+T% of endopterygote insects occurred in Upper Carboniferous and coincided with a significant acceleration of PPC divergence. However, acceleration of PPC divergence is not significantly correlated with the increase of the A+T% (P > 0.1). The high A+T%, the biased codon usage, and the increased PPC divergence of Hymenoptera can in that respect most easily be explained by directional mutation pressure which began in the Upper Carboniferous and still occurs in most members of the order. Given the roughly identical A+T% of the cyt b and CO-II genes from the other insects whose DNA sequences are known (A. mellifera, D. melanogaster, and D. yakuba), it seems most likely that the A+T% of T. rufoniger declined secondarily within the last 100 Myr as a result of a reduced directional mutation pressure.

Amino Acid Sequence↗

Evolutionary lability of context-dependent codon bias in bacteria.

In bacteria, synonymous codon usage can be considerably affected by base composition at neighboring sites. Such context-dependent biases may be caused by either selection against specific nucleotide motifs or context-dependent mutation biases. Here we consider the evolutionary conservation of context-dependent codon bias across 11 completely sequenced bacterial genomes. In particular, we focus on two contextual biases previously identified in Escherichia coli; the avoidance of out-of-frame stop codons and AGG motifs. By identifying homologues of E. coli genes, we also investigate the effect of gene expression level in Haemophilus influenzae and Mycoplasma genitalium. We find that while context-dependent codon biases are widespread in bacteria, few are conserved across all species considered. Avoidance of out-of-frame stop codons does not apply to all stop codons or amino acids in E. coli, does not hold for different species, does not increase with gene expression level, and is not relaxed in Mycoplasma spp., in which the canonical stop codon, TGA, is recognized as tryptophan. Avoidance of AGG motifs shows some evolutionary conservation and increases with gene expression level in E. coli, suggestive of the action of selection, but the cause of the bias differs between species. These results demonstrate that strong context-dependent forces, both selective and mutational, operate on synonymous codon usage but that these differ considerably between genomes.

Codon↗

Effect of codon optimization on expression levels of a functionally folded malaria vaccine candidate in prokaryotic and eukaryotic expression systems.

We have produced two synthetic genes that code for the F2 domain located within region II of the 175-kDa Plasmodium falciparum erythrocyte binding antigen (EBA-175) to determine the effects of codon alteration on protein expression in homologous and heterologous host systems. EBA-175 plays a key role in the process of merozoite invasion into erythrocytes through a specific receptor-ligand interaction. The F2 domain of EBA-175 is the ligand that binds to the glycophorin A receptor on human erythrocytes and is therefore a target of vaccine development efforts. We designed synthetic genes based on P. falciparum, Escherichia coli, and Pichia codon usage and expressed recombinant F2 in E. coli and Pichia pastoris. Compared to the expression of the native F2 sequence, conversion to prokaryote (E. coli)- or eukaryote (Pichia)-based codon usage dramatically improved the levels of recombinant protein expression in both E. coli and P. pastoris. The majority of the protein expressed in E. coli, however, was produced as inclusion bodies. The protein expressed in P. pastoris, on the other hand, was expressed as a secreted, soluble protein. The P. pastoris-produced protein was superior to that produced in E. coli based on its ability to bind to red blood cells. Consistent with these observations, the antibodies generated against the Pichia-produced protein prevented the binding of recombinant EBA to red blood cells. These antibodies recognize EBA-175 present on merozoites as well as in sporozoites by immunofluorescence. Our results suggest that the Pichia-based EBA-F2 vaccine construct has further potential to be developed for clinical use.

Animals↗

A model of protein translation including codon bias, nonsense errors, and ribosome recycling.

We present and analyse a model of protein translation at the scale of an individual messenger RNA (mRNA) transcript. The model we develop is unique in that it incorporates the phenomena of ribosome recycling and nonsense errors. The model conceptualizes translation as a probabilistic wave of ribosome occupancy traveling down a heterogeneous medium, the mRNA transcript. Our results show that the heterogeneity of the codon translation rates along the mRNA results in short-scale spikes and dips in the wave. Nonsense errors attenuate this wave on a longer scale while ribosome recycling reinforces it. We find that the combination of nonsense errors and codon usage bias can have a large effect on the probability that a ribosome will completely translate a transcript. We also elucidate how these forces interact with ribosome recycling to determine the overall translation rate of an mRNA transcript. We derive a simple cost function for nonsense errors using our model and apply this function to the yeast (Saccharomyces cervisiae) genome. Using this function we are able to detect position dependent selection on codon bias which correlates with gene expression levels as predicted a priori. These results indirectly validate our underlying model assumptions and confirm that nonsense errors can play an important role in shaping codon usage bias.

Animals↗

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family↗

Chloramphenicol resistance in Campylobacter coli: nucleotide sequence, expression, and cloning vector construction.

A chloramphenicol-resistance determinant (CmR), originally cloned from Campylobacter coli plasmid pNR9589 in Japan, was isolated and the nucleotide sequence determined, which contained an open reading frame of 621 bp. The gene product was identified as Cm acetyltransferase (CAT), which had a putative amino acid sequence that showed 43% to 57% identity with other CAT proteins of both Gram+ and Gram- origin. Although expression of the cat gene was constitutive in both C. coli and Escherichia coli, results of primer extension experiments indicated that transcription was initiated at different sites in these two species. A kanamycin-resistance determinant, identified as the aphA-3 gene, was located downstream from the cat gene. The codon usage of the cat gene is very different from that used in E. coli, however, the CAT polypeptide was synthesized in large amounts in E. coli maxicells. Therefore, the codon usage bias is not one of the obstacles which affects Campylobacter spp. gene expression in E. coli. New Campylobacter cloning vectors were constructed in this study.

Amino Acid Sequence↗

Theory of degenerate coding and informational parameters of protein coding genes.

The theory of degenerate coding is presented in a way enabling further application to molecular biology. There are two kinds of redundancy of a degenerate code. The first is due to the excess in codon length and the second to the code degeneracy. If the code is asymmetrically degenerate, the second kind of redundancy can be profitable for control of error rate. This control can be performed just by selective synonymous codon usage. Utilisation of the genetic code is partially influenced by this theoretical possibility. In particular the degree of error protectivity is well correlated with deviation from equiprobability in synonymous codon usage. The biological significance of this fact is discussed.

Animals↗

A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.

BACKGROUND: Correlations between genome composition (in terms of GC content) and usage of particular codons and amino acids have been widely reported, but poorly explained. We show here that a simple model of processes acting at the nucleotide level explains codon usage across a large sample of species (311 bacteria, 28 archaea and 257 eukaryotes). The model quantitatively predicts responses (slope and intercept of the regression line on genome GC content) of individual codons and amino acids to genome composition. RESULTS: Codons respond to genome composition on the basis of their GC content relative to their synonyms (explaining 71-87% of the variance in response among the different codons, depending on measure). Amino-acid responses are determined by the mean GC content of their codons (explaining 71-79% of the variance). Similar trends hold for genes within a genome. Position-dependent selection for error minimization explains why individual bases respond differently to directional mutation pressure. CONCLUSIONS: Our model suggests that GC content drives codon usage (rather than the converse). It unifies a large body of empirical evidence concerning relationships between GC content and amino-acid or codon usage in disparate systems. The relationship between GC content and codon and amino-acid usage is ahistorical; it is replicated independently in the three domains of living organisms, reinforcing the idea that genes and genomes at mutation/selection equilibrium reproduce a unique relationship between nucleic acid and protein composition. Thus, the model may be useful in predicting amino-acid or nucleotide sequences in poorly characterized taxa.

Amino Acids↗

Mutation and selection at silent and replacement sites in the evolution of animal mitochondrial DNA.

Two patterns are presented that illustrate the interaction of mutation and selection in the evolution of animal mtDNA: 1) variation among taxa in the ratio of polymorphism to divergence (rpd) at silent and replacement sites in protein-coding genes, and 2) strand-differences in polymorphism and divergence at 'silent' sites that suggest a mutation-selection balance in the evolution of codon usage. Cytochrome b data from GenBank show that about half of the species pairs tested have a significant excess of amino acid polymorphism, relative to divergence. The remaining half of species pairs do not depart from neutrality, but generally do show an excess of amino acid polymorphism. Sequences from Drosophila pseudoobscura displaying a signature of an expanding population show a slight, but non-significant, deficiency of amino acid polymorphism suggestive of recently intensified selection on mildly deleterious mutations. Genes whose reading frames lie on the major coding strand of Drosophila mtDNA show a preponderance of T- > C substitutions, while genes encoded on the minor strand experience more A- > G than T- > C substitutions between species at both silent and replacement sites. However, silent mutations at third codon positions are introduced into the population in proportions opposite to those observed as fixed differences between species (e.g., an excess of T- > C polymorphisms are found at the ND5 gene on the minor coding strand). The high A + T content of insect mtDNAs imposes strong codon usage bias favoring A-ending and T-ending codons resulting in a distinct mutation-selection balance for genes encoded on opposites strands. Thus, at both replacement and silent sites, mutations that appear to be constrained in terms of divergence between species are in excess within species. The data suggest that mildly deleterious mutations are common in mitochondrial genes. A test of this, and a competing, hypothesis is proposed that requires additional sequence surveys of polymorphism and divergence. An important challenge is to tease apart the impact of mutation and selection on levels of polymorphism versus divergence in a genome that does not generally recombine.

Animals↗

Codon catalog usage is a genome strategy modulated for gene expressivity.

The nucleic acid sequence bank now contains 161 mRNAs, 43 new genes are added. One sequence, that of B. mori fibroin, is dropped due to uncertainty on the starting point for translation. Frequencies of all codons are given for each gene added and for each genome type in the total bank. A new series of correspondence analyses on codon use is presented, substantiating the genome hypothesis. Internal regulation of mRNA expression by different third base choices between quartet and duet codons is proposed for bacterial genes.

Amino Acid Sequence↗

The targeting of somatic hypermutation.

Somatic hypermutation does not occur randomly within immunoglobulin V genes but, rather, is preferentially targeted to certain nucleotide positions (hot spots) and away from others (cold spots). Cold spots often coincide with residues essential for V gene folding. Hotspots, which appear to be strategically located to favour affinity maturation, are most frequently located in the CDRs (particularly CDR1) though conserved hotspots are also found at the base of FR3. Hotspots are in part created by local DNA sequence and the strong biases of codon usage in V genes indicate that the genes have evolved such that somatic hypermutation is targeted to those parts of the V where it is likely to prove most useful. These features of mutational hotspots and biased codon usage are also evident in V genes of lower animals suggesting that diversification by strategic targeting of non-templated mutation may have evolved early in antigen receptor evolution.

Animals↗

Simultaneous horizontal gene transfer of a gene coding for ribosomal protein l27 and operational genes in Arthrobacter sp.

Phylogenetic analysis of bacterial L27 ribosomal proteins showed that, against taxonomy, the L27 protein from the Actinobacteria Arthrobacter sp. clusters with protein sequences from the Bacillus group. The L27 gene clusters in the Arthrobacter sp. genome with six genes responsible for creatinine and sarcosine degradation. Phylogenetic analyses of orthologue proteins encoded by three of these genes also showed a phylogenetic relationship with Bacillus species. Comparisons between the synonymous codon usage of the Arthrobacter sp. genes and those from complete genomes showed that Arthrobacter genes encoding the L27 ribosomal protein and the proteins responsible for the degradation of creatinine and sarcosine have a codon usage that is more similar to that of Bacillus species than that of Arthrobacter. We suggest that the Arthrobacter sp. genes encoding the L27 ribosomal protein and the proteins responsible for the degradation of creatinine and sarcosine were acquired simultaneously through horizontal gene transfer from an unknown Bacillus species.

Amino Acid Sequence↗

Genome variability and capsid structural constraints of hepatitis a virus.

The number of synonymous mutations per synonymous site (K(s)), the number of nonsynonymous mutations per nonsynonymous site (K(a)), and the codon usage statistic (N(c)) were calculated for several hepatitis A virus (HAV) isolates. While K(s) was similar to those of poliovirus (PV) and foot-and-mouth disease virus (FMDV), K(a) was 1 order of magnitude lower. The N(c) parameter provides information on codon usage bias and decreases when bias increases. The N(c) value in HAV was about 38, while in PV and FMDV, it was about 53. The emergence of 22 rare codons in front of 8 in PV and 7 in FMDV was detected. Most of the conserved rare codons of the P1 region were strategically located at the carboxy borders of beta barrels and alpha helices, their potential function being the assurance of proper folding of the capsid proteins through a decrease in the translation speed. This strategic location was not observed for amino acids encoded by the conserved rare codons of the 3D region. The percentage of bases with low pairing number values was higher in the latter region, suggesting a role of the conserved rare codons in the maintenance of RNA structure. Many of the rare codons in HAV are among the most frequent in humans, unlike in PV or in FMDV. This fact may be explained by the lack of cellular shutoff in HAV. One hypothesis is that HAV has evolved in order to avoid competition with its host for cellular tRNAs.

Amino Acid Sequence↗

Contrasting patterns of evolutionary divergence within the Acinetobacter calcoaceticus pca operon.

The six enzymes required for catabolism of protocatechuate to succinate and acetylCoA are encoded by the pca genes in the Gram-bacterium, Acinetobacter calcoaceticus. The clustered A. calcoaceticus cat genes encode an analogous set of enzymes associated with the metabolic dissimilation of catechol. The nucleotide (nt) sequences of pcaIJFB and pcaK, reported here, complete evidence showing that all of the pca structural genes are tightly grouped in the order pcaIJFBDKCHG within a single operon. The pcaIJF region is nearly identical in nt sequence to the A. calcoaceticus catIDJF region which exhibits a G+C content and a codon usage pattern exceptional for A. calcoaceticus. In contrast, pcaD, pcaC, pcaH and pcaG have diverged substantially from their evolutionary counterparts in the cat region; all of these divergent genes exhibit G+C contents and codon usage patterns that are typical for A. calcoaceticus. The pcaIJF and catIJF regions are known to exchange DNA sequence information, and this property may have contributed to their nt sequence conservation. The pcaK gene has no counterpart among known cat genes. The deduced amino-acid sequence of PcaK indicates that it may be a transmembrane protein associated with transport.

Acetyl Coenzyme A↗

Co-influence of transgene expression in mammalian cells. Mutual influence of transgenes on their expression in mammalian cells.

It becomes increasingly clear that therapeutic gene delivery should provide not only for the sustained high level of gene expression but also, in most cases, for the regulated expression of transgenes as much as it occurs under natural conditions. Over the past few years a variety of different systems have been developed in order to regulate the amounts of transcribed RNA upon administration of exogenous agents, or in autoregulated manner. While efforts were focused on optimizing gene expression at the transcriptional level, other levels are still overlooked. In the meantime, regulation of gene expression is not restricted to transcription, but is also executed at the post-transcriptional level, i.e. mRNA stability, processing, transport, translation, protein stability, and modification. Codon usage is considered to be one of the critical factors that limit the expression rate of heterologous genes in different organisms at the posttranscriptional level. HIV-1 structural genes gag, pol, and env represent one of the most extensively utilized models for studying codon usage-mediated effects on transgene expression. In the current work we demonstrate that the codon content affects not only CMV-driven HIV-1 gag expression but also the expression of luciferase reporter gene transcribed independently from the SV40 promoter. The expression levels of both transgenes co-transfected into the human H1299 were inversely co-dependent. The observed phenomenon may be described as sequence-independent post-transcriptional gene silencing, which reflects the existing limitation of transgene expression in mammalian cells at the post-transcriptional level. Optimization of the codon usage may provide for the additional level of regulation of transgene expression in gene transfer experiments in order to maintain the concentration of the protein at the therapeutic levels.

Cell Line, Tumor↗