PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Suppression of the negative effect of minor arginine codons on gene expression; preferential usage of minor codons within the first 25 codons of the Escherichia coli genes.

AGA and AGG codons for arginine are the least used codons in Escherichia coli, which are encoded by a rare tRNA, the product of the dnaY gene. We examined the positions of arginine residues encoded by AGA/AGG codons in 678 E. coli proteins. It was found that AGA/AGG codons appear much more frequently within the first 25 codons. This tendency becomes more significant in those proteins containing only one AGA or AGG codon. Other minor codons such as CUA, UCA, AGU, ACA, GGA, CCC and AUA are also found to be preferentially used within the first 25 codons. The effects of the AGG codon on gene expression were examined by inserting one to five AGG codons after the 10th codon from the initiation codon of the lacZ gene. The production of beta-galactosidase decreased as more AGG codons were inserted. With five AGG codons, the production of beta-galactosidase (Gal-AGG5) completely ceased after a mid-log phase of cell growth. After 22 hr induction of the lacZ gene, the overall production of Gal-AGG5 was 11% of the control production (no insertion of arginine codons). When five CGU codons, the major arginine codon were inserted instead of AGG, the production of beta-galactosidase (Gal-CGU5) continued even after stationary phase and the overall production was 66% of the control. The negative effect of the AGG codons on the Gal-AGG5 production was found to be dependent upon the distance between the site of the AGG codons and the initiation codon. As the distance was increased by inserting extra sequences between the two codons, the production of Gal-AGG5 increased almost linearly up to 8 fold. From these results, we propose that the position of the minor codons in an mRNA plays an important role in the regulation of gene expression possibly by modulating the stability of the initiation complex for protein synthesis.

Amino Acid Sequence↗

The complete mitochondrial genome of Tupaia belangeri and the phylogenetic affiliation of scandentia to other eutherian orders.

The complete mitochondrial genome of Tupaia belangeri, a representative of the eutherian order Scandentia, was determined and compared with full-length mitochondrial sequences of other eutherian orders described to date. The complete mitochondrial genome is 16, 754 nt in length, with no obvious deviation from the general organization of the mammalian mitochondrial genome. Thus, features such as start codon usage, incomplete stop codons, and overlapping coding regions, as well as the presence of tandem repeats in the control region, are within the range of mammalian mitochondrial (mt) DNA variation. To address the question of a possible close phylogenetic relationship between primates and Tupaia, the evolutionary affinities among primates, Tupaia and bats as representatives of the Archonta superorder, ferungulates, guinea pigs, armadillos, rats, mice, and hedgehogs were examined on the basis of the complete mitochondrial DNA sequences. The opossum sequence was used as an outgroup. The trees, estimated from 12 concatenated genes encoded on the mitochondrial H-strand, add further molecular evidence against an Archonta monophyly. With the new data described in this paper, most of both the mitochondrial and the nuclear data point away from Scandentia as the closest extant relatives to primates. Instead, the complete mitochondrial data support a clustering of Scandentia with Lagomorpha connecting to the branch leading to ferungulates. This closer phylogenetic relationship of Tupaia to rabbits than to primates first received support from several analyses of nuclear and partial mitochondrial DNA data sets. Given that short sequences are of limited use in determining deep mammalian relationships, the partial mitochondrial data available to date support this hypothesis only tentatively. Our complete mitochondrial genome data therefore add considerably more evidence in support of this hypothesis.

Animals↗

Shortening of the symptom-free period in rhesus macaques is associated with decreasing nonsynonymous variation in the env variable regions of simian immunodeficiency virus SIVsm during passage.

During six blood passages of simian immunodeficiency virus SIVsm in rhesus macaques, the asymptomatic period shortened from 18 months to 1 month. To study SIVsm envelope gene (env) evolution during passage in rhesus macaques, the C1 to CD4 binding regions of multiple clones were sequenced at seroconversion and again at death. The env variation found during adaptation was almost completely confined to the variable regions. Intrasample sequence variation among clones at seroconversion was lower than the variation among clones at death. Intrasample variation among clones from a single time point as well as intersample variation decreased during the passage. In the variable regions, the mean number of intrasample nonsynonymous nucleotide substitutions decreased from the first passage (5.26 x 10(-2) +/- 0.6 x 10(-2) per site) to the fifth passage (2.24 x 10(-2) +/- 0.4 x 10(-2) per site), whereas in the constant regions, the mean number of intrasample nonsynonymous nucleotide substitutions differed less between the first and fifth passages (1. 14 x 10(-2) +/- 0.27 x 10(-2) and 0.80 x 10(-2) +/- 0.24 x 10(-2) per site). Shortening of the asymptomatic period coincided with a rise in the Ks/Ka ratio (ratio between the number of synonymous [Ks] and the number of nonsynonymous [Ka] substitutions) from 1.080 in passage one to 1.428 in passage five and mimicked the difference seen in the intrahost evolution between asymptomatic and fast-progressing individuals infected with human immunodeficiency virus type 1. The distribution of nonsynonymous substitutions was biphasic, with most of the adaptation of env variable regions occurring in the first three passages. This phase, in which the symptom-free period fell to 4 months, was followed by a plateau phase of apparently reduced adaptation. Analysis of codon usage revealed decreased codon redundancy in the variable regions. Overall, the results suggested a biphasic pattern of adaptation and evolution, with extremely rapid selection in the first three passages followed by an equilibrium or stabilization of the variation between env clones at different time points in passages four to six.

Animals↗

Codon utilization, DNA landscaping and fractal analysis in bacteriophage phi(adh).

The bacteriophage phi(adh) has a low G+C content and encodes its protein products using a restricted number of the codons, which it could theoretically use. Investigated were (i) the restricted codon usage by determining codon indices and codon distances for various genes and ORFs, (ii) distribution of purines and pyrimidines on the two strands of the double-stranded genome and within all genes and ORFs, and (iii) nucleotide positional bias within the genome. The genes and ORFs can be clustered into four groups, based on codon distance analysis. The genome landscape showed that the plus strand was more purine-rich than the negative one and that the only area of the genome where the landscape was located in the pyrimidine-rich region was at the start of the sequence which was also the only region of the genome where ORFs were found on the negative strand. The nucleotide composition of the genome, examined by fractal analysis showed little, if any, DNA positional bias, as opposed to overall compositional bias with a self-similarity profile. The ORFs showed a bias in favour of purines on the coding strand.

Amino Acids↗

Cluster analysis of the codon use frequency of MHC genes from different species.

The relative synonymous codon use frequency of 135 MHC genes from four mammal species (Homo sapiens, Pan troglodyte, Macaca mulanta and Rattus norvegicus) is analyzed using a hierarchical cluster method. The result suggests that gene function is the dominant factor that determines codon usage bias, while species is a minor factor that determines further difference in codon usage bias for genes with similar functions. The conclusion may be useful in gene classification and gene function prediction.

Animals↗

A type II protein secretory pathway required for levansucrase secretion by Gluconacetobacter diazotrophicus.

The endophytic diazotroph Gluconacetobacter diazotrophicus secretes a constitutively expressed levansucrase (LsdA, EC 2.4.1.10) to utilize plant sucrose. LsdA, unlike other extracellular levansucrases from gram-negative bacteria, is transported to the periplasm by a signal-peptide-dependent pathway. We identified an unusually organized gene cluster encoding at least the components LsdG, -O, -E, -F, -H, -I, -J, -L, -M, -N, and -D of a type II secretory system required for LsdA translocation across the outer membrane. Another open reading frame, designated lsdX, is located between the operon promoter and lsdG, but it was not identified in BLASTX searches of the DDBJ/EMBL/GenBank databases. The lsdX, -G, and -O genes were isolated from a cosmid library of strain SRT4 by complementation of an ethyl methanesulfonate mutant unable to transport LsdA across the outer membrane. The downstream genes lsdE, -F, -H, -I, -J, -L, -M, -N, and -D were isolated through chromosomal walking. The high G+C content (64 to 74%) and the codon usage of the genes identified are consistent with the G+C content and codon usage of the standard G. diazotrophicus structural gene. Sequence analysis of the gene cluster indicated that a polycistronic transcript is synthesized. Targeted disruption of lsdG, lsdO, or lsdF blocked LsdA secretion, and the bacterium failed to grow on sucrose. Replacement of Cys(162) by Gly at the C terminus of the pseudopilin LsdG abolished the protein functionality, suggesting that there is a relationship with type IV pilins. Restriction fragment length polymorphism analysis revealed conservation of the type II secretion operon downstream of the levansucrase-levanase (lsdA-lsdB) locus in 14 G. diazotrophicus strains representing 11 genotypes recovered from four different host plants in diverse geographical regions. To our knowledge, this is the first report of a type II pathway for protein secretion in the Acetobacteraceae.

Amino Acid Sequence↗

Intercodon dinucleotides affect codon choice in plant genes.

In this work, 710 CDSs corresponding to over 290 000 codons equally distributed between Brassica napus, Arabidopsis thaliana, Lycopersicon esculentum, Nicotiana tabacum, Pisum sativum, Glycine max, Oryza sativa, Triticum aestivum, Hordeum vulgare and Zea mays were considered. For each amino acid, synonymous codon choice was determined in the presence of A, G, C or T as the initial nucleotide of the subsequent triplet; data were statistically analysed under the hypothesis of an independent assortment of codons. In 33.4% of cases, a frequency significantly (P: = 0.01) different from that expected was recorded. This was mainly due to a pervasive intercodon TpA and CpG deficiency. As a general rule, intercodon TpAs and CpGs were preferably replaced by CpAs and TpGs, respectively. In several instances, codon frequencies were also modified to avoid homotetramer and homotrimer formation, to reduce intercodon ApCs downstream (1,2) GG or AG dinucleotides, as well as to increase GpA or ApG intercodons under certain contexts. Since TpA, CpG and homotetra(tri)mer deficiency directly or indirectly accounted for 77% of significant variation in the codon frequency, it can be concluded that codon usage mirrors precise needs at the DNA structure level. Plant species exhibited a phylogenetically-related adaptation to structural constraints. Codon usage flexibility was reflected in strikingly different arrays of optimum codons for probe design.

Base Composition↗

Primary structure of the tolC gene that codes for an outer membrane protein of Escherichia coli K12.

We present the nucleotide sequence of the tolC gene of Escherichia coli K12, and the amino acid sequence of the TolC protein (an outer membrane protein) as deduced from it. The mature TolC protein comprises 467 amino acid residues, and, as previously reported (1), a signal sequence of 22 amino acid residues is attached to the N-terminus. The C-terminus of the gene is followed by a stem-loop structure (8 base pair stem, 4 base loop) which may be a rho-independent termination signal. The codon usage of the gene is nonrandom; the major isoaccepting species of tRNA are preferentially utilised, or, among synonomous codons recognized by the same tRNA, those codons are used which can interact better with the anticodon (2,3). In contrast to the codon usage for other outer membrane proteins of E. coli (4) the rare arginine codons AGA and AGG are used once and twice respectively.

Bacterial Outer Membrane Proteins↗

[Role of the code redundancy in determining cotranslational protein folding].

It has been demonstrated earlier in our laboratory that rare codon clusters can determine the boundaries of the polypeptide chain fragments of the same secondary structure type during the co-translational protein folding. According to this data, co-translational protein folding can occur under condition of a correlation between the frequency of codon choice in mRNAs and the relative abundance of their isoaccepting tRNAs. The alterations in the spectrum and concentrations of the isoaccepting tRNAs in different cells were demonstrated by many authors. The existence of a mechanism of the coordinate regulation of the levels (activities) of the isoaccepting tRNAs, corresponding aminoacyl-tRNA synthetases and mRNAs predominantly translated at a given moment of time can be suggested. Such a mechanism can ensure the needed accuracy of the protein folding process. Analysis of gene sequences of various pro- and eukaryotic organisms carried out in the present work revealed that the codon usage frequency spectra of simultaneously synthesized proteins are similar. The relative appearance of the most rare and frequent codons in investigated gene sequences displays a high degree of conservatism. It has also been found that structural-homologous proteins from different organisms (cytochromes c, myoglobins) have very similar codon frequency distribution profiles. This property retains despite the significant variations in the codon usage spectra in the investigated gene sequences. The data obtained indicate that the codon distribution in mRNAs whose diversity is mainly conditioned by the genetic code redundance is a program that determines translational rates of different mRNA parts thus controlling the spatial folding of the synthesized peptide chain.

Animals↗

Codon optimization of papillomavirus genes.

Early and late genes of human and animal papillomaviruses show a codon composition seemingly unfavorable for expression in mammalian cells. It remains unclear how the viruses manage to achieve high levels of late gene expression during the viral life cycle. One possible solution could be that the availability of certain t-RNAs changes with progressing stages of cellular differentiation. Previous studies have demonstrated that modification of codon usage of papillomavirus late (L1 and L2) and early genes (E7) can overcome poor expression of these proteins both in transient and in stable expression systems. This was shown not only for human but also for plant cells. Two strategies can be employed to alter codon usage: elimination of only those codons that are rarely used in a particular expression system, or exchange of all possible codons by the ones most frequently used. Currently, there are two protocols for codon modification--a template-less polymerase chain reaction (PCR)-based protocol, in which very long overlapping oligodeoxynucleotides are used in an overlap-extension reaction, or a ligase chain reaction, in which shorter oligodeoxynucleotides are fused together after an annealing procedure. Both methods are presented and discussed.

Amino Acids↗

A review of protein structure and gene organisation for proteins associated with mineralised tissue and calcium phosphate stabilisation encoded on human chromosome 4.

Several proteins associated with mineralised tissue (teeth and bone) or involved in calcium phosphate stabilisation in the body fluids, milk and saliva have been mapped to the q arm of human chromosome 4. These include the dentine/bone proteins dentine sialophosphoprotein (DSPP), dentine matrix protein 1 (DMP1), bone sialoprotein (BSP), matrix extracellular phosphoglycoprotein, osteopontin (OPN), enamelin, ameloblastin, milk caseins, salivary statherin, and proline-rich proteins. The proposed function of those that are multiphosphorylated is: (i) the stabilisation of calcium phosphate in solution (e.g. casein, statherin) preventing spontaneous precipitation and seeded-crystal growth or (ii) promoting biomineralisation (e.g. the phosphophoryn domain of DSPP), where the protein described as a template macromolecule, is proposed to act as a nucleator/promoter of crystal growth. The genes of these proteins have been subjected to conserved chromosomal synteny during mammalian evolution. The multiphosphorylated proteins statherin, caseins, phosphophoryn, BSP and OPN have been characterised as intrinsically disordered. The codon usage patterns for the amino acid serine reveal a bias for AGC and AGT codons within the human genes dspp, dmp1 and bsp, mouse dspp and dmp1 but not significantly for statherin or caseins. This pattern was also observed in the gene encoding hen phosvitin that also contains stretches of multiphosphorylated serines and in the dmp1 gene sequences of mammalian, reptilian and avian classes. In conclusion, these intrinsically disordered multiphosphorylated proteins are the translation products of genes displaying examples of codon usage bias, internal repeats and conserved chromosomal synteny within the mammalian class.

Animals↗

Comparison of the nucleoside sequence of trpA and sequences immediately beyond the trp operon of Klebsiella aerogenes. Salmonella typhimurium and Escherichia coli.

The nucleotide sequence of trpA of Klebsiella aerogenes is presented and compared with the trpA sequences of Salmonella typhimurium and Escherichia coli. The majority of the approximately 200 differences between each pair of trpA's are single nucleotide pair changes that do not alter the amino acid sequence. Codon usage conforms to the general patterns revealed by examination of other prokaryotic gene sequences. However, codon usage in K. aerogenes trpA reflects the high G+C content of the genome of this organism. The DNA sequences just beyond trpA, the presumed transcription termination region, are also compared for the three species. Perusal of these sequences indicates that the secondary structure of the transcript segment just beyond trpA has been preserved, while the primary sequence has diverged appreciably.

Amino Acid Sequence↗

Complete nucleotide sequence and genetic organization of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens.

The complete nucleotide sequence of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens has been determined. The plasmid genome comprises 10,207 bp and has a dA + dT content of 75%. Functions have been tentatively assigned to 6 of the 10 open reading frames and an origin-like region of repeated sequence identified. The codon usage of this extremely dA + dT rich plasmid is highly unusual and displays a pronounced preference for codons with the lowest dG + dC content. Only one of the genes from pIP404 was expressed at a significant level in Escherichia coli, suggesting that the atypical codon usage could represent a major obstacle to heterologous gene expression.

Base Sequence↗

A periodic pattern of mRNA secondary structure created by the genetic code.

Single-stranded mRNA molecules form secondary structures through complementary self-interactions. Several hypotheses have been proposed on the relationship between the nucleotide sequence, encoded amino acid sequence and mRNA secondary structure. We performed the first transcriptome-wide in silico analysis of the human and mouse mRNA foldings and found a pronounced periodic pattern of nucleotide involvement in mRNA secondary structure. We show that this pattern is created by the structure of the genetic code, and the dinucleotide relative abundances are important for the maintenance of mRNA secondary structure. Although synonymous codon usage contributes to this pattern, it is intrinsic to the structure of the genetic code and manifests itself even in the absence of synonymous codon usage bias at the 4-fold degenerate sites. While all codon sites are important for the maintenance of mRNA secondary structure, degeneracy of the code allows regulation of stability and periodicity of mRNA secondary structure. We demonstrate that the third degenerate codon sites contribute most strongly to mRNA stability. These results convincingly support the hypothesis that redundancies in the genetic code allow transcripts to satisfy requirements for both protein structure and RNA structure. Our data show that selection may be operating on synonymous codons to maintain a more stable and ordered mRNA secondary structure, which is likely to be important for transcript stability and translation. We also demonstrate that functional domains of the mRNA [5'-untranslated region (5'-UTR), CDS and 3'-UTR] preferentially fold onto themselves, while the start codon and stop codon regions are characterized by relaxed secondary structures, which may facilitate initiation and termination of translation.

3' Untranslated Regions↗

[Transformation of Chlamydomonas reinhardtii CW-15 with the hygromycin phosphotransferase gene as a selective marker].

To transform Chlamydomonas reinhardtii Dang. Cells, plasmid pCTVHyg was constructed with the use of the Escherichia coli hygromycin phosphotransferase gene (hpt) controlled by the SV40 early promoter. Cells of the CW-15 mutant strain were transformed by electroporation, with the yield reaching 10(3) hygromycin-resistant (HygR) clones per 10(6) recipient cells. The exogenous DNA integrated in the Ch. reinhardtii nuclear genome showed stable transmission for approximately 350 cell generations, while hygromycin resistance was expressed as an unstable character. Codon usage was compared for the hpt gene and Ch. reinhardtii nuclear genes. The results testified that codon usage bias, which is characteristic of Ch. reinhardtii, is not the major factor affecting foreign gene expression. The advantages of the selective system for studying Ch. reinhardtii transformation with heterologous genes are discussed.

Animals↗

Preferential use of A- and U-rich codons for Mycoplasma capricolum ribosomal proteins S8 and L6.

The nucleotide sequence of the 1.3 kilobase-pair DNA segment, which contains the genes for ribosomal proteins S8 and L6, and a part of L18 of Mycoplasma capricolum, has been determined and compared with the corresponding sequence in Escherichia coli (Cerretti et al., Nucl. Acids Res. 11, 2599, 1983). Identities of the predicted amino acid sequences of S8 and L6 between the two organisms are 54% and 42%, respectively. The A + T content of the M. capricolum genes is 71%, which is much higher than that of E. coli (49%). Comparisons of codon usage between the two organisms have revealed that M. capricolum preferentially uses A- and U-rich codons. More than 90% of the codon third positions and 57% of the first positions in M. capricolum is either A or U, whereas E. coli uses A or U for the third and the first positions at a frequency of 51% and 36%, respectively. The biased choice of the A- and U-rich codons in this organism has been also observed in the codon replacements for conservative amino acid substitutions between M. capricolum and E. coli. These facts suggest that the codon usage of M. capricolum is strongly influenced by the high A + T content of the genome.

Adenine↗

The sequence of the chloroplast atpB gene and its flanking regions in Chlamydomonas reinhardtii.

The chloroplast (cp)-encoded CF1 ATPase beta-subunit gene (atpB) of Chlamydomonas reinhardtii and its flanking regions have been sequenced. The derived amino acid (aa) sequence is highly homologous to that of the beta-subunit gene in Escherichia coli, bovine heart mitochondria, and higher plant cp. In contrast to all other cp genomes, the CF1 epsilon subunit gene (atpE) does not lie at the 3' end of the atpB gene but maps to a position 92 kb away in the other single-copy region. Northern blots confirm that the beta subunit is not encoded as part of a dicistronic message as it is in higher plants. The region just upstream from the atpB gene in C. reinhardtii contains two small open reading frames (ORFs) and not the gene for the large subunit of ribulose-1,5-bisphosphate carboxylase/oxygenase as is found in cp genomes of higher plants. No transcripts for either ORF were detected, but the codon usage in these ORFs as well as in the atpB gene follows the unique pattern of codon usage previously seen in other cp genes in C. reinhardtii.

Amino Acid Sequence↗