PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Genetic plasticity of V genes under somatic hypermutation: statistical analyses using a new resampling-based methodology.

Evidence for somatic hypermutation of immunoglobulin genes has been observed in all of the species in which immunoglobulins have been found. Previous studies have suggested that codon usage in immunoglobulin variable (V) region genes is such that the sequence-specificity of somatic hypermutation results in greater mutability in complementarity-determining regions of the gene than in the framework regions. We have developed a new resampling-based methodology to explore genetic plasticity in individual V genes and in V gene families in a statistically meaningful way. We determine what factors contribute to this mutability difference and characterize the strength of selection for this effect. We find that although the codon usage in immunoglobulin V genes renders them distinct among translationally equivalent sequences with random codon usage, they are nevertheless not optimal in this regard. We find that the mutability patterns in a number of species are similar to those we find for human sequences. Interestingly, sheep sequences show extremely strong mutability differences, consistent with the role of somatic hypermutation in the diversification of primary antibody repertoire in these animals. Human TCR V(beta) sequences resemble immunoglobulin in mutability pattern, suggesting one of several alternatives, that hypermutation is functionally operating in TCR, that it was once operating in TCR or in the common precursor of TCR and immunoglobulin, or that the hypermutation mechanism has evolved to exploit the codon usage in immunoglobulin (and fortuitously, TCR) rather than vice-versa. Our findings provide support to the hypothesis that somatic hypermutation appeared very early in the phylogeny of immune systems, that it is, to a large extent, shared between species, and that it makes an essential contribution to the generation of the antibody repertoire.

Base Sequence↗

A nucleotide polymorphism in ERCC1 in human ovarian cancer cell lines and tumor tissues.

We studied the DNA sequence of the entire coding region of ERCC1 gene, in five cell lines established from human ovarian cancer (A2780, A2780/CP70, MCAS, OVCAR-3, SK-OV-3), 29 human ovarian cancer tumor tissue specimens, one human T-lymphocyte cell line (H9), and non-malignant human ovary tissue (NHO). Samples were assayed by PCR-SSCP and DNA sequence analyses. A silent mutation at codon 118 (site for restriction endonuclease MaeII) in exon 4 of the gene was detected in MCAS, OVCAR-3 and SK-OV-3 cells, and NHO. This mutation was a C-->T transition, that codes for the same amino acid: asparagine. This transition converts a common codon usage (AAC) to an infrequent codon usage (AAT), whereas frequency of use is reduced two-fold. This base change was associated with a detectable band shift on SSCP analysis. For the 29 ovarian cancer specimens, the same base change was observed in 15 tumor samples and was associated with the same band shift in exon 4. Cells and tumor tissue specimens that did not contain the C-->T transition, did not show the band shift in exon 4. Our data suggest that this alteration at codon 118 within the ERCC1 gene, may exist in platinum-sensitive and platinum-resistant ovarian cancer tissues.

Antineoplastic Agents↗

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena↗

On the origin of Ser/Thr kinases in a prokaryote.

The family of Ser/Thr and/or Tyr kinases and that of His kinases play essential roles in signal transduction. For a long time, the former has been found in eukaryotes, the latter in prokaryotes. Studies in the last decade have shown, however, that most bacteria possess from one to more than 10 genes encoding Ser/Thr kinases. This observation raises an important question concerning the evolutionary origin of Ser/Thr kinases found in bacteria. To answer this question, we have analyzed a family of 11 genes encoding Ser/Thr kinases in the cyanobacterium Synechocystis sp. PCC 6803. This bacterium contains the largest number of Ser/Thr kinases among all bacteria whose genomic sequences have been released so far. In this study, we have developed a user-friendly computer program for statistical analysis of codon usages and GC content. The results demonstrate that Ser/Thr kinases have similar codon usages and GC contents as the average of all possible open reading frames (ORFs) deduced from the genome. In contrast, ORFs encoding transposases, as a control in our analysis, display a disparity in both codon usage and GC content, confirming their multiple origin and genetic promiscuity. In light of our results, we propose that Ser/Thr kinases existed before the divergence between prokaryotes and eukaryotes during evolution, or were laterally transferred into prokaryotes at the early stages of bacterial evolution. If Ser/Thr kinases have persisted ever since in prokaryotes under evolutionary pressure, it is then expected that they play important, possibly even essential roles in regulating bacterial activities as do their counterparts in eukaryotes.

Base Composition↗

Poly(3-hydroxyvalerate) depolymerase of Pseudomonas lemoignei.

Pseudomonas lemoignei is equipped with at least five polyhydroxyalkanoate (PHA) depolymerase structural genes (phaZ1 to phaZ5) which enable the bacterium to utilize extracellular poly(3-hydroxybutyrate) (PHB), poly(3-hydroxyvalerate) (PHV), and related polyesters consisting of short-chain-length hxdroxyalkanoates (PHA(SCL)) as the sole sources of carbon and energy. Four genes (phaZ1, phaZ2, phaZ3, and phaZ5) encode PHB depolymerases C, B, D, and A, respectively. It was speculated that the remaining gene, phaZ4, encodes the PHV depolymerase (D. Jendrossek, A. Frisse, A. Behrends, M. Andermann, H. D. Kratzin, T. Stanislawski, and H. G. Schlegel, J. Bacteriol. 177:596-607, 1995). However, in this study, we show that phaZ4 codes for another PHB depolymeraes (i) by disagreement of 5 out of 41 amino acids that had been determined by Edman degradation of the PHV depolymerase and of four endoproteinase GluC-generated internal peptides with the DNA-deduced sequence of phaZ4, (ii) by the lack of immunological reaction of purified recombinant PhaZ4 with PHV depolymerase-specific antibodies, and (iii) by the low activity of the PhaZ4 depolymerase with PHV as a substrate. The true PHV depolymerase-encoding structural gene, phaZ6, was identified by screening a genomic library of P. lemoignei in Escherichia coli for clearing zone formation on PHV agar. The DNA sequence of phaZ6 contained all 41 amino acids of the GluC-generated peptide fragments of the PHV depolymerase. PhaZ6 was expressed and purified from recombinant E. coli and showed immunological identity to the wild-type PHV depolymerase and had high specific activities with PHB and PHV as substrates. To our knowledge, this is the first report on a PHA(SCL) depolymerase gene that is expressed during growth on PHV or odd-numbered carbon sources and that encodes a protein with high PHV depolymerase activity. Amino acid analysis revealed that PhaZ6 (relative molecular mass [M(r)], 43,610 Da) resembles precursors of other extracellular PHA(SCL) depolymerases (28 to 50% identical amino acids). The mature protein (M(r), 41,048) is composed of (i) a large catalytic domain including a catalytic triad of S(136), D(211), and H(269) similar to serine hydrolases; (ii) a linker region highly enriched in threonine residues and other amino acids with hydroxylated or small side chains (Thr-rich region); and (iii) a C-terminal domain similar in sequence to the substrate-binding domain of PHA(SCL) depolymerases. Differences in the codon usage of phaZ6 for some codons from the average codon usage of P. lemoignei indicated that phaZ6 might be derived from other organisms by gene transfer. Multialignment of separate domains of bacterial PHA(SCL) depolymerases suggested that not only complete depolymerase genes but also individual domains might have been exchanged between bacteria during evolution of PHA(SCL) depolymerases.

Acyltransferases↗

The Synthetic Gene Designer: a flexible web platform to explore sequence manipulation for heterologous expression.

"Codon optimization" is a general approach to improving heterologous expression where genes are moved from their native genomes into alternatives that exhibit different patterns of codon usage. However, despite reports of successful manipulations and the existence of stand-alone codon optimization software packages or commercial services that offer to redesign genes, the scientific community lacks any systematic understanding of what exactly it means to optimize codon usage. Thus we present a bona fide web application, the "Synthetic Gene Designer," which contrasts with existing software by providing a centralized, free, and transparent platform for the broader scientific community to develop knowledge about synthetic gene design. Consistent with this goal, our software is associated with a moderated e-forum that promotes discussion of synthetic gene design and offers technical support. In addition, the Synthetic Gene Designer presents enhanced functionality over existing software options: for example, it enables users to work with non-standard genetic codes, with user-defined patterns of codon usage and an expanded range of methods for codon optimization. The Synthetic Gene Designer, together with on-line tutorials and the forum, is available at .

Animals↗

Codon catalog usage and the genome hypothesis.

Frequencies for each of the 61 amino acid codons have been determined in every published mRNA sequence of 50 or more codons. The frequencies are shown for each kind of genome and for each individual gene. A surprising consistency of choices exists among genes of the same or similar genomes. Thus each genome, or kind of genome, appears to possess a "system" for choosing between codons. Frameshift genes, however, have widely different choice strategies from normal genes. Our work indicates that the main factors distinguishing between mRNA sequences relate to choices among degenerate bases. These systematic third base choices can therefore be used to establish a new kind of genetic distance, which reflects differences in coding strategy. The choice patterns we find seem compatible with the idea that the genome and not the individual gene is the unit of selection. Each gene in a genome tends to conform to its species' usage of the codon catalog; this is our genome hypothesis.

Animals↗

Origin and evolution of genes specifying resistance to macrolide, lincosamide and streptogramin antibiotics: data and hypotheses.

Resistance to macrolide, lincosamide and streptogramin antibiotics is due to alteration of the target site or detoxification of the antibiotic. Postranscriptional methylation of 23S ribosomal rRNA confers resistance to macrolide (M), lincosamide (L) and streptogramin (S) B-type antibiotics, the so-called MLSB phenotype. Several classes of rRNA methylases conferring resistance to MLSB antibiotics have been characterized in Gram-positive cocci, in Bacillus spp, and in strains of actinomycetes producing erythromycin. The enzymes catalyze N6-dimethylation of an adenine residue situated in a highly conserved region of prokaryotic 23S rRNA. In this review, we compare the amino acid sequences of the rRNA methylases and analyze the codon usage in the corresponding erm (erythromycin resistance methylase) genes. The homology detected at the protein level is consistent with the notion that an ancestor of the erm genes was implicated in erythromycin resistance in a producing strain. However, the rRNA methylases of producers and non-producers present substantial sequence diversity. In Gram-positive bacteria the preferential codon usage in the erm genes reflects the guanosine plus cytosine content of the chromosome of the host. These observations suggest that the presence of erm genes in these micro-organisms is ancient. By contrast, it would appear that enterobacteria have acquired only recently an rRNA methylase gene of the ermB class from a Gram-positive coccus since the genes isolated in Escherichia coli and in Gram-positive cocci are highly homologous (homology greater than 98%) and present a codon usage typical of the latter micro-organisms. As opposed to the MLSB phenotype which results from a single biochemical mechanism, inactivation of structurally related antibiotics of the MLS group involves synthesis of various other enzymes. In enterobacteria, resistance to erythromycin and oleandomycin is due to production of erythromycin esterases which hydrolyze the lactone ring of the 14-membered macrolides. We recently reported the nucleotide sequence of ereA and ereB (erythromycin resistance esterase) genes which encode erythromycin esterases type I and II, respectively. The amino acid sequences of the two isozymes do not exhibit statistically significant homology. Analysis of codon usage in both genes suggests that esterase type I is indigenous to E. coli, whereas the type II enzyme was acquired by E. coli from a phylogenetically remote micro-organism. Inactivation of lincosamides, first reported in staphylococci and lactobacilli of animal origin, was also recently detected in Gram-positive cocci isolated from humans.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

The correlation between synonymous and nonsynonymous substitutions in Drosophila: mutation, selection or relaxed constraints?

Codon usage bias, the preferential use of particular codons within each codon family, is characteristic of synonymous base composition in many species, including Drosophila, yeast, and many bacteria. Preferential usage of particular codons in these species is maintained by natural selection acting largely at the level of translation. In Drosophila, as in bacteria, the rate of synonymous substitution per site is negatively correlated with the degree of codon usage bias, indicating stronger selection on codon usage in genes with high codon bias than in genes with low codon bias. Surprisingly, in these organisms, as well as in mammals, the rate of synonymous substitution is also positively correlated with the rate of nonsynonymous substitution. To investigate this correlation, we carried out a phylogenetic analysis of substitutions in 22 genes between two species of Drosophila, Drosophila pseudoobscura and D. subobscura, in codons that differ by one replacement and one synonymous change. We provide evidence for a relative excess of double substitutions in the same species lineage that cannot be explained by the simultaneous mutation of two adjacent bases. The synonymous changes in these codons also cannot be explained by a shift to a more preferred codon following a replacement substitution. We, therefore, interpret the excess of double codon substitutions within a lineage as being the result of relaxed constraints on both kinds of substitutions in particular codons.

Animals↗

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms↗

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts↗

Minor shift in background substitutional patterns in the Drosophila saltans and willistoni lineages is insufficient to explain GC content of coding sequences.

BACKGROUND: Several lines of evidence suggest that codon usage in the Drosophila saltans and D. willistoni lineages has shifted towards a less frequent use of GC-ending codons. Introns in these lineages show a parallel shift toward a lower GC content. These patterns have been alternatively ascribed to either a shift in mutational patterns or changes in the definition of preferred and unpreferred codons in these lineages. RESULTS AND DISCUSSION: To gain additional insight into this question, we quantified background substitutional patterns in the saltans/willistoni group using inactive copies of a novel, Q-like retrotransposable element. We demonstrate that the pattern of background substitutions in the saltans/willistoni lineage has shifted to a significant degree, primarily due to changes in mutational biases. These differences predict a lower equilibrium GC content in the genomes of the saltans/willistoni species compared with that in the D. melanogaster species group. The magnitude of the difference can readily account for changes in intronic GC content, but it appears insufficient to explain changes in codon usage within the saltans/willistoni lineage. CONCLUSION: We suggest that the observed changes in codon usage in the saltans/willistoni clade reflects either lineage-specific changes in the definitions of preferred and unpreferred codons, or a weaker selective pressure on codon bias in this lineage.

Animals↗

Variation in synonymous codon use and DNA polymorphism within the Drosophila genome.

A strong negative correlation between the rate of amino-acid substitution and codon usage bias in Drosophila has been attributed to interference between positive selection at nonsynonymous sites and weak selection on codon usage. To further explore this possibility we have investigated polymorphism and divergence at three kinds of sites: synonymous, nonsynonymous and intronic in relation to codon bias in D. melanogaster and D. simulans. We confirmed that protein evolution is one of the main explicative parameters for interlocus codon bias variation (r(2) approximately 40%). However, intron or synonymous diversities, which could have been expected to be good indicators of local interference [here defined as the additional increase of drift due to selection on tightly linked sites, also called 'genetic draft' by Gillespie (2000)] did not covary significantly with codon bias or with protein evolution. Concurrently, levels of polymorphism were reduced in regions of low recombination rates whereas codon bias was not. Finally, while nonsynonymous diversities were very well correlated between species, neither synonymous nor intron diversities observed in D. melanogaster were correlated with those observed in D. simulans. All together, our results suggest that the selective constraint on the protein is a stable component of gene evolution while local interference is not. The pattern of variation in genetic draft along the genome therefore seems to be instable through evolutionary times and should therefore be considered as a minor determinant of codon bias variance. We argue that selective constraints for optimal codon usage are likely to be correlated with selective constraints on the protein, both between codons within a gene, as previously suggested, and also between genes within a genome.

Animals↗

Expression of the genomic form of the bovine viral diarrhea virus E2 ORF in a bovine herpesvirus-1 vector.

Bovine viral diarrhea virus (BVDV) is a ubiquitous pathogen of cattle with a world-wide distribution. Recently, the possibility of using recombinant virus vectors to immunize cattle against selected BVDV genes has gained widespread interest. Among the virus vectors tested, bovine herpesvirus-1 (BHV1) provides many unique advantages. However, results of recent studies have raised the possibility that the codon usage pattern required for optimal expression in a BHV1-infected cell may be incompatible with the codon usage pattern of BVDV. If true, use of BHV1 to express BVDV proteins would require construction of synthetic BVDV genes that have been modified to resemble the codon pattern of BHV1. To explore this possibility, we constructed a BHV1 recombinant containing the genomic form of the BVDV (NADL) E2 ORF and compared expression of the E2 protein with that of the endogenous BHV1 gD protein. We observed that E2 was expressed at a significant rate compared to that of the gD protein. We conclude that codon usage problems are unlikely to constitute a serious problem for expression of BVDV proteins in BHV1 vectors.

Animals↗

Compositional properties of nuclear genes from Plasmodium falciparum.

We have analyzed the compositional distributions of coding sequences and their different codon positions, as well as the codon usage of the nuclear genes of Plasmodium falciparum, a parasite characterized by an extremely GC-poor genome. As expected, coding sequences are AT-rich, codon usage is strongly biased towards A or T in third codon positions, and some particular amino acids (aa) are especially abundant in the encoded proteins. Remarkably, however, no difference was detected between housekeeping (HK) and antigen (Ag) genes, in spite of differences in expression level and evolutionary constraints. Moreover, all the features found in P. falciparum are very similar to those found in a bacterium characterized by a very GC-poor genome, Staphylococcus aureus. These findings stress the importance of compositional constraints in determining codon usage and aa utilisation.

Amino Acids↗

Comparison of three actin-coding sequences in the mouse; evolutionary relationships between the actin genes of warm-blooded vertebrates.

We have determined the sequences of three recombinant cDNAs complementary to different mouse actin mRNAs that contain more than 90% of the coding sequences and complete or partial 3' untranslated regions (3'UTRs): pAM 91, complementary to the actin mRNA expressed in adult skeletal muscle (alpha sk actin); pAF 81, complementary to an actin mRNA that is accumulated in fetal skeletal muscle and is the major transcript in adult cardiac muscle (alpha c actin); and pAL 41, identified as complementary to a beta nonmuscle actin mRNA on the basis of its 3'UTR sequence. As in other species, the protein sequences of these isoforms are highly (greater than 93%) conserved, but the three mRNAs show significant divergence (13.8-16.5%) at silent nucleotide positions in their coding regions. A nucleotide region located toward the 5' end shows significantly less divergence (5.6-8.7%) among the three mouse actin mRNAs; a second region, near the 3' end, also shows less divergence (6.9%), in this case between the mouse beta and alpha sk actin mRNAs. We propose that recombinational events between actin sequences may have homogenized these regions. Such events distort the calculated evolutionary distances between sequences within a species. Codon usage in the three actin mRNAs is clearly different, and indicates that there is no strict relation between the tissue type, and hence the tRNA precursor pool, and codon usage in these and other muscle mRNAs examined. Analysis of codon usage in these coding sequences in different vertebrate species indicates two tendencies: increases in bias toward the use of G and C in the third codon position in paralogous comparisons (in the order alpha c less than beta less than alpha sk), and in orthologous comparisons (in the order chicken less than rodent less than man). Comparison of actin-coding sequences between species was carried out using the Perler method of analysis. As one moves backward in time, changes at silent sites first accumulate rapidly, then begin to saturate after -(30-40) million years (MY), and actually decrease between -400 and -500 MY. Replacements or silent substitutions therefore cannot be used as evolutionary clocks for these sequences over long periods. Other phenomena, such as gene conversion or isochore compartmentalization, probably distort the estimated divergence time.

Actins↗

Nucleotide sequence of the structural gene for tryptophanase of Escherichia coli K-12.

The tryptophanase structural gene, tnaA, of Escherichia coli K-12 was cloned and sequenced. The size, amino acid composition, and sequence of the protein predicted from the nucleotide sequence agree with protein structure data previously acquired by others for the tryptophanase of E. coli B. Physiological data indicated that the region controlling expression of tnaA was present in the cloned segment. Sequence data suggested that a second structural gene of unknown function was located distal to tnaA and may be in the same operon. The pattern of codon usage in tnaA was intermediate between codon usage in four of the ribosomal protein structural genes and the structural genes for three of the tryptophan biosynthetic proteins.

Amino Acid Sequence↗

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis↗