PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Predicted highly expressed genes in archaeal genomes.

Based primarily on 16S rRNA sequence comparisons, life has been broadly divided into the three domains of Bacteria, Archaea, and Eukarya. Archaea is further classified into Crenarchaea and Euryarchaea. Archaea generally thrive in extreme environments as assessed by temperature, pH, and salinity. For many prokaryotic organisms, ribosomal proteins (RP), transcription/translation factors, and chaperone genes tend to be highly expressed. A gene is predicted highly expressed (PHX) if its codon usage is rather similar to the average codon usage of at least one of the RP, transcription/translation factors, and chaperone gene classes and deviates strongly from the average gene of the genome. The thermosome (Ths) chaperonin family represents the most salient PHX genes among Archaea. The chaperones Trigger factor and HSP70 have overlapping functions in the folding process, but both of these proteins are lacking in most archaea where they may be substituted by the chaperone prefoldin. Other distinctive PHX proteins of Archaea, absent from Bacteria, include the proliferating cell nuclear antigen PCNA, a replication auxiliary factor responsible for tethering the catalytic unit of DNA polymerase to DNA during high-speed replication, and the acidic RP P0, which helps to initiate mRNA translation at the ribosome. Other PHX genes feature Cell division control protein 48 (Cdc48), whereas the bacterial septation proteins FtsZ and minD are lacking in Crenarchaea. RadA is a major DNA repair and recombination protein of Archaea. Archaeal genomes feature a strong Shine-Dalgarno ribosome-binding motif more pronounced in Euryarchaea compared with Crenarchaea.

Archaea↗

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence↗

Divergent evolutionary constraints on mitochondrial and nuclear genomes of malaria parasites.

Genetic variation among malaria parasites has important consequences with regard to drug resistance, pathogenicity, immunity, transmission, and speciation. In this regard, malaria parasites have been shown to display a high degree of inter- and intra-species genetic divergence. The nuclear genomes of Plasmodium falciparum, Plasmodium yoelii, and Plasmodium gallinaceum are vastly divergent yet share a similar codon usage and total A/T content of approximately 82%. This is in contrast to other primate-specific species including P. vivax which have an A/T content of approximately 67%. To assess the effects of this evolutionary divergence on the conservation of gene content, organization, and codon usage in the mitochondrial DNA (mtDNA) of malaria parasites, we have cloned and sequenced the mitochondrial genome of Plasmodium vivax, and compared it with the mtDNAs of P. falciparum, P. yoelii, and P. gallinaceum. The P. vivax mitochondrial genome was found to be 5990 base pairs in length, and displayed a gene organization identical to that of P. falciparum, P. yoelii, and P. gallinaceum. Furthermore, there was a remarkable 90% conservation of sequence identity between the mitochondrial genomes of all four species. As an example of intra-species conservation, comparison of mtDNAs from two independently cloned P. falciparum isolates, Malay Camp and C10, revealed only a single nucleotide substitution. A/T content of the P. vivax mitochondrial genome was found to be identical to other species of Plasmodium, hence, we have postulated that the mitochondrial genomes of malaria parasites were refractory to the evolutionary shifts in nucleotide content seen among the nuclear genomes of malaria parasites. Among different Plasmodium species, the second position of mitochondrial codons were found to be the least prone to substitutions and displayed a significant bias in pyrimidines. These aspects of mitochondrial codon usage were distinct from the nuclear genome and may reflect functional aspects of decoding by the mitochondrial translational system.

Amino Acid Sequence↗

Effect of a rare leucine codon, TTA, on expression of a foreign gene in Streptomyces lividans.

Streptomyces are bacteria with a very high chromosomal G+C composition (> 70 mol%) and extremely biased codon usage. In order to investigate the relationship between codon usage and gene expression in Streptomyces, we used ssi (Streptomyces subtilisin inhibitor) as a reporter gene and monitored its secretory expression in S. lividans. In consequence of alteration of the native codons of Leu, Lys and Ser of ssi to minor ones by site-directed mutagenesis, i.e., Leu79-Leu80: CTG-CTC to TTA-TTA, Lys89: AAG to AAA, Ser108-Ser109: TCG-AGC to TCT-TCT, respectively, the production of SSI was reduced remarkably in the case of TTA codons, while it was slightly increased in the case of AAA and almost the same in TCT codons. This conspicuous decrease found for Leu codon replacement was probably due to the low availability of intracellular tRNA(Leu) (UUA), a product of bldA which has been reported to be expressed only during the late stage of growth.

Amino Acid Sequence↗

The complete mitochondrial genome of the stomatopod crustacean Squilla mantis.

BACKGROUND: Animal mitochondrial genomes are physically separate from the much larger nuclear genomes and have proven useful both for phylogenetic studies and for understanding genome evolution. Within the phylum Arthropoda the subphylum Crustacea includes over 50,000 named species with immense variation in body plans and habitats, yet only 23 complete mitochondrial genomes are available from this subphylum. RESULTS: I describe here the complete mitochondrial genome of the crustacean Squilla mantis (Crustacea: Malacostraca: Stomatopoda). This 15994-nucleotide genome, the first described from a hoplocarid, contains the standard complement of 13 protein-coding genes, 22 transfer RNA genes, two ribosomal RNA genes, and a non-coding AT-rich region that is found in most other metazoans. The gene order is identical to that considered ancestral for hexapods and crustaceans. The 70% AT base composition is within the range described for other arthropods. A single unusual feature of the genome is a 230 nucleotide non-coding region between a serine transfer RNA and the nad1 gene, which has no apparent function. I also compare gene order, nucleotide composition, and codon usage of the S. mantis genome and eight other malacostracan crustaceans. A translocation of the histidine transfer RNA gene is shared by three taxa in the order Decapoda, infraorder Brachyura; Callinectes sapidus, Portunus trituberculatus and Pseudocarcinus gigas. This translocation may be diagnostic for the Brachyura. For all nine taxa nucleotide composition is biased towards AT-richness, as expected for arthropods, and is within the range reported for other arthropods. Codon usage is biased, and much of this bias is probably due to the skew in nucleotide composition towards AT-richness. CONCLUSION: The mitochondrial genome of Squilla mantis contains one unusual feature, a 230 base pair non-coding region has so far not been described in any other malacostracan. Comparisons with other Malacostraca show that all nine genomes, like most other mitochondrial genomes, share a bias toward AT-richness and a related bias in codon usage. The nine malacostracans included in this analysis are not representative of the diversity of the class Malacostraca, and additional malacostracan sequences would surely reveal other unusual genomic features that could be useful in understanding mitochondrial evolution in this taxon.

Animals↗

Molecular cloning, heterologous expression, and primary structure of the structural gene for the copper enzyme nitrous oxide reductase from denitrifying Pseudomonas stutzeri.

The nos genes of Pseudomonas stutzeri are required for the anaerobic respiration of nitrous oxide, which is part of the overall denitrification process. A nos-coding region of ca. 8 kilobases was cloned by plasmid integration and excision. It comprised nosZ, the structural gene for the copper-containing enzyme nitrous oxide reductase, genes for copper chromophore biosynthesis, and a supposed regulatory region. The location of the nosZ gene and its transcriptional direction were identified by using a series of constructs to transform Escherichia coli and express nitrous oxide reductase in the heterologous background. Plasmid pAV5021 led to a nearly 12-fold overexpression of the NosZ protein compared with that in the P. stutzeri wild type. The complete sequence of the nosZ gene, comprising 1,914 nucleotides, together with 282 nucleotides of 5'-flanking sequences and 238 nucleotides of 3'-flanking sequences was determined. An open reading frame coded for a protein of 638 residues (Mr, 70,822) including a presumed signal sequence of 35 residues for protein export. The presequence is in conformity with the periplasmic location of the enzyme. Another open reading frame of 2,097 nucleotides, in the opposite transcriptional direction to that of nosZ, was excluded by several criteria from representing the coding region for nitrous oxide reductase. Codon usage for nosZ of P. stutzeri showed a high G + C content in the degenerate codon position (83.9% versus an average of 60.2%) and relaxed codon usage for the Glu codon, characteristic features of Pseudomonas genes from other species. E. coli nitrous oxide reductase was purified to homogeneity. It had the Mr of the P. stutzeri enzyme but lacked the copper chromophore.

Amino Acid Sequence↗

Relationships between transcriptional and translational control of gene expression in Saccharomyces cerevisiae: a multiple regression analysis.

Natural selection for an increased translation efficiency has been proposed as the main determinant for the bias in codon usage observed in many genes of Saccharomyces cerevisiae. Recently, the efficiency of transcription of a large number of yeast genes has been determined, based on the cellular content of the respective mRNAs: this provides an additional dimension to the study of the multisep process of gene expression. Using a representative set of yeast genes with a known level of transcription, the relationship between transcriptional and translational steps was evaluated by a multiple linear regression model. This analysis demonstrated a positive correlation between the amount of transcript, given as the number of mRNA copies per cell for each individual gene, and indices evaluating the effects of translational selection on the corresponding codon usage pattern. This finding suggests a close association of the cellular mRNA content, regulated also at the transcriptional level, to its efficiency of translation, mediated by a fine-tuning of codon usage strategy. Moreover, multiple regression analysis demonstrated that the transcription level of a gene can be approximately predicted using indices of bias deriving from its nucleotide sequence. This allowed for an extensive investigation of uncharacterized regions of the complete genome sequence of S. cerevisiae, to detect new potential short protein coding genes that were not considered by previous searching procedures. Several small open reading frames exhibiting a statistically significant coding potential were thus identified as good candidates for functional analysis.

Amino Acid Sequence↗

Optimization of the synthesis of porcine somatotropin in Escherichia coli.

We report on the influence of choice of promoter and RNA polymerase, 5'-untranslated regions and ribosome binding sites, codon usage, leader peptide coding sequences and poly A tail in the 3'-untranslated region on the synthesis of porcine somatotropin (PST) in Escherichia coli. A total of 12 different constructs were tested in this study for the production of porcine somatotropin (PST) in E. coli. Several factors have significant effects on PST synthesis. In the presence of a strong promoter and a strong ribosome binding site, the next most important factor seems to be the combination of sequences at the 5'-end of the mRNA including both the 5'-untranslated region and the start of the coding sequence. Codon usage in the 5'-coding sequence per se is not important in determining the level of PST synthesis where high level expression is achieved from a strong ribosome binding site. However, where low level synthesis of recombinant PST (rPST) is achieved, codon usage in the 5'-coding sequence is important in determining the level of PST synthesis. Leader sequences dramatically reduce the level of PST synthesis. The presence of a poly A tail in the 3'-untranslated region has no significant effect on PST synthesis.

Animals↗

The 'effective number of codons' revisited.

Frank Wright [Gene 87 (1990) 23] derived a formula for calculation of a quantity termed the 'effective number of codons' (Nc) based on codon homozygosities. This quantity is a number between 20 and 61 and tells to what degree the codon usage in a gene is biased, i.e., it approaches 20 codons for the extremely biased genes, and approaches 61 for the genes where all possible codons are used with no preference. Among the different measures of codon bias Nc is considered the most useful and has found widespread use in papers dealing with codon usage phenomena. In this paper, the mathematical behaviours of codon homozygosities and Nc are evaluated, using Escherichia coli as the model organism. The results indicate that the classical formula for calculation of Nc could appropriately be substituted under circumstances, where there is bias discrepancy, i.e., when one amino acid (or more) within a degeneracy group is associated with strong codon bias while at the same time others in the same degeneracy group have little bias. An alternative estimator, termed Nc, is proposed and tested against Nc, and performs better when there is such bias discrepancy.

Codon↗

Chromosomal localization of the human hexabrachion (tenascin) gene and evidence for recent reduplication within the gene.

Using analysis of rodent-human somatic cell hybrids as well as in situ hybridization of hexabrachion cDNA probes to normal human metaphase chromosomes, we have localized the human hexabrachion gene to chromosome 9, bands q32-q34. We also put forward the hypothesis that there has been a recent reduplication of a small segment of the human hexabrachion gene. We support this hypothesis by comparison of codon usage in this segment of the gene to codon usage in the remainder of the gene. This hypothesis is also supported by comparison of the sequence of human hexabrachion to that of the chicken hexabrachion. In addition, the latter comparison shows that the reduplication most likely occurred after the divergence of mammalian and avian species.

Amino Acid Sequence↗

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4↗

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence↗

The contributions of replication orientation, gene direction, and signal sequences to base-composition asymmetries in bacterial genomes.

Asymmetries in base composition between the leading and the lagging strands have been observed previously in many prokaryotic genomes. Since a majority of genes is encoded on the leading strand in these genomes, previous analyses have not been able to determine the relative contribution to the base composition skews of replication processes and transcriptional and/or translational forces. Using qualitative graphical presentations and quantitative statistical analyses (analysis of variance), we have found that a significant proportion of the GC and AT skews can be attributed to replication orientation, i.e., the sequence of a gene is influenced by whether it is encoded on the leading or lagging strand. This effect of replication orientation on skews is independent of, and can be opposite in sign to, the effects of transcriptional or translational processes, such as selection for codon usage, amino acid preferences, expression levels (inferred from codon adaptation index), or potential short signal sequences (e.g., chi sequences). Mutational differences between the leading and the lagging strands are the most likely explanation for a significant proportion of the base composition skew in these bacterial genomes. The finding that base composition skews due to replication orientation are independent of those due to selection for function of the encoded protein may complicate the interpretation of phylogenetic relationships, conserved positions in nucleotide or amino acid sequence alignments, and codon usage patterns.

Analysis of Variance↗

A systematic method to identify genomic islands and its applications in analyzing the genomes of Corynebacterium glutamicum and Vibrio vulnificus CMCP6 chromosome I.

MOTIVATION: Some genomic islands contain horizontally transferred genes, which play critical roles in altering the genotypes and phenotypes of organisms, and horizontal gene transfer has been recognized as a universal event throughout bacterial evolution. A windowless method to display the distribution of genomic GC content, the cumulative GC profile, is proposed to identify genomic islands in genomes whose complete genome sequences are available. Two new indices are proposed to assess the codon usage bias and amino acid usage bias in genomic islands. RESULTS: A 211 kb genomic island (CGGI-1) has been identified in the genome of Corynebacterium glutamicum, and three genomic islands VVGI-1, VVGI-2 and VVGI-3, with lengths 167, 40 and 33 kb, respectively, have been identified in the genome of Vibrio vulnificus CMCP6 chromosome I. The CGGI-1 is flanked by two approximately 500 bp direct repeats, and utilizes a Val-tRNA as the integration site. For the VVGI-1 and VVGI-2, each has an integrase gene at 5' junction. All the identified genomic islands show unusual GC content, codon usage and amino acid usage, compared with the rest of the genomes. In addition, it is found that genomic islands are fairly homogenous in terms of GC content variation. An index, h, to quantify the homogeneity of GC content for genomic islands is proposed, and it is shown that h is less than 0.1 for all the genomic islands analyzed. The cumulative GC profile, as well as various indices to assess the codon usage bias, amino acid usage bias and homogeneity of the genomic islands, will be useful in the analysis of other genomes. AVAILABILITY: Programs used in this work and numerical results are available upon request.

Algorithms↗

Detecting anomalous gene clusters and pathogenicity islands in diverse bacterial genomes.

A gene in a genome is defined as putative alien (pA) if its codon usage difference from the average gene exceeds a high threshold and codon usage differences from ribosomal protein genes, chaperone genes and protein-synthesis-processing factors are also high. pA gene clusters in bacterial genomes are relevant for detecting genomic islands (GIs), including pathogenicity islands (PAIs). Four other analyses appropriate to this task are G+C genome variation (the standard method); genomic signature divergences (dinucleotide bias); extremes of codon bias; and anomalies of amino acid usage. For example, the cagA domain of Helicobacter pylori is highly deviant in its genome signature and codon bias from the rest of the genome. Using these methods we can detect two potential PAIs in the Neisseria meningitidis genome, which contain hemagglutinin and/or hemolysin-related genes. Additionally, G+C variation and genome signature differences of the Mycobacterium tuberculosis genome indicate two pA gene clusters.

Bacteria↗

Rare codons in E. coli and S. typhimurium signal sequences.

Codon usage has been examined in the signal sequences of 27 genes encoding proteins which possess leader peptides, and are inner-membrane located or exported. The results have been compared with codon usage in the corresponding coding sequences of most of the mature proteins. A bias is observed in the usage of rare codons for two of the three hydrophobic amino acids for which there are rare codons. Since hydrophobic residues are predominant in leader peptides, we suggest that a resulting concentration of rare codons in the signal sequence may play a role (or have played a role in the evolutionary past) in the secretion process by delaying translation.

Base Sequence↗

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence↗

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals↗