PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Cloning and hemolysin-mediated secretory expression of a codon-optimized synthetic human interleukin-6 gene in Escherichia coli.

Previously, we constructed human interleukin-6 (hIL-6)-secreting Escherichia coli and Salmonella typhimurium strains by fusion of the hIL-6 cDNA to the HlyA(s) secretional signal, utilizing the hemolysin export apparatus for extracellular delivery of a bioactive hIL-6-hemolysin (hIL-6-HlyA(s)) fusion protein. Molecular analysis of the secretion process revealed that low secretion levels were due to inefficient gene expression. To adapt the codon usage in hIL-6 cDNA to the E. coli codon bias, a synthetic hIL-6Ec gene variant was constructed from 20 overlapping oligonucleotides, yielding a 561-bp fragment, which comprises the complete hIL-6 cDNA sequence. Genetic fusion of the hIL-6Ec gene with the hlyA(s) secretional signal as an integral part of the hemolysin operon resulted in 3-fold higher hIL-6-HlyA(s) secretion levels in E. coli, compared to a strain expressing the original hIL-6-hlyA(s) fusion gene. An increase in the electrophoretic mobility of secreted hIL-6-HlyA(s) in non-reducing SDS-PAGE, similar to that found for recombinant mature hIL-6, and the absence of such a mobility shift in the intracellular hIL-6-HlyA(s) protein fraction indicated that in hIL-6-HlyA(s) most probably correct intramolecular disulfide bond formation occurred during the secretion step. To confirm the disulfide bond formation, hIL-6-HlyA(s) was purified by a single-step immunoaffinity chromatography from culture supernatant in yields of 18 microg/L culture supernatant with purity in the range of 60%. These results demonstrate that codon usage has an impact on the hemolysin-mediated secretion of hIL-6 and, furthermore, provide evidence that the hemolysin system enables secretory delivery of disulfide-bridged proteins.

Amino Acid Sequence↗

Genomic background drives the divergence of duplicated amylase genes at synonymous sites in Drosophila.

In some Drosophila species, there are two types of greatly diverged amylase (Amy) genes (Amy clusters 1 and 2), each encoding active amylase isozymes. Cluster 1 is located at the middle of its chromosomal arm, and the region has a normal local recombination rate. However, cluster 2 is near the centromere, and this region is known to have a reduced recombination rate. Although nonsynonymous substitutions follow a molecular clock, synonymous substitutions were accelerated in cluster 2 after gene duplications. This resulted in a higher GC content at the third codon position (GC3) and codon usage bias in cluster 1, and lower GC3 content and codon usage bias in the cluster 2. However, no systematic difference in GC content was observed in the first and second codon positions or the 3'-flanking regions. Therefore, differences in local recombination rate rather than mutation bias might explain the divergence at synonymous sites between the two Amy clusters within species (Hill-Robertson effect). Alternatively, the different patterns and levels of expression between the two clusters may imply that the reduced expression level in cluster 2 caused by chromatin potentiation decreased the codon bias. Both of these hypotheses imply the importance of the genomic background as a driving force of divergence between non-tandemly duplicated genes.

3' Flanking Region↗

Molecular characterization of the tdc operon of Escherichia coli K-12.

The nucleotide sequence of a 2-kilobase DNA fragment of the tdc region of Escherichia coli K-12, previously cloned in this laboratory, revealed two open reading frames, tdcC and ORFX, downstream from the tdcB gene (formerly designated tdc) encoding biodegradative threonine dehydratase. A 24-base-pair sequence separated tdcC from the dehydratase coding region, and an untranslated region of 60 nucleotides, which contains a recognizable -10 consensus sequence, was found between tdcC and ORFX. The deduced amino acid sequence of tdcC showed it to be a large hydrophobic polypeptide of 431 amino acid residues, whereas ORFX coded for a small 135-residue polypeptide lacking glutamine and tryptophan. A computer-assisted sequence analysis revealed no similarity among the tdcB, tdcC, and ORFX polypeptides, and a search of the GenBank database failed to detect similarity with any other known proteins. The tdc genes and ORFX showed similar codon usage and, in analogy with other bacterial genes, showed codon usage typical for genes expressed at an intermediate level. Transcriptional analysis with S1 nuclease indicated two distinct transcription start sites upstream of the tdcB gene in regions previously identified as promoterlike elements P1 and P2. Interestingly, expression of tdcB and tdcC, but not ORFX, was contingent upon the presence of P1. These results taken together tend to suggest that the biodegradative threonine dehydratase is the second gene in a polycistronic transcription unit constituting a novel operon (tdcABC) in E. coli implicated in anaerobic threonine metabolism.

Amino Acid Sequence↗

[Cloning and identification of the priming glycosyltransferase gene involved in exopolysaccharide 139A biosynthesis in Streptomyces].

Recently in our laboratory, Streptomyces sp. 139 has been identified to produce a new exopolysaccharide designated EPS 139A that shows anti-rheumatic arthritis activity. The strategy of studying EPS 139A biosynthesis is to clone the key gene in the EPS biosynthesis pathway, i.e. the priming glycosyltransferase gene catalyzing the first step of nucleotide sugar transfer. Degenerate primers-based PCR approach was adopted to isolate the putative priming glycosyltransferase gene in Streptomyces sp. 139. According to the genes encoding the priming glycosyltransferases that have been identified in several microorganisms, a multiple alignment of the amino acid sequences of these genes was used to identify regions conserved between all genes. To clone the priming glycosyltransferase gene in Streptomyces sp. 139, degenerate primers were designed from these conserved regions taking into account information on Streptomyces codon usage to amplify an internal DNA fragment of this gene. A distinctive PCR product with the expected size of 0.3 kb was amplified from Streptomyces sp. 139 total genomic DNA. Sequence analysis showed that it is part of a putative priming glycosyltransferase gene and contains the predicted conserved domain B. To isolate the complete priming glycosyltransferase gene, a Streptomyces sp. 139 genomic library was constructed in the E. coli--Streptomyces shuttle vector pOJ446. Using the 0.3 kb PCR product of priming glycosyltransferase gene as a probe, 17 positive colonies were isolated by colony hybridization. A 4.0 kb BamHI fragment from all positive cosmids that hybridized to this probe was sequenced, which revealed the complete priming glycosyltransferase gene. The priming glycosyltransferase gene ste5 (GenBank under accession number AY131229) most likely begins with GTG, preceded by a probable ribosome binding site (RBS), GGGGA. It encodes a 492-amino-acid protein with molecular weight of 54 kDa and isoelectric point of 10.6. The G + C content of ste5 is 73%, close to the average of G + C content (74%) for Streptomyces. Moreover, the preference usage of G or C as third base of codons are found in the ste5, which is in accordance with the Streptomyces codon usage. A BlastP search showed that the C-terminal region of Ste5 shows highly homology with a number of priming glycosyltransferases from many different organisms. Ste5 contains two putative catalytic residues, Glu and Asp (residues 423 and 474) with a spacing of approximately 50 amino acids that conserved in various beta-glycosyltransferases. Moreover, the C-terminal one third of Ste5 contains three domains, A, B and C that is reported to be common to glycosyltransferases. By hydrophilicity plot prediction, the N-terminal two thirds of Ste5 exhibits 5 putative transmembrane domains. To investigate the involvement of the identified polysaccharide gene cluster in EPS 139A biosynthesis, the gene ste5 encoding priming glycosyltransferase was insertionally disrupted by a single-crossover homologous recombination event. A 0.85 kb internal fragment of ste5 was cloned into vector pKC1139 to yield pLY5015 that was transduced into Streptomyces sp. 139. Correct integration in Streptomyces LY1001 ste5- mutant strain was confirmed by Southern hybridization. After fermentation, no EPS 139A could be detected in the cultures of ste5- mutant strain Streptomyces LY1001. Therefore, the gene ste5 identified in this work is involved in the synthesis of the Streptomyces sp. 139 EPS.

Amino Acid Sequence↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

Organization and expression of algal (Chlamydomonas reinhardtii) mitochondrial DNA.

The mitochondrial genome of Chlamydomonas reinhardtii, a unicellular green alga, is a linear 15.8 kilobase pair (kbp) molecule. In gene arrangement and mode of expression, as well as in size, it differs radically from the large (200-2400 kbp) mitochondrial genomes of higher plants. Heterologous hybridization experiments and nucleotide sequence analysis have revealed that C. reinhardtii mitochondrial DNA (mtDNA) is a compactly organized genome specifying at least eight proteins, a minimum of three transfer RNAs, and large subunit (LS) and small subunit (SS) ribosomal RNAs. Both strands of the mtDNA encode genetic information, with genes organized into perhaps a single transcriptional unit on each strand. Stable transcripts have been identified by Northern hybridization analysis, and transcript termini have been mapped by primer extension and S1 nuclease protection experiments. The results suggest that mature RNAs, which virtually saturate the genome, are generated by precise endonucleolytic cleavage of long precursors, with specific motifs (both primary sequence and secondary structure) implicated as processing signals. Codon usage in C. reinhardtii mitochondria is highly biased, with eight codons entirely absent from all protein-coding genes; however, even though codon usage is restricted, it appears that C. reinhardtii mtDNA cannot encode the minimum number of tRNAs needed to support mitochondrial protein synthesis. The most striking feature of C. reinhardtii mtDNA is the division of SS and LS rRNA genes into a number of separate subgenic coding segments ('modules') that are interspersed with one another and with protein-coding and tRNA genes. We have identified abundant small RNAs, transcribed from these modules, that approximate to the latter in size. This indicates that splicing of rRNA 'pieces' does not occur in this system. Rather, the mature rRNAs apparently exist and function as non-covalent complexes of small RNAs (four in SS rRNA, at least eight in LS rRNA), held together by intermolecular base pairing. These complexes contain all the conserved elements of the minimal secondary structures that define the functional core of conventional LS and SS rRNAs.

Base Sequence↗

A survey of codon and amino acid frequency bias in microbial genomes focusing on translational efficiency.

Unequal use of synonymous codons has been found in several prokaryotic and eukaryotic genomes. This bias has been associated with translational efficiency. The prevalence of this bias across lineages is currently unknown. Here, a new method (GCB) to measure codon usage bias is presented. It uses an iterative approach for the determination of codon scores and allows the computation of an index of codon bias suitable for interspecies comparison. A server to calculate GCB-values of individual genes as well as a list of compiled results are available at www.g21.bio.uni-goettingen.de. The method was applied to complete bacterial genomes. The relation of codon usage bias with amino acid composition and the choice of stop codons were determined and discussed.

Amino Acids↗

Statistical, computational and visualization methodologies to unveil gene primary structure features.

OBJECTIVES: Gene sequence features such as codon bias, codon context, and codon expansion (e.g. trinucleotide repeats) can be better understood at the genomic scale level by combining statistical methodologies with advanced computer algorithms and data visualization through sophisticated graphical interfaces. This paper presents the ANACONDA system, a bioinformatics application for gene primary structure analysis. METHODS: Codon usage tables using absolute metrics and software for multivariate analysis of codon and amino acid usage are available in public databases. However, they do not provide easy computational and statistical tools to carry out detailed gene primary structure analysis on a genomic scale. We propose the usage of several statistical methods--contingency table analysis, residual analysis, multivariate analysis (cluster analysis)--to analyze the codon bias under various aspects (degree of association, contexts and clustering). RESULTS: The developed solution is a software application that provides a user-guided analysis of codon sequences considering several contexts and codon usage on a genomic scale. The utilization of this tool in our molecular biology laboratory is focused on particular genomes, especially those from Saccharomyces cerevisiae, Candida albicans and Escherichia coli. In order to illustrate the applicability and output layouts of the software these species are herein used as examples. CONCLUSIONS: The statistical tools incorporated in the system are allowing to obtain global views of important sequence features. It is expected that the results obtained will permit identification of general rules that govern codon context and codon usage in any genome. Additionally, identification of genes containing expanded codons that arise as a consequence of erroneous DNA replication events will permit uncovering new genes associated with human disease.

Algorithms↗

Optimizing heterologous expression in dictyostelium: importance of 5' codon adaptation.

Expression of heterologous proteins in Dictyostelium discoideum presents unique research opportunities, such as the functional analysis of complex human glycoproteins after random mutagenesis. In one study, human chorionic gonadotropin (hCG) and human follicle stimulating hormone were expressed in Dictyostelium. During the course of these experiments, we also investigated the role of codon usage and of the DNA sequence upstream of the ATG start codon. The Dictyostelium genome has a higher AT content than the human, resulting in a different codon preference. The hCG-beta gene contains three clusters with infrequently used codons that were changed to codons that are preferred by Dictyostelium. The results reported here show that optimizing the first 5-17 codons of the hCG gene contributes to 4- to 5-fold increased expression levels, but that further optimization has no significant effect. These observations suggest that optimal codon usage contributes to ribosome stabilization, but does not play an important role during the elongation phase of translation. Furthermore, adapting the 5'-sequence of the hCG gene to the Dictyostelium 'Kozak'-like sequence increased expression levels approximately 1.5-fold. Thus, using both codon optimization and 'Kozak' adaptation, a 6- to 8-fold increase in expression levels could be obtained for hCG.

Amino Acid Sequence↗

The Vitreoscilla hemoglobin gene: molecular cloning, nucleotide sequence and genetic expression in Escherichia coli.

Vitreoscilla hemoglobin is involved in oxygen metabolism of this bacterium, possibly in an unusual role for a microbe. We have isolated the Vitreoscilla hemoglobin structural gene from a pUC19 genomic library using mixed oligodeoxy-nucleotide probes based on the reported amino acid sequence of the protein. The gene is expressed in Escherichia coli from its natural promoter as a major cellular protein. The nucleotide sequence, which is in complete agreement with the known amino acid sequence of the protein, suggests the existence of promoter and ribosome binding sites with a high degree of homology to consensus E. coli upstream sequences. In the case of at least some amino acids, a codon usage bias can be detected which is different from the biased codon usage pattern in E. coli. The downstream sequence exhibits homology with the 3' end sequences of several plant leghemoglobin genes. E. coli cells expressing the gene contain greater than fivefold more heme than controls.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of the gene of the molybdenum-containing aldehyde oxido-reductase of Desulfovibrio gigas. The deduced amino acid sequence shows similarity to xanthine dehydrogenase.

In this report, we describe the isolation of a 4020-bp genomic PstI fragment of Desulfovibrio gigas harboring the aldehyde oxido-reductase gene. The aldehyde oxido-reductase gene spans 2718 bp of genomic DNA and codes for a protein with 906 residues. The protein sequence shows an average 52% (+/- 1.5%) similarity to xanthine dehydrogenase from different organisms. The codon usage of the aldehyde oxidoreductase is almost identical to a calculated codon usage of the Desulfovibrio bacteria.

Aldehyde Oxidoreductases↗

[Analysis of apolipoprotein gene family in codon space--non-random selection of nucleotide changes in evolution].

The choice of nucleotide changes in DNA evolution can be either selectively neutral or biased. To study how apolipoprotein gene selects the nucleotide substitutions in the course of evolution, a codon space is constructed in which its DNA sequence can be mapped as a matrix of nucleotide frequencies in three codon positions. Accordingly, a number of methods that measure the nonrandomness of nucleotide distribution in codon space are developed based on maximum entropy techniques to define the nature of nucleotide change selection in evolution. By these methods, we demonstrated that the nucleotide composition in 1st and 3rd codon position of apolipoprotein genes is highly nonrandom, which appears to be a result of non-neutral selection of codon positions by adenosine and thymidine. In addition, this paper is also concerned in the divergence of synonymos codon usage and its correlation to taxonomic distances among species. As a result, a codon usage clock was reported in apolipoprotein A-I. Our studies suggest that non-random selection of nucleotide changes in codon space may represent an evolutionary characteristics of apolipoprotein genes.

Animals↗

A DnaB intein in Rhodothermus marinus: indication of recent intein homing across remotely related organisms.

A dnaB gene encoding a homologue of the Escherichia coli DNA helicase DnaB was cloned and sequenced in the thermophilic eubacterium Rhodothermus marinus, predicting a DnaB protein that harbors an intein. This DnaB intein is 428 amino acid residues long, has several putative intein sequence motifs (including two putative endonuclease motifs), and is capable of protein splicing when produced in E. coli cells. The R. marinus DnaB intein is a close homologue of a DnaB intein in the cyanobacterium Synechocystis sp. strain PCC6803. The two inteins are positioned identically in their respective DnaB proteins. They also share a 54% sequence identity (74% sequence similarity) that is markedly higher than the 37% sequence identity shared by the extein sequences of the two DnaB proteins. Horizontal intein transfer (homing) is therefore invoked to relate these two DnaB inteins. The codon usage of R. marinus DnaB intein coding sequence differs markedly from the codon usages of its flanking extein coding sequences and other genes in the same genome, suggesting more recent acquisition of the DnaB intein in this organism.

Amino Acid Sequence↗

How mitochondria redefine the code.

Annotated, complete DNA sequences are available for 213 mitochondrial genomes from 132 species. These provide an extensive sample of evolutionary adjustment of codon usage and meaning spanning the history of this organelle. Because most known coding changes are mitochondrial, such data bear on the general mechanism of codon reassignment. Coding changes have been attributed variously to loss of codons due to changes in directional mutation affecting the genome GC content (Osawa and Jukes 1988), to pressure to reduce the number of mitochondrial tRNAs to minimize the genome size (Anderson and Kurland 1991), and to the existence of transitional coding mechanisms in which translation is ambiguous (Schultz and Yarus 1994a). We find that a succession of such steps explains existing reassignments well. In particular, (1) Genomic variation in the prevalence of a codon's third-position nucleotide predicts relative mitochondrial codon usage well, though GC content does not. This is because A and T, and G and C, are uncorrelated in mitochondrial genomes. (2) Codons predicted to reach zero usage (disappear) do so more often than expected by chance, and codons that do disappear are disproportionately likely to be reassigned. However, codons predicted to disappear are not significantly more likely to be reassigned. Therefore, low codon frequencies can be related to codon reassignment, but appear to be neither necessary nor sufficient for reassignment. (3) Changes in the genetic code are not more likely to accompany smaller numbers of tRNA genes and are not more frequent in smaller genomes. Thus, mitochondrial codons are not reassigned during demonstrable selection for decreased genome size. Instead, the data suggest that both codon disappearance and codon reassignment depend on at least one other event. This mitochondrial event (leading to reassignment) occurs more frequently when a codon has disappeared, and produces only a small subset of possible reassignments. We suggest that coding ambiguity, the extension of a tRNA's decoding capacity beyond its original set of codons, is the second event. Ambiguity can act alone but often acts in concert with codon disappearance, which promotes codon reassignment.

Base Composition↗

Structure and expression of the gene encoding the periplasmic arylsulfatase of Chlamydomonas reinhardtii.

Chlamydomonas reinhardtii produces a periplasmic arylsulfatase in response to sulfur deprivation. We have isolated and sequenced arylsulfatase cDNAs from a lambda gt11 expression library. The amino acid sequence of the protein, as deduced from the nucleotide sequence, has features characteristic of secreted proteins, including a signal sequence and putative glycosylation sites. The gene has a broad codon usage with seven codons, all having A residues in the third position, not previously observed in C. reinhardtii genes. Arylsulfatase transcription is tightly regulated by sulfur availability. The approximately 2.7 kb arylsulfatase transcript is very susceptible to degradation, disappearing in less than an hour after sulfur starved cells are administered either sulfate or alpha-amanitin. The accumulation of the arylsulfatase transcript is also suppressed by the addition of cycloheximide. Transcription initiation from the arylsulfatase gene occurs approximately 100 bp upstream of the initiation codon, in a region that is 5' to a 43 bp imperfect inverted repeat. Preceding the transcription start site are sequences similar to those present in promoter regions of other genes from C. reinhardtii.

Amino Acid Sequence↗

Development of Polymorphic EST Markers Suitable for Genetic Linkage Mapping of Catfish.

: Expressed sequence tag (EST) markers are important for gene mapping and for marker-assisted selection (MAS). To develop EST markers for use in catfish gene mapping, 100 randomly picked complementary DNAs from the channel catfish (Ictalurus punctatus) pituitary library were sequenced. The EST sequences were used to design primers to amplify channel catfish and blue catfish (I. furcatus) genomic DNAs. Polymerase chain reaction products of the ESTs were analyzed to determine length polymorphism between the channel catfish and blue catfish. Eleven polymorphic EST markers were identified. Five of the 11 EST markers were from known genes and the other six were from unidentified ESTs. Seven ESTs were found to be associated with microsatellite sequences. Analysis of channel catfish gene sequences indicated highly biased codon usage, with 16 codons being preferably used. These codons were more preferably used in highly expressed ribosomal protein genes and in highly expressed pituitary hormone genes. G/C-rich codons are less used in channel catfish than those in other vertebrates suggesting AT-richness of the channel catfish genome.

Journal Article↗

Catalyzing bacterial speciation: correlating lateral transfer with genetic headroom.

Unlike crown eukaryotic species, microbial species are created by continual processes of gene loss and acquisition promoted by horizontal genetic transfer. The amounts of foreign DNA in bacterial genomes, and the rate at which this is acquired, are consistent with gene transfer as the primary catalyst for microbial differentiation. However, the rate of successful gene transfer varies among bacterial lineages. The heterogeneity in foreign DNA content is directly correlated with amount of genetic headroom intrinsic to a bacterial species. Genetic headroom reflects the amount of potentially dispensable information--reflected in codon usage bias and codon context bias--that can be transiently sacrificed to allow experimentation with functions introduced by gene transfer. In this way, genetic headroom offers a potential metric for assessing the propensity of a lineage to speciate.

Bacteria↗

Molecular population genetics and evolution of a prion-like protein in Saccharomyces cerevisiae.

The prion-like behavior of Sup35p, the eRF3 homolog in the yeast Saccharomyces cerevisiae, mediates the activity of the cytoplasmic nonsense suppressor known as [PSI(+)]. Sup35p is divided into three regions of distinct function. The N-terminal and middle (M) regions are required for the induction and propagation of [PSI(+)] but are not necessary for translation termination or cell viability. The C-terminal region encompasses the termination function. The existence of the N-terminal region in SUP35 homologs of other fungi has led some to suggest that this region has an adaptive function separate from translation termination. To examine this hypothesis, we sequenced portions of SUP35 in 21 strains of S. cerevisiae, including 13 clinical isolates. We analyzed nucleotide polymorphism within this species and compared it to sequence divergence from a sister species, S. paradoxus. The N domain of Sup35p is highly conserved in amino acid sequence and is highly biased in codon usage toward preferred codons. Amino acid changes are under weak purifying selection based on a quantitative analysis of polymorphism and divergence. We also conclude that the clinical strains of S. cerevisiae are not recently derived and that outcrossing between strains in S. cerevisiae may be relatively rare in nature.

Amino Acid Sequence↗