PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The mouse int-2 gene exhibits basic fibroblast growth factor activity in a basic fibroblast growth factor-responsive cell line.

The int-2 protein is related to basic fibroblast growth factor (bFGF) by amino acid sequence homology. To assess its biological activity, we constructed retroviral vectors containing four variants of mouse int-2 complementary DNA under the transcriptional control of the beta-actin promoter and tested their effects on human SW13 adrenal cortical tumor cells. This cell line specifically requires bFGF, interleukin 1, or transforming growth factor e for anchorage-independent growth in soft agar. Despite encoding a signal sequence that should direct the protein to the secretory pathway, vectors containing unmodified int-2 complementary DNA, or a form optimized for translation initiation at the AUG codon, were incapable of inducing SW13 growth in soft agar. However, SW13 transfectants expressing a construct (pSP1), in which a mouse immunoglobulin signal peptide sequence is linked to the int-2 coding sequences, grew well in soft agar. The concentrated conditioned medium from these pSP1-transfected cells supported anchorage-independent growth of SW13 indicator cells and competed with bFGF for binding to receptors. Western blot analysis with an int-2-specific antiserum detected Mr 30,000-32,000 int-2 products in cell extracts and conditioned medium from pSP1-transfected clones, whereas the conditioned medium from these and other SW13 clones contained only low levels of bFGF as measured in a specific radioimmunoassay. These data suggest that the product of the int-2 gene can functionally replace bFGF in modulating the anchorage-independent growth of SW13 cells.

Amino Acid Sequence↗

[Expression and purification of recombinant human interleukin 4 in Escherichia coli].

Human interleukin 4 (IL-4) cDNA was optimized and synthesized according to E. coli preferred codon. A recombinant expression plasmid pET-30a (+)/rhIL-4 was constructed with the target cDNA inserted between Nde I and EcoR I sites, which can translate the mature IL-4 protein with an extra methionine residue at N-terminal. The expression vector was transformed into E. coli BL21 (DE3). The rhIL-4 protein was expressed in the inclusion body. By using the optimized fermentation conditions, the high expression level was achieved with the expression level as high as 35% of total protein obtained. A purification strategy has been designed which includes Q-Sepharose and SP-Sepharose ion-exchange chromatography and dialysis renaturation. The rhIL-4 was purified with the purity more than 98% and the yield of 40 mg per liter fermentation culture achieved. Western blot proved that the purified protein is IL-4. Amino acid sequencing revealed that N-terminal 16 residue sequence is identical to the theoretical sequence. Biological activity assay on TF-1 cells demonstrated that the rhIL-4 is active with an activity of 2.5 x 10(6) AU/mg. This study promises large scale production of rhIL-4.

Amino Acid Sequence↗

Multi-line split DNA synthesis: a novel combinatorial method to make high quality peptide libraries.

BACKGROUND: We developed a method to make a various high quality random peptide libraries for evolutionary protein engineering based on a combinatorial DNA synthesis. RESULTS: A split synthesis in codon units was performed with mixtures of bases optimally designed by using a Genetic Algorithm program. It required only standard DNA synthetic reagents and standard DNA synthesizers in three lines. This multi-line split DNA synthesis (MLSDS) is simply realized by adding a mix-and-split process to normal DNA synthesis protocol. Superiority of MLSDS method over other methods was shown. We demonstrated the synthesis of oligonucleotide libraries with 1016 diversity, and the construction of a library with random sequence coding 120 amino acids containing few stop codons. CONCLUSIONS: Owing to the flexibility of the MLSDS method, it will be able to design various "rational" libraries by using bioinformatics databases.

Amino Acid Sequence↗

Involvement of the 5' proximal coding sequences of hepatitis C virus with internal initiation of viral translation.

The 5' nontranslated region (NTR) of hepatitis C virus (HCV) consists of 341 nucleotides (nt). This region comprises the majority of the internal ribosome entry site (IRES) which controls the efficiency of viral translation. Previous studies of the 3' boundary of the HCV IRES yielded conflicting data regarding the involvement of viral coding sequences in IRES activity. We therefore studied the functional significance of the 5' proximal coding sequences of the HCV core gene on IRES activity. We constructed monocistronic and bicistronic DNAs that contained either a chloramphenicol acetyl transferase (CAT) gene or a luciferase (Luc) gene as the reporter. Results from both in vitro and in vivo experiments indicated that the optimal IRES ranged within nt 1-371. Further mutational analyses of sequences surrounding the initiation codon revealed that primary sequences downstream of the AUG initiator rather than the secondary structure are important in regulating optimal IRES function. We are also able to demonstrate that a non-AUG codon could be used to initiate the synthesis of a reporter protein, albeit with lower efficiency. These findings bear important implications for the HCV IRES secondary structures.

5' Untranslated Regions↗

Combinatorial codons: a computer program to approximate amino acid probabilities with biased nucleotide usage.

Using techniques from optimization theory, we have developed a computer program that approximates a desired probability distribution for amino acids by imposing a probability distribution on the four nucleotides in each of the three codon positions. These base probabilities allow for the generation of biased codons for use in mutational studies and in the design of biologically encoded libraries. The dependencies between codons in the genetic code often makes the exact generation of the desired probability distribution for amino acids impossible. Compromises are often necessary. The program, therefore, not only solves for the "optimal" approximation to the desired distribution (where the definition of "optimal" is influenced by several types of parameters entered by the user), but also solves for a number of "sub-optimal" solutions that are classified into families of similar solutions. A representative of each family is presented to the program user, who can then choose the type of approximation that is best for the intended application. The Combinatorial Codons program is available for use over the web from http://www.wi.mit.edu/kim/computing.html.

Amino Acids↗

Point mutations define a sequence flanking the AUG initiator codon that modulates translation by eukaryotic ribosomes.

By analyzing the effects of single base substitutions around the ATG initiator codon in a cloned preproinsulin gene, I have identified ACCATGG as the optimal sequence for initiation by eukaryotic ribosomes. Mutations within that sequence modulate the yield of proinsulin over a 20-fold range. A purine in position -3 (i.e., 3 nucleotides upstream from the ATG codon) has a dominant effect; when a pyrimidine replaces the purine in position -3, translation becomes more sensitive to changes in positions -1, -2, and +4. Single base substitutions around an upstream, out-of-frame ATG codon affect the efficiency with which it acts as a barrier to initiating at the downstream start site for preproinsulin. The optimal sequence for initiation defined by mutagenesis is identical to the consensus sequence that emerged previously from surveys of translational start sites in eukaryotic mRNAs. The mechanism by which nucleotides flanking the ATG codon might exert their effect is discussed.

Animals↗

Growth-rate-dependent accumulation of twelve tRNA species in Escherichia coli.

We have previously shown that in Escherichia coli the accumulation of five leucine and three methionine tRNA species is regulated so that those tRNA species that translate major codons increase while those that translate minor codons decrease as the growth rate increases. Here, we have analyzed the growth-rate-dependence of another 12 tRNA species. We find that the level of three tRNA species cognate to the major glycine, proline and arginine codons, respectively, increase with increasing growth rates. Conversely, four tRNAs that are cognate to minor codons within the same amino acid families decrease with increasing growth rates. In addition, the glutamyl as well as the phenylalanyl isoacceptor species are accumulated in proportion to the content of these two amino acids in the proteins produced at different growth rates. In summary, the patterns of the growth-rate-dependence for the accumulation of these 17 tRNA species support the interpretation that the major codon preference is an arrangement to maximize the growth rates of bacteria in rich media by optimizing the kinetic efficiency of translation. In contrast, we find that three minor tRNA species cognate to two rare arginine codons and one minor glycine codon, respectively, increase with increasing growth rate. Such findings suggest that there are additional constraints on the accumulation of these tRNA species that may be distinct from those required to optimize the kinetic efficiency of translation.

Base Sequence↗

Genomic mapping and sequence analysis of the fowl adenovirus serotype 10 hexon gene.

The gene for the major capsid protein (hexon) of fowl adenovirus serotype 10 (FAV-10) has been identified by the use of the expression vector pGEX and rabbit polyclonal antisera raised against FAV-10. The nucleotide sequence of the entire hexon gene has been determined. Sequence analysis revealed an open reading frame of 2808 bp coding for a putative polypeptide 936 amino acids long with a molecular mass of 105.5 kDa. The translation initiation codon has a local sequence which conforms with the optimal translation start sequence of CC(A/G)CCATGG. The location of the hexon gene in the FAV genome was from 46.85 to 52.81 map units, which is to the left of the hexon gene in the genomes of both bovine and human adenovirus (52.4 to 60.5 map units.). A splice acceptor site was identified 12 bp upstream of the initiation codon by using mRNA and PCR. It had the sequence TAGG which conforms to the consensus sequence of (C/T)AGG. Comparison of the amino acid sequence of the FAV-10 hexon with those of the bovine, human and murine hexon gene products revealed highest levels of identity occurring in the regions corresponding to the pedestals which form the base region of the hexon, and the lowest levels of identity in the regions corresponding to the loops which are exposed to the external environment.

Amino Acid Sequence↗

Codon usage pattern in alpha 2(I) chain domain of chicken type I collagen and its implications for the secondary structure of the mRNA and the synthesis pauses of the collagen.

A stability map of local secondary structure of the mRNA of the triple-helical alpha 2(I) chain domain of chicken type I collagen was obtained by plotting the free energy of the optimal secondary structure of a local segment in mRNA against the segment position along a base sequence of the mRNA. It was found that the positions of the minima of free energy in the plot coincide with the positions where synthesis pauses of the alpha-chain polypeptides of the corresponding sizes translated from the mRNA have been reported to occur (1). The codon usage pattern of each of the three major amino acids of the alpha-chain domain of the collagen, Gly, Pro and Ala, fluctuates considerably along the base sequence segments of the mRNA and a deviation of the pattern from that of the average of the whole alpha 2(I) chain domain mRNA, particularly for Gly codons, leads to a loss of the stability of the local secondary structure of the mRNA. The results suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA of the chicken collagen alpha 2(I) chain domain which leads to a nonuniform polypeptide elongation pattern.

Animals↗

Biological and molecular evidence for the transgenosis of genes from bacteria to plant cells.

Specialized transducing phages (lambda and varphi80) have been used as vectors in the transfer of genes (wild type and mutant) from the bacterium Escherichia coli to haploid cell lines of the plants Lycopersicon esculentum and Arabidopsis thaliana. The overall phenomenon of transfer, gene maintenance, transcription, translation, and function has been termed transgenosis. Transgenosis of galactose and lactose operon genes was detected by survival and growth of the plant cells on defined medium with galactose and lactose as sole sources of bulk carbon. Phages carrying a defective operon, unrelated bacterial genes, or no bacterial genes, do not affect the normal result of death on these media, nor do they prevent growth on optimal growth media. Transgenosis of the E. coli gene z (lac operon) was confirmed by a biochemical-immunological test specific for E. coli beta-galactosidase. Plant cells were unable to effectively suppress an E. coli nonsense mutation. The E. coli mutant suppressor gene, supF(+), specifies insertion of tyrosine at amber (UAG) nonsense codons. Introduction of supF(+) results in a lethal transgenosis on medium normally optimal for plant cell growth. It is concluded that amber codons are vital to the life of plant cells. Differentiating cells of A. thaliana were not affected by supF(+).

Journal Article↗

[Physico-chemical basis of the genetic code origin: stereochemical analysis of interactions of amino acids and nucleotides based on the progene hypothesis].

A progene hypothesis has been proposed earlier to explain the mechanism of origin of the self-reproducing genetic system. Progenes (precursors of the genetic system) are mixed anhydrides of an amino acid and deoxyribotrinucleotide at the 3'-gamma-terminal phosphate (NpNpNppp-AA); they are produced from dinucleotides (NpNp) and 3'-gamma-aminoacylnucleotidylates (Nppp-AA) as a result of specific interaction between amino acid and dinucleotide. The postulated mechanism of progene formation accounts for the selection of substances, including chirality, the origin of the genetic code as well as for the mechanisms of formation, self-reproduction and evolution of the simpliest genetic system ("gene--polypeptide"). A stereochemical analysis of the progene formation mechanism has allowed us to support the main statements of the hypothesis that relate to the origin of the genetic code and to selection of substances. Atomic groups that could be responsible for the specificity of interaction between dinucleotides and amino acids in progene formation have been revealed. Stereochemical evidence for the physicochemical basis of the origin of the existing genetic code have been produced: 1) a special role of the second nucleotide in the codon is demonstrated in amino acid coding by the progene hypothesis principle; 2) an advantage of T against U in such coding is demonstrated; 3) for 16 amino acids out of 20 an agreement has been obtained between the optimal dinucleotide as revealed by the stereochemical analysis and the codon dinucleotides; 4) an explanation for the third nucleotide selection mechanism is offered. A restoration of the prebiotic code, based on these results, has indicated that the code contains 32 codons, is statistical and group-wise. It encodes 7 groups of isofunctional amino acids: 3 overlapping groups of non-polar amino acids 1) medium-size hydrophobic amino acids (chiefly Val, n-Val and a-But), 2) small and medium-size non-polar amino acids (chiefly Ala Val, n-Val a-But and Gly), 3) small non-polar amino acids (Gly, Ala, a-But) and 4 groups of polar amino acids--1) hydroxy--+dicarbonic (Asp, Glu, Ser and Thr), 2) dicarbonic (Asp and Glu), 3) hydroxy (Ser and Thr) and 4) basic (Arg and Lys). The code includes about 20 amino acids among which are 15-17 canonical and a few common non-canonical. The prebiotic code explains many properties of the existing genetic code and is capable of evolving into the latter by way of a gradual replacement of the physicochemical coding mechanism by the enzymatic coding mechanism.

Amino Acids↗

[tRNA adaptation and the optimization of translation].

The intracellular level of each tRNA species is adjusted to the codon frequency of the mRNA being decoded. This was first observed in such highly differentiated cells as the silk gland of Bombyx mori, which produces fibroin and sericin, and the rabbit reticulocyte. tRNA adaptation also occurs in other cell types from E. coli to mammalian cells. Regardless of the mechanism regulating tRNA biosynthesis, we believe that tRNA adaptation is the basic step optimizing chain elongation at the ribosomal level. We propose the system of trial and error as a working model for the ribosome. This model clarifies the correlations between iso-accepting tRNA levels and codon frequencies, as well as the effect of tRNA pool balance on mean elongation rate and non-uniform individual elongation rate (depending on whether codons are rare or abundant) for fibroin mRNA translated in a reticulocyte cell-free system.

Adaptation, Physiological↗

Overexpression in Escherichia coli and characterization of the chloroplast fructose-1,6-bisphosphatase from wheat.

An important Calvin cycle enzyme, chloroplast fructose-1, 6-bisphosphatase (FBPase) from wheat, has been cloned and expressed up to 15% of the total cell protein using a pPLc expression vector in Escherichia coli by replacing the codons in the 5'-terminal encoding sequence with optimal and A/T-rich ones. The overexpressed wheat FBPase is soluble, fully active, and heat stable. It can be purified by chromatography in turn on DEAE-Sepharose and Sephacryl S-200, and around 15 mg of purified enzymes (>95%) is obtained from 1 liter of cultured bacteria. Its special activity is 8.8 u/mg, K(cat) is 22.9/S, K(m) is 121 microM, and V(max) is 128 micromol/min. mg. The recombinant FBPase can be activated by DTT, Na(+), or low concentrations of Li(+), Ca(2+), Zn(2+), GuHCl, and urea, while it can be inhibited by K(+) or NH(+)(4).

Amino Acid Sequence↗

Modified bacteriophage lambda promoter vectors for overproduction of proteins in Escherichia coli.

A new series of expression vectors that direct high-level overproduction of gene products in Escherichia coli is described. All contain strong bacteriophage lambda promoters, PR and PL, arranged in tandem so that both promote transcription into genes inserted into or between unique restriction sites. The vectors also direct expression of the lambda cI857 gene (from its natural promoter, PM), which enables their use in any E. coli host strain to effect controlled expression by shifting the temperature of cultures from 30 to 42 degrees C. The vectors pCE30, pND201, pPT150 and pMA200U are derivatives of the high-copy-number plasmid pUC9. Vector pCE33 is an analogous derivative of the heat-inducible runaway-replication plasmid, pMOB45, and directs overproduction of proteins by virtue of increase in both gene dosage and transcription following treatment at 42 degrees C. The vectors pND201 and pPT150 bear a ribosome-binding site (RBS) perfectly complementary to the 3' end of E. coli 16-S rRNA a few bp upstream from a unique HpaI site. Ways in which they may be used to improve the efficiency of translation of mRNA by substitution of a natural RBS with selection for optimal spacing from an ATG (or GTG) start codon are described. The phagemid vector pMA200U is a direct analog of pCE30 designed to facilitate preparation of single-stranded DNA templates for use in oligodeoxyribonucleotide-directed mutagenesis of overexpressed genes.

Bacteriophage lambda↗

Translational features of human alpha 2b interferon production in Escherichia coli.

The yield of human alpha 2b interferon in Escherichia coli was optimized by replacement of low-usage arginine codons located in the mRNA 5' end. The differences observed among the various gene variants suggest that codon usage, Shine-Dalgarno-like sequences, and mRNA secondary structure contribute to the performance of E. coli translation machinery.

Amino Acid Sequence↗

Optimization of the expression of equistatin in Pichia pastoris.

To improve the expression of equistatin, a proteinase inhibitor from the sea anemone Actinia equina, in the yeast Pichia pastoris, we prepared gene variants with yeast-preferred codon usage and lower repetitive AT and GC content. The full gene optimization approximately doubled the level of steady-state mRNA and protein accumulated in the culture medium. The removal of a short stretch of 12 additional nucleotides from the multiple cloning site (MCS) sequence in the vector pPIC9 had an enhancement effect similar to full gene optimization (factor 1.5) at the mRNA level. However, at the protein level, this increase was 4- to 10-fold. The optimized gene without the MCS sequence yielded 1.66 g/L active protein in a bioreactor and was purified by a new two-step procedure with a recovery of activity that was >95%. This production level constitutes an overall improvement of about 20-fold relative to our previously published results. The characteristics of the MCS sequence element are discussed in the light of its apparent ability to act as negative expression regulator.

Amino Acid Sequence↗

A hardware interpretation of the evolution of the genetic code.

A quantitative rationale for the evolution of the genetic code is developed considering the principle of minimal hardware. This principle defines an optimal code as one that minimizes for a given amount of information encoded, the product of the number of physical devices used by the average complexity of each device. By identifying the number of different amino acids, number of nucleotide positions per codon and number of base types that can occupy each such position with, respectively, the amount of information, number of devices and the complexity, we show that optimal codes occur for 3, 7 and 20 amino acids with codons having a single, two and three base positions per codon, respectively. The advantage of a code of exactly 4 symbols is deduced, as well as a plausible evolutionary pathway from a code of doublets to triplets. The present day code of 20 amino acids encoded by 64 codons is shown to be the most optimal in an absolute sense. Using a tetraplet code further evolution to a code in which there would be 55 amino acids is in principle possible, but such a code would deviate slightly more than the present day code from the minimal hardware configuration. The change from a triplet code to a tetraplet code would occur at about 32 amino acids. Our conclusions are independent of, but consistent with, the observed physico-chemical properties of the amino acids and codon structures. These correlations could have evolved within the constrains imposed by the minimal hardware principle.

Amino Acids↗