PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

[Context organization of mRNA 5'-untranslated regions of higher plants].

Computer analysis of nucleotide sequences of 5'-untranslated regions (5'-UTR) of higher plants mRNA adopted from the EMBL nucleotide sequence databank was carried out. It was demonstrated that the average nucleotide frequencies of the leader sequences and adjacent regions of basal promoters are similar, whereas introns and 3'-UTR have a higher content of T and a lower content of C. A particular 5'-UTR contextual feature is a misbalance in the content of complementary nucleotides; probably a stable secondary structure adversely affects the translation properties of the leader sequence. About 20% of 5'-UTR contain AUG triplets, which is twice the earlier estimate. Considered are the properties of leader open reading frames (uORF), the possible causes of their high content in 5'-UTRs of eukaryotic mTNAs, and correlations between the features of uORFs and of the protein-encoding sequence of the gene. It is demonstrated that in effectively translated mRNAs the leader AUG triplets are more frequently located in a nonoptimal context, whereas terminating codons of uORFs more frequently exist in the optimal one. A hypothesis is put forward that the efficiency of termination at the uORF stop codon might substantially interfere with the mRNA translation activity.

5' Untranslated Regions↗

[Rare initiation codons are regulators of expression of the rpoC gene].

Translation of the rpoC genes in Escherichia coli and Salmonella typhimurium is known to start from the GUG codon. Now, using toeprint analysis we have shown UUG to be the initiation codon of the Pseudomonas putida rpoC gene. IF3 does not seem to proofread initiation at the UUG codon. The rpoC genes of P. putida, E. coli, and S. typhimurium, which use rare start codons, have strong SD-domains AGGAGG (P. p.) and GGGAG (E. c., S. t.), optimal seven-nucleotide spacing between SD and start codons, and good second codon AAA. We suggest that rpoC presents an infrequent case of the regulation of translation initiation by selecting the start codon.

Base Sequence↗

Ribosome-initiator tRNA complex as an intermediate in translation initiation in Escherichia coli revealed by use of mutant initiator tRNAs and specialized ribosomes.

For functional studies of mutant Escherichia coli initiator tRNAs in vivo, we previously described a strategy based on the use of tRNA genes carrying an anticodon sequence change from CAU to CUA along with a mutant chloramphenicol acetyltransferase (CAT) gene carrying an initiation codon change from AUG to UAG. Surprisingly, under conditions where the mutant initiator tRNA is optimally active, the CAT gene with the UAG initiation codon produced more CAT protein (3- to 9-fold more depending on the conditions) than the wild-type CAT gene. Here we show that two new mutant CAT genes having GUC and AUC initiation codons also produce more of the CAT protein in the presence of the corresponding mutant initiator tRNAs. These results are most easily understood if assembly of the 30S ribosome-initiator tRNA-mRNA initiation complex in vivo proceeds with the 30S ribosome binding first to the initiator tRNA and then to the mRNA. In cells overproducing the mutant initiator tRNAs, most ribosomes would carry the mutant initiator tRNA and these ribosomes would select the mutant CAT mRNA over the other mRNAs.

Acylation↗

Genetic plasticity of V genes under somatic hypermutation: statistical analyses using a new resampling-based methodology.

Evidence for somatic hypermutation of immunoglobulin genes has been observed in all of the species in which immunoglobulins have been found. Previous studies have suggested that codon usage in immunoglobulin variable (V) region genes is such that the sequence-specificity of somatic hypermutation results in greater mutability in complementarity-determining regions of the gene than in the framework regions. We have developed a new resampling-based methodology to explore genetic plasticity in individual V genes and in V gene families in a statistically meaningful way. We determine what factors contribute to this mutability difference and characterize the strength of selection for this effect. We find that although the codon usage in immunoglobulin V genes renders them distinct among translationally equivalent sequences with random codon usage, they are nevertheless not optimal in this regard. We find that the mutability patterns in a number of species are similar to those we find for human sequences. Interestingly, sheep sequences show extremely strong mutability differences, consistent with the role of somatic hypermutation in the diversification of primary antibody repertoire in these animals. Human TCR V(beta) sequences resemble immunoglobulin in mutability pattern, suggesting one of several alternatives, that hypermutation is functionally operating in TCR, that it was once operating in TCR or in the common precursor of TCR and immunoglobulin, or that the hypermutation mechanism has evolved to exploit the codon usage in immunoglobulin (and fortuitously, TCR) rather than vice-versa. Our findings provide support to the hypothesis that somatic hypermutation appeared very early in the phylogeny of immune systems, that it is, to a large extent, shared between species, and that it makes an essential contribution to the generation of the antibody repertoire.

Base Sequence↗

High-level expression of the neutralizing epitope of porcine epidemic diarrhea virus by a tobacco mosaic virus-based vector.

Porcine epidemic diarrhea virus (PEDV) causes acute enteritis in pigs of all ages and is often fatal for neonates. A tobacco mosaic virus (TMV)-based vector was utilized for the expression of a core neutralizing epitope of PEDV (COE) for the development of a plant-based vaccine. In this study, the coding sequence of a COE gene was optimized based on the modification of codon usage in tobacco plant genes and the removal of mRNA-destabilizing sequences. The native and synthetic COE genes were cloned into TMV-based vectors and expressed in tobacco plants. The recombinant COE protein constituted up to 5.0% of the total soluble protein in the leaves of tobacco plants infected with the TMV-based vector containing synthetic COE gene, which was approximately 30-fold higher than that in tobacco plants infected with TMV-based vector containing a native COE gene. Therefore, this result indicates that the plant viral expression system with a synthetic gene optimized for plant expression is suitable to produce a large amount of antigen for the development of plant-based vaccine rapidly.

Amino Acid Sequence↗

Codon evolution and conservation of the reading phase in genetic code translation.

The description of the optimized evolution of a code based on 4 nucleotides involves a sequential transition of codons, formed firstly by monomers evolving to dimers and then to triplets, in accordance with the progressive increase of the number of amino acids to be coded. The successive increase in the size of these codons during evolution implies changes in the phase reading of the genetic message, which could become chaotic. In order to overcome this constraint, this paper proposes a codon evolution where two things occur simultaneously: codons change in size and there is an alternation of the molecule which holds the information. For example, the nucleotides of the original oligonucleotide are read as monomers when they are translated to an oligopeptide, but further on, this oligopeptide which is read as amino acid dimers, is translated to a nucleotide form (oligonucleotide). Finally, amino acids conforming a peptide are translated from this oligonucleotide, through a reading of triplets. Although plausible, this evolution is a low-probability process due to the fact that it requires a singular sequence of the oligonucleotide and oligopeptide involved. An alternative hypothesis of evolution is also discussed. It proposes that with the exclusion of the establishment of monomer and dimer codons, there is a direct generation of a code of trinucleotides which arises only when a certain number of amino acids has already been generated. Both hypotheses are discussed in terms of the development of a code in which an optimized hardware is maintained through out its evolution.

Amino Acids↗

Molecular evolution before the origin of species.

Amino acids at conserved sites in the residue sequence of 10 ancient proteins, from 844 phylogenetically diverse sources, were used to specify their time of origin in the interval before species divergence from the last common ancestor (LCA). The order of amino acid addition to the genetic code, based on biosynthesis path length and other molecular evidence, provided a reference for evaluating the 'code age' of each residue profile examined. Significantly earlier estimates were obtained for conserved amino acid residues in these proteins than non-conserved residues. Evidence from the primary structure of 'fossil' proteins thus corroborated the biosynthetic order of amino acid addition to the code.Low potential ferredoxin (Fdxn) had the earliest residue profile among the proteins in this study. A phylogenetic tree for 82 prokaryote Fdxn sequences was rooted midway between bacteria and archaea branches. LCA Fdxn had a 23-residue antecedent whose residue profile matched mid-expansion phase codon assignments and included an amide residue. It contained a highly acidic N-terminal region and a non-charged C-terminal region, with all four cysteine residues. This small protein apparently anchored a [4Fe-4S] cluster, ligated by C-terminal cysteines, to a positively charged mineral surface, consistent with mediating e(-) transfer in a primordial surface system before cells appeared. Its negatively charged N-terminal 'attachment site' was highly mutable during evolution of ancestral Fdxn for Bacteria and Archaea, consistent with a loss of function after cell formation. An initial glutamate to lysine substitution may link 'attachment site' removal to early post-expansion phase entry of basic amino acids to the code. As proteins evidently anchored non-charged amide residues initially, surface attachment of cofactors and other functional groups emerges as a general function of pre-cell proteins.A phylogenetic tree of 107 proteolipid (PL) helix-1 sequences from H(+)-ATPase of bacteria, archaea and eukaryotes had its root between prokaryote branches. LCA PL h1 residue profile optimally fit a late expansion phase codon array. Sequence repeats in transmembrane PL helices h1 and h2 indicated formation of the archetypal PL hairpin structure involved successive tandem duplications, initiated within the gene for an 11-residue (or 4-residue) hydrophobic peptide. Ancestral PL h1 lacked acidic residues, in a fundamental departure from the prototype pre-cell protein. By this stage, proteins with a hydrophobic domain had evolved. Its non-polar, late expansion phase residue profile point to ancestral PL being a component of an early permeable cell membrane. Other indicators of cell formation about this stage of code evolution include phospholipid biosynthesis path length, FtsZ residue profile, and late entry of basic amino acids into the genetic code. Estimates based on conserved residues in prokaryote cell septation protein, FtsZ, and proteins involved with synthesis, transcription and replication of DNA revealed FtsZ, ribonucleotide reductase, RNA polymerase core subunits and 5'-->3' flap exonuclease, FEN-1, originated soon after cells putatively evolved. While reverse transcriptase and topoisomerase I, Topo I, appeared late in the pre-divergence era, when the genetic code was essentially complete. The transition from RNA genes to a DNA genome seemingly proceeded via formation of a DNA-RNA heteroduplex. These results suggest formation of DNA awaited evolution of a catalyst with a hydrophobic domain, capable of sequestering radical bearing intermediates in its synthesis from ribonucleotide precursors. Late formation of topology altering protein, Topo I, further suggests consolidation of genes into chromosomes followed synthesis of comparatively thermostable DNA strands.

Amino Acid Sequence↗

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals↗

The three in-frame ATG, clustered in the translation initiation sequence of human factor IX gene, are required for an optimal protein production.

Three in-frame potential methionine codons have been identified in human factor IX gene and are clustered at amino acids -46, -41 and -39. In view of initiating a gene therapy approach, human factor IX production has been evaluated after modifications of these first three in-frame translation start sites. To characterize the most efficient translation initiation context, five factor IX cDNA expression vectors directed by CMV promoter-enhancer were generated. These vectors contained different starting site combinations including one, two or three ATG. A quantitative analysis of factor IX production in stably transfected CHO cells and in a rabbit reticulocyte lysate cell free system revealed the ability of all single site to generate fully active factor IX. However, the factor IX production level increased with the ATG number and the wild type (WT) cDNA bearing the 3 ATG induced the highest protein production. A truncated intron I of factor IX, previously suggested of having an expression-augmenting activity, was also placed in the WT factor IX cDNA. In stably transfected CHO cells, a 8-fold increase in protein production was measured. These results show that at least in vitro, the presence of the three ATG seems to be crucial for a maximal factor IX production. The data also suggest that both the three ATG and the truncated intron I are required for an optimal factor IX production in a perspective of a human gene therapy of haemophilia.

Animals↗

Identification of mutations in the connexin 26 gene that cause autosomal recessive nonsyndromic hearing loss.

Mutations in the Cx26 gene have been shown to cause autosomal recessive nonsyndromic hearing loss (ARNSHL) at the DFNB1 locus on chromosome 13q12. Using direct sequencing, we screened the Cx26 coding region of affected and nonaffected members from seven ARNSHL families either linked to the DFNB1 locus or in which the ARNSHL phenotype cosegregated with markers from chromosome 13q12. Cx26 mutations were found in six of the seven families and included two previously described mutations (W24X and W77X) and two novel Cx26 mutations: a single base pair deletion of nucleotide 35 resulting in a frameshift and a C-to-T substitution at nucleotide 370 resulting in a premature stop codon (Q124X). We have developed and optimized allele-specific PCR primers for each of the four mutations to rapidly determine carrier and noncarrier status within families. We also have developed a single stranded conformational polymorphism (SSCP) assay which covers the entire Cx26 coding region. This assay can be used to screen individuals with nonsyndromic hearing loss for mutations in the CX26 gene.

Alleles↗

Efficient gene expression in mammalian cells from a dicistronic transcriptional unit in an improved retroviral vector.

We have studied the properties of dicistronic transcriptional units in retroviral vectors. In these vectors, the promoter in the 5' retroviral long terminal repeat (LTR) controls expression of both an upstream cistron (luc) encoding firefly luciferase and a downstream cistron (neo), a selectable marker encoding neomycin phosphotransferase (NPTII). By assaying for simultaneous expression of luc and neo after transfection or infection of hamster BHK, rat 208F, and mouse retroviral packaging cell lines, we have identified important factors that affect expression from the downstream cistron, including the presence of intercistronic ATG sequences, the length of the intercistronic sequence and conformity of the sequence surrounding the downstream start codon to the eukaryotic consensus sequence. Optimized dicistronic vectors produced amounts of NPTII comparable to a vector in which neo was driven by a strong internal promoter consisting of a modified Rous sarcoma virus LTR. Additionally, they produced higher virus titers and demonstrated improved stability of gene expression in the absence of selection. By virtue of their physical compactness and elimination of the need for a separate promoter for every gene, dicistronic transcriptional units allow the introduction of larger genes into retroviral vectors and may allow for more than two genes to be placed in a single vector.

Avian Sarcoma Viruses↗

Expression, purification of human vasostatin120-180 in Escherichia coli, and its anti-angiogenic characterization.

According to codon preference of Escherichia coli, the optimized coding sequence of human vasostatin120-180aa (VAS) was obtained by chemical synthesis and molecular cloning methods. Using PCR and enzyme digestion, the full encoding sequence for VAS was cloned into the E. coli expression vector pALEX and expressed as a GST fusion protein in BL21 (DE3) strain. GST-VAS protein approximately accounted for 45% of the total bacterial proteins. Most of target protein existed in inclusion body. To improve the solubility of GST-VAS, the contribution of low temperature and molecular chaperone co-expression to the solubility of GST-VAS was tested. The results showed that co-expression with chaperons, TF and GroES/GroEL, and low expression temperature cooperatively improved the solubility of GST-VAS from 10 to 85%, and the yield of soluble GST-VAS was sixfold increased. When purified by GST affinity chromatography, 50 mg GST-VAS was obtained with purity over 85% from 1 L culture. Intact VAS was released by enterokinase digestion and further purified by Sephadex G50 gel filtration chromatography. About 7.2 mg intact homogeneous VAS protein was finally produced from 1L bacterial culture. The identity of GST-VAS and VAS was validated by Western blotting analysis. Recombinant VAS protein displayed distinct inhibition of endothelial cell proliferation and anti-angiogenic activity by chick embryo chorioallantoic membrane assay.

Angiogenesis Inhibitors↗

Eukaryotic mRNAs encoding abundant and scarce proteins are statistically dissimilar in many structural features.

It is well known that non-coding mRNA sequences are dissimilar in many structural features. For individual mRNAs correlations were found for some of these features and their translational efficiency. However, no systematic statistical analysis was undertaken to relate protein abundance and structural characteristics of mRNA encoding the given protein. We have demonstrated that structural and contextual features of eukaryotic mRNAs encoding high- and low-abundant proteins differ in the 5' untranslated regions (UTR). Statistically, 5' UTRs of low-expression mRNAs are longer, their guanine plus cytosine content is higher, they have a less optimal context of the translation initiation codons of the main open reading frames and contain more frequently upstream AUG than 5' UTRs of high-expression mRNAs. Apart from the differences in 5' UTRs, high-expression mRNAs contain stronger termination signals. Structural features of low- and high-expression mRNAs are likely to contribute to the yield of their protein products.

5' Untranslated Regions↗

Melibiose permease and alpha-galactosidase of Escherichia coli: identification by selective labeling using a T7 RNA polymerase/promoter expression system.

Identification and selective labeling of the melibiose permease and alpha-galactosidase in Escherichia coli, which are encoded by the melB and melA genes, respectively, have been accomplished by selectively labeling the two gene products with a T7 RNA polymerase expression system [Tabor, S., & Richardson, C. C. (1985) Proc. Natl. Acad. Sci. U.S.A. 82, 1074]. Following generation of a novel EcoRI restriction site in the intergenic sequence between the two genes of the mel operon by oligonucleotide-directed, site-specific mutagenesis, melA and melB were separately inserted into plasmid pT7-6 of the T7 expression system. Expression of melB was markedly enhanced by placing a strong, synthetic ribosome binding site at an optimal distance upstream from the initiation codon of melB. Expression of cloned gene products was characterized functionally and by performing autoradiographic analysis on total cell, inner membrane, and cytoplasmic proteins from cells pulse labeled with (35S)methionine in the presence of rifampicin and resolved by sodium dodecyl sulfate/polyacrylamide gel electrophoresis. The results first confirm that alpha-galactosidase is a cytoplasmic protein with an Mr of 50K; in contrast, the membrane-bound melibiose permease is identified as a protein with an apparent Mr of 39K, a value significantly higher than that of 30K previously suggested [Hanatani et al. (1984) J. Biol. Chem. 259, 1807].

Animals↗

Binding of the bacteriophage T4 regA protein to mRNA targets: an initiator AUG is required.

Bacteriophage T4 regA protein translationally represses the synthesis of a subset of early phage-induced proteins. The protein binds to the translation initiation site of at least two mRNAs and prevents formation of the initiation complex. We show here that the protein binds to the translation initiation sites of other regA-sensitive mRNAs. Analysis of mRNA binding by filtration and nuclease protection assays shows that AUG is necessary but not sufficient for specific binding of regA protein to its mRNA targets. Anticipating the need for large quantities of regA protein for structural studies to further define the regA protein-RNA ligand interaction, we also report cloning the regA gene into a T4 overexpression system. The expression of regA protein in uninfected E. coli is lethal, so in our system regA driven by a strong T7 promoter is sequestered in a T4 phage until 'induction' by phage infection is desired. We have replaced the regA sensitive wild-type ribosome binding site with a strong insensitive ribosome binding site at an optimal distance from the regA initiation codon for maximizing expression. We have obtained large amounts of regA protein.

Base Sequence↗

Complete sequence and comparative analysis of the genome of herpes B virus (Cercopithecine herpesvirus 1) from a rhesus monkey.

The complete DNA sequence of herpes B virus (Cercopithecine herpesvirus 1) strain E2490, isolated from a rhesus macaque, was determined. The total genome length is 156,789 bp, with 74.5% G+C composition and overall genome organization characteristic of alphaherpesviruses. The first and last residues of the genome were defined by sequencing the cloned genomic termini. There were six origins of DNA replication in the genome due to tandem duplication of both oriL and oriS regions. Seventy-four genes were identified, and sequence homology to proteins known in herpes simplex viruses (HSVs) was observed in all cases but one. The degree of amino acid identity between B virus and HSV proteins ranged from 26.6% (US5) to 87.7% (US15). Unexpectedly, B virus lacked a homolog of the HSV gamma(1)34.5 gene, which encodes a neurovirulence factor. Absence of this gene was verified in two low-passage clinical isolates derived from a rhesus macaque and a zoonotically infected human. This finding suggests that B virus most likely utilizes mechanisms distinct from those of HSV to sustain efficient replication in neuronal cells. Despite the considerable differences in G+C content of the macaque and B virus genes (51% and 74.2%, respectively), codons used by B virus are optimal for the tRNA population of macaque cells. Complete sequence of the B virus genome will certainly facilitate identification of the genetic basis and possible molecular mechanisms of enhanced B virus neurovirulence in humans, which results in an 80% mortality rate following zoonotic infection.

Animals↗

Avidin expressed in transgenic rice confers resistance to the stored-product insect pests Tribolium confusum and Sitotroga cerealella.

Rice (Oryza sativa var. Nipponbare) was transformed with an artificial avidin gene. The features of this construct are as follows: (1) a signal peptide sequence derived from barley alpha amylase was added at the N-terminal region, (2) codon usage of the gene was optimized for rice, and (3) the gene was driven by rice glutelin GluB-1, an endosperm-specific promoter. Avidin was produced in the grain of the transgenic rice but not in the leaves. The concentration of avidin in the kernels was about 1,800 ppm. All larvae of the confused flour beetle (Tribolium confusum) and Angoumois grain moth (Sitotroga cerealella) died when fed transgenic avidin rice powder or kernels, respectively, whereas most of the test insects developed into adults when they were fed a nontransgenic rice control diet. Avidin extracted from the transgenic rice kernel lost most biotin-binding activity after 5 min heating at 95 degrees C.

Amino Acid Sequence↗

Immunobead-PCR: a technique for the detection of circulating tumor cells using immunomagnetic beads and the polymerase chain reaction.

The presence of tumor cells in the circulation may predict disease recurrence and metastases. We have developed a sensitive technique for the detection of carcinoma cells in blood, using immunomagnetic beads to enrich for epithelial cells and the polymerase chain reaction to identify a tumor marker. The colon carcinoma cell line SW480, homozygous for a K-ras codon 12 mutation, was used to establish optimal conditions. The SW480 cells were serially diluted in normal blood and incubated with immunomagnetic beads labeled with a monoclonal antibody specific for epithelial cells. Cells bound to the beads were retrieved using a magnetic field and the presence of K-ras codon 12 mutations determined by a polymerase chain reaction based analysis. SW480 cells could be detected in dilutions up to 1 SW480 cell/10(5) leukocytes in whole blood.

Base Sequence↗