PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Bacterial start site prediction.

With the growing number of completely sequenced bacterial genes, accurate gene prediction in bacterial genomes remains an important problem. Although the existing tools predict genes in bacterial genomes with high overall accuracy, their ability to pinpoint the translation start site remains unsatisfactory. In this paper, we present a novel approach to bacterial start site prediction that takes into account multiple features of a potential start site, viz., ribosome binding site (RBS) binding energy, distance of the RBS from the start codon, distance from the beginning of the maximal ORF to the start codon, the start codon itself and the coding/non-coding potential around the start site. Mixed integer programing was used to optimize the discriminatory system. The accuracy of this approach is up to 90%, compared to 70%, using the most common tools in fully automated mode (that is, without expert human post-processing of results). The approach is evaluated using Bacillus subtilis, Escherichia coli and Pyrococcus furiosus. These three genomes cover a broad spectrum of bacterial genomes, since B.subtilis is a Gram-positive bacterium, E.coli is a Gram-negative bacterium and P. furiosus is an archaebacterium. A significant problem is generating a set of 'true' start sites for algorithm training, in the absence of experimental work. We found that sequence conservation between P. furiosus and the related Pyrococcus horikoshii clearly delimited the gene start in many cases, providing a sufficient training set.

Algorithms↗

Maximum likelihood estimation on large phylogenies and analysis of adaptive evolution in human influenza virus A.

Algorithmic details to obtain maximum likelihood estimates of parameters on a large phylogeny are discussed. On a large tree, an efficient approach is to optimize branch lengths one at a time while updating parameters in the substitution model simultaneously. Codon substitution models that allow for variable nonsynonymous/synonymous rate ratios (omega = d(N)/d(S)) among sites are used to analyze a data set of human influenza virus type A hemagglutinin (HA) genes. The data set has 349 sequences. Methods for obtaining approximate estimates of branch lengths for codon models are explored, and the estimates are used to test for positive selection and to identify sites under selection. Compared with results obtained from the exact method estimating all parameters by maximum likelihood, the approximate methods produced reliable results. The analysis identified a number of sites in the viral gene under diversifying Darwinian selection and demonstrated the importance of including many sequences in the data in detecting positive selection at individual sites.

Algorithms↗

PCR-based gene synthesis as an efficient approach for expression of the A+T-rich malaria genome.

The A+T-rich genome of the human malaria parasite Plasmodium falciparum encodes genes of biological importance that cannot be expressed efficiently in heterologous eukaryotic systems, owing to an extremely biased codon usage and the presence of numerous cryptic polyadenylation sites. In this work we have optimized an assembly polymerase chain reaction (PCR) method for the fast and extremely accurate synthesis of a 2.1 kb Plasmodium falciparum gene (pfsub-1) encoding a subtilisin-like protease. A total of 104 oligonucleotides, designed with the aid of dedicated computer software, were assembled in a single-step PCR. The assembly was then further amplified by PCR to produce a synthetic gene which has been cloned and successfully expressed in both Pichia pastoris and recombinant baculovirus-infected High Five(TM) cells. We believe this strategy to be of special interest as it is simple, accessible and has no limitation with respect to the size of the gene to be synthesized. Used as a systematic approach for the malarial genome or any other A + T-rich organism, the method allows the rapid synthesis of a nucleotide sequence optimized for expression in the system of choice and production of sufficiently large amounts of biological material for complete molecular and structural characterization.

Amino Acid Sequence↗

Enhanced heterologous expression of two Streptomyces griseolus cytochrome P450s and Streptomyces coelicolor ferredoxin reductase as potentially efficient hydroxylation catalysts.

The herbicide-inducible, soluble cytochrome P450s CYP105A1 and CYP105B1 and their adjacent ferredoxins, Fd1 and Fd2, of Streptomyces griseolus were expressed in Escherichia coli to high levels. Conditions for high-level expression of active enzyme able to catalyze hydroxylation have been developed. Analysis of the expression levels of the P450 proteins in several different E. coli expression hosts identified E. coli BL21 Star(DE3)pLysS as the optimal host cell to express CYP105B1 as judged by CO difference spectra. Examination of the codons used in the CYP1051A1 sequence indicated that it contains a number of codons corresponding to rare E. coli tRNA species. The level of its expression was improved in the modified forms of E. coli BL21(DE3), which contain extra copies of rare codon E. coli tRNA genes. The activity of correctly folded cytochrome P450s was further enhanced by cloning a ferredoxin reductase from Streptomyces coelicolor downstream of CYP105A1 and CYP105B1 and their adjacent ferredoxins. Expression of CYP105A1 and CYP105B1 was also achieved in Streptomyces lividans 1326 by cloning the P450 genes and their ferredoxins into the expression vector pBW160. S. lividans 1326 cells containing CYP105A1 or CYP105B1 were able efficiently to dealkylate 7-ethoxycoumarin.

Bacterial Proteins↗

Cardiac troponin I sense-antisense RNA duplexes in the myocardium.

Natural antisense RNA is now thought to regulate, at least in part, a growing number of eukaryotic genes. It is becoming increasingly apparent that such endogenous antisense RNA molecules may modulate gene expression in a manner analogous to synthetic oligomers. Here, we report the detection of antisense-orientated RNA transcripts of cardiac specific troponin I in rat and human myocardium. Interestingly, the different sizes of the rat and human antisense cTNI transcripts suggest species-specific reverse transcription initiation sites. Moreover, for the first time in cardiomyocytes, we could demonstrate in vivo duplex formation between sense and antisense transcripts. The existence of antisense-sense duplexes represents compelling evidence and a potential mechanism for endogenous antisense transcript-mediated modulation of mRNA translation. The potential effect of attenuating translation was illustrated by in vitro and in vivo model systems. Testing several oligonucleotides based on the natural antisense sequences, the optimal region for inhibition of translation was identified as being close to the translational start codon.

Adult↗

Possibility of genetic coding of amino acid sequences by coherent electronic states in nucleotide chains.

The concept of coherent electronic states and coherent interactions in supramolecular structures is applied to the process of genetic information coding and its transcription from DNA to mRNA. A new genetic code is proposed based on the assumption of coherent electron states in linear chains of nucleotide bases. A new interpretation of codon equivalency (redundancy) is given. The number of existing amino acids is derived from the optimalization principle applied to the physical system storing the genetic information in the new code. The proposed code uses a variable number of positions or nucleotide bases along the DNA-mRNA structure to code a single amino acid in a protein. The average of this variable number must be equal to the base of natural logarithms (e = 2.7 . . .) in order to minimize the number of nucleotides required to code a sequence of amino acids.

Amino Acid Sequence↗

Inhibition of influenza virus replication in cultured cells by RNA-cleaving DNA enzyme.

Influenza virus replication has been effectively inhibited by antisense phosphothioate oligonucleotides targeting the AUG initiation codon of PB2 mRNA. We designed RNA-cleaving DNA enzymes from 10-23 catalytic motif to target PB2-AUG initiation codon and measured their RNA-cleaving activity in vitro. Although the RNA-cleaving activity was not optimal under physiological conditions, DNA enzymes inhibited viral replication in cultured cells more effectively than antisense phosphothioate oligonucleotides. Our data indicated that DNA enzymes could be useful for the control of viral infection.

Animals↗

Protein evolution drives the evolution of the genetic code and vice versa.

A model for the developmental pathway of the genetic code, grounded on group theory and the thermodynamics of codon-anticodon interaction is presented. At variance with previous models, it takes into account not only the optimization with respect to amino acid attributes but, also physicochemical constraints and initial conditions. A 'simple-first' rule is introduced after ranking the amino acids with respect to two current measures of chemical complexity. It is shown that a primeval code of only seven amino acids is enough to build functional proteins. It is assumed that these proteins drive the further expansion of the code. The proposed primeval code is compared with surrogate codes randomly generated and with another proposal for primeval code found in the literature. The departures from the 'universal' code, observed in many organisms and cellular compartments, fit naturally in the proposed evolutionary scheme. A strong correlation is found between, on one side, the two classes of aminoacyl-tRNA synthetases, and on the other, the amino acids grouped by end-atom-type and by codon type. An inverse of Davydov's rules, to associate the amino acid end atoms (O/N and non-O/non-N) of 18 amino acids with codons containing a weak base (A/U), extended to the 20 amino acids, is derived.

Amino Acid Sequence↗

The phylogenetic utility of the codon-degeneracy model.

The codon-degeneracy model (CDM) predicts relative frequencies of substitution for any set of homologous protein-coding DNA sequences based on patterns of nucleotide degeneracy, codon composition, and the assumption of selective neutrality. However, at present, the CDM is reliant on outside estimates of transition bias. A new method by which the power of the CDM can be used to find a synonymous transition bias that is optimal for any given phylogenetic tree topology is presented. An example is illustrated that utilizes optimized transition biases to generate CDM GF-scores for every possible phylogenetic tree for pocket gophers of the genus Orthogeomys. The resulting distribution of CDM GF-scores is compared and contrasted with the results of maximum parsimony and maximum likelihood methods. Although convergence on a single tree topology by the CDM and another method indicates greater support for that particular tree, the value of CDM GF-score as the sole optimality criterion for phylogeny reconstruction remains to be determined. It is clear, however, that the a priori estimation of an optimum transition bias from codon composition has a direct application to differentiating between alternative trees.

Animals↗

A plasmid system for optimization of Fab' production in Escherichia coli: importance of balance of heavy chain and light chain synthesis.

We demonstrate the importance of optimizing the balance of light chain (LC) and heavy chain (HC) expression to achieve high level production of Fab' fragments in the Escherichia coli periplasm. The LC:HC balance has been controlled by varying the codon usage of the signal peptide (SP) and 5' mature domain coding regions. Different SP coding regions have been identified from a codon wobble-based library using alkaline phosphatase (AP) as a reporter gene. A plasmid system that enables random combination of these variant SP coding regions is used to construct optimized Fab' expression plasmids. These small plasmid libraries facilitated selection of optimal Fab' expression plasmids and resulted in increases of periplasmic yield, up to 580 mgL(-1) from E. coli fermentations and will enable rapid variable region subcloning and selection of future Fab(') expression plasmids.

Base Sequence↗

Effects of codon usage versus putative 5'-mRNA structure on the expression of Fusarium solani cutinase in the Escherichia coli cytoplasm.

Matching the codon usage of recombinant genes to that of the expression host is a common strategy for increasing the expression of heterologous proteins in bacteria. However, while developing a cytoplasmic expression system for Fusarium solani cutinase in Escherichia coli, we found that altering codons to those preferred by E. coli led to significantly lower expression compared to the wild-type fungal gene, despite the presence of several rare E. coli codons in the fungal sequence. On the other hand, expression in the E. coli periplasm using a bacterial PhoA leader sequence resulted in high levels of expression for both the E. coli optimized and wild-type constructs. Sequence swapping experiments as well as calculations of predicted mRNA secondary structure provided support for the hypothesis that differential cytoplasmic expression of the E. coli optimized versus wild-type cutinase genes is due to differences in 5(') mRNA secondary structures. In particular, our results indicate that increased stability of 5(') mRNA secondary structures in the E. coli optimized transcript prevents efficient translation initiation in the absence of the phoA leader sequence. These results underscore the idea that potential 5(') mRNA secondary structures should be considered along with codon usage when designing a synthetic gene for high level expression in E. coli.

Amino Acid Sequence↗

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing↗

Conformational preferences of the base substituent in hypermodified nucleotide queuosine 5'-monophosphate 'pQ' and protonated variant 'pQH+'.

Conformational preferences of the base substituent in hypermodified nucleotide queuosine 5'-monophosphate 'pQ' and its protonated form 'pQH+' have been studied using quantum chemical Perturbative Configuration Interaction with Localized Orbitals PCILO method. The salient points have also been examined using molecular mechanics force field MMFF, parameterized modified neglect of differential overlap PM3 and Hartree Fock-Density Functional Theory HF DFT (pBP/DN*) approaches. Aqueous solvation of pQ and pQH+ has also been studied using molecular dynamics simulations. Consistent with the observed crystal structure, in isolated protonated form pQH+, the quaternary amine HN(13)(+)H, of the sidechain having 7-aminomethyl linkage, hydrogen bonds with the carbonyl oxygen O(10) of the base. However, N(13)H-O(10) hydrogen bonding is not preferred for unprotonated pQ, whether isolated or hydrated. Interaction between the 5'-phosphate and the 7-aminomethyl group is more likely for isolated pQ. The cyclopentenediol hydroxyl group O4"H may hydrogen bond with the O(10) in isolated pQ as well as in pQH+. The O4"H may hydrogen bond with the 5'-phosphate as well. The presence of -CH2-NH- and O"H groups in pQ and pQH+ allows interesting possibilities for intranucleotide hydrogen bonds and interactions across the anticodon loop. Simultaneous hydrogen bonds O2P-HN(13)+H-O(10) are indicated for hydrated pQH+. Unlike weak involvement of O4"H, these interactions also persist in hydrated pQH+ and may much reduce backbone flexibility. Resulting sub-optimal Q:C base pairing leads to unbiased reading of U or C as the third codon letter. Cyclopentenediol hydroxyl groups may interact with other biomolecules, allowing specific recognition. Prospective pQ(34) and pQ(34)H+ sites for codon-anticodon base pairing remain unhindered, but non canonical Q:G base pairing (amber-suppression) is ruled out.

Anticodon↗

cDNA cloning and functional expression in yeast Saccharomyces cerevisiae of beta-naphthoflavone-induced rabbit liver P-450 LM4 and LM6.

A cDNA library was constructed from liver mRNA of a beta-naphthoflavone-induced rabbit. Two clones pLM4-1 and pLM6-1 containing 2.2-kbp inserts that hybridized at low stringincy with a mouse P1 P-450 probe were selected. The clone pLM4-1 was fully sequenced and found to contain a full-length cDNA coding for cytochrome P-450 LM4. Partial sequence and restriction mapping made it possible to identify pLM6-1 as coding for the major part of cytochrome P-450 LM6. Cloned LM4-1 cDNA was reformed by deletion of the 5' and 3' non-coding regions before insertion into yeast expression vectors PYe DP1/10. A similar operation was performed on pLM6-1 cDNA after replacement of the missing N-terminus-coding sequences by homologous sequences form the pLM4-1 clone resulting in a chimeric cytochrome P-450 coding sequence. Expression of cloned rabbit cytochrome P-450 into transformed yeast was optimized by studying the effect of the nature of the DNA sequence just preceding the initiation codon on the level of cytochrome P-450 production. Yeast synthesized cytochromes P-450 were characterized by immunoblotting, spectra and catalytic activity determinations. Cloned cytochrome P-450 LM4 was found by all criteria to be identical to the authentic rabbit one. The chimeric cytochrome P-450 that contains the 143 N-terminal amino acids of cytochrome P-450 LM4 and the remaining 375 amino acids of cytochrome P-450 LM6 was found to exhibit most of the authentic cytochrome P-450 LM6 catalytic properties. Enzymatic and evolutionary implications of these results are discussed.

Amino Acid Sequence↗

Pullulanase type I from Fervidobacterium pennavorans Ven5: cloning, sequencing, and expression of the gene and biochemical characterization of the recombinant enzyme.

The gene encoding the type I pullulanase from the extremely thermophilic anaerobic bacterium Fervidobacterium pennavorans Ven5 was cloned and sequenced in Escherichia coli. The pulA gene from F. pennavorans Ven5 had 50.1% pairwise amino acid identity with pulA from the anaerobic hyperthermophile Thermotoga maritima and contained the four regions conserved among all amylolytic enzymes. The pullulanase gene (pulA) encodes a protein of 849 amino acids with a 28-residue signal peptide. The pulA gene was subcloned without its signal sequence and overexpressed in E. coli under the control of the trc promoter. This clone, E. coli FD748, produced two proteins (93 and 83 kDa) with pullulanase activity. A second start site, identified 118 amino acids downstream from the ATG start site, with a Shine-Dalgarno-like sequence (GGAGG) and TTG translation initiation codon was mutated to produce only the 93-kDa protein. The recombinant purified pullulanases (rPulAs) were optimally active at pH 6 and 80 degrees C and had a half-life of 2 h at 80 degrees C. The rPulAs hydrolyzed alpha-1,6 glycosidic linkages of pullulan, starch, amylopectin, glycogen, alpha-beta-limited dextrin. Interestingly, amylose, which contains only alpha-1,4 glycosidic linkages, was not hydrolyzed by rPulAs. According to these results, the enzyme is classified as a debranching enzyme, pullulanase type I. The extraordinary high substrate specificity of rPulA together with its thermal stability makes this enzyme a good candidate for biotechnological applications in the starch-processing industry.

Amino Acid Sequence↗

Mutations affecting translational coupling between the rep genes of an IncB miniplasmid.

The nature of translational coupling between repB and repA, the overlapping rep genes of the IncB plasmid pMU720, was examined. Mutations in the start codon of the promoter proximal gene, repB, reduced the efficiency of translation of both rep genes. Moreover, there was no independent initiation of repA translation in the absence of repB translation. The position of the repB stop codon was crucial for the efficient expression of repA, with the wild-type positioning being optimal. Translational coupling was found to be totally dependent on the formation of a pseudoknot structure. A model which invokes formation of a pseudoknot to facilitate initiation of repA is proposed.

Bacterial Proteins↗

Characterization of an alpha 1----3-galactosyltransferase homologue on human chromosome 12 that is organized as a processed pseudogene.

UDP-Gal:Gal beta 1----4GlcNAc alpha 1----3-galactosyltransferase is a terminal glycosyltransferase that is widely expressed in a variety of mammalian species, with the notable exception of man, apes, and Old World monkeys. We recently reported the isolation of a bovine cDNA clone that contains the complete coding sequence for this enzyme (Joziasse, D. H., Shaper, J. H., Van den Eijnden, D. H., Van Tunen, A. J., and Shaper, N. L. (1989) J. Biol. Chem. 264, 14290-14297). Using this cDNA as a probe, we have demonstrated that, although transcripts cannot be detected in a variety of established human cell lines by Northern blot analysis, homologous sequences are present in human genomic DNA. To establish that these sequences represent a human homologue of alpha 1----3-galactosyltransferase, we have used the bovine cDNA as a probe to isolate two nonoverlapping clones (HGT-2 and HGT-10) from a human genomic DNA library. Clone HGT-2 contains a 1.5-kilobase uninterrupted linear sequence similar to bovine alpha 1----3-galactosyltransferase that is organized as a processed pseudogene. This sequence, flanked by Alu type repeats, contains a short 5'- and 3'-untranslated region and a complete recognizable coding region that is 81% similar at the nucleotide level to bovine alpha 1----3-galactosyltransferase. This putative coding region contains multiple frameshift mutations and nonsense codons in all three reading frames which precludes the synthesis of a functional enzyme. Nevertheless, after optimal alignment, translation predicts a polypeptide that is 68% similar at the amino acid level to the bovine enzyme. Based on Southern analysis and limited sequence analysis, clone HGT-10 contains coding sequences similar to the NH2-terminal region of bovine alpha 1----3-galactosyltransferase. By analysis of panels of human-rodent somatic cell hybrids we have established that the nonfunctional, processed pseudogene and the human homologue represented by HGT-10 are located on human chromosomes 12 and 9, respectively. Interestingly, a comparison of the predicted amino acid sequence of the carboxyl-terminal two-thirds of human alpha 1----3-galactosyltransferase, with the corresponding region of the human blood group A, UDP-GalNAc:[Fuc alpha 1----2]Gal beta 1----4GlcNAc alpha 1----3-GalNAc-transferase (Yamamoto, F., Marken, J., Tsuji, T., White, T., Clausen, H., and Hakomori, S. (1990a) J. Biol. Chem. 265, 1146-1151), reveals a significant similarity (39%) suggesting that these two enzymes may have arisen from the same ancestral gene as a result of gene duplication and subsequent divergence.

Amino Acid Sequence↗

mRNA sequences influencing translation and the selection of AUG initiator codons in the yeast Saccharomyces cerevisiae.

The secondary structure and sequences influencing the expression and selection of the AUG initiator codon in the yeast Saccharomyces cerevisiae were investigated with two fused genes, which were composed of either the CYC7 or CYC1 leader regions, respectively, linked to the lacZ coding region. In addition, the strains contained the upf1-delta disruption, which stabilized mRNAs that had premature termination codons, resulting in wild-type levels. The following major conclusions were reached by measuring beta-galactosidase activities in yeast strains having integrated single copies of the fused genes with various alterations in the 89 and 38 nucleotide-long untranslated CYC7 and CYC1 leader regions, respectively. The leader region adjacent to the AUG initiator codon was dispensable, but the nucleotide preceding the AUG initiator at position -3 modified the efficiency of translation by less than twofold, exhibiting an order of preference A > G > C > U. Upstream out-of-frame AUG triplets diminished initiation at the normal site, from essentially complete inhibition to approximately 50% inhibition, depending on the position of the upstream AUG triplet and on the context (-3 position nucleotides) of the two AUG triplets. In this regard, complete inhibition occurred when the upstream and downstream AUG triplets were closer together, and when the upstream and downstream AUG triplets had, respectively, optimal and suboptimal contexts. Thus, leaky scanning occurs in yeast, similar to its occurrence in higher eukaryotes. In contrast, termination codons between two AUG triplets causes reinitiation at the downstream AUG in higher eukaryotes, but not generally in yeast. Our results and the results of others with GCN4 mRNA and its derivatives indicate that reinitiation is not a general phenomenon in yeast, and that special sequences are required.

Base Sequence↗