PubMed HealthSearch

SEARCH · PubMed Health

Results for “codon optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A computer program for the design of optimal synthetic oligonucleotide probes for protein coding genes.

A computer program has been written in FORTRAN 77 to locate on a protein sequence a region with optimum length and limited degeneracy in order to design artificial oligonucleotide probes for use in molecular cloning. In addition the program checks for regions of homology between this probe and any other base sequence found in nucleotide sequence data banks. There are options in the program to eliminate rare codons or to make preferential choices of bases in order to minimize the degeneracy of probes.

Algorithms

Nucleotide sequence of a cDNA clone encoding the entire glycoprotein from the New Jersey serotype of vesicular stomatitis virus.

The nucleotide sequence of the mRNA encoding the glycoprotein from the New Jersey serotype of vesicular stomatitis virus (VSV) was determined from a cDNA clone containing the entire coding region. The sequence of 12 5'-terminal noncoding nucleotides present in the mRNA but not in the cDNA clone was determined from a primer extended to the 5' terminus of the mRNA. The mRNA is 1,573 nucleotides long (excluding polyadenylic acid) and encodes a protein of 517 amino acids. Only six nucleotides occur between the translation termination codon and the polyadenylic acid. Short homologies between the untranslated termini of this mRNA and the mRNAs of the Indiana serotype were found. The predicted protein sequence was compared with that of the glycoprotein of the Indiana serotype of VSV and with the glycoprotein of rabies virus, using a computer program which determines optimal alignment. An amino acid identity of 50.9% was found for the two VSV serotypes. Approximately 20% identity was found between the rabies virus and VSV New Jersey glycoproteins. The positions and sizes of the transmembrane domains, the signal sequences, and the glycosylation sites are identical in both VSV serotypes. Two of five serine residues which were possible esterification sites for palmitate in the glycoprotein from the Indiana serotype are changed to glycine residues in the glycoprotein from the New Jersey serotype. Because the glycoprotein of the New Jersey serotype does not contain esterified palmitate, we suggest that one or both of these residues are the probable esterification sites in the glycoprotein from the Indiana serotype.

Base Sequence

Mature apolipoprotein AI and its precursor proApoAI: influence of the sequence at the 5' end of the gene on the efficiency of expression in Escherichia coli.

Apolipoprotein AI (ApoAI) plays a central role in the regulation of lipid metabolism. Initial attempts to express human apoAI cDNA in Escherichia coli did not yield detectable levels of the mature protein. By analyzing the efficiency of expression of apoAI-lacZ gene fusions, we have been able to show that the sequence at the 5' end of the ApoAI-coding region is a critical parameter. Indeed, silent changes in the codons for the first 8 residues of ApoAI, which did not alter the amino acid sequence, affected expression dramatically. Analysis of the corresponding mRNA steady-state levels suggested a role for differential mRNA stability in the control of apoAI expression in this system. Among all the possible alternative sequences, we have identified an optimal sequence which, when reinserted in the original expression plasmid, yields high level production of mature ApoAI. This procedure has been extended to the production of the natural variant ApoAI-Milano and the precursor proApoAI. Availability of these recombinant molecules would allow the investigation of their structural and biological features. In addition, the methodology used to optimize ApoAI expression is of general interest in assuring high expression of heterologous proteins in E. coli.

Amino Acid Sequence

Kinetics of translation of gamma B crystallin and its circularly permutated variant in an in vitro cell-free system: possible relations to codon distribution and protein folding.

Analysis of nascent gamma B-crystallin peptides accumulating during in vitro translation in a rabbit reticulocyte lysate cell-free system was carried out. As a consequence of the irregular distribution of rare codons along the polypeptide chain of gamma B-crystallin, translation of the two-domain protein is a non-uniform process characterized by specific pauses. One of the major delays occurs during the translation of the connecting peptide between the domains. Comparing the kinetics of translation of natural gamma B-crystallin and its circularly permutated variant (with the order of the N- and C-terminal domains exchanged) reveals that the natural N-terminal domain is translated faster than the C-terminal one. Since the N-terminal domain in natural gamma B-crystallin is known to be more stable and to fold faster than the C-terminal one [E.-M. Mayr et al. (1994) J. Mol. Biol. 235, 84-88], the present data suggest that the translation rates are optimized to tune the synthesis and folding of the nascent polypeptide chain. In this connection, the pause in the linker region between the domains provides a delay allowing the correct folding of the N-terminal domain and its subsequent assistance in the stabilization of the C-terminal one.

Animals

A generalized information function applied to the genetic code.

The problem of the partitioning of the degeneracy of the codons in the genetic code is considered in the framework of a generalized information function IG = c sigma kpk(ln pk + G(Ek] where k represents the number of codons in a specific degeneracy class and G(Ek) is an arbitrary real valued function. For G(Ek) = 0 the Shannon information function is recovered. For a particular choice of G(Ek) that takes the dominance of even degeneracies into account, it is found by direct numerical calculations that the correct degeneracy partitioning appears as optimal values of the Ig function. This results is also supported by optimization calculations in which the generalized information function is regarded as a continuous function in the degeneracy variables.

Amino Acids

Detection of hepatitis B pre-core mutant by allele specific polymerase chain reaction.

AIM: Development of a specific polymerase chain reaction (PCR) assay for detection of the pre-core, stop codon, mutant of hepatitis B virus (HBV). METHODS: PCR primers, specific at the 3'-end for nucleotide 1896 of either the pre-core, stop codon, mutant or wild type HBV, were synthesised using published sequence data. Positive control templates for both types of virus were synthesised by the PCR, incorporating sequences specific for each virus type at the appropriate position. These templates were used to optimise the specificity of the procedure. Formalin fixed, paraffin wax embedded human tissue from acute or fulminant HBV hepatitis from Hong Kong or Oxford was then investigated for presence of mutant or wild type virus. The HBV DNA was amplified from this tissue using a two step procedure, with an initial amplification phase followed by a second diagnostic phase on optimally diluted target DNA. RESULTS: Specific detection of mutant or wild type HBV was achieved. An important factor in determining specificity was the temperature of annealing, 70 degrees C proving to be highly specific. To overcome the inherent variation of target copy number in clinical samples and to provide an intrinsic positive control, it was important to generate and standardise the amount of target HBV used for the specific PCR. Two cases of fulminant hepatitis and four cases of acute hepatitis from Hong Kong, and one case of fulminant hepatitis from Oxford, contained only wild type HBV, with no evidence of a mutant virus. CONCLUSION: This method can be applied to FFPE tissues. It is rapid, non-radioactive, and specific for the stop codon mutation at nucleotide 1896 of HBV. Preliminary investigation of a small number of cases of fulminant hepatitis from Oxford and Hong Kong showed only wild type virus. The result differs from results published from Japan and Israel.

Alleles

Nucleotide sequence of the Synechococcus sp. PCC7942 branching enzyme gene (glgB): expression in Bacillus subtilis.

The nucleotide sequence of the Synechococcus sp. PCC7942 glgB gene has been determined. The gene contains a single open reading frame (ORF) of 2322 bp encoding a polypeptide of 774 amino acids (aa) with an Mr of 89,206. Extensive sequence similarity exists between the deduced aa sequence of the Synechococcus sp. glgB gene product and that of the Escherichia coli branching enzyme in the middle portions of the proteins (62% identical aa). In contrast, the N-terminal portions shared little homology. The sequenced region which follows glgB contains an ORF encoding 79 aa of the N terminus of a polypeptide that shares extensive sequence similarity (41% identical aa) with human and rat uroporphyrinogen decarboxylase. This suggests that the region downstream from glgB contains the hemE gene and, therefore, that the organization of genes involved in glycogen biosynthesis in Synechococcus sp. is different from that described for E. coli. A fusion gene was constructed between the 5' end of the Bacillus licheniformis penP gene and the Synechococcus sp. glgB gene. The fusion gene was efficiently expressed in the Gram+ micro-organism Bacillus subtilis and specified a branching enzyme with an optimal temperature for activity similar to the wild-type enzyme.

1,4-alpha-Glucan Branching Enzyme

A general approach to isolating Plasmodium falciparum genes using non-redundant oligonucleotides inferred from protein sequences of other organisms.

We have constructed a number of oligonucleotide probes and tested their utility in identifying various genes in Plasmodium falciparum. The probe sequences were based on known conserved regions of proteins from other organisms, coupled with an analysis of the codon usage of the parasite. By using long single oligonucleotides, we have successfully isolated the DHFR-TS gene, two actin genes and two tubulin genes from the K1 (Thailand) isolate of P. falciparum. We compare these single probes to multiply-redundant short oligonucleotide probes and to heterologous probes. We also present a detailed quantitative analysis of optimal probe design, and of how this approach can best be implemented as a general method of isolating plasmodial genes.

Animals

A novel cis element essential for stimulated transcription of the p41 promoter of human herpesvirus 6.

The p41 DNA-binding protein of human herpesvirus 6 is an apparent processivity factor important for viral DNA replication. The p41 promoter was characterized to understand how this processivity factor is regulated. A single transcription start site and a functional TATA box are located 48 and 74 bp, respectively, upstream of the start codon. A reporter construct containing 1,027 bp of the sequence upstream of the p41 start codon was inactive in uninfected T cells but functioned as a strong promoter in human herpesvirus 6-infected cells. Mutational analysis identified a 21-bp element (the EA site) which is located at -73 to -52 bp relative to the transcription start site and is essential for promoter activity. The ability of the EA site to stimulate transcription optimally appears to be strictly dependent upon its distance from the p41 basal promoter. The EA site contains three overlapping sequences, a CAAT-enhancer-binding protein (C/EBP) transcription factor recognition site and two repeat elements. Mobility shift assays using the EA site identified four binding activities (C1 to C4). C1 and C2 are present in both uninfected and infected cells and do not contain C/EBP factors. In infected cells, point mutation of the EA site abrogates C1 and C2 binding activities and destroys transcriptional activity of the p41 promoter. C3 and C4 are present in uninfected cells only and were found to contain C/EBP factors. These findings indicate that in infected cells, transcriptional stimulation of the p41 promoter by the EA site requires C1 and C2 binding activities. These results further suggest that transcriptional activity may also depend upon the elimination of C3 and C4 binding activities.

Amino Acid Sequence

The D-E region of the D1 protein is involved in multiple quinone and herbicide interactions in photosystem II.

The region between helices D and E (D-E region) of the D1 protein of photosystem II (PSII) is exposed at the stromal side of the photosynthetic membrane, contains the secondary plastoquinone (QB) binding niche, and is involved in processes at the reducing side of PSII. The role of the D-E region was studied in 27 site-directed mutants generated in the psbAII gene of the cyanobacterium Synechocystis sp. PCC 6803. The photochemical performance of the modified PSII reaction centers was assessed with respect to photoautotrophic growth, oxygen evolution, fluorescence induction, and herbicide inhibition. A few mutations, located at positions presumably involved in essential interactions in the QB binding niche, greatly interfered with PSII performance. On the other hand, mutations in the presumptive loop region between helices D and de resulted in relatively minor effects, indicating a flexible region not critical for photochemical function. Indeed, although more than 80% of the D-E region is phylogenetically invariant, the bulk of the mutations affected the measured parameters only moderately. The significance of the conserved residues appears to be in subtle interactions that optimize the thermodynamic balance between some of the redox components of PSII, as indicated by mild changes in the steady state fluorescence. Many mutations modified tolerances to PSII herbicides. The dispersion of these mutations throughout the D-E region indicates the complex nature of the interactions, direct and indirect, affecting herbicide binding in the QB niche. Mutation of codons Ser221 and Ser222 to Leu221 and Ala222 revealed a new location coordinating the herbicide diuron in the D1 protein.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Characterization of an Escherichia coli gene encoding betaine aldehyde dehydrogenase (BADH): structural similarity to mammalian ALDHs and a plant BADH.

An open reading frame of 1476 nucleotides, cloned from a region of the Escherichia coli genome encoding betaine biosynthesis functions, was shown to encode a betaine aldehyde dehydrogenase (BADH; EC 1.2.1.8). Either of two adjacent codons (5'-GTGATG) could function as a start codon, producing a presumptive polypeptide of 491 or 490 amino acids. The deduced primary structure of the E. coli BADH showed 39-43% positional identity, over its entire length, to aldehyde dehydrogenases (ALDH: EC 1.2.1.3) of mammalian origin. This similarity increased to 75-77% when conservative aa substitutions were also taken into consideration. Spinach BADH was also similar to the bacterial BADH, showing 38% identity and 80% overall similarity. Other homologs included a fungal and a putative bacterial ALDH. Although E. coli BADH was specific for the substrate, betaine aldehyde, it showed the highest levels of similarity to the prototype human ALDH-2. Only one gap in each sequence had to be introduced for optimal alignment. The conservation between E. coli BADH and the ALDHs was also evident in the predicted secondary structures and hydrophilicity profiles of the polypeptides, suggesting a similarity in the overall folding patterns of ALDH and BADH. These observations suggest a common ancestry for BADH and ALDH, preceding prokaryote-eukaryote divergence.

Aldehyde Dehydrogenase

Renaturation of recombinant human pro-urokinase expressed in Escherichia coli.

A synthetic gene encoding human pro-urokinase (pro-UK) with E. coli-favored codon usage was cloned into plasmid pET-3d and expressed in E. coli BL21(DE3) LysS strain. The expressed products, which accumulated as inactive inclusion bodies, were denatured and renatured in vitro. A broad range of parameters such as pH, protein concentration, denaturant concentration, the use of cosolvent polyethylene glycol and presence of basic or acidic amino acid was examined. At optimal renaturation condition, pro-UK activity of more than 1000I.U was obtained from 1 milliliter cell culture.

Cloning, Molecular

Cloning and sequence determination of a cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase.

A partial cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase (ACMPK) was isolated from a lambda ZAP expression library by immunoselection using monospecific polyclonal antibodies to the enzyme. The sequence of both strands of the cDNA (CMKa) was determined. The deduced amino acid (aa) sequence contained all eleven consensus domains found in serine/threonine protein kinases [Hanks et al., Science 241 (1988) 42-52], as well as a putative calmodulin-binding domain. The cDNA contained an intron, lacked an in-frame start codon, and was not polyadenylated. A full-length copy of CMKa was subsequently isolated from a lambda gt10 library of A. nidulans cDNA using a restriction fragment of the first clone as a probe. It contained an in-frame start codon, an open reading frame (ORF) of 1242 bp and was polyadenylated. The ORF encoded a protein of 414 aa residues with an M(r) of 46,895 and an isoelectric point pI = 6.4. These values are in good agreement with that observed for the native enzyme [Bartelt et al., Proc. Natl. Acad. Sci. USA 85 (1988) 3279-3283]. When aligned to optimize homology, 29% of the predicted aa sequence of ACMPK is identical to that of the alpha-subunit of rat brain calmodulin-dependent protein kinase II. ACMPK shares 40 and 44% identity in aa sequence with YCMK1 and YCMK2, respectively, two Ca2+/calmodulin-dependent protein kinases recently cloned from Saccharomyces cerevisiae [Pausch et al., EMBO J. 10 (1991) 1511-1522]. Results of Southern analysis of restriction digests of genomic DNA indicate that ACMPK is encoded by a single-copy gene.

Amino Acid Sequence

Triphasic concentration effects of gentamicin on activity and misreading in protein synthesis.

Gentamicin is shown to exert a triphasic concentration effect on peptide synthesis in vitro with natural messengers. Low concentrations (up to 2 micron) caused slowing and a decrease in total synthesis, but little misreading (assayed with extracts lacking Glu-tRNA); the inhibition was greater with an initiating system (with phage RNA as messenger) than with pure chain elongation on purified endogenous polysomes of Escherichia coli. Moderate concentrations (up to 100 micron) slowed synthesis less, markedly increased its duration in the noninitiating system, and strongly stimulated misreading; at optimal concentrations total synthesis was even greater than normal. Moreover, with phage RNA these concentrations increased the synthesis of large polypeptides. We conclude that binding of gentamicin to its first site causes inhibition but little misreading; binding to additional site(s) partly reverses the inhibition by first-site binding and markedly stimulates misreading, and the misreading appears to favor "readthrough" of termination codons. In the third phase (greater than 100 micron) synthesis is slowed again but the pattern of misreading does not appear to be altered; this effect need not involve a specific further action on the ribosome.

Bacterial Proteins

The androgen receptor in LNCaP cells contains a mutation in the ligand binding domain which affects steroid binding characteristics and response to antiandrogens.

The human prostate tumor cell line LNCaP contains an abnormal androgen receptor system with broad steroid binding specificity. Progestagens, estradiol and several antiandrogens compete with androgens for binding to the androgen receptor in the cells to a higher extent than in other androgen sensitive systems. Optimal growth of LNCaP cells is observed after addition of the synthetic androgen R1881 (0.1 nM). In addition, estrogens, progestagens and several antiandrogens do not inhibit androgen responsive growth, but have striking growth stimulatory effects and increase EGF receptor level and acid phosphatase secretion. We have found that the androgen receptor in the LNCaP cells contains a single point mutation changing the sense of codon 868 (Thr to Ala) in the ligand binding domain. Expression vectors containing the normal or mutated androgen receptor sequence were transfected into COS or HeLa cells. Androgens, progestagens, estrogens and several antiandrogens bind the mutated androgen receptor protein and activate the expression of an androgen-regulated reporter gene (GRE-tk-CAT), indicating that the mutation directly affects both binding specificity and the induction of gene expression. Interestingly, the antiandrogen casodex showed antiandrogenic properties in growth studies of LNCaP cells and did not induce reporter gene activity in Hela cells transfected with the mutant receptor. The mutated androgen receptor of LNCaP cells is therefore a useful tool in the elucidation of different levels of action of steroids and antisteroids.

Binding Sites

High-throughput method for determination of apolipoprotein E genotypes with use of restriction digestion analysis by microplate array diagonal gel electrophoresis.

Molecular epidemiological research has identified the association of a common apolipoprotein E (apo E) isoform (E4 as opposed to E3), with risk both of coronary artery disease and of Alzheimer dementia. In addition, the role of apo E genotype (usually E2/E2) in Type III hyperlipidemia is well known. However, both for diagnostic and research purposes, apo E genotyping is cumbersome. The preferred approach is electrophoretic sizing of restriction digestion fragments, enabling simultaneous analysis of the two codons (112 and 158) that represent the six common genotypes (E2/E2; E2/E3; E2/E4; E3/E3; E3/E4; E4/E4). However, the consequent demands of high-yield PCR, high-resolution, high-throughput electrophoresis, and sufficient detection sensitivity have left shortfalls in published protocols. In conjunction with a high-throughput electrophoresis system we described recently, microplate array diagonal gel electrophoresis (MADGE), we have constructed extensively optimized, simplified protocols for DNA isolation from mouthwash samples for PCR setup and high-yield PCR, for restriction digestion, and for subsequent MADGE gel image analysis. The integral system enables one worker to readily undertake apo E genotyping of as many as hundreds of DNA samples per day, without special equipment.

Apolipoproteins E

Cloning and expression of the branching enzyme gene (glgB) from the cyanobacterium Synechococcus sp. PCC7942 in Escherichia coli.

Using the glgB gene from Escherichia coli as a hybridization probe, the gene encoding the branching enzyme of the cyanobacterium Synechococcus sp. PCC7942 has been identified on a 3.9-kb PstI fragment which was cloned into plasmid pUC9. Two types of plasmids have been isolated. Plasmid pKVN1 was expressing the Synechococcus sp. gene as was shown by complementation of the glgB mutation of E. coli KV832. Plasmid pKVN2, which carried the same insert in the opposite orientation was unable to complement E. coli KV832, indicating that the promoter of the cloned gene was either absent or was not recognized in E. coli. Determination of branching activity in extracts of Synechococcus sp. and E. coli KV832[pKVN1] showed that the enzyme was optimally active at approximately 35 degrees C. No significant activity was present at temperatures higher than 55 degrees C, reflecting the mesophilic nature of the cloned enzyme. In a cell-free coupled transcription-translation system the cloned gene specified two proteins of 84 kDa and 72 kDa, respectively, which are probably translated independently from the same gene by initiation at two different start codons.

1,4-alpha-Glucan Branching Enzyme

A standardized vector system for manipulation and enhanced expression of genes in Escherichia coli.

Different families of cloning and expression vectors were engineered on a standard plasmid. They contain several regulatory signals for transcription and/or translation initiation and termination. The plasmids in each series differ only in the number, type, and order of unique restriction cleavage sites clustered in front of a transcription terminator. The pLK30 plasmids are general cloning vectors and the corresponding pLK50 plasmids carry the lambda pL promoter. The pLK60 vectors carry the lambda pR promoter and translation initiation signals of the cro gene containing the Shine-Dalgarno sequence and initiation codon. The pLK70 series is similar to pLK60 except that additional 5'-translated cro sequences are included. The pLK80 plasmids have a lacZ gene fragment suitable for the construction of hybrid genes. The presence of translational stop signals in the pLK90 series facilitates the manipulation of genes truncated at the 3' end. This standardized pLK vector system offers great versatility in gene manipulation and in optimization of gene expression under the control of strong regulatable promoters. Measurement of expression levels under repressed conditions permits the identification of optimal promoter-gene configurations in constructions directing high-level expression.

Base Sequence