PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon optimization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Optimality of the genetic code with respect to protein stability and amino-acid frequencies.

BACKGROUND: The genetic code is known to be efficient in limiting the effect of mistranslation errors. A misread codon often codes for the same amino acid or one with similar biochemical properties, so the structure and function of the coded protein remain relatively unaltered. Previous studies have attempted to address this question quantitatively, by estimating the fraction of randomly generated codes that do better than the genetic code in respect of overall robustness. We extended these results by investigating the role of amino-acid frequencies in the optimality of the genetic code. RESULTS: We found that taking the amino-acid frequency into account decreases the fraction of random codes that beat the natural code. This effect is particularly pronounced when more refined measures of the amino-acid substitution cost are used than hydrophobicity. To show this, we devised a new cost function by evaluating in silico the change in folding free energy caused by all possible point mutations in a set of protein structures. With this function, which measures protein stability while being unrelated to the code's structure, we estimated that around two random codes in a billion (109) are fitter than the natural code. When alternative codes are restricted to those that interchange biosynthetically related amino acids, the genetic code appears even more optimal. CONCLUSIONS: These results lead us to discuss the role of amino-acid frequencies and other parameters in the genetic code's evolution, in an attempt to propose a tentative picture of primitive life.

Amino Acid Substitution↗

Nucleotide sequence and structural analysis of the rat RT1.Eu and RT1.Aw3l genes, and of genes related to RT1.O and RT1.C.

A cDNA library was constructed using mRNA isolated from the R21 strain of rats which have the major histocompatibility complex (MHC) haplotype RT1.AlBlDlEu and the growth and reproduction complex (grc) genotype grc+. The cDNA clones that hybridized with the class I probes pAG64c and pARI.5 and were 1.3-1.7 kilobases were selected. Full-length clones were identified by sequencing partially the 5' and 3' ends of each clone, by the presence of a start codon at the 5' end, and by a polyadenylation sequence at the 3' end. The full-length cDNA clones were examined for in vitro transcription by transfection into human CIR cells using electroporation, and expression was detected by flow cytometry using monoclonal antibodies specific to the heavy chains and polyclonal antibody to beta 2-microglobulin. The RT1.Eu gene was transcribed and expressed optimally, and its nucleotide and deduced amino acid sequences differed significantly from the RT1.Aa, RT1.A(l), RT.Au, LW2, and 11/3R genes but only slightly from the RT1.K gene. The high level of sequence similarity between RT1.Eu and RT1.K suggests that the two genes may have originated from a common ancestral gene. In addition, three new genes (RT1.Aw3l, RT1.C-type, and RT1.O-type) were identified. The RT1.Aw3l gene is almost identical to RT1.A(l) with the exception of an in frame deletion of 21 nucleotides in exon 2 leading to a 7 amino acid deletion in the alpha 1 domain of the deduced amino acid sequence and 11 nucleotide substitutions and insertions in the rest of the sequence. It transcribed optimally, but no significant expression was detected. The RT1.C-type gene 119 is very similar (97%) to the LW2 gene in the 3' untranslated region, which suggests that it is in the RT1.C region. It transcribed optimally, but no significant expression was detected. The RT1.O-type gene 149 has all the features of a class Ib gene, but a premature stop codon in the alpha 1 domain causes incomplete translation. Its in vitro transcription was very low, and no expression was detected. These studies, combined with previous work, indicate that in the MHC of the R21 strain three class Ia genes (Eu, A(l), Aw3l) and three class Ib genes (C-type, O-type, N) are transcribed but only two class Ia genes (Eu, A(l)) are expressed.

Amino Acid Sequence↗

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment↗

Restructuring the translation initiation region of the human parathyroid hormone gene for improved expression in Escherichia coli.

Overexpression of native human parathyroid hormone in Escherichia coli was achieved by a modification of the 5' end of the genomic gene sequence, thereby adapting this part of the translation initiation region to the bacterial host. Some simple rules abstracted from optimization studies of translation initiation of a beta-interferon gene were applied. These included (a) extending complementarity of the mRNA to the anticodon loop of tRNAfMet by use of a codon with a purine nucleotide directly following the ATG, (b) avoidance of stable secondary structure in the mRNA by use of synonymous A/U-rich codons, (c) elimination of a potential second Shine-Dalgarno sequence. The appropriate silent changes led to a 20-fold increase in parathyroid hormone production resulting in 4.3% of total soluble protein. This result proves the validity of our simple approach for optimization of foreign gene expression in E. coli.

Base Sequence↗

Inducible expression vectors incorporating the Escherichia coli atpE translational initiation region.

New expression vectors were constructed for use in strains of Escherichia coli. Their most important feature is a polylinker system that facilitates the insertion of a gene in an optimal relationship to the highly efficient E. coli atpE translational initiation region (from nucleotide -50 to the start codon). Three ATG-containing restriction endonuclease sites can be used for the insertion of the 5' end of a gene at, or near to, its translational initiation codon. These sites may alternatively be used for the creation of a suitable translational start codon. Transcription is started by the bacteriophage lambda major promoters pR and pL in tandem and terminated by the bacteriophage fd terminator. Transcriptional initiation is very effectively repressed at 28-30 degrees C by the product of the bacteriophage lambda cIts857 gene, which is also present on the vectors. Full induction is achieved by shifting the incubation temperature to 42 degrees C. The combination of highly efficient transcriptional and translational signals on these vectors allowed high-level expression of sequences encoding human interferon beta and interleukin 2 and of the E. coli atpA, sucC and sucD genes.

DNA Restriction Enzymes↗

On the information content of the genetic code.

In living organisms 20 amino acids along with the terminator value(s) are encoded by 64 codons giving a degeneracy of the codons as described by the genetic code. A basic theoretical problem of genetic codes is to explain the particular distribution of degeneracies of partitions involved in the codes. In this work the degeneracy problem is considered in the framework of information theory. It is shown by direct numerical evaluation of a certain degeneracy information function associated with the genetic code that the degeneracy of the codes is observed to be related to the optimization of this function.

Amino Acids↗

Clinical value of K-ras codon 12 analysis and endobiliary brush cytology for the diagnosis of malignant extrahepatic bile duct stenosis.

Extrahepatic biliary stenosis can be caused by benign and malignant disorders. In most cases, a tissue diagnosis is needed for optimal management of patients, but the sensitivity of biliary cytology for the diagnosis of a malignancy is relatively low. The additional diagnostic value of K-ras mutational analysis of endobiliary brush cytology was assessed. Endobiliary brush cytology specimens obtained during endoscopic retrograde cholangiopancreaticography were prospectively collected from 312 consecutive patients with extrahepatic biliary stenosis. The results of conventional light microscopic cytology and K-ras codon 12 mutational analysis were compared and evaluated in view of the final diagnosis made by histological examination of the stenotic lesion and/or patient follow-up. The sensitivities of cytology and mutational analysis to detect malignancy were 36 and 42%, respectively. When both tests were combined, the sensitivity increased to 62%. The specificity of cytology was 98%, and the specificity of the mutational analysis and of both tests combined was 89%. Positive predictive values for cytology, mutational analysis, and both tests combined were 98, 92, and 94%, whereas the corresponding negative predictive values were 34, 34, and 44%, respectively. The sensitivity of K-ras mutational analysis was 63% for pancreatic carcinomas compared to 27% for bile duct, gallbladder, and ampullary carcinomas. K-ras mutational analysis can be considered supplementary to conventional light microscopy of endobiliary brush cytology to diagnose patients with malignant extrahepatic biliary stenosis, particularly in the case of pancreatic cancer. The presence of a K-ras codon 12 mutation in endobiliary brush cytology per se supports a clinical suspicion of malignancy, even when the conventional cytology is negative or equivocal.

Bile Duct Neoplasms↗

Cooperative effects by the initiation codon and its flanking regions on translation initiation.

The purine-rich Shine-Dalgarno (SD) sequence located a few bases upstream of the mRNA initiation codon supports translation initiation by complementary binding to the anti-SD in the 16S rRNA, close to its 3' end. AUG is the canonical initiation codon but the weaker UUG and GUG codons are also used for a minority of genes. The codon sequence of the downstream region (DR), including the +2 codon immediately following the initiation codon, is also important for initiation efficiency. We have studied the interplay between these three initiation determinants on gene expression in growing Escherichia coli. One optimal SD sequence (SD(+)) and one lacking any apparent complementarity to the anti-SD in 16S rRNA (SD(-)) were analyzed. The SD(+) and DR sequences affected initiation in a synergistic manner and large differences in the effects were found. The gene expression level associated with the most efficient of these DRs together with SD(-) was comparable to that of other DRs together with SD(+). The otherwise weak initiation codon UUG, but not GUG, was comparable with AUG in strength, if placed in the context of two of the DRs. The +2 codon was one, but not the only, determinant for this unexpectedly high efficiency of UUG.

Base Sequence↗

Folding of the MS2 coat protein in Escherichia coli is modulated by translational pauses resulting from mRNA secondary structure and codon usage: a hypothesis.

Possible translational pauses within the coat protein of the RNA bacteriophage MS2 were located on the basis of a distribution plot of rare codons and RNA secondary structure. It appeared that the position of certain codon pauses corresponds with the size of some nascent polypeptide intermediates, which have been isolated from MS2-infected cells. Other accumulated polypeptide intermediates seemed to be related to RNA regions, where double-stranded secondary structures occur, which probably impede the movement of ribosomes during chain elongation. We assume that a discontinuous translation rate is designed to allow optimal folding of this (and other) polypeptide(s).

Capsid↗

Rapid evolution of translational control mechanisms in RNA genomes.

We have introduced 13 base substitutions into the coat protein gene of RNA bacteriophage MS2. The mutations, which are clustered ahead of the overlapping lysis cistron, do not change the amino acid sequence of the coat protein, but they disrupt a local hairpin, which is needed to control translation of the lysis gene. The mutations decreased the phage titer by four orders of magnitude but, upon passaging, the virus accumulated suppressor mutations that raised the fitness to almost wild-type level. Analysis of the pseudorevertants showed that the disruption of the local hairpin, controlling expression of the lysis gene, had apparently been so complete that its restoration by chance mutations could not be achieved. Instead, alternative foldings initiated by the starting mutations were further stabilized and optimized. Strikingly, in the pseudorevertants analyzed, translational control of the lysis gene had been restored. This feat was accomplished by, on average, four suppressor mutations that generally occurred at codon wobble positions. We also introduced 11 mutations in a hairpin more upstream in the coat protein gene and not implicated in lysis control. Here the titer dropped by three logs, but pseudorevertants with a fitness close to wild-type were soon generated. These pseudorevertants again were the result of the optimization of alternative foldings induced by the mutations. The transition of the secondary structure from wild-type to pseudorevertant could be visualized by structure probing. Our study shows that the folding of the RNA is an important phenotypic property of RNA viruses. However, its distortion can easily be overcome by optimizing alternative base-pairings. These new structures are not qualitatively equivalent to the original one, since they do not successfully compete with the wild-type.

Base Sequence↗

Regulation of T-cell antigen receptor (TCR) alpha-chain expression by TCR beta-chain transcripts.

The TCR is an alpha beta heterodimer, a part of the multimeric structure through which physiological T-cell activation occurs. The expression of TCR alpha chain is greatly diminished in a beta-chain-deficient mutant Jurkat cell line (J.RT3-T3.5). The relationship between the expression of the TCR alpha and beta chains has been examined by stable transfection of a series of TCR beta-chain mutant constructs into this mutant cell line. The level of alpha-chain transcript was dramatically upregulated by the expression of the beta chain and specifically by a transcript of the beta-chain variable region alone, including a transcript in which the ATG start codon was mutated. The downregulation of the endogenous alpha-chain transcripts in mutants cells lacking complete beta-chain transcripts occurred primarily at the posttranscriptional level. This evidence for a regulatory function of the TCR beta-chain gene represents an unusual regulatory pathway in which the transcript of one gene is required for the optimal expression of another gene.

Amino Acid Sequence↗

A cell-free protein synthesis system as an investigational tool for the translation stop processes.

Using Escherichia coli cell-free protein synthesis system and aminoacylated amber suppressor tRNA, we successfully inserted an unnatural amino acid S-(2-nitrobenzyl)cysteine into human erythropoietin. Three different types of translation stop suppression were observed and each of the three types was easily discerned with SDS-PAGE. Optimal conditions were established for correct stop and programmed suppressions. Since this system differentiates proteins produced by misreading of codons from those produced by programmed suppression, we conclude that this cell-free translation system that we describe in this paper will be of a great use for future investigations on translation stop processes.

Cell-Free System↗

Recombination and chimeragenesis by in vitro heteroduplex formation and in vivo repair.

We describe a simple method for creating libraries of chimeric DNA sequences derived from homologous parental sequences. A heteroduplex formed in vitro is used to transform bacterial cells where repair of regions of non-identity in the heteroduplex creates a library of new, recombined sequences composed of elements from each parent. Heteroduplex recombination provides a convenient addition to existing DNA recombination methods ('DNA shuffling') and should be particularly useful for recombining large genes or entire operons. This method can be used to create libraries of chimeric polynucleotides and proteins for directed evolution to improve their properties or to study structure-function relationships. We also describe a simple test system for evaluating the performance of DNA recombination methods in which recombination of genes encoding truncated green fluorescent protein (GFP) reconstructs the full-length gene and restores its characteristic fluorescence. Comprising seven truncated GFP constructs, this system can be used to evaluate the efficiency of recombination between mismatches separated by as few as 24 bp and as many as 463 bp. The optimized heteroduplex recombination protocol is quite efficient, generating nearly 30% fluorescent colonies for recombination between two genes containing stop codons 463 bp apart (compared to a theoretical limit of 50%).

Cloning, Molecular↗

Screening of a mutant plasmid with high expression efficiency of GC-rich leuB gene of an extreme thermophile, Thermus thermophilus, in Escherichia coli.

A mutant plasmid with elevated expression efficiency of GC-rich Thermus thermophilus leuB gene was screened in Escherichia coli. A wild-type plasmid pHB2 carrying T. thermophilus leuB gene was introduced into leuB-deficient E. coli C600 cells. During successive cultures of the transformant in leucine-free medium, the original plasmid was spontaneously replaced by a mutant plasmid. The expression efficiency of the leuB gene on the mutant plasmid was 4.8-fold higher than that of the wild-type plasmid. Sequencing of the mutant plasmid revealed that the open reading frame (ORF1) in front of the leuB gene was shortened from 822 to 306 bp. Several expression vectors were constructed to investigate the effect of the length of ORF1, and the optimal length for the expression of the following leuB gene was determined. It was also shown that the stop codon of ORF1 should be overlapped with the initiation codon of leuB gene for the highest efficiency.

Amino Acid Sequence↗

The Mycobacterium tuberculosis katG promoter region contains a novel upstream activator.

An Escherichia coli-mycobacterial shuttle vector, pJCluc, containing a luciferase reporter gene, was constructed and used to analyse the Mycobacterium tuberculosis katG promoter. A 1.9 kb region immediately upstream of katG promoted expression of the luciferase gene in E. coli and Mycobacterium smegmatis. A smaller promoter fragment (559 bp) promoted expression with equal efficiency, and was used in all further studies. Two transcription start sites were mapped by primer extension analysis to 47 and 56 bp upstream of the GTG initiation codon. Putative promoters associated with these show similarity to previously identified mycobacterial promoters. Deletions in the promoter fragment, introduced with BAL-31 nuclease and restriction endonucleases, revealed that a region between 559 and 448 bp upstream of the translation initiation codon, designated the upstream activator region (UAR), is essential for promoter activity in E. coli, and is required for optimal activity in M. smegmatis. The katG UAR was also able to increase expression from the Mycobacterium paratuberculosis P(AN) promoter 15-fold in E. coli and 12-fold in M. smegmatis. An alternative promoter is active in deletion constructs in which either the UAR or the katG promoters identified here are absent. Expression from the katG promoter peaks during late exponential phase, and declines during stationary phase. The promoter is induced by ascorbic acid, and is repressed by oxygen limitation and growth at elevated temperatures. The promoter constructs exhibited similar activities in Mycobacterium bovis BCG as they did in M. smegmatis.

Bacterial Proteins↗

Tracking and quantitation of retroviral-mediated transfer using a completely humanized, red-shifted green fluorescent protein gene.

We have developed murine retroviral vectors (RVs) containing an optimized green fluorescent protein (GFP) gene to study retroviral gene transfer and expression in living cells. We used the codon "humanized", "red-shifted" GFP gene, hGFP-S65T, a gain of function variant of the wild-type GFP from the jellyfish Aequorea victoria. We cloned the hGFP-S65T gene into the RV plasmid pLNCX (pLNChG65T). A stable amphotropic RV-producer cell line (VPC), designated LNChG65T VPC, was generated that exhibited bright fluorescence in greater than 95% of the cells. Human A375 melanoma cells and IGROV ovarian carcinoma cells transduced from LNCh-G65T VPC demonstrated high levels of fluorescence. The expression of a single integrated hGFP-S65T gene in eukaryotic cells provides a powerful tool to study gene transfer, expression and functional studies in vitro and in vivo.

Animals↗

Genotypic classification of colorectal adenocarcinoma. Biologic behavior correlates with K-ras-2 mutation type.

BACKGROUND: New measures enabling better prediction of biologic behavior of large bowel cancer are highly desirable. One hundred ninety-four consecutive primary, recurrent, and metastatic colorectal adenocarcinomas, accessioned during 1991 at Rhode Island Hospital, were classified according to the presence and specific type of K-ras-2 point mutation. METHODS: An integrated histopathologic-genetic approach was used to detect mutations starting with minute, topographically selected, tissue samples from formaldehyde-fixed, paraffin-embedded specimens. RESULTS: Each colorectal adenocarcinoma exhibited either no or only one of seven specific types of K-ras-2 mutation. The mutation type of each primary tumor was present consistently in its metastatic deposits. Thirty-five percent of primary colorectal adenocarcinomas were found to be mutated (42 of 119). A significantly higher mutation rate (65%) was seen in lymphogenous-hematogenous metastases as a group (35 of 54; P < 0.005). By contrast, 22% of anastomotic recurrences and transcoelomic metastasis were mutated (4 of 18). Twenty-eight percent of adenocarcinomas with invasion limited to muscularis propria (Tis, T1, T2) were mutated (16 of 57), compared to 41% for more deeply invasive tumors (T3, T4; 26 of 63). When colorectal adenocarcinomas were analyzed by specific K-ras-2 mutation type, it was found that codon 13 mutated tumors did not progress to local or distant metastasis (P < 0.01). Tumors having a codon 12 valine substitution did not metastasize beyond pericolonic-perirectal lymph nodes. In contrast, colorectal cancers with codon 12 aspartic acid substitutions accounted for most of the distant hematogenous deposits (P < 0.01). Tumors with normal K-ras-2 accounted for most intraperitoneal deposits. CONCLUSIONS: Genotyping of colorectal adenocarcinoma by K-ras-2 status can identify subsets of patients likely to pursue indolent and aggressive forms of disease. The integrated histopathologic-genetic approach outlined is feasible for use in diagnostic pathology, providing information that together with clinicopathologic staging may individualize and optimize treatment.

Adenine↗

Line probe assay for rapid detection of drug-selected mutations in the human immunodeficiency virus type 1 reverse transcriptase gene.

Upon prolonged treatment with various antiretroviral nucleoside analogs such as 3'-azido-3'-deoxythymidine, 2',3'-dideoxyinosine, 2',3'-dideoxycytidine, (-)- beta-L-2', 3'dideoxy-3'thiacytidine and 2',3'-didehydro-3'-deoxythymidine, selection of human immunodeficiency virus type 1 (HIV-1) strains with mutations in the reverse transcriptase (RT) gene has been reported. We designed a reverse hybridization line probe assay (LiPA) for the rapid and simultaneous characterization of the following variations in the RT gene: M41 or L41; T69, N69, A69, or D69; K70 or R70; L74 or V74; V75 or T75; M184, I184, or V184; T215, Y215, or F215; and K219, Q219, or E219. Nucleotide polymorphisms for codon L41 (TTG or CTG), T69 (ACT or ACA), V75 (GTA or GTG), T215 (ACC or ACT), and Y215 (TAC or TAT) could be detected. In addition to the codons mentioned above, several third-letter polymorphisms in the direct vicinity of the target codons (E40, E42, K43, K73, D76, Q182, Y183, D185, G213, F214, and L214) were found, and specific probes were selected. In total, 48 probes were designed and applied to the LiPA test strips and optimized with a well-characterized and representative reference panel. Plasma samples from 358 HIV-infected patients were analyzed with all 48 probes. The amino acid profiles could be deduced by LiPA hybridization in an average of 92.7% of the samples for each individual codon. When combined with changes in viral load and CD4+ T-cell count, this LiPA approach proved to be useful in studying genetic resistance in follow-up samples from antiretroviral agent-treated HIV-1-infected individuals.

Acquired Immunodeficiency Syndrome↗