PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A human testis-specific mRNA for phosphoribosylpyrophosphate synthetase that initiates from a non-AUG codon.

Two highly homologous subunits for phosphoribosylpyrophosphate synthetase are encoded by human X-linked genes, PRPS1 and PRPS2 (Taira, M., Kudoh, J., Minoshima, S., Iizasa, T., Shimada, H., Shimizu, Y., Tatibana, M., and Shimizu, N. (1989b) Somat. Cell Mol. Genet. 15, 29-37). These genes are expressed in most tissues, whereas an additional unique mRNA (1.4 kilobases) is present in the testes of rats as well as mice and humans (Taira, M., Iizasa, T., Yamada, K., Shimada, H., and Tatibana, M. (1989a) Biochim. Biophys. Acta 1007, 203-208). In this paper, cDNA cloning revealed that the human testis-specific mRNA was encoded by an autosomal gene, termed PRPS3. RNA blot analysis showed that the expression of this gene began at 4 weeks of age in rats, coinciding with the reported appearance of primary spermatocytes. A cDNA clone of PRPS3 was sequenced and found to encode a predicted product of 317 amino acids which was highly homologous to those of PRPS1 and PRPS2 (94.3% and 91.2% identities, respectively). However, the PRPS3 cDNAs lacked an ATG initiator for translation at the expected position, and instead contained an ACG triplet. In vitro transcription/translation studies, combined with in vitro site-directed mutagenesis experiments, suggested that the ACG codon at this position did serve as a start codon. Analysis of amino-terminal sequence of the radiolabeled PRPS3 product, prepared by in vitro translation, supported the predicted sequence starting with Pro-1, and, in addition, this product was labeled with N-formyl[35S]methionyl-tRNAi. These results suggested that the synthesis of the nascent polypeptide could initiate with methionine at the position corresponding to the ACG codon.

Amino Acid Sequence↗

Use of silent mutations in cDNA encoding human glutathione transferase M2-2 for optimized expression in Escherichia coli.

Heterologous expression of human glutathione transferase M2-2 (GST M2-2) using Escherichia coli was improved 140-fold by mutating the cDNA expressing the enzyme. Expression of GST M2-2 from this cDNA clone, pKHXhGM2, generated approximately 190 mg protein per liter of bacterial culture, corresponding to approximately 12% of the total amount of soluble protein. The high-level-expressing cDNA was generated by oligonucleotide-directed mutagenesis introducing alternative silent mutations into the third nucleotide of codons 2, 4-7, and 10-14 in the 5' end of the cDNA coding region. The choice of alternative codons was restricted to those naturally occurring in highly biased genes in E. coli. Furthermore, the wild-type TAG stop codon at the 3' end was replaced with the two stop codons TAA and TGA in tandem to increase translation termination efficiency. The resulting partially randomized cDNA library was assayed for high-level expression using immunoscreening. Sequence similarities between the constructed high-level-expressing GST M2-2 cDNA and a similarly designed cDNA encoding the closely related human GST M1-1 suggest that the codons in the region immediately following the start codon are influential in achieving high-level expression. Pyrimidines seem to be more favorable than purines in the third position of codons in optimizing the expression of these enzymes in E. coli.

Amino Acid Sequence↗

The L1 family of long interspersed repetitive DNA in rabbits: sequence, copy number, conserved open reading frames, and similarity to keratin.

The L1 family of long interspersed repetitive DNA in the rabbit genome (L1Oc) has been studied by determining the sequence of the five L1 repeats in the rabbit beta-like globin gene cluster and by hybridization analysis of other L1 repeats in the genome. L1Oc repeats have a common 3' end that terminates in a poly A addition signal and an A-rich tract, but individual repeats have different 5' ends, indicating a polar truncation from the 5' end during their synthesis or propagation. As a result of the polar truncations, the 5' end of L1Oc is present in about 11,000 copies per haploid genome, whereas the 3' end is present in at least 66,000 copies per haploid genome. One type of L1Oc repeat has internal direct repeats of 78 bp in the 3' untranslated region, whereas other L1Oc repeats have only one copy of this sequence. The longest repeat sequenced, L1Oc5, is 6.5 kb long, and genomic blot-hybridization data using probes from the 5' end of L1Oc5 indicate that a full length L1Oc repeat is about 7.5 kb long, extending about 1 kb 5' to the sequenced region. The L1Oc5 sequence has long open reading frames (ORFs) that correspond to ORF-1 and ORF-2 described in the mouse L1 sequence. In contrast to the overlapping reading frames seen for mouse L1, ORF-1 and ORF-2 are in the same reading frame in rabbit and human L1s, resulting in a discistronic structure. The region between the likely stop codon for ORF-1 and the proposed start codon for ORF-2 is not conserved in interspecies comparisons, which is further evidence that this short region does not encode part of a protein. ORF-1 appears to be a hybrid of sequences, of which the 3' half is unique to and conserved in mammalian L1 repeats. The 5' half of ORF-1 is not conserved between mammalian L1 repeats, but this segment of L1Oc is related significantly to type II cytoskeletal keratin.

Animals↗

Nucleotide sequence of the Drosophila glucose-6-phosphate dehydrogenase gene and comparison with the homologous human gene.

Glucose-6-phosphate dehydrogenase (G6PD) has a major role in NADPH production and is found in almost all cell types. The structural gene for G6PD is X-linked in Drosophila melanogaster, as it is in most eukaryotic organisms, and due to its ubiquitous expression, it can be considered a typical 'housekeeping' gene. Here we present the complete nucleotide (nt) sequence of G6PD cDNAs as well as the genomic copy of the G6PD gene. The G6PD gene has three introns so that the protein-coding region is divided into four segments. The 5'-end of mature G6PD mRNA is located 289 +/- 1 nt upstream from the start codon. The sequence upstream from the transcription start point is G + T-rich and contains no commonly found transcription regulatory elements, such as a TATA box or GGGCGG sequence. D. melanogaster G6PD is 65% homologous with the human G6PD protein but has no homology with the human sequence for the first 42 amino acid residues. The G6PD gene was shown to be active when transduced to autosomal positions. For each transformant, G6PD activity in both male and female adults was not significantly different, indicating that the transduced gene, unlike the resident G6PD, is not dosage-compensated in males.

Amino Acid Sequence↗

Characterization of an Escherichia coli gene encoding betaine aldehyde dehydrogenase (BADH): structural similarity to mammalian ALDHs and a plant BADH.

An open reading frame of 1476 nucleotides, cloned from a region of the Escherichia coli genome encoding betaine biosynthesis functions, was shown to encode a betaine aldehyde dehydrogenase (BADH; EC 1.2.1.8). Either of two adjacent codons (5'-GTGATG) could function as a start codon, producing a presumptive polypeptide of 491 or 490 amino acids. The deduced primary structure of the E. coli BADH showed 39-43% positional identity, over its entire length, to aldehyde dehydrogenases (ALDH: EC 1.2.1.3) of mammalian origin. This similarity increased to 75-77% when conservative aa substitutions were also taken into consideration. Spinach BADH was also similar to the bacterial BADH, showing 38% identity and 80% overall similarity. Other homologs included a fungal and a putative bacterial ALDH. Although E. coli BADH was specific for the substrate, betaine aldehyde, it showed the highest levels of similarity to the prototype human ALDH-2. Only one gap in each sequence had to be introduced for optimal alignment. The conservation between E. coli BADH and the ALDHs was also evident in the predicted secondary structures and hydrophilicity profiles of the polypeptides, suggesting a similarity in the overall folding patterns of ALDH and BADH. These observations suggest a common ancestry for BADH and ALDH, preceding prokaryote-eukaryote divergence.

Aldehyde Dehydrogenase↗

Presence of splice variant forms of cytochrome P4502D1 in rat brain but not in liver.

Cytochromes P450 (P450), a family of heme-containing proteins, is involved in the oxidative metabolism of both foreign and endogenous compounds. Although liver is quantitatively the major organ involved in the metabolism of most xenobiotics, there is increasing evidence that these enzymes are present in extrahepatic tissues, such as lung, kidney, brain, etc and they may contribute to the in situ metabolism of xenobiotics in these organs. The possible relationship between genetic polymorphism seen in P4502D6 and incidence of neurodegenerative diseases, such as Parkinson's disease, has prompted the characterization of P4502D enzymes in rat brain. In the present study, we demonstrate that P4502D1 (the rat homologue of human P4502D6) is constitutively expressed in rat brain and the mRNA and protein are localized predominantly in neuronal cell population in the olfactory bulb, cortex, cerebellum, and hippocampus. An alternate spliced transcript of CYP2D1 having exon 3 deletion was detected in rat brain but not in liver. Deletion of exon 3 causes frame shift and generates a stop codon at 391 bp relative to the start codon ATG leading to premature termination of translation. Thus, Northern blotting and in situ hybridization represent contributions from functional transcripts and alternate spliced variants that do not translate into functional protein. Further, the splice variant having partial inclusion of intron 6 detected in human brain was not detected in rat brain indicating that alternate spliced gene products of P450 enzymes are generated in species-specific and tissue-specific manner.

Alternative Splicing↗

Genomic organization and characterization of the mouse ELYS gene.

Differentiation of hematopoietic stem cells into blood cells is controlled by several transcription factors. Recently, we identified a putative transcription factor, ELYS (for embryonic large molecule derived from yolk sac), using a subtraction strategy. During mouse embryogenesis, ELYS transcripts were predominantly expressed in hematopoietic tissues, such as the yolk sac, aorta-gonad-mesonephros (AGM), and liver. Here, we report the cloning and characterization of the mouse ELYS gene. The ELYS gene spanned approximately 60kb encoding 36 exons, and was assigned between D1Mit315 and D1Mit458 markers in chromosome 1. The transcription initiation site was identified as the G residue located 670bp upstream of the translation start codon. A region downstream of the transcriptional start site contributed to high promoter activity. This region contained potential DNA elements for transcription factors such as GATA-1, -2, -3, heat shock factor (HSF) 2, and NF-kappaB, which are known to play important roles in hematopoietic events.

Animals↗

Preference for guanosine at first codon position in highly expressed Escherichia coli genes. A relationship with translational efficiency.

The variation in base composition at the three codon sites in relation to gene expressivity, the latter estimated by the Codon Adaptation Index, has been studied in a sample of 1371 Escherichia coli genes. Correlation and regression analyses show that increasing expression levels are accompanied by higher frequencies of base G at first, of base A at second and of base C at third codon positions. However, correlation between expressivity and base compositional biases at each codon site was only significant and positive at first codon position. The preference for G-starting codons as gene expression level increases is discussed in terms of translational optimization.

Amino Acids↗

Complete nucleotide sequence and transcription of ermF, a macrolide-lincosamide-streptogramin B resistance determinant from Bacteroides fragilis.

DNA sequence analysis of a portion of an EcoRI fragment of the Bacteroides fragilis R plasmid pBF4 has allowed us to identify the macrolide-lincosamide-streptogramin B resistance (MLSr) gene, ermF. ermF had a relative moles percent G + C of 32, was 798 base pairs in length, and encoded a protein of approximately 30,360 daltons. Comparison between the deduced amino acid sequence of ermF and six other erm genes from gram-positive bacteria revealed striking homologies among all of these determinants, suggesting a common origin. Based on these and other data, we believe that ermF codes for an rRNA methylase. Analysis of the nucleotide sequences upstream and downstream from the ermF gene revealed the presence of directly repeated sequences, now identified as two copies of the insertion element IS4351. One of these insertion elements was only 26 base pairs from the start codon of ermF and contained the transcriptional start signal for this gene as judged by S1 nuclease mapping experiments. Additional sequence analysis of the 26 base pairs separating ermF and IS4351 disclosed strong similarities between this region and the upstream regulatory control sequences of ermC and ermA (determinants of staphylococcal origin). These results suggested that ermF was not of Bacteroides origin and are discussed in terms of the evolution of ermF and the expression of drug resistance in heterologous hosts.

Amino Acid Sequence↗

Expression of dicistronic transcriptional units in transgenic tobacco.

We investigated whether the two cistrons of a dicistronic mRNA can be translated in plants to yield both gene products. The coding sequences of various reporter genes were combined in dicistronic units, and their expression was analyzed in stably transformed tobacco plants at the RNA and protein levels. The presence of an upstream cistron resulted in all cases in a drastically reduced expression of the downstream cistron. The translational efficiency of the gene located downstream in the dicistronic units was 500- to 1,500-fold lower than that in a monocistronic control; a 500-fold lower value was obtained with a dicistronic unit in which both cistrons were separated by 30 nucleotides, whereas a 1,500-fold lower value was obtained with a dicistronic unit in which the stop codon of the upstream cistron and the start codon of the downstream cistron overlapped. As a strategy to select indirectly for transformants with enhanced levels of expression of a gene which is by itself nonselectable, the gene of interest can be cloned upstream from a selectable marker in a dicistronic configuration. This strategy can be used provided that the amount of dicistronic mRNA is high. If, on the other hand, the expression of the dicistronic unit is too low, selection of the downstream cistron will primarily give clones with rearranged dicistronic units.

Base Sequence↗

Isolation and characterization of the mouse gene for the type 3 iodothyronine deiodinase.

The type 3 iodothyronine deiodinase (D3) is a selenoenzyme that inactivates thyroid hormones by removing a iodine from the 5-position of the tyrosyl ring. D3 is highly expressed in many tissues during the early stages of development, and its activity is regulated by selected growth factors and various hormones. To gain further insights into the structure, functional role, and regulation of this enzyme, we screened a mouse liver genomic library with a rat D3 complementary DNA probe and isolated a 12-kb clone coding for the Dio3. Restriction analysis followed by Southern blotting and nucleotide sequencing demonstrated that the Dio3 contains a single exon, 1853 bp in length, that encodes the entire length of the messenger RNA expressed in murine placenta and neonatal skin. Primer extension experiments identified two potential transcriptional start sites located 77 and 60 nt upstream of the ATG translational start codon. The region immediately 5' to the start sites contains consensus TATA, CAAT, and GC elements. Furthermore, a 526-nucleotide genomic fragment from this region was demonstrated to efficiently drive a luciferase reporter construct when transfected into COS-7, XTC-2, or XL-2 cells or into primary cultures of rat preadipocytes derived from neonatal brown fat. In conclusion, D3 transcripts in the placenta and skin are encoded by the Dio3 gene from a single exon whose expression is regulated by an upstream region that contains several consensus promoter elements. Further characterization of this gene will provide new insights into the factors regulating the unique pattern of D3 expression during development.

Amino Acid Sequence↗

High-level expression vectors to synthesize unfused proteins in Escherichia coli.

A new class of plasmid vectors (pANK-12, pANH-1, and pPL2) for synthesizing unfused proteins was constructed by inserting synthetic linkers at the NdeI site (CATATG) of plasmid pJL6, which contains the lambda cII gene initiator codon. These expression vectors contain the lambda pL promoter, the cII ribosome-binding site, cII start codon and unique restriction sites (KpnI, Asp718, HpaI, BamHI) downstream from the initiator ATG for expression of unfused proteins. The main advantage of these vectors is that any DNA fragment with an open reading frame that does not possess a start and/or a stop codon can be directed to overproduce protein in an unfused form.

Cloning, Molecular↗

Purified Escherichia coli F-factor TraY protein binds oriT.

The traY gene of the Escherichia coli F plasmid has been shown by genetic studies (R. Everett and N. Willetts, J. Mol. Biol. 136:129-150, 1980) to be involved in the site-specific nicking reaction at oriT required for the initiation of DNA transfer during bacterial conjugation. In order to assign a biochemical function to TraY protein, the traY gene was cloned in a plasmid vector which utilizes the strong T7 phi 10 promoter to overproduce the protein. The plasmid-encoded TraY protein was specifically labeled with [35S]methionine, and purification of the polypeptide was accomplished by monitoring the radioactive label. Purified TraY protein had a relative molecular mass of approximately 17,000, as determined by polyacrylamide gel electrophoresis in the presence of sodium dodecyl sulfate. The amino terminus of the purified protein was sequenced to confirm that the protein was encoded by the traY gene. The protein sequence revealed that the start codon for the TraY protein was a UUG codon 36 base pairs upstream of the AUG start site originally deduced from the DNA sequence (T. Fowler, L. Taylor, and R. Thompson, Gene 26:79-89, 1983). This start sequence confirmed the premise of Inamoto et al. that the F-plasmid TraY polypeptide-coding sequence would begin with UUG, creating a reading frame which renders a large degree of amino acid sequence identity with the TraY polypeptide from R100 (S. Inamoto, Y. Yoshioka, and E. Ohtsubo, J. Bacteriol. 170:2749-2757, 1988). The purified TraY protein from F bound specifically to the origin of transfer region of the F plasmid. However, no nicking activity was detected at oriT by using TraY protein or TraY protein in conjunction with helicase I.

Bacterial Proteins↗

Construction of a new shuttle expression vector for Bacillus subtilis and Escherichia coli by using a polycistronic system.

A shuttle vector has been constructed by fusing the Bacillus subtilis trimethoprim-resistance-carrying (TpR) plasmid pNC601 with the Escherichia coli plasmid pBR322. The resultant plasmid pNBL1 can replicate in both B. subtilis and E. coli, conferring Tp resistance on both cells and ampicillin resistance (ApR) on E. coli. The B. subtilis dihydrofolate reductase operon (dfr) on pNC601 and therefore on pNBL1 consists of the thymidylate synthase B gene (thyB) and the TpR-dihydrofolate reductase gene lacking the C-terminal seven codons (designated as drfA' as compared with the complete dfrA gene). A direct-expression vector pNBL3 has been constructed by inserting synthetic oligodeoxynucleotides containing a Bacillus ribosome-binding site (RBS) and the ATG codon downstream from dfrA' on pNBL1. When the E. coli lacZ gene was placed downstream from the dfrA' gene in pNBL3, efficient synthesis of beta-galactosidase was observed in both cells, showing that the polycistronic expression system is suitable for directing expression of heterologous genes. Translational efficiency of the lacZ gene on pNBL3 was further examined in B. subtilis by changing the sequence upstream from lacZ. Unlike the results previously reported [Sprengel et al., Nucleic Acids Res. 13 (1985) 893-909], when RBS was present, the high level of lacZ expression was preserved irrespective of spacing between the stop codon of the upstream dfrA' gene and the start codon of the downstream lacZ gene. However, in the absence of RBS, the spacing between both genes affected lacZ expression. That is, translational coupling of dfrA'-lacZ was observed, although the translational efficiency was very low.

Ampicillin Resistance↗

Production of human prolyl 4-hydroxylase in Escherichia coli.

Prolyl 4-hydroxylase (P4H) catalyzes the post-translational hydroxylation of proline residues in collagen strands. The enzyme is an alpha2beta2 tetramer in which the alpha subunits contain the catalytic active sites and the beta subunits (protein disulfide isomerase) maintain the alpha subunits in a soluble and active conformation. Heterologous production of the native alpha2beta2 tetramer is challenging and had not been reported previously in a prokaryotic system. Here, we describe the production of active human P4H tetramer in Escherichia coli from a single bicistronic vector. P4H production requires the relatively oxidizing cytosol of Origami B(DE3) cells. Induction of the wild-type alpha(I) cDNA in these cells leads to the production of a truncated alpha subunit (residues 235-534), which assembles with the beta subunit. This truncated P4H is an active enzyme, but has a high Km value for long substrates. Replacing the Met235 codon with one for leucine removes an alternative start codon and enables production of full-length alpha subunit and assembly of the native alpha2beta2 tetramer in E. coli cells to yield 2 mg of purified P4H per liter of culture (0.2 mg/g of cell paste). We also report a direct, automated assay of proline hydroxylation using high-performance liquid chromatography. We anticipate that these advances will facilitate structure-function analyses of P4H.

Catalysis↗

Molecular characterization of a HMW glutenin subunit allele providing evidence for silencing of x-type gene on Glu-B1.

Understanding the molecular structure of high-molecular-weight glutenin subunit (HMW-GS) may provide useful evidence for the study on the improvement of quality of cultivated wheat and the evolution of Glu-1 alleles. Sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) shows that the subunits encoded by Glu-B1 were null, named 1Bxm, in a Triticum turgidum var. dicoccoides line PI94640. Primers based on the conserved regions in wheat HMW-GS gene promoter and coding sequences were used to amplify the genomic DNA of line PI94640. The PCR products were sequenced, and the total nucleotide sequence of 3,442 bp including upstream sequence of 1,070 bp was obtained. Compared with the reported gene sequences of Glu-1Bx alleles, the promoter region of the Glu-1Bxm showed close resemblance to 1Bx7. The Glu-1Bxm coding region differs from the other Glu-1Bx alleles for a deduced mature protein with only 212 residues, and a stop codon (TAA) at 637 bp downstream from the start codon was present, which was probably responsible for the silencing of x-type subunit genes at the Glu-B1 locus. Phylogenetic tree based on the nucleotide sequence alignment of HMW glutenin subunit genes showed that 1Bxm was the most ancient type of Glu-B1 alleles, suggesting that the evolution rates are different among Glu-1Bx genes. Further study on the contribution of the unique silenced Glu-B1 alleles to quality improvement was also discussed.

Alleles↗

Transcription of three sets of genes coding for the core light-harvesting proteins in the purple sulfur bacterium, Allochromatium vinosum.

The nucleotide sequence of the puf operon coding for the subunits of the photosynthetic reaction center and the core light-harvesting complex (LH1) of the purple sulfur bacterium, Allochromatium (A.) vinosum (formally Chromatium vinosum), was completely determined. Unlike other known puf operons, which contain only one set of genes coding for the LH1 apoproteins, pufB and pufA, the A. vinosum puf operon included three sets of pufB and pufA genes with a gene order of pufB (1) A (1) LMCB (2) A (2) B (3) A (3). Northern hybridization analysis suggested that all of the nine puf genes are co-transcribed as a 4.43 kb mRNA. Three small mRNAs corresponding to pufB (2) A (2) B (3) A (3), pufB (2) A (2) B (3), and pufB (2) A (2) were detected, as well as two small mRNAs covering pufB (1) A (1). Analysis of the nucleotide sequence of the puf operon, including the flanking regions and 5'-ends of the six mRNAs, suggested that the transcription of the A. vinosum puf operon is initiated at 74 bp downstream from the bchZstop codon (295 bp upstream from the pufB (1) start codon), and regulated by a promoter located at its direct upstream. The possible promoter is overlapped with a binding motif of a repressor protein for pigment-biosynthesis genes, PpsR or CrtJ, known in other purple bacteria. No other possible promoters were found within the puf genes. These findings indicate that three sets of pufA and pufB genes of A. vinosum are co-transcribed as a long mRNA containing all the puf genes, and, from this long mRNA, the five short mRNAs are possibly derived by post-transcriptional modifications.

Journal Article↗