PubMed HealthSearch

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Signals determining translational start-site recognition in eukaryotes and their role in prediction of genetic reading frames.

A special methionyl-tRNA (RNAi) is universally required to initiate translation. The conversation of this reactant throughout evolution, as well as its unusual decoding properties, suggested an alternate mechanism for tRNA-mRNA interactions at initiation. We have reported that the sequence of bases neighboring the start codons of many eubacterial genes are complementary not only to the 16S rRNA 3' end and to the anticodon of tRNAi, but, also, have the potential to base-pair the D, T or extended anticodon loops of this tRNAi. The coding properties of tRNAi and mutations that affect translation suggest that these signals may function. This hypothesis explains the observation that unusual triplets can start prokaryotic and mitochondrial genes and predicts the occurrence of other reading frames. Furthermore, it suggests a unifying model of chain initiation based on RNA-RNA contacts and displacements. Here we examine the start domain of 290 eukaryotic genes for their ability to base-pair the tRNAi loops and the 18S rRNA. We observe that both methionine start, and methionine coding regions have the potential to pair with the 18S rRNA, but that the nucleotide distribution about start codons strongly favoured such pairings over that near internal AUGs. The 5' extended anticodon of tRNAi is methylated, and was not represented in the mRNA with high frequency. However, the tetramer AUGg did occur with high frequency in the start domain. A modification of the tRNAi T loop also decreases its base-pairing potential. Interestingly, complementarity to the T loop did not occur with high frequency in the start sites. The early coding region, 10 to 34 nucleotides 3' to the initiator AUG, is complementary to the tRNAi D loop in many cases, while no such affinity is found near internal AUGs. The nucleotides around initiator AUGs were heavily biassed toward the sequence gccaccAUGgcg. No such tendency was noted around internal AUGs. Although the role of this sequence bias is unclear, the sequence gccaccAUGg has been shown by Kozak to promote initiation. Another distinguishing feature was a C-rich tract 7 to 34 nucleotides 5' to the initiator AUGs. Ability to pair with more than eight bases of the start consensus sequence, matching of 6 or 7 nucleotides to the D loop on the 3' side, an C-richness on the 5' side were used as criteria for distinguishing start AUGs.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals

Cloning and sequence determination of a cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase.

A partial cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase (ACMPK) was isolated from a lambda ZAP expression library by immunoselection using monospecific polyclonal antibodies to the enzyme. The sequence of both strands of the cDNA (CMKa) was determined. The deduced amino acid (aa) sequence contained all eleven consensus domains found in serine/threonine protein kinases [Hanks et al., Science 241 (1988) 42-52], as well as a putative calmodulin-binding domain. The cDNA contained an intron, lacked an in-frame start codon, and was not polyadenylated. A full-length copy of CMKa was subsequently isolated from a lambda gt10 library of A. nidulans cDNA using a restriction fragment of the first clone as a probe. It contained an in-frame start codon, an open reading frame (ORF) of 1242 bp and was polyadenylated. The ORF encoded a protein of 414 aa residues with an M(r) of 46,895 and an isoelectric point pI = 6.4. These values are in good agreement with that observed for the native enzyme [Bartelt et al., Proc. Natl. Acad. Sci. USA 85 (1988) 3279-3283]. When aligned to optimize homology, 29% of the predicted aa sequence of ACMPK is identical to that of the alpha-subunit of rat brain calmodulin-dependent protein kinase II. ACMPK shares 40 and 44% identity in aa sequence with YCMK1 and YCMK2, respectively, two Ca2+/calmodulin-dependent protein kinases recently cloned from Saccharomyces cerevisiae [Pausch et al., EMBO J. 10 (1991) 1511-1522]. Results of Southern analysis of restriction digests of genomic DNA indicate that ACMPK is encoded by a single-copy gene.

Amino Acid Sequence

Expression of recombinant growth hormone in Escherichia coli: effect of the region between the Shine-Dalgarno sequence and the ATG initiation codon.

We constructed a synthetic Escherichia coli expression system in which various promoter elements can be changed easily. In this study we investigated the effect of a number of portable Shine-Dalgarno regions (SD regions) on the synthesis of two modified recombinant human growth hormones (hGH). The production of these modified hGH was measured during exponential growth and after the bacteria had reached stationary phase. The results show that the optimal distance between the SD region (AGGAGG) and the ATG start codon is approximately 11 nucleotides. However, the nucleotide sequence in this region also influences expression: 6-10 adenines result in comparable expression levels despite the varying lengths. Two overlapping SD regions reduce expression of the growth hormones considerably, whereas two potential ATG start codons do not affect expression. Having a SD-ATG region partly or totally complementary to the 5' end of the 16S ribosomal RNA does not alter translation efficiency. Estimation of the delta G values for the association between the 16S rRNA and the ribosome-binding region suggests that these are not indicators of expression efficiency.

Base Sequence

T7 RNA polymerase can direct expression of influenza virus cap-binding protein (PB2) in Escherichia coli.

Influenza virus cap-binding protein (PB2; Mr 85,000) is made in Escherichia coli when the cloned cDNA is transcribed by T7 RNA polymerase. Translation begins at the probable natural start codon and also from at least five internal sites in the same reading frame. The eukaryotic initiation site is not typical of protein initiation sites of E. coli, in that the closest potential Shine-Dalgarno sequence is far (15 nucleotides) from the start codon. Nevertheless, protein synthesis initiates efficiently at this site even in competition with a strong upstream prokaryotic initiation site. PB2 is somewhat unstable in the cell, but accumulates to a level where it is easily detectable in electrophoresis patterns of total cell protein. The full-length protein and various subfragments of it are insoluble in crude extracts, but have been useful for producing antibodies.

Base Sequence

Histidine ammonia-lyase from Streptomyces griseus.

Histidine ammonia-lyase (histidase; HutH) has been purified to homogeneity from Streptomyces griseus and the N-terminal amino acid (aa) sequence used to clone the histidase-encoding structural gene, hutH. The purified enzyme shows typical saturation kinetics and is inhibited competitively by D-histidine and histidinol phosphate. High concentrations of K.cyanide inactivate HutH unless the enzyme is protected by the substrate or histidinol phosphate. On the basis of the nucleotide sequence, the hutH structural gene would encode a protein of 53 kDa with an N terminus identical to that determined for the purified enzyme. Immediately upstream from hutH is a region that strongly resembles a class of Streptomyces promoters active during vegetative growth; however, there is no obvious ribosome-binding site adjacent to the hutH translation start codon. The deduced aa sequence of an upstream partial open reading frame shows no similarity with other proteins, including HutP of Bacillus subtilis and HutU of Pseudomonas putida. Promoter-probe analysis indicates that promoter activity maps within the DNA surrounding the hutH start codon. Pairwise comparisons of the primary structures of bacterial and mammalian histidases, together with the unique kinetic properties and gene organization, suggest that streptomycete histidase may represent a distinct family of histidases.

Amino Acid Sequence

Nucleotide sequence of the leader region of the phenylalanine operon of Escherichia coli.

The pheA structural gene of the phenylalanine operon of Escherichia coli is preceded by a transcribed leader region of about 170 nucleotide pairs. In vitro transcription of plasmids and restriction fragments containing the phe promoter and leader region yields a major RNA transcript about 140 nucleotides in length. This transcript, pheA leader RNA, has the following features: (i) a potential ribosome binding site and AUG translation start codon about 20 nucleotides from its 5' end; (ii) 14 additional in phase amino acid codons and a UGA stop codon after the AUG; 7 of these 14 are Phe codons; (iii) a 3'-OH terminus about 140 nucleotides from the 5' end (transcription termination occurs in an A.T-rich region which is subsequent to a G.C-rich region; just beyond the site of transcription termination there is a sequence corresponding to a ribosome binding site and the AUG translation start codon of the pheA structural gene); (iv) a sequence which would permit extensive intrastrand stable hydrogen bonding. In addition to G.C-rich stem structures, highly analogous to those proposed for the leader RNAs of the tryptophan operons of E. coli and Salmonella typhimurium [Lee, F. & Yanofsky, C. (1977) Proc. Natl. Acad. Sci. USA 74, 4365-4369], there is also extensive base-pairing possible between the phe codon region and a more distal region of the leader transcript. The roles of synthesis of the Phe-rich leader peptide and secondary structure of the leader transcript in the regulation of transcription termination at the attenuator of the phe operon are discussed.

Base Sequence

Interactions of plasmid-encoded replication initiation proteins with the origin of DNA replication in the broad host range plasmid RK2.

The TrfA proteins, encoded by the broad host range plasmid RK2, are required for replication of this plasmid in a variety of Gram-negative bacteria. Two TrfA proteins, 33 and 44 kDa in molecular mass (designated TrfA-33 and TrfA-44, respectively), are expressed from the trfA gene of RK2 through the use of two alternative in-frame start codons within the same open reading frame. The two proteins have been purified from Escherichia coli to near homogeneity as a mixture of wild-type TrfA-44/33, as TrfA-33 alone and as a functional variant form of TrfA-44, designated TrfA-44(98L), which contains a leucine in place of the TrfA-33 methionine start codon. Cross-linking experiments demonstrated that TrfA-33 can multimerize in solution. By using gel mobility shift and DNase I footprinting techniques the binding properties of TrfA-33, TrfA-44(98L), and TrfA-44/33 to the origin of replication of plasmid RK2 were analyzed. All three protein preparations were able to bind very specifically to the cluster of five direct repeats (iterons) contained in the minimal origin of replication. Each protein preparation produced a ladder of TrfA/minimal oriV complexes of decreasing electrophoretic mobility. The DNase I protection pattern on the five iterons was identical for all three protein preparations and extended from the beginning of the first iteron to 5 base pairs upstream of the fifth iteron. Studies on the affinity of the proteins for DNA fragments containing one, two, or all five iterons of the origin revealed a strong preference of TrfA protein for DNA containing at least two iterons. To study the stability of TrfA.DNA complexes, association and dissociation rates of TrfA-33 and DNA fragments with one, two, or five iterons were measured. This analysis showed that unlike complexes involving two or five iterons the TrfA/one iteron complexes were highly unstable, suggesting some form of cooperativity between proteins or iterons in the formation of stable complexes and/or the requirement of specific sequences bordering the iterons at the RK2 origin of replication for the stabilization of TrfA/DNA complexes.

Bacterial Proteins

Enrichment for 5'-TG termini: a method for subcloning structural genes into expression vectors.

We describe a method for creating a population of randomly digested, blunt-ended DNA fragments with 5'-TG... 3'-AC... (TG) at their termini. When these fragments were ligated to a blunt-ended vector that contains ...CATA-3' ...GTAT-5' at its termini, as high as 84% of the clones obtained after transformation contained plasmids with a reconstructed Nde I site ... CATATG... ...GTATAC.... When the DNA vector is prepared from an appropriate plasmid [Gross et al., Mol. Cell. Biol. 5 (1985) 1015-1024; Kotewicz et al., Gene 35 (1985) 249-258], the ATG within the restriction site corresponds to a start codon positioned downstream from a strong ribosome-binding site and controllable promoter. If a gene has been digested to the TG contained within its authentic initiation codon, the endogenous translation-initiation control sequences are deleted and expression can be controlled using the plasmid-derived promoter. In addition, a gene digested to other in-frame TGs can potentially express proteins with altered N termini. Using this method, we have placed the structural gene of SP6 RNA polymerase, trimmed precisely to its authentic start codon, under the control of the tac promoter.

Base Sequence

Cloning and characterization of the MboII restriction-modification system.

The two genes encoding the class IIS restriction-modification system MboII from Moraxella bovis were cloned separately in two compatible plasmids and expressed in E. coli RR1 delta M15. The nucleotide sequences of the MboII endonuclease (R.MboII) and methylase (M.MboII) genes were determined and the putative start codon of R.MboII was confirmed by amino acid sequence analysis. The mboIIR gene specifies a protein of 416 amino acids (MW: 48,617) while the mboIIM gene codes for a putative 260-residue polypeptide (MW: 30,077). Both genes are aligned in the same orientation. The coding region of the methylase gene ends 11 bp upstream of the start codon of the restrictase gene. Comparing the amino acid sequence of M.MboII with sequences of other N6-adenine methyltransferases reveals a significant homology to M.RsrI, M.HinfI and M.DpnA. Furthermore, M.MboII shows homology to the N4-cytosine methyltransferase BamHI.

Amino Acid Sequence

Conformational alteration of mRNA structure and the posttranscriptional regulation of erythromycin-induced drug resistance.

The DNA sequence of the ermC gene of plasmid pE194 is presented. This determinant is responsible for erythromycin-induced resistance to the macrolide-lincosamide-streptogramin B group of antibiotics and specifies a 29,000 dalton inducible protein. The locations of the ermC promoter, as well as that of a probable transcriptional terminator, are established both from the sequence and by transcription mapping. The sequence contains an open reading frame sufficient to encode the previously identified 29,000 dalton ermC protein. Between the promoter and the putative ATG start codon is a 141 base pair leader sequence, within which several regulatory (constitutive) mutations have been mapped and sequenced. The leader has a second open reading frame, sufficient to encode a 19 amino acid peptide. It is suggested that induction by erythromycin involves a shift between alternative ribosome-bound mRNA conformations, so that the ribosome binding sequence and the start codon for synthesis of the 29K protein are unmasked in the presence of inducer. Possible active and inactive folded configuration of the leader sequence are presented, as well as the effects on these configurations of regulatory mutations.

Bacillus subtilis

Nucleotide sequence and transcript analysis of three photosystem II genes from the cyanobacterium Synechococcus sp. PCC7942.

The genome of the cyanobacterium Synechococcus sp. PCC7942 contains two genes encoding the D2 polypeptide of photosystem II (PSII), which are designated here as psbDI and psbDII. The psbDI gene, like the psbD gene of plant chloroplasts, is cotranscribed with and overlaps the open reading frame of the psbC gene, encoding the PSII protein CP43. The psbDII gene is not linked to psbC, and appears to be transcribed as a monocistronic message. The two psbD genes encode identical polypeptides of 352 amino acids, which are 86% conserved with the D2 polypeptide of spinach. In plants, the translational start codon of the psbC gene has been reported to be an ATG codon 50 bp upstream from the end of the psbD gene. This triplet is not present in the psbDI sequence of Synechococcus sp., but is replaced by ACG, a codon which is very unlikely to initiate translation. Translation of the psbC gene may begin at a GTG codon which overlap the psbDI open reading frame by 14 bp and is preceded by a block of homology to the 3' end of the 16S ribosomal RNA, a potential ribosome-binding site. There are only two bp differences between the sequences of the two psbD genes; one of these results in substitution in psbDII of GCG for the presumed GTG start codon in psbDI.

Amino Acid Sequence

DNA sequence of the Escherichia coli gene, gnd, for 6-phosphogluconate dehydrogenase.

Expression of gnd of Escherichia coli, which encodes 6-phosphogluconate dehydrogenase, an enzyme of the hexose monophosphate shunt, is subject to growth rate-dependent regulation and is gene dosage-dependent: the level of the enzyme increases in direct proportion to the cellular growth rate at both low and high gene copy numbers. We have determined the nucleotide sequence of gnd and flanking control regions, the 5'-end of in vivo gnd mRNA, and the start codon of the structural gene. Analysis of the sequence indicated that: (i) the gnd promoter is typical of other E. coli promoters and the structural gene is followed by a rho-independent transcription termination signal; (ii) the 56-nucleotide leader of gnd mRNA does not contain a rho-independent transcription termination signal, so growth rate-dependent regulation of 6-phosphogluconate dehydrogenase level is not carried out by an attenuation mechanism analogous to the one that controls expression of the E. coli ampC gene; (iii) the codon composition of the structural gene resembles that of other highly expressed E. coli genes and thus is not responsible for the regulation either; (iv) the structural gene is preceded at an optimal distance by a strong Shine-Dalgarno (SD) sequence, AGGAG ; (v) the leader region of the mRNA contains regions of dyad symmetry that have the potential to sequester the SD sequence and the start codon. This latter feature of the gene suggests that growth rate-dependent regulation may involve regulation of translation initiation frequency.

Amino Acid Sequence

Antisense oligonucleotide inhibition of encephalomyocarditis virus RNA translation.

We report the inhibition of encephalomyocarditis virus (EMCV) RNA translation in cell-free rabbit reticulocyte lysates by antisense oligonucleotides (13-17-base oligomers) complementary to (a) the viral 5' non-translated region, (b) the AUG start codon and (c) the coding sequence. Our results demonstrate that the extent of translation inhibition is dependent on the region where the complementary oligonucleotides bind. Non-complementary and 3'-non-translated-region-specific oligonucleotides had no effect on translation. A significant degree of translation inhibition was obtained with oligonucleotides complementary to the viral 5' non-translated region and AUG initiation codon. Digestion of the oligonucleotide:RNA hybrid by RNase H did not significantly increase translation inhibition in the case of 5'-non-translated-region-specific and initiator-AUG-specific oligonucleotides; in contrast, RNase H digestion was necessary for inhibition by the coding-region-specific oligonucleotide. We propose that (a) 5'-non-translated-region-specific oligonucleotides inhibit translation by affecting the 40S ribosome binding and/or passage to the AUG start codon, (b) AUG-specific oligonucleotides inhibit translation initiation by inhibiting the formation of an active 80S ribosome and (c) the coding-region-specific oligonucleotide does not prevent protein synthesis because the translating 80S ribosome can dislodge the oligonucleotide from the EMCV RNA template.

Animals

Different promoters of SHV-2 and SHV-2a beta-lactamase lead to diverse levels of cefotaxime resistance in their bacterial producers.

Clinical Klebsiella pneumoniae isolates as well as Escherichia coli transformants producing the beta-lactamases SHV-2 or SHV-2a demonstrate MIC values for cefotaxime of 4 mg l-1 or 64 to greater than 128 mg l-1, respectively. The beta-lactamases differ by one possibly insignificant amino acid exchange at position number 10 of the mature protein; their kinetic parameters are rather similar. The 5' untranslated regions of both corresponding genes show no homology starting 74 nucleotides upstream to the start codon. Hybridization of intragenically annealing oligonucleotides to dot-blotted serial dilutions of total cellular RNA from E. coli transformants harbouring these genes cloned into the same vector plasmid gave a positive signal down to 1.2 micrograms (SHV-2) and 0.32 to 0.16 micrograms (SHV-2a), indicating a four to eight times higher amount of specific transcript in the case of SHV-2a. By primer extension analysis and S1 nuclease digestion the starting point to transcription was located 100 nucleotides (SHV-2) and 50 nucleotides (SHV-2a) in front of the start codon. No other transcripts of different length could be detected after prolonged exposure. Northern blot analysis demonstrated the length of the beta-lactamase mRNA to be about 1.6 kb in both cases, thus comprising a potential open reading frame downstream of the two enzymes' genes. Selective PCR amplification of both promoter regions and of the structural gene of SHV-2 and subsequent combined cloning of each of the promoters and the SHV-2 gene into pBGS19 using a BamHI restriction site introduced by three point mutations into the cloned sequences was employed to transforms E. coli DH5 alpha.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence

Alternative initiation of translation determines cytoplasmic or nuclear localization of basic fibroblast growth factor.

Three forms of basic fibroblast growth factor (bFGF), initiated at an AUG (18 kDa) and two CUG (21 and 22.5 kDa) start codons, were produced following transfection of COS cells with human hepatoma bFGF cDNA. The subcellular localization of the different forms was investigated directly or by using chimeric genes constructed by fusion of the bFGF and chloramphenicol acetyltransferase open reading frames. The AUG-initiated proteins were cytoplasmic, while the CUG-initiated forms were nuclear. The signal sequence responsible for the nuclear localization of bFGF is contained within 37 amino acid residues between the second CUG and the AUG start codons. Alternative initiation of translation regulates the subcellular localization of bFGF and thus could modulate its role in cell growth and differentiation control.

Amino Acid Sequence

Conserved reiterated domains in Clostridium thermocellum endoglucanases are not essential for catalytic activity.

The complete nucleotide sequence of the Clostridium thermocellum celE gene, coding for an endo-beta-1,4-glucanase (endoglucanase E; EGE) with xylan-hydrolysing activity has been determined. The structural gene consists of an open reading frame (ORF) of 2442 bp commencing with a GTG start codon and followed by a TAA stop codon. The nucleotide sequence obtained has been confirmed by comparing the predicted amino acid sequence with that derived by N-terminal amino acid sequencing of the purified protein. The EGE sequence contains a region homologous to the reiterated domain found at the C terminus of other endoglucanases from the same organism. BAL 31 deletions of the structural gene have revealed the extent to which this conserved sequence is necessary for endoglucanase and xylanase activity. A region of DNA, upstream from the structural gene has also been sequenced and a ribosome-binding site and putative promoter sequences have been identified. A second ORF which ends 349 bp 5' to the GTG start codon of the celE gene has also been identified. The encoded product contains a C terminus homologous to other C. thermocellum endoglucanases.

Amino Acid Sequence

Cytomegalovirus assembly protein nested gene family: four 3'-coterminal transcripts encode four in-frame, overlapping proteins.

The genomic region encoding the assembly protein of simian cytomegalovirus (CMV) strain Colburn has been cloned, sequenced, and found to be organized as a nested set of four in-frame, 3'-coterminal genes, each with its own TATA promoter element and translational start codon, and all using a single 3' polyadenylation signal. The 3' end of the longest open reading frame (1.770 bp) was identical to the 930-bp sequence coding for the assembly protein precursor, as determined from a cDNA clone. The assembly protein coding region of human CMV strain AD169 was similarly organized, suggesting that both viral genomes could give rise to four independently transcribed 3'-coterminal RNAs coding for four overlapping, in-frame, carboxy-coterminal proteins. These predictions were tested and confirmed. Four mRNAs corresponding in size and sequence to those predicted were identified in both human and simian CMV-infected cells by using transcript-specific antisense oligonucleotide probes in Northern (RNA blot) assays. The 5' ends of the three largest of these Colburn transcripts were determined by S1 nuclease protection assays and found to map between the anticipated TATA sequences and corresponding translational start codons. The four predicted overlapping proteins were identified by immunoassays in lysates of simian and human CMV-infected cells by using an antiserum specific for the carboxyl end of the assembly protein precursor. The structural relationship of both sets of proteins was verified by comparing their peptide patterns following protein cleavage at tryptophan residues by N-chlorosuccinimide. The similar organization of the homologous coding regions in other herpesviruses into at least two nested, in-frame, 3'-coterminal genes is discussed.

Amino Acid Sequence

SV40 deletion mutant (d1861) with agnoprotein shortened by four amino acids.

d1861 is an SV40 deletion mutant which was thought to lack the agnoprotein coding region and was used to verify the role of agnoprotein in the life cycle of SV40. In the present study the region flanking the deletion was sequenced and, in contrast to the available information, it was found that d1861 lacks 12 nt in phase, downstream from the AUG start codon of agnoprotein (residues 347-358). Using the runoff protocol with viral transcriptional complexes (VTC), that in vitro elongate the in vivo preinitiated nascent RNA, it was found that in vivo the major initiation site for late transcription is at residue 325, the same as in wild type (WT). In comparison with WT, d1861 encodes information for agnoprotein shortened by four amino acids and it has been identified in d1861 infected cells. However, pulse-chase experiments indicated that the rate of synthesis of d1861 agnoprotein is slower than that of WT agnoprotein and that it has a turnover rate of 1 hr as compared to 3 hr of WT agnoprotein. The reduced rate of synthesis of d1861 agnoprotein can be explained by nuclease S1 analyses in which the major leader of d1861 16 S RNA, that encodes the agnoprotein, appeared in significantly lower amounts as compared to the major leader of WT 16 S RNA. Furthermore, analysis of the potential secondary structures at the 5' end of the leader of d1861 16 S RNA has revealed stable structures in which the start codon of agnoprotein is sequestered in a stem. The involvement of RNA secondary structures in regulating the synthesis of agnoprotein is discussed.

Amino Acid Sequence