PubMed HealthSearch

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Paramecium mitochondrial DNA sequences and RNA transcripts for cytochrome oxidase subunit I, URF1, and three ORFs adjacent to the replication origin.

A 2-kb region adjacent to the replication origin (ori) and a 3-kb region located between the small and large ribosomal RNAs of Paramecium mitochondrial (mt) DNA have been sequenced and the locations of their transcripts determined. The ori segment contains four transcripts, some of which are overlapping, which encode a known protein and two other open reading frames. The other segment encodes, on separate transcripts, the cytochrome c oxidase subunit one gene (COI) and the URF1 gene (ND1) common to most mt genomes. All these genes have the same orientation and do not contain introns. The COI gene is the most divergent of those known and has an internal 108 amino acid 'insert' not found in COI genes from other organisms. With these data it is possible to define a probable Paramecium mt genetic code. With the exception that TGA codes for tryptophan and the use of different start codons, Paramecium mtDNA appears to follow the universal code. GTA possibly can be used as a start codon.

Amino Acid Sequence

A genomic clone containing the promoter for the gene encoding the human lysosomal enzyme, alpha-galactosidase A.

We have isolated and characterized a human genomic clone for a lysosomal enzyme gene. The start point of transcription was identified using primer extension of poly(A)+ mRNA. This genomic clone is specific for human alpha-galactosidase A, and it includes sequences for the promoter, complete signal peptide, first exon, and part of the first intron. Direct and inverted repeat elements of 10, 11, 16, 19, and 22 nucleotides (nt) flank the promoter site. A (GA)n repeat element of approx. 60 nt with strong homology to similar elements identified in several species is located upstream from the promoter. A GGGCGG site specific for DNA-binding protein Sp1 is located near a CAAT box, and the CCGCCC inverted repeat of the Sp1 binding sequence is located by the TATA box. The sequence immediately flanking the ATG start codon of the human alpha-galactosidase A is highly homologous to sequences flanking the ATG start codons of the other human lysosomal hydrolases for which sequence information is available (beta-glucocerebrosidase, cathepsin B, cathepsin D, and beta-hexosaminidase alpha chain), but not for any of the other 133 human signal peptides examined. Our analysis also reveals that conversion of the propeptide to the mature enzyme involves cleavage of a C-terminal rather than an N-terminal fragment. This information about the normal alpha-galactosidase A gene will be useful for comparison to data obtained from patients with Fabry disease, who are characterized by a deficiency of this enzyme. This is the first genomic clone described to date for any lysosomal enzyme, and it establishes a reference for future analyses of the molecular events that mediate the expression of this important class of enzymes.

Amino Acid Sequence

Structural organization of the genes for murine and human leukemia inhibitory factor. Evolutionary conservation of coding and non-coding regions.

Leukemia inhibitory factor, LIF, is a glycoprotein with multiple activities in both the adult and the embryo. LIF appears to be encoded by a unique gene in both mouse and man, although the 3'-untranslated region of the mouse LIF gene gives a complex hybridization pattern on Southern blots. The complete nucleotide sequences of both the murine and human LIF genes and their flanking regions (8.7 and 7.6 kilobase pairs, respectively) were determined and compared. Both genes comprise three exons, two introns and an unusually long 3'-untranslated region (3.2 kilobase pairs), specificying a mRNA of approximately 4.1 kilobases. Two start sites of LIF-transcription were determined, by S1-nuclease protection and by a novel approach involving the polymerase chain reaction. S1-nuclease protection revealed a start site 60-64 base pairs upstream of the translational start codon and immediately downstream of a TATA box (TATATAAAT). The PCR approach identified a second transcriptional start site 160 base pairs 5' of the start codon and adjacent to a "TATA-like" element (CATAATTT). A comparison of the murine and human LIF gene sequences revealed a high degree of conservation in the coding regions and in segments of the untranslated and flanking regions. Seven segments displaying greater than 75% homology were identified, with the 5' and 3' ends of the transcription unit revealing the highest degree of homology. These conserved regions represents potential cis-acting control elements.

Amino Acid Sequence

Molecular structure and transformation of the glucose dehydrogenase gene in Drosophila melanogaster.

We have precisely mapped and sequenced the three 5' exons of the Drosophila melanogaster Gld gene and have identified the start sites for transcription and translation. The first exon is composed of 335 nucleotides and does not contain any putative translation start codons. The second exon is separated from the first exon by 8 kb and contains the Gld translation start codon. The inferred amino acid sequence of the amino terminus contains two unusual features: three tandem repeats of serine-alanine, and a relatively high density of cysteine residues. P element-mediated transformation experiments demonstrated that a 17.5-kb genomic fragment contains the functional and regulatory components of the Gld gene.

Animals

The parainfluenza virus type 1 P/C gene uses a very efficient GUG codon to start its C' protein.

Parainfluenza virus type 1 (PIV1) and Sendai virus (SEN) are very closely related, but the PIV1 P/C gene does not contain the ACG codon which initiates the SEN C' protein. Nevertheless, a protein corresponding to the PIV1 C' protein was observed both in vivo and in vitro. The initiation site of this protein maps upstream of the PIV1 C protein AUG in a region that does not contain an AUG codon. We have used site-directed mutagenesis to demonstrate that the PIV1 C' protein initiates from a GUG codon, four codons upstream of where the ACG is found in SEN. Remarkably, this GUG appears to initiate in vivo almost as frequently as AUG in the same context. However, whereas GUG permits downstream expression of the P and C proteins, AUG in this context does not. The conservation of an upstream non-AUG initiation codon for C' among PIV1 and SEN suggests that it is important for virus replication, even though some paramyxoviruses express only the C protein and others have no C open reading frame at all.

Base Sequence

Scanning model for translational reinitiation in eubacteria.

Premature termination of translation in eubacteria, like Escherichia coli, often leads to reinitiation at nearby start codons. Restarts also occur in response to termination at the end of natural coding regions, where they serve to enforce translational coupling between adjacent cistrons. Here, we present a model in which the terminated but not released ribosome reaches neighboring initiation codons by lateral diffusion along the mRNA. The model is based on the finding that introduction of an additional start codon between the termination and the reinitiation site consistently obstructs ribosomes to reach the authentic restart site. Instead, the ribosome now begins protein synthesis at this newly introduced AUG codon. This ribosomal scanning-like movement is bidirectional, has a radius of action of more than 40 nucleotides in the model system used, and activates the first encountered restart site. The ribosomal reach in the upstream direction is less than in the downstream one, probably due to dislodging by elongating ribosomes. The proposed model has parallels with the scanning mechanism postulated for eukaryotic translational initiation and reinitiation.

Bacteriolysis

Regulatory elements and transcriptional regulation by testosterone and retinoic acid of the rat nerve growth factor receptor promoter.

The low-affinity nerve growth factor receptor (LNGFR) is a membrane-associated glycoprotein which is thought to participate in some of the biological activities of nerve growth factor (NGF). Expression of the LNGFR gene is known to be regulated both during development and in response to various agents in cell culture. However, molecular mechanisms responsible for the regulation have not been described. We report here an analysis of a 4.8-kb sequence from the 5'-flanking region of the rat LNGFR gene. Several regulatory elements were identified in this region by transfection of plasmid constructs containing sequences from LNGFR fused to a bacterial cat reporter gene. The proximal part of the promoter region (0.4-kb) was shown to be sufficient to support cat expression in all cell types used. A silencer element located between -1.5 kb and -1.8 kb from the start of translation, as well as an enhancer element in more upstream regions of the promoter, were identified in the phaeochromocytoma cell line, PC12, and in the Sertoli cell line, TM4, that express the LNGFR gene. Treatment of TM4 cells with retinoic acid (RA) increases the level of LNGFR mRNA twofold, while testosterone treatment results in a tenfold decrease. Regions of the promoter responsive to testosterone and RA in TM4 cells were found at -610 to -860 bp and -1840 to -4800 bp upstream from the translation start codon, respectively. A RA-responsive element active in PC12 cells is located between bp -610 to -860 from the start codon.

Animals

Transcriptional analysis of the restriction and modification genes of bacteriophage P1.

Bacteriophage P1 res and mod genes encode the restriction and modification polypeptides of the Type III restriction enzyme EcoP1. Northern blot analysis using res- and mod-specific probes revealed the presence of two separate transcripts in strains harbouring the EcoP1 restriction and modification genes. Furthermore, by constructing a series of fusions with a promoter less lacZ gene, we show that both the res and mod genes are transcribed from separate promoters. A more detailed investigation of the mod promoter region revealed two promoters located some 70 and 140bp upstream from the translational start codon. In addition, another pair of promoters and a further separate promoter are located more than 500bp upstream from this start codon. Two short open reading frames are located between these distal and proximal promoter clusters. Transcription of the res gene is initiated from within the mod open reading frame from two adjacent promoters. In addition a functional promoter is located on the antisense strand close to the res promoter region. The relationship between the transcription units of the res and mod genes is discussed.

Amino Acid Sequence

Autogenous regulation of the gene for transcription termination factor rho in Escherichia coli: localization and function of its attenuators.

We present evidence that the expression of rho is regulated by rho-dependent attenuation of transcription. Gene fusion analysis with nested series of deletions of rho indicated that the transcription of rho is attenuated in a rho-dependent manner in the leader region and that neither a read-through transcription from the upstream gene, trxA, nor a modulation of transcription initiation of the rho promoter is involved in the self-control of rho. S1 mapping and Northern hybridization analyses localized at least six transcription attenuation or termination sites in the region ranging from the 3' end of the trxA structural gene to the middle of the rho structural gene. Among them, the most upstream site overlapping the rho promoter sequence was assigned to the terminator for the trxA gene, and the second and third sites, mapping about 80 and 50 nucleotides upstream from the start codon of rho, were suggested to function as the major attenuation sites for regulation of the rho expression. Further, the start points of the trxA and rho RNAs were determined in an in vitro transcription system to be located 111 nucleotides (U) and 255 nucleotides (G) upstream from their respective start codons. These results necessitate revisions of previous predictions on the sites of transcriptional signals in the trxA and rho genes (S. Brown, B. Albrechtsen, S. Pedersen, and P. Klemm, J. Mol. Biol. 162:283-298, 1982; C.-J. Lim, D. Geraghty, and J. A. Fuchs, J. Bacteriol. 163:311-316, 1985; B.J. Wallace and S.R. Kushner, Gene 32:399-408, 1984).

Chromosome Deletion

Overexpression of the Thermus aquaticus B malate dehydrogenase-encoding gene in Escherichia coli.

Expression of the Thermus aquaticus B malate dehydrogenase (MDH)-encoding gene (mdh), cloned in Escherichia coli, was initially at a relatively low level (0.1% of soluble cell protein) and was effected by read-through from the tac promoter in the plasmid vector used. An enhancement in expression to 0.4% of soluble cell protein was achieved by shortening the intervening sequence between the promoter and the translation start codon of mdh. An NdeI restriction site (5'-CAT-ATG-3') was engineered in the shortened fragment, which also changed the start codon from GTG to ATG. This resulted in an eightfold increase in expression, to 3.2% of soluble cell protein. Expression was further increased by subcloning the mdh gene via the engineered NdeI site, into two plasmid expression vectors, one carrying the E. coli trpP promoter and the other the E. coli mdhP promoter. In both these expression systems, 40-50% of the soluble cell protein was T. aquaticus MDH. This suggests that expression of the cloned T. aquaticus mdh in E. coli is enhanced predominantly by the optimisation of transcription and translation initiation signals. Moreover, the base composition of the coding region and the pattern of codon usage dictated by it appear to have little effect on expression. Heat treatment of the cell extract at 85 degrees C further effected purification of T. aquaticus MDH to over 80% of the soluble cell protein. The MDHs purified to homogeneity from the high-expression clones were identical with the MDH isolated from T. aquaticus B cells with respect to all measured parameters.

Base Sequence

Molecular cloning and nucleotide sequence of the aminopeptidase T gene of Thermus aquaticus YT-1 and its high-level expression in Escherichia coli.

Aminopeptidase T (AP-T) is a metallo-dependent dimeric enzyme of Thermus aquaticus YT-1, an extremely thermophilic bacterium. We cloned the AP-T gene from T. aquaticus YT-1 into Escherichia coli using a synthetic oligonucleotide as a hybridization probe. The nucleotide sequence of the AP-T gene was found to encode 408 amino acid residues with GTG as a start codon. The molecular weight was calculated to be 44,820. The AP-T was overproduced in E. coli (about 5% of total soluble protein) when the start codon of the gene was changed from GTG to ATG, and the gene was downstream from the tac promoter. The AP-T expressed in E. coli was heat stable and easily purified by heat treatment (80 degrees C, 30 min). The N-terminal amino acid sequence of AP-T showed similarity with that of aminopeptidase II from Bacillus stearothermophilus.

Amino Acid Sequence

Dual translational initiation sites control function of the lambda S gene.

Lysis gene S of phage lambda has a 107 codon reading frame beginning with the codons Met1-Lys2-Met3. Genetic data have suggested that translational initiation occurs at both Met1 and Met3, generating two polypeptides, S107 and S105 respectively. We have proposed a model in which the proper scheduling of lysis depends on the partition of translational initiations between the two start codons. Here, using in vitro methods, we show that two stem-loop structures, one immediately upstream of the reading frame and a second approximately 10 codons within the gene, control the partitioning event. Utilizing primer-extension inhibition or 'toeprinting', we show that the two S start codons are served by two adjacent Shine-Dalgarno sequences. Moreover, the timing of lysis supported by the wild-type and a number of mutant alleles in vivo can be correlated with the ratio of ternary complex formation over Met1 and Met3 in vitro. Thus the regulation of the S gene is unique in that the products of two adjacent in-frame initiation events have opposing function.

Bacteriolysis

High molecular mass forms of basic fibroblast growth factor are initiated by alternative CUG codons.

A 6.75-kilobase human hepatoma-derived basic fibroblast growth factor (bFGF) cDNA was cloned and sequenced. An amino-terminal sequence generated from a purified hepatoma bFGF was found to correspond to the nucleotide sequence and to begin 8 amino acids upstream from the putative methionine start codon thought to initiate a 154-amino acid bFGF translation product. This sequence suggests that a form of bFGF of at least 163 amino acids exists. The hepatoma cDNA was transcribed in vitro into RNA; in vitro translation of this RNA generated three forms of bFGF with molecular masses of 18, 21, and 22.5 kDa. By use of in vitro mutagenesis, it was found that the 22.5-kDa bFGF and possibly the 21-kDa form were initiated with CUG start codons. The 18-kDa bFGF was initiated with an AUG codon. By transfecting into COS cells human hepatoma bFGF cDNA and a construct from which the AUG initiator was eliminated, it was found that the higher molecular mass forms of bFGF were as biologically active as the 18-kDa form.

Amino Acid Sequence

Coordinate increase in major transcripts from the high pI alpha-amylase multigene family in barley aleurone cells stimulated with gibberellic acid.

The purpose of this study was to identify specifically genes and transcripts for the high pI isozyme of barley alpha-amylase. From hybridization of coding sequence probes to blots of genomic DNA digested with restriction enzymes that do not cut within our cloned high pI alpha-amylase cDNA, it is estimated that about 7 alpha-amylase genes or pseudogenes exist. No difference could be detected between barley aleurone cell and sprout DNAs. Experiments using probes from the 5' and 3' untranslated sequences of the high pI alpha-amylase cDNA clone identified three HindIII fragments that probably carry high pI sequences. Primer extension experiments used as a primer the terminal 5' coding sequence from our cDNA clone; this primer would not cross-hybridize to low pI alpha-amylase transcripts. Two major transcripts were identified. These shared a conserved 23-base sequence immediately 5' to the ATG start codon, although a C----G transversion and a 3-base deletion were present within this sequence. An unusual 8-base pair GC palindrome was present in the conserved region immediately preceding the ATG start codon. Distal to the conserved sequence there was no apparent homology. One transcript carrying a 97-base untranslated region was identical to our high pI cDNA clone E. The gene for the other was recovered from a lambda phage genomic library. The 5' coding sequence was very similar, but not identical to clone E, demonstrating that these transcripts arise from separate genes. The two transcripts increased coordinately in aleurone cells stimulated with gibberellic acid. These data indicate that there is a high pI alpha-amylase multigene family with at least two active members, both of which are regulated in some manner by the plant hormone gibberellic acid.

Base Sequence

The complete nucleotide sequence of the TL-DNA of the Agrobacterium tumefaciens plasmid pTiAch5.

We have determined the complete primary structure (13 637 bp) of the TL-region of Agrobacterium tumefaciens octopine plasmid pTiAch5 . This sequence comprises two small direct repeats which flank the TL-region at each extremity and are involved in the transfer and/or integration of this DNA segment in plants. TL-DNA specifies eight open-reading frames corresponding to experimentally identified transcripts in crown gall tumor tissue. The eight coding regions are not interrupted by intervening sequences and are separated from each other by AT-rich regions. Potential transcriptional control signals upstream of the 5' and 3' ends of all the transcribed regions resemble typical eukaryotic signals: (i) transcriptional initiation signals ('TATA' or Goldberg- Hogness box) are present upstream to the presumed translational start codons; (ii) ' CCAAT ' sequences are present upstream of the proposed 'TATA' box; (iii) polyadenylation signals are present in the 3'-untranslated regions. Furthermore, no Shine-Dalgarno sequences are present upstream of the presumed translational start codons.

Amino Acid Oxidoreductases

Amino acid sequence of the L-lactate dehydrogenase of Bacillus caldotenax deduced from the nucleotide sequence of the cloned gene.

The Bacillus caldotenax L-lactate dehydrogenase gene (lct) has been cloned into Escherichia coli, using the Bacillus stearothermophilus lct gene as a hybridisation probe, and its complete nucleotide sequence determined. The lct structural gene consists of an open reading frame of 951 base pairs commencing with an ATG start codon and followed by a TAA stop codon. Upstream of the gene are putative transcriptional promoter -35 and -10 regions; a ribosome binding site with a predicted delta G of -66.9 kJ/mol is also present six base pairs upstream of the ATG start codon. The B. caldotenax lct gene is highly homologous to the B. stearothermophilus lct gene displaying a DNA sequence homology of 89.7%. Examination of the DNA sequence 3' of the lct gene revealed the presence of two further open reading frames. This suggests that the lct gene may be the first gene of an operon. The deduced amino acid sequence of the L-lactate dehydrogenase (LDH) from B. caldotenax predicted a protein of 317 amino acid residues; comparison with the B. stearothermophilus enzyme revealed only 30 amino acid differences between the two enzymes; thus the enzymes are 90.4% homologous. These amino acid differences must account for the different thermostabilities of the two enzymes. The B. caldotenax lct gene was efficiently expressed in E. coli and the original lct-containing plasmid construct isolated (pKD1) induced the synthesis of LDH at a level of 4.5% of the E. coli soluble cell protein whilst a SmaI subfragment of this clone, (pKD2) produced LDH at a level of 6.9% of the E. coli soluble cell protein. LDH isolated from E. coli cells had the same thermal stability properties as LDH isolated from B. caldotenax cells.

Amino Acid Sequence

Characterization of a gene encoding a manganese peroxidase from Phanerochaete chrysosporium.

The complete nucleotide (nt) sequence of a gene (mnp-1) encoding manganese peroxidase isozyme 1 (MnP-1) (pI = 4.9) from Phanerochaete chrysosporium has been determined. The sequence of 2539 bp includes 526 bp of 5'-flanking sequence and 368 bp 3' to the poly(A) site. Comparison of cDNA and genomic sequences indicates six introns varying in size from 57-72 bp. Intron splice-junction sequences all adhere to the GT---AG rule. The positions of the introns show little similarity to the intron positions in the closely related lignin peroxidase-encoding genes. The 5' upstream region of the mnp-1 gene contains a TATAA element and three inverted CCAAT elements (ATTGG) at nt positions -81, -181, -195, and -304, respectively, relative to the start codon. In addition, the mnp-1 gene contains three putative heat-shock (HS) elements similar to the consensus C--GAA--TTC--G sequence, and two consensus metal response elements located within 500 bp upstream from the start codon. Furthermore, Northern-blot analysis demonstrates that mnp gene transcription is regulated by HS.

Amino Acid Sequence

Human spermidine synthase: cloning and primary structure.

Using a synthetic deoxyoligonucleotide mixture constructed for a tryptic peptide of the bovine enzyme as a probe, cDNA coding for the full-length subunit of spermidine synthase was isolated from a human decidual cDNA library constructed on phage lambda gt11. After subcloning into the Eco RI site of pBR322 and propagation, both strands of the insert were sequenced using a shotgun strategy. Starting from the first start codon, which was immediately preceded by a GC-rich region including four overlapping CCGCC consensus sequences, an open reading frame for a 302-amino-acid polypeptide was resolved. This peptide had an Mr of 33,827, started with methionine, and ended with serine. The identity of the isolated cDNA was confirmed by comparison of the deduced amino acid sequence with resolved sequences of the tryptic peptides of bovine spermidine synthase. The coding strand of the cDNA revealed no special regulatory or ribosome-binding signals within 82 nucleotides preceding the start codon and no polyadenylation signal within 247 nucleotides following the stop codon. The coding region, containing a 13-nucleotide repeat close to the 5' end, was longer than, and very different from, that of the bacterial counterpart. This region seems to be of retroviral origin and shows marked homology with sequences found in a variety of human, mammalian, avian, and viral genes and mRNAs. By computer analysis, the first 200 nucleotides of the 5' end of the coding strand appear able to form a very stable secondary structure with a free energy change of -157.6 kcal/mole.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence