PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Expression of the Saccharomyces cerevisiae PIS1 gene is modulated by multiple ATGs in the promoter.

The PIS1 gene encodes a key branchpoint phospholipid biosynthetic enzyme, phosphatidylinositol synthase. The PIS1 promoter contains the unusual feature of three ATG codons (ATGs1, 2, and 3) in-frame with three stop codons, located just before the authentic start codon (ATG4). Using a PIS1(promoter)-lacZ reporter expression system and site-directed mutagenesis, we investigated the role the "upstream" ATG codons play in modulation of PIS1 expression. Of the single codon changes, mutation of the first ATG (ATG1) resulted in the largest increase of the reporter gene PIS1(promoter)-lacZ expression. All combinations of altered upstream ATG codons also resulted in greater reporter expression. Reverse transcription-PCR revealed that at least some PIS1 transcripts include all AUG codons, and their synthesis is probably directed by a second TATA box upstream of the putative TATA box. These results indicate that the multiple upstream AUG codons are present in at least some PIS1 transcripts and negatively impact PIS1 expression.

Base Sequence↗

Tissue specific glucocorticoid receptor expression, a role for alternative first exon usage?

The CpG island upstream of the GR is highly structured and conserved at least in all the animal species that have been investigated. Sequence alignment of these CpG islands shows inter-species homology ranging from 64 to 99%. This 3.1kb CpG rich region upstream of the GR exon 2 encodes 5' untranslated mRNA regions. These CpG rich regions are organised into multiple first exons and, as we and others have postulated, each with its own promoter region. Alternative mRNA transcript variants are obtained by the splicing of these alternative first exons to a common acceptor site in the second exon of the GR. Exon 2 contains an in-frame stop codon immediately upstream of the ATG start codon to ensure that this 5' heterogeneity remains untranslated, and that the sequence and structure of the GR is unaffected. Tissue specific differential usage of exon 1s has been observed in a range of human tissues, and to a lesser extent in the rat and mouse. The GR expression level is tightly controlled within each tissue or cell type at baseline and upon stimulation. We suggest that no single promoter region may be capable of containing all the necessary promoter elements and yet preserve the necessary proximity to the transcription initiation site to produce such a plethora of responses. Thus we further suggest that alternative first exons each under the control of specific transcription factors control both the tissue specific GR expression and are involved in the tissue specific GR transcriptional response to stimulation. Spreading the necessary promoter elements over multiple promoter regions, each with an associated alternative transcription initiation site would appear to vastly increase the capacity for transcriptional control of GR.

5' Untranslated Regions↗

The 5'-upstream region of human programmed cell death 5 gene contains a highly active TATA-less promoter that is up-regulated by etoposide.

The PDCD5 (programmed cell death 5), a novel apoptosis related gene, is functionally associated with cell apoptosis, exhibits a ubiquitous expression pattern and is up-regulated in some types of tumor cells undergoing apoptosis. To study the transcriptional regulation of the PDCD5 gene, we have cloned 1.1 kb of its 5'-upstream region. The DNA sequencing analysis revealed a major transcriptional start site at 72 base pairs in front of the ATG translational start codon. The upstream of the transcriptional start site lacks a canonical TATA box and CAAT box. Transient transfection and luciferase assay demonstrate that this region presents extremely strong promoter activity. The 5'-deleted sequences fused to a luciferase reporter gene demonstrated that the -555/-383 region from the transcription start site is crucial for transcriptional regulation, and the luciferase reporter gene's expression significantly increased in the early stage of cell apoptosis induced by etoposide. These results imply that the PDCD5 gene may be a target gene under the control of some important apoptosis-related transcriptional factors during the cell apoptosis.

5' Flanking Region↗

Analysis of full length ADAMTS6 transcript reveals alternative splicing and a role for the 5' untranslated region in translational control.

The ADAMTS (A Disintegrin and metalloproteinase with thrombospondin-1 type repeats) family of enzymes have been implicated in turnover of extracellular matrix. We previously showed that levels of ADAMTS6 mRNA in ARPE-19 cells were markedly increased following treatment with tumour necrosis factor alpha (TNFalpha). This study shows that the ADAMTS6 transcript contains unusually large untranslated regions (UTRs) at both the 5' and 3'end. The 5'UTR contains 11 AUG codons upstream of the predicted ADAMTS6 start codon and potently inhibits translation of a downstream reporter gene. However some translation can be restored by truncating the 5'UTR from the 5'end. The 5'UTR was tested for internal ribosome entry site activity using a bicistronic luciferase reporter plasmid, but none was detected. Using the 5' and 3'UTR sequences to screen the GenBank database we identified a full length ADAMTS6 cDNA of 7262 bp. This transcript is alternatively spliced at the 3'end of the open reading frame (ORF), resulting in an extended ORF containing 3 additional tsp-1 type repeats. Quantitative RT-PCR showed that the long and short form of the ADAMTS6 ORF are co-expressed in ARPE-19 cells, but the relative levels of the two forms is modulated by TNFalpha. The region of the transcript encoding the catalytic domain also contains several notable differences compared to the previously published ADAMTS6 cDNA sequence, including a redefinition of the predicted active site motif.

5' Untranslated Regions↗

Identification and molecular characterisation of a peritrophin gene, peritrophin-48, from the myiasis fly Chrysomya bezziana.

The peritrophic matrix lines the midgut of most insects and has important roles in digestion, protection of the midgut from mechanical damage and invasion by micro-organisms. Although a few intrinsic peritrophic matrix proteins have been characterised, no direct homologues of any of these proteins have been found in other insect species, even closely related species, suggesting that the peritrophic matrix proteins show considerable sequence divergence. We now report the identification of the cDNA and genomic DNA sequences of a Chrysomya bezziana homologue of the Lucilia cuprina intrinsic peritrophic matrix protein, peritrophin-48. The gene for C. bezziana peritrophin-48 spans 1315 bp and consists of three exons (65, 560 and 690 bp, respectively) separated by introns of 566 and 72 bp. The transcriptional start site, identified by a consensus of cDNA clones and primer extension analysis, is probably located 58 bp upstream from the start codon. However, there may be multiple start sites for transcription. Two potential TATA boxes and a consensus arthropod transcription initiator are located within 134 bp of sequence upstream of the putative transcriptional start site suggesting that this region contains the gene promoter. Immuno-fluorescence localization demonstrated that C. bezziana peritrophin-48 was localised to the larval peritrophic matrix. Protein fold recognition analysis indicated structural similarities between peritrophin-48 and wheatgerm lectin. As wheatgerm lectin binds chitin, this result suggested that C. bezziana peritrophin-48 may also bind chitin, a constituent of the peritrophic matrix. Chitin binding studies with a recombinant peritrophin-48 protein confirmed that it binds chitin. A Drosophila melanogaster homologue of peritrophin-48 encoded in an EST and a genomic sequence was also identified. The pairwise percentage identities of the deduced amino acid sequences for the peritrophin-48 homologues from the three higher Dipteran species were relatively low, ranging between 32 and 42%. Despite this sequence variability, the predicted structure of these proteins, dictated by five domains, each containing a characteristic distribution of six cysteines, was strictly conserved. It is concluded that considerable sequence variation can be tolerated in this protein because of the constraints imposed on the structure of the protein by an extensive disulphide bonded framework.

Amino Acid Sequence↗

Expression of tylM genes during tylosin production: phantom promoters and enigmatic translational coupling motifs.

In the genome of Streptomyces fradiae, the three tylM genes are codirectional with the upstream gene, tylGV. Although the introduction of transcriptional blocks into the tylM genes revealed that they are normally cotranscribed, expression of tylMI still persisted (albeit at a very low level) when either of the upstream genes, tylMII or tylMIII, was disrupted. Such expression apparently resulted from transcriptional initiation at spurious sites that probably contribute insignificantly, if at all, to promote activity in the wild type. Prior to the onset of tylosin production, tylMIII is transcribed independently of tylGV from an authentic promoter buried within tylGV. This latter observation is interesting given that the TGA stop codon of tylGV overlaps the GTG start codon of tylMIII. Evidently, terminally overlapping genes are not always translationally coupled.

Anti-Bacterial Agents↗

Complete sequence of the Drosophila nonmuscle myosin heavy-chain transcript: conserved sequences in the myosin tail and differential splicing in the 5' untranslated sequence.

We have sequenced a cDNA that encodes the nonmuscle myosin heavy chain from Drosophila melanogaster. An alternatively spliced exon at the 5' end generates two distinct heavy-chain transcripts: the longer transcripts inserts an additional start codon upstream of the primary translation start site and encodes a myosin heavy chain with a 45-residue extension at its amino terminus. The remainder of the coding sequence reveals extensive homology with other conventional myosins, especially metazoan nonmuscle and smooth muscle myosin isoforms. Comparisons among available myosin heavy-chain sequences establish that characteristic differences in sequence throughout the length of both the globular myosin head and extended rod-like tail readily distinguish nonmuscle and smooth muscle myosins from striated muscle isoforms and predict a basis for their functional diversity.

Amino Acid Sequence↗

Nucleotide sequence of sporulation locus spoIIA in Bacillus subtilis.

We have determined a sequence of 2073 bp from two recombinant plasmids carrying the whole spoIIA locus from Bacillus subtilis, the expression of which is required for spore formation. The sequence contains three long open reading frames (ORFs), each of them being preceded by a ribosome binding site. These three putative proteins (mol. wts 13100, 16300 and 22200) are likely to be expressed and are probably encoded on the same mRNA. The stop codon of ORF1 overlaps with the start codon of ORF2 suggesting that there might be translational coupling between the two ORFs. Although some known promoter sequences were found, the only one upstream from the first open reading frame is about 260 bp from it.

Bacillus subtilis↗

Molecular analysis of a Clostridium butyricum NCIMB 7423 gene encoding 4-alpha-glucanotransferase and characterization of the recombinant enzyme produced in Escherichia coli.

An Escherichia coli clone was detected in a Clostridium butyricum NCIMB 7423 plasmid library capable of degrading soluble amylose. Deletion subcloning of its recombinant plasmid indicated that the gene(s) responsible for amylose degradation was localized on a 1.8 kb NspHI-Scal fragment. This region was sequenced in its entirety and shown to encompass a large ORF capable of encoding a protein with a calculated molecular mass of 57,184 Da. Although the deduced amino acid sequence showed only weak similarity with known amylases, significant sequences identity was apparent with the 4-alpha-glucano-transferase enzymes of Streptococcus pneumoniae (46.9%), potato (42.9%) and E. coli (16.2%). The clostridial gene (designated maIQ) was followed by a second ORF which, through its homology to the equivalent enzymes of E. coli and S. pneumoniae, was deduced to encode maltodextrin phosphorylase (MaIP). The translation stop codon of MaIQ overlapped the translation start codon of the putative maIP gene, suggesting that the two genes may be both transcriptionally and translationally coupled. 4-alpha-Glucanotransferase catalyses a disproportionation reaction in which single or multiple glucose units from oligosaccharides are transferred to the 4-hydroxyl group of acceptor sugars. Characterization of the recombinant C. butyricum enzyme demonstrated that glucose, maltose and maltotriose could act as acceptor, whereas of the three only maltotriose could act as donor. The enzyme therefore shares properties with the E. coli MaIQ protein, but differs significantly from the glucanotransferase of Thermotoga maritima, which is unable to use maltotriose as donor or glucose as acceptor. Physiologically, the concerted action of 4-alpha-glucanotransferase and maltodextrin phosphorylase provides C. butyricum with a mechanism of utilizing amylose/maltodextrins with little drain on cellular ATP reserves.

Amino Acid Sequence↗

Gene cluster for dissimilatory nitrite reductase (nir) from Pseudomonas aeruginosa: sequencing and identification of a locus for heme d1 biosynthesis.

The primary structure of an nir gene cluster necessary for production of active dissimilatory nitrite reductase was determined from Pseudomonas aeruginosa. Seven open reading frames, designated nirDLGHJEN, were identified downstream of the previously reported nirSMCF genes. From nirS through nirN, the stop codon of one gene and the start codon of the next gene were closely linked, suggesting that nirSMCFDLGHJEN are expressed from a promoter which regulates the transcription of nirSM. The amino acid sequences deduced from the nirDLGH genes were homologous to each other. A gene, designated nirJ, which encodes a protein of 387 amino acids, showed partial identity with each of the nirDLGH genes. The nirE gene encodes a protein of 279 amino acids homologous to S-adenosyl-L-methionine:uroporphyrinogen III methyltransferase from other bacterial strains. In addition, NirE shows 21.0% identity with NirF in the N-terminal 100-amino-acid residues. A gene, designated nirN, encodes a protein of 493 amino acids with a conserved binding motif for heme c (CXXCH) and a typical N-terminal signal sequence for membrane translocation. The derived NirN protein shows 23.9% identity with nitrite reductase (NirS). Insertional mutation and complementation analyses showed that all of the nirFDLGHJE genes were necessary for the biosynthesis of heme d1.

Amino Acid Sequence↗

Codon replacement in the PGK1 gene of Saccharomyces cerevisiae: experimental approach to study the role of biased codon usage in gene expression.

The coding sequences of genes in the yeast Saccharomyces cerevisiae show a preference for 25 of the 61 possible coding triplets. The degree of this biased codon usage in each gene is positively correlated to its expression level. Highly expressed genes use these 25 major codons almost exclusively. As an experimental approach to studying biased codon usage and its possible role in modulating gene expression, systematic codon replacements were carried out in the highly expressed PGK1 gene. The expression of phosphoglycerate kinase (PGK) was studied both on a high-copy-number plasmid and as a single copy gene integrated into the chromosome. Replacing an increasing number (up to 39% of all codons) of major codons with synonymous minor ones at the 5' end of the coding sequence caused a dramatic decline of the expression level. The PGK protein levels dropped 10-fold. The steady-state mRNA levels also declined, but to a lesser extent (threefold). Our data indicate that this reduction in mRNA levels was due to destabilization caused by impaired translation elongation at the minor codons. By preventing translation of the PGK mRNAs by the introduction of a stop codon 3' and adjacent to the start codon, the steady-state mRNA levels decreased dramatically. We conclude that efficient mRNA translation is required for maintaining mRNA stability in S. cerevisiae. These findings have important implications for the study of the expression of heterologous genes in yeast cells.

Amino Acid Sequence↗

Cloning and characterization of the murine coagulation factor X gene.

The gene encoding murine coagulation factor X (fX) was isolated and characterized from a lamdaFIX II library generated from murine genomic DNA. The 20130 bp sequence contains 18049 nucleotides that extend from the initiating methionine to the polyadenylation site. 1056 nucleotides 5' of the start codon were determined and contain putative start sites for the FX mRNA as well as sites for binding of putative transcription factors. The sequence extends 1024 3' of the polyadenylation site. The gene contains 8 exons and 7 introns which were determined by comparing the mouse FX cDNA and gene sequences. The exonic structure of the gene is similar to that of the other mammalian vitamin K-dependent serine proteases of the coagulation system. These include an exon encoding the prepropepetide, the gla-domain, a short helical stack, two exons for the two EGF domains, the activation pepetide, and two exons encoding the serine protease domain. The 5' sequence of the mouse FX gene overlaps with the 3' region of the FVII gene indicating that the murine FVII and FX gene are arranged in a head to tail arrangement as they are in humans.

Animals↗

Translational pathophysiology: a novel molecular mechanism of human disease.

In higher eukaryotes, the expression of about 1 gene in 10 is strongly regulated at the level of messenger RNA (mRNA) translation into protein. Negative regulatory effects are often mediated by the 5'-untranslated region (5'-UTR) and rely on the fact that the 40S ribosomal subunit first binds to the cap structure at the 5'-end of mRNA and then scans for the first AUG codon. Self-complementary sequences can form stable stem-loop structures that interfere with the assembly of the preinitiation complex and/or ribosomal scanning. These stem loops can be further stabilized by the interaction with RNA-binding proteins, as in the case of ferritin. The presence of AUG codons located upstream of the physiological start site can inhibit translation by causing premature initiation and thereby preventing the ribosome from reaching the physiological start codon, as in the case of thrombopoietin (TPO). Recently, mutations that cause disease through increased or decreased efficiency of mRNA translation have been discovered, defining translational pathophysiology as a novel mechanism of human disease. Hereditary hyperferritinemia/cataract syndrome arises from various point mutations or deletions within a protein-binding sequence in the 5'-UTR of the L-ferritin mRNA. Each unique mutation confers a characteristic degree of hyperferritinemia and severity of cataract in affected individuals. Hereditary thrombocythemia (sometimes called familial essential thrombocythemia or familial thrombocytosis) can be caused by mutations in upstream AUG codons in the 5'-UTR of the TPO mRNA that normally function as translational repressors. Their inactivation leads to excessive production of TPO and elevated platelet counts. Finally, predisposition to melanoma may originate from mutations that create translational repressors in the 5'-UTR of the cyclin-dependent kinase inhibitor-2A gene.

5' Untranslated Regions↗

Structure and regulation of the glpFK operon encoding glycerol diffusion facilitator and glycerol kinase of Escherichia coli K-12.

The glpFK operon maps near minute 88 on the linkage map of Escherichia coli K-12 with glpF promoter proximal. The glpF gene encodes a cytoplasmic membrane protein which facilitates the diffusion of glycerol into the cell. The glpK gene encodes glycerol kinase. In the present work, the nucleotide sequence of the 5'-end of the operon, including the control region, the glpF gene, and part of the glpK gene, was determined. The facilitator was predicted to contain 281 amino acids with a calculated molecular weight of 29,780. It is a highly hydrophobic protein with a minimum of six potential transmembrane alpha helices. The transcription start site for the glpFK operon was located 71 base pairs upstream from the proposed translation start codon for glpF. Preceding the transcription start site were sequences similar to the -10 and -35 consensus sequences for bacterial promoters. Binding sites for the cAMP-cAMP receptor protein (CRP) complex and the glp repressor were identified by DNase I footprinting. The region protected by the cAMP.CRP complex contained tandem sequences resembling the consensus sequence for CRP binding. The CRP sites were centered at 37.5 and 60.5 base pairs upstream of the start of transcription. The glp repressor protected an extensive area (-89 to -7 relative to the start point of transcription), sufficient for the binding of four repressor tetramers. Two additional binding sites for the repressor were identified within the glpK coding region. The DNA containing these two operators synergistically increased the apparent affinity of glp repressor for DNA fragments containing the four operators in the promoter region of the glpFK operon. With this study, a total of 13 operators for the glp regulon have been characterized. Comparison of these operators revealed the consensus 5'-WATGTTCGWT-3' for the operator half-site (W = A or T). The relative affinity of the glp repressor for the various glp operators was assessed in vivo using a promoter-probe vector. The relative apparent affinity of the control regions for glp repressor was glpFK greater than glpD greater than glpACB greater than glpTQ. The degree of catabolite repression for each of the operons was assessed using a similar system. In this case, the relative sensitivity of the glp operons to catabolite repression was glpTQ greater than glpFK greater than glpACB greater than glpD.

Amino Acid Sequence↗

Structure and organization of hip, an operon that affects lethality due to inhibition of peptidoglycan or DNA synthesis.

High-frequency persistence to the lethal effects of inhibition of either DNA or peptidoglycan synthesis, the Hip phenotype, results from mutations at the hip locus of Escherichia coli K-12. The nucleotide sequence of DNA fragments which complement these mutations revealed an operon consisting of a possible regulatory region, including sequences with modest homology to an E. coli promoter, and two open reading frames which are translated both in vitro and in vivo. The stop codon of a 264-bp open reading frame, hipB, and the start codon of a 1,320-bp open reading frame, hipA, share an adenine residue. Assays of promoter strength, the location of the probable promoter with respect to the start of transcription, and codon usage all indicate that hipB and hipA are weakly expressed genes. The activity of the promoter is impaired by an adjacent downstream sequence which includes the coding region of hipB. The impairment is partially relieved by insertion of a premature translation termination signal within the coding region of hipB, suggesting involvement of the HipB protein in the regulation of this promoter. The arrangement of hipB and hipA within the operon and the toxicity of hipA for strains defective in or lacking hipB suggest an important interaction between the products of these genes.

Amino Acid Sequence↗

The chloroplast infA gene with a functional UUG initiation codon.

All chloroplast genes reported so far possess ATG start codons and sometimes GTGs as an exception. Sequence alignments suggested that the chloroplast infA gene encoding initiation factor 1 in the green alga Chlorella vulgaris has TTG as a putative initiation codon. This gene was shown to be transcribed by RT-PCR analysis. The infA mRNA was translated accurately from the UUG codon in a tobacco chloroplast in vitro translation system. Mutation of the UUG codon to AUG increased translation efficiency approximately 300-fold. These results indicate that the UUG is functional for accurate translation initiation of Chlorella infA mRNA but it is an inefficient initiation codon.

Algal Proteins↗

Nucleotide sequence of an immediate-early frog virus 3 gene.

We have used "gene walking" with synthetic oligonucleotides and M13 dideoxynucleotide sequencing techniques to obtain the complete coding and flanking sequences of the gene encoding a major immediate-early RNA (molecular weight, 169,000) of frog virus 3. R-loop mapping of the cloned XbaI K fragment of frog virus 3 DNA with immediate-early RNA from infected cells showed that an RNA of approximately 500 to 600 nucleotides (the right size to code for the immediate-early viral 18-kilodalton protein of unknown function) hybridized to a region within 100 base pairs of one end of the XbaI K fragment; no evidence for splicing was observed in the electron microscope or by single-strand nuclease analysis. Further restriction mapping narrowed the location of the gene to the XbaI end of a 2-kilobase-pair XbaI-Bg/II fragment, which was bidirectionally subcloned into the bacteriophage pair mp10 and mp11 for sequencing. Mung bean nuclease mapping was used to identify both the 5' and the 3' ends of the mRNA. The 5' end mapped within an AT-rich region 19 base pairs upstream from two in-phase AUG start codons that were immediately followed by an open reading frame of 157 amino acids. Another AT-rich sequence was found at -29 base pairs from the 5' end of the mRNA start site; this sequence may function as a TATA box. The 3' end of the message displayed considerable microheterogeneity, but clearly terminated within a third AT-rich region 50 to 60 base pairs from the translation stop codon. The eucaryotic polyadenylic acid addition signal (AATAAA) was not present, a finding to be expected since frog virus 3 mRNA is not polyadenylated. Both the single-stranded mp10 clone of the XbaI-Bg/II fragment and a 15-base oligonucleotide complementary to the region flanking the two AUG translation start codons inhibited translation of the immediate-early 18-kilodalton protein in vitro, confirming the identity of the sequenced gene. As the regulatory sequences of this gene did not resemble those of known eucaryotic genes or of the cytoplasmic vaccinia virus, we conclude that frog virus 3 has evolved unique signals for the initiation and termination of transcription.

Base Sequence↗

cis- and trans-acting suppressors of a translation initiation defect at the cyc1 locus of Saccharomyces cerevisiae.

The cyc1-362 mutant of Saccharomyces cerevisiae is deficient in iso-1-cytochrome c as a consequence of an aberrant ATG codon that initiates a short open reading frame (uORF) in the cyc1 transcribed leader region. We have isolated and characterized functional revertants of cyc1-362 in an effort to define cis- and trans-acting factors that can suppress the effect of the uORF. Genetic and DNA sequence analyses have defined three classes of revertants: (i) those that acquired point mutations in the upstream ATG (uATG), restoring iso-1-cytochrome c to its normal level; (ii) substitution of the normal A residue at position -1 relative to the uATG by either C or T, enhancing iso-1-cytochrome c production from approximately 2% to 6% (C) or 10% (T) of normal, indicating that the nucleotide immediately preceding the initiator codon can affect the efficiency of AUG start codon recognition and that purines are preferred over pyrimidines at this site; and (iii) extragenic suppressors that enhance iso-1-cytochrome c expression to 10-40% of normal while retaining the uATG. These suppressors are represented by five different genes, designated sua1-sua4 and sua6. In contrast to the previously described sua7 and sua8 suppressors, they do not compensate for the uATG by affecting cyc1 transcription start site selection. Potential suppressor mechanisms are discussed.

Amino Acid Sequence↗