PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Nucleotide sequence of an mRNA transcribed in latent growth-transforming virus infection indicates that it may encode a membrane protein.

The most abundant Epstein-Barr virus mRNA in a latently infected cell line, IB4, established by in vitro growth transformation with virus, was a 2,8-kilobase RNA encoded by largely unique DNA near the right end of the genome. The RNA was transcribed from right to left, and two introns were spliced out. This region of the genome was sequenced, and the exons of the RNA were identified by S1 analysis of DNA-RNA hybrids and primer extension. The first start codon in the RNA was 40 nucleotides from its 5' end. Beginning with the start codon, there was a 1,158-nucleotide open reading frame which crossed both introns. The important characteristics of the translated protein were as follows. (i) The amino terminus was highly charged and not suggestive of a leader sequence. (ii) There were six markedly hydrophobic alpha-helical domains, each having 21 amino acids and connected by 5 to 7 amino acid segments predicted to be reverse turns. (iii) The carboxy-terminal 200 amino acids were markedly acidic, containing 6 glutamic and 37 aspartic acids. The hydrophobic region is predicted to form six membrane-spanning regions, leaving the short charged amino terminus and long acidic carboxy terminus on the inside of the membrane. This protein could be responsible for the new antigen detected in the plasma membrane of Epstein-Barr virus-transformed cells, lymphocyte-determined membrane antigen. There were two other open reading frames in the RNA.

Amino Acid Sequence↗

Conservation in the 5' flanking sequences of transcribed members of the Caenorhabditis elegans major sperm protein gene family.

The major sperm proteins (MSPs) are encoded in the Caenorhabditis genome by a multigene family with more than 50 genes dispersed in small clusters at three chromosomal loci. In spite of their dispersed locations, all of the MSP genes appear to be expressed at the same time exclusively in the testis, indicating co-ordinate temporal and spatial regulation of these dispersed genes. Many of the MSP genes must be transcribed, because RNA hybridization with gene-specific probes showed that individual genes each contribute less than 3% to the total poly(A)+ RNA, and 13 out of 14 sequenced cDNAs came from different genes. Primer extension assays from MSP mRNA showed that most of the MSP mRNAs must be initiated at position -35 from the translation start codon. Extensive similarity was found in the first 100 nucleotides of genomic sequence flanking the start codons of ten MSP genes from different chromosomal locations. All MSP genes contained a consensus ribosome binding site, a consensus TATA homology 27 nucleotides distal to the site of mRNA initiation, and ten highly conserved nucleotides adjacent to the site of initiation. All the MSP genes contained the sequence AGATCT located approximately 65 nucleotides upstream from the transcriptional start, but little or no similarity was found more distal to this. Some of these conserved sequences may be cis-acting control elements that ensure the cell and temporal specificity of transcription of these co-ordinately regulated genes.

Animals↗

Characterization of the 5' flanking region of the gene encoding rat liver glycogen phosphorylase.

A genomic region encompassing 800 bp of the promoter-regulatory region and exon 1 of the gene (LGP) encoding rat liver glycogen phosphorylase has been isolated and characterized. Transcripts of the LGP gene initiate predominantly within an 8-bp region 48-bp upstream from the start codon. Additional transcripts were detected that initiate as far as 95 bp upstream from the start codon. To identify cis-acting sequences involved in regulating transcription, HepG2 cells were transfected with vectors containing serial deletions of the promoter-regulatory region of LGP ligated to the cat reporter gene. Two upstream regions were found to enhance transcription. One of these regions contains an alternating purine-pyrimidine sequence. LGP, which lacks a consensus TATA sequence, is like TATA-less and CAAT-less housekeeping genes in that it contains G + C-rich domains upstream from multiple transcription start points. Nuclear proteins from adult rat tissues bound in a tissue-specific fashion to one of these G + C-rich regions.

Amino Acid Sequence↗

Herpes simplex virus type 1 (HSV-1) uracil-DNA glycosylase: functional expression in Escherichia coli, biochemical characterization, and selective inhibition by 6-(p-n-octylanilino)uracil.

The Herpes simplex virus type 1 (HSV-1) uracil-DNA glycosylase (UDG) is encoded by the UL2 gene. The translation from the first putative start codon of UL2 predicts a polypeptide of 334 residues, while the translation from the second start codon predicts a polypeptide of 244 residues. We have cloned and expressed the two forms of UDG, by means of the prokaryotic expression vector pMAL-c2, and both of them were enzymatically active. Furthermore, the enzymatic properties of the recombinant UDGs and of the enzyme purified from HSV-1-infected cells were similar. The two UDG polypeptides have molecular weights of 27 and 37 kDa, respectively. The 37-kDa form of recombinant UDG is consistent with the reported molecular mass of 37 kDa for the enzyme purified from HSV-1-infected cells. Both recombinant UDGs were as sensitive as UDG purified from HSV-1-infected cells to 6-(p-n-octylanilino)uracil, the most potent of a series of uracil analogs that inhibit the viral enzyme.

Base Sequence↗

cDNA sequence and expression of a phosphoenolpyruvate carboxylase gene from soybean.

A full-length cDNA encoding a subunit of phosphoenolpyruvate carboxylase (PEPC) was isolated from a developing seed expression library of the C3 plant Glycine max. The corresponding mRNA is present at similar levels in leaf, stem, root and developing seed. Two potential start codons exist, and the activity of protein initiated from the first such codon could be subject to regulation by protein kinase. Sequence comparison shows a similar upstream start codon in the case of the Ppc2 gene from Mesembryanthemum crystallinum, previously assumed to lack the sequences necessary for phosphorylation. The soybean encoded protein tends to resemble other 'C3-type' PEPC proteins more closely than those implicated in C4 or crassulacean acid metabolism.

Amino Acid Sequence↗

The evolutionarily conserved eukaryotic arginine attenuator peptide regulates the movement of ribosomes that have translated it.

Translation of the upstream open reading frame (uORF) in the 5' leader segment of the Neurospora crassa arg-2 mRNA causes reduced initiation at a downstream start codon when arginine is plentiful. Previous examination of this translational attenuation mechanism using a primer-extension inhibition (toeprint) assay in a homologous N. crassa cell-free translation system showed that arginine causes ribosomes to stall at the uORF termination codon. This stalling apparently regulates translation by preventing trailing scanning ribosomes from reaching the downstream start codon. Here we provide evidence that neither the distance between the uORF stop codon and the downstream initiation codon nor the nature of the stop codon used to terminate translation of the uORF-encoded arginine attenuator peptide (AAP) is important for regulation. Furthermore, translation of the AAP coding region regulates synthesis of the firefly luciferase polypeptide when it is fused directly at the N terminus of that polypeptide. In this case, the elongating ribosome stalls in response to Arg soon after it translates the AAP coding region. Regulation by this eukaryotic leader peptide thus appears to be exerted through a novel mechanism of cis-acting translational control.

Amino Acid Sequence↗

PKD1 upstream open reading frames affect Polycystin-1 expression and polycystic kidney disease phenotypes.

Autosomal dominant polycystic kidney disease (ADPKD) accounts for 5%-10% of prevalent end-stage kidney failure (ESKD). ADPKD cysts result from a loss of sufficient functional expression of PKD1/Polycystin-1 (PC1) in approximately 80% of families. Kidney disease severity correlates with the extent to which PC1 dosage is reduced below a critical level, and evidence suggests therapeutic benefit from increasing PC1 expression in these conditions. Upstream open reading frame (uORF) translation can reduce translation of a protein's coding sequence. Ribosome profiling data and bioinformatic predictions suggested the presence of conserved PKD1 uORFs, so we sought to explore their biological role. We generated luciferase reporters and two humanized PKD1 5' UTR mouse models with or without single nucleotide edits removing uORF start codons (ΔuORF) to define active uORFs and test their impact on PC1 translation. PKD1 uORF start codons can robustly initiate translation, and ΔuORF conveys a 2-4 fold increase in PC1 protein expression and resultant prevention of kidney cysts in Dnajb11 as well as in Pkd1 missense models. PKD1 uORF1-blocking steric antisense oligonucleotides (ASOs) substantially increase PC1 expression in vitro. PKD1 uORFs play an important role in the low basal expression of WT PKD1, and their inhibition represents an opportunity to therapeutically increase PC1 translation in polycystic kidney and liver disease resulting from reduced dosage of PC1.

Animals↗

Characterization of binding sequences for butyrolactone autoregulator receptors in streptomycetes.

BarA of Streptomyces virginiae is a specific receptor protein for a member of butyrolactone autoregulators which binds to an upstream region of target genes to control transcription, leading to the production of the antibiotic virginiamycin M(1) and S. BarA-binding DNA sequences (BarA-responsive elements [BAREs]), to which BarA binds for transcriptional control, were restricted to 26 to 29-nucleotide (nt) sequences on barA and barB upstream regions by the surface plasmon resonance technique, gel shift assay, and DNase I footprint analysis. Two BAREs (BARE-1 and BARE-2) on the barB upstream region were located 57 to 29 bp (BARE-1) and 268 to 241 bp (BARE-2) upstream from the barB translational start codon. The BARE located on the barA upstream region (BARE-3) was found 101 to 76 bp upstream of the barA start codon. High-resolution S1 nuclease mapping analysis revealed that BARE-1 covered the barB transcription start site and BARE-3 covered an autoregulator-dependent transcription start site of the barA gene. Deletion and mutation analysis of BARE-2 demonstrated that at least a 19-nt sequence was required for sufficient BarA binding, and A or T residues at the edge as well as internal conserved nucleotides were indispensable. The identified binding sequences for autoregulator receptor proteins were found to be highly conserved among Streptomyces species.

Autoreceptors↗

Three variant introns of the same general class in the mitochondrial gene for cytochrome oxidase subunit 1 in Aspergillus nidulans.

The oxiA gene of Aspergillus nidulans, coding for cytochrome oxidase subunit 1, is shown by DNA sequencing to contain three introns. An AUG start codon is not present at the beginning of the sequence, suggesting that either another codon, possibly the four base codon AUGA, is used for initiation or there is a further short intron between the true start codon and the beginning of the recognisable coding region. The second and third introns have long open reading frames, which could code for maturase proteins. The lack of conservation of amino acid sequence in the putative region of proteolytic cleavage for maturase formation suggests that the first conserved decapeptide may act as the recognition signal for protein processing. The third intron is remarkably (70%) homologous to the second intron of the cytochrome oxidase subunit 1 gene of Schizosaccharomyces pombe and both are located in exactly the same position. The third Aspergillus intron has an in-frame insertion of a 37-bp GC-rich DNA sequence which is now flanked by a 5-bp repeat, a well-known feature of transposable elements. All three introns in the oxiA gene have a 'core' RNA secondary structure found in a class of introns fitting the RNA splicing model of Davies et al. (1982). This core RNA structure may play a catalytic as well as a structural role in intron splicing. A sequence within the intron could act as a guide to align the splice sites of two of the introns in accordance with the model of Davies et al.

Aspergillus nidulans↗

In vitro transcriptional and translational block of the bcl-2 gene operated by peptide nucleic acid.

The antisense and antigene activity of peptide nucleic acid (PNA) targeted to the human B-cell lymphoma (bcl)-2 gene was evaluated in vitro. Several PNAs complementary to different sequences of bcl-2, including the start codon and the 5'-untranslated region (5'-UTR), were tested. One PNA directed against the AUG start codon and another recognizing the 5'-UTR were able to specifically reduce Bcl-2 protein synthesis in a cell-free system; however, only partial inhibition (80 and 54%, respectively) was obtained when they were used singularly. Complete translation block was obtained with the simultaneous presence of both PNAs. A triplex-forming bis-PNA was targeted to a homopurine sequence on the coding strand of the bcl-2 cDNA. In an in vitro transcription assay this PNA specifically inhibited the transcription of bcl-2 at concentrations as low as 300 nM, with the concomitant appearance of a truncated 200-base-long product. These results demonstrate the ability of PNA to selectively modulate both translation and transcription of bcl-2 in vitro and suggest its potential use as an antisense and an antigene agent in order to downregulate bcl-2 expression in tumors.

Dose-Response Relationship, Drug↗

Paramecium mitochondrial DNA sequences and RNA transcripts for cytochrome oxidase subunit I, URF1, and three ORFs adjacent to the replication origin.

A 2-kb region adjacent to the replication origin (ori) and a 3-kb region located between the small and large ribosomal RNAs of Paramecium mitochondrial (mt) DNA have been sequenced and the locations of their transcripts determined. The ori segment contains four transcripts, some of which are overlapping, which encode a known protein and two other open reading frames. The other segment encodes, on separate transcripts, the cytochrome c oxidase subunit one gene (COI) and the URF1 gene (ND1) common to most mt genomes. All these genes have the same orientation and do not contain introns. The COI gene is the most divergent of those known and has an internal 108 amino acid 'insert' not found in COI genes from other organisms. With these data it is possible to define a probable Paramecium mt genetic code. With the exception that TGA codes for tryptophan and the use of different start codons, Paramecium mtDNA appears to follow the universal code. GTA possibly can be used as a start codon.

Amino Acid Sequence↗

A genomic clone containing the promoter for the gene encoding the human lysosomal enzyme, alpha-galactosidase A.

We have isolated and characterized a human genomic clone for a lysosomal enzyme gene. The start point of transcription was identified using primer extension of poly(A)+ mRNA. This genomic clone is specific for human alpha-galactosidase A, and it includes sequences for the promoter, complete signal peptide, first exon, and part of the first intron. Direct and inverted repeat elements of 10, 11, 16, 19, and 22 nucleotides (nt) flank the promoter site. A (GA)n repeat element of approx. 60 nt with strong homology to similar elements identified in several species is located upstream from the promoter. A GGGCGG site specific for DNA-binding protein Sp1 is located near a CAAT box, and the CCGCCC inverted repeat of the Sp1 binding sequence is located by the TATA box. The sequence immediately flanking the ATG start codon of the human alpha-galactosidase A is highly homologous to sequences flanking the ATG start codons of the other human lysosomal hydrolases for which sequence information is available (beta-glucocerebrosidase, cathepsin B, cathepsin D, and beta-hexosaminidase alpha chain), but not for any of the other 133 human signal peptides examined. Our analysis also reveals that conversion of the propeptide to the mature enzyme involves cleavage of a C-terminal rather than an N-terminal fragment. This information about the normal alpha-galactosidase A gene will be useful for comparison to data obtained from patients with Fabry disease, who are characterized by a deficiency of this enzyme. This is the first genomic clone described to date for any lysosomal enzyme, and it establishes a reference for future analyses of the molecular events that mediate the expression of this important class of enzymes.

Amino Acid Sequence↗

The genomic structure and chromosomal location of the human TR2 orphan receptor, a member of the steroid receptor superfamily.

Human TR2 orphan receptor, isolated from the testis and prostate, is a member of the steroid/thyroid hormone receptor superfamily. With the screening of a human genomic library and the combination of primer walking and PCR sequencing, we found that the entire TR2 orphan receptor gene coding region and 5'-untranslated region feature 13 introns and 14 exons, and that the consensus splice sequences (GT-AG) are present in all intron-exon boundaries. Within the region that codes for the DNA binding domain, TR2 orphan receptor gene has a distinct intron-exon junction. Whereas all other known steroid receptors have one splice site that separates their first and second zinc fingers in the DNA binding domain, TR2 orphan receptor has a rare splice site located in the middle of its first zinc finger. The identification of specific junction sequences for potential alternative splicing sites helps to explain the existence of multiple forms of TR2 orphan receptor cDNA (TR2-5, 7, 9, 11). The S1 nuclease protection assay for TR2 message revealed that there are multiple transcription initiations, and that the major cap site surrounded by an initiator-like sequence is located at the 104th nucleotide upstream from the translation start codon. Sequence analysis of a 2.7-kb DNA fragment upstream of the TR2 orphan receptor translation start codon unveiled several potential cis-acting elements, such as AP-1, HNF-5, GATA1 binding sites, and GC boxes. Using fluorescence in situ hybridization combined with a high-resolution G-banding technique, we found that the TR2 orphan receptor gene was mapped to human chromosome 12 at band q22, whereas the structurally closely related TR4 orphan receptor gene was mapped to human chromosome 3 at band q24.3.

Amino Acid Sequence↗

Structural organization of the genes for murine and human leukemia inhibitory factor. Evolutionary conservation of coding and non-coding regions.

Leukemia inhibitory factor, LIF, is a glycoprotein with multiple activities in both the adult and the embryo. LIF appears to be encoded by a unique gene in both mouse and man, although the 3'-untranslated region of the mouse LIF gene gives a complex hybridization pattern on Southern blots. The complete nucleotide sequences of both the murine and human LIF genes and their flanking regions (8.7 and 7.6 kilobase pairs, respectively) were determined and compared. Both genes comprise three exons, two introns and an unusually long 3'-untranslated region (3.2 kilobase pairs), specificying a mRNA of approximately 4.1 kilobases. Two start sites of LIF-transcription were determined, by S1-nuclease protection and by a novel approach involving the polymerase chain reaction. S1-nuclease protection revealed a start site 60-64 base pairs upstream of the translational start codon and immediately downstream of a TATA box (TATATAAAT). The PCR approach identified a second transcriptional start site 160 base pairs 5' of the start codon and adjacent to a "TATA-like" element (CATAATTT). A comparison of the murine and human LIF gene sequences revealed a high degree of conservation in the coding regions and in segments of the untranslated and flanking regions. Seven segments displaying greater than 75% homology were identified, with the 5' and 3' ends of the transcription unit revealing the highest degree of homology. These conserved regions represents potential cis-acting control elements.

Amino Acid Sequence↗

The gene for ribosomal protein L7a-1 in Schizosaccharomyces pombe contains an intron after the initiation codon.

The gene encoding ribosomal protein L7a-1 in the fission yeast Schizosaccharomyces pombe is identified by the similarity of its open reading frame to the respective gene in Saccharomyces cerevisiae. The L7a gene is encoded in two different genomic environments as frequently found for ribosomal protein genes in this organism. One of these genes, L75a-1, is located on chromosome 2. The two consensus promoter elements homol D and homol E are both identified upstream of the start codon of this gene. The ATG start codon is separated from the main reading frame by an intron of 66 nucleotides.

Amino Acid Sequence↗

A multifactor complex of eukaryotic initiation factors, eIF1, eIF2, eIF3, eIF5, and initiator tRNA(Met) is an important translation initiation intermediate in vivo.

Translation initiation factor 2 (eIF2) bound to GTP transfers the initiator methionyl tRNA to the 40S ribosomal subunit. The eIF5 stimulates GTP hydrolysis by the eIF2/GTP/Met-tRNA(i)(Met) ternary complex on base-pairing between Met-tRNA(i)(Met) and the start codon. The eIF2, eIF5, and eIF1 all have been implicated in stringent selection of AUG as the start codon. The eIF3 binds to the 40S ribosome and promotes recruitment of the ternary complex; however, physical contact between eIF3 and eIF2 has not been observed. We show that yeast eIF5 can bridge interaction in vitro between eIF3 and eIF2 by binding simultaneously to the amino terminus of eIF3 subunit NIP1 and the amino-terminal half of eIF2beta, dependent on a conserved bipartite motif in the carboxyl terminus of eIF5. Additionally, the amino terminus of NIP1 can bind concurrently to eIF5 and eIF1. These findings suggest the occurrence of an eIF3/eIF1/eIF5/eIF2 multifactor complex, which was observed in cell extracts free of 40S ribosomes and found to contain stoichiometric amounts of tRNA(i)(Met). The multifactor complex was disrupted by the tif5-7A mutation in the bipartite motif of eIF5. Importantly, the tif5-7A mutant is temperature sensitive and displayed a substantial reduction in translation initiation at the restrictive temperature. We propose that the multifactor complex is an important intermediate in translation initiation in vivo.

Binding Sites↗

Amino acid substitutions in membrane-spanning domains of Hol1, a member of the major facilitator superfamily of transporters, confer nonselective cation uptake in Saccharomyces cerevisiae.

Selection for the ability of Saccharomyces cerevisiae cells to take up histidinol, the biosynthetic precursor to histidine, results in dominant mutations at HOL1. The DNA sequence of HOL1 was determined, and it predicts a 65-kDa protein related to the major facilitator family (drug resistance subfamily) of putative transport proteins. Two classes of mutations were obtained: (i) those that altered the coding region of HOL1, conferring the ability to take up histidinol; and (ii) cis-acting mutations (selected in a mutant HOL1-1 background) that increased expression of the Hol1 protein. The ability to transport histidinol and other cations was conferred by single amino acid substitutions at any of three sites located within putative membrane-spanning domains of the transporter. These mutations resulted in the conversion of a small hydrophobic amino acid codon to a phenylalanine codon. Selection for spontaneous mutations that increase histidinol uptake by such HOL1 mutants resulted in mutations that abolish the putative start codon of a six-codon open reading frame located approximately 171 nucleotides downstream of the transcription initiation site and 213 nucleotides upstream of the coding region of HOL1. This single small upstream open reading frame (uORF) confers translational repression upon HOL1; genetic disruption of the putative start codon of the uORF results in a 5- to 10-fold increase in steady-state amounts of Hol1 protein without significantly affecting the level of HOL1 mRNA expression.

Alleles↗

On biased distribution of introns in various eukaryotes.

We conducted comprehensive analyses on intron positions in the Mus musculus genome by comparing genomic sequences in the GenBank database and cDNA sequences in the mouse cDNA library recently developed by Riken Genomic Sciences Center. Our results confirm that introns have a tendency to be located toward the 5' end of the gene. The same type of analysis was conducted in the coding region of seven eukaryotes (Saccharomyces cerevisiae, Plasmodium falciparum, Caenorhabditis elegans, Drosophila melanogaster, M. musculus, Homo sapiens, Arabidopsis thaliana). Introns in genes with a single intron have a locational bias toward the 5' end in all species except A. thaliana. We also measured the distance from the start codon to the position of the intron, and found that single introns prefer the location immediately after the start codon in S. cerevisiae and P. falciparum. We discuss three possible explanations for these findings: (1) they are the consequence of intron loss by reverse-transcriptase; (2) they are necessary to accommodate the function; and (3) they are concerned with the mechanism of pre-mRNA splicing.

Animals↗