PubMed HealthSearch

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The reverse transcriptase gene of cauliflower mosaic virus is translated separately from the capsid gene.

Cauliflower mosaic virus (CaMV) possesses start codons at the beginning of its reverse transcriptase (RT) gene (ORF V) suggesting that, unlike in retroviruses and retrotransposons, it is translated independently from the capsid gene (ORF IV). To test this hypothesis a mutational analysis of the CaMV ORF IV/V overlapping region was performed. Mutants in which both ORFs are separated by stop codons in all three reading frames are viable and stable, while mutations affecting the first two AUG codons of ORF V are either lethal or unstable, giving rise to true and second site reversions. Mutants in which the AUG codons were replaced by ACG or AAG reverted only slowly and ACG could direct the synthesis of small amounts of reporter enzyme in transfected plant protoplasts, showing that this codon can act in plant cells as a weak start codon. CaMV has apparently developed a strategy for translation of the RT gene different from that in retroviruses and retrotransposons, but similar to that of hepadnaviruses, another group of pararetroviruses. The separate translation of the RT gene as a common feature of pararetroviruses might reflect the difference in their life cycle in comparison with retroviruses.

Base Sequence

Bacteriophage T4 gene 21 encodes two proteins essential for phage maturation.

The T4 prohead protease (T4 PPase) is the key enzyme in the morphopoietic pathway of the T4 phage head. It is responsible for the proteolytic processing of all head proteins allowing protein rearrangement and head expansion. To study its biochemistry and gene regulation, T4 gene 21 was cloned into an expression vector under the control of the inducible tac promoter. Two proteins of apparent molecular weights of 21.5 and 27.5 kDa were detected after induction. These proteins are synthesized using two different start codons in the same reading frame. Destruction of either start codon resulted in the loss of the respective protein. Complementation experiments with bacteriophage T4 21(-)-infected cells showed that both proteins are functional in vivo and essential for T4 phage assembly.

Base Sequence

Purification and characterization of an endoglucanase from Streptomyces lividans 66 and DNA sequence of the gene.

The endoglucanase isolated from culture filtrates of Streptomyces lividans IAF74 was shown to have an Mr of 46,000 and a pI of 3.3. The specific enzyme activity of 539 IU/mg, determined by the reducing assay method on carboxymethyl cellulose, is among the highest reported in the literature. The cellulase showed typical endo-type activity when reacting on oligocellodextrins. Optimal enzyme activity was obtained at 50 degrees C and pH 5.5. The kinetic constants for this endoglucanase, determined with carboxymethyl cellulose as the substrate, were a Vmax of 24.9 IU/mg of enzyme and a Km of 4.2 mg/ml. Activity was found against neither methylumbelliferyl- nor p-nitrophenyl-cellobiopyranoside nor with xylan. The DNA sequence contains one possible reading frame validated by the N terminus of the mature purified protein. However, neither ATG nor GTG starting codons were identified near the ribosome-binding site. A putative TTG codon was found as a good candidate for the start codon. Comparison of the primary amino acid sequence of the endoglucanase of S. lividans revealed that the N terminus contains a bacterial cellulose-binding domain. The catalytic domain at the C terminus showed similarity to endoglucanases from a Bacillus sp. Thus, the endoglucanase CelA belongs to family A of cellulases as described before (N. R. Gilkes, B. Henrissat, D. G. Kilburn, R. C. Miller, Jr., and R. A. J. Warren, Microbiol. Rev. 55:303-315, 1991.

Amino Acid Sequence

Cloning and nucleotide sequence of the Bacillus subtilis hom gene coding for homoserine dehydrogenase. Structural and evolutionary relationships with Escherichia coli aspartokinases-homoserine dehydrogenases I and II.

The Bacillus subtilis hom gene, encoding homoserine dehydrogenase (L-homoserine:NADP+ oxidoreductase, EC 1.1.1.3) has been cloned and its nucleotide sequence determined. The B. subtilis enzyme expressed in Escherichia coli is sensitive by inhibition by threonine and allows complementation of a strain lacking homoserine dehydrogenases I and II. Nucleotide sequence analysis indicates that the hom stop codon overlaps the start codon of thrC (threonine synthase) suggesting that these genes, as well as thrB (homoserine kinase) located downstream from thrC, belong to the same transcription unit. The deduced amino acid sequence of the B. subtilis homoserine dehydrogenase shows extensive similarity with the C-terminal part of E. coli aspartokinases-homoserine dehydrogenases I and II; this similarity starts at the exact point where the similarity between E. coli or B. subtilis aspartokinases and E. coli aspartokinases-homoserine dehydrogenases stops. These data suggest that the E. coli bifunctional polypeptide could have resulted from the direct fusion of ancestral aspartokinase and homoserine dehydrogenase. The B. subtilis homoserine dehydrogenase has a C-terminal extension of about 100 residues (relative to the E. coli enzymes) that could be involved in the regulation of the enzyme activity.

Alcohol Oxidoreductases

Nucleotide sequence, transcriptional mapping, and temporal expression of the gene encoding p39, a major structural protein of the multicapsid nuclear polyhedrosis virus of Orgyia pseudotsugata.

The gene encoding the 39-kDa major structural protein (p39) of Orgyia pseudotsugata nuclear polyhedrosis virus (OpMNPV) was sequenced and transcriptionally mapped, and its expression was examined at various times postinfection. By Northern hybridization, primer extension, and S1 nuclease analysis, we identified p39 mRNAs of approximately 2600 nt. By primer extension analysis, we identified two major sets of transcripts which initiated around -48 and -96 nt upstream of the translation start codon. The transcription start sites were located within the conserved baculovirus late gene consensus sequence, ATAAG, which is duplicated in the p39 5' flanking region. In OpMNPV-infected Lymantria dispar cells, the p39 mRNAs were expressed abundantly at 24 and 36 hr p.i. but were present in lower quantities at 48 hr p.i. The p39 gene contained an open reading frame of 1053 nt which encodes a predicted protein of 351 amino acids with an estimated molecular weight of 39.5 kDa. Three repeats of the amino acid sequence Ala-Pro-Ala-Ala-Pro were identified at the C-terminus of the predicted p39 protein.

Amino Acid Sequence

Development of a prokaryotic expression vector that exploits dicistronic gene organization.

For unknown reasons, levels of expression of foreign genes inserted into expression vectors in Escherichia coli have frequently been undetectable. The most critical step in the successful production of foreign proteins seems to be the initiation of translation. Since most prokaryotic genes are transcribed in a polycistronic form, we have devised a new prokaryotic expression system utilizing dicistronic gene organization. Downstream from a strong promoter and the gene encoding glutathione S-transferase from Schistosoma japonicum, various foreign genes were connected via a ribosome-binding site, a stop codon and a start codon. The VH domain of an immunoglobulin fused to the alpha subunit of tryptophan synthase, FK506-binding protein, cyclophilin, and a domain of a major histocompatibility complex antigen were successfully produced in E. coli as discrete polypeptides by this method.

Amino Acid Isomerases

Characterization of the human N-CAM promoter.

In contrast with the complex series of splicing choices that generate the various membrane-associated isoforms of the neural cell-adhesion molecule alternative splicing of 5' exons does not contribute to additional molecular diversity. A single regulatory unit in genomic DNA, mapping to a 5 kb restriction-endonuclease-HindIII fragment, controls the expression of all major RNA size classes. DNA sequence analysis of a 2 kb fragment spanning the two major identified transcriptional initiation sites (194 and 188 bp from the ATG codon) and translation start codon indicates that the regulatory unit does not possess classical TATA or CCAAT motifs. The region of the putative promoter exhibits a GC-rich content and a high frequency of the dinucleotide CpG, both characteristics of a HTF(HpaII tiny fragments)-island. Introduction of deletion-mutant chimaeric-gene constructs into human and rodent N-CAM-expressing cell lines defines an active promoter region of 467 bp (-144 to -611 bp from the ATG codon). This region of genomic DNA contains consensus sites for the interaction of known transcriptional factors.

Animals

Rhizobium meliloti fixGHI sequence predicts involvement of a specific cation pump in symbiotic nitrogen fixation.

We present genetic and structural analyses of a fix operon conserved among rhizobia, fixGHI from Rhizobium meliloti. The nucleotide sequence of the operon suggests it may contain a fourth gene, fixS. Adjacent open reading frames of this operon showed an overlap between TGA stop codons and ATG start codons in the form of an ATGA motif suggestive of translational coupling. All four predicted gene products contained probable transmembrane sequences. FixG contained two cysteine clusters typical of iron-sulfur centers and is predicted to be involved in a redox process. FixI was found to be homologous with P-type ATPases, particularly with K+ pumps from Escherichia coli and Streptococcus faecalis but also with eucaryotic Ca2+, Na+/K+, H+/K+, and H+ pumps, which implies that FixI is a pump of a specific cation involved in symbiotic nitrogen fixation. Since prototrophic growth of fixI mutants appeared to be unimpaired, the predicted FixI cation pump probably has a specifically symbiotic function. We suggest that the four proteins FixG, FixH, FixI, and FixS may participate in a membrane-bound complex coupling the FixI cation pump with a redox process catalyzed by FixG.

Adenosine Triphosphatases

Cloning, sequencing, and mapping of the bacterioferritin gene (bfr) of Escherichia coli K-12.

The bacterioferritin (BFR) of Escherichia coli K-12 is an iron-storage hemoprotein, previously identified as cytochrome b1. The bacterioferritin gene (bfr) has been cloned, sequenced, and located in the E. coli linkage map. Initially a gene fusion encoding a BFR-lambda hybrid protein (Mr 21,000) was detected by immunoscreening a lambda gene bank containing Sau3A restriction fragments of E. coli DNA. The bfr gene was mapped to 73 min (the str-spc region) in the physical map of the E. coli chromosome by probing Southern blots of restriction digests of E. coli DNA with a fragment of the bfr gene. The intact bfr gene was then subcloned from the corresponding lambda phage from the gene library of Kohara et al. (Y. Kohara, K. Akiyama, and K. Isono, Cell 50:495-508, 1987). The bfr gene comprises 474 base pairs and 158 amino acid codons (including the start codon), and it encodes a polypeptide having essentially the same size (Mr 18,495) and N-terminal sequence as the purified protein. A potential promoter sequence was detected in the 5' noncoding region, but it was not associated with an "iron box" sequence (i.e., a binding site for the iron-dependent Fur repressor protein). BFR was amplified to 14% of the total protein in a bfr plasmid-containing strain. An additional unidentified gene (gen-64), encoding a relatively basic 64-residue polypeptide and having the same polarity as bfr, was detected upstream of the bfr gene.

Amino Acid Sequence

Conservation in the 5' flanking sequences of transcribed members of the Caenorhabditis elegans major sperm protein gene family.

The major sperm proteins (MSPs) are encoded in the Caenorhabditis genome by a multigene family with more than 50 genes dispersed in small clusters at three chromosomal loci. In spite of their dispersed locations, all of the MSP genes appear to be expressed at the same time exclusively in the testis, indicating co-ordinate temporal and spatial regulation of these dispersed genes. Many of the MSP genes must be transcribed, because RNA hybridization with gene-specific probes showed that individual genes each contribute less than 3% to the total poly(A)+ RNA, and 13 out of 14 sequenced cDNAs came from different genes. Primer extension assays from MSP mRNA showed that most of the MSP mRNAs must be initiated at position -35 from the translation start codon. Extensive similarity was found in the first 100 nucleotides of genomic sequence flanking the start codons of ten MSP genes from different chromosomal locations. All MSP genes contained a consensus ribosome binding site, a consensus TATA homology 27 nucleotides distal to the site of mRNA initiation, and ten highly conserved nucleotides adjacent to the site of initiation. All the MSP genes contained the sequence AGATCT located approximately 65 nucleotides upstream from the transcriptional start, but little or no similarity was found more distal to this. Some of these conserved sequences may be cis-acting control elements that ensure the cell and temporal specificity of transcription of these co-ordinately regulated genes.

Animals

cDNA sequence and expression of a phosphoenolpyruvate carboxylase gene from soybean.

A full-length cDNA encoding a subunit of phosphoenolpyruvate carboxylase (PEPC) was isolated from a developing seed expression library of the C3 plant Glycine max. The corresponding mRNA is present at similar levels in leaf, stem, root and developing seed. Two potential start codons exist, and the activity of protein initiated from the first such codon could be subject to regulation by protein kinase. Sequence comparison shows a similar upstream start codon in the case of the Ppc2 gene from Mesembryanthemum crystallinum, previously assumed to lack the sequences necessary for phosphorylation. The soybean encoded protein tends to resemble other 'C3-type' PEPC proteins more closely than those implicated in C4 or crassulacean acid metabolism.

Amino Acid Sequence

PKD1 upstream open reading frames affect Polycystin-1 expression and polycystic kidney disease phenotypes.

Autosomal dominant polycystic kidney disease (ADPKD) accounts for 5%-10% of prevalent end-stage kidney failure (ESKD). ADPKD cysts result from a loss of sufficient functional expression of PKD1/Polycystin-1 (PC1) in approximately 80% of families. Kidney disease severity correlates with the extent to which PC1 dosage is reduced below a critical level, and evidence suggests therapeutic benefit from increasing PC1 expression in these conditions. Upstream open reading frame (uORF) translation can reduce translation of a protein's coding sequence. Ribosome profiling data and bioinformatic predictions suggested the presence of conserved PKD1 uORFs, so we sought to explore their biological role. We generated luciferase reporters and two humanized PKD1 5' UTR mouse models with or without single nucleotide edits removing uORF start codons (ΔuORF) to define active uORFs and test their impact on PC1 translation. PKD1 uORF start codons can robustly initiate translation, and ΔuORF conveys a 2-4 fold increase in PC1 protein expression and resultant prevention of kidney cysts in Dnajb11 as well as in Pkd1 missense models. PKD1 uORF1-blocking steric antisense oligonucleotides (ASOs) substantially increase PC1 expression in vitro. PKD1 uORFs play an important role in the low basal expression of WT PKD1, and their inhibition represents an opportunity to therapeutically increase PC1 translation in polycystic kidney and liver disease resulting from reduced dosage of PC1.

Animals

Paramecium mitochondrial DNA sequences and RNA transcripts for cytochrome oxidase subunit I, URF1, and three ORFs adjacent to the replication origin.

A 2-kb region adjacent to the replication origin (ori) and a 3-kb region located between the small and large ribosomal RNAs of Paramecium mitochondrial (mt) DNA have been sequenced and the locations of their transcripts determined. The ori segment contains four transcripts, some of which are overlapping, which encode a known protein and two other open reading frames. The other segment encodes, on separate transcripts, the cytochrome c oxidase subunit one gene (COI) and the URF1 gene (ND1) common to most mt genomes. All these genes have the same orientation and do not contain introns. The COI gene is the most divergent of those known and has an internal 108 amino acid 'insert' not found in COI genes from other organisms. With these data it is possible to define a probable Paramecium mt genetic code. With the exception that TGA codes for tryptophan and the use of different start codons, Paramecium mtDNA appears to follow the universal code. GTA possibly can be used as a start codon.

Amino Acid Sequence

A genomic clone containing the promoter for the gene encoding the human lysosomal enzyme, alpha-galactosidase A.

We have isolated and characterized a human genomic clone for a lysosomal enzyme gene. The start point of transcription was identified using primer extension of poly(A)+ mRNA. This genomic clone is specific for human alpha-galactosidase A, and it includes sequences for the promoter, complete signal peptide, first exon, and part of the first intron. Direct and inverted repeat elements of 10, 11, 16, 19, and 22 nucleotides (nt) flank the promoter site. A (GA)n repeat element of approx. 60 nt with strong homology to similar elements identified in several species is located upstream from the promoter. A GGGCGG site specific for DNA-binding protein Sp1 is located near a CAAT box, and the CCGCCC inverted repeat of the Sp1 binding sequence is located by the TATA box. The sequence immediately flanking the ATG start codon of the human alpha-galactosidase A is highly homologous to sequences flanking the ATG start codons of the other human lysosomal hydrolases for which sequence information is available (beta-glucocerebrosidase, cathepsin B, cathepsin D, and beta-hexosaminidase alpha chain), but not for any of the other 133 human signal peptides examined. Our analysis also reveals that conversion of the propeptide to the mature enzyme involves cleavage of a C-terminal rather than an N-terminal fragment. This information about the normal alpha-galactosidase A gene will be useful for comparison to data obtained from patients with Fabry disease, who are characterized by a deficiency of this enzyme. This is the first genomic clone described to date for any lysosomal enzyme, and it establishes a reference for future analyses of the molecular events that mediate the expression of this important class of enzymes.

Amino Acid Sequence

Structural organization of the genes for murine and human leukemia inhibitory factor. Evolutionary conservation of coding and non-coding regions.

Leukemia inhibitory factor, LIF, is a glycoprotein with multiple activities in both the adult and the embryo. LIF appears to be encoded by a unique gene in both mouse and man, although the 3'-untranslated region of the mouse LIF gene gives a complex hybridization pattern on Southern blots. The complete nucleotide sequences of both the murine and human LIF genes and their flanking regions (8.7 and 7.6 kilobase pairs, respectively) were determined and compared. Both genes comprise three exons, two introns and an unusually long 3'-untranslated region (3.2 kilobase pairs), specificying a mRNA of approximately 4.1 kilobases. Two start sites of LIF-transcription were determined, by S1-nuclease protection and by a novel approach involving the polymerase chain reaction. S1-nuclease protection revealed a start site 60-64 base pairs upstream of the translational start codon and immediately downstream of a TATA box (TATATAAAT). The PCR approach identified a second transcriptional start site 160 base pairs 5' of the start codon and adjacent to a "TATA-like" element (CATAATTT). A comparison of the murine and human LIF gene sequences revealed a high degree of conservation in the coding regions and in segments of the untranslated and flanking regions. Seven segments displaying greater than 75% homology were identified, with the 5' and 3' ends of the transcription unit revealing the highest degree of homology. These conserved regions represents potential cis-acting control elements.

Amino Acid Sequence

Molecular structure and transformation of the glucose dehydrogenase gene in Drosophila melanogaster.

We have precisely mapped and sequenced the three 5' exons of the Drosophila melanogaster Gld gene and have identified the start sites for transcription and translation. The first exon is composed of 335 nucleotides and does not contain any putative translation start codons. The second exon is separated from the first exon by 8 kb and contains the Gld translation start codon. The inferred amino acid sequence of the amino terminus contains two unusual features: three tandem repeats of serine-alanine, and a relatively high density of cysteine residues. P element-mediated transformation experiments demonstrated that a 17.5-kb genomic fragment contains the functional and regulatory components of the Gld gene.

Animals

The parainfluenza virus type 1 P/C gene uses a very efficient GUG codon to start its C' protein.

Parainfluenza virus type 1 (PIV1) and Sendai virus (SEN) are very closely related, but the PIV1 P/C gene does not contain the ACG codon which initiates the SEN C' protein. Nevertheless, a protein corresponding to the PIV1 C' protein was observed both in vivo and in vitro. The initiation site of this protein maps upstream of the PIV1 C protein AUG in a region that does not contain an AUG codon. We have used site-directed mutagenesis to demonstrate that the PIV1 C' protein initiates from a GUG codon, four codons upstream of where the ACG is found in SEN. Remarkably, this GUG appears to initiate in vivo almost as frequently as AUG in the same context. However, whereas GUG permits downstream expression of the P and C proteins, AUG in this context does not. The conservation of an upstream non-AUG initiation codon for C' among PIV1 and SEN suggests that it is important for virus replication, even though some paramyxoviruses express only the C protein and others have no C open reading frame at all.

Base Sequence

Scanning model for translational reinitiation in eubacteria.

Premature termination of translation in eubacteria, like Escherichia coli, often leads to reinitiation at nearby start codons. Restarts also occur in response to termination at the end of natural coding regions, where they serve to enforce translational coupling between adjacent cistrons. Here, we present a model in which the terminated but not released ribosome reaches neighboring initiation codons by lateral diffusion along the mRNA. The model is based on the finding that introduction of an additional start codon between the termination and the reinitiation site consistently obstructs ribosomes to reach the authentic restart site. Instead, the ribosome now begins protein synthesis at this newly introduced AUG codon. This ribosomal scanning-like movement is bidirectional, has a radius of action of more than 40 nucleotides in the model system used, and activates the first encountered restart site. The ribosomal reach in the upstream direction is less than in the downstream one, probably due to dislodging by elongating ribosomes. The proposed model has parallels with the scanning mechanism postulated for eukaryotic translational initiation and reinitiation.

Bacteriolysis