PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Nucleotide sequence and transcriptional analysis of the pif gene of Spodoptera frugiperda nucleopolyhedrovirus (SfMNPV).

Defective viruses, not transmissible alone, increase the transmissibility of complete genotypes in natural populations of Spodoptera frugiperda multicapsid nucleopolyhedrovirus (SfMNPV). The defective phenotype is associated with a 15 kb deletion, which includes the pif (per os infectivity factor) gene. The sequence of a 2.4 kb fragment that includes pif was determined. Multiple transcripts encompassing pif were detected by Northern blot analysis. RT-PCR and nuclease protection analysis demonstrated the presence of run-through transcripts starting upstream of pif. A 2.0 kb messenger started from a CTAAG promoter motif located 11 nt upstream of the pif start codon, and ended 450 nt downstream from the pif stop codon. This pif mRNA included a small downstream ORF (homologous to Se37). A transcript of 0.8 kb was detected that may correspond to a specific transcript from this small ORF. This transcript would start at a late consensus motif internal to pif coding sequences, ending at the same polyadenylation signal as the pif transcript. These transcription features resemble those of pif transcription in Spodoptera littoralis NPV, although the genomic location of pif is not equivalent in the two viruses. SfMNPV pif can encode a protein of 529 amino acids, closely related to Spodoptera exigua MNPV PIF.

Amino Acid Sequence↗

Purification of bovine lysosomal alpha-mannosidase, characterization of its gene and determination of two mutations that cause alpha-mannosidosis.

Bovine kidney lysosomal alpha-mannosidase was purified to homogeneity and the gene was cloned. The gene was organized in 24 exons that spanned 16 kb and its corresponding cDNA contained an open reading frame of 2997 bp beginning from a putative ATG start codon. The deduced amino acid sequence contained a signal peptide of 50 amino acids adjacent to a protein sequence of 949 amino acids that was cleaved into five peptides in the mature enzyme; starting with the peptide derived from the N-terminal part of this precursor, their molecular masses were 35/38 (peptide a), 11/13 (peptide b), 22 (peptide c), 38 (peptide d) and 13/15 kDa (peptide e). Variation in the degree of N-glycosylation accounts for molecular mass heterogeneities of peptides a, b and e. Peptides a, b and c were disulphide-linked. A T961-->C transition, resulting in Phe321-->Leu substitution, was identified in the cDNA of alpha-mannosidosis-affected Angus cattle. In affected Galloway cattle, a G662-->A transition that causes Arg221-->His substitution was identified. Phe321 and Arg221 are conserved among the alpha-mannosidase class-2 family, indicating that the substitutions resulted from disease-causing mutations in these breeds.

Amino Acid Sequence↗

Characterization of the gene encoding human platelet glycoprotein IX.

Glycoprotein IX is a relatively small (M(r) 20,000) surface glycoprotein of human platelets; one of three (Ib alpha, Ib beta, IX) polypeptide chains present in the glycoprotein Ib-IX complex that functions as the von Willebrand factor receptor and mediates platelet adhesion in the arterial circulation. Using a cDNA for human glycoprotein IX as the probe, clones were isolated from a human genomic library, and the genomic sequence for glycoprotein IX (3.2 kilobases) was determined. The transcriptional start site was located by RNase protection and primer extension experiments. The gene includes three exons and two introns within 1.6 kilobases of DNA, and the entire open reading frame for glycoprotein IX is included within the third exon. The genes for glycoproteins IX and Ib alpha share similar exon sequences on the 5' side of their ATG start codons, and both genes possess introns in this region. The glycoprotein IX gene contains two consensus regulatory sequences (GATA and ACTTCCT [ets]) in its promoter region (5' flank, within 67 bases of the start site) that are also present in similar sites in the previously described "megakaryocyte-platelet" genes (glycoprotein IIb, platelet factor 4, beta-thromboglobulin: human and rat). Thus, the glycoprotein IX gene shares structural features with other megakaryocyte-platelet genes and contains at least two consensus cis-acting regulatory elements that may govern gene expression.

Amino Acid Sequence↗

Analysis of the cytosolic proteome of Halobacterium salinarum and its implication for genome annotation.

The halophilic archaeon Halobacterium salinarum (strain R1, DSM 671) contains 2784 protein-coding genes as derived from the genome sequence. The cytosolic proteome containing 2042 proteins was separated by two-dimensional gel electrophoresis (2-DE) and systematically analyzed by a semi-automatic procedure. A reference map was established taking into account the narrow isoelectric point (pI) distribution of halophilic proteins between 3.5 and 5.5. Proteins were separated on overlapping gels covering the essential areas of pI and molecular weight. Every silver-stained spot was analyzed resulting in 661 identified proteins out of about 1800 different protein spots using matrix-assisted laser desorption/ionization time of flight mass spectrometry (MALDI-TOF MS) peptide mass fingerprinting (PMF). There were 94 proteins that were found in multiple spots, indicating post-translational modification. An additional 141 soluble proteins were identified on 2-D gels not corresponding to the reference map. Thus about 40% of the cytosolic proteome was identified. In addition to the 2784 protein-coding genes, the H. salinarum genome contains more than 6000 spurious open reading frames longer than 100 codons. Proteomic information permitted an improvement in genome annotation by validating and correcting gene assignments. The correlation between theoretical pI and gel position is exceedingly good and was used as a tool to improve start codon assignments. The fraction of identified chromosomal proteins was much higher than that of those encoded on the plasmids. In combination with analysis of the GC content this observation permitted an unambiguous identification of an episomal insert of 60 kbp ("AT-rich island") in the chromosome, as well as a 70 kbp region from the chromosome that has integrated into one of the megaplasmids and carries a series of essential genes. About 63% of the chromosomally encoded proteins larger than 25 kDa were identified, proving the efficacy of 2-DE MALDI-TOF MS PMF technology. The analysis of the integral membrane proteome by tandem mass spectrometric techniques added another 141 identified proteins not identified by the 2-DE approach (see following paper).

Bacterial Proteins↗

Systematic alteration of the nucleotide sequence preceding the translation initiation codon and the effects on bacterial expression of the cloned SV40 small-t antigen gene.

In the preceding paper (Derom et al., 1981) we described the cloning in bacterial plasmids of the simian virus 40 (SV40) small-t antigen gene under transcriptional control of the bacteriophage lambda pL promoter. Systematic variation of the distance and/or nucleotide sequence between the Shine-Dalgarno ribosome interaction sequence and the small-t translation initiation codon leads to considerable differences in production of small-t by the different plasmids. Secondary structure models derived for the different mRNAs confirm our previous conclusions about the requirement first for an accessible start codon and second for an accessible ribosome interaction site for efficient translation initiation. Secondary structure models for mRNAs from plasmids containing the small-t gene under control of the lac promoter are in agreement with these conclusions.

Antigens, Viral↗

Promoter analysis of the membrane protein gp64 gene of the cellular slime mold Polysphondylium pallidum.

We cloned a genomic fragment of the membrane protein gp64 gene of the cellular slime mold Polysphondylium pallidum by inverse PCR. Primer extension analysis identified a major transcription start site 65 bp upstream of the translation start codon. The promoter region of the gp64 gene contains sequences homologous to a TATA box at position -47 to -37 and to an initiator (Inr, PyPyCAPyPyPyPy) at position -3 to +5 from the transcription start site. Successively truncated segments of the promoter were tested for their ability to drive expression of the beta-galactosidase reporter gene in transformed cells; also the difference in activity between growth conditions was compared. The results indicated that there are two positive vegetative regulatory elements extending between -187 and -62 bp from the transcription start site of the gp64 promoter; also their activity was two to three times higher in the cells grown with bacteria in shaken suspension than in the cells grown in an axenic medium.

Animals↗

Isolation and characterization of a new FHL1 variant (FHL1C) from porcine skeletal muscle.

Four and a half LIM domain protein 1 (FHL1) was initially described as an abundant skeletal muscle protein with four LIM domains and a GATA like zinc finger. FHL1 was shown to be expressed in skeletal muscle as well as in a variety of other tissues. Recently, alternatively spliced FHL1 mRNAs were identified coding for C-terminal truncated proteins. The tissue distribution of these variants is more restricted and their functional properties seem to be different. We have isolated and characterized a new variant of FHL1 from porcine skeletal muscle (FHL1C). FHL1C is characterized by a newly identified start codon resulting in a 16 amino acids longer N- terminal region. We have isolated and characterized the porcine FHL1C gene spanning approximately 14 kb and harboring six exons. Using primer extension analysis, the transcription start site of FHL1C was mapped, indicating that FHL1C is regulated by an alternative promoter. The tissue distribution of FHL1C expression was studied by RT-PCR. The porcine FHL1C gene was assigned to the distal part of the long arm of the X chromosome by fluorescence in situ hybridization and screening of a somatic porcine/rodent cell hybrid panel.

Alternative Splicing↗

The human glucocerebrosidase gene has two functional ATG initiator codons.

Gaucher disease is due to a deficiency in the activity of the enzyme glucocerebrosidase. Glucocerebrosidase is a lysosomal enzyme that presumably requires a signal peptide for transport across the membrane of the rough endoplasmic reticulum and glycosylation for transport into lysosomes. Human glucocerebrosidase cDNA contains two potential ATG start codons in its long open reading frame. The signal peptides that are initiated from each ATG are quite different in their hydrophobicity. We demonstrate that either ATG can function independently to produce active glucocerebrosidase enzyme in cultured fibroblasts. The glucocerebrosidase activity produced from translation products initiated at either ATG is found predominantly in the lysosomes.

Amino Acid Sequence↗

Initiator tRNA may recognize more than the initiation codon in mRNA: a model for translational initiation.

A special methionine tRNA (tRNAi) is universally required to initiate translation. Amongst species a tRNAi structural conservation is most apparent in the anticodon and T arms of the molecule but extends into the variable loop and the 3' strand of the D stem. This suggested that they could share a similar ancestral or current function in initiation of translation. We report that the sequence of bases neighboring the translational start codons of many eubacterial genes are complementary not only to the extended anticodon but also to the D and T loops of tRNAi. Study of the coding properties of tRNAi and of mutations that affect translation suggests that the translational start domain can be a mosaic of signals complementary to the loops of tRNAi. The hypothesis of multiple loop recognition suggests that unusual triplets can start prokaryotic and mitochondrial genes and predicts the occurrence of other reading frames. Furthermore, it suggests a unifying model for chain initiation based on RNA contacts and displacements.

Animals↗

Molecular identification of the long isoform of the human neuropeptide Y Y5 receptor and pharmacological comparison with the short Y5 receptor isoform.

The neuropeptide Y Y5 receptor gene generates two splice variants, referred to here as Y5(L) (long isoform) and Y5(S) (short isoform). Y5(L) mRNA differs from Y5(S) mRNA in its 5' end, generating a putative open reading frame with 30 additional nucleotides upstream of the initiator AUG compared with the Y5(S) mRNA. The purpose of the present work was to investigate the existence of the Y5(L) mRNA. The authenticity of this transcript was confirmed by isolating part of its 5' untranslated region through 5' rapid amplification of cDNA ends and analysing its tissue distribution. To study the initiation of translation on Y5(L) mRNA, we cloned the Y5(L) cDNA and two Y5(L) cDNA mutants lacking the first or the second putative initiation start codon. Transient expression of the three plasmids in COS-7 cells and saturation binding experiments using (125)I-labelled polypeptide YY (PYY) as a ligand showed that initiation of translation on Y5(L) mRNA could start at the first AUG, giving rise to a Y5(L) receptor with an N-terminal 10-amino-acid extension when compared with the Y5(S) receptor. The human Y5(L) and Y5(S) receptor isoforms displayed similar affinity constants (1.3 nM and 1.5 nM respectively). [(125)I]PYY binding to COS-7 cells expressing either the Y5(L) or the Y5(S) isoform was inhibited with the same rank order of potency by a selection of six chemically diverse compounds: PYY>neuropeptide Y>pancreatic polypeptide>CGP71683A>Synaptic 34>Banyu 6. Comparison of the tissue distribution of Y5(L) and Y5(S) mRNAs, as determined by reverse transcription-PCR analysis, indicated that expression of Y5(L) mRNA occurs in a tissue-specific manner. Finally, we have shown that the two AUG triplets contained in the 5' untranslated region of Y5(L) mRNA did not affect receptor expression.

5' Untranslated Regions↗

The mouse chondroadherin gene: characterization and chromosomal localization.

The mouse chondroadherin gene was isolated from a cosmid genomic library by the use of a rat chondroadherin cDNA probe. Southern blot analysis of mouse genomic DNA revealed a simple pattern of hybridization indicating a single copy gene for chondroadherin. The mouse chondroadherin gene encompasses 4.1 kb and consists of four exons separated by one large intron of 1929 bp followed by two smaller introns of 247 and 225 bp, respectively. Most of the translated region, including the start codon and the main part of a leucine-rich region, is contained within the first exon. Two small exons of 164 and 146 bp encode the rest of the protein. Interestingly, 4 bases from the stop codon, in the 3'-UTR, a third intron is located. A putative promoter region of 669 bp was sequenced and shown to contain a potential TATAA-box signal 29 bp upstream of the transcription start site and several recognition sites for transcription factors. The exon/intron organization of the chondroadherin gene differs from those of the other known genes of the leucine-rich repeat (LRR) family in the extracellular matrix. Taken together with comparison of protein sequences of other members of the LRR family in the extracellular matrix, the data suggest that chondroadherin has evolved along a different pathway. The chondroadherin gene was mapped to mouse chromosome 11, near D11Mit14, by single-strand conformation polymorphism linkage analysis.

Amino Acid Sequence↗

Identification and characterization of upstream open reading frames (uORF) in the 5' untranslated regions (UTR) of genes in Saccharomyces cerevisiae.

We have taken advantage of recently sequenced hemiascomycete fungal genomes to computationally identify additional genes potentially regulated by upstream open reading frames (uORFs). Our approach is based on the observation that the structure, including the uORFs, of the post-transcriptionally uORF regulated Saccharomyces cerevisiae genes GCN4 and CPA1 is conserved in related species. Thirty-eight candidate genes for which uORFs were found in multiple species were identified and tested. We determined by 5' RACE that 15 of these 38 genes are transcribed. Most of these 15 genes have only a single uORF in their 5' UTR, and the length of these uORFs range from 3 to 24 codons. We cloned seven full-length UTR sequences into a luciferase (LUC) reporter system. Luciferase activity and mRNA level were compared between the wild-type UTR construct and a construct where the uORF start codon was mutated. The translational efficiency index (TEI) of each construct was calculated to test the possible regulatory function on translational level. We hypothesize that uORFs in the UTR of RPC11, TPK1, FOL1, WSC3, and MKK1 may have translational regulatory roles while uORFs in the 5' UTR of ECM7 and IMD4 have little effect on translation under the conditions tested.

5' Untranslated Regions↗

Cloning, expression, purification, and characterization of Nocardia sp. GTP cyclohydrolase I.

The sequence of the gene from Nocardia sp. NRRL 5646 encoding GTP cyclohydrolase I (GCH), gch, and its adjacent regions was determined. The open reading frame of Nocardia gch contains 684 nucleotides, and the deduced amino acid sequence represents a protein of 227 amino acid residues with a calculated molecular mass of 24,563Da. The uncommon start codon TTG was identified by matching the N-terminal amino acid sequence of purified Nocardia GCH with the deduced amino acid sequence. A likely ribosomal binding site was identified 9bp upstream of the translational start site. The 3' end flank region encodes a peptide that shares high homology with dihydropteroate synthases. Nocardia GCH has 73 and 60% identity to the proteins encoded by the putative gch of Mycobacterium tuberculosis and Streptomyces coelicolor, respectively. Nocardia GCH was highly expressed in Escherichia coli cells carrying a pHAT10 based expression vector, and moderately expressed in Mycobacterium smegmatis cells carrying a pSMT3 based expression vector. Enterokinase digestion of recombinant Nocardia GCH, and in-gel digestion of Nocardia GCH and recombinant GCH followed by MALDI-TOF-MS analysis, confirmed that the actual subunit size of the enzyme was 24.5kDa. Thus, we conclude that the active form of native Nocardia GCH is a decamer. Our earlier incorrect conclusion was that the native enzyme was an octamer derived from the anomalous SDS-PAGE migration of the subunit.

Amino Acid Sequence↗

Characterization of a rabbit gene encoding a clofibrate-inducible fatty acid omega-hydroxylase: CYP4A6.

CYP4A6 mRNAs are induced in the rabbit liver and kidney following treatment with the antihyperlipidemic drug clofibrate. As a first step toward the elucidation of the mechanism controlling the induction of this and other CYP4A genes by clofibrate and other peroxisome proliferators, we have cloned and characterized the CYP4A6 gene. Genomic DNA containing the first 12 exons encoding CYP4A6 was isolated as three recombinant lambda phage, two of which were overlapping. The sequence of more than 1000 bp of the 5' upstream region as well as of the first 12 exons has been determined. These 12 exons encode all but approximately 80 bp at the 3' terminus of CYP4A6. Intron/exon junctions within the coding region of the gene are conserved relative to the rat CYP4A1 and CYP4A2 genes. Primer extension analysis indicates that transcription is initiated 33 bp upstream of the start codon. The CYP4A6 promoter region, like that of the rat CYP4A1 and CYP4A2 genes, does not contain a consensus TATA box. However, a consensus Sp1 recognition element is apparent at -46 bp upstream of the transcription start site. In addition, a sequence related to one of two regulatory elements that control the induction of the rat acyl-CoA oxidase gene by ciprofibrate is present upstream of the CYP4A6 promoter.

Acyl-CoA Oxidase↗

Cloning and sequencing of the dnaK region of Streptomyces coelicolor A3(2).

The dnaK homologue of Streptomyces coelicolor A3(2) strain M145 has been cloned and sequenced. Nucleotide sequence analysis of a 2.5-kb region revealed an open reading frame (ORF) encoding a predicted DnaK protein of 618 amino acids (M(r) = 66,274). The dnaK coding sequence displays extreme codon bias and shows a strong preference for CGY and GGY, for Arg and Gly codons, respectively. The predicted DnaK sequence has a high Lys:Arg ratio which is not typical of streptomycete proteins. The region immediately downstream from dnaK contains an ORF for a GrpE-like protein; the predicted start codon of grpE overlaps the last two codons of dnaK, indicating that the two genes are translationally coupled. This organisation differs from that reported for other prokaryotes.

Amino Acid Sequence↗

Human mitochondrial C1-tetrahydrofolate synthase: gene structure, tissue distribution of the mRNA, and immunolocalization in Chinese hamster ovary calls.

C1-tetrahydrofolate (THF) synthase is a trifunctional enzyme found in eukaryotes that contains the activities 10-formyl-THF synthetase, 5,10-methenyl-THF cyclohydrolase, and 5,10-methylene-THF dehydrogenase. The cytoplasmic isozyme of C1-THF synthase is well characterized in a number of mammals, including humans; but a mitochondrial isozyme has been previously identified only in the yeast Saccharomyces. Here, we report the identification and characterization of the human gene encoding a functional mitochondrial C1-THF synthase. The gene spans 236 kilobase pairs on chromosome 6 and consists of 28 exons plus one alternative exon. The gene encodes a protein of 978 amino acids, including an N-terminal mitochondrial targeting sequence. The mitochondrial isozyme is 61% identical to the human cytoplasmic isozyme. Expression of the gene was detected in most human tissues, but transcripts were highest in placenta, thymus, and brain. Two mRNAs were detected, a 3.6-kb transcript and a 1.1-kb transcript, and both transcripts were observed in varying ratios in each tissue. The shorter transcript results from an alternative splicing event, where exon 7 is spliced to exon 8a instead of exon 8. Exon 8a is derived from an exonized Alu sequence, sharing no homology with exon 8 of the long transcript, and encodes just 15 amino acids followed by a stop codon and a polyadenylation signal. This short transcript potentially encodes a bifunctional enzyme lacking 10-formyl-THF synthetase activity. Both transcripts initiate at the same 5'-site, 107 nucleotides up-stream of the ATG start codon. The full-length (2934 bp) cDNA fused to a C-terminal V5 epitope tag was expressed in Chinese hamster ovary cells. Immunoblots of subfractionated cells revealed a 107-kDa protein only in the mitochondrial fractions of these cells, confirming the mitochondrial localization of the protein. Yeast cells expressing the full-length human cDNA exhibited elevated 10-formyl-THF synthetase activity, confirming its identification as the human mitochondrial C1-THF synthase.

Alternative Splicing↗

First restriction and genetic mapping of the genomic DNA of urease-positive thermophilic campylobacters (UPTC), and small restriction fragment sequencing.

A restriction and genetic map of urease-positive thermophilic campylobacter (UPTC) CF89-12 genome DNA is constructed using a pulsed-field gel electrophoresis procedure after digestion with SalI and SmaI and Southern blot hybridisation. Each of the six gene fragments (flaA, glyA, lysS, recA, sodB and ureAB) selected are mapped in only a fragment on the restriction map. Three DNA fragments for rrn operon probes are mapped in multiple regions on the map. When two SmaI-digested neighbouring small fragments hybridised with rrn probes are cloned and sequenced, a total sequence length of 7487 bp is determined. In the sequence, part of the pnp gene (734 bp) bearing a p-independent transcriptional termination region, a cluster of five tRNA genes including the putative promoter region, a hypothetical Cj0171-like 507-bp sequence containing an internal termination codon, and a part of the rrn operon including the putative promoter region (4700 bp) are identified. The 507 bp sequence carried both putative transcriptional promoter sequences, including a ribosome binding site upstream of the ATG start codon and a characteristic G9 structure, and a possible p-independent transcriptional termination region. A hypothetical Cj0170-like 204-bp sequence containing an internal termination codon also occurred, overlapping partly with the Cj0171-like sequence. Based on nucleotide sequence alignment analysis between the UPTC rrn operon examined here and the previously reported one, two different 16S-23S ribosomal DNA (rDNA) internal spacer regions are shown to exist.

Base Sequence↗

Structure of the parsley caffeoyl-CoA O-methyltransferase gene, harbouring a novel elicitor responsive cis-acting element.

The sequence of the S-adenosyl-L-methionine:trans-caffeoyl-CoA O-methyltransferase (CCoAOMT, EC2.1.1.104) gene, including the 5'-flanking region of 5 kb, was determined from parsley (Petroselinum crispum) plants. The enzyme appears to be encoded by one or two genes, and the ORF is arranged in five exons spaced by introns from 107 to 263 bp in length. The genomic sequence matches the ORF of the cDNA previously reported from elicited parsley cell cultures, showing only three base changes that do not affect the enzyme polypeptide sequence. S1 nuclease protection assays and primer extension analyses with genomic and cDNA templates revealed the transcription start site 67 bp upstream of the translation start codon, indicating a shorter 5'-UTR than reported previously for the transcript. Promoter regulatory consensus elements such as two 'CAAT' boxes and one 'TATA' box were identified at -196, -127 and -31, respectively, relative to the transcription start site, and an SV 40-like enhancer element is located 347 bp upstream. Most notably, three putative cis-regulatory elements were recognized by sequence alignments, which represent motifs recurring in the promoters of several genes of the stress-inducible phenylpropanoid pathway (boxes P, A and L). Transient expression assays with a set of 5'-truncated promoter-GUS fusions show that significant promoter activity is retained in a 354 bp promoter fragment. In vitro DNase 1 footprint experiments and electrophoretic mobilty shift assays (EMSA) identified in this fragment a unique sequence motif with elicitor-inducible trans-factor binding activity, which was unrelated to boxes P, A, or L. This novel cis-regulatory element, designated box E, appears to be conserved in the TATA-proximal regions of other stress-inducible phenylpropanoid genes, and in vitro binding of nuclear protein was confirmed in EMSA assays for such an element from the PAL-1 promoter (-54 to -45). Moreover, the deletion of box E reduced the activity and erased the elicitor-responsiveness of the CCoAOMT promoter in transient expression assays. The results corroborate the proposed physiological function of CCoAOMT in elicited plant cells and may shed new light on the sequential action of trans-active factors in the regulation of phenylpropanoid genes.

Amino Acid Sequence↗