PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Frameshifting at the internal stop codon within the mRNA for bacterial release factor-2 on eukaryotic ribosomes.

A translational frameshift is necessary in the synthesis of Escherichia coli release factor 2 (RF-2) to bypass an in-frame termination codon within the coding sequence. High-efficiency frameshifting around this codon can occur on eukaryotic ribosomes as well as prokaryotic ribosomes. This was determined from the relative efficiency of translation of RF-2 RNA compared with that for the other release factor RF-1, which lacks the in-frame premature stop codon. Since the termination product is unstable an absolute measure of the efficiency of frameshifting has not been possible. A gene fusion between trpE and RF-2 was carried out to give a stable termination product as well as the frameshift product, thereby allowing a direct determination of frameshifting efficiency. The extension of RF-2 RNA near its start codon with a fragment of the trpE gene, while still allowing high efficiency frameshifting on prokaryotic ribosomes, surprisingly gives a different estimate of frameshifting on the eukaryotic ribosomes than that obtained with RF-2 RNA alone. This paradox may be explained by long distance context effects on translation rates in the frameshift region created by the trpE sequences in the gene fusion, and may reflect that pausing and translation rate are fundamental factors in determining the efficiency of frameshifting.

Base Sequence↗

Characterization of the succinate dehydrogenase-encoding gene cluster (sdh) from the rickettsia Coxiella burnetii.

We have identified and sequenced four genes that encode the protein subunits comprising the succinate dehydrogenase enzyme complex (Sdh) of the rickettsia Coxiella burnetii. The Sdh-encoding gene cluster (sdhCDAB) begins 3326 bp upstream from the citrate synthase-encoding gene (gltA) start codon and is read with opposite polarity. An open reading frame encoding the N-terminal 280 amino acids (aa) of 2-oxoglutarate dehydrogenase (SucA) begins 24 bp downstream from the stop codon of the gene specifying the iron-sulfur subunit (sdhB) of Sdh. The deduced aa sequence of Sdh subunits and the N-terminal portion of SucA revealed significant aa identity with the Esherichia coli homologues ranging from a low of 36.6% for SdhD to a high of 61.2% for SdhA and SdhB. Primer extension identified transcription start points (tsp) for sdh and sucA. The region upstream from the sdh tsp, but not the sucA tsp, displayed homology to promoter consensus sequences of E. coli. Further evidence that sucA transcription can occur independent of sdh transcription was provided by demonstrating that a TnphoA insertion disrupting sdhB had no effect on the production of SucA by an E. coli cell-extract-directed in vitro transcription/translation system. The plasmid clone pLPM60, which carries the C. burnetii sdhCDAB coding and upstream regulatory regions, rescued an E. coli sdhA mutant (MOB252), indicating functional expression of the rickettsial locus. A cell extract of MOB252 transformed with pLPM60 showed a sixfold greater level of Sdh enzyme activity over the E. coli wild type. A plasmid clone lacking the sdh upstream regulatory region did not complement nor produce sdh mRNA by dot blot analysis.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Evolutionarily conserved non-AUG translation initiation in NAT1/p97/DAP5 (EIF4G2).

Only a few cases of exclusive translation initiation at non-AUG codons have been reported. We recently demonstrated that mammalian NAT1 mRNA, encoded by EIF4G2, uses GUG as its only translation initiation codon. In this study, we identified NAT1 orthologs from chicken, Xenopus, and zebrafish and found that in all species, the GUG codon also serves as the initiation codon. In all species, the GUG codon fulfilled the reported requirements for non-AUG initiation: an optimal Kozak motif and a downstream hairpin structure. Site-directed mutagenesis showed that nucleotides at positions -3 and +4 are critical for the GUG-mediated translation initiation in vitro. We found that NAT1 orthologs in Drosophila melanogaster and Halocynthia roretzi also use non-AUG start codons, demonstrating evolutionary conservation of the noncanonical translation initiation.

Amino Acid Sequence↗

Complete mitochondrial genome sequence of Urechis caupo, a representative of the phylum Echiura.

BACKGROUND: Mitochondria contain small genomes that are physically separate from those of nuclei. Their comparison serves as a model system for understanding the processes of genome evolution. Although hundreds of these genome sequences have been reported, the taxonomic sampling is highly biased toward vertebrates and arthropods, with many whole phyla remaining unstudied. This is the first description of a complete mitochondrial genome sequence of a representative of the phylum Echiura, that of the fat innkeeper worm, Urechis caupo. RESULTS: This mtDNA is 15,113 nts in length and 62% A+T. It contains the 37 genes that are typical for animal mtDNAs in an arrangement somewhat similar to that of annelid worms. All genes are encoded by the same DNA strand which is rich in A and C relative to the opposite strand. Codons ending with the dinucleotide GG are more frequent than would be expected from apparent mutational biases. The largest non-coding region is only 282 nts long, is 71% A+T, and has potential for secondary structures. CONCLUSIONS: Urechis caupo mtDNA shares many features with those of the few studied annelids, including the common usage of ATG start codons, unusual among animal mtDNAs, as well as gene arrangements, tRNA structures, and codon usage biases.

Amino Acid Sequence↗

A cis-acting element in the BCL-2 gene controls expression through translational mechanisms.

The bcl-2 gene becomes activated in many types of human cancers and contributes to neoplastic cell expansion, as well as to resistance to radiation and chemotherapy, by blocking programmed cell death or apoptosis. The expression of this proto-oncogene is regulated at both the transcriptional and post-transcriptional levels. DNA sequence comparisons of human, mouse, rat and chicken bcl-2 cDNAs revealed the presence of an open reading frame (ORF) [correction of (OFR)] located upstream of the normal coding region. Because upstream ORFs (uORFs) have been associated with translational repression, we analysed the functional significance of the 11 amino-acid uORF in the human BCL-2 gene (-119 to -84 bp). Deletion of this uORF from chloramphenicol acetyltransferase (CAT) reporter gene constructs that contained the bcl-2 promoter and entire 5'-untranslated region (5'-UTR), as well as introduction of an A-->T mutation at position -119 bp that destroyed the AUG-initiation codon, significantly increased CAT activity in HeLa, CEM, and other cell lines, without producing a corresponding elevation in CAT mRNA levels. Positioning this uORF, together with its accompanying Kozak sequences, between a heterologous promoter from SV40 and a CAT reporter gene resulted in marked inhibition of CAT protein production without a decrease in CAT mRNA. Mutation of the start codon (ATG-->TTG) of this uORF completely abolished its inhibitory activity, consistent with a translational mechanism. Taken together, these findings suggest that the uORF located within the 5'UTR of the bcl-2 gene is necessary and sufficient for translational regulation of bcl-2 gene expression.

Animals↗

Sequence analysis and expression of the aspartokinase and aspartate semialdehyde dehydrogenase operon from rifamycin SV-producing amycolatopsis mediterranei.

A approximately 4.8 kb KpnI fragment, from the upstream region of the methylmalonyl-CoA mutase gene (mutAB) of rifamycin SV-producing Amycolatopsis mediterranei, was cloned and partially sequenced. Codon preference analysis showed three complete ORFs. ORF2 is internal to ORF1, and encodes a polypeptide corresponding to 172 amino acids, whereas ORF1 encodes a polypeptide of 421 amino acids. They were identified as the encoding genes of aspartokinase alpha- and beta-subunits by comparing the amino acid sequences with those in the database. The downstream ORF3, whose start codon was overlapped with the stop codon of both ORF1 and ORF2 by 1 bp, was identified as the aspartate semialdehyde dehydrogenase gene (asd), encoding a polypeptide of 346 amino acids. Subclones containing either the ask gene or the asd gene were constructed, in which the genes could be expressed under Lac promoters. Two subclones could transform E. coli CGSC 5074 (ask-) and E. coli X6118 (asd-) to prototrophy, supporting the functional assignments. Southern hybridisation indicated that the approximately 4.8 kb sequenced region represented a continuous segment in the A. mediterranei chromosome. It is concluded that ask and asd genes are present in an operon in A. mediterranei, and therefore that organisation of these two genes is the same as in most gram-positive bacteria, such as Mycobacteria, Corynebacterium glutamicum and Bacillus subtilis, but is different from Streptomyces akiyoshiensis.

Actinobacteria↗

Hibiscus chlorotic ringspot virus p27 and its isoforms affect symptom expression and potentiate virus movement in kenaf (Hibiscus cannabinus L.).

Hibiscus chlorotic ringspot virus (HCRSV), a member of the genus Carmovirus, encodes p27 (27-kDa protein) and two other in-frame isoforms (p25 and p22.5) that are coterminal at the carboxyl end. Only p27, which initiates at the 2570CUG codon, was detected in transfected kenaf (Hibiscus cannabinus L.) protoplasts through fusion to a Flag tag at either its N or C terminus. Subcellular localization of a p27-green fluorescent fusion protein in kenaf epidermal cells showed that it was localized to membrane structures close to cell walls. To study the functions of these proteins, a number of start codon mutants and premature translation termination mutants were constructed. Phenotypic differences were observed between the wild-type virus and these mutants during infection. Infectivity assays on plants indicated that p27 is a determinant of symptom severity. Without p25, appearance of symptoms on systemically infected kenaf leaves was delayed by 4 to 8 days. In a timecourse analysis, Western blot assays revealed that the delay corresponded to retardation in virus systemic movement, which suggested that p25 is probably involved in virus systemic movement. Mutations disrupting expression of p22.5 did not affect symptoms or virus movement.

3' Untranslated Regions↗

Structure of the human gene encoding the protein repair L-isoaspartyl (D-aspartyl) O-methyltransferase.

The protein L-isoaspartyl/D-aspartyl O-methyltransferase (EC 2.1.1.77) catalyzes the first step in the repair of proteins damaged in the aging process by isomerization or racemization reactions at aspartyl and asparaginyl residues. A single gene has been localized to human chromosome 6 and multiple transcripts arising through alternative splicing have been identified. Restriction enzyme mapping, subcloning, and DNA sequence analysis of three overlapping clones from a human genomic library in bacteriophage P1 indicate that the gene spans approximately 60 kb and is composed of 8 exons interrupted by 7 introns. Analysis of intron/exon splice junctions reveals that all of the donor and acceptor splice sites are in agreement with the mammalian consensus splicing sequence. Determination of transcription initiation sites by primer extension analysis of poly(A)+ mRNA from human brain identifies multiple start sites, with a major site 159 nucleotides upstream from the ATG start codon. Sequence analysis of the 5'-untranslated region demonstrates several potential cis-acting DNA elements including SP1, ETF, AP1, AP2, ARE, XRE, CREB, MED-1, and half-palindromic ERE motifs. The promoter of this methyltransferase gene lacks an identifiable TATA box but is characterized by a CpG island which begins approximately 723 nucleotides upstream of the major transcriptional start site and extends through exon 1 and into the first intron. These features are characteristic of housekeeping genes and are consistent with the wide tissue distribution observed for this methyltransferase activity.

Base Sequence↗

Yeast promoters URA1 and URA3. Examples of positive control.

Transcription of the two unlinked structural genes URA1 and URA3 of Saccharomyces cerevisiae is positively regulated by the gene product PPR1. We have used S1 digestion and primer extension mapping to investigate the RNAs produced in different genetic backgrounds: wild-type, ppr1 deletion mutants, constitutively induced and non-inducible ppr1 mutants. Results show that each structural gene specifies multiple messenger RNA classes with different 5'-terminal sequences. The basal level of these transcripts does not require a functional PPR1 gene. Induction of URA1 results from an even increase of the level of synthesis of all the transcripts in contrast to that of URA3 which is effected by selectively increasing the levels of synthesis of one subset of transcripts. The PPR1-mediated control was also studied in the foreign genetic background of Schizosaccharomyces pombe using autonomously replicating hybrid plasmids carrying the gene URA1 or URA3 along with the regulatory gene PPR1, either in a constitutive or non-inducible allelic form. The 5' ends of the transcripts URA1 and URA3 made in S. pombe map upstream from the initiation sites used in S. cerevisiae. In contrast to S. cerevisiae, in S. pombe the URA3 but not URA1 transcripts respond to the PPR1-induction. We have identified a minimal control region for the PPR1-specific induction of URA1, that includes sequences located between the T-A-T-A box and the translation start codon. This region contains sequence features in common with URA3. There is an extensive alternating Pu:Py region including the T-A-T-A box of both promoters and an eight base-pair exact homology; further downstream, there is another 11 base-pair highly conserved sequence which either overlaps or lies in close proximity to the unregulated start sites of URA1 in S. pombe and of URA3 in S. cerevisiae. A positive regulatory model taking into accounts all these observations is presented.

Base Sequence↗

Structural and functional analysis of the promoter region involved in full expression of the cryIIIA toxin gene of Bacillus thuringiensis.

The promoter region of the cryIIIA toxin gene of Bacillus thuringiensis is composed of at least three domains: an upstream region extending from nucleotide positions -635 to -553 (with reference to the translational start codon of cryIIIA), an internal region extending from nucleotide positions -553 to -367, and a downstream region extending from nucleotide position -367 to +18. Deletion analysis and transcriptional fusions to the lacZ gene indicate that full expression of cryIIIA requires the association of the upstream and the downstream region. Primer extension experiments reveal a major cryIIIA transcript (designated T-129) starting at nucleotide position -129 and another transcript (designated T-558) starting at nucleotide position -558. Mutation in the -35 region of the promoter responsible for the initiation of T-558 indicates that the upstream promoter is essential for full expression of cryIIIA, although not sufficient. Deletion of the DNA region carrying the previously described cryIIIA promoter does not affect full expression of cryIIIA and does not modify the 5' end of T-129. Taken together, these results indicate that the 5' end of T-129 is not a trnascriptional start site. Therefore, we propose that T-129 results from the processing of the mRNA initiated at the upstream promoter (T-558), generating a stable mRNA with a 5' extremity at nucleotide position -129. From primer extension analysis and transcriptional fusions to lacZ, it appears that the upstream promoter is weakly but significantly expressed during the vegetative phase of growth, is activated at the onset of sporulation and remains active at least until t5. However, unlike the promoters of other cry genes, this promoter is similar to sigma A-dependent promoters rather than sporulation-specific promoters. This promoter may therefore be transcribed by the E sigma A form of RNA polymerase. Activation at the onset of sporulation could result from the disappearance of a repressor, or the appearance of a stationary-phase-specific activator.

Bacillus thuringiensis↗

Expression of chicken hepatic type I and type III iodothyronine deiodinases during embryonic development.

In embryonic chicken liver (ECL) two types of iodothyronine deiodinases are expressed: D1 and D3. D1 catalyzes the activation as well as the inactivation of thyroid hormone by outer and inner ring deiodination, respectively. D3 only catalyzes inner ring deiodination. D1 and D3 have been cloned from mammals and amphibians and shown to contain a selenocysteine (Sec) residue. We characterized chicken D1 and D3 complementary DNAs (cDNAs) and studied the expression of hepatic D1 and D3 messenger RNAs (mRNAs) during embryonic development. Oligonucleotides based on two amino acid sequences strongly conserved in the different deiodinases (NFGSCTSecP and YIEEAH) were used for reverse transcription-PCR of poly(A+) RNA isolated from embryonic day 17 (E17) chicken liver, resulting in the amplification of two 117-bp DNA fragments. Screening of an E17 chicken liver cDNA library with these probes led to the isolation of two cDNA clones, ECL1711 and ECL1715. The ECL1711 clone was 1360 bp long and lacked a translation start site. Sequence alignment showed that it shared highest sequence identity with D1s from other vertebrates and that the coding sequence probably lacked the first five nucleotides. An ATG start codon was engineered by site-directed mutagenesis, generating a mutant (ECL1711M) with four additional codons (coding for MGTR). The open reading frame of ECL1711M coded for a 249-amino acid protein showing 58-62% identity with mammalian D1s. An in-frame TGA codon was located at position 127, which is translated as Sec in the presence ofa Sec insertion sequence (SECIS) identified in the 3'-untranslated region. Enzyme activity expressed in COS-1 cells by transfection with ECL1711M showed the same catalytic, substrate, and inhibitor specificities as native chicken D1. The ECL1715 clone was 1366 bp long and also lacked a translation start site. Sequence alignment showed that it was most homologous with D3 from other species and that the coding sequence lacked approximately the first 46 nucleotides. The deduced amino acid sequence showed 62-72% identity with the D3 sequences from other species, including a putative Sec residue at a corresponding position. The 3'-untranslated region of ECL1715 also contained a SECIS element. These results indicate that ECL1711 and ECL1715 are near-full-length cDNA clones for chicken D1 and D3 selenoproteins, respectively. The ontogeny of D1 and D3 expression in chicken liver was studied between E14 and 1 day after hatching (C1). D1 activity showed a gradual increase from E14 until C1, whereas D1 mRNA level remained relatively constant. D3 activity and mRNA level were highly significantly correlated, showing an increase from E14 to E17 and a strong decrease thereafter. These results suggest that the regulation of chicken hepatic D3 expression during embryonic development occurs predominantly at the pretranslational level.

Amino Acid Sequence↗

M161Ag is a potent cytokine inducer with complement activating function (review).

We discovered a membrane-associated novel gene product expressed on some malignant human cells/cell lines undergoing apoptosis. This protein, named M161Ag, activated human complement and efficiently induced the pro-inflammatory cytokines IL-1 Beta , TNF-alpha and IL-6, and also IL-10 and IL-12 in human peripheral blood monocytes. M161Ag was a 43 kDa palmitoylated protein containing five amino acids encoded by TGA codons. These TGA codons were found to be translated into Trp, consistent with expression in prokaryotes including mitochondria and mycoplasma. The amino-terminal lipid was characteristic of prokaryote proteins participating in membrane anchoring. The M161Ag genomic clone contained a Pribnow box at the -35 and -10 promoter portions and the Shine-Dalgarno ribosomal binding site approximately 10 bp upstream of the translational start codon. The Mycoplasma fermentans origin of this protein was then confirmed by genomic Southern analysis; cells only infected with M. fermentans were positive for M161Ag. Thus, latent infection with M. fermentans allows tumor cells to produce M161Ag leading to activation of the host immune system. Here, we summarize bacterial proteins resembling M161Ag which may be candidates for therapeutic use if they exert immuno-regulatory functions.

Amino Acid Sequence↗

Mutations affecting translation of the bacteriophage T4 rIIB gene cloned in Escherichia coli.

Mutant ribosome binding sites of the bacteriophage T4 rIIB gene, resident on an 873 bp DNA fragment, were cloned into a plasmid vector as in-frame fusions to a reporter gene, beta-galactosidase. The collection of mutations included changes in the region 5' to the Shine/Dalgarno sequence, a mutation of the Shine/Dalgarno sequence, the alternate initiation codons GUG, AUA and ACG, and mutants in which several closely spaced initiation codons compete with each other on the same mRNA. The results show that the secondary structure variations we have installed 5' to the Shine/Dalgarno sequence have little effect on translation. GUG is essentially as good an initiator of translation as AUG when they are assayed on separate messages, but is outcompeted at least 50-fold in the sequence AUGUG. AUA and ACG are poor start codons, and are temperature sensitive. The initiation codon pair AUGAUA, in which the AUG is only two nucleotides from the Shine/Dalgarno sequence, displays a novel cold-sensitive phenotype.

Base Sequence↗

Efficiency of reinitiation of translation on human immunodeficiency virus type 1 mRNAs is determined by the length of the upstream open reading frame and by intercistronic distance.

In this study, we examined the mechanism of translation of the human immunodeficiency virus type 1 tat mRNA in eucaryotic cells. This mRNA contains the tat open reading frame (ORF), followed by rev and nef ORFs, but only the first ORF, encoding tat, is efficiently translated. Introduction of premature stop codons in the tat ORF resulted in efficient translation of the downstream rev ORF. We show that the degree of inhibition of translation of rev is proportional to the length of the upstream tat ORF. An upstream ORF spanning 84 nucleotides was predicted to inhibit 50% of the ribosomes from initiating translation at downstream AUGs. Interestingly, the distance between the upstream ORF and the start codon of the second ORF also played a role in efficiency of downstream translation initiation. It remains to be investigated if these conclusions relate to translation of mRNAs other than human immunodeficiency virus type 1 mRNAs. The strong inhibition of rev translation exerted by the presence of the tat ORF may reflect the different roles of Tat and Rev in the viral life cycle. Tat acts early to induce high production of all viral mRNAs. Rev induces a switch from the early to the late phase of the viral life cycle, resulting in production of viral structural proteins and virions. Premature Rev production may result in entrance into the late phase in the presence of suboptimal levels of viral mRNAs coding for structural proteins, resulting in inefficient virus production.

Base Sequence↗

Regulation of the Escherichia coli lrp gene.

Lrp (leucine-responsive regulatory protein) is a major Escherichia coli regulatory protein which regulates expression of a number of operons, some negatively and some positively. This work relates to a characterization of lrp, the gene encoding Lrp. Nucleotide sequencing established that the coding regions of lrp and trxB (encoding thioredoxin reductase) are separated by 543 bp and that the two genes are transcribed in opposite directions. In addition, we used primer extension, deletion analyses, and lrp-lacZ transcriptional fusions to delineate the promoter and regulatory region of the lrp operon. The lrp promoter is located 267 nucleotides upstream of the translational start codon of the lrp gene. In comparison with a wild-type strain, expression of the lrp operon was increased about 3-fold in a strain lacking Lrp and decreased about 10-fold in a strain overproducing Lrp. As observed from DNA mobility shift and DNase I footprinting analyses, Lrp binds to one or more sites within the region -80 to -32 relative to the start point of lrp transcription. A mutational analysis indicated that this same region is at least partly required for repression of lrp expression in vivo. These results demonstrate that autogenous regulation of lrp involves Lrp acting directly to cause repression of lrp transcription.

Bacterial Proteins↗

Regulation of a Bacteroides operon that controls excision and transfer of the conjugative transposon CTnDOT.

CTnDOT is a conjugative transposon (CTn) that is found in many Bacteroides strains. Transfer of CTnDOT is stimulated 100- to 1,000-fold if the cells are first exposed to tetracycline (TET). Both excision and transfer of CTnDOT are stimulated by TET. An operon that contains a TET resistance gene, tetQ, and two regulatory genes, rteA and rteB, is essential for control of excision and transfer functions. At first, it appeared that RteA and RteB, which are members of a two-component regulatory system, might be directly responsible for the TET effect. We show here, however, that neither RteA nor RteB affected expression of the operon. TetQ, a ribosome protection type of TET resistance protein, actually reduced operon expression, possibly by interacting with ribosomes that are translating the tetQ message. Fusions of tetQ with a reporter gene, uidA, were only expressed at a high level when the fusion was cloned in frame with the first six codons of tetQ. However, out of frame fusions or fusions ending at the other five codons of tetQ showed much lower expression of the uidA gene. Moreover, reverse transcription-PCR amplification of tetQ mRNA revealed that despite the fact that the uidA gene product, beta-glucuronidase (GUS), was produced only when the cells were exposed to TET, tetQ mRNA was produced in both the presence and absence of TET. Computer analysis of the region upstream of the tetQ start codon predicted that the mRNA in this region could form a complex RNA hairpin structure that would prevent access of ribosomes to the ribosome binding site. Mutations that abolished base pairing in the stem that formed the base of this putative hairpin structure made GUS production as high in the absence of TET as in TET-stimulated cells. Compensatory mutations that restored the hairpin structure led to a return of regulated production of GUS. Thus, the tetQ-rteA-rteB operon appears to be regulated by a translational attenuation mechanism.

Amino Acid Sequence↗

Accuracy improvement for identifying translation initiation sites in microbial genomes.

MOTIVATION: At present the computational gene identification methods in microbial genomes have a high prediction accuracy of verified translation termination site (3' end), but a much lower accuracy of the translation initiation site (TIS, 5' end). The latter is important to the analysis and the understanding of the putative protein of a gene and the regulatory machinery of the translation. Improving the accuracy of prediction of TIS is one of the remaining open problems. RESULTS: In this paper, we develop a four-component statistical model to describe the TIS of prokaryotic genes. The model incorporates several features with biological meanings, including the correlation between translation termination site and TIS of genes, the sequence content around the start codon; the sequence content of the consensus signal related to ribosomal binding sites (RBSs), and the correlation between TIS and the upstream consensus signal. An entirely non-supervised training system is constructed, which takes as input a set of annotated coding open reading frames (ORFs) by any gene finder, and gives as output a set of organism-specific parameters (without any prior knowledge or empirical constants and formulas). The novel algorithm is tested on a set of reliable datasets of genes from Escherichia coli and Bacillus subtillis. MED-Start may correctly predict 95.4% of the start sites of 195 experimentally confirmed E.coli genes, 96.6% of 58 reliable B.subtillis genes. Moreover, the test results indicate that the algorithm gives higher accuracy for more reliable datasets, and is robust to the variation of gene length. MED-Start may be used as a postprocessor for a gene finder. After processing by our program, the improvement of gene start prediction of gene finder system is remarkable, e.g. the accuracy of TIS predicted by MED 1.0 increases from 61.7 to 91.5% for 854 E.coli verified genes, while that by GLIMMER 2.02 increases from 63.2 to 92.0% for the same dataset. These results show that our algorithm is one of the most accurate methods to identify TIS of prokaryotic genomes. AVAILABILITY: The program MED-Start can be accessed through the website of CTB at Peking University: http://ctb.pku.edu.cn/main/SheGroup/MED_Start.htm.

Algorithms↗

Lactococcus lactis glyceraldehyde-3-phosphate dehydrogenase gene, gap: further evidence for strongly biased codon usage in glycolytic pathway genes.

The gene gap, encoding glyceraldehyde-3-phosphate dehydrogenase (EC 1.2.1.12), was isolated from a genomic library of Lactococcus lactis LM0230 DNA. Plasmids containing the L. lactis gene were able to complement a gap mutant of Escherichia coli. The nucleotide sequence of gap predicted a polypeptide chain of 337 amino acids for the enzyme and a subunit molecular mass of 36,043. The codon usage in gap and four other glycolytic genes from L. lactis showed a high degree of bias, when compared with 84 other chromosomal genes. Northern blot analysis of total L. lactis RNA showed that gap hybridized strongly with a 1.3 kb transcript. The 5' end of the transcript was determined by primer extension analysis to be a C located 35 bp upstream from the gap start codon. These transcript analyses, and the orientation of the open reading frames in the DNA flanking gap, indicated that in L. lactis gap is expressed on a monocistronic transcript. Nucleotide sequencing indicated that the DNA adjacent to gap did not encode other glycolytic pathway enzymes. The DNA sequence flanking gap contained two open reading frames (ORF156 and ORF211) of unknown function. The 3' end of a clpA homologue was identified in the sequence upstream of ORF156. The location of gap on the L. lactis DL11 chromosome map was determined to be between map coordinates 0.530 and 0.660.

Amino Acid Sequence↗