PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bacterial coding sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Easy selection of recombinant clones.

We have developed a polymerase chain reaction (PCR)-based procedure to facilitate the selection of recombinant clones. The insert to be cloned is ligated to an antibiotic resistance marker. The ligation product is amplified by PCR, followed by standard cloning procedure into a bacterial vector. The selection for the antibiotic resistance coded by the PCR product ensures 100% insertion frequency, eliminating the screening of the transformants.

Bacillus subtilis↗

Porcine growth hormone: molecular cloning of cDNA and expression in bacterial and mammalian cells.

Porcine growth hormone (PGH) precursor cDNAs were cloned from a pituitary cDNA library constructed in lambda gt11 by immunoscreening. One of the three clones characterized contained an entire nucleotide sequence for the 216-amino-acid precursor molecule. The deduced amino-acid sequence of PGH confirmed the sequence previously reported for that of the genomic DNA of PGH except for one base difference in the coding sequence. Expression of the full-length PGH cDNA was achieved in bacteria and mammalian cells. The mammalian cell line, COS-1, produced the GH molecule which processed the signal peptide and had the same molecular weight as standard PGH, in contrast to the higher molecular weight of the bacterial product. Radioimmunoassay of the recombinant PGH produced in COS-1 cells also revealed an inhibition curve similar to that of the standard PGH.

Amino Acid Sequence↗

Analysis of Bacillus megaterium lipolytic system and cloning of LipA, a novel subfamily I.4 bacterial lipase.

The lipolytic system of Bacillus megaterium 370 was investigated, showing the existence of at least two secreted lipases and a cell-bound esterase. A gene coding for an extracellular lipase was isolated and cloned in Escherichia coli. The cloned enzyme displayed high activity on short to medium chain length (C(4)-C(8)) substrates, and poor activity on C(18) substrates. On the basis of amino acid sequence homology, the cloned lipase was classified into subfamily I.4 of bacterial lipases.

Bacillus megaterium↗

Sequence analysis of the DNA encoding the Eco RI endonuclease and methylase.

The Eco RI endonuclease and methylase recognize the same hexanucleotide substrate sequence. We have determined the sequence of a fragment of DNA which encodes these enzymes using the chain-termination method of Sanger (Sanger, F., Nicklen, S., and Coulson, A. R. (1977) Proc. Natl. Acad. Sci. U. S. A. 74, 5463-5467). The amino acid sequences of both enzymes were derived from the DNA sequence. The coding regions selected include the only open translational frames of sufficient length to accommodate the enzymes. They coincide with previously established gene boundaries and orientation. The predicted amino acid sequences correlate well with analyses of the purified protein. Comparison of the nucleotide and protein sequences reveals no homology between the endonuclease and methylase which might provide insight into the origin of the restriction-modification system or the mechanism of common substrate recognition. Based on secondary structure predictions, the two enzymes also have grossly different molecular architecture. The base composition of the sequence is 65% A + T, and the codon usage is significantly different from that observed in several Escherichia coli chromosomal genes. In some cases, frequently selected codons are recognized by minor tRNA species. A spontaneous mutation in the endonuclease gene was isolated. Serine replaces arginine at residue 187. In crude extracts, Eco RI specific cleavage is approximately 0.3% wild type.

Amino Acid Sequence↗

Identification of a second flagellin gene and functional characterization of a sigma70-like promoter upstream of a Leptospira borgpetersenii flaB gene.

Leptospira borgpetersenii, one of the causative agents of leptospirosis in both animals and humans, is a bacterial pathogen with characteristic motility that is mediated by the rotation of two periplasmic flagella (PF). The flaB gene coding for a core polypeptide subunit of PF was previously characterized by sequence analysis of its open reading frame (ORF) (M. Lin, J Biochem Mol Biol Biophys 2:181-187, 1999). The present study was undertaken to isolate and clone the uncharacterized sequence upstream of the flaB gene by using a PCR-based genome walking procedure. This has resulted in a 1470-bp genomic DNA sequence in which an 846-bp ORF coding for a 281-amino acid polypeptide (31.3 kDa) is identified 455 bp upstream from the flaB start codon. The encoded protein exhibits 72% amino acid identity to the deduced FlaB protein sequence of L. borgpetersenii and a high degree of sequence homology to the FlaB proteins of other spirochaetes. This has demonstrated for the first time that a second flaB gene homolog is present in a Leptospira species. The newly identified gene is designated flaB1, and the previously cloned flaB renamed flaB2. Within the intergenic sequence between flaB1 and flaB2, a potential stem-loop structure (12-bp inverted repeats) was identified 25 bp downstream of the flaB1 stop codon; this could serve as a transcription terminator for the flaB1 mRNA. Three E. coli-like promoter regions (I, II, and III) for binding Esigma(70), a regulatory sequence uncommonly found in flagellar genes, were predicted upstream of the flaB2 ORF. Only promoter region II contains a promoter that is functional in E. coli, as revealed at phenotypic and transcriptional levels by its capability of directing the expression of the chloramphenicol acetyltransferase (CAT) gene in the promoter probe vector pKK232-8. These observations may suggest that flaB1 and flaB2 are transcribed separately and do not form a transcriptional operon controlled by a single promoter.

Bacterial Proteins↗

The napEDABC gene cluster encoding the periplasmic nitrate reductase system of Thiosphaera pantotropha.

The napEDABC locus coding for the periplasmic nitrate reductase of Thiosphaera pantotropha has been cloned and sequenced. The large and small subunits of the enzyme are coded by napA and napB. The sequence of NapA indicates that this protein binds the GMP-conjugated form of the molybdopterin cofactor. Cysteine-181 is proposed to ligate the molybdenum atom. It is inferred that the active site of the periplasmic nitrate reductase is structurally related to those of the molybdenum-dependent formate dehydrogenases and bacterial assimilatory nitrate reductases, but is distinct from that of the membrane-bound respiratory nitrate reductases. A four-cysteine motif at the N-terminus of NapA binds a [4Fe-4S] cluster. The DNA- and protein-derived primary sequence of NapB confirm that this protein is a dihaem c-type cytochrome and, together with spectroscopic data, indicate that both NapB haems have bis-histidine ligation. napC is predicted to code for a membrane-anchored tetrahaem c-type cytochrome that shows sequence similarity to the NirT cytochrome c family. NapC may be the direct electron donor to the NapAB complex. napD is predicted to encode a soluble cytoplasmic protein and napE a monotopic integral membrane protein, napDABC genes can be discerned at the aeg-46.5 locus of Escherichia coli K-12, suggesting that this operon encodes a periplasmic nitrate reductase system, while napD and napC are identified adjacent to the napAB genes of Alcaligenes eutrophus H16.

Amino Acid Sequence↗

Mitochondrial and cytoplasmic fumarases in Saccharomyces cerevisiae are encoded by a single nuclear gene FUM1.

Respiratory defective pet mutants of Saccharomyces cerevisiae assigned to complementation group G5 are deficient in fumarase. A representative mutant from this complementation group was used to clone a nuclear gene (FUM1) whose sequence encodes a protein homologous to bacterial fumarase. Based on the primary structure homology and the elevated levels of fumarase in transformants harboring FUM1 on a multicopy plasmid, this gene is concluded to code for yeast fumarase. In wild type yeast, fumarase is detected in both mitochondria and the soluble postribosomal protein fraction. Several lines of evidence indicate that the two compartmentally distinct fumarases are isoenzyme products of FUM1. Mutations in FUM1 simultaneously abolish both activities. Transformation of a fumarase mutant with a plasmid containing FUM1 leads to increased fumarase activity in mitochondria and in the postribosomal supernatant fraction. Transformation of the same mutant with a plasmid construct in which the region of FUM1 coding for the amino-terminal 17 amino acids of fumarase is deleted results in a preferential increase of nonmitochondrial fumarase. Northern and S1 nuclease analysis of fumarase transcripts in wild type yeast and in a mutant transformed with FUM1 on an episomal plasmid indicate that the gene is transcribed from multiple start sites, some of which are located inside the coding sequence. The major transcript presumed to code for mitochondrial fumarase has a 5'-untranslated leader of 185 nucleotides. The most abundant shorter transcripts have 5' termini from 57 to 68 nucleotides downstream of the first ATG; their translation products lacking the amino-terminal mitochondrial import signal are proposed to target fumarase to the cytoplasm.

Amino Acid Sequence↗

Flexibility of the genetic code with respect to DNA structure.

MOTIVATION: The primary function of DNA is to carry genetic information through the genetic code. DNA, however, contains a variety of other signals related, for instance, to reading frame, codon bias, pairwise codon bias, splice sites and transcription regulation, nucleosome positioning and DNA structure. Here we study the relationship between the genetic code and DNA structure and address two questions. First, to which degree does the degeneracy of the genetic code and the acceptable amino acid substitution patterns allow for the superimposition of DNA structural signals to protein coding sequences? Second, is the origin or evolution of the genetic code likely to have been constrained by DNA structure? RESULTS: We develop an index for code flexibility with respect to DNA structure. Using five different di- or tri-nucleotide models of sequence-dependent DNA structure, we show that the standard genetic code provides a fair level of flexibility at the level of broad amino acid categories. Thus the code generally allows for the superimposition of any structural signal on any protein-coding sequence, through amino acid substitution. The flexibility observed at the level of single amino acids allows only for the superimposition of punctual and loosely positioned signals to conserved amino acid sequences. The degree of flexibility of the genetic code is low or average with respect to several classes of alternative codes. This result is consistent with the view that DNA structure is not likely to have played a significant role in the origin and evolution of the genetic code.

Amino Acids↗

Folding of an enzyme into an active conformation while bound as peptidyl-tRNA to the ribosome.

Rhodanese bound to bacterial ribosomes as peptidyl-tRNA can be folded into an enzymatically active conformation by generating C-terminal extensions of the wild-type enzyme. Rhodanese was synthesized by coupled transcription/translation in a cell-free Escherichia coli system from plasmids containing the coding sequences for the wild-type enzyme or its C-terminally extended mutants. Two proteins with extensions of 23 amino acids or longer were enzymatically active while bound to the ribosomes whereas wild-type protein and a 13-amino acid extension were not. All forms of the enzyme were active after termination and release of the full-length protein from the ribosomes. All five of the bacterial chaperones were required to substantially increase the specific enzymatic activity of the extended rhodanese while the nascent protein was bound to ribosomes. The results provide direct support for the hypothesis that proteins acquire tertiary structure as they are formed in ribosomes.

Cell-Free System↗

Biosynthesis of bacterial glycogen. Primary structure of Escherichia coli 1,4-alpha-D-glucan:1,4-alpha-D-glucan 6-alpha-D-(1, 4-alpha-D-glucano)-transferase as deduced from the nucleotide sequence of the glg B gene.

The nucleotide sequence of the glg B gene, coding for branching enzyme (EC 2.4.1.18), was elucidated. It consists of 2181 base pairs specifying a protein of 727 amino acids. The deduced amino acid sequence was consistent with the amino acid analysis that was obtained with the pure protein as well as with the molecular weight determined from sodium dodecyl sulfate-gel electrophoresis. The deduced amino acid sequence was also consistent with the amino-terminal amino acid sequence and the amino acid sequence analysis of various peptides obtained from CNBr degradation of purified branching enzyme.

1,4-alpha-Glucan Branching Enzyme↗

Modeling and predicting transcriptional units of Escherichia coli genes using hidden Markov models.

MOTIVATION: The hidden Markov model (HMM) is a valuable technique for gene-finding, especially because its flexibility enables the inclusion of various sequence features. Recent programs for bacterial gene-finding include the information of ribosomal binding site (RBS) to improve the recognition accuracy of the start codon, using this feature. We report here our attempt to extend the model into the total transcriptional unit, enabling the prediction of operon structures. RESULTS: First, we improved the prediction accuracy of coding sequences (CDSs) by employing the models of 'typical', 'atypical' and 'negative (false-positive)' classes as well as the models of RBS and its downstream spacer. The sensitivity of exactly predicting the 204 experimentally confirmed CDSs reached 90.2% in an objective test. Based on the prediction result of CDSs, the positions of the promoters and terminators were predicted. Our model could exactly recognize 60% of 390 known transcriptional units. Thus, the accuracy and significance of this prediction problem is far from trivial. We would like to propose this problem as an open theme in bioinformatics because the ongoing or planned post-sequencing projects will produce much data for future improvements.

Algorithms↗

Low-resolution sequencing of Rhodobacter sphaeroides 2.4.1T: chromosome II is a true chromosome.

The photosynthetic bacterium Rhodobacter sphaeroides 2.4.1T has two chromosomes, CI (approximately 3.0 Mb) and CII (approximately 0.9 Mb). In this study a low-redundancy sequencing strategy was adopted to analyse 23 out of 47 cosmids from an ordered CII library. The sum of the lengths of these 23 cosmid inserts was approximately 495 kb, which comprised approximately 417 kb of unique DNA. A total of 1145 sequencing runs was carried out, with each run generating 559 +/- 268 bases of sequence to give approximately 640 kb of total sequence. After editing, approximately 2.8% bases per run were estimated to be ambiguous. After the removal of vector and Escherichia coli sequences, the remaining approximately 565 kb of R. sphaeroides sequences were assembled, generating approximately 291 kb of unique sequences. BLASTX analysis of these unique sequences suggested that approximately 131 kb (45% of the unique sequence) had matches to either known genes, or database ORFs of hypothetical or unknown function (dORFs). A total of 144 strong matches to the database was found; 101 of these matches represented genes encoding a wide variety of functions, e.g. amino acid biosynthesis, photosynthesis, nutrient transport, and various regulatory functions. Two rRNA operons (rrnB and rrnC) and five tRNAs were also identified. The remaining 160 kb of DNA sequence which did not yield database matches was then analysed using CODONPREFERENCE from the GCG package. This analysis suggested that 122 kb (42% of the total unique DNA sequence) could encode putative ORFs (pORFs), with the remaining 38 kb (13%) possibly representing non-coding intergenic DNA. From the data so far obtained, CII does not appear to be specialized for encoding any particular metabolic function, physiological state or growth condition. These data suggest that CII contains genes which are functionally as diverse as those found on any other bacterial chromosome and also contains sequences (pORFs), which may prove to be unique to this organism.

Base Sequence↗

Bacterial molecular phylogeny using supertree approach.

It has been claimed that complete genome sequences would clarify phylogenetic relationships between organisms but, up to now, no satisfying approach has been proposed to use efficiently these data. For instance, if the coding of presence or absence of genes in complete genomes gives interesting results, it does not take into account the phylogenetic information contained in sequences and ignores hidden paralogy by using a similarity-based definition of orthology. Also, concatenation of sequences of different genes takes hardly in consideration the specific evolutionary rate of each gene. At last, building a consensus tree is strongly limited by the low number of genes shared among all organisms. Here, we use a new method based on supertree construction, which permits to cumulate in one supertree the information and statistical support of hundreds of trees from orthologous gene families and to build the phylogeny of 33 prokaryotes and four eukaryotes with completely sequenced genomes. This approach gives a robust supertree, which demonstrates that a phylogeny of prokaryotic species is conceivable and challenges the hypothesis of a thermophilic origin of bacteria and present-day life. The results are compatible with the hypothesis of a core of genes for which lateral transfers are rare but they raise doubts on the widely admitted "complexity hypothesis" which predicts that this core is mainly implicated in informational processes.

Bacteria↗

Cloning and sequence analysis of the gene encoding invertase (INV1) from the yeast Candida utilis.

The gene INV1 encoding invertase from the yeast Candida utilis has been cloned using a homologous PCR hybridization probe, amplified with two sets of degenerate primers designed considering sequence comparisons between yeast invertases. The cloned gene was sequenced and found to encode a polypeptide of 533 amino acids that contain a 26 amino-acid signal peptide and 12 potential N-glycosylation sites. The nucleotide sequences of the 5' and 3' non-coding regions were found to contain motifs probably involved in initiation, regulation and termination of gene transcription. The amino-acid sequence shows significant identity with other yeast, bacterial and plant beta-fructofuranosidases. The INV1 gene from C. utilis was able to complement functionally the suc2 mutation of S. cerevisiae.

Amino Acid Sequence↗

Alteration of chain length selectivity of a Rhizopus delemar lipase through site-directed mutagenesis.

The coding sequences of the Rhizopus delemar lipase and prolipase were altered by oligonucleotide-directed mutagenesis to introduce amino acid substitutions. The resulting mutant enzymes, synthesized by the bacterial host Escherichia coli BL21 (DE3), were tested for their ability to hydrolyze the triglycerides triolein (TO), tricaprylin (TC) and tributyrin (TB). Mutagenesis and lipase gene expression were carried out using plasmid vectors derived from previously described recombinant plasmids [Joerger and Haas (1993) Lipids 28, 81-88] by introduction of the origin of replication of bacteriophage f1. Substitution of threonine 83 (thr83), a residue thought to be involved in oxyanion binding, by alanine essentially eliminated lipolytic activity toward all substrates examined (TB, TO and TC). Replacement of thr83 with serine caused from two- to sevenfold reductions in the activity toward these substrates. Introduction of tryptophan (trp) at position 89, where such a residue is found in closely related fungal lipases, reduced the specific activity toward the three triglyceride substrates. For the mutagenesis of residues in the predicted acyl chain binding groove, mutagenic primers were designed to cause the replacement of a specific codon within the prolipase gene with codons for all other amino acids. Phenylalanine 95 (phe95), phe112, valine 206 (val206) and val209, were targeted. A phenotypic screen was successfully employed to identify cells producing prolipase with altered preference for olive oil, TC or TB. In assays involving equimolar mixtures of the three triglycerides, a prolipase with a phe95-->aspartate mutation showed an almost twofold increase in the relative activity toward TC. Substitution of trp for phe112 caused an almost threefold decrease in the relative preference for TC, but elevated relative TB hydrolysis. Replacement of val209 with trp resulted in an enzyme with a two- and fourfold enhanced preference for TC and TB, respectively.

Base Sequence↗

The Alcaligenes eutrophus hemN gene encoding the oxygen-independent coproporphyrinogen III oxidase, is required for heme biosynthesis during anaerobic growth.

The insertion mutant HF231 of Alcaligenes eutrophus H16 failed to grow anaerobically on nitrate and nitrite. When grown under oxygen limitation, mutant HF231 specifically excreted coproporphyrin III, an intermediate of heme biosynthesis. With the help of a Tn5-labeled fragment, we identified and cloned the corresponding wild-type fragment. Sequence analysis of the mutant locus revealed an open reading frame consisting of 1,473 bp, predicting a protein of 491 amino acids that corresponds to a size of 54.2 kDa. In the non-coding upstream region, consensus elements that are indicative for binding sites of the anaerobic transcriptional regulator Fnr were identified. The deduced polypeptide showed extensive sequence similarity with various bacterial oxygen-independent coproporphyrinogen III oxidases designated HemN. HemN catalyzes the oxidative decarboxylation of coproporphyrinogen III to yield protoporphyrinogen IX. Anaerobic growth on nitrate and nitrite of mutant HF231 was restored by introducing the hemN gene of A. eutrophus or of Pseudomonas aeruginosa on a broad-host-range vector. Likewise, the A. eutrophus hemN complemented heme biosynthesis of a Salmonella typhimurium hemF/hemN double mutant during anaerobic and aerobic growth. Analysis of a transcriptional lacZ gene fusion showed that expression of hemN in A. eutrophus is nitrate-independent and repressed by oxygen.

Alcaligenes↗

Differential excision patterns of the En-transposable element at the A2 locus in maize relate to the insertion site.

Defined mutant alleles with resident transposons display characteristic patterns of germinal and somatic reversion, and heritable changes in the timing and frequency of reversions, which have been termed "change of state" by McClintock, constantly arise. Several mechanisms were proposed to account for these changes. They may be ascribed to the structure and composition of the elements themselves (composition hypothesis) or to their location (position hypothesis). In the current study, insertion positions were determined for three autonomous En-controlled mutable alleles of the A2 locus in maize that show different somatic reversion patterns. A relationship was observed between En insertion positions in the single coding region of the intronless A2 gene and anthocyanin variegation patterns in the aleurone. An insertion in the 5' region of the coding sequence produced a very late somatic variegation pattern, whereas two early variegation patterns were caused by En insertions in the 3' region of the coding sequence.

Alleles↗

Opacity genes in Neisseria gonorrhoeae: control of phase and antigenic variation.

The chromosome of N. gonorrhoeae contains several complete expression genes coding for variant opacity proteins. DNA sequence analysis of two opacity genes derived from the same locus (opaE1) of two isogenic gonococcal variants reveals common and variable regions in these genes. Genomic blotting experiments using synthetic probes suggest gene conversion as a principle for the assembly of variant sequence information in opacity genes. The 5' region of opacity genes is composed of identical pentameric pyrimidine units (CTCTT) encoding the hydrophobic portion of the opacity leader peptide. This coding repeat is variable in a given locus with respect of the number of pentameric units. While all expression loci in a single cell are constitutively transcribed, the production of opacity proteins is determined by the coding repeat sequence on the translational level.

Amino Acid Sequence↗