PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bacterial coding sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Genome sequence of Symbiobacterium thermophilum, an uncultivable bacterium that depends on microbial commensalism.

Symbiobacterium thermophilum is an uncultivable bacterium isolated from compost that depends on microbial commensalism. The 16S ribosomal DNA-based phylogeny suggests that this bacterium belongs to an unknown taxon in the Gram-positive bacterial cluster. Here, we describe the 3.57 Mb genome sequence of S.thermophilum. The genome consists of 3338 protein-coding sequences, out of which 2082 have functional assignments. Despite the high G + C content (68.7%), the genome is closest to that of Firmicutes, a phylum consisting of low G + C Gram-positive bacteria. This provides evidence for the presence of an undefined category in the Gram-positive bacterial group. The presence of both spo and related genes and microscopic observation indicate that S.thermophilum is the first high G + C organism that forms endospores. The S.thermophilum genome is also characterized by the widespread insertion of class C group II introns, which are oriented in the same direction as chromosomal replication. The genome has many membrane transporters, a number of which are involved in the uptake of peptides and amino acids. The genes involved in primary metabolism are largely identified, except those that code several biosynthetic enzymes and carbonic anhydrase. The organism also has a variety of respiratory systems including Nap nitrate reductase, which has been found only in Gram-negative bacteria. Overall, these features suggest that S.thermophilum is adaptable to and thus lives in various environments, such that its growth requirement could be a substance or a physiological condition that is generally available in the natural environment rather than a highly specific substance that is present only in a limited niche. The genomic information from S.thermophilum offers new insights into microbial diversity and evolutionary sciences, and provides a framework for characterizing the molecular basis underlying microbial commensalism.

Actinobacteria↗

A new series of trpE vectors that enable high expression of nonfusion proteins in bacteria.

Expression of recombinant proteins in bacteria has facilitated the characterization of many gene products. However, the biochemical characterization of recombinant proteins is limited since the bacterially expressed proteins are often synthesized as fusion polypeptides. The presence of bacterial sequences in fusion proteins further limits the use of these proteins for generating antibodies since the bacterial sequences are also antigenic. We describe two new bacterial expression vectors based on the pATH series of plasmids. These vectors were made by precisely deleting all of the trpE coding sequences found in pATH. The new vectors have enabled us to express eukaryotic genes as nonfusion polypeptides. These altered plasmids can be used to insert any DNA sequence of interest through a multiple cloning site located just 3' of an ATG start codon. Protein expression is still under the control of the trp operon and is carried out at great efficiency when the bacteria are tryptophan deprived. Studies presented here test the expression system with neurofilament subunits, NF-L and NF-H. Large amounts of recombinant nonfusion proteins were produced. Also, a time course of induction shows that the production of the nonfusion proteins was under the control of the trp operon which is readily inducible after tryptophan starvation and addition of indoleacrylic acid. These vectors may be useful for the overexpression of many proteins in a form closely approximating their native state.

Amino Acid Sequence↗

[Characteristics of the context shift in the frequency of synonymic codons in Escherichia coli].

We have demonstrated, that coding regions of E. coli DNA exhibit the non-random shifts of codon usage frequency depending both on the type of the 3' nucleotide adjacent to the codon and the degree of gene expression. Analysis of primary data--statistics of tetranucleotide occurrences--was performed by the techniques of contingency tables. The results of the investigation allowed us to suggest that the phenomenon observed is connected with the influence of the 3'-adjacent nucleotide on the level of missens-errors and another types of inaccuracy during translation. The specific advantages of such mutations in the third position of the codon are based on the adaptation of the codon to the 3' context in order to increase the efficiency of translation.

Base Sequence↗

The Saccharomyces cerevisiae MGT1 DNA repair methyltransferase gene: its promoter and entire coding sequence, regulation and in vivo biological functions.

We previously cloned a yeast DNA fragment that, when fused with the bacterial lacZ promoter, produced O6-methylguanine DNA repair methyltransferase (MGT1) activity and alkylation resistance in Escherichia coli (Xiao et al., EMBO J. 10,2179). Here we describe the isolation of the entire MGT1 gene and its promoter by sequence directed chromosome integration and walking. The MGT1 promoter was fused to a lacZ reporter gene to study how MGT1 expression is controlled. MGT1 is not induced by alkylating agents, nor is it induced by other DNA damaging agents such as UV light. However, deletion analysis defined an upstream repression sequence, whose removal dramatically increased basal level gene expression. The polypeptide deduced from the complete MGT1 sequence contained 18 more N-terminal amino acids than that previously determined; the role of these 18 amino acids, which harbored a potential nuclear localization signal, was explored. The MGT1 gene was also cloned under the GAL1 promoter, so that MTase levels could be manipulated, and we examined MGT1 function in a MTase deficient yeast strain (mgt1). The extent of resistance to both alkylation-induced mutation and cell killing directly correlated with MTase levels. Finally we show that mgt1 S.cerevisiae has a higher rate of spontaneous mutation than wild type cells, indicating that there is an endogenous source of DNA alkylation damage in these eukaryotic cells and that one of the in vivo roles of MGT1 is to limit spontaneous mutations.

Amino Acid Sequence↗

Molecular microbial diversity in soils from eastern Amazonia: evidence for unusual microorganisms and microbial population shifts associated with deforestation.

Although the Amazon Basin is well known for its diversity of flora and fauna, this report represents the first description of the microbial diversity in Amazonian soils involving a culture-independent approach. Among the 100 sequences of genes coding for small-subunit rRNA obtained by PCR amplification with universal small-subunit rRNA primers, 98 were bacterial and 2 were archaeal. No duplicate sequences were found, and none of the sequences had been previously described. Eighteen percent of the bacterial sequences could not be classified in any known bacterial kingdom. Two sequences may represent a unique branch between the vast majority of bacteria and the deeply branching, predominantly thermophilic bacteria. Five sequences formed a clade that may represent a novel group within the class Proteobacteria. In addition, rRNA intergenic spacer analysis was used to show significant microbial population differences between a mature forest soil and an adjacent pasture soil.

Bacteria↗

Efficient bacterial expression of bovine and porcine growth hormones.

cDNAs prepared using poly(A)mRNA from pituitaries and containing the coding sequences for bovine and porcine growth hormones (bGH and pGH) were cloned in bacteria. The primary structures of the peptide hormones derived from the nucleotide sequences of the respective cDNAs show approximately 90% homology. The cloned cDNAs were modified using synthetic DNA to construct expression vectors for efficient bacterial production of the mature animal growth hormones.

Animals↗

A minimal gene set for cellular life derived by comparison of complete bacterial genomes.

The recently sequenced genome of the parasitic bacterium Mycoplasma genitalium contains only 468 identified protein-coding genes that have been dubbed a minimal gene complement [Fraser, C.M., Gocayne, J.D., White, O., Adams, M.D., Clayton, R.A., et al. (1995) Science 270, 397-403]. Although the M. genitalium gene complement is indeed the smallest among known cellular life forms, there is no evidence that it is the minimal self-sufficient gene set. To derive such a set, we compared the 468 predicted M. genitalium protein sequences with the 1703 protein sequences encoded by the other completely sequenced small bacterial genome, that of Haemophilus influenzae. M. genitalium and H. influenzae belong to two ancient bacterial lineages, i.e., Gram-positive and Gram-negative bacteria, respectively. Therefore, the genes that are conserved in these two bacteria are almost certainly essential for cellular function. It is this category of genes that is most likely to approximate the minimal gene set. We found that 240 M. genitalium genes have orthologs among the genes of H. influenzae. This collection of genes falls short of comprising the minimal set as some enzymes responsible for intermediate steps in essential pathways are missing. The apparent reason for this is the phenomenon that we call nonorthologous gene displacement when the same function is fulfilled by nonorthologous proteins in two organisms. We identified 22 nonorthologous displacements and supplemented the set of orthologs with the respective M. genitalium genes. After examining the resulting list of 262 genes for possible functional redundancy and for the presence of apparently parasite-specific genes, 6 genes were removed. We suggest that the remaining 256 genes are close to the minimal gene set that is necessary and sufficient to sustain the existence of a modern-type cell. Most of the proteins encoded by the genes from the minimal set have eukaryotic or archaeal homologs but seven key proteins of DNA replication do not. We speculate that the last common ancestor of the three primary kingdoms had an RNA genome. Possibilities are explored to further reduce the minimal set to model a primitive cell that might have existed at a very early stage of life evolution.

Amino Acid Sequence↗

Long-range correlations between DNA bending sites: relation to the structure and dynamics of nucleosomes.

It has been established that the precise positioning of nucleosomes on genomic DNA can be achieved, at least for a minority of them, through sequence-dependent processes. However, to what extent DNA sequences play a role in the positioning of the major part of nucleosomes is still debated. The aim of the present study is to examine to what extent long-range correlations (LRC) are related to the presence of nucleosomes. Using the wavelet transform technique, we perform a comparative analysis of the DNA text and of the corresponding bending profiles generated with curvature tables based on nucleosome positioning data. The exploration of a number of eukaryotic and bacterial genomes through the optics of the so-called "wavelet transform microscope" reveals a characteristic scale of 100-200 bp that separates two regimes of different LRC. Here, we focus on the existence of LRC in the small-scale regime (10-200 bp) which are actually observed in eukaryotic genomes, in contrast to their absence in eubacterial genomes. Analysis of viral DNA genomes shows that, like their host's genomes, eukaryotic viruses present LRC but eubacterial viruses do not. There is one exception for genomes of poxviruses (Vaccinia and Melamoplus sanguinipes) which do not replicate in the cell nucleus and do not exhibit LRC. No small-scale LRC are detected in the genomes of all examined RNA viruses, with the exception of retroviruses. These results together with the observation of LRC between particular sequence motifs known to participate in the formation of nucleosomes (e.g. AA dinucleotides) strongly suggest that the 10-200 bp LRC are a signature of the sequence-dependence of nucleosome positioning. Finally, we discuss possible interpretations of these LRC in terms of the physical mechanisms that might govern the positioning and the dynamics of the nucleosomes along the DNA chain through cooperative processes.

Bacteria↗

Sequence analysis of the gtfB gene from Streptococcus mutans.

The nucleotide sequence of the gtfB gene from Streptococcus mutans GS-5, coding for glucosyltransferase I activity, was determined. The gene codes for a strongly hydrophilic protein with a molecular size of 165,800 daltons. The deduced amino acid sequence revealed a typical gram-positive bacterial signal sequence at the NH2 terminus of the protein and 3.5 direct repeating units (each containing 65 amino acids) at the COOH terminus. Nucleotide sequencing of the region immediately downstream from the gtfB gene revealed the presence of a putative gene coding for an extracellular protein. This open reading frame is partially homologous to the gtfB gene.

Amino Acid Sequence↗

Cloning and expression of a gene encoding Sm16, an anti-inflammatory protein from Schistosoma mansoni.

The gene encoding Sm16, an anti-inflammatory, immunomodulatory protein present abundantly in secretions of the infective stages of Schistosoma mansoni was cloned and partially characterized. A data base analysis showed sequence homology to an earlier reported schistosomular stathmin-like gene sequences reported in dbEST and Genbank. The putative gene coding for Sm16 is of 500 bp with an open reading frame of 117 aa that included an N-terminal signal peptide sequence of 18 aa. There are three potential sites for phosphorylation (two serine and one tyrosine residue) but no glycosylation sites in the sequence. The coding region of Sm16 was amplified from cercarial cDNA, cloned and expressed in bacterial and insect expression systems. The purified recombinant protein showed strong immunoreactivity with a polyclonal rabbit anti-Sm16 antibody raised against the native anti-inflammatory protein Sm16. Contrary to earlier report, this gene appears to be not stage-specific. Metabolic labeling studies suggested that Sm16 is phosphorylated and is synthesized by both cercariae and schistosomula of S. mansoni. Sequence homology with human stathmin, a cell cycle regulatory phospho protein, was 30%. However, when probed with specific antibodies, no cross reactivity was observed between Sm16 and human stathmin.

Amino Acid Sequence↗

Avoidance of DNA methylation. A virus-encoded methylase inhibitor and evidence for counterselection of methylase recognition sites in viral genomes.

The ocr+ gene of bacterial virus T7 codes for the first protein recognized to inhibit a specific group of DNA methylases. The recognition sequences of several other DNA methylases, not susceptible to Ocr inhibition, are significantly suppressed in the virus genome. The bacterial virus T3 encodes an Ado-Met hydrolase, destroying the methyl donor and causing T3 DNA to be totally unmethylated. These observations could stimulate analogous investigations into the regulation of DNA methylation patterns of eukaryotic viruses and cells. For instance, an underrepresentation of methylation sites (5'-CG) is also true for animal DNA viruses. Moreover, we were able to disclose some novel properties of DNA restriction-modification enzymes concerning the protection of DNA recognition sequences in which only one strand can be methylated (e.g., type III enzyme EcoP15) and the primary resistance of (unmethylated) DNA recognition sites towards type II restriction endonuclease EcoRII.

Base Sequence↗

Heterochromatic sequences in a Drosophila whole-genome shotgun assembly.

BACKGROUND: Most eukaryotic genomes include a substantial repeat-rich fraction termed heterochromatin, which is concentrated in centric and telomeric regions. The repetitive nature of heterochromatic sequence makes it difficult to assemble and analyze. To better understand the heterochromatic component of the Drosophila melanogaster genome, we characterized and annotated portions of a whole-genome shotgun sequence assembly. RESULTS: WGS3, an improved whole-genome shotgun assembly, includes 20.7 Mb of draft-quality sequence not represented in the Release 3 sequence spanning the euchromatin. We annotated this sequence using the methods employed in the re-annotation of the Release 3 euchromatic sequence. This analysis predicted 297 protein-coding genes and six non-protein-coding genes, including known heterochromatic genes, and regions of similarity to known transposable elements. Bacterial artificial chromosome (BAC)-based fluorescence in situ hybridization analysis was used to correlate the genomic sequence with the cytogenetic map in order to refine the genomic definition of the centric heterochromatin; on the basis of our cytological definition, the annotated Release 3 euchromatic sequence extends into the centric heterochromatin on each chromosome arm. CONCLUSIONS: Whole-genome shotgun assembly produced a reliable draft-quality sequence of a significant part of the Drosophila heterochromatin. Annotation of this sequence defined the intron-exon structures of 30 known protein-coding genes and 267 protein-coding gene models. The cytogenetic mapping suggests that an additional 150 predicted genes are located in heterochromatin at the base of the Release 3 euchromatic sequence. Our analysis suggests strategies for improving the sequence and annotation of the heterochromatic portions of the Drosophila and other complex genomes.

Algorithms↗

Four inteins and three group II introns encoded in a bacterial ribonucleotide reductase gene.

A bacterial ribonucleotide reductase gene was found to encode four inteins and three group II introns in the oceanic N2-fixing cyanobacterium Trichodesmium erythraeum. The 13,650-bp ribonucleotide reductase gene is divided into eight extein- or exon-coding sequences that together encode a 768-amino acid mature ribonucleotide reductase protein, with 83% of the gene sequence encoding introns and inteins. The four inteins are encoded on the second half of the gene, and each has conserved sequence motifs for a protein-splicing domain and an endonuclease domain. These four inteins, together with known inteins, define five intein insertion sites in ribonucleotide reductase homologues. Two of the insertion sites are 10 amino acids apart and next to key catalytic residues of the enzyme. Protein-splicing activities of all four inteins were demonstrated in Escherichia coli. The four inteins coexist with three group II introns encoded on the first half of the same gene, which suggests a breakdown of the presumed barrier against intron insertion in this bacterial conserved protein-coding gene.

Amino Acid Motifs↗

Mannitol-specific enzyme II of the bacterial phosphotransferase system. III. The nucleotide sequence of the permease gene.

The nucleotide sequence of the mtlA gene, which codes for the mannitol-specific Enzyme II of the Escherichia coli phosphotransferase system, is presented. From the gene sequence, the primary translation product is predicted to consist of 637 amino acids (Mr = 67,893). This result is compared to the amino acid composition and molecular weight of the purified mannitol Enzyme II protein. The hydrophobic and hydrophilic properties of the enzyme were evaluated along its amino acid sequence using a computer program (Kyte, J., and Doolittle, R. F. (1982) J. Mol. Biol. 157, 105-132). The computer analysis predicts that the NH2-terminal half of the enzyme resides within the membrane, whereas the COOH-terminal half of the enzyme has the properties of a soluble protein. The possible functions of such a protein structure are discussed. RNA mapping has identified the promoter and mRNA start point for the mtl operon.

Amino Acids↗

Structure and variation of three canine genes involved in serotonin binding and transport: the serotonin receptor 1A gene (htr1A), serotonin receptor 2A gene (htr2A), and serotonin transporter gene (slc6A4).

Aggressive behavior is the most frequently encountered behavioral problem in dogs. Abnormalities in brain serotonin metabolism have been described in aggressive dogs. We studied canine serotonergic genes to investigate genetic factors underlying canine aggression. Here, we describe the characterization of three genes of the canine serotonergic system: the serotonin receptor 1A and 2A gene (htr1A and htr2A) and the serotonin transporter gene (slc6A4). We isolated canine bacterial artificial chromosome clones containing these genes and designed oligonucleotides for genomic sequencing of coding regions and intron-exon boundaries. Golden retrievers were analyzed for DNA sequence variations. We found two nonsynonymous single nucleotide polymorphisms (SNPs) in the coding sequence of htr1A; one SNP close to a splice site in htr2A; and two SNPs in slc6A4, one in the coding sequence and one close to a splice site. In addition, we identified a polymorphic microsatellite marker for each gene. Htr1A is a strong candidate for involvement in the domestication of the dog. We genotyped the htr1A SNPs in 41 dogs of seven breeds with diverse behavioral characteristics. At least three SNP haplotypes were found. Our results do not support involvement of the gene in domestication.

Amino Acid Sequence↗

Listeriolysin O is essential for virulence of Listeria monocytogenes: direct evidence obtained by gene complementation.

The role of listeriolysin O in the intracellular multiplication of Listeria monocytogenes and, therefore, its pathogenicity was questioned through a genetic complementation study. A nonhemolytic mutant was generated by inserting a single copy of transposon Tn917 in the bacterial chromosome. This insertion was localized by DNA sequence analysis in hlyA, the gene coding for listeriolysin O. As was another mutant that we previously characterized, this mutant was avirulent in the mouse. It was transformed with a plasmid carrying only hlyA, able to replicate in L. monocytogenes, and stably maintained in vitro and in vivo. The complemented strain displayed a hemolytic phenotype identical to that of the wild-type strain and was fully virulent, therefore attributing a crucial role to listeriolysin O in virulence and excluding the hypothesis of a polar effect of the transposon insertion on genes adjacent to hlyA and possibly involved in virulence.

Animals↗

Physical identification of an internal promoter, ilvAp, in the distal portion of the ilvGMEDA operon.

It has been previously demonstrated that the ilvGMEDA operon is expressed in vivo from the promoters ilvGp2 and ilvEp. An additional internal promoter is identified and designated ilvAp. This internal promoter, which allows independent expression of ilvA, has been analyzed both in vivo and in vitro. Our results indicate that: (1) ilvAp exists in both Escherichia coli K-12 and Salmonella typhimurium, as demonstrated by fusion to the galK reporter gene; (2) ilvAp is located in the distal coding sequence of ilvD; (3) the ilvAp sequences are not identical for these two bacterial species; (4) transcription from ilvAp of E. coli K-12 was demonstrated; (5) expression from ilvAp responds to the availability of oxygen; (6) potential 3' 5'-cyclic AMP receptor protein binding sites exist adjacent to ilvAp.

Base Sequence↗