PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Small genes/gene-products in Escherichia coli K-12.

Forty-two protein spots of observed M(r) 6-15 kDa were resolved by two-dimensional gel electrophoresis, stained by Coomassie blue and subjected to Edman microsequencing. All of the proteins could be related back to their encoding open reading frames, thereby vindicating the bioinformatic tools currently utilised in their identification. However, only 14/42 gene-products were expressed as annotated. Translation was confirmed for 14 open reading frames with no attributed function (EcoGene Y-entries), while N-terminal sequence allowed the start codon to be accurately annotated for the genes yigF, yccU, yqiC, ynfD, and yeeX. The methionine start codon was cleaved in 11 gene-products (AtpE, Hns, RpoZ, RplL, CspC, YccJ, YggX, YjgF, HimA, InfA, RpsQ) and a further five showed loss of a signal peptide (PspE, HdeB, HdeA, YnfD, YkfE). Internal (Tig, AtpA, TufA) and N-terminal fragmentation (CspD, RpsF, AtcU) of much larger proteins was also detected, which may have resulted from physiological or translational processes. M(r) and pI isoforms were detected respectively for PtsH and GatB, each being phosphoproteins, as well as RplY which manifested differences with respect to predicted M(r) and pI. In addition, YjgF was shown to belong to a small gene family of unknown function with ancient conserved regions across procaryotes and eucaryotes. YgiN was revealed to have a paralogue and orthologues in Bacillus subtilis, Synechocystis sp., Mycobacterium tuberculosis, Neisseria gonorrhoea, and Rhodococcus erythropolis. Orthologues are also reported for YihD, YccU and YeeX. Of the 14 Y-genes, only YkfE possessed no detectable orthologues. These results highlight the need to complement genomic analysis with detailed proteomics in order to gain a better understanding of cellular molecular biology, while the confirmation of the open reading frame start codon using Edman degradation protein microsequencing has yet to be superseded by recent advances in mass spectrometry.

Amino Acid Sequence↗

Identification, DNA sequence, and distribution of IS981, a new, high-copy-number insertion sequence in lactococci.

An insertion in the lactococcal plasmid pGBK17, which inactivated the gene(s) encoding resistance to the prolate-headed phage c2, was cloned, sequenced, and identified as a new lactococcal insertion sequence (IS). IS981 was 1,222 bp in size and contained two open reading frames, one large enough to encode a transposase. IS981 ended in imperfect inverted repeats of 26 of 40 bp and generated a 5-bp direct repeat of target DNA at the site of insertion. IS981 was present on the chromosome of Lactococcus lactis subsp. lactis LM0230 from where it transposed to pGBK17 during transformation. Twenty-three strains of lactococci examined for the presence of IS981 by Southern hybridization showed 4 to 26 copies per genome, with L. lactis subsp. cremoris strains containing the highest number of copies. Comparison of the DNA sequence and the amino acid sequence of the long open reading frame to other known sequences showed that IS981 is related to a family of IS elements that includes IS2, IS3, IS51, IS150, IS600, IS629, IS861, IS904, and ISL1.

Amino Acid Sequence↗

Nucleotide sequence and organization of the upstream region of the Corynebacterium glutamicum lysA gene.

Maximum expression of the Corynebacterium glutamicum lysA gene is dependent upon the presence of a 2.3 kb region immediately 5' of the lysA reading frame. Subcloning and functional analysis of the upstream region implied that this region contained the lysA promoter. Sequence determination of the upstream region revealed a single open reading frame, orfX, in the same orientation as lysA. The orfX coding sequence exhibited all the sequence characteristics of a gene with the potential for a 550-amino-acid polypeptide product. Expression of lysA is coupled to that of orfX via a common promoter located immediately 5' of orfX. The RNA start site has been determined by S1 nuclease mapping. Both the orfX and the lysA gene are expressed as a single 3.0 kb RNA transcript. These data indicate that orfX and lysA are genes within a two-gene operon. Expression of the lysA gene is not subject to regulation by lysine. The orfX gene product was shown not to be directly linked to the lysine biosynthetic pathway, nor is it the enzyme incorporating DAP into the peptidoglycan precursor.

Amino Acid Sequence↗

Distribution of RNA editing sites in Oenothera mitochondrial mRNAs and rRNAs.

To investigate whether RNA editing in plant mitochondria modifies structural RNAs as well as protein-coding RNAs we compared the genomic-encoded information with the respective transcripts of several genes in Oenothera. The genes analysed are the 5S, 18S and 26 S rRNAs, the alpha-subunit of ATPase (atpA), cytochrome b (cytb), orfB, which is located upstream of cytochrome oxidase subunit III, and the respective leader, trailer and spacer sequences. All open reading frames were found to be edited to some degree. The atpA coding region has the least edited mRNA in Oenothera mitochondria, with only four nucleotides altered in the 1533 nucleotide open reading frame. From this analysis we conclude that frequent RNA editing is indicative of functional protein coding regions in plant mitochondria. The extensive editing in orfB, for example, suggests that this orf codes for a mitochondrial protein. No RNA editing event was found in the 5S rRNA or in the 1824 nucleotides analysed of the 18S rRNA, but two nucleotides were found to be altered in the 1970 nucleotides compared for the 26S rRNA. One nucleotide alteration has changed C to U, the other in reverse U to C. However, only one of five cDNA clones covering this region shows the modifications, similar to many silent editing events in open reading frames. RNA editing in the structural RNAs thus does not seem to be essential for their function in the mitochondrial ribosome.

Adenosine Triphosphatases↗

Nucleotide sequence of cDNA encoding the coat protein of beet yellows virus.

A cDNA clone of beet yellows viral RNA expressed the viral coat protein gene in E. coli. The sequence of the 2724 nucleotide insert revealed three open reading frames, the 3' of which was shown to be the coat protein cistron. This cistron is expressed in E. coli, in spite of there being no obvious ribosome binding site upstream.

Amino Acid Sequence↗

Nucleotide sequence of a small plasmid isolated from Acetobacter pasteurianus.

A 1440-bp plasmid named pAP12875 was isolated from Acetobacter pasteurianus and its nucleotide sequence determined. An open reading frame was found capable of coding for a protein that has similarity with the replication protein of pVT736-1 from Actinobacillus actinomycetemcomitans and the 32-kDa protein of phage Pf3 from Pseudomonas aeruginosa.

Acetobacter↗

Cloning, sequencing and analysis of the structural gene and regulatory region of the Pseudomonas aeruginosa chromosomal ampC beta-lactamase.

The chromosomal gene from Pseudomonas aeruginosa encoding beta-lactamase has been cloned, and the sequence determined and compared with corresponding sequences of beta-lactamases from members of the enterobacteriaceae. Upstream of the beta-lactamase gene is an open reading frame which we postulate encodes a regulatory protein, AmpR. We identified a helix-turn-helix region in AmpR and a putative AmpR-binding site.

Amino Acid Sequence↗

Characterization of the tol-pal and cyd region of Escherichia coli K-12: transcript analysis and identification of two new proteins encoded by the cyd operon.

Sequence analysis showed that the cyd operon is immediately upstream of the tol-pal region. Northern (RNA) blot analysis demonstrated that the transcript for the cyd operon terminates just before the promoter for transcription of the tol genes. The cyd transcript contains cydA cydB followed by two open reading frames: orfC, encoding a 37-residue peptide, and orfD, encoding a 97-residue peptide. Both OrfC and OrfD are synthesized in minicells.

Amino Acid Sequence↗

Nucleotide sequence of the gene for a thermostable esterase from Pseudomonas putida MR-2068.

The esterase gene (est) of Pseudomonas putida MR-2068 was cloned into Escherichia coli JM109. An 8-kb inserted DNA directed synthesis of an esterase in E. coli. The esterase gene was in a 1.1-kb PstI-ClaI fragment within the insert DNA. The complete nucleotides of the DNA fragment containing the esterase gene were sequenced and found to include a single open reading frame of 828 bp coding for a protein of 276 amino acid residues. The open reading frame was confirmed by N-terminal amino acid sequence analysis of the purified esterase. A potential Shine-Dalgarno sequence is followed by the open reading frame. The esterase activity of the recombinant E. coli was more than 200 times higher than that of parental strain, P. putida MR-2068.

Amino Acid Sequence↗

Sequence analysis of the genes encoding the phosphoprotein of recent isolates of canine distemper virus in Japan.

The nucleotide sequences of the phosphoprotein (P) of canine distemper virus (CDV) strains isolated between 1992 and 1996 in Japan were determined. This is the first report of the complete sequences of the P genes of recently prevalent CDV strains. The deduced amino acid sequences of the P, C and V proteins showed that in the new Japanese isolates, these proteins have approximately 93%, 90-91% and 92% identities with those of the Onderstepoort vaccine strain, respectively. The predicted functional regions were conserved. RNA editing resulting in a shift to the open reading frame (ORF) of the V protein was shown to occur with the same efficiency in both the field isolates and vaccine strain.

Amino Acid Sequence↗

Primary structure of the novel bacterial rhodopsin from extremely halophilic archaeon Haloarcula japonica strain TR-1.

A novel bacterial rhodopsin was identified in Haloarcula japonica strain TR-1. The gene encoding the bacterial rhodopsin was cloned and sequenced. The structural gene consisted of an open reading frame of 750 nucleotides encoding 250 amino acids. The deduced amino acid sequence of the Ha. japonica bacterial rhodopsin showed the highest homology to those of cruxrhodopsins.

Amino Acid Sequence↗

Identification of a novel 23kDa protein encoded by putative open reading frame 2 of TT virus (TTV) genotype 1 different from the other genotypes.

We report the entire open reading frames (ORFs) sequences of four TT virus (TTV) isolates, one genotype 2 (G2) and three G4 isolates. Despite a DNA virus, TTV possesses high rate of amino acid (aa) substitution: the aa sequence homology of ORF1 and 2 is lower than the nucleotide homology. The partial 'N22' region of ORF1 is suitable for genotyping of 'prototype TTV' isolates, because the phylogenetic tree from partial 'N22' sequence is consistent with that from the entire ORF1. Based on our sequence data, ORF2 from most isolates excluding G1 encode truncated 49 aa (pORF2a) because of an in-frame stop codon, although ORF2s from most G1 isolates encode 202 aa (pORF2ab). Just downstream the stop codon, another ORF encoding a protein of approximately 150 aa (pORF2b) is found, whose homology is quite low among these genotypes. Our in vitro transcription/translation study supports that all G1a and a part of G b without an in-frame stop codon dominantly encode pORF2ab, a novel 23 kDa protein, whereas the other genotypes with an in-frame stop codon encode pORF2b (17 kDa). Our data indicate TTV G1a and a part of G1b should have different characteristics from the other genotypes.

Amino Acid Sequence↗

Odocoileus hemionus deer adenovirus is related to the members of Atadenovirus genus.

The Odocoileus hemionus deer adenovirus (OdAdV-1) causes systemic and local vasculitis and proves extremely lethal for mule deer. To characterize the virus, part of the genome flanking the fiber gene was cloned and sequenced. The sequence revealed two open-reading frames that mapped to pVIII hexon-associated protein precursor and fiber protein of several other adenoviruses. The highest amino acid homology for pVIII and fiber was found with the members of the proposed Atadenovirus genus: ovine adenovirus isolate 287 (OAdV-287), bovine adenovirus 4 (BAdV-4) and duck adenovirus 1 (DAdV-1). The homology with bovine adenovirus type 3 (BAdV-3) proved low. The E3 region was not found between the gene for pVIII and fiber. These data suggest that OdAdV-1 is a member of the Atadenovirus genus.

Adenoviridae↗

Cloning and nucleotide sequence analysis of the Streptococcus mutans membrane-bound, proton-translocating ATPase operon.

The function of the membrane-bound ATPase in S. mutans is to regulate cytoplasmic pH values for the purpose of maintaining delta pH. Previous studies have shown that as part of its acid-adaptive ability, S. mutans is able to increase H(+)-ATPase levels in response to acidification. As part of the study of ATPase regulation in S. mutans, we have cloned the ATPase operon and determined its genetic organization. The structural genes from S. mutans were found to be in the order: c, a, b, delta, alpha, gamma, beta, and epsilon; where c and a were reversed from the more typical bacterial organization. The operon contained no I gene homologue but was preceded by a 239-bp intergenic space. Deduced aa sequences from open reading frames indicated that genes encoding homologues of glycogen phosphorylase and nonphosphorylating, NADP-dependent glyceraldehyde-3-phosphate dehydrogenase flank the H(+)-ATPase operon, 5' and 3' respectively. Sequence analysis indicated the presence of three inverted-repeat nt sequences in the glgP-uncE intergenic space. Primer extension analysis of mRNAs prepared from batch-grown or steady-state cultures demonstrated that the transcriptional start site did not change as a function of culture pH value. The data suggest that potential stem-and-loop structures in the promoter region of the operon do not function to alter the starting position of ATPase-specific mRNA transcription.

Amino Acid Sequence↗

Cloning and functional characterization of the human fractalkine receptor promoter regions.

We have previously shown that reduced expression of the fractalkine receptor, CX3CR1, is correlated with rapid HIV disease progression and with reduced susceptibility to acute coronary events. In order to elucidate the mechanisms underlying transcriptional regulation of CX3CR1 expression, we structurally and functionally characterized the CX3CR1 gene. It consists of four exons and three introns spanning over 18 kb. Three transcripts are produced by splicing the three untranslated exons with exon 4, which contains the complete open reading frame. The transcript predominantly found in leucocytes corresponds to the splicing of exon 2 with exon 4. Transcripts corresponding to splicing of exons 1 and 4 are less abundant in leucocytes and splicing of exons 3 and 4 are rare longer transcripts. A constitutive promoter activity was found in the regions extending upstream from untranslated exons 1 and 2. Interestingly, exons 1 and 2 enhanced the activity of their respective promoters in a cell-specific manner. These data show that the CX3CR1 gene is controlled by three distinct promoter regions, which are regulated by their respective untranslated exons and that lead to the transcription of three mature messengers. This highly complex regulation may allow versatile and precise expression of CX3CR1 in various cell types.

Alternative Splicing↗

Defective Marek's disease virus DNA contains a gene encoding a potential nuclear DNA binding protein and a HSV a-like sequence.

Four RNA transcripts from chicken embryo fibroblast cells infected with Marek's disease virus (MDV) strain 281Ml/1 hybridized to the 4-kbp MDV replicon DNA. In an attempt to identify open reading frames coding for the four transcripts, we determined the nucleotide sequences of 4-kbp replicon DNA (represents a single monomeric repeat unit of defective MDV genome). Computer analysis indicates that the 4-kbp MDV replicon DNA contains two intact open reading frames (ORFs) with common promoter regulatory elements. ORF-A codes for a putative 204 amino acid protein that shares 21 and 36% amino acid sequence identity to nuclear DNA binding proteins such as the EBNA-1 of Epstein-Barr virus and galline, a chicken sperm histone protein, respectively. ORF-B encodes for a potential 350 amino acid protein, which did not show any significant amino acid sequence identity to known protein sequences within Swiss-Protein data base. ORF-B may, therefore, encode a MDV specific protein. The 5'-region of MDV replicon DNA revealed seven reiterated copies of an 11-bp motif sharing 8 out of 11 nucleotide sequence identity to DR2 elements of the herpes simplex virus strain USA-8 a sequence.

Amino Acid Sequence↗

A Rev protein is expressed in caprine arthritis encephalitis virus (CAEV)-infected cells and is required for efficient viral replication.

Caprine arthritis encephalitis virus (CAEV) is a lentivirus that is closely related to visna virus and more distantly related to the human lentivirus human immunodeficiency virus 1 (HIV-1). Like other lentiviruses, the genome of CAEV contains multiple small ORFs that encode viral regulatory proteins. Sequence analysis of the CAEV genome and cDNAs generated from mRNA in infected cells has suggested that one of these ORFs encodes a protein (Rev-C) that is analogous to Rev of visna virus and HIV. Antibodies generated to a carboxy-terminal peptide of the rev ORF immunoprecipitate an 18-kDa protein from cells transfected with the Rev cDNA clone. Immunoprecipitation and immunofluorescence analysis of CAEV-infected ovine primary cells show that the product of the rev ORF is expressed during infection and localizes to the nucleolus of infected cells. Also, sera from CAEV-infected goats specifically immunoprecipitates an in vitro-translated product from the full-length Rev cDNA clone as well as that from the unique second open reading frame of Rev-C which shows that the Rev-C protein is expressed during natural CAEV infection of animals. Insertion of either a mutation that creates two stop codons in the unique second open reading frame of Rev-C or a mutation in the basic domain of Rev-C into the CAEV infectious molecular clone renders the virus unable to replicate in primary goat synovial membrane cells. Analysis of the RNA and proteins produced from both Rev-deficient clones indicates that they are defective in the accumulation of structural gene mRNAs in the cytoplasm as well as in synthesis of structural proteins compared to the wild-type CAEV clone. These data indicate that CAEV encodes a Rev protein that is required for efficient viral replication in culture.

Amino Acid Sequence↗

The complete DNA sequence and genome organization of the avian adenovirus, hemorrhagic enteritis virus.

Hemorrhagic enteritis virus (HEV) belongs to the Adenoviridae family, a subgroup of adenoviruses (Ads) that infect avian species. In this article, the complete DNA sequence and the genome organization of the virus are described. The full-length of the genome was found to be 26,263 bp, shorter than the DNA of any other Ad described so far. The G + C content of the genome is 34.93%. There are short terminal repeats (39 bp), as described for other Ads. Genes were identified by comparison of the DNA and predicted amino acid sequences with published sequences of other Ads. The organization of the genome in respect to late genes (52K, IIIa, penton base, core protein, hexon, endopeptidase, 100K, pVIII, and fiber), early region 2 genes (polymerase, terminal protein, and DNA binding protein), and intermediate gene IVa2 was found to be similar to that of other human and avian Ad genomes. No sequences similar to E1 and E4 regions were found. Very low similarity to ovine E3 region was found. Open reading frames were identified with no similarity to any published Ad sequence.

Adenoviridae↗