PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Sequence and function analysis of a 9.74 kb fragment of Saccharomyces cerevisiae chromosome X including the BCK1 gene.

In the framework of the European BIOTECH project for sequencing the Saccharomyces cerevisiae genome, we have determined the nucleotide sequence of the cosmid clone 233 provided by F. Galibert (Rennes Cedex, France). We present here 9743 base pairs of sequence derived from the left arm of chromosome X. This sequence reveals three new open reading frames and includes the published sequence (5' end and open reading frame) of the gene BCK1/SLK1/SSP31 also identified as ORFAA. Deletion mutants of two earlier unknown open reading frames J0840 and J0904 are viable and the open reading frame J0902 is essential for yeast growth.

Amino Acid Sequence↗

Genes of the R-phycocyanin II locus of marine Synechococcus spp., and comparison of protein-chromophore interactions in phycocyanins differing in bilin composition.

R-phycocyanin II (RPCII) is a recently discovered member of the phycocyanin family of photosynthetic light-harvesting proteins. Genes encoding the alpha and beta subunits of RPCII were cloned and sequenced from marine Synechococcus sp. strains WH8020 and WH8103. The deduced amino acid sequences of RPCII were compared to two other types of phycocyanin, C-phycocyanin (CPC) and phycoerythrocyanin (PEC). These three types vary in the composition of their covalently bound bilin prosthetic groups. In terms of amino acid sequence identity RPCII is highly homologous to CPC and PEC, suggesting that the known three-dimensional structures of the latter two are representative of RPCII. Thus the amino acid residues contacting the three bilins of RPCII could be inferred and compared to those in CPC and PEC. Certain residues were identified among the three phycocyanins as possibly correlating with specific bilin isomers. In overall sequence RPCII and CPC are more homologous to one another than either is to PEC. This probably reflects functional homology in the roles of RPCII and CPC in the transfer of light energy to the core of the phycobilisome, a function not attributed to PEC. The genomes of Synechococcus sp. strains WH8020, WH8103 and WH7803 share homologous open reading frames in the vicinity of RPCII genes. The nucleotide sequence extending 3' from RPCII genes in strain WH8020 revealed two open reading frames homologous to components of an alpha CPC phycocyanobilin lyase. These open reading frames may encode a lyase specific for the attachment of phycoerythrobilin to alpha RPCII.

Amino Acid Sequence↗

The hygromycin-resistance-encoding gene as a selection marker for vaccinia virus recombinants.

Hygromycin B (Hy), an inhibitor of RNA translation, was shown to block the replication of vaccinia virus (VV) in cultured cell lines. Insertion of the Escherichia coli Hy resistance-encoding gene (hph) into the VV genome under control of early or late synthetic VV promoters could overcome inhibition of viral replication. When hph was inserted into VV in tandem with the human papillomavirus type 16 (HPV16) L1 open reading frame, hph recombinant viruses could be selected which expressed HPV16 L1.

Base Sequence↗

Revised nucleotide sequence of an archaeal insertion element (ISH28) reveals a putative transposase gene.

The published sequence of the insertion element ISH28 contained many small ORFs that were difficult to interpret. We resequenced the entire element and found seven nucleotide differences. The corrected sequence of ISH28 is 938 bp long, and now reveals a single open reading frame of 828 bp. The putative protein is highly similar (49% aa identity) to the predicted transposase of ISH1.

Amino Acid Sequence↗

Genome-wide analysis of the Emigrant family of MITEs of Arabidopsis thaliana.

Miniature inverted-repeat transposable elements (MITEs) are structurally similar to defective class II elements, but their high copy number and the size and sequence conservation of most MITE families suggest that they can be amplified by a replicative mechanism. Here we present a genome-wide analysis of the Emigrant family of MITEs from Arabidopsis thaliana. In order to be able to detect divergent ancient copies, and low copy number subfamilies with a different internal sequence we have developed a computer program to look for Emigrant elements based solely on the terminal inverted-repeat sequence. We have detected 151 Emigrant elements of different subfamilies. Our results show that different bursts of amplification, probably of few active, or master, elements, have occurred at different times during Arabidopsis evolution. The analysis of the insertion sites of the Emigrant elements shows that recently inserted Emigrant elements tend to be located far from open reading frames, whereas more ancient Emigrant subfamilies are preferentially found associated to genes.

Amino Acid Sequence↗

Organization and expression of the polynucleotide phosphorylase gene (pnp) of Streptomyces: Processing of pnp transcripts in Streptomyces antibioticus.

We have examined the expression of pnp encoding the 3'-5'-exoribonuclease, polynucleotide phosphorylase, in Streptomyces antibioticus. We show that the rpsO-pnp operon is transcribed from at least two promoters, the first producing a readthrough transcript that includes both pnp and the gene for ribosomal protein S15 (rpsO) and a second, Ppnp, located in the rpsO-pnp intergenic region. Unlike the situation in Escherichia coli, where observation of the readthrough transcript requires mutants lacking RNase III, we detect readthrough transcripts in wild-type S. antibioticus mycelia. The Ppnp transcriptional start point was mapped by primer extension and confirmed by RNA ligase-mediated reverse transcription-PCR, a technique which discriminates between 5' ends created by transcription initiation and those produced by posttranscriptional processing. Promoter probe analysis demonstrated the presence of a functional promoter in the intergenic region. The Ppnp sequence is similar to a group of promoters recognized by the extracytoplasmic function sigma factors, sigma-R and sigma-E. We note a number of other differences in rspO-pnp structure and function between S. antibioticus and E. coli. In E. coli, pnp autoregulation and cold shock adaptation are dependent upon RNase III cleavage of an rpsO-pnp intergenic hairpin. Computer modeling of the secondary structure of the S. antibioticus readthrough transcript predicts a stem-loop structure analogous to that in E. coli. However, our analysis suggests that while the readthrough transcript observed in S. antibioticus may be processed by an RNase III-like activity, transcripts originating from Ppnp are not. Furthermore, the S. antibioticus rpsO-pnp intergenic region contains two open reading frames. The larger of these, orfA, may be a pseudogene. The smaller open reading frame, orfX, also observed in Streptomyces coelicolor and Streptomyces avermitilis, may be translationally coupled to pnp and the gene downstream from pnp, a putative protease.

Amino Acid Sequence↗

Molecular cloning and characterization of bacteriophage P2 genes R and S involved in tail completion.

The sequences of two previously known tail genes, R and S, of the temperate bacteriophage P2 and the sequence of an additional open reading frame (orf-30) located between S and V, were determined. Amber mutations mapping within R and S, Ram3, Ram42, Ram23, Sam75, and Sam89 were sequenced and found to be within their corresponding open reading frames. We constructed overproducing plasmids for R and S and identified these proteins by SDS-PAGE of whole-cell lysates and Coomassie blue staining. The predicted molecular masses of proteins R and S were M(r) 17,400 and 17,300, respectively, although both polypeptides migrated more slowly during gel electrophoresis than would be expected from the sequence data. orf-30 occupies the strand opposite from RS and V and is preceded by several weak potential sigma 70-RNA polymerase promoters, some of which overlap with the V promoter. A construct that had the putative orf-30 promoter region upstream of the lacZ gene produced low levels of beta-galactosidase activity in vivo. Expression from the orf-30 promoter was not stimulated by the phage P4 transcriptional activator protein, delta, which acts at all the known P2 and P4 late promoters. Insertion mutagenesis showed that orf-30 was not an essential gene for P2 growth in Escherichia coli. None of the gene or protein sequences exhibited extensive homology to sequences in the nucleic acid and protein databases. However, the R protein contains a small region homologous to one in the phage T4 tail protein gp15, which is required for T4 tails to bind heads. We propose that R and S are tail completion proteins that are essential for stable head joining.

Amino Acid Sequence↗

Analysis of genomic rearrangement and subsequent gene deletion of the attenuated Orf virus strain D1701.

The orf virus (OV) strain D1701 belongs to the genetically heterogenous parapoxvirus (PPV) genus of the family Poxviridae. The attenuated OV D1701 has been licensed as a live vaccine against contagious ecthyma in sheep. Detailed knowledge on the genetic structure and organization of this PPV vaccine strain is an important prerequisite to reveal possible genetic mechanisms of PPV attenuation. The present study demonstrates a genomic map of the approximately 158 kbp DNA of OV D1701 established by hybridization studies of cloned restriction fragments covering the complete viral genome. The results show an enlargement of the inverted terminal repeats (ITR) to up to 18 kbp due to recombination between nonhomologous sequences during cell culture adaptation. DNA sequencing of the region adjacent to the ITR junction revealed the absence of one open reading frame designated E2L. In contrast to a transposition-deletion variant of the New Zealand OV strain NZ2 (Fleming et al., 1995) the two genes E3L (a homologue of dUTPase) and G1L neighbouring E2L are retained in OV D1701. DNA and RNA analyses proved the presence of E2L gene in wild-type OV isolated directly from scab material. The data presented indicate that the E2L gene is nonessential for virus replication in vitro and in vivo, and may represent one important viral gene in determining virulence and pathogenesis of OV.

Amino Acid Sequence↗

Structure and expression of the TREX1 and TREX2 3' --> 5' exonuclease genes.

The TREX1 and TREX2 genes encode mammalian 3'-->5' exonucleases. Expression of the TREX genes in human cells was investigated using a reverse transcription-polymerase chain reaction strategy. Our results show that TREX1 and TREX2 are expressed in all tissues tested, providing direct evidence for the expression of these genes in human cells. Potential transcription start sites are identified for the TREX genes using rapid amplification of cDNA ends to recover the 5'-flanking regions of the TREX transcripts. The 5'-flanking sequences indicate transcription initiation from consensus putative promoters identified -140 and -650 base pairs upstream of the TREX1 open reading frame (ORF) and -623 and -753 base pairs upstream of the TREX2 ORF. Novel TREX1 and TREX2 cDNAs are identified that contain protein-coding sequences generated from exons positioned in genomic DNA up to 18 kilobases 5' to the TREX1 ORF and up to 25 kilobases 5' to the TREX2 ORF. These novel cDNAs and sequences in the GenBank data base indicate that transcripts containing the TREX1 and TREX2 ORFs are produced using a variety of mechanisms that include alternate promoter usage, alternative splicing, and varied sites for 3' cleavage and polyadenylation. These initial studies have revealed previously unrecognized complexities in the structure and expression of the TREX1 and TREX2 genes.

Amino Acid Sequence↗

High variability of human cytomegalovirus UL150 open reading frame in low-passaged clinical isolates.

OBJECTIVE: To investigate the polymorphism of human cytomegalovirus (HCMV) UL150 open reading frame (ORF) in low-passaged clinical isolates, and to study the relationship between the polymorphism and different pathogenesis of congenital HCMV infection. METHODS: PCR was performed to amplify the entire HCMV UL150 ORF region of 29 clinical isolates, which had been proven containing detectable HCMV-DNA using fluorescence quantitative PCR. PCR amplification products were sequenced directly, and the data were analyzed. RESULTS: Totally 25 among 29 isolates were amplified, and 18 isolates were sequenced successfully. HCMV UL150 ORF sequences derived from congenitally infected infants were high variability. The UL150 ORF in all 18 clinical isolates shifted backward by 8 nucleotides leading to frame-shift, and contained a single nucleotide deletion at nucleotide position 226 compared with that of Toledo strain. The nucleotide diversity was 0.1% to 6.8% and the amino acid diversity was 0.2% to 19.2% related to Toledo strain. However, the nucleotide diversity was 0.1% to 6.4% and amino acid diversity was 0.2% to 8.3% by compared with Merlin strain. Compared with Toledo, 4 new cysteine residues and 13 additional posttranslational modification sites were observed in UL150 putative proteins of clinical isolates. Moreover, the UL150 putative protein contained an additional transmembrane helix at position of 4-17 amino acid related to Toledo. CONCLUSION: HCMV UL150 ORF and deduced amino acid sequences of clinical strains are hypervariability. No obvious linkage between the polymorphism and different pathogenesis of congenital HCMV infection is found.

Amino Acid Sequence↗

Location and characterization of the bovine herpesvirus type 4 thymidine kinase gene; comparison with thymidine kinase genes of other herpesviruses.

The location and nucleotide sequence of the bovine herpesvirus type 4 (BHV-4) thymidine kinase (TK) gene was determined. The coding region of the TK gene is 1335 nucleotides long and corresponds to a polypeptide of 445 amino acids. Comparison of TK amino acid sequences of BHV-4 and 16 herpesvirus TKs reveals a greater homology to those of the gammaherpesviruses EBV and specially HVS, than to those of alphaherpesviruses. The open reading frames detected in the vicinity of TK gene were homologous to the corresponding ones in other herpesviruses.

Amino Acid Sequence↗

In vivo biotinylated proteins as targets for phage-display selection experiments.

Screening phage-displayed combinatorial libraries represents an attractive method for identifying affinity reagents to target proteins. Two critical components of a successful selection experiment are having a pure target protein and its immobilization in a native conformation. To achieve both of these requirements in a single step, we have devised cytoplasmic expression vectors for expression of proteins that are tagged at the amino- or carboxy-terminus (pMCSG16 and 15) via the AviTag, which is biotinylated in vivo with concurrent expression of the BirA biotin ligase. To facilitate implementation in high-throughput applications, the engineered vectors, pMCSG15 and pMCSG16, also contain a ligase-independent cloning site (LIC), which permits up to 100% cloning efficiency. The expressed protein can be purified from bacterial cell lysates with immobilized metal affinity chromatography or streptavidin-coated magnetic beads, and the beads used directly to select phage from combinatorial libraries. From selections using the N-terminally biotinylated version of one target protein, a peptide ligand (Kd= 9 microM) was recovered that bound in a format-dependent manner. To demonstrate the utility of pMCSG16, a set of 192 open reading frames were cloned, and protein was expressed and immobilized for use in high-throughput selections of phage-display libraries.

Bacterial Proteins↗

Molecular and functional characterization of the murine glucocerebrosidase gene.

A genomic clone of glucocerebrosidase (D-glucosyl-N-acyl-sphingosine glucohydrolase; E.C. 3.2.1.45) purified from a genomic library derived from a Balb/c mouse was analyzed by restriction mapping and nucleotide sequencing of its promoter and protein coding regions. Promoter activity was functionally assessed by ligation of a 2 kb glucocerebrosidase fragment to the protein coding segment of a bacterial neomycin resistance gene. Smaller segments of the 5' flanking sequence were then analyzed for their ability to initiate transcription of the chloramphenicol acetyltransferase reporter gene. A 319 bp Eco RI-Bgl II fragment (containing 259 bp upstream of the cDNA 5' limit) ligated to the chloramphenicol acetyltransferase open reading frame produced considerable activity.

3T3 Cells↗

Adaptive evolution in LINE-1 retrotransposons.

We traced the sequence evolution of the active lineage of LINE-1 (L1) retrotransposons over the last approximately 25 Myr of human evolution. Five major families (L1PA5, L1PA4, L1PA3B, L1PA2, and L1PA1) of elements have succeeded each other as a single lineage. We found that part of the first open-reading frame (ORFI) had a higher rate of nonsynonymous (amino acid replacement) substitution than synonymous substitution during the evolution of the ancestral L1PA5 through the L1PA3B families. This segment encodes the coiled coil region of the protein-protein interaction domain of the ORFI protein (ORFIp). Statistical analysis of these changes indicates that positive selection had been acting on this region. In contrast, the coiled coil segment hardly changed during the evolution of the L1PA3B to the present L1PA1 family. Therefore, selective pressure on the coiled coil segment has changed over time. We suggest that the fast rate of amino acid replacement in the coiled coil segment reflects the adaptation of L1 either to a changing genomic environment or to host repression factors. In contrast, the second open-reading frame and the nucleic acid-binding domain of the first open-reading frame are extremely well conserved, attesting to the strong purifying selection acting on these regions.

Amino Acid Sequence↗

Isolation and characterization of a carboxysome shell gene from Thiobacillus neapolitanus.

The gene coding for the major carboxysome shell peptide (csoS1) from Thiobacillus neapolitanus has been isolated and sequenced. Oligonucleotide primers for polymerase chain reaction (PCR) amplification of the 5' end of the gene were made possible by amino acid sequencing of the N-terminal residues of the shell peptide. A 41 bp PCR product was used as a probe to isolate the gene. The deduced amino acid composition of the 216 bp gene shows a high degree of hydrophobicity. The gene is located within a series of three repeated regions of DNA and appears to have arisen via gene duplication. The transcript of csoS1 is approximately 400 bases in length. The shell peptide shares significant homology with Synechococcus open reading frames implicated in carboxysome structure/assembly. These open reading frames and csoS1 are related and are probably members of a carboxysome gene family.

Amino Acid Sequence↗

Novel resistance-nodulation-cell division efflux system AdeDE in Acinetobacter genomic DNA group 3.

Resistance-nodulation-cell division type efflux pump AdeDE was identified in acinetobacters belonging to genomic DNA group 3. Inactivation of adeE showed that it may be responsible for reduced susceptibility to amikacin, ceftazidime, chloramphenicol, ciprofloxacin, erythromycin, ethidium bromide, meropenem, rifampin, and tetracycline. However, unlike what was found for other similar efflux systems, the open reading frame for the outer membrane component was not found downstream of the adeDE gene cluster.

Acinetobacter↗

Characterization of the mobilization region of a Bacteroides insertion element (NBU1) that is excised and transferred by Bacteroides conjugative transposons.

Many Bacteroides clinical isolates carry large conjugative transposons that, in addition to transferring themselves, excise, circularize, and transfer smaller, unlinked chromosomal DNA segments called NBUs (nonreplicating Bacteroides units). We report the localization and DNA sequence of a region of one of the NBUs, NBU1, that was necessary and sufficient for mobilization by Bacteroides conjugative transposons and by IncP plasmids. The fact that the mobilization region was internal to NBU1 indicates that the circular form of NBU1 is the form that is mobilized. The NBU1 mobilization region contained a single large (1.4-kbp) open reading frame (ORF1), which was designated mob. The oriT was located within a 220-bp region upstream of mob. The deduced amino acid sequence of the mob product had no significant similarity to those of mobilization proteins of well-characterized Escherichia coli group plasmids such as RK2 or of either of the two mobilization proteins of Bacteroides plasmid pBFTM10. There was, however, a high level of similarity between the deduced amino acid sequence of the mob product and that of the product of a Bacteroides vulgatus cryptic open reading frame closely linked to a cefoxitin resistance gene (cfxA).

Amino Acid Sequence↗

Construction of a vector plasmid for use in Gluconobacter oxydans.

A host vector system in Gluconobacter oxydans was constructed. An Acetobacter-Escherichia coli shuttle vector was introduced with the efficiency of 10(4) transformants/microg of DNA. Next, aiming for a self-cloning vector, we found a cryptic plasmid (which we named pAG5) of 5648 bp in G. oxydans strain IFO 3171, and sequenced the nucleotides. The plasmid seemed to have only one open reading flame (ORF) for a possible replication protein. Shuttle vectors of Gluconobacter-E. coli were constructed with the plasmid pAG5 and an E. coli vector, pUC18.

Escherichia coli↗