PubMed HealthSearch

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Analysis of a region from the bacteriophage resistance plasmid pCI528 involved in its conjugative mobilization between Lactococcus strains.

A 10-kb HindIII fragment of pCI528 cloned into the nonconjugative shuttle vector pCI3340 could be transferred by conjugative mobilization from Lactococcus lactis subsp. lactis MG1363, whereas other HindIII fragments of pCI528 or the vector alone were nonmobilizable. Subcloning of this 10-kb region identified a 4.4-kb BglII-EcoRI fragment which contained all the DNA essential for transfer. Sequence analysis of a 2-kb region within this 4.4 kb-segment revealed a region rich in inverted repeats and two potential overlapping open reading frames, one of which demonstrated homology to mobilization proteins of two nonconjugative staphylococcal plasmids.

Amino Acid Sequence

Identification of a T4 gene required for bacteriophage mRNA processing.

A ribonucleolytic activity that cleaves within the Shine/Dalgarno sequences of the bacteriophage T4 motA and ORF2 mRNAs was recently described. We have identified additional sites of processing within several other ribosome binding sites, including two sites in the polycistronic frd transcript. Deletion mutants (farP) that overproduce the product of frd are defective in this mRNA processing. The mutants were used to identify processing events dependent on the T4 activity including attack at nuclease-sensitive sites within the coding sequences of some genes and within the intercistronic region 5' of gene 43. All known processing sites lie within similar sequences. Another mutant in mRNA processing carries a point mutation in one of the open reading frames (orf61.9) removed by the farP deletions. Introduction of a cloned copy of this open reading frame into a unique site in the chromosome of farP phage is sufficient to restore mRNA processing capability. The open reading frame probably encodes the T4 regB protein.

Base Sequence

Isolation of a yeast tropomyosin-related cDNA clone that encodes a novel transmembrane protein having a C-terminal highly basic region.

A novel cDNA clone was isolated from a yeast Saccharomyces cerevisiae lambda gt11 cDNA library using rabbit anti-rat tropomyosin (TM30nm) polyclonal antibody (RTM8-2). It consists of an open reading frame of 951 bp, encoding 317 amino acid residues. The putative sequence recognized by RTM8-2 was present in Asn-235 to Thr-250. The deduced amino acid sequence and hydropathy plot suggested that this protein has a tropomyosin-homologous sequence, a predicted transmembrane, and a C-terminal basic region. A search of the data bases (EMBL and GenBank) revealed that the 128bp sequence in the 3' untranslated region (3'UTR) is almost identical (96.9%) to the human cDNA clone 54E05 sequence (EMBL accession number Z15978).

Amino Acid Sequence

Distribution and characterization of plasmid-related sequences in the chromosomal DNA of different thermophilic Methanobacterium strains.

The genomes of several thermophilic members of the genus Methanobacterium were analyzed for homology to the related restriction-modification plasmids pFV1 and pFZ1 from M. thermoformicicum strains THF and Z-245, respectively. Two plasmid regions, designated FR-I and FR-II, could be identified with chromosomal counterparts in six Methanobacterium strains. Multiple copies of the pFV1-specific element FR-I were detected in the M. thermoformicicum strains CSM3, FF1, FF3 and M. thermoautotrophicum delta H. Sequence analysis showed that one FR-I element had been integrated in almost identical sequence contexts into the chromosomes of the strains CSM3 and delta H. Comparison of the FR-I elements from these strains with that from pFV1 revealed that they consisted of two subfragments, boxI (1118 bp) and boxII (383 bp), the order of which is variable. Each subfragment was identical on the sequence level with the corresponding plasmid-borne element and was flanked by terminal direct repeats with the consensus sequence A(A/T)ATTT. These results suggest that FR-I represents a mobile element. FR-II was located on both plasmids pFV1 and pFZ1, and on the chromosome of M. thermoformicicum strains THF, CSM3 and HN4. Comparison of the nucleotide sequences of the two plasmid FR-II copies and that from the chromosome of strain CSM3 showed that the FR-II segments were approximately 2.5-3.0 kb in size and contained large open reading frames (ORFs) that may encode highly related proteins with an as yet unknown function.

Amino Acid Sequence

IS870 requires a 5'-CTAG-3' target sequence to generate the stop codon for its large ORF1.

The TB regions of the Agrobacterium vitis octopine/cucumopine Ti plasmids constitute a family of related structures. All contain a bacterial insertion element downstream of the TB-iaaM gene, IS870.1. Whereas 43 isolates with octopine/cucumopine Ti plasmids carry only one IS870 copy, strain Ag57 carries a second copy (IS870.2) 3.9 kb to the right of IS870.1 and part of the same TB region. Two other octopine/cucumopine strains carry an IS870 copy on their chromosome (IS870.3). A study of the unmodified insertion sites of IS870.2 and IS870.3, cloned from closely related strains, enabled us to delimit the IS870 elements. IS870 has a size of 1,152 bp and is terminated by inverted repeats. It contains a large open reading frame without a stop codon. However, a stop codon is generated by insertion into the target sequence 5'-CTAG-3'. IS870 is related to five other insertion sequence elements. For two of these, the stop codon of the largest open reading frame is also created by insertion into a CTAG target site.

Amino Acid Sequence

[Features of the structure of the 7K-copy of Drosophila MDG4 (gypsy) retrotransposon provides evidence that the 7K-subfamily of MDG4 is potentially capable of transposition].

The complete nucleotide sequence of the 7K variant of the gypsy retrotransposon of Drosophila melanogaster was determined. This variant belongs to the 7K subfamily of gypsy, which was previously considered inactive. All differences found in the sequenced 7K copy compared to the transpositionally active 6K variants were point mutations. These nucleotide substitutions account for about 1% of the total base pair number of gypsy. Long terminal repeats (LTR) have the highest rate of nucleotide substitutions However, changes in nontranslated regions did not involve the promoter region and other supposed cis-acting elements. Sixteen amino acid substitutions were found in the coding region of gypsy. These substitutions were mainly located on the boarders of potential functional domains and within the third open reading frame (ORF3). Comparative analysis of structures of these two variants of gypsy suggests the potential ability of 7K copies to transpose.

Animals

Human papillomavirus type 13 and pygmy chimpanzee papillomavirus type 1: comparison of the genome organizations.

Human papillomavirus type 13(HPV-13) is associated with oral focal epithelial hyperplasia (FEH) in humans. A recent epidemic of a FEH-like disease in a pygmy chimpanzee (Pan paniscus) colony allowed us to clone a novel papillomavirus genome. To assess the homology between HPV-13 and the pygmy chimpanzee papillomavirus type 1 (PCPV-1), the complete nucleotide sequences of both FEH-related viruses were determined. In both viruses, all eight major open reading frames were located on one strand and the genomic organization was similar to that of other mucosal papillomaviruses. The genomes of PCPV-1 and HPV-13 showed extensive overall sequence homology (85%). They could be classified, using phylogenetic analysis, together with HPV types 6, 11, 43, and 44 in a group associated with benign orogenital lesions. These data indicate that two phylogenetically related papillomaviruses can elicit similar pathology in different primate host species, reflecting viral genomic similarities.

Animals

The COQ7 gene encodes a protein in saccharomyces cerevisiae necessary for ubiquinone biosynthesis.

Ubiquinone (coenzyme Q) is a lipid that transports electrons in the respiratory chains of both prokaryotes and eukaryotes. Mutants of Saccharomyces cerevisiae deficient in ubiquinone biosynthesis fail to grow on nonfermentable carbon sources and have been classified into eight complementation groups (coq1 coq8; Tzagoloff, A., and Dieckmann, C. L.(1990) Microbiol. Rev. 54, 211-225). In this study we show that although yeast coq7 mutants lack detectable ubiquinone, the coq7 1 mutant does synthesize demethoxyubiquinone (2-hexaprenyl-3-methyl-6-methoxy-1,4-benzoquinone), a ubiquinone biosynthetic intermediate. The corresponding wild-type COQ7 gene was isolated, sequenced, and found to restore growth on nonfermentable carbon sources and the synthesis of ubiquinone. The sequence predicts a polypeptide of 272 amino acids which is 40% identical to a previously reported Caenorhabditis elegans open reading frame. Deletion of the chromosomal COQ7 gene generates respiration defective yeast mutants deficient in ubiquinone. Analysis of several coq7 deletion strains indicates that, unlike the coq7 1 mutant, demethoxyubiquinone is not produced. Both coq7 1 and coq7 deletion mutants, like other coq mutants, accumulate an early intermediate in the ubiquinone biosynthetic pathway, 3-hexaprenyl-4-hydroxybenzoate. The data suggest that the yeast COQ7 gene may encode a protein involved in one or more monoxygenase or hydroxylase steps of ubiquinone biosynthesis.

Amino Acid Sequence

In vivo restriction by LlaI is encoded by three genes, arranged in an operon with llaIM, on the conjugative Lactococcus plasmid pTR2030.

The LlaI restriction and modification (R/M) system is encoded on pTR2030, a 46.2-kb conjugative plasmid from Lactococcus lactis. The llaI methylase gene, sequenced previously, encodes a functional type IIS methylase and is located approximately 5 kb upstream from the abiA gene, encoding abortive phage resistance. In this study, the sequence of the region between llaIM and abiA was determined and revealed four consecutive open reading frames (ORFs). Northern (RNA) analysis showed that the four ORFs were part of a 7-kb operon with llaIM and the downstream abiA gene on a separate transcriptional unit. The deduced protein sequence of ORF2 revealed a P-loop consensus motif for ATP/GTP-binding sites and a three-part consensus motif for GTP-binding proteins. Data bank searches with the deduced protein sequences for all four ORFs revealed no homology except for ORF2 with MerB, in three regions that coincided with the GTP-binding motifs in both proteins. To phenotypically analyze the llaI operon, a 9.0-kb fragment was cloned into a high-copy-number lactococcal shuttle vector, pTRKH2. The resulting construct, pTRK370, exhibited a significantly higher level of in vivo restriction and modification in L. lactis NCK203 than the low-copy-number parental plasmid, pTR2030. A combination of deletion constructions and frameshift mutations indicated that the first three ORFs were involved in LlaI restriction, and they were therefore designated llaI.1, llaI.2, and llaI.3. Mutating llaI.1 completely abolished restriction, while disrupting llaI.2 or llaI.3 allowed an inefficient restriction of phage DNA to occur, manifested primarily by a variable plaque phenotype. ORF4 had no discernible effect on in vivo restriction. A frameshift mutation in llaIM proved lethal to L. lactis NCK203, implying that the restriction component was active without the modification subunit. These results suggested that the LlaI R/M system is unlike any other R/M system studied to date and has diverged from the type IIS class of restriction enzymes by acquiring some characteristics reminiscent of type I enzymes.

Amino Acid Sequence

Nucleotide sequence of rice dwarf phytoreovirus genome segment 2: completion of sequence analyses of rice dwarf virus.

The complete nucleotide sequence of rice dwarf phytoreovirus genome segment 2 (S2) was determined to be 3,512 nucleotides long with one open reading frame initiating at nucleotide 15 and terminating at nucleotide 3363. The encoded polypeptide was predicted to have 1,116 residues with a relative molecular weight of 123 kD. Comparison of S2 of two isolates showed they had identical lengths and 97 and 98.3% nucleotide and amino acid sequence identities, respectively. A search of the Swiss-Prot data base (R 22.0) failed to find any proteins with significant homology to the S2-encoded protein. Determination of the nucleotide sequence of the S2 has completed the sequence determination of the genome of rice dwarf virus. Homology searches made for proteins encoded by each of the genomic segments showed that the polypeptide encoded by S11 has similarity to histone H1 protein and VP6 of blue tongue virus, indicating it might possess nucleic acid binding properties.

Amino Acid Sequence

Endogenous D-type (HERV-K) related sequences are packaged into retroviral particles in the placenta and possess open reading frames for reverse transcriptase.

All primates studied to date produce retroviral-like particles in their placentae. We have purified these particles from two primate species, one Old World (human) and one New World (marmoset), and have identified the retroviral sequences which are packaged into these particles. Three families of sequences have been detected in these particles in human, all of which have the highest homology to B- and D-type retroviruses and to the human endogenous retrovirus HERV-K10. Previous studies have reported that the New World monkeys do not possess sequences with homology to HERV-K10. We have identified a new family of low-copy-number sequences which are present in New World monkeys and which possess 70% homology to the HERV-K family. Particles from both species possess reverse transcriptase activity and we have found that some of these retroviral particles package sequences which encode long open reading frames in pol, as revealed by expression cloning in Escherichia coli. These open reading frames could encode the reverse transcriptase enzyme activity found in the particles.

Animals

Gorilla and orangutan c-myc nucleotide sequences: inference on hominoid phylogeny.

The nucleotide sequences of the gorilla and orangutan myc loci have been determined by the dideoxy nucleotide method. As previously observed in the human and chimpanzee sequences, an open reading frame (ORF) of 188 codons overlapping exon 1 could be deduced from the gorilla sequence. However, no such ORF appeared in the orangutan sequence. The two sequences were aligned with those of human and chimpanzee as hominoids and of gibbon and marmoset as outgroups of hominoids. The branching order in the evolution of primates was inferred from these data by different methods: maximum parsimony and neighbor-joining. Our results support the view that the gorilla lineage branched off before the human and chimpanzee diverged and strengthen the hypothesis that chimpanzee and gorilla are more related to human than is orangutan.

Amino Acid Sequence

Polymorphisms in tandemly repeated sequences of Saccharomyces cerevisiae mitochondrial DNA.

A spontaneously arising mitochondrial DNA (mtDNA) variant of Saccharomyces cerevisiae has been formed by two extra copies of a 14-bp sequence (TTAATTAAATTATC) being added to a tandem repeat of this unit. Similar polymorphisms in tandemly repeated sequences have been found in a comparison between mtDNAs from our strain and others. In 5850 bp of intergenic mtDNA sequence, polymorphisms in tandemly repeated sequences of three or more base pairs occur approximately every 400-500 bp whereas differences in 1-2 bp occur approximately every 60 bp. Some polymorphisms are associated with optional G + C-rich sequences (GC clusters). Two such optional GC clusters and one A + T repeat polymorphism have been discovered in the tRNA synthesis locus. In addition, the variable presence of large open reading frames are documented and mechanisms for generating intergenic sequence diversity in S. cerevisiae mtDNA are discussed.

Base Composition

The gamma-tubulin-encoding gene from the basidiomycete fungus, Ustilago violacea, has a long 5'-untranslated region.

The gene (gamma-tub) encoding gamma-tubulin (gamma-Tub) was isolated from a cosmid library constructed for Ustilago violacea by using a PCR-amplified DNA fragment as a probe. About 2.8 kb of DNA sequence was analyzed and found to encode a protein of 469 amino acids highly homologous to the gamma-Tub from other organisms. There were eight introns interrupting the coding sequence. A 'TATA'-like sequence was found 389 bp upstream from the initial Met codon. No polyadenylation signal was found in the 3' non-coding region. Southern blot analyses indicated that gamma-tub is a single-copy gene. Northern blot analyses indicated that a 1.81-kb RNA species was transcribed. Primer extension experiments determined that the transcription start point (tsp) is at 58 bp downstream from the putative TATA box, with another possible tsp at 95 bp downstream. The long 5' non-coding sequence of the RNA contained several small open reading frames; their possible roles in the regulation of gamma-tub translation are discussed.

Amino Acid Sequence

Human herpesvirus 6 (strain U1102) encodes homologues of the conserved herpesvirus glycoprotein gM and the alphaherpesvirus origin-binding protein.

The nucleotide sequence of 3,134 bp from the genome of human herpesvirus 6 (HHV-6) strain U1102 was determined. The sequence overlaps and is contiguous with the 21,858 bp nucleotide sequence published by us previously. The sequence reported here encodes two open reading frames, named 18L and 19R. The protein encoded by 18L shares amino acid sequence similarity with the multiply hydrophobic glycoprotein M conserved in the genomes of all herpesviruses sequenced to date. ORF 19R encodes a protein which shares a significant degree of amino acid sequence conservation with the origin-binding protein homologues encoded by members of the alphaherpesvirus subgroup, but does not share detectable amino acid sequence homology with positionally analogous open reading frames present in the genomes of other betaherpesviruses or in the genomes of gammaherpesviruses.

Amino Acid Sequence

[CHL15--a new gene controlling the replication of chromosomes in saccharomycetes yeast: cloning, physical mapping, sequencing, and sequence analysis].

We have analyzed the CHL15 gene, earlier identified in a screen for yeast mutants with increased loss of chromosome III and artificial circular and linear chromosomes in mitosis. Mutations in the CHL15 gene lead to a 100-fold increase in the rate of chromosome III loss per cell division and a 200-fold increase in the rate of marker homozygosis on this chromosome by mitotic recombination. Analysis of segregation of artificial circular minichromosome and artificially generated nonessential marker chromosome fragment indicated that sister chromatid loss (1:0 segregation) is a main reason of chromosome destabilization in the chl15-1 mutant. A genomic clone of CHL15 was isolated and used to map its physical position on chromosome XVI. Nucleotide sequence analysis of CHL15 revealed a 2.8-kb open reading frame with a 105-kD predicted protein sequence. At the N-terminal region of the protein sequences potentially able to form DNA-binding domains defined as zinc-fingers were found. The C-terminal region of the predicted protein displayed a similarity to sequence of regulatory proteins known as the helix-loop-helix (HLH) proteins. Data on partial deletion analysis suggest that the HLH domain is essential for the function of the CHL15 gene product. Analysis of the upstream untranslated region of CHL15 revealed the presence of the hexamer element, ACGCGT (an MluI restriction site) controlling both the periodic expression and coordinate regulation of the DNA synthesis genes in budding yeast. Deletion in the RAD52 gene, the product of which is involved in double-strand break/recombination repair and replication, leads to a considerable decrease in the growth rate of the chl15 mutant. We suggest that CHL15 is a new DNA synthesis gene in the yeast Saccharomyces cerevisiae.

Amino Acid Sequence

Sequence and function analysis of a 9.74 kb fragment of Saccharomyces cerevisiae chromosome X including the BCK1 gene.

In the framework of the European BIOTECH project for sequencing the Saccharomyces cerevisiae genome, we have determined the nucleotide sequence of the cosmid clone 233 provided by F. Galibert (Rennes Cedex, France). We present here 9743 base pairs of sequence derived from the left arm of chromosome X. This sequence reveals three new open reading frames and includes the published sequence (5' end and open reading frame) of the gene BCK1/SLK1/SSP31 also identified as ORFAA. Deletion mutants of two earlier unknown open reading frames J0840 and J0904 are viable and the open reading frame J0902 is essential for yeast growth.

Amino Acid Sequence

Genes of the R-phycocyanin II locus of marine Synechococcus spp., and comparison of protein-chromophore interactions in phycocyanins differing in bilin composition.

R-phycocyanin II (RPCII) is a recently discovered member of the phycocyanin family of photosynthetic light-harvesting proteins. Genes encoding the alpha and beta subunits of RPCII were cloned and sequenced from marine Synechococcus sp. strains WH8020 and WH8103. The deduced amino acid sequences of RPCII were compared to two other types of phycocyanin, C-phycocyanin (CPC) and phycoerythrocyanin (PEC). These three types vary in the composition of their covalently bound bilin prosthetic groups. In terms of amino acid sequence identity RPCII is highly homologous to CPC and PEC, suggesting that the known three-dimensional structures of the latter two are representative of RPCII. Thus the amino acid residues contacting the three bilins of RPCII could be inferred and compared to those in CPC and PEC. Certain residues were identified among the three phycocyanins as possibly correlating with specific bilin isomers. In overall sequence RPCII and CPC are more homologous to one another than either is to PEC. This probably reflects functional homology in the roles of RPCII and CPC in the transfer of light energy to the core of the phycobilisome, a function not attributed to PEC. The genomes of Synechococcus sp. strains WH8020, WH8103 and WH7803 share homologous open reading frames in the vicinity of RPCII genes. The nucleotide sequence extending 3' from RPCII genes in strain WH8020 revealed two open reading frames homologous to components of an alpha CPC phycocyanobilin lyase. These open reading frames may encode a lyase specific for the attachment of phycoerythrobilin to alpha RPCII.

Amino Acid Sequence