PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Sequence and function analysis of a 9.46 kb fragment of Saccharomyces cerevisiae chromosome X.

In the framework of the European yeast genome sequencing project, we have determined the nucleotide sequence of the cosmid clone 233 provided by F. Galibert (Rennes Cedex, France). We present here 9464 base pairs of this cosmid located on the left arm of Saccharomyces cerevisiae chromosome X. This sequence contains two new open reading frames and includes the published sequences of the RADH gene (also identified as SRS2/HPR5) and the 3'-end of the gene BCK1/SLK1/SSP31. Deletion mutants of the two unknown genes J0909 and J0911 are viable.

Amino Acid Sequence↗

A 12.5 kb fragment of the yeast chromosome II contains two adjacent genes encoding ribosomal proteins and six putative new genes, one of which encodes a putative transcriptional factor.

The nucleotide sequence of a 12.5 kb fragment localized to the right arm of chromosome II of Saccharomyces cerevisiae has been determined. The sequence contains eight putative genes. Two of them are contiguous and represent two ribosomal protein genes: SUP46 and URP1. SUP46 is implicated in translation fidelity and encodes the ribosomal protein S13. URP1 is homologous to the rat ribosomal protein gene L21. The open reading frame (ORF) YBR1245 is similar in its N-terminal part to transcription factors like SRF and MCM1. The ORF YBR1308 shows homology with proteins of the AAA-family (ATPases Associated with diverse cellular Activities). Two genes are predicted to encode putative membrane proteins.

Amino Acid Sequence↗

Nucleotide sequence analysis of four genes, hupC, hupD, hupF and hupG, downstream of the hydrogenase structural genes in Bradyrhizobium japonicum.

The nucleotide sequence of a 2.2 kb region downstream of the hydrogenase structural genes in Bradyrhizobium japonicum was determined. Four genes encoding predicted polypeptides of 27.8 (HupC), 21.4 (HupD), 10.6 (HupF) and 15.8 (HupG) kDa were identified, of which the first three probably belong to the same operon as the hup structural genes, hupS and hupL. HupC is homologous to the hydrophobic polypeptides with four potential transmembrane regions that are encoded by open reading frames following the hydrogenase structural genes in Rhodobacter capsulatus, Escherichia coli, Azotobacter vinelandii, Wolinella succinogenes, Rhizobium leguminosarum and Alcaligenes eutrophus. Also HupD, HupF and HupG are homologous to genes involved in processing, maturation, functioning and regulation of hydrogenase activity in various hydrogen-oxidizing bacteria.

Amino Acid Sequence↗

Iron transport genes of the pJM1-mediated iron uptake system of Vibrio anguillarum are included in a transposonlike structure.

The pJM1 genes encoding the proteins involved in iron transport in the anguibactin iron uptake system were found to be flanked by insertion sequences in a composite transposonlike structure. These Vibrio anguillarum insertion sequences, ISV-A1 and ISV-A2, are related to IS903, IS102, and the ISVs found in Vibrio parahaemoliticus, Vibrio mimicus, and non-O1 Vibrio cholerae flanking various tdh (thermostable direct hemolysin) genes. The inverted repeats at the ends of ISV-A1 and ISV-A2 have no more than three mismatches when compared to the inverted repeats of the other ISVs or IS903 and IS102. ISV-A1 and ISV-A2 are flanked by 9-bp direct repeats, which is the number of bases that are duplicated upon IS903 or IS102 transposition. The similarities found between the V. anguillarum ISVs and the other ISVs as well as IS903 and IS102 suggest that they derive from a common ancestral insertion sequence. At the end of ISV-A1 there is a -35 sequence region followed by a -10 sequence found in the pJM1 sequence immediately outside the ISV. This promoter region is followed by an open reading frame with the potential to encode a polypeptide of 26,985 Da whose function is still unknown. The functionality of this promoter has been demonstrated and expression analysis showed that the promoter is regulated by the iron concentration of the media.

Amino Acid Sequence↗

Sequence analysis of an Erwinia stewartii plasmid, pSW100.

The nucleotide sequence of the smallest plasmid of Erwinia stewartii SW2 was determined. This plasmid, pSW100 (4272 bp), consists of a 702-bp region homologous to the origins of replication of plasmids p15A, ColE1, and ColA. Plasmid pSW100 also contains sequences homologous to the bom region and mobCABD genes of ColE1, except that a single-base deletion in mobA has been detected. This deletion did not affect the mobilization ability of pSW100 by an endogenous conjugative plasmid of E. stewartii SW2, pDC250. Plasmid pSW100 also has a 1596-bp open reading frame with five 132-bp perfect repeats. However, this open reading frame and the repeats are not required for plasmid replication. Furthermore, curing of pSW100 did not cause any phenotypic change, suggesting that this plasmid is not essential for the survival of the organism.

Amino Acid Sequence↗

The complete nucleotide sequence of prune dwarf ilarvirus RNA 3: implications for coat protein activation of genome replication in ilarviruses.

The complete nucleotide sequence of prune dwarf ilarvirus (PDV) RNA 3 has been determined from cloned viral cDNAs. The PDV RNA 3 is 2129 nucleotides and contains two large open reading frames (ORFs) separated by an intergenic region of 72 nucleotides. The 5' proximal ORF (ORF-1) is 882 nucleotides, encoding a gene product which shares homology with putative cell-to-cell movement proteins of related viruses, including tobacco streak virus (TSV) and alfalfa mosaic virus (AIMV). The downstream ORF (ORF-2) is 657 nucleotides and encodes a gene product which shares primary sequence homology and structural features with AIMV coat protein. Furthermore, when expressed in bacteria, this ORF produces a polypeptide which comigrates with authentic PDV coat protein and reacts with PDV coat protein antiserum. Hybridization data suggest that the genomic organization of PDV RNAs 3 and 4 is similar to that of TSV, the only other ilarvirus for which sequence information is published. The 3' untranslated region (UTR) of PDV RNA 3 is similar to that of TSV and AIMV, containing a potential stem-loop structure followed by the sequence AUGC, a structure which may signal binding of coat protein and activation of genome replication. However, a striking feature of the deduced PDV coat protein sequence is the absence of a "zinc-finger" motif thought to function in binding of the coat protein to the 3'-UTR in ilarviruses and AIMV. This result suggests that the zinc-finger motif is not a required aspect of coat protein activation of replication in ilarviruses.

Amino Acid Sequence↗

Cloning and sequencing of an Escherichia coli K12 gene which encodes a polypeptide having similarity to the human ferritin H subunit.

Using lambda phage clones containing segments of the Escherichia coli K12 chromosome as hybridization probes, we found one gene at 42 min on the E. coli chromosome map, the expression of which was affected by RNase III. The sequence of the DNA fragment containing this gene (gen-165) revealed the presence of an open reading frame encoding a polypeptide of 165 amino acid residues. The amino acid sequence deduced from the nucleotide sequence exhibited a remarkable similarity to that of the human ferritin H chain.

Amino Acid Sequence↗

Cloning and sequencing of the pac gene encoding the penicillin G acylase of Bacillus megaterium ATCC 14945.

The pac gene encoding the penicillin G acylase (PGA) of Bacillus megaterium ATCC 14945 has been cloned in Escherichia coli HB101 (proA, leuB) using a selective minimal medium containing phenylacetyl-L-leucine instead of L-leucine. The nucleotide sequence of this gene has been determined and contains an open reading frame of 2406 nucleotides. The deduced amino acid sequence shows significant similarity with other beta-lactam acylases. Although the PGA of B. megaterium is extracellular, the enzyme produced in E. coli appears to have a cytoplasmic localization.

Amino Acid Sequence↗

Solenopsis invicta virus-1A (SINV-1A): distinct species or genotype of SINV-1?

We have cloned and sequenced a 2845 bp cDNA representing the 3'-end of either a new picorna-like virus species or genotype of Solenopsis invicta virus-1 (SINV-1). Analysis of the nucleotide sequence revealed 1 large open reading frame. The amino acid sequence of the translated open reading frame was most identical to structural proteins of SINV-1 (97%), followed by the Kashmir bee virus (KBV, 30%), and acute bee paralysis virus (ABPV, 29%). A PCR-based survey for SINV-1 and the new species or genotype (tentatively named S. invicta virus-1A, SINV-1A) using RNA extracts of S. invicta collected around Gainesville, Florida, revealed a mean colony infestation rate of 25% by SINV-1 and 55% by SINV-1A. Both SINV-1 and SINV-1A were found to co-infect 17.5% of the nests surveyed. Although the data preclude definitive species or genotype assignment, there is no doubt that SINV-1A is distinct from SINV-1, identifiable, and infects S. invicta. We provide a simple RT-PCR technique capable of discerning SINV-1 and SINV-1A infection of S. invicta.

Amino Acid Sequence↗

Characterisation of the 11 Kb DNA region adjacent to the gene encoding Desulfovibrio gigas flavoredoxin.

Flavoredoxin is an FMN binding protein that functions as an electron carrier in the sulphate metabolism of Desulfovibrio gigas. The neighbouring DNA regions of the gene encoding flavoredoxin were sequenced and characterised. Transcript analysis of the flavoredoxin gene resulted in a positive band corresponding to the size of the coding region, suggesting that flavoredoxin is encoded by a monocystronic unit, as previously suggested by sequence analysis. Analysis of the adjacent DNA regions revealed several interesting genes. The sequenced DNA regions contain nine open reading frames (ORFs) organised in two polycystronic and two monocystronic units. These genes encode proteins involved in different metabolic pathways, namely in DNA methylation, tRNA and rRNA modification, mRNA metabolism, cell division, CoA synthesis and lipoprotein transport across the membrane.

Blotting, Northern↗

Characterization of the ToxB gene from Pyrenophora tritici-repentis.

The ToxB gene was cloned and characterized from a race 5 isolate of Pyrenophora tritici-repentis from North Dakota. ToxB contains a 261-bp open reading frame that encodes a 23 amino acid putative signal peptide and a 64 amino acid host-selective toxin, Ptr ToxB. Analysis of Ptr ToxB from heterologous expression in Pichia pastoris confirms that ToxB encodes a host-selective toxin.

Amino Acid Sequence↗

Nucleotide sequence of the genes encoding the canine herpesvirus gB, gC and gD homologues.

The nucleotide sequence of the genes encoding the canine herpesvirus (CHV) gB, gC and gD homologues was determined. These genes are predicted to encode polypeptides of 879, 459 and 345 amino acids, respectively. Comparison of the predicted amino acid sequences of CHV gB, gC and gD with the homologous sequences from other herpesviruses indicates that CHV is an alphaherpesvirus, a conclusion that is consistent with the previous classification of this virus according to biological properties. Alignment of the homologous gB, gC and gD amino acid sequences indicates that most of the cysteine residues are conserved, suggesting that these glycoproteins possess similar tertiary structures. The nucleotide sequence of the open reading frame downstream from the CHV gC gene was also determined. The predicted amino acid sequence of this putative polypeptide appears to be homologous to a family of proteins encoded downstream from the gC gene in most, although not all, alphaherpesviruses.

Amino Acid Sequence↗

Characterization of a novel genital human papillomavirus by overlapping PCR: candHPV86 identified in cervicovaginal cells of a woman with cervical neoplasia.

A novel human papillomavirus (HPV), candHPV86, was cloned and characterized from cervicovaginal cells obtained from a 37-year-old Hispanic woman with cervical intraepithelial neoplasia grade 1 (CIN1) using an overlapping PCR technique. Primers were designed by phylogenetic alignment of closely related HPV genomes using the L1 fragment sequence amplified by GP5+/6+. The 7983 bp complete nucleotide sequence of the HPV genome was determined by sequence walking. A basic local alignment sequence tool (BLAST) homology search using the L1 open reading frame demonstrated that this HPV was most closely related to HPVHAN2294 (GenBank, AJ400628; 86% homology) and HPV84 (84% homology). candHPV86 was placed in the HPV genome homology group A3 by phylogenetic analyses. The overlapping PCR technique is applicable for characterizing the complete spectrum and variation of HPVs in a population.

Adult↗

Cloning and functional analysis of cDNAs with open reading frames for 300 previously undefined genes expressed in CD34+ hematopoietic stem/progenitor cells.

Three hundred cDNAs containing putatively entire open reading frames (ORFs) for previously undefined genes were obtained from CD34+ hematopoietic stem/progenitor cells (HSPCs), based on EST cataloging, clone sequencing, in silico cloning, and rapid amplification of cDNA ends (RACE). The cDNA sizes ranged from 360 to 3496 bp and their ORFs coded for peptides of 58-752 amino acids. Public database search indicated that 225 cDNAs exhibited sequence similarities to genes identified across a variety of species. Homology analysis led to the recognition of 50 basic structural motifs/domains among these cDNAs. Genomic exon-intron organization could be established in 243 genes by integration of cDNA data with genome sequence information. Interestingly, a new gene named as HSPC070 on 3p was found to share a sequence of 105bp in 3' UTR with RAF gene in reversed transcription orientation. Chromosomal localizations were obtained using electronic mapping for 192 genes and with radiation hybrid (RH) for 38 genes. Macroarray technique was applied to screen the gene expression patterns in five hematopoietic cell lines (NB4, HL60, U937, K562, and Jurkat) and a number of genes with differential expression were found. The resource work has provided a wide range of information useful not only for expression genomics and annotation of genomic DNA sequence, but also for further research on the function of genes involved in hematopoietic development and differentiation.

Alternative Splicing↗

Primary structure and processing of the Candida tsukubaensis alpha-glucosidase. Homology with the rabbit intestinal sucrase-isomaltase complex and human lysosomal alpha-glucosidase.

The nucleotide sequence of a 4.39-kb DNA fragment encoding the alpha-glucosidase gene of Candida tsukubaensis is reported. The cloned gene contains a major open reading frame (ORF 1) which encodes the alpha-glucosidase as a single precursor polypeptide of 1070 amino acids with a predicted molecular mass of 119 kDa. N-terminal amino acid sequence analysis of the individual subunits of the purified enzyme, expressed in the recombinant host Saccharomyces cerevisiae, confirmed that the alpha-glucosidase precursor is proteolytically processed by removal of an N-terminal signal peptide to yield the two peptide subunits 1 and 2, of molecular masses 63-65 kDa and 50-52 kDa, respectively. Both subunits are secreted by the heterologous host S. cerevisiae in a glycosylated form. Coincident with its efficient expression in the heterologous host, the C. tsukubaensis alpha-glucosidase gene contains many of the canonical features of highly expressed S. cerevisiae genes. There is considerable sequence similarity between C. tsukubaensis alpha-glucosidase, the rabbit sucrase-isomaltase complex (proSI) and human lysosomal acid alpha-glucosidase. The cloned DNA fragment from C. tsukubaensis contains a second open reading frame (ORF 2) which has the capacity to encode a polypeptide of 170 amino acids. The function and identity of the polypeptide encoded by ORF 2 is not known.

Amino Acid Sequence↗

a1/EBP: a leucine zipper protein that binds CCAAT/enhancer elements in the avian leukosis virus long terminal repeat enhancer.

Avian leukosis virus (ALV) induces bursal lymphoma in chickens after integration of proviral long terminal repeat (LTR) enhancer sequences next to the c-myc proto-oncogene. Labile LTR-binding proteins appear to be essential for c-myc hyperexpression, since both LTR-enhanced transcription and the activities of LTR-binding proteins are specifically decreased after inhibition of protein synthesis (A. Ruddell, M. Linial, W. Schubach, and M. Groudine, J. Virol. 62:2728-2735, 1988). This lability is restricted to hematopoietic cells from ALV-susceptible chicken strains, suggesting that the labile proteins play an important role in lymphomagenesis. The major labile activity binding to the a1 LTR region (A. Ruddell, M. Linial, and M. Groudine, Mol. Cell. Biol. 12:5660-5668, 1989) was purified from bursal lymphoma cells by conventional and oligonucleotide affinity chromatography, yielding three proteins of 35, 40, and 42 kDa. More than one of these species binds the a1 LTR region, as judged by gel shift analysis. A gene encoding an a1-binding protein (designated a1/EBP) was cloned by screening a bursal lymphoma cDNA library for fusion proteins binding the a1 LTR site. DNase I footprinting and gel shift assays indicate that the a1/EBP fusion protein binds multiple LTR CCAAT/enhancer elements in a pattern similar to that of the purified B-cell protein. DNA sequence analysis shows that this 2.2-kb cDNA encodes a 209-amino-acid open reading frame containing carboxy-terminal basic and leucine zipper motifs, indicating that a1/EBP encodes a novel member of the leucine zipper family of transcription factors.

Amino Acid Sequence↗

Ribin, a protein encoded by a message complementary to rRNA, modulates ribosomal transcription and cell proliferation.

The control of rRNA transcription, tightly coupled to the cell cycle and growth state of the cell, is a key process for understanding the mechanisms that drive cell proliferation. Here we describe a novel protein, ribin, found in rodents, that binds to the rRNA promoter and stimulates its activity. The protein also interacts with the basal rRNA transcription factor UBF. The open reading frame encoding ribin is 96% complementary to a central region of the large rRNA. This demonstrates that ribosomal DNA-related sequences in higher eukaryotes can be expressed as protein-coding messages. Ribin contains two predicted nuclear localization sequence elements, and green fluorescent protein-ribin fusion proteins localize in the nucleus. Cell lines overexpressing ribin exhibit enhanced rRNA transcription and faster growth. Furthermore, these cells significantly overcome the suppression of rRNA synthesis caused by serum deprivation. On the other hand, the endogenous ribin level correlates positively with the amount of serum in the medium. The data show that ribin is a limiting stimulatory factor for rRNA synthesis in vivo and suggest its involvement in the pathway that adapts ribosomal transcription and cell proliferation to physiological changes.

Amino Acid Sequence↗

Repetition, conservation, and variation: the multiple cp32 plasmids of Borrelia species.

Members of the spirochete genus Borrelia contain large numbers of extrachromosomal DNAs. Sequence analysis of the B. burgdorferi strain B31 genome indicated that its many plasmids contain large quantities of repeated sequences, the most obvious of which are the cp32 plasmid family. Individual spirochetes may carry nine or more different, but homologous, cp32 plasmids. Every other species of Borrelia examined thus far also contains multiple plasmids related to the B. burgdorferi cp32s. These plasmids are arguably the best characterized of all the borrelial plasmids, and epitomize the apparent redundancy evident in the many plasmids carried by these bacteria. Despite their extensive similarities, cp32 plasmids contain some open reading frames whose sequences often vary between plasmids, and which encode proteins synthesized by the bacteria during vertebrate infection. In this review, we analyze the hypervariable and conserved regions of the cp32 plasmid family, and discuss possible reasons why borreliae harbor multiple gene paralogs.

Amino Acid Sequence↗