PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Identification, sequences, and expression of Treponema pallidum chemotaxis genes.

Treponema pallidum, the agent of syphilis, is a pathogenic spirochete that has no known mechanisms of genetic exchange and cannot be continuously cultivated in vitro. A probe based on the nucleotide sequence of the T. pallidum cheA gene was used to screen a T. pallidum genomic DNA library. A treponemal DNA region containing four open reading frames (orfs) was identified. The proteins encoded by these orfs have significant homology with proteins involved in bacterial chemotaxis. The orfs have been designated cheA, cheW, cheX, and cheY. The cheA, cheW, and cheY genes were individually-cloned and expressed in vitro. The observed molecular mass of each protein correlated well with its predicted molecular mass. Reverse transcriptase-PCR data indicate that cheA through cheY are co-transcribed. The organization of these genes suggests that they comprise an operon. We hypothesize that the ability to sense and respond to nutrient gradients is important for the survival and dissemination of T. pallidum in vivo. The presence of a putative che operon strongly suggests that T. pallidum has the potential for a chemotactic response.

Amino Acid Sequence↗

Genome-wide atlas of gene expression in the adult mouse brain.

Molecular approaches to understanding the functional circuitry of the nervous system promise new insights into the relationship between genes, brain and behaviour. The cellular diversity of the brain necessitates a cellular resolution approach towards understanding the functional genomics of the nervous system. We describe here an anatomically comprehensive digital atlas containing the expression patterns of approximately 20,000 genes in the adult mouse brain. Data were generated using automated high-throughput procedures for in situ hybridization and data acquisition, and are publicly accessible online. Newly developed image-based informatics tools allow global genome-scale structural analysis and cross-correlation, as well as identification of regionally enriched genes. Unbiased fine-resolution analysis has identified highly specific cellular markers as well as extensive evidence of cellular heterogeneity not evident in classical neuroanatomical atlases. This highly standardized atlas provides an open, primary data resource for a wide variety of further studies concerning brain organization and function.

Animals↗

Involvement of upstream open reading frames in regulation of rat V(1b) vasopressin receptor expression.

The V(1b) vasopressin receptor, expressed mainly in the corticotroph of the anterior pituitary, mediates the stimulatory effect of vasopressin on ACTH release. To clarify the regulation of receptor expression, we cloned, sequenced (up to approximately 5 kb from the translation start site), and characterized the 5'-flanking region of the rat V(1b) receptor gene. We identified the transcription start site by amplification of cDNA ends and found a new intron within the 5'-untranslated region (5'-UTR) by comparing the sequence with that of cDNA. We then confirmed that the obtained promoter indeed has transcriptional activity by use of the luciferase reporter in AtT-20 mouse corticotroph cells. Interestingly, there were five short upstream open reading frames (uORFs) located within the 5'-UTR that were found to suppress V(1b) expression. Subsequent mutational analyses showed that the two downstream uORFs have an inhibitory effect on expression in both homologous and heterologous contexts. Furthermore, the inhibition did not accompany a parallel decrease in mRNA, suggesting that the suppressive effect occurs at a level downstream of transcription. Taken together, our data strongly suggest that the expression of the V(1b) receptor is regulated at the posttranscriptional as well as transcriptional level through uORFs within the 5'-UTR region of the mRNA. Whether the uORF-mediated regulation of V(1b) expression is functionally linked to any intracellular and/or extracellular factor(s) awaits further research.

5' Untranslated Regions↗

Subfamilies of CR1 non-LTR retrotransposons have different 5'UTR sequences but are otherwise conserved.

CR1 elements and CR1-related (CR1-like) elements are a novel family of non-LTR retrotransposons that are found in all vertebrates (reptilia, amphibia, fish, and mammals), whereas more distantly related elements are found in several invertebrate species. CR1 elements have several features that distinguish them from other non-LTR retrotransposons. Most notably, their 3' termini lack a polyadenylic acid (poly A) tail and instead contain 2-4 copies of a unique 8 bp repeat. CR1 elements are present at approximately 100,000 copies in the chicken genome. The vast majority of these elements are severely 5' truncated and mutated; however, six subfamilies (CR1-A through CR1-F) are resolved by sequence comparisons. One of these subfamilies (i.e. CR1-B) previously was analyzed in detail. In the present study, we identified several full-length elements from the CR1-F subfamily. Although regions within the open reading frames and 3' untranslated regions of CR1-F and CR1-B elements are well conserved, their respective 5' untranslated regions are unrelated. Thus, our results suggest that new CR1 subfamilies form when elements with intact open reading frames acquire new 5' UTRs, which could, in principle, function as promoters.

Amino Acid Sequence↗

Sequence of the HindIII T fragment of human cytomegalovirus, which encodes a DNA helicase.

The DNA sequence of the HindIII T fragment of human cytomegalovirus strain AD169 has been determined. This 6225 bp sequence has been analysed for transcription signals and probable open reading frames. Similarities with herpes simplex virus, varicella-zoster virus and Epstein-Barr virus genes were observed for three of the predicted open reading frames; a virion protein and a unique DNA helicase are believed to be the functional products of two of these open reading frames. Two other open reading frames are novel in that no homologues could be found, either in the known herpesvirus sequences or in the Protein Identification Resource database. Both of these open reading frames also lie in the genomic coding region of a 5.0 kb RNA which is transcribed throughout the infectious cycle.

Amino Acid Sequence↗

Pox proteomics: mass spectrometry analysis and identification of Vaccinia virion proteins.

BACKGROUND: Although many vaccinia virus proteins have been identified and studied in detail, only a few studies have attempted a comprehensive survey of the protein composition of the vaccinia virion. These projects have identified the major proteins of the vaccinia virion, but little has been accomplished to identify the unknown or less abundant proteins. Obtaining a detailed knowledge of the viral proteome of vaccinia virus will be important for advancing our understanding of orthopoxvirus biology, and should facilitate the development of effective antiviral drugs and formulation of vaccines. RESULTS: In order to accomplish this task, purified vaccinia virions were fractionated into a soluble protein enriched fraction (membrane proteins and lateral bodies) and an insoluble protein enriched fraction (virion cores). Each of these fractions was subjected to further fractionation by either sodium dodecyl sulfate-polyacrylamide gel electophoresis, or by reverse phase high performance liquid chromatography. The soluble and insoluble fractions were also analyzed directly with no further separation. The samples were prepared for mass spectrometry analysis by digestion with trypsin. Tryptic digests were analyzed by using either a matrix assisted laser desorption ionization time of flight tandem mass spectrometer, a quadrupole ion trap mass spectrometer, or a quadrupole-time of flight mass spectrometer (the latter two instruments were equipped with electrospray ionization sources). Proteins were identified by searching uninterpreted tandem mass spectra against a vaccinia virus protein database created by our lab and a non-redundant protein database. CONCLUSION: Sixty three vaccinia proteins were identified in the virion particle. The total number of peptides found for each protein ranged from 1 to 62, and the sequence coverage of the proteins ranged from 8.2% to 94.9%. Interestingly, two vaccinia open reading frames were confirmed as being expressed as novel proteins: E6R and L3L.

Amino Acid Sequence↗

RIRE2, a novel gypsy-type retrotransposon from rice.

The 441-bp DNA segment in a PCR-amplified fragment from Oryza sativa cv. IR36 was found to have a sequence with features characteristic of LTRs of retroelements, which was named RIRE2 (Rice retroelement #2) and further analyzed. Cloning and sequencing analyses of the DNA segments connected to LTR-like sequence showed that RIRE2 has a long internal region almost 10 kb long that is flanked by LTR-like sequences. This internal region carries a primer binding site (PBS) and polypurine tract (PPT) which are necessary for cDNA synthesis of retroelements. The PBS sequence is complementary to the 3' end region of tRNA(Arg). The internal region has an rt gene homologous to that of gypsy-type retrotransposons, evidence that RIRE2 is indeed a retrotransposon related to gypsy from Drosophila. RIRE2 has an extra sequence more than 4 kb long in the region downstream of gag-pol. Phylogenetic analysis of the putative amino-acid sequences of the rt gene as well as the int gene showed that RIRE2 is related to a group of gypsy-type retrotransposons of a large size that include Grande1-4 of teosinte, Tat4-1 and Athila1-1 of Arabidopsis thaliana, and Cyclops-2 of pea, but distantly related to any other group of gypsy-type retrotransposons, including RIRE3 and RIRE8 of rice. RIRE2 and Grande1-4 had the highest homology in the gag-pol region, but the nucleotide sequences of the LTR regions differed. Both elements had significant homology in the middle area of the extra regions downstream of gag-pol, in which they had an open reading frame encoding a protein with no known function on the opposite strand from that coding for gag-pol.

Amino Acid Sequence↗

The complete sequence of a 6794 bp segment located on the right arm of chromosome II of Saccharomyces cerevisiae. Finding of a putative dUTPase in a yeast.

The DNA sequence of a 6794 bp fragment located at about 100 kb from the right telomere of chromosome II from Saccharomyces cerevisiae has been determined. Sequence analysis reveals five open reading frames. One is the ARO4 gene encoding the 3-deoxy-D-arabinoheptulosonate 7-phosphate synthase. Another presents strong homology with the S5 ribosomal protein from bacteria. The open reading frame YBR1705 shows significant homology with dUTPase, suggesting for the first time the existence of such an enzyme in S. cerevisiae.

Amino Acid Sequence↗

Structure and expression of a root-specific rice gene.

Two rice cDNA clones (COS6 and COS9) were isolated, corresponding to genes that were highly expressed in roots from seedlings and mature plants. A genomic clone (GOS9) corresponding to cDNA clone COS9 was isolated and the intron/exon structure was determined by comparing the nucleotide sequences of the mRNA and the genomic clone. 5' ends and 3' ends of the mRNA were determined by primer extension and S1-nuclease mapping respectively. The open reading frame present in GOS9 potentially encodes a protein (14 kDa) that does not show any significant homology to other proteins in databases.

Amino Acid Sequence↗

Variants of the 5'-untranslated sequence of human growth hormone receptor mRNA.

The human growth hormone receptor (GHR) gene was proposed to contain multiple 5'-noncoding exons (Leung et al., 1987). The exact number and structure of these exons are unknown. As a first step in investigating this point more closely, we decided to clone alternative 5'-noncoding sequences of human liver GHR mRNA. The ligation-mediated single-sided polymerase chain reaction (PCR) was applied for selective amplification of 5'-terminal sequences of human liver GHR cDNA. PCR products were cloned and sequenced. Eight different sequence variants diverging in the 5'-untranslated regions beginning 12 base pairs upstream from the initiating ATG codon were found. One variant seems to represent unspliced or partially spliced GHR mRNA. The remaining variants probably correspond to multiple alternatively spliced forms of GHR mRNA. Homologs for three of these variants were found among previously published 5'-noncoding sequences of GHR cDNA obtained from other species by conventional cDNA cloning. Most of the cloned human liver GHR cDNA variants contain one or more ATG preceding the main GHR open reading frame start of translation. Thus, the GHR genes appeared to be a striking example of a very complex transcription unit.

Animals↗

Sequence analysis of the sbsA gene encoding the 130-kDa surface-layer protein of Bacillus stearothermophilus strain PV72.

Bacillus stearothermophilus (Bs) contains a surface-layer (S-layer) protein (SbsA), which forms a hexagonal array on the cell wall. In order to understand the structural/functional relationship of SbsA from Bs PV72, the entire nucleotide (nt) sequence of the sbsA gene was determined from three overlapping fragments. The 3'-end was cloned and expressed in Escherichia coli, whereas the 5'-region was amplified from the genome of Bs PV72 by the polymerase chain reaction using two overlapping fragments. The open reading frame (3684 nt) of sbsA is predicted to encode a protein of 1228 amino acids (aa). The SbsA is synthesized with a leader sequence of 30 aa. The predicted SbsA aa profile was similar to most other sequenced S-layer proteins, containing more acidic than basic aa (pI 5.1) and a very low amount of sulfur-containing aa. Based on aa sequence data, SbsA has weak homology of with the S-layer proteins from B. sphaericus, Rickettsia rickettsii, B. brevis HPD31 and B. brevis 47 (OWP).

Amino Acid Sequence↗

Molecular characterization of a prophage of Salmonella enterica serotype Typhimurium DT104.

Isolates of the Salmonella enterica serotype Typhimurium definitive phage type (DT104) were found to contain the same prophage (designated phage ST104). The complete sequence of the DNA genome of prophage ST104 was determined. The entire DNA sequence consisted of 41,391 bp, including 64 open reading frames, and exhibited high similarity to P22 and to phage type conversion phage ST64T.

Genome, Viral↗

Sequence analysis of the putative E3 region of bovine adenovirus type 2.

Adenoviruses are nonenveloped icosahedral-shaped particles. The double-stranded viral DNA genome contains four major early transcription units, designated E1 (a and b), E2 (a and b), E3 and E4, which are expressed in a regulated manner soon after infection. The gene products of the region E3, shown to be nonessential for viral replication in vitro, are believed to be involved in counteracting host immunosurveillance. Human adenovirus type 5 DNA sequences of transcription units L4 and L5 adjacent to E3 were used to localize E3 within the bovine adenovirus type 2. The DNA sequences between 74.8 and 84.4 mu containing E3 and the fiber gene were determined. The E3 region was found to consist of about 2.3 kb pairs and to encode four proteins longer than 60 amino acids. However, these four open reading frames did not show significant homology to any other known adenovirus DNA or protein sequence.

Adenovirus E3 Proteins↗

The marine pathogen Vibrio vulnificus encodes a putative homologue of the Vibrio harveyi regulatory gene, luxR: a genetic and phylogenetic comparison.

Vibrio vulnificus is an opportunistic pathogen that exhibits numerous virulence factors, including the secretion of a zinc metalloprotease and the production of a capsule. We have cloned and sequenced a gene from V. vulnificus that is a homologue of the positive transcriptional regulator, luxR, of the lux operon in Vibrio harveyi. This gene encodes a putative, single complete open reading frame designated smcR, which shares greater than 75% nucleotide identity with luxR of V. harveyi. The deduced amino acid sequence of the putative SmcR protein is more than 90% identical and 95% similar to that of LuxR of V. harveyi, suggesting that V. vulnificus possesses a member of the family of signal-response genes recently described in Vibrio cholerae and in Vibrio parahaemolyticus. Our data also demonstrate that, in addition to V. vulnificus, all six Vibrio spp. tested contained genes that hybridized with the luxR probe. We also present evidence that this regulatory protein was inherited from a common ancestor, and that the gene is ancient and widespread in marine Vibrio spp.

Amino Acid Sequence↗

Swt1, a novel yeast protein, functions in transcription.

The conserved TREX complex couples transcription to nuclear mRNA export. Here, we report that the uncharacterized open reading frame YOR166c genetically interacts with TREX complex components and encodes a novel protein named Swt1 for "synthetically lethal with TREX." Co-immunoprecipitation experiments show that Swt1 also interacts with the TREX complex biochemically. Consistent with a potential role in transcription as suggested by its interaction with TREX, Swt1 localizes mainly to the nucleus. Importantly, deletion of Swt1 leads to decreased transcription. Taken together, these data suggest that Swt1 functions in gene expression in conjunction with the TREX complex.

Amino Acid Sequence↗

Molecular analysis of the En/Spm transposable element system of Zea mays.

The nucleotide sequence of the autonomous transposable element En-1 isolated from the wx-844::En-1 allele has been determined. En-1 is 8287 bp long. The structure of the mosaic gene 1, coding for the major En transcript, has been established. The promoter gene 1 is located in the highly structured left end of the element and the gene spans almost the entire length of En-1. The first intron of gene 1 is 4434 nucleotides long and contains two large open reading frames, 2714 bp and 761 bp in size, which hybridize to minor RNA species in Northern blot experiments.

Alleles↗

Sorghum mitochondrial atp6: divergent amino extensions to a conserved core polypeptide.

Sorghum mitochondrial atp6 occurs as one copy in the line Tx398 and as two copies in IS1112C. In IS1112C a repeated sequence diverged within the atp6 open reading frames. The two open reading frames (1137 bp, atp6-1; 1002 bp, atp6-2) share an identical conserved region of 756 bp but are flanked 5' by divergent extensions of 246 (atp6-1) or 381 bp (atp6-2). Tx398 carried only atp6-2. The breakpoint of the repeated sequence of the conserved core region corresponds to the amino acid sequence Ser-Pro-Leu-Asp, which is the amino terminus of the proteolytically processed yeast ATP6. The 5' extensions of atp6-1 and atp6-2 were similar to those of rice and maize, respectively. Each open reading is transcribed, however nuclear background influenced transcriptional patterns of atp6-2 in IS1112C.

Amino Acid Sequence↗