PubMed HealthSearch

PubMed · 9399793

The Genome Sequence DataBase (GSDB): improving data quality and data access.

Abstract

In 1997 the primary focus of the Genome Sequence DataBase (GSDB; www. ncgr.org/gsdb ) located at the National Center for Genome Resources was to improve data quality and accessibility. Efforts to increase the quality of data within the database included two major projects; one to identify and remove all vector contamination from sequences in the database and one to create premier sequence sets (including both alignments and discontiguous sequences). Data accessibility was improved during the course of the last year in several ways. First, a graphical database sequence viewer was made available to researchers. Second, an update process was implemented for the web-based query tool, Maestro. Third, a web-based tool, Excerpt, was developed to retrieve selected regions of any sequence in the database. And lastly, a GSDB flatfile that contains annotation unique to GSDB (e.g., sequence analysis and alignment data) was developed. Additionally, the GSDB web site provides a tool for the detection of matrix attachment regions (MARs), which can be used to identify regions of high coding potential. The ultimate goal of this work is to make GSDB a more useful resource for genomic comparison studies and gene level studies by improving data quality and by providing data access capabilities that are consistent with the needs of both types of studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

C Harger, M Skupski, J Bingham, A Farmer, S Hoisie, P Hraber, D Kiphart, L Krakowski, M McLeod, J Schwertfeger, G Seluja, A Siepel, G Singh, D Stamper, P Steadman, N Thayer, R Thompson, P Wargo, M Waugh, J J Zhuang, P A Schad. 1998-01-01. The Genome Sequence DataBase (GSDB): improving data quality and data access.. https://doi.org/10.1093/nar%2F26.1.21

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence

Identification and characterization of Prp45p and Prp46p, essential pre-mRNA splicing factors.

Through exhaustive two-hybrid screens using a budding yeast genomic library, and starting with the splicing factor and DEAH-box RNA helicase Prp22p as bait, we identified yeast Prp45p and Prp46p. We show that as well as interacting in two-hybrid screens, Prp45p and Prp46p interact with each other in vitro. We demonstrate that Prp45p and Prp46p are spliceosome associated throughout the splicing process and both are essential for pre-mRNA splicing. Under nonsplicing conditions they also associate in coprecipitation assays with low levels of the U2, U5, and U6 snRNAs that may indicate their presence in endogenous activated spliceosomes or in a postsplicing snRNP complex.

Base Sequence

Solvent organization in an oligonucleotide crystal. The structure of d(GCGAATTCG)2 at atomic resolution.

We describe the crystal structure of d(GCGAATTCG) determined by x-ray diffraction at atomic resolution level (0.89 A). The duplex structure is practically identical to that described at 2.05 A resolution (Van Meervelt, L., Vlieghe, D., Dautant, A., Gallois, B., Précigoux, G., and Kennard, O. (1995) Nature 374, 742-744), however about half of the phosphate groups show multiple conformations. The crystal has three regions with different solvent structure. One of them contains several ordered Mg(+2) ions and can be considered as an ionic crystal. A second region is formed by a network of ordered water molecules with a polygonal organization that binds three duplexes. The third region is formed by channels of solvent in which very few ordered solvent molecules are visible. The less ordered phosphates are found facing this channel. The latter region provides a view of DNA with highly movable charges, both negative phosphates and counterions, without a precise location.

Base Sequence