PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Database Management Systems”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Chemical effects in biological systems--data dictionary (CEBS-DD): a compendium of terms for the capture and integration of biological study design description, conventional phenotypes, and 'omics data.

A critical component in the design of the Chemical Effects in Biological Systems (CEBS) Knowledgebase is a strategy to capture toxicogenomics study protocols and the toxicity endpoint data (clinical pathology and histopathology). A Study is generally an experiment carried out during a period of time for the purpose of obtaining data, and the Study Design Description captures the methods, timing, and organization of the Study. The CEBS Data Dictionary (CEBS-DD) has been designed to define and organize terms in an attempt to standardize nomenclature needed to describe a toxicogenomics Study in a structured yet intuitive format and provide a flexible means to describe a Study as conceptualized by the investigator. The CEBS-DD will organize and annotate information from a variety of sources, thereby facilitating the capture and display of toxicogenomics data in biological context in CEBS, i.e., associating molecular events detected in highly-parallel data with the toxicology/pathology phenotype as observed in the individual Study Subjects and linked to the experimental treatments. The CEBS-DD has been developed with a focus on acute toxicity studies, but with a design that will permit it to be extended to other areas of toxicology and biology with the addition of domain-specific terms. To illustrate the utility of the CEBS-DD, we present an example of integrating data from two proteomics and transcriptomics studies of the response to acute acetaminophen toxicity (A. N. Heinloth et al., 2004, Toxicol. Sci. 80, 193-202).

Acetaminophen↗

'PACLIMS': a component LIM system for high-throughput functional genomic analysis.

BACKGROUND: Recent advances in sequencing techniques leading to cost reduction have resulted in the generation of a growing number of sequenced eukaryotic genomes. Computational tools greatly assist in defining open reading frames and assigning tentative annotations. However, gene functions cannot be asserted without biological support through, among other things, mutational analysis. In taking a genome-wide approach to functionally annotate an entire organism, in this application the approximately 11,000 predicted genes in the rice blast fungus (Magnaporthe grisea), an effective platform for tracking and storing both the biological materials created and the data produced across several participating institutions was required. RESULTS: The platform designed, named PACLIMS, was built to support our high throughput pipeline for generating 50,000 random insertion mutants of Magnaporthe grisea. To be a useful tool for materials and data tracking and storage, PACLIMS was designed to be simple to use, modifiable to accommodate refinement of research protocols, and cost-efficient. Data entry into PACLIMS was simplified through the use of barcodes and scanners, thus reducing the potential human error, time constraints, and labor. This platform was designed in concert with our experimental protocol so that it leads the researchers through each step of the process from mutant generation through phenotypic assays, thus ensuring that every mutant produced is handled in an identical manner and all necessary data is captured. CONCLUSION: Many sequenced eukaryotes have reached the point where computational analyses are no longer sufficient and require biological support for their predicted genes. Consequently, there is an increasing need for platforms that support high throughput genome-wide mutational analyses. While PACLIMS was designed specifically for this project, the source and ideas present in its implementation can be used as a model for other high throughput mutational endeavors.

Algorithms↗

PASBio: predicate-argument structures for event extraction in molecular biology.

BACKGROUND: The exploitation of information extraction (IE), a technology aiming to provide instances of structured representations from free-form text, has been rapidly growing within the molecular biology (MB) research community to keep track of the latest results reported in literature. IE systems have traditionally used shallow syntactic patterns for matching facts in sentences but such approaches appear inadequate to achieve high accuracy in MB event extraction due to complex sentence structure. A consensus in the IE community is emerging on the necessity for exploiting deeper knowledge structures such as through the relations between a verb and its arguments shown by predicate-argument structure (PAS). PAS is of interest as structures typically correspond to events of interest and their participating entities. For this to be realized within IE a key knowledge component is the definition of PAS frames. PAS frames for non-technical domains such as newswire are already being constructed in several projects such as PropBank, VerbNet, and FrameNet. Knowledge from PAS should enable more accurate applications in several areas where sentence understanding is required like machine translation and text summarization. In this article, we explore the need to adapt PAS for the MB domain and specify PAS frames to support IE, as well as outlining the major issues that require consideration in their construction. RESULTS: We introduce PASBio by extending a model based on PropBank to the MB domain. The hypothesis we explore is that PAS holds the key for understanding relationships describing the roles of genes and gene products in mediating their biological functions. We chose predicates describing gene expression, molecular interactions and signal transduction events with the aim of covering a number of research areas in MB. Analysis was performed on sentences containing a set of verbal predicates from MEDLINE and full text journals. Results confirm the necessity to analyze PAS specifically for MB domain. CONCLUSIONS: At present PASBio contains the analyzed PAS of over 30 verbs, publicly available on the Internet for use in advanced applications. In the future we aim to expand the knowledge base to cover more verbs and the nominal form of each predicate.

Cybernetics↗

Sequence database searches via de novo peptide sequencing by tandem mass spectrometry.

A method is described for searching protein sequence databases using tandem mass spectra of tryptic peptides. The approach uses a de novo sequencing algorithm to derive a short list of possible sequence candidates which serve as query sequences in a subsequent homology-based database search routine. The sequencing algorithm employs a graph theory approach similar to previously described sequencing programs. In addition, amino acid composition, peptide sequence tags and incomplete or ambiguous Edman sequence data can be used to aid in the sequence determinations. Although sequencing of peptides from tandem mass spectra is possible, one of the frequently encountered difficulties is that several alternative sequences can be deduced from one spectrum. Most of the alternative sequences, however, are sufficiently similar for a homology-based sequence database search to be possible. Unfortunately, the available protein sequence database search algorithms (e.g. Blast or FASTA) require a single unambiguous sequence as input. Here we describe how the publicly available FASTA computer program was modified in order to search protein databases more effectively in spite of the ambiguities intrinsic in de novo peptide sequencing algorithms.

Algorithms↗

Evaluation of narrative text for case finding: the need for accuracy measurement.

This article reviews the analysis of a narrative text electronic search technique being used in the insurance industry. We reviewed a previously published study of motor vehicle crashes in roadway construction workzones as well as additional data supplied by the authors with respect to the methods of keyword selection. The narrative text search technique was evaluated with decision statistics and was found to have a sensitivity of 92.3%, 95% confidence interval 67.5%-99.6%. This range of sensitivity, at its most extreme value, led to a 32.5% underestimation of claims prevalence. Furthermore, because the electronic search developed two classification categories from a limited text field (approximately 20 words), only half of the cases had at least one classification. Systematic error estimates were used to obtain true population proportions for crash characteristics, revealing significant underestimations in costs. This analysis highlights the need for investigators to apply decision statistics to narrative text searching techniques when they are used essentially as diagnostic test procedures on insurance claims datasets.

Accidents, Occupational↗

UMD (Universal mutation database): a generic software to build and analyze locus-specific databases.

The human genome is thought to contain about 80,000 genes and presently only 3,000 are known to be implicated in genetic diseases. In the near future, the entire sequence of the human genome will be available and the development of new methods for point mutation detection will lead to a huge increase in the identification of genes and their mutations associated with genetic diseases as well as cancers, which is growing in frequency in industrial states. The collection of these mutations will be critical for researchers and clinicians to establish genotype/phenotype correlations. Other fields such as molecular epidemiology will also be developed using these new data. Consequently, the future lies not in simple repositories of locus-specific mutations but in dynamic databases linked to various computerized tools for their analysis and that can be directly queried on-line. To meet this goal, we devised a generic software called UMD (Universal Mutation Database). It was developed as a generic software to create locus-specific databases (LSDBs) with the 4(th) Dimension(R) package from ACI. This software includes an optimized structure to assist and secure data entry and to allow the input of various clinical data. Thanks to the flexible structure of the UMD software, it has been successfully adapted to nine genes either involved in cancer (APC, P53, RB1, MEN1, SUR1, VHL, and WT1) or in genetic diseases (FBN1 and LDLR). Four new LSDBs are under construction (VLCAD, MCAD, KIR6, and COL4A5). Finally, the data can be transferred to core databases.

Chromosome Mapping↗

Database and software for the analysis of mutations at the lacZ locus in transgenic rodents.

The use of transgenic rodents is becoming increasingly widespread in genetic toxicology. In an effort to centralize and standardize the information regarding mutations in rodents bearing the lacZ transgene, we have created a computerized database that contains published information about DNA sequence alterations on over 100 mutants. Information on the literature citation, mutagenic conditions, organs from specific animals, mutation frequency in each organ, specific mutation, amino acid change, and other data are provided for each mutant. We have also produced a software package for the analysis of the lacZ database. Routines have been developed for the analysis of single base substitutions, including programs to 1) determine whether two mutational spectra are statistically different, 2) determine whether mutations show a DNA strand bias, 3) determine the frequency of transitions and transversions, 4) display the number and kind of mutations observed at each base in the coding region, 5) perform nearest-neighbor analysis, and 6) display mutable amino acids in the lacZ protein. The software runs only on IBM-compatible machines running Microsoft Windows. The software and lacZ database are freely available via the Internet (http:@sunsite.unc.edu/dnam/mainpage.ht ml). These programs simplify the analysis of the rapidly increasing information about lacZ mutation. The programs permit the facile comparison between different lacZ data sets as well as the identification of mutational patterns that may be of importance to experimenters studying the mechanisms of mutation and mutational spectra in transgenic animals.

Amino Acids↗

The breast cancer information core: database design, structure, and scope.

The Breast Cancer Information Core (BIC) is an open access, on-line mutation database for breast cancer susceptibility genes. In addition to creating a catalogue of all mutations and polymorphisms in breast cancer susceptibility genes, a principle aim of the BIC is to facilitate the detection and characterization of these genes by providing technical support in the form of mutation detection protocols, primer sequences, and reagent access. Additional information at the site includes a literature review compiled from published studies, links to other internet-based, breast cancer information and research resources, and an interactive discussion forum which enables investigators to post or respond to questions and/or comments on a bulletin board. Hum Mutat 16:123-131, 2000. Published 2000 Wiley-Liss, Inc.

Algorithms↗

Capturing cases in workers' compensation databases: the example of neck pain.

BACKGROUND: There is a need to more accurately enumerate workers with musculoskeletal injuries who make lost-time claims to workers compensation boards. The objective of this study is to develop an approach to more accurately enumerate these workers. METHODS: Lost-time claims to the Ontario Workplace Safety & Insurance Board (WSIB) were reviewed. Using neck pain as an example, nature of injury and part of body codes were identified to classify cases. Claims of a random sample of 434 claimants were reviewed. The proportion of claimants classified as having neck pain was computed. RESULTS: The proportion of claimants classified with soft-tissue injuries to the neck varied from 0.88 for codes including "neck/cervical region," 0.69 for "back region" to 0.05 for those coded as "shoulder/upper arm." CONCLUSIONS: Restricting the enumeration of injuries to specific part of body codes can lead to a gross underestimation of the magnitude of soft-tissue disorders in epidemiological studies using workers' compensation data. The proposed approach leads to more accurate enumeration.

Accidents, Occupational↗

Environmental conditions and transcriptional regulation in Escherichia coli: a physiological integrative approach.

Bacteria develop a number of devices for sensing, responding, and adapting to different environmental conditions. Understanding within a genomic perspective how the transcriptional machinery of bacteria is modulated, as a response for changing conditions, is a major challenge for biologists. Knowledge of which genes are turned on or turned off under specific conditions is essential for our understanding of cell behavior. In this study we describe how the information pertaining to gene expression and associated growth conditions (even with very little knowledge of the associated regulatory mechanisms) is gathered from the literature and incorporated into RegulonDB, a database on transcriptional regulation and operon organization in E. coli. The link between growth conditions, signal transduction, and transcriptional regulation is modeled in the database in a simple format that highlights biological relevant information. As far as we know, there is no other database that explicitly clarifies the effect of environmental conditions on gene transcription. We discuss how this knowledge constitutes a benchmark that will impact future research aimed at integration of regulatory responses in the cell; for instance, analysis of microarrays, predicting culture behavior in biotechnological processes, and comprehension of dynamics of regulatory networks. This integrated knowledge will contribute to the future goal of modeling the behavior of E. coli as an entire cell. The RegulonDB database can be accessed on the web at the URL: http://www.cifn.unam.mx/Computational_Biology/regulondb/.

Adaptation, Physiological↗

CASRdb: calcium-sensing receptor locus-specific database for mutations causing familial (benign) hypocalciuric hypercalcemia, neonatal severe hyperparathyroidism, and autosomal dominant hypocalcemia.

Familial hypocalciuric hypercalcemia (FHH) is caused by heterozygous loss-of-function mutations in the calcium-sensing receptor (CASR), in which the lifelong hypercalcemia is generally asymptomatic. Homozygous loss-of-function CASR mutations manifest as neonatal severe hyperparathyroidism (NSHPT), a rare disorder characterized by extreme hypercalcemia and the bony changes of hyperparathyroidism, which occur in infancy. Activating mutations in the CASR gene have been identified in several families with autosomal dominant hypocalcemia (ADH), autosomal dominant hypoparathyroidism, or hypocalcemic hypercalciuria. Individuals with ADH may have mild hypocalcemia and relatively few symptoms. However, in some cases seizures can occur, especially in younger patients, and these often happen during febrile episodes due to intercurrent infection. Thus far, 112 naturally-occurring mutations in the human CASR gene have been reported, of which 80 are unique and 32 are recurrent. To better understand the mutations causing defects in the CASR gene and to define specific regions relevant for ligand-receptor interaction and other receptor functions, the data on mutations were collected and the information was centralized in the CASRdb (www.casrdb.mcgill.ca), which is easily and quickly accessible by search engines for retrieval of specific information. The information can be searched by mutation, genotype-phenotype, clinical data, in vitro analyses, and authors of publications describing the mutations. CASRdb is regularly updated for new mutations and it also provides a mutation submission form to ensure up-to-date information. The home page of this database provides links to different web pages that are relevant to the CASR, as well as disease clinical pages, sequence of the CASR gene exons, and position of mutations in the CASR. The CASRdb will help researchers to better understand and analyze the mutations, and aid in structure-function analyses.

Database Management Systems↗

Common interchange standards for proteomics data: Public availability of tools and schema.

The Proteomics Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparision, exchange and verification. To this end, a Level 1 Molecular Interaction XML data exchange format has been developed which has been accepted for publication and is freely available at the PSI website (http.//psidev.sf.net/). Several major protein interaction databases are already making data available in this format. A draft XML interchange format for mass spectrometry data has been written and is currently undergoing evaluation whilst work is ongoing to develop a proteomics data integration model, MIAPE.

Computational Biology↗

Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories.

With the human Plasma Proteome Project (PPP) pilot phase completed, the largest and most ambitious proteomics experiment to date has reached its first milestone. The correspondingly impressive amount of data that came from this pilot project emphasized the need for a centralized dissemination mechanism and led to the development of a detailed, PPP specific data gathering infrastructure at the University of Michigan, Ann Arbor as well as the protein identifications database project at the European Bioinformatics Institute as a general proteomics data repository. One issue that crept up while discussing which data to store for the PPP concerns whether the raw, binary data coming from the mass spectrometers should be stored, or rather the more compact and already significantly processed peak lists. As this debate is not restricted to the PPP but relates to the proteomics community in general, we will attempt to detail the relative merits and caveats associated with centralized storage and dissemination of raw data and/or peak lists, building on the extensive experience gained during the PPP pilot phase. Finally, some suggestions are made for both immediate and future storage of MS data in public repositories.

Computational Biology↗

PRIDE: the proteomics identifications database.

The advent of high-throughput proteomics has enabled the identification of ever increasing numbers of proteins. Correspondingly, the number of publications centered on these protein identifications has increased dramatically. With the first results of the HUPO Plasma Proteome Project being analyzed and many other large-scale proteomics projects about to disseminate their data, this trend is not likely to flatten out any time soon. However, the publication mechanism of these identified proteins has lagged behind in technical terms. Often very long lists of identifications are either published directly with the article, resulting in both a voluminous and rather tedious read, or are included on the publisher's website as supplementary information. In either case, these lists are typically only provided as portable document format documents with a custom-made layout, making it practically impossible for computer programs to interpret them, let alone efficiently query them. Here we propose the proteomics identifications (PRIDE) database (http://www.ebi.ac.uk/pride) as a means to finally turn publicly available data into publicly accessible data. PRIDE offers a web-based query interface, a user-friendly data upload facility, and a documented application programming interface for direct computational access. The complete PRIDE database, source code, data, and support tools are freely available for web access or download and local installation.

Computational Biology↗

PhosphaBase: an ontology-driven database resource for protein phosphatases.

PhosphaBase is an ontology-driven database resource containing information on the protein phosphatase family. It is the first public resource dedicated to protein phosphatases, which are enzymes that perform dephosphorylation reactions. In conjunction with the phosphorylation action of protein kinases, phosphatases are involved in important control and communication mechanisms in the cell. They have also been implicated in many human diseases, including diabetes and obesity, cancers, and neurodegenerative conditions. PhosphaBase aims to centralize the growing base of knowledge in the phosphatase research domain. The resource is built around a formal, domain-specific DAML+OIL ontology, and the data are collected from heterogeneous biological sources using Gene Ontology terms as a means of data extraction. The overall ontology-driven architecture provides a robust structure with distinct advantages for sustainability and provides the potential for the development of diagnostic tools, as well as a data repository.

Animals↗

MS1, MS2, and SQT-three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications.

As the speed with which proteomic labs generate data increases along with the scale of projects they are undertaking, the resulting data storage and data processing problems will continue to challenge computational resources. This is especially true for shotgun proteomic techniques that can generate tens of thousands of spectra per instrument each day. One design factor leading to many of these problems is caused by storing spectra and the database identifications for a given spectrum as individual files. While these problems can be addressed by storing all of the spectra and search results in large relational databases, the infrastructure to implement such a strategy can be beyond the means of academic labs. We report here a series of unified text file formats for storing spectral data (MS1 and MS2) and search results (SQT) that are compact, easily parsed by both machine and humans, and yet flexible enough to be coupled with new algorithms and data-mining strategies.

Database Management Systems↗

Adopting a database as a solution to managing electron image data.

A database was used for data management and interprogram communication in an image processing and three-dimensional reconstruction program suite for biological bundles. The programs were modified from the MRC crystallographic package. The database server works with local and remote programs and data sets, allows simultaneous requests from multiple clients, and maintains multiple databases and data tables within them. It has built-in security for the data access. Several graphical user interfaces are available to view and/or edit data tables. In addition, FORTRAN interface and function libraries are written to communicate with image processing software. The data management overhead is inexpensive, requiring only narrow bandwidth from the network. It easily handles several data tables with over 1000 entries.

Computer Communication Networks↗