PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Database Management Systems”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Evaluation of two dependency parsers on biomedical corpus targeted at protein-protein interactions.

We present an evaluation of Link Grammar and Connexor Machinese Syntax, two major broad-coverage dependency parsers, on a custom hand-annotated corpus consisting of sentences regarding protein-protein interactions. In the evaluation, we apply the notion of an interaction subgraph, which is the subgraph of a dependency graph expressing a protein-protein interaction. We measure the performance of the parsers for recovery of individual dependencies, fully correct parses, and interaction subgraphs. For Link Grammar, an open system that can be inspected in detail, we further perform a comprehensive failure analysis, report specific causes of error, and suggest potential modifications to the grammar. We find that both parsers perform worse on biomedical English than previously reported on general English. While Connexor Machinese Syntax significantly outperforms Link Grammar, the failure analysis suggests specific ways in which the latter could be modified for better performance in the domain.

Abstracting and Indexing↗

The genome-enabled electronic medical record.

The integration of patient-specific genomic information into the electronic medical record (EMR) will create many opportunities to improve patient care. Key to the successful incorporation of genomic information into the EMR will be the development of laboratory information systems capable of appropriately formatting molecular diagnostic and cytogenetic findings in the EMR. Due to the lack of granular genomics-related content in existing medical vocabularies, the adoption of new standards for describing clinically significant genomic information will be an important step toward recognizing the genome-enabled EMR. Appropriate capture of patient-specific genomic results in the EMR will generate new opportunities to utilize this information in clinical decision support, including automated response to pharmacogenomic-based risks.

Computational Biology↗

Local structure-based sequence profile database for local and global protein structure predictions.

MOTIVATION: A large body of evidence suggests that protein structural information is frequently encoded in local sequences-sequence-structure relationships derived from local structure/sequence analyses could significantly enhance the capacities of protein structure prediction methods. In this paper, the prediction capacity of a database (LSBSP2) that organizes local sequence-structure relationships encoded in local structures with two consecutive secondary structure elements is tested with two computational procedures for protein structure prediction. The goal is twofold: to test the folding hypothesis that local structures are determined by local sequences, and to enhance our capacity in predicting protein structures from their amino acid sequences. RESULTS: The LSBSP2 database contains a large set of sequence profiles derived from exhaustive pair-wise structural alignments for local structures with two consecutive secondary structure elements. One computational procedure makes use of the PSI-BLAST alignment program to predict local structures for testing sequence fragments by matching the testing sequence fragments onto the sequence profiles in the LSBSP2 database. The results show that 54% of the test sequence fragments were predicted with local structures that match closely with their native local structures. The other computational procedure is a filter system that is capable of removing false positives as possible from a set of PSI-BLAST hits. An assessment with a large set of non-redundant protein structures shows that the PSI-BLAST + filter system improves the prediction specificity by up to two-fold over the prediction specificity of the PSI-BLAST program for distantly related protein pairs. Tests with the two computational procedures above demonstrate that local sequence-structure relationships can indeed enhance our capacity in protein structure prediction. The results also indicate that local sequences encoded with strong local structure propensities play an important role in determining the native state folding topology.

Algorithms↗

MaGe: a microbial genome annotation system supported by synteny results.

Magnifying Genomes (MaGe) is a microbial genome annotation system based on a relational database containing information on bacterial genomes, as well as a web interface to achieve genome annotation projects. Our system allows one to initiate the annotation of a genome at the early stage of the finishing phase. MaGe's main features are (i) integration of annotation data from bacterial genomes enhanced by a gene coding re-annotation process using accurate gene models, (ii) integration of results obtained with a wide range of bioinformatics methods, among which exploration of gene context by searching for conserved synteny and reconstruction of metabolic pathways, (iii) an advanced web interface allowing multiple users to refine the automatic assignment of gene product functions. MaGe is also linked to numerous well-known biological databases and systems. Our system has been thoroughly tested during the annotation of complete bacterial genomes (Acinetobacter baylyi ADP1, Pseudoalteromonas haloplanktis, Frankia alni) and is currently used in the context of several new microbial genome annotation projects. In addition, MaGe allows for annotation curation and exploration of already published genomes from various genera (e.g. Yersinia, Bacillus and Neisseria). MaGe can be accessed at http://www.genoscope.cns.fr/agc/mage.

Computational Biology↗

Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievals.

BACKGROUND: Searching for small tandem/disperse repetitive DNA sequences streamlines many biomedical research processes. For instance, whole genomic array analysis in yeast has revealed 22 PHO-regulated genes. The promoter regions of all but one of them contain at least one of the two core Pho4p binding sites, CACGTG and CACGTT. In humans, microsatellites play a role in a number of rare neurodegenerative diseases such as spinocerebellar ataxia type 1 (SCA1). SCA1 is a hereditary neurodegenerative disease caused by an expanded CAG repeat in the coding sequence of the gene. In bacterial pathogens, microsatellites are proposed to regulate expression of some virulence factors. For example, bacteria commonly generate intra-strain diversity through phase variation which is strongly associated with virulence determinants. A recent analysis of the complete sequences of the Helicobacter pylori strains 26695 and J99 has identified 46 putative phase-variable genes among the two genomes through their association with homopolymeric tracts and dinucleotide repeats. Life scientists are increasingly interested in studying the function of small sequences of DNA. However, current search algorithms often generate thousands of matches -- most of which are irrelevant to the researcher. RESULTS: We present our hash function as well as our search algorithm to locate small sequences of DNA within multiple genomes. Our system applies information retrieval algorithms to discover knowledge of cross-species conservation of repeat sequences. We discuss our incorporation of the Gene Ontology (GO) database into these algorithms. We conduct an exhaustive time analysis of our system for various repetitive sequence lengths. For instance, a search for eight bases of sequence within 3.224 GBases on 49 different chromosomes takes 1.147 seconds on average. To illustrate the relevance of the search results, we conduct a search with and without added annotation terms for the yeast Pho4p binding sites, CACGTG and CACGTT. Also, a cross-species search is presented to illustrate how potential hidden correlations in genomic data can be quickly discerned. The findings in one species are used as a catalyst to discover something new in another species. These experiments also demonstrate that our system performs well while searching multiple genomes -- without the main memory constraints present in other systems. CONCLUSION: We present a time-efficient algorithm to locate small segments of DNA and concurrently to search the annotation data accompanying the sequence. Genome-wide searches for short sequences often return hundreds of hits. Our experiments show that subsequently searching the annotation data can refine and focus the results for the user. Our algorithms are also space-efficient in terms of main memory requirements. Source code is available upon request.

Algorithms↗

Data modeling for immunological and clinical data of leukemia and myasthenia patients.

In this study it was investigated whether and to what extent semantic data models and their methods for data modeling are useful for adequate representation and integration of immunological and clinical data. To that end the special research program in leukemia research and immunogenetics (SFB 120) of the University of Tübingen was taken as an example. Based on the semantic data model RM/T we propose the design of a database system, report on its realization, and discuss this approach. Using a semantic data model, the quality of data increased considerably. This means, for instance, that the integration of the molecular-biological knowledge allows a better control of the person-related results. Hence, the decisions based of these data may have greater validity and the treatment on leukemia patients can be improved. Furthermore, the elucidation of immune mechanisms concerning auto-immune diseases could be improved.

Autoimmune Diseases↗

The logical structure of the VIDEOFAR drug data base.

The quality of the analyses that can be carried out by a Drug Prescription Monitoring System depends on the completeness and accuracy of the information on drugs. Data quality depends also on the organization of the data base that must be designed to allow higher retrieval power with regard to the information level theoretically contained in stored data. Considering these requirements we developed inside the VIDEOFAR project, starting from the preceding experiences at the Istituto Superiore di Sanità, an activity of design and realization of a Drug Data Base whose structure, both logical and physical, is described.

Database Management Systems↗

A World Wide Web (WWW) server database engine for an organelle database, MitoDat.

We describe a simple database search engine "dbEngine" which may be used to quickly create a searchable database on a World Wide Web (WWW) server. Data may be prepared from spreadsheet programs (such as Excel, etc.) or from tables exported from relationship database systems. This Common Gateway Interface (CGI-BIN) program is used with a WWW server such as available commercially, or from National Center for Supercomputer Algorithms (NCSA) or CERN. Its capabilities include: (i) searching records by combinations of terms connected with ANDs or ORs; (ii) returning search results as hypertext links to other WWW database servers; (iii) mapping lists of literature reference identifiers to the full references; (iv) creating bidirectional hypertext links between pictures and the database. DbEngine has been used to support the MitoDat database (Mendelian and non-Mendelian inheritance associated with the Mitochondrion) on the WWW.

Computer Communication Networks↗

Gene expression databases and data mining.

The DNA microarray technology has arguably caught the attention of the worldwide life science community and is now systematically supporting major discoveries in many fields of study. The majority of the initial technical challenges of conducting experiments are being resolved, only to be replaced with new informatics hurdles, including statistical analysis, data visualization, interpretation, and storage. Two systems of databases, one containing expression data and one containing annotation data are quickly becoming essential knowledge repositories of the research community. This present paper surveys several databases, which are considered "pillars" of research and important nodes in the network. This paper focuses on a generalized workflow scheme typical for microarray experiments using two examples related to cancer research. The workflow is used to reference appropriate databases and tools for each step in the process of array experimentation. Additionally, benefits and drawbacks of current array databases are addressed, and suggestions are made for their improvement.

Breast Neoplasms↗

ProMode: a database of normal mode analyses on protein molecules with a full-atom model.

MOTIVATION: Although information from protein dynamics simulation is important to understand principles of architecture of a protein structure and its function, simulations such as molecular dynamics and Monte Carlo are very CPU-intensive. Although the ability of normal mode analysis (NMA) is limited because of the need for a harmonic approximation on which NMA is based, NMA is adequate to carry out routine analyses on many proteins to compute aspects of the collective motions essential to protein dynamics and function. Furthermore, it is hoped that realistic animations of the protein dynamics can be observed easily without expensive software and hardware, and that the dynamic properties for various proteins can be compared with each other. RESULTS: ProMode, a database collecting NMA results on protein molecules, was constructed. The NMA calculations are performed with a full-atom model, by using dihedral angles as independent variables, faster and more efficiently than the calculations using Cartesian coordinates. In ProMode, an animation of the normal mode vibration is played with a free plug-in, Chime (MDL Information Systems, Inc.). With the full-atom model, the realistic three-dimensional motions at an atomic level are displayed with Chime. The dynamic domains and their mutual screw motions defined from the NMA results are also displayed. Properties for each normal mode vibration and their time averages, e.g. fluctuations of atom positions, fluctuations of dihedral angles and correlations between the atomic motions, are also presented graphically for characterizing the collective motions in more detail. AVAILABILITY: http://promode.socs.waseda.ac.jp

Computer Graphics↗

A program for storage and retrieval of demographic, sample, clinical laboratory, and restriction endonuclease map data.

A program, written in dBASE, is described that manages demographic, sample, clinical laboratory, and restriction endonuclease map data. The program exports migration distances of DNA fragments resulting from specified endonuclease digests to a commercially available, but modified, curve fitting program (CURVE-FITTER) where the fragment sizes in base pairs are calculated. The calculated values are imported back into dBASE files for report generation or later analysis. The program will produce hard copy reports for single or multiple individuals.

Clinical Laboratory Information Systems↗

BIAS: Bioinformatics Integrated Application Software.

MOTIVATION: We introduce a development platform especially tailored to Bioinformatics research and software development. BIAS (Bioinformatics Integrated Application Software) provides the tools necessary for carrying out integrative Bioinformatics research requiring multiple datasets and analysis tools. It follows an object-relational strategy for providing persistent objects, allows third-party tools to be easily incorporated within the system and supports standards and data-exchange protocols common to Bioinformatics. AVAILABILITY: BIAS is an OpenSource project and is freely available to all interested users at http://www.mcb.mcgill.ca/~bias/. This website also contains a paper containing a more detailed description of BIAS and a sample implementation of a Bayesian network approach for the simultaneous prediction of gene regulation events and of mRNA expression from combinations of gene regulation events. CONTACT: hallett@mcb.mcgill.ca.

Computational Biology↗

Architectural design and tools to support the transparent access to hospital information systems, radiology information systems, and picture archiving and communication systems.

The fragmentation of the electronic patient record among hospital information systems (HIS), radiology information systems (RIS), and picture archiving and communication systems (PACS) makes the viewing of the complete medical patient record inconvenient. The purpose of this report is to describe the system architecture, development tools, and implementation issues related to providing transparent access to HIS, RIS, and PACS information. A client-mediator-server architecture was implemented to facilitate the gathering and visualization of electronic medical records from these independent heterogeneous information systems. The architecture features intelligent data access agents, run-time determination of data access strategies, and an active patient cache. The development and management of the agents were facilitated by data integration CASE (computer-assisted software engineering) tools. HIS, RIS, and PACS data access and translation agents were successfully developed. All pathology, radiology, medical, laboratory, admissions, and radiology reports for a patient are available for review from a single integrated workstation interface. A data caching system provides fast access to active patient data. New network architectures are evolving that support the integration of heterogeneous software subsystems. Commercial tools are available to assist in the integration procedure.

Computer Systems↗

Graph visualization techniques for web clustering engines.

One of the most challenging issues in mining information from the World Wide Web is the design of systems that present the data to the end user by clustering them into meaningful semantic categories. We show that the analysis of the results of a clustering engine can significantly take advantage of enhanced graph drawing and visualization techniques. We propose a graph-based user interface for Web clustering engines that makes it possible for the user to explore and visualize the different semantic categories and their relationships at the desired level of detail.

Algorithms↗

TSGDB: a database system for tumor suppressor genes.

UNLABELLED: A Web-based database system was constructed and implemented that contains 174 tumor suppressor genes. The database homepage was created to accommodate these genes in a pull-down window so that each gene can be viewed individually in a separate Web page. Information displayed on each page includes gene name, aliases, source organism, chromosome location, expression cells/tissues, gene structure, protein size, gene functions and major reference sources. Queries to the database can be conducted through a user-friendly interface, and query results are returned in the HTML format on dynamically generated web pages. AVAILABILITY: The database is available at http://www.cise.ufl.edu/~yy1/HTML-TSGDB/Homepage.html (data files also at http://www.patcar.org/Databases/Tumor_Suppressor_Genes)

Abstracting and Indexing↗

PACS and patient data management systems.

It is important for a PACS to have access to the patient data, as well as to the images themselves, for the purpose of sophisticated image archiving, retrieving, viewing and interpretation. There are many kinds of patient data concerning image examinations (i.e., patient name, ID, age, examination date and time, examined regions, methods, findings on images, diagnoses or diagnostic impressions, etc.). Some of them are acquired from image examination apparatus, some are supplied by diagnostic radiologists, while some need be retrieved from the radiology and hospital information systems. To facilitate this data exchange, a PACS-RIS-HIS coupling is required. The author has constructed at Tokyo University Hospital a small PACS called TRACS, which adopts one of the possible PACS-RIS-HIS coupling configurations.

Computer Systems↗

Constructing biological networks through combined literature mining and microarray analysis: a LMMA approach.

MOTIVATION: Network reconstruction of biological entities is very important for understanding biological processes and the organizational principles of biological systems. This work focuses on integrating both the literatures and microarray gene-expression data, and a combined literature mining and microarray analysis (LMMA) approach is developed to construct gene networks of a specific biological system. RESULTS: In the LMMA approach, a global network is first constructed using the literature-based co-occurrence method. It is then refined using microarray data through a multivariate selection procedure. An application of LMMA to the angiogenesis is presented. Our result shows that the LMMA-based network is more reliable than the co-occurrence-based network in dealing with multiple levels of KEGG gene, KEGG Orthology and pathway. AVAILABILITY: The LMMA program is available upon request.

Abstracting and Indexing↗

A patterned approach for linking knowledge-based systems to external resources.

Knowledge-based systems (KBSs) have been developed and used in industry and government as assistance systems, voting partner systems, and embedded applications. As web-based systems change the face of software implementations, these closed, internal KBSs need to be integrated into multicomponent applications that provide updated and extensible services. Therefore, KBSs must be adapted to an environment in which data and control are exchanged with external processes and resources; complementing other participating systems or using them to refine its own results. This integration can be a daunting task. If improperly done, it can result in an inefficient and unmanageable composite application. One approach to simplifying this task is the use of architectural patterns for integration. These patterns are assembled from functional entities that resolve component interoperability conflicts. In this paper, we describe an architectural pattern called the Knowledge Director pattern, which directs the integration of a closed KBS into a broader application environment.

Artificial Intelligence↗