PubMed Health⌕ Search

PubMed · 16160315

Semantic challenges in database Federation: lessons learned.

Abstract

In this project an integrated analysis of data from disparate surgery and anaesthesiology departmental information systems was carried out. Due to the lack of shared primary keys, a multi-stage "soft" matching method was implemented. Results of the matching steps are described in detail. Inconsistencies were shown to exist for identifying data, semantic definition of documentation content and documented data itself. Minimum requirements for interdisciplinary documentation in autonomous systems should include shared semantic definitions of documentation content as well as robust and regularly validated interfaces for identifying data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thomas Ganslandt, Udo Kunzmann, Katharina Diesch, Péter Pálffy, Hans-Ulrich Prokosch. 2005. Semantic challenges in database Federation: lessons learned.. https://pubmed.ncbi.nlm.nih.gov/16160315/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A high level interface to SCOP and ASTRAL implemented in python.

BACKGROUND: Benchmarking algorithms in structural bioinformatics often involves the construction of datasets of proteins with given sequence and structural properties. The SCOP database is a manually curated structural classification which groups together proteins on the basis of structural similarity. The ASTRAL compendium provides non redundant subsets of SCOP domains on the basis of sequence similarity such that no two domains in a given subset share more than a defined degree of sequence similarity. Taken together these two resources provide a 'ground truth' for assessing structural bioinformatics algorithms. We present a small and easy to use API written in python to enable construction of datasets from these resources. RESULTS: We have designed a set of python modules to provide an abstraction of the SCOP and ASTRAL databases. The modules are designed to work as part of the Biopython distribution. Python users can now manipulate and use the SCOP hierarchy from within python programs, and use ASTRAL to return sequences of domains in SCOP, as well as clustered representations of SCOP from ASTRAL. CONCLUSION: The modules make the analysis and generation of datasets for use in structural genomics easier and more principled.

Database Management Systems↗

The Gene Ontology (GO) project in 2006.

The Gene Ontology (GO) project (http://www.geneontology.org) develops and uses a set of structured, controlled vocabularies for community use in annotating genes, gene products and sequences (also see http://song.sourceforge.net/). The GO Consortium continues to improve to the vocabulary content, reflecting the impact of several novel mechanisms of incorporating community input. A growing number of model organism databases and genome annotation groups contribute annotation sets using GO terms to GO's public repository. Updates to the AmiGO browser have improved access to contributed genome annotations. As the GO project continues to grow, the use of the GO vocabularies is becoming more varied as well as more widespread. The GO project provides an ontological annotation system that enables biologists to infer knowledge from large amounts of data.

Database Management Systems↗

Development and validation of queries using structured query language (SQL) to determine the utilization of comparison imaging in radiology reports stored on PACS.

The purpose of this research was to develop queries that quantify the utilization of comparison imaging in free-text radiology reports. The queries searched for common phrases that indicate whether comparison imaging was utilized, not available, or not mentioned. The queries were iteratively refined and tested on random samples of 100 reports with human review as a reference standard until the precision and recall of the queries did not improve significantly between iterations. Then, query accuracy was assessed on a new random sample of 200 reports. Overall accuracy of the queries was 95.6%. The queries were then applied to a database of 1.8 million reports. Comparisons were made to prior images in 38.69% of the reports (693,955/1,793,754), were unavailable in 18.79% (337,028/1,793,754), and were not mentioned in 42.52% (762,771/1,793,754). The results show that queries of text reports can achieve greater than 95% accuracy in determining the utilization of prior images.

Database Management Systems↗