The use of Dublin Core metadata in a structured health resource guide on the internet.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Access to current clinical information involves searches of bibliographic databases, such as MEDLINE, and subsequent evaluation of retrieval results for relevance to a specific clinical situation and quality of the reported research. We establish the amount of information that needs to be provided by an information retrieval system to assist healthcare practitioners in identifying clinically relevant information and evaluating its potential strength of evidence. We find 92% of titles informative enough for a practitioner to correctly classify publications as clinical, but not sufficient for classification of research quality. We suggest automatic organization of retrieval results into strength of evidence categories to supplement title-based judgments and provide quick access to the abstracts of the most promising articles. We find information in the abstracts sufficient to identify articles potentially immediately useful for clinical decision support. These findings are important to the design of information retrieval systems supporting small, low-bandwidth handheld computers.
Explore the source record for details and available documents.
OBJECTIVES: HealthCyberMap (HCM-http://healthcybermap.semanticweb.org) is a web-based service for healthcare professionals and librarians, patients and the public in general that aims at mapping parts of the health information resources in cyberspace in novel ways to improve their retrieval and navigation. METHODS AND SERVICE DESCRIPTION: HCM adopts a clinical metadata framework built upon a clinical coding ontology for the semantic indexing, classification and browsing of Internet health information resources. A resource metadata base holds information about selected resources. HCM then uses GIS (Geographic Information Systems) spatialization methods to generate interactive navigational cybermaps from the metadata base. These visual cybermaps are based on familiar medical metaphors. CONCLUSIONS: HCM cybermaps can be considered as semantically spatialized, ontology-based browsing views of the underlying resource metadata base. Using a clinical coding scheme as a metric for spatialization ('semantic distance') is unique to HCM and is very much suited for the semantic categorization and navigation of Internet health information resources. Clinical codes ensure reliable and unambiguous topical indexing of these resources. HCM also introduces a useful form of cyberspatial analysis for the detection of topical coverage gaps in the resource metadata base using choropleth (shaded) maps of human body systems.
A wide variety of data sets produced by individual investigators are now synthesized to address ecological questions that span a range of spatial and temporal scales. It is important to facilitate such syntheses so that "consumers" of data sets can be confident that both input data sets and synthetic products are reliable. Necessary documentation to ensure the reliability and validation of data sets includes both familiar descriptive metadata and formal documentation of the scientific processes used (i.e., process metadata) to produce usable data sets from collections of raw data. Such documentation is complex and difficult to construct, so it is important to help "producers" create reliable data sets and to facilitate their creation of required metadata. We describe a formal representation, an "analytic web," that aids both producers and consumers of data sets by providing complete and precise definitions of scientific processes used to process raw and derived data sets. The formalisms used to define analytic webs are adaptations of those used in software engineering, and they provide a novel and effective support system for both the synthesis and the validation of ecological data sets. We illustrate the utility of an analytic web as an aid to producing synthetic data sets through a worked example: the synthesis of long-term measurements of whole-ecosystem carbon exchange. Analytic webs are also useful validation aids for consumers because they support the concurrent construction of a complete, Internet-accessible audit trail of the analytic processes used in the synthesis of the data sets. Finally we describe our early efforts to evaluate these ideas through the use of a prototype software tool, SciWalker. We indicate how this tool has been used to create analytic webs tailored to specific data-set synthesis and validation activities, and suggest extensions to it that will support additional forms of validation. The process metadata created by SciWalker is readily adapted for inclusion in Ecological Metadata Language (EML) files.
SemanticEye, an ontology with associated tools, improves the classification and open accessibility of chemical information in electronic publishing. In a manner analogous to digital music management, RDF metadata encoded as Adobe XMP can be extracted from a variety of document formats, such as PDF, and managed in an RDF repository called Sesame. Users upload electronic documents containing XMP to a central server by "dropping" them into WebDAV folders. The documents can then be navigated in a Web browser via their metadata, and multiple documents containing identical metadata can then be aggregated. SemanticEye does not actually store any documents. By including unique identifiers within the XMP, such as the DOI, associated documents can be retrieved from the Web with the help of resolving agents. The power of this metadata driven approach is illustrated by including, within the XMP, InChI identifiers for molecular structures and finding relationships between articles based on their InChIs. SemanticEye will become increasingly more comprehensive as usage becomes more widespread. Furthermore, following the Semantic Web architecture enables the reuse of open software tools, provides a "semantically intuitive" alternative to search engines, and fosters a greater sense of trust in Web-based scientific information.
MOTIVATION: A Robot Scientist is a physically implemented robotic system that can automatically carry out cycles of scientific experimentation. We are commissioning a new Robot Scientist designed to investigate gene function in S. cerevisiae. This Robot Scientist will be capable of initiating >1,000 experiments, and making >200,000 observations a day. Robot Scientists provide a unique test bed for the development of methodologies for the curation and annotation of scientific experiments: because the experiments are conceived and executed automatically by computer, it is possible to completely capture and digitally curate all aspects of the scientific process. This new ability brings with it significant technical challenges. To meet these we apply an ontology driven approach to the representation of all the Robot Scientist's data and metadata. RESULTS: We demonstrate the utility of developing an ontology for our new Robot Scientist. This ontology is based on a general ontology of experiments. The ontology aids the curation and annotating of the experimental data and metadata, and the equipment metadata, and supports the design of database systems to hold the data and metadata. AVAILABILITY: EXPO in XML and OWL formats is at: http://sourceforge.net/projects/expo/. All materials about the Robot Scientist project are available at: http://www.aber.ac.uk/compsci/Research/bio/robotsci/.
HealthCyberMap (http://healthcybermap.semanticweb.org) is a Semantic Web project that aims at mapping selected parts of health information resources in cyberspace in novel semantic ways to improve their retrieval and navigation. This paper describes HealthCyberMap semantic subject search engine methodology and early prototype which attempt to overcome the limitations of conventional free text search engines. Explicit concepts in resource metadata map onto a brokering domain ontology (a clinical terminology or classification) allowing a Semantic Web search engine to infer implicit meanings (synonyms and semantic relationships) not directly mentioned in either the resource or its metadata. Similarly, user queries would map to the same ontology allowing the search engine to infer the implicit semantics of user queries and use them to optimise retrieval. Related issues of metadata, clinical terminologies and automatic vs. manual indexing of medical Web resources are also discussed, together with future methodological directions, which include the use of a true terminology server as an intelligent broker between user queries and HealthCyberMap pool of resource metadata. A comparative evaluation of the new engine based on relevance metrics is also proposed.
Larger cohorts improve the power of tumor gene expression analysis, but the signal is muddied if datasets are processed using different methods or have inaccurate metadata. Here we present five compendia containing consistently processed gene expression data derived from 16,446 diverse RNA sequencing datasets. To create the compendia, we obtained access to RNA sequence data from repositories containing public data as well as clinical partners with access to non-published data. We then assessed the quality, quantified gene expression, harmonized clinical metadata, and released the expression values and metadata without access restrictions. These datasets have been used for diverse projects ranging from identifying similarities between tumor types to assessing how well cell lines recapitulate tumors. They have also been used for n-of-1 analysis to identify genes with unusual expression patterns in a single sample and to infer molecular diagnosis. The comparison to new data is enabled by our dockerized, freely available pipeline. The compendia have been cited in at least 20 publications.
The web has become such an extensive health information repository in the world that it is increasingly difficult to search for relevant medical information. Most medical information available on the web is not peer reviewed, and is retrieved imprecisely by current web search mechanisms (i.e. based on keywords). This paper presents the MedISeek metadata model that allows one to describe medical visual information (i.e. medical images) of different modalities, including their properties, components, relationships and authorship. The model uses the web architecture and supports the international classification of diseases and related health problems (i.e. ICD-10). An RDF schema (Resource Description Framework (RDF), http://www.w3.org/RDF/.) derived from this metadata model is integrated to each medical image, and specifies the semantics of each property in the image. Thus, relevant information can be extracted directly from the images, and data integrity is better preserved in the web. A prototype, presented here, has been built to validate the metadata model, and the mechanism for medical visual information exchange on the web. Our preliminary experimental results indicate that authorized users of our system have been able to describe, store and retrieve medical images and their associated diagnostic information.
Reliable evolutionary inference increasingly depends on public genome resources, and the effects of uneven assembly quality, incomplete metadata, and biased taxonomic sampling remain poorly quantified. Using the species-rich fungal lineage Nectriaceae as a model system, we analysed 1530 genome sequence assemblies to assess metadata completeness, sampling representation, and genome quality. One-third of the assemblies lacked essential metadata, sequencing was heavily skewed toward a few agriculturally important lineages, and sampling of many genera was limited or nonexistent. BUSCO and QUAST metrics revealed substantial heterogeneity in assembly quality, with widespread fragmentation and numerous assemblies falling outside expected quality thresholds. From 763 single-copy orthologs identified in 576 higher-quality genomes, we reconstructed a phylogenomic backbone and quantified gene- and site-level concordance across the tree. Although major clades were broadly recovered, extensive gene-tree discordance and a polyphyletic Fusarium nisikadoi species complex revealed unresolved boundaries and conflict among loci. These results show how data quality, incomplete sampling, and discordant genomic histories can constrain phylogenomic resolution, and provide a general framework for improving comparative genomic resources and large-scale evolutionary inference.
One goal of eScience is to enable the end-to-end publication of experiments and results. In the Combechem project we have developed an innovative human-centred system which captures the process of a chemistry experiment from plan to execution. The system comprises an electronic lab book replacement, which has been successfully trialled in a synthetic organic chemistry laboratory, and a flexible back-end storage system. Working closely with the users, we found that a light touch and a high degree of flexibility was required in the user interface. In this paper, we concentrate on the representation and storage of human-scale experiment metadata, introducing an ontology to describe the record of an experiment, and a storage system for the data from our lab book software. Just as the interfaces need to be flexible to cope with whatever a chemist wishes to record, so the back end solutions need to be similarly flexible to store any metadata that may be created. The storage system is based on Semantic Web technologies, such as RDF, and Web Services. It gives a much higher degree of flexibility to the type of metadata it can store, compared to the use of rigid relational databases.
Using the video metadata descriptors and data model defined in the accompanying paper (Shotton, D. M. et al. (2002) A metadata classification schema for semantic content analysis of videos. J. Microsc. 205, 33-42), we discuss how analysis of the content of scientific videos, and subsequent query by content of the resulting semantic metadata, can be enhanced by the use of an object-relational database. We illustrate this by describing VANQUIS, a Web-based prototype video analysis and query interface system for the interactive spatio-temporal analysis and subsequent query by content of videos. Using VANQUIS to generate standard SQL (structured query language) statements that address complex data types stored in an object-relational database, relationships between characters and events contained within and between videos can be identified, and the appropriate video segments containing these characters and events can be retrieved for viewing. We give examples of analysis and query implementation by using VANQUIS to analyse a biological microscopy video, and discuss the wider potential of this methodology for the analysis and query by content of videos containing more general subject matter.
MOTIVATION: Sites with substantive bioinformatics operations are challenged to build data processing and delivery infrastructure that provides reliable access and enables data integration. Locally generated data must be processed and stored such that relationships to external data sources can be presented. Consistency and comparability across data sets requires annotation with controlled vocabularies and, further, metadata standards for data representation. Programmatic access to the processed data should be supported to ensure the maximum possible value is extracted. Confronted with these challenges at the National Cancer Institute Center for Bioinformatics, we decided to develop a robust infrastructure for data management and integration that supports advanced biomedical applications. RESULTS: We have developed an interconnected set of software and services called caCORE. Enterprise Vocabulary Services (EVS) provide controlled vocabulary, dictionary and thesaurus services. The Cancer Data Standards Repository (caDSR) provides a metadata registry for common data elements. Cancer Bioinformatics Infrastructure Objects (caBIO) implements an object-oriented model of the biomedical domain and provides Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. caCORE has been used to develop scientific applications that bring together data from distinct genomic and clinical science sources. AVAILABILITY: caCORE downloads and web interfaces can be accessed from links on the caCORE web site (http://ncicb.nci.nih.gov/core). caBIO software is distributed under an open source license that permits unrestricted academic and commercial use. Vocabulary and metadata content in the EVS and caDSR, respectively, is similarly unrestricted, and is available through web applications and FTP downloads. SUPPLEMENTARY INFORMATION: http://ncicb.nci.nih.gov/core/publications contains links to the caBIO 1.0 class diagram and the caCORE 1.0 Technical Guide, which provide detailed information on the present caCORE architecture, data sources and APIs. Updated information appears on a regular basis on the caCORE web site (http://ncicb.nci.nih.gov/core).
High-throughout genomic data provide an opportunity for identifying pathways and genes that are related to various clinical phenotypes. Besides these genomic data, another valuable source of data is the biological knowledge about genes and pathways that might be related to the phenotypes of many complex diseases. Databases of such knowledge are often called the metadata. In microarray data analysis, such metadata are currently explored in post hoc ways by gene set enrichment analysis but have hardly been utilized in the modeling step. We propose to develop and evaluate a pathway-based gradient descent boosting procedure for nonparametric pathways-based regression (NPR) analysis to efficiently integrate genomic data and metadata. Such NPR models consider multiple pathways simultaneously and allow complex interactions among genes within the pathways and can be applied to identify pathways and genes that are related to variations of the phenotypes. These methods also provide an alternative to mediating the problem of a large number of potential interactions by limiting analysis to biologically plausible interactions between genes in related pathways. Our simulation studies indicate that the proposed boosting procedure can indeed identify relevant pathways. Application to a gene expression data set on breast cancer distant metastasis identified that Wnt, apoptosis, and cell cycle-regulated pathways are more likely related to the risk of distant metastasis among lymph-node-negative breast cancer patients. Results from analysis of other two breast cancer gene expression data sets indicate that the pathways of Metalloendopeptidases (MMPs) and MMP inhibitors, as well as cell proliferation, cell growth, and maintenance are important to breast cancer relapse and survival. We also observed that by incorporating the pathway information, we achieved better prediction for cancer recurrence.
We have implemented a pair of database projects, one serving cortical electrophysiology and the other invertebrate neurones and recordings. The design for each combines aspects of two proven schemes for information interchange. The journal article metaphor determined the type, scope, organization and quantity of data to comprise each submission. Sequence databases encouraged intuitive tools for data viewing, capture, and direct submission by authors. Neurophysiology required transcending these models with new datatypes. Time-series, histogram and bivariate datatypes, including illustration-like wrappers, were selected by their utility to the community of investigators. As interpretation of neurophysiological recordings depends on context supplied by metadata attributes, searches are via visual interfaces to sets of controlled-vocabulary metadata trees. Neurones, for example, can be specified by metadata describing functional and anatomical characteristics. Permanence is advanced by data model and data formats largely independent of contemporary technology or implementation, including Java and the XML standard. All user tools, including dynamic data viewers that serve as a virtual oscilloscope, are Java-based, free, multiplatform, and distributed by our application servers to any contemporary networked computer. Copyright is retained by submitters; viewer displays are dynamic and do not violate copyright of related journal figures. Panels of neurophysiologists view and test schemas and tools, enhancing community support.
The dominant paradigm for searching and browsing large data stores is text-based: presenting a scrollable list of search results in response to textual search term input. While this works well for the Web, there is opportunity for improvement in the domain of personal information stores, which tend to have more heterogeneous data and richer metadata. In this paper, we introduce FacetMap, an interactive, query-driven visualization, generalizable to a wide range of metadata-rich data stores. FacetMap uses a visual metaphor for both input (selection of metadata facets as filters) and output. Results of a user study provide insight into tradeoffs between FacetMap's graphical approach and the traditional text-oriented approach.
Query Integrator System (QIS) is a database mediator framework intended to address robust data integration from continuously changing heterogeneous data sources in the biosciences. Currently in the advanced prototype stage, it is being used on a production basis to integrate data from neuroscience databases developed for the SenseLab project at Yale University with external neuroscience and genomics databases. The QIS framework uses standard technologies and is intended to be deployable by administrators with a moderate level of technological expertise: It comes with various tools, such as interfaces for the design of distributed queries. The QIS architecture is based on a set of distributed network-based servers, data source servers, integration servers, and ontology servers, that exchange metadata as well as mappings of both metadata and data elements to elements in an ontology. Metadata version difference determination coupled with decomposition of stored queries is used as the basis for partial query recovery when the schema of data sources alters.