PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Metadata”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The use of interactive graphical maps for browsing medical/health Internet information resources.

As online information portals accumulate metadata descriptions of Web resources, it becomes necessary to develop effective ways for visualising and navigating the resultant huge metadata repositories as well as the different semantic relationships and attributes of described Web resources. Graphical maps provide a good method to visualise, understand and navigate a world that is too large and complex to be seen directly like the Web. Several examples of maps designed as a navigational aid for Web resources are presented in this review with an emphasis on maps of medical and health-related resources. The latter include HealthCyberMap maps http://healthcybermap.semanticweb.org/, which can be classified as conceptual information space maps, and the very abstract and geometric Visual Net maps of PubMed http://pubmed.antarcti.ca/start. Information resources can be also organised and navigated based on their geographic attributes. Some of the maps presented in this review use a Kohonen Self-Organising Map algorithm, and only HealthCyberMap uses a Geographic Information System to classify Web resource data and render the maps. Maps based on familiar metaphors taken from users' everyday life are much easier to understand. Associative and pictorial map icons that enable instant recognition and comprehension are preferred to geometric ones and are key to successful maps for browsing medical/health Internet information resources.

Journal Article↗

PrimeAnswers: A practical interface for answering primary care questions.

This paper describes an institutional approach taken to build a primary care reference portal. The objective for the site is to make access to and use of clinical reference faster and easier and to facilitate the use of evidence-based answers in daily practice. Reference objects were selected and metadata applied to a core set of sources. Metadata were used to search, sort, and filter results and to define deep-linked queries and structure the interface. User feedback resulted in an expansion in the scope of reference objects to meet the broad spectrum of information needs, including patient handouts and interactive risk management tools. RESULTS of a user satisfaction survey suggest that a simple interface to customized content makes it faster and easier for primary care clinicians to find information during the clinic day and to improve care to their patients. The PrimeAnswers portal is a first step in creating a fast search of a customized set of reference objects to match a clinician's patient care questions in the clinic. The next step is developing methods to solve the problem of matching a clinician's question to a specific answer through precise retrieval from reference sources; however, lack of internal structure and Web service standards in most clinical reference sources is an unresolved problem.

Consumer Behavior↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

A new module for on-line manipulation and display of molecular information in the brain architecture management system.

A new "Molecules" module of the Brain Architecture Management System (BAMS; http://brancusi.usc.edu/bkms) is described. With this module, BAMS becomes the first online knowledge management system to handle central nervous system (CNS) region and celltype chemoarchitectonic data in the context of axonal connections between regions and cell types, in multiple species. The "Molecules" module implements a general knowledge representation schema for data and metadata collated from published and unpublished material, and allows insertion of complex reports about the presence of molecules collated from the literature. For different CNS neural regions and cell types, the module's database structure includes representation of molecule expression revealed by various techniques including in situ hybridization and immunohistochemistry, molecule coexpression and time-dependent level changes, and physiological state of subjects. The metadata representation allows online comparison and evaluation of inserted experiments, and "Molecules"structure allows rapid development of data transfer protocols enabling neuroinformatics visualization tools to display gene expression patterns residing in BAMS, in terms of levels of expressed molecules and in situ hybridization data. The module's web interface allows users to construct lists of CNS regions containing a molecule (depending on physiological state), retrieve further details about inserted records, compare time-dependent data within and across experiments, reconstruct gene expression patterns, and construct complex reports from individual experiments.

Animals↗

Generic design of Web-based clinical databases.

BACKGROUND: The complexity and the rapid evolution and expansion of the domain of clinical information make development and maintenance of clinical databases difficult. Whenever new data types are introduced or existing types are modified in a conventional relational database system, the physical design of the database must be changed accordingly. For this reason, it is desirable that a clinical database be flexible and allow for modifications and for addition of new types of data without having to change the physical database schema. The ideal clinical database would therefore implement a highly-detailed logical database schema in a completely-generic physical schema that stores the wide variety of clinical data in a small and constant number of tables. OBJECTIVE: The objective was to review the medical literature regarding generic design of clinical databases. METHODS: A search strategy was devised for PubMed and Google to get the best match of peer-reviewed articles and free Web resources on the subject. RESULTS: Eight peer reviewed articles and a Web tutorial were found. All the resources described the so-called Entity-Attribute-Value (EAV) design as a means of simplifying the physical layout of data tables in a clinical database. In Entity-Attribute-Value design all data can be stored in a single generic table with conceptually 3 columns: 1 for entity (eg, patient identification), 1 for attribute (eg, name), and 1 for value (eg, "Jens Hansen"). To add more descriptive fields to the entity class, all that is necessary is to add attribute values to be stored in the attribute field. The main advantages of the Entity-Attribute-Value design are flexibility and effective entity-centered data retrieval. The main disadvantages are complicated front-end programming needed to display data in a conventional layout that the user understands and less-efficient attribute-centered queries. The Internet offers unique opportunities for database deployment, eliminating problems of user-interface deployment. Furthermore, Web forms may be generated in a completely-generic fashion during run time from metadata describing the semantic structure of clinical information stored in the database. CONCLUSIONS: The Entity-Attribute-Value model is useful for generic design of clinical databases. Depending on the specific requirements of the application, more or less complex metadata models may be applied.

Databases, Factual↗

Organizing medical networked information (OMNI).

The Internet has become a major source of biomedical information over the last 5 years. Several projects have recently been established to help users find respectable information sources quickly. OMNI (Organizing Medical Networked Information) is one such filtering and indexing project. OMNI has focused on the quality of information and the application to Internet resources of standard tools for organizing information such as the National Library of Medicine's Medical Subject Headings and the Dublin Core metadata format. Now two years old, the OMNI project fulfils a valuable role for the UK biomedical community, through its gateway service (http:@omni.ac.uk), its printed resource guides and its training workshop programme. OMNI is also a focus for biomedical metadata activities in the UK. The gateway continues to grow in size and further work on information quality issues and integration is planned.

Abstracting and Indexing↗

Performance Profiles of Short DNA Barcode Segments for Family Level Detection of Asteraceae Within Asterales.

Short DNA barcodes may facilitate sequence recovery from degraded material, but their ability to retain target-family identity while excluding related taxa varies among genomic regions. We computationally evaluated 16 nuclear, plastid, and mitochondrial marker regions from 11 Asterales families using 279,956 NCBI locus-record matches and an accession-disjoint discovery/test design. Thirty-one candidate segments of 50-200 bp (mean, 98.55 bp) were screened in discovery data and evaluated for within-Asteraceae sequence recall, differentiation from non-Asteraceae Asterales, in silico primer behavior, phylogenetic placement, and exploratory matching across 808 metadata-defined metagenomic samples. Conserved regions such as matR and rbcL showed high within-Asteraceae identity, whereas ITS1, ITS, and trnH-psbA showed larger differences from related-family backgrounds; ITS2 and ycf1 showed intermediate profiles. Candidate segments were placed within or immediately adjacent to Asteraceae reference branches in segment-specific maximum-likelihood analyses, although support and topology varied among regions. Metadata-defined target-containing groups had higher mean query coverage and identity than background groups; because target presence was not independently verified and no classifier was fitted, these comparisons were descriptive and did not estimate diagnostic accuracy. Definitionally linked sequence statistics were interpreted as structural associations rather than evidence of causal evolutionary mechanisms. These results provide a family-level computational comparison of candidate short segments for Asteraceae detection within Asterales. Species identification, operational marker combinations, threshold robustness, and laboratory performance require validation using taxonomically dense, voucher-linked, and experimentally characterized datasets.

Asteraceae↗

Labeling and filtering of medical information on the Internet.

Internet information undergoes no quality controls and virtually anybody can publish anything. Because of this, it is difficult for searchers to take information retrieved from the Internet at face value. A related problem is the uncontrolled promotion of medical products on the Internet. A further problem of today's Internet is that authors use no uniform keywords and other descriptive labels, which deteriorates the quality of search results. A solution for all these problems could be widespread use of descriptive and evaluative metainformation associated with medical Internet information. Our concept is based on a recently established infrastructure for assigning metadata to Internet information, the so-called PICS Standard (Platform for Internet Content Selection). We prototyped a PICS-based rating vocabulary for medical information (med-PICS), containing descriptive and evaluative categories, to be used by the webauthor and third-party label services (such as medical associations), respectively. We propose an international effort to assign metadata to medical Internet information.

Humans↗

Neuronal database integration: the Senselab EAV data model.

We discuss an approach towards integrating heterogeneous nervous system data using an augmented Entity-Attribute-Value (EAV) schema design. This approach, widely used in implementing electronic patient record systems (EPRSs), allows the physical schema of the database to be relatively immune to changes in domain knowledge. This is because new kinds of facts are added as data (or as metadata) rather than hard-coded as the names of newly created tables or columns. Because the domain knowledge is stored as metadata, a framework developed in one scientific domain can be ported to another with only modest revision. We describe our progress in creating a code framework that handles browsing and hyperlinking of the different kinds of data.

Databases, Factual↗

Representation by standard terminologies of health status concepts contained in two health status assessment instruments used in rheumatic disease management.

Health and functional status data have been shown to have clinical utility in predicting outcome. Various metadata registries in the form of patient self-administered health assessment questionnaires have been incorporated into routine clinical care and clinical research of patients with rheumatic disease. Examples of such health assessment instruments are the Clinical Health Assessment Questionnaire (CLINHAQ) and the Modified Health Assessment Questionnaire (MHAQ). These instruments contain concepts that are an integral part of the health and functional status domain. Using an automated indexing tool we examined the clinical content coverage by SNOMED RT and the Unified Medical Language System (UMLS) Metathesaurus for health and functional status concepts identified in the MHAQ and CLINHAQ. Significant differences existed between the overall representational ability of SNOMED and UMLS for concepts identified in the MHAQ (49%, vs. 77% respectively, p < .005) and for concepts identified in the CLINHAQ (30% vs. 64% respectively p < .005). Representational capability by SNOMED-RT and UMLS for concepts in a given health assessment instrument was carried across four semantic classes of "attitudes", "symptoms", "activities", and "social attributes". The conceptual content coverage of health status assessment concepts contained in the MHAQ and CLINHAQ by SNOMED-RT and UMLS was incomplete but better for UMLS with its panoply of vocabulary sources. This observed overall improved representation by UMLS appeared to be due to better representation of concepts in "activities" and "social attributes" semantic classes. Representation of health or functional status concepts in a computerized medical record should be founded on a universally agreed concept model of that domain. Established functional and health status metadata registries can serve as important sources for concepts and candidate classes within that domain.

Health Status↗

IML: An image markup language.

Image Markup Language is an extensible markup language (XML) schema used to describe both image metadata and annotations. It describes both data pertaining to an entire image, and data that are tied to specific regions or features of the image. Developed for a specific domain in Medical Education, this pa-per describes extensions to take advantage of the Dublin Core metadata standard, and of an XML schema for vector graphics representation. We have developed a prototype system of open source tools implementing an authoring system, a client system, and an image annotation database which can be queried though the Web.

Diagnostic Imaging↗

MeSHmap: a text mining tool for MEDLINE.

Our research goal is to explore text mining from the metadata included in MEDLINE documents. We present MeSHmap our prototype text mining system that exploits the MeSH indexing accompanying MEDLINE records. MeSHmap supports searches via PubMed followed by user driven exploration of the MeSH terms and subheadings in the retrieved set. The potential of the system goes beyond text retrieval. It may also be used to compare entities of the same type such as pairs of drugs or pairs of procedures etc. In addition there is the potential to generate maps of entities (drugs or diseases etc.) such that the strength of the link between two entities in the map represents their similarity as expressed in the MeSH metadata of the MEDLINE documents. Higher level operators have been proposed to support these comparison and mapping functions. This paper motivates and describes MeSHmap. Future work will include user evaluations of the system.

Abstracting and Indexing↗

Managing troubled data: coastal data partnerships smooth data integration.

Understanding the ecology, condition, and changes of coastal areas requires data from many sources. Broad-scale and long-term ecological questions, such as global climate change, biodiversity, and cumulative impacts of human activities, must be addressed with databases that integrate data from several different research and monitoring programs. Various barriers, including widely differing data formats, codes, directories, systems, and metadata used by individual programs, make such integration troublesome. Coastal data partnerships, by helping overcome technical, social, and organizational barriers, can lead to a better understanding of environmental issues, and may enable better management decisions. Characteristics of successful data partnerships include a common need for shared data, strong collaborative leadership, committed partners willing to invest in the partnership, and clear agreements on data standards and data policy. Emerging data and metadata standards that become widely accepted are crucial. New information technology is making it easier to exchange and integrate data. Data partnerships allow us to create broader databases than would be possible for any one organization to create by itself.

Conservation of Natural Resources↗

Quality assurance of medical ontologies.

OBJECTIVE: To review the literature concerning the quality assurance of medical ontologies. METHODS: scholar.google.com was searched using the search strings (+ontology +"quality assurance") and (+ontology +"evaluation/evaluating"). Relevant publications were selected by manual review. Other work already familiar to the author, or suggested by other researchers contacted by the author, were included. The papers were analysed for common themes. RESULTS: Four broad properties of an ontology were identified that may be quality-assured: philosophical validity, compliance with meta-ontological commitments, 'content correctness', and fitness for purpose. Each published methodology addressed only a subset of these properties. 'Content' may be divided into domain knowledge content, and metadata describing either the provenance of domain knowledge content, or relationships between it and lexical information (e.g. for display and retrieval). 'Correctness' (whether of domain knowledge content or metadata) may also be further subdivided into truth, completeness, parsimony and internal consistency. CONCLUSIONS: Understanding of how to assure the quality of ontologies, or evaluate their fitness for specific purposes, is improving but remains poor. A combination of methodologies is required, but tools to support a comprehensive quality assurance programme remain lacking. Perfect quality of an ontology is not provable and may not be desirable: an ontology compliant with all current philosophical theories, following necessary ontological commitments, and with entirely 'correct' content, may be too complex to be directly usable or useful. The extent to which an ontology's fitness for purpose is predicted or influenced by its other properties remains to be determined. Field studies of ontologies in use, including interrater effects, are required.

Medical Informatics↗

Automating identification of adverse events related to abnormal lab results using standard vocabularies.

Laboratory data need to be imported automatically into central Clinical Study Data Management Systems (CSDMSs), and abnormal laboratory data need to be linked to clinically related adverse events. This import of laboratory data can be automated through mapping to standard vocabularies with HL7/LOINC mapping to the metadata within a CSDMS. We have designed a system that uses the UMLS metathesaurus as a common source to map or link abnormal laboratory values to adverse event CTCAE coded terms and grades in the metadata of TrialDB, a generic CSDMS.

Clinical Laboratory Information Systems↗

The Common Data Elements for cancer research: remarks on functions and structure.

OBJECTIVES: The National Cancer Institute (NCI) has developed the Common Data Elements (CDE) to serve as a controlled vocabulary of data descriptors for cancer research, to facilitate data interchange and inter-operability between cancer research centers. We evaluated CDE's structure to see whether it could represent the elements necessary to support its intended purpose, and whether it could prevent errors and inconsistencies from being accidentally introduced. We also performed automated checks for certain types of content errors that provided a rough measure of curation quality. METHODS: Evaluation was performed on CDE content downloaded via the NCI's CDE Browser, and transformed into relational database form. Evaluation was performed under three categories: 1) compatibility with the ISO/IEC 11179 metadata model, on which CDE structure is based, 2) features necessary for controlled vocabulary support, and 3) support for a stated NCI goal, set up of data collection forms for cancer research. RESULTS: Various limitations were identified both with respect to content (inconsistency, insufficient definition of elements, redundancy) as well as structure--particularly the need for term and relationship support, as well as the need for metadata supporting the explicit representation of electronic forms that utilize sets of common data elements. CONCLUSIONS: While there are numerous positive aspects to the CDE effort, there is considerable opportunity for improvement. Our recommendations include review of existing content by diverse experts in the cancer community; integration with the NCI thesaurus to take advantage of the latter's links to nationally used controlled vocabularies, and various schema enhancements required for electronic form support.

Biomedical Research↗

SQLGEN: a framework for rapid client-server database application development.

SQLGEN is a framework for rapid client-server relational database application development. It relies on an active data dictionary on the client machine that stores metadata on one or more database servers to which the client may be connected. The dictionary generates dynamic Structured Query Language (SQL) to perform common database operations; it also stores information about the access rights of the user at log-in time, which is used to partially self-configure the behavior of the client to disable inappropriate user actions. SQLGEN uses a microcomputer database as the client to store metadata in relational form, to transiently capture server data in tables, and to allow rapid application prototyping followed by porting to client-server mode with modest effort. SQLGEN is currently used in several production biomedical databases.

Computer Communication Networks↗

Automatic query mapping among genomic databases: a pilot exploration.

As databases in the human genome project proliferate, it is important for users of one genomic database to identify similar or inconsistent data in other autonomously developed genomic databases. To do so, the user needs to issue the same query across multiple databases. We describe an approach that allows a query issued against one database to be automatically mapped to an equivalent query against another structurally different database. Our approach features two components: 1) a database designed to capture knowledge (metadata) that describes the correspondences among individual database components and 2) a module that utilizes the metadata to perform query mappings. As a demonstration, we apply our query mapping approach to two chromosome map databases (DB/12 and GDB).

Algorithms↗