PubMed Health⌕ Search

PubMed · 12463837

Using the metaschema to audit UMLS classification errors.

Abstract

The Unified Medical Language System integrates about 800,000 concepts from 99 biomedical terminologies. Each concept is assigned to at least one semantic type of the Semantic Network. During the integration, it is unavoidable that some classification errors and inconsistencies will be introduced. In this paper, we present an auditing technique to find such errors and inconsistencies. Our technique is based on an expert reviewing the pure intersections of meta-semantic types of the metaschema, a compact abstract view of the Semantic Network. Results regarding the pure intersections are reported. The analysis results for pure intersections with 1 to 6 concepts are presented. Various kinds of errors are identified.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Huanying Helen Gu, Hua Min, Yi Peng, Li Zhang, Yehoshua Perl. 2002. Using the metaschema to audit UMLS classification errors.. https://pubmed.ncbi.nlm.nih.gov/12463837/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation.

MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).

Semantics↗

Enriching the structure of the UMLS semantic network.

The Unified Medical Language System's (UMLS's) Semantic Network (SN)---consisting of a network of semantic types---has a two-tree structure, where each semantic type has at most one parent semantic type. This arrangement is restrictive because some semantic types are, by their definition, specializations of several parents. As a proposed enhancement to the SN, its semantic types have previously been partitioned into groups, each of which contains semantic types of some specific area. However, some groups of this proposed partition contain forest (i.e., multiple-tree) structures or even isolated semantic types. Both situations imply a disconnected internal structure. Connectivity is actually one way to assess the proposed "semantic validity" principle for partitions. It is a desired, although not required, property. In this paper, we introduce a methodology for identifying "missing" IS-A links and adding them to the SN. This process transforms the SN into a Directed Acyclic Graph (DAG) structure, with semantic types permitted to have multiple parents. A result of our methodology is the transformation of the proposed SN partition into groups satisfying the connectivity property.

Semantics↗