PubMed · 9357693
Corpus-based identification and refinement of semantic classes.
Abstract
Medical Language Processing (MLP), especially in specific domains, requires fine-grained semantic lexica. We examine whether robust natural language processing tools used on a representative corpus of a domain help in building and refining a semantic categorization. We test this hypothesis with ZELLIG, a corpus analysis tool. The first clusters we obtain are consistent with a model of the domain, as found in the SNOMED nomenclature. They correspond to coarse-grained semantic categories, but isolate as well lexical idiosyncrasies belonging to the clinical sub-language. Moreover, they help categorize additional words.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
A Nazarenko, P Zweigenbaum, J Bouaud, B Habert. 1997. Corpus-based identification and refinement of semantic classes.. https://pubmed.ncbi.nlm.nih.gov/9357693/
Cite the original work for its findings. Save a collection to share your selection of sources.