PubMed Health⌕ Search

Biomedical subjects

P Zweigenbaum

Publications and source records attributed to P Zweigenbaum.

16 recordsLinked to original sources

Looking for French-English translations in comparable medical corpora.

Cross-language retrieval of medical information needs to translate input queries into target language queries. It must be prepared to cope with 'new' words not yet listed in a multilingual lexicon. We address the issue of finding translational equivalents of such 'unknown' words from French to English in the medical domain. We rely on non-parallel, comparable corpora and an initial bilingual medical lexicon. We compare the distributional contexts of source and target words, testing several weighting factors and similarity measures. For the best combination (the Jaccard similarity measure with or without weighting), the correct translation is found in the top 10 candidates for more than 60% of the test words. This shows the potential of this technique to help extending bilingual medical lexicons.

Information Storage and Retrieval↗

An assessment of the visibility of MeSH-indexed medical web catalogs through search engines.

Manually indexed Internet health catalogs such as CliniWeb or CISMeF provide resources for retrieving high-quality health information. Users of these quality-controlled subject gateways are most often referred to them by general search engines such as Google, AltaVista, etc. This raises several questions, among which the following: what is the relative visibility of medical Internet catalogs through search engines? This study addresses this issue by measuring and comparing the visibility of six major, MeSH-indexed health catalogs through four different search engines (AltaVista, Google, Lycos, Northern Light) in two languages (English and French). Over half a million queries were sent to the search engines; for most of these search engines, according to our measures at the time the queries were sent, the most visible catalog for English MeSH terms was CliniWeb and the most visible one for French MeSH terms was CISMeF.

Abstracting and Indexing↗

Building a text corpus for representing the variety of medical language.

Medical language processing has focused until recently on a few types of textual documents. However, a much larger variety of document types are used in different settings. It has been showed that Natural Language Processing (NLP) tools can exhibit very different behavior on different types of texts. Without better informed knowledge about the differential performance of NLP tools on a variety of medical text types, it will be difficult to control the extension of their application to different medical documents. We endeavored to provide a basis for such informed assessment: the construction of a large corpus of medical text samples. We propose a framework for designing such a corpus: a set of descriptive dimensions and a standardized encoding of both meta-information (implementing these dimensions) and content. We present a proof of concept demonstration by encoding an initial corpus of text samples according to these principles.

Documentation↗

The contribution of morphological knowledge to French MeSH mapping for information retrieval.

MeSH-indexed Internet health directories must provide a mapping from natural language queries to MeSH terms so that both health professionals and the general public can query their contents. We describe here the design of lexical knowledge bases for mapping French expressions to MeSH terms, and the initial evaluation of their contribution to Doc'CISMeF, the search tool of a MeSH-indexed directory of French-language medical Internet resources. The observed trend is in favor of the use of morphological knowledge as a moderate (approximately 5%) but effective factor for improving query to term mapping capabilities.

Algorithms↗

A general method for sifting linguistic knowledge from structured terminologies.

Morphological knowledge is useful for medical language processing, information retrieval and terminology or ontology development. We show how a large volume of morphological associations between words can be learnt from existing medical terminologies by taking advantage of the semantic relations already encoded between terms in these terminologies: synonymy, hierarchy and transversal relations. The method proposed relies on no a priori linguistic knowledge. Since it can work with different relations between terms, it can be applied to any structured terminology. Tested on SNOMED and ICD in French and English, it proves to identify fairly reliable morphological relations (precision > 90%) with a good coverage (over 88% compared to the UMLS lexical variant generation program). For English words with a stem longer than 3 characters, recall reaches 98.8% for inflection and 94.7% for derivation.

Linguistics↗

Identifying proper names in parallel medical terminologies.

We propose several criteria to identify proper names in biomedical terminologies. Traditional, pattern-based methods that rely on the immediate context of a proper name are not applicable. However, the availability of translations of some terminologies supports methods based on invariant words instead. A combination of five criteria achieved 86% precision and 88% recall on the 16,401 word forms of the International Classification of Diseases.

Disease↗

Language-independent automatic acquisition of morphological knowledge from synonym pairs.

Medical words exhibit a rich and productive morphology. Beyond simple inflection, derivation and composition are a common way to form new words. Morphological knowledge is therefore very important for any medical language processing application. Whereas rich morphological resources are available for the English medical language with the UMLS Specialist Lexicon, no such resources are publicly available for French or most other languages. We propose a simple and powerful method to help acquire automatically such knowledge. This method takes advantage of the synonym terms present in medical terminologies. In a bootstrapping step, it detects morphologically related words from which it learns "derivation rules". In an expansion step, it then applies these rules to the whole vocabulary available. Our goal is to acquire data for French and other languages for which they are not available. However, to evaluate the efficiency of the method, we tested it on English in a setting which is close to that prevailing for French, and we confronted its results to those obtained with the Specialist lexical variant generation tool.

Electronic Data Processing↗

Acquisition of lexical resources from SNOMED for medical language processing.

Medical language processing depends on large-coverage, fine-grained specialized lexicons. The vast majority of existing electronic lexicons concern the English language; for other languages such as French, resources are scarce. In contrast, large medical thesauri exist in numerous languages, including French. Our goal was to study what kind of linguistic information could be extracted from thesauri into a lexicon, in which places human intervention is necessary, and what kind of issues arise in this process. We designed in this purpose a method to build a semantic lexicon from a subset of the SNOMED axes in their French translation.

France↗

From text to knowledge: a unifying document-centered view of analyzed medical language.

Although medical language processing (MLP) has achieved some success, the actual use and dissemination of data extracted from free text by MLP systems is still very limited. We claim that the adoption of an 'enriched-document' paradigm (or 'document-centered' view) can help to address this issue. We present this paradigm and explain how it can be implemented, then discuss its expected benefits both for end-users and MLP researchers.

Artificial Intelligence↗

Hospitexte: towards a document-based hypertextual electronic medical record.

The patient record is a repository for knowledge about a patient. Work in Artificial Intelligence and knowledge representation has evidenced the intrinsic difficulty of formalizing knowledge for computer processing. It is therefore not a surprise that most attempts at computerizing the patient record have only had a limited degree of success or applicability. We claim that this is due to the fact that medicine is an empirical domain, and thus fundamentally resists formalization. Therefore, the only way medical knowledge can be fully expressed is through natural languages which is indeed what clinicians actually use. We proposed and designed an electronic medical record which adheres to this hypothesis and where structured documents play a prominent role.

Artificial Intelligence↗

Corpus-based identification and refinement of semantic classes.

Medical Language Processing (MLP), especially in specific domains, requires fine-grained semantic lexica. We examine whether robust natural language processing tools used on a representative corpus of a domain help in building and refining a semantic categorization. We test this hypothesis with ZELLIG, a corpus analysis tool. The first clusters we obtain are consistent with a model of the domain, as found in the SNOMED nomenclature. They correspond to coarse-grained semantic categories, but isolate as well lexical idiosyncrasies belonging to the clinical sub-language. Moreover, they help categorize additional words.

Classification↗

Evaluating a normalized conceptual representation produced from natural language patient discharge summaries.

The Menelas project aimed to produce a normalized conceptual representation from natural language patient discharge summaries. Because of the complex and detailed nature of conceptual representations, evaluating the quality of output of such a system is difficult. We present the method designed to measure the quality of Menelas output, and its application to the state of the French Menelas prototype as of the end of the project. We examine this method in the framework recently proposed by Friedman and Hripcsak. We also propose two conditions which enable to reduce the evaluation preparation workload.

Abstracting and Indexing↗

A multi-lingual architecture for building a normalised conceptual representation from medical language.

The overall goal of MENELAS is to provide better access to the information contained in natural language patient discharge summaries (PDSs), through the design and implementation of a prototype able to analyse medical texts. The approach taken by MENELAS is based on the following key principles: (i) to maximise the usefulness of natural language analysis and the usability of its results, the output of natural language analysis must be a normalised conceptual representation of medical information; and (ii) to maximise the reuse of resources, language analysis should be domain-independent and conceptual representation should be language-independent. This paper discusses the results obtained and the issues raised when implementing these principles during the project.

Artificial Intelligence↗

Issues in the structuring and acquisition of an ontology for medical language understanding.

Medical natural language understanding basically aims at representing the contents of medical texts in a formal, conceptual representation. The understanding process itself increasingly relies on a body of domain knowledge, generally expressed in the same conceptual formalism. The design of such a conceptual representation is a key knowledge-acquisition issue. When representing knowledge, the most important point is to ensure that the formal exploitation of the knowledge representation conforms to its meaning in the domain. We examined some methodological and theoretical principles to enforce this conformity. These principles result from our experience in MENELAS, a medical language understanding project.

Artificial Intelligence↗

MENELAS: an access system for medical records using natural language.

The overall goal of MENELAS is to provide better access to the information contained in natural language patient discharge summaries, through the design and implementation of a pilot system able to access medical reports through natural languages. A first, experimental version of the MENELAS indexing prototype for French has been assembled. Its function is to encode free text PDSs into both an internal representation and ICD-9-CM nomenclature codes. A preliminary evaluation shows the potential for reasonable coverage and precision. The MENELAS prototype will be enhanced and extended into a pilot system which will be tested in two hospital sites.

Abstracting and Indexing↗

Structuration and acquisition of medical knowledge. Using UMLS in the conceptual graph formalism.

The use of a taxonomy, such as the concept type lattice (CTL) of Conceptual Graphs, is a central structuring piece in a knowledge-based system. The knowledge it contains is constantly used by the system, and its structure provides a guide for the acquisition of other pieces of knowledge. We show how UMLS can be used as a knowledge resource to build a CTL and how the CTL can help the process of acquisition for other kinds of knowledge. We illustrate this method in the context of the MENELAS natural language understanding project.

Artificial Intelligence↗