PubMed Health⌕ Search

Biomedical subjects

Robert Baud

Publications and source records attributed to Robert Baud.

17 recordsLinked to original sources

Health search engine with e-document analysis for reliable search results.

OBJECTIVE: After a review of the existing practical solution available to the citizen to retrieve eHealth document, the paper describes an original specialized search engine WRAPIN. METHOD: WRAPIN uses advanced cross lingual information retrieval technologies to check information quality by synthesizing medical concepts, conclusions and references contained in the health literature, to identify accurate, relevant sources. Thanks to MeSH terminology [1] (Medical Subject Headings from the U.S. National Library of Medicine) and advanced approaches such as conclusion extraction from structured document, reformulation of the query, WRAPIN offers to the user a privileged access to navigate through multilingual documents without language or medical prerequisites. RESULTS: The results of an evaluation conducted on the WRAPIN prototype show that results of the WRAPIN search engine are perceived as informative 65% (59% for a general-purpose search engine), reliable and trustworthy 72% (41% for the other engine) by users. But it leaves room for improvement such as the increase of database coverage, the explanation of the original functionalities and an audience adaptability. CONCLUSION: Thanks to evaluation outcomes, WRAPIN is now in exploitation on the HON web site (http://www.healthonnet.org), free of charge. Intended to the citizen it is a good alternative to general-purpose search engines when the user looks up trustworthy health and medical information or wants to check automatically a doubtful content of a Web page.

Europe↗

Methodology to ease the construction of a terminology of problems.

INTRODUCTION: Problem lists summarize an aspect of the patient's medical history and provide an important way to implement entry points for clinical pathways and guideline-oriented care. However, in order to automate processes based on problem lists, the use of controlled vocabularies is required. We developed a methodology to extract a collection of standardized problem-related terms from medical documents entered in free text by physicians. METHODS: We extracted a corpus of sentences describing problems from a randomized selection of admission notes collected at the University Hospitals of Geneva. Theses sentences underwent manual and automatic normalization processes, and a statistical clustering, in order to build a set of terms. RESULTS: We obtained 17,805 sentences from 5000 admission notes. We refined them into 1546 terms, 88.6% of which could be related to a relevant problem statement. DISCUSSION: A clinically relevant problems terminology was derived from clinical admission notes in free-text using a few methodical steps with a reasonable investment of human resources. Such an approach will ease the development and the use of problem lists better suited to user needs.

Medical Records, Problem-Oriented↗

Amplification of Terminologia anatomica by French language terms using Latin terms matching algorithm: a prototype for other language.

OBJECTIVE: Terminologia anatomica is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subject to the addition of terms from others languages. On the other hand, Nomina anatomica, the previous standard, has been widely translated. Aim of this work was to append foreign terms to Terminologia by using similarity-matching algorithm between its Latin terms and those from Nomina. METHODS: A semi-automatic matching of Latin terms from Terminologia with those of Nomina was performed using a string-to-string distance algorithm and manual assessment. We used a French-Latin version of Nomina together with Terminologia and we suggested French terms for Terminologia. Coverage was evaluated by the number of exact and approximate matches. A target of 78% was set due to the higher number of terms in Terminologia compared to Nomina. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5982 terms (76.5%) of Terminologia. Our results indicated that more than 75% of the terms from Terminologia came from Nomina, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 78% target. The method is based only on Latin terms and can be used for other languages. We consider this work as a starting point for adding terms to other knowledge sources, such as the foundational model of anatomy or the Unified Medical Language System (UMLS).

Algorithms↗

Recent advances in natural language processing for biomedical applications.

We survey a set a recent advances in natural language processing applied to biomedical applications, which were presented in Geneva, Switzerland, in 2004 at an international workshop. While text mining applied to molecular biology and biomedical literature can report several interesting achievements, we observe that studies applied to clinical contents are still rare. In general, we argue that clinical corpora, including electronic patient records, must be made available to fill the gap between bioinformatics and medical informatics.

Abstracting and Indexing↗

UMLF: a unified medical lexicon for French.

Medical Informatics has a constant need for basic medical language processing tasks, e.g. for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Abstracting and Indexing↗

Towards a multilingual version of terminologia anatomica.

OBJECTIVE: Terminologia Anatomica (TA) is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subdued to the addition of terms from others languages. On the other hand Nomina Anatomica (NA), the previous standard, has been widely translated. Aim of this work was to append foreign terms to TA by using similarity matching algorithm between its Latin terms and those from NA. METHODS: A semi-automatic matching of Latin terms from TA with those of NA was performed using a string-to-string distance algorithm and manual assessment. We used a French - Latin version of NA together with TA and we suggested French terms for TA. Coverage was evaluated by the number of exact and approximate matches. A target of 80% was set due to the superior number of terms in TA compared to NA. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5,982 terms (76.5%) of TA. Our results outlined that more than 75% of the terms from TA came from NA, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 80% target. The method is based only on Latin terms and can be used for other languages and for others terminologies including Latin terms.

Algorithms↗

Predicting Lexical Relations between Biomedical Terms: towards a Multilingual Morphosemantics-based System.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This issue is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morpho-semantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a la morphosemantic parser, (2) a multilingual table defining basic relations between word roots, and (3) a set of language-independant rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3,000 lexical families.

Language↗

Extracting key sentences with latent argumentative structuring.

PROBLEM: Key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles. Abstracts are also used for this purpose. However with usually more than 300 words, abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into 4 argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION. METHODS: We rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Bayes Theorem↗

A natural language based search engine for ICD10 diagnosis encoding.

We have developed a multiple step process for implementing an ICD10 search engine. The complexity of the task has been shown and we recommend collecting adequate expertise before starting any implementation. Underestimation of the expert time and inadequate data resources are probable reasons for failure. We also claim that when all conditions are met in term of resource and availability of the expertise, the benefits of a responsive ICD10 search engine will be present and the investment will be successful.

Databases as Topic↗

XML as standard for communicating in a document-based electronic patient record: a 3 years experiment.

During the past few years, the eXtensible Markup Language (XML) has progressively become a gold standard for accessing, representing and exchanging information, especially in the health care environment. This paper presents an implementation of the use of XML for the electronic patient record (EPR) and discusses more specifically its growing use in two areas of the EPR: first, as a format for the exchange of structured messages, and second, as a comprehensible way of representing patient documents. These statements rely on a 3 years experiment conducted at the Geneva University Hospital as part of its document-centered EPR.

Delivery of Health Care↗

Towards a unified medical lexicon for French.

Medical Informatics has a constant need for basic Medical Language Processing tasks, e.g., for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Algorithms↗

A frame-based representation of ICD-10.

UNLABELLED: Physicians are required to code information concerning a patient's stay in order to measure the medical activity in hospitals. They use the International Statistical Classification of Diseases and Related Health Problems, Tenth Revision (ICD-10). Coding is usually performed manually and computerized tools may be useful in speeding up and facilitating the tedious task of coding patient information. The aim of this work is to build a surface semantic model of ICD-10 in order to ameliorate a coding help system. METHODS: This work was focused on chapter XI of the ICD-10, Diseases of the Digestive System. Each term from both analytical and alphabetical indexes about this chapter were submitted to a morphological analysis in order to extract the medical concepts within. After a statistical analysis of these concepts and the way they connect themselves, a semantic model based on a "semantic frame" approach was built. RESULTS: Although this model could represent a reasonable amount of medical knowledge within chapter XI of the ICD-10 in a quite satisfactory way, it shows lack of efficiency for some other chapters. CONCLUSION: Difficulties have to be overcome when modelling a classification meant for manual utilisation, and a lot of work still has to be done to obtain an effective coding help system using the ICD-10.

Forms and Records Control↗

Clinical documents: attribute-values entity representation,context, page layout and communication.

This paper presents how acquisition, storage and communication of clinical documents is implemented at the University Hospitals of Geneva. Careful attention has been given to user-interfaces, in order to support complex layouts, spell checking, and templates management with automatic prefilling. A dual architecture has been developed for storage using an entity-attribute-value unified database and a consolidated, patient-centered, layout-respectful file-based storage, providing both representation power and speed of access. This architecture allows a great flexibility for storing a continuum of data types, ranging from simple typed values to complex clinical reports. Finally, communication is entirely based on HTTP-XML internally, and a HL-7 CDA interface V2 is currently studied for external communication. Some of the problems encountered, mostly related to the typology of documents and the ontology of clinical attributes are evoked.

Computer Systems↗

UMLF: a Unified Medical Lexicon for French.

Lexical resources for medical language, such as lists of words with inflectional and derivational information, are publicly available for the English lantuate with the UMLS Specialist Lexicon. The goal of the UMLF project is to pool and unify existing resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a Unified Medical Lexicon for French. We present here the current status of the project.

France↗

XML as standard for communicating in a document-based electronic patient record: a three years experiment.

During the past few years, the eXtensible Markup Language (XML) has experienced a growing use for accessing, representing and exchanging information, especially in the health care environment. This paper discusses the potentials of the use of XML for the electronic patient record (EPR) in two ways: first, as a format for the exchange of structured messages, and second, as a comprehensible way of representing patient documents. These statements rely on a three years experiment conducted at the Geneva University Hospital as part of its document-centred EPR.

Medical Records Systems, Computerized↗

Using lexical disambiguation and named-entity recognition to improve spelling correction in the electronic patient record.

In this article, we show how a set of natural language processing (NLP) tools can be combined to improve the processing of clinical records. The study concentrates on improving spelling correction, which is of major importance for quality control in the electronic patient record (EPR). As first task, we report on the design of an improved interactive tool for correcting spelling errors. Unlike traditional systems, the linguistic context (both semantic and syntactic) is used to improve the correction strategy. The system is organized along three modules. Module 1 is based on a classical spelling checker, it means that it is context-independent and simply measures a string-edit-distance between a misspelled word and a list of well-formed words. Module 2 attempts to rank more relevantly the set of candidates provided by the first module using morpho-syntactic disambiguation tools. Module 3 processes words with the same part-of-speech (POS) and apply word-sense (WS) disambiguation in order to rerank the set of candidates. As second task, we show how this improved interactive spell checker can be cast as a fully automatic system by adjunction of another NLP module: a named-entity (NE) extractor, i.e. a tool able to identify words as such patient and physician names. This module is used to avoid replacement of named-entities when the system is not used in an interactive mode. Results confirm that using the linguistic context can improve interactive spelling correction, and justify the use of named-entity recognizer to conduct fully automatic spelling correction. It is concluded that NLP is mature enough to help information processing in EPR.

Artificial Intelligence↗