PubMed Health⌕ Search

Biomedical subjects

Patrick Ruch

Publications and source records attributed to Patrick Ruch.

16 recordsLinked to original sources

Using argumentation to extract key sentences from biomedical abstracts.

PROBLEM: key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles and to help retrieving biomedical articles. Abstracts are also used for this purpose. However with usually more than 300 words, MEDLINE abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into four argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION METHODS: we rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification using Bayesian learners is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Abstracting and Indexing↗

Advancing biomedical image retrieval: development and analysis of a test collection.

OBJECTIVE: Develop and analyze results from an image retrieval test collection. METHODS: After participating research groups obtained and assessed results from their systems in the image retrieval task of Cross-Language Evaluation Forum, we assessed the results for common themes and trends. In addition to overall performance, results were analyzed on the basis of topic categories (those most amenable to visual, textual, or mixed approaches) and run categories (those employing queries entered by automated or manual means as well as those using visual, textual, or mixed indexing and retrieval methods). We also assessed results on the different topics and compared the impact of duplicate relevance judgments. RESULTS: A total of 13 research groups participated. Analysis was limited to the best run submitted by each group in each run category. The best results were obtained by systems that combined visual and textual methods. There was substantial variation in performance across topics. Systems employing textual methods were more resilient to visually oriented topics than those using visual methods were to textually oriented topics. The primary performance measure of mean average precision (MAP) was not necessarily associated with other measures, including those possibly more pertinent to real users, such as precision at 10 or 30 images. CONCLUSIONS: We developed a test collection amenable to assessing visual and textual methods for image retrieval. Future work must focus on how varying topic and run types affect retrieval performance. Users' studies also are necessary to determine the best measures for evaluating the efficacy of image retrieval systems.

Abstracting and Indexing↗

Health search engine with e-document analysis for reliable search results.

OBJECTIVE: After a review of the existing practical solution available to the citizen to retrieve eHealth document, the paper describes an original specialized search engine WRAPIN. METHOD: WRAPIN uses advanced cross lingual information retrieval technologies to check information quality by synthesizing medical concepts, conclusions and references contained in the health literature, to identify accurate, relevant sources. Thanks to MeSH terminology [1] (Medical Subject Headings from the U.S. National Library of Medicine) and advanced approaches such as conclusion extraction from structured document, reformulation of the query, WRAPIN offers to the user a privileged access to navigate through multilingual documents without language or medical prerequisites. RESULTS: The results of an evaluation conducted on the WRAPIN prototype show that results of the WRAPIN search engine are perceived as informative 65% (59% for a general-purpose search engine), reliable and trustworthy 72% (41% for the other engine) by users. But it leaves room for improvement such as the increase of database coverage, the explanation of the original functionalities and an audience adaptability. CONCLUSION: Thanks to evaluation outcomes, WRAPIN is now in exploitation on the HON web site (http://www.healthonnet.org), free of charge. Intended to the citizen it is a good alternative to general-purpose search engines when the user looks up trustworthy health and medical information or wants to check automatically a doubtful content of a Web page.

Europe↗

Automatic assignment of biomedical categories: toward a generic approach.

MOTIVATION: We report on the development of a generic text categorization system designed to automatically assign biomedical categories to any input text. Unlike usual automatic text categorization systems, which rely on data-intensive models extracted from large sets of training data, our categorizer is largely data-independent. METHODS: In order to evaluate the robustness of our approach we test the system on two different biomedical terminologies: the Medical Subject Headings (MeSH) and the Gene Ontology (GO). Our lightweight categorizer, based on two ranking modules, combines a pattern matcher and a vector space retrieval engine, and uses both stems and linguistically-motivated indexing units. RESULTS AND CONCLUSION: Results show the effectiveness of phrase indexing for both GO and MeSH categorization, but we observe the categorization power of the tool depends on the controlled vocabulary: precision at high ranks ranges from above 90% for MeSH to <20% for GO, establishing a new baseline for categorizers based on retrieval methods.

Abstracting and Indexing↗

Methodology to ease the construction of a terminology of problems.

INTRODUCTION: Problem lists summarize an aspect of the patient's medical history and provide an important way to implement entry points for clinical pathways and guideline-oriented care. However, in order to automate processes based on problem lists, the use of controlled vocabularies is required. We developed a methodology to extract a collection of standardized problem-related terms from medical documents entered in free text by physicians. METHODS: We extracted a corpus of sentences describing problems from a randomized selection of admission notes collected at the University Hospitals of Geneva. Theses sentences underwent manual and automatic normalization processes, and a statistical clustering, in order to build a set of terms. RESULTS: We obtained 17,805 sentences from 5000 admission notes. We refined them into 1546 terms, 88.6% of which could be related to a relevant problem statement. DISCUSSION: A clinically relevant problems terminology was derived from clinical admission notes in free-text using a few methodical steps with a reasonable investment of human resources. Such an approach will ease the development and the use of problem lists better suited to user needs.

Medical Records, Problem-Oriented↗

Using argumentation to retrieve articles with similar citations: an inquiry into improving related articles search in the MEDLINE digital library.

The aim of this study is to investigate the relationships between citations and the scientific argumentation found abstracts. We design a related article search task and observe how the argumentation can affect the search results. We extracted citation lists from a set of 3200 full-text papers originating from a narrow domain. In parallel, we recovered the corresponding MEDLINE records for analysis of the argumentative moves. Our argumentative model is founded on four classes: PURPOSE, METHODS, RESULTS and CONCLUSION. A Bayesian classifier trained on explicitly structured MEDLINE abstracts generates these argumentative categories. The categories are used to generate four different argumentative indexes. A fifth index contains the complete abstract, together with the title and the list of Medical Subject Headings (MeSH) terms. To appraise the relationship of the moves to the citations, the citation lists were used as the criteria for determining relatedness of articles, establishing a benchmark; it means that two articles are considered as "related" if they share a significant set of co-citations. Our results show that the average precision of queries with the PURPOSE and CONCLUSION features is the highest, while the precision of the RESULTS and METHODS features was relatively low. A linear weighting combination of the moves is proposed, which significantly improves retrieval of related articles.

Abstracting and Indexing↗

Recent advances in natural language processing for biomedical applications.

We survey a set a recent advances in natural language processing applied to biomedical applications, which were presented in Geneva, Switzerland, in 2004 at an international workshop. While text mining applied to molecular biology and biomedical literature can report several interesting achievements, we observe that studies applied to clinical contents are still rare. In general, we argue that clinical corpora, including electronic patient records, must be made available to fill the gap between bioinformatics and medical informatics.

Abstracting and Indexing↗

Data-poor categorization and passage retrieval for gene ontology annotation in Swiss-Prot.

BACKGROUND: In the context of the BioCreative competition, where training data were very sparse, we investigated two complementary tasks: 1) given a Swiss-Prot triplet, containing a protein, a GO (Gene Ontology) term and a relevant article, extraction of a short passage that justifies the GO category assignment; 2) given a Swiss-Prot pair, containing a protein and a relevant article, automatic assignment of a set of categories. METHODS: Sentence is the basic retrieval unit. Our classifier computes a distance between each sentence and the GO category provided with the Swiss-Prot entry. The Text Categorizer computes a distance between each GO term and the text of the article. Evaluations are reported both based on annotator judgements as established by the competition and based on mean average precision measures computed using a curated sample of Swiss-Prot. RESULTS: Our system achieved the best recall and precision combination both for passage retrieval and text categorization as evaluated by official evaluators. However, text categorization results were far below those in other data-poor text categorization experiments The top proposed term is relevant in less that 20% of cases, while categorization with other biomedical controlled vocabulary, such as the Medical Subject Headings, we achieved more than 90% precision. We also observe that the scoring methods used in our experiments, based on the retrieval status value of our engines, exhibits effective confidence estimation capabilities. CONCLUSION: From a comparative perspective, the combination of retrieval and natural language processing methods we designed, achieved very competitive performances. Largely data-independent, our systems were no less effective that data-intensive approaches. These results suggests that the overall strategy could benefit a large class of information extraction tasks, especially when training data are missing. However, from a user perspective, results were disappointing. Further investigations are needed to design applicable end-user text mining tools for biologists.

Computational Biology↗

UMLF: a unified medical lexicon for French.

Medical Informatics has a constant need for basic medical language processing tasks, e.g. for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Abstracting and Indexing↗

Extracting key sentences with latent argumentative structuring.

PROBLEM: Key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles. Abstracts are also used for this purpose. However with usually more than 300 words, abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into 4 argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION. METHODS: We rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Bayes Theorem↗

Coping with the variability of medical terms.

OBJECTIVES: To cope with medical terms, which present a high variability of expression through a single natural language, in the sense that any term may be reformulated in hundred of different ways. METHODS: A typology of term variants is presented as a systematic approach in order to favour the implementation of an exhaustive solution. Then, an algorithm able to handle all variants is designed. RESULTS: Using MetaMap, single terms are analyzed with a success rate varying between 68 and 88 %; the algorithm presented in this paper improves this situation. CONCLUSIONS: This experience shows that a semantic driven method, based on a thesaurus, provides a satisfactory solution to the problem of variability of a single term. The presented typology is representative of most variants in a language.

Algorithms↗

Towards a unified medical lexicon for French.

Medical Informatics has a constant need for basic Medical Language Processing tasks, e.g., for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Algorithms↗

A frame-based representation of ICD-10.

UNLABELLED: Physicians are required to code information concerning a patient's stay in order to measure the medical activity in hospitals. They use the International Statistical Classification of Diseases and Related Health Problems, Tenth Revision (ICD-10). Coding is usually performed manually and computerized tools may be useful in speeding up and facilitating the tedious task of coding patient information. The aim of this work is to build a surface semantic model of ICD-10 in order to ameliorate a coding help system. METHODS: This work was focused on chapter XI of the ICD-10, Diseases of the Digestive System. Each term from both analytical and alphabetical indexes about this chapter were submitted to a morphological analysis in order to extract the medical concepts within. After a statistical analysis of these concepts and the way they connect themselves, a semantic model based on a "semantic frame" approach was built. RESULTS: Although this model could represent a reasonable amount of medical knowledge within chapter XI of the ICD-10 in a quite satisfactory way, it shows lack of efficiency for some other chapters. CONCLUSION: Difficulties have to be overcome when modelling a classification meant for manual utilisation, and a lot of work still has to be done to obtain an effective coding help system using the ICD-10.

Forms and Records Control↗

UMLF: a Unified Medical Lexicon for French.

Lexical resources for medical language, such as lists of words with inflectional and derivational information, are publicly available for the English lantuate with the UMLS Specialist Lexicon. The goal of the UMLF project is to pool and unify existing resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a Unified Medical Lexicon for French. We present here the current status of the project.

France↗

Using lexical disambiguation and named-entity recognition to improve spelling correction in the electronic patient record.

In this article, we show how a set of natural language processing (NLP) tools can be combined to improve the processing of clinical records. The study concentrates on improving spelling correction, which is of major importance for quality control in the electronic patient record (EPR). As first task, we report on the design of an improved interactive tool for correcting spelling errors. Unlike traditional systems, the linguistic context (both semantic and syntactic) is used to improve the correction strategy. The system is organized along three modules. Module 1 is based on a classical spelling checker, it means that it is context-independent and simply measures a string-edit-distance between a misspelled word and a list of well-formed words. Module 2 attempts to rank more relevantly the set of candidates provided by the first module using morpho-syntactic disambiguation tools. Module 3 processes words with the same part-of-speech (POS) and apply word-sense (WS) disambiguation in order to rerank the set of candidates. As second task, we show how this improved interactive spell checker can be cast as a fully automatic system by adjunction of another NLP module: a named-entity (NE) extractor, i.e. a tool able to identify words as such patient and physician names. This module is used to avoid replacement of named-entities when the system is not used in an interactive mode. Results confirm that using the linguistic context can improve interactive spelling correction, and justify the use of named-entity recognizer to conduct fully automatic spelling correction. It is concluded that NLP is mature enough to help information processing in EPR.

Artificial Intelligence↗