PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Grepator: accents & case mix for thesaurus.

There is a real need among researchers and students for pedagogical resources. In France, information retrieval techniques have been developed, for example in the Doc'CISMeF web site. As Pubmed, documents are indexed with (French) MeSH terms, one of the problems discovered, in quality studies, is the inadequacies between the user requests and the MeSH controlled vocabulary. Moreover, French (but also Greek or Spanish), pose specific problems for indexing, due to the diacritic characters.In this article, we present the Grepator project. The main goal is to transform any thesaurus (or any entry) in case mix and accentuated characters, for a specific domain. Furthermore, Grepator has to complete MeSH terms according to their usual form in natural language and finally, to correct user spelling mistakes. Grepator is based on a statistical approach. A large French medical corpus has been constituted from pedagogical resources indexed in CISMeF. Using regular expressions, Grepator searches the more usual ways to spell the word.. Seventy five percent of MeSH terms are found in the corpus, using this method, with less than one mistake for a hundred words. This first evaluation of the tools is analyzed and we discuss further steps that might be developed.

Diagnosis-Related Groups↗

Predicting Lexical Relations between Biomedical Terms: towards a Multilingual Morphosemantics-based System.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This issue is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morpho-semantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a la morphosemantic parser, (2) a multilingual table defining basic relations between word roots, and (3) a set of language-independant rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3,000 lexical families.

Language↗

Simplified representation of concepts and relations on screen.

The fully automated generation of diagnostic codes requires a knowledge-based system which is capable of interpreting noun phrases. The sense content of the words must be analysed and represented for this purpose. The codes are then generated based on this representation.In comparison with other knowledge-based systems, a system of this kind places the emphasis on the data structures and not on the calculus; coding itself is a simple matter compared to the much more difficult task of incorporating the complex information contained in the words used in natural language in a systematic data model. Initial attempts were based on the assumption that each word was linked to one conceptual meaning, whereas such a naive viewpoint certainly no longer applies today. The notation of concepts and their relations is the task at hand.Existing notation methods include predicate logic, conceptual graphs (CGs) as proposed by J. F. Sowa [2], GRAIL as used by the GALEN Project [1] and methods developed as part of the WWW consortium, e.g. RDF's (Resource Description Frameworks). For the purpose of coding, we developed a notation system using "concept particles" back in 1989 [3]. In 1996, the resulting experience led us to represent "concept molecules" (CM), with which both complex data structures and multi-branched rules can be denoted in a simple manner [4]. In this paper we shall explain the principles behind this notation and compare it with another modern concept representation system, conceptual graphs.

Expert Systems↗

Breaking the language barrier: machine assisted diagnosis using the medical speech translator.

In this paper, we describe and evaluate an Open Source medical speech translation system (MedSLT) intended for safety-critical applications. The aim of this system is to eliminate the language barriers in emergency situation. It translates spoken questions from English into French, Japanese and Finnish in three medical subdomains (headache, chest pain and abdominal pain), using a vocabulary of about 250-400 words per sub-domain. The architecture is a compromise between fixed-phrase translation on one hand and complex linguistically-based systems on the other. Recognition is guided by a Context Free Grammar Language Model compiled from a general unification grammar, automatically specialised for the domain. We present an evaluation of this initial prototype that shows the advantages of this grammar-based approach for this particular translation task in term of both reliability and use.

Communication Barriers↗

Automatic lexicon acquisition for a medical cross-language information retrieval system.

We present a method for the automated acquisition of a multilingual medical lexicon (for Spanish and Swedish) to be used within the framework of a medical cross-language text retrieval system. We incorporate seed lexicons and parallel corpora derived from the UMLS Metathesaurus. The seed lexicons for Spanish and Swedish are automatically generated from (previously manually constructed) Portuguese, German and English sources. Lexical and semantic hypotheses are then validated making iterative use of co-occurrence patterns of hypothesized translation synonyms in the parallel corpora.

Humans↗

Extracting key sentences with latent argumentative structuring.

PROBLEM: Key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles. Abstracts are also used for this purpose. However with usually more than 300 words, abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into 4 argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION. METHODS: We rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Bayes Theorem↗

[Continuous speech recognition system for radiological reporting: comparison with experience of dictation].

PURPOSE: To compare rates of accuracy of recognition between experienced dictators and inexperienced ones in using an enrollment-less continuous speech recognition (CSR) system of radiological reporting, and to evaluate the usefulness of the system. MATERIALS AND METHODS: Twenty board-certified radiologists were classified into 2 groups: a group of 10 members with more than 6 years' experience of conventional dictation by transcriptionist (group A) and a group of 10 members with no experience of dictation (group B). All radiologists created fresh radiological reports on sets of images using free-style dictation in the reports. We counted errors and total words in the reports individually, and compared the rates of accuracy of word recognition in the two groups. We used a CSR system AmiVoice (Advanced Media, Inc., Tokyo, Japan). RESULTS: The average rate of accuracy of word recognition was 96.42 +/- 1.68% in group A and 95.92 +/- 1.15% in group B. There was no significant difference in accuracy rate between the two groups. CONCLUSION: The accuracy of word recognition was independent of the experience of dictation, and the enrollment-less CSR system of radiological reporting was considered convenient and useful.

Humans↗

A terminological and ontological analysis of the NCI Thesaurus.

OBJECTIVE: The National Cancer Institute Thesaurus is described by its authors as "a biomedical vocabulary that provides consistent, unambiguous codes and definitions for concepts used in cancer research" and which "exhibits ontology-like properties in its construction and use". We performed a qualitative analysis of the Thesaurus in order to assess its conformity with principles of good practice in terminology and ontology design. MATERIALS AND METHODS: We used both the on-line browsable version of the Thesaurus and its OWL-representation (version 04.08b, released on August 2, 2004), measuring each in light of the requirements put forward in relevant ISO terminology standards and in light of ontological principles advanced in the recent literature. RESULTS: We found many mistakes and inconsistencies with respect to the term-formation principles used, the underlying knowledge representation system, and missing or inappropriately assigned verbal and formal definitions. CONCLUSION: Version 04.08b of the NCI Thesaurus suffers from the same broad range of problems that have been observed in other biomedical terminologies. For its further development, we recommend the use of a more principled approach that allows the Thesaurus to be tested not just for internal consistency but also for its degree of correspondence to that part of reality which it is designed to represent.

Computational Biology↗

MorphoSaurus--design and evaluation of an interlingua-based, cross-language document retrieval engine for the medical domain.

OBJECTIVES: We propose an interlingua-based indexing approach to account for the particular challenges that arise in the design and implementation of cross-language document retrieval systems for the medical domain. METHODS: Documents, as well as queries, are mapped to a language-independent conceptual layer on which retrieval operations are performed. We contrast this approach with the direct translation of German queries to English ones which, subsequently, are matched against English documents. RESULTS: We evaluate both approaches, interlingua-based and direct translation, on a large medical document collection, the OHSUMED corpus. A substantial benefit for interlingua-based document retrieval using German queries on English texts is found, which amounts to 93% of the (monolingual) English baseline. CONCLUSIONS: Most state-of-the-art cross-language information retrieval systems translate user queries to the language(s) of the target documents. In contra-distinction to this approach, translating both documents and user queries into a language-independent, concept-like representation format is more beneficial to enhance cross-language retrieval performance.

Abstracting and Indexing↗

Discovering compact and highly discriminative features or feature combinations of drug activities using support vector machines.

Nowadays, high throughput experimental techniques make it feasible to examine and collect massive data at the molecular level. These data, typically mapped to a very high dimensional feature space, carry rich information about functionalities of certain chemical or biological entities and can be used to infer valuable knowledge for the purposes of classification and prediction. Typically, a small number of features or feature combinations may play determinant roles in functional discrimination. The identification of such features or feature combinations is of great importance. In this paper, we study the problem of discovering compact and highly discriminative features or feature combinations from a rich feature collection. We employ the support vector machine as the classification means and aim at finding compact feature combinations. Comparing to previous methods on feature selection, which identify features solely based on their individual roles in the classification, our method is able to identify minimal feature combinations that ultimately have determinant roles in a systematic fashion. Experimental study on drug activity data shows that our method can discover descriptors that are not necessarily significant individually but are most significant collectively.

Algorithms↗

IAIMS and UMLS at Columbia-Presbyterian Medical Center.

The authors use an example to illustrate combining Integrated Academic Information Management System (IAIMS) components (applications) into an integral whole, to facilitate using the components simultaneously or in sequence. They examine a model for classifying IAIMS systems, proposing ways in which the United Medical Language System (UMLS) can be exploited them.

Computer Peripherals↗

Assessing the difficulty and time cost of de-identification in clinical narratives.

OBJECTIVE: To characterize the difficulty confronting investigators in removing protected health information (PHI) from cross-discipline, free-text clinical notes, an important challenge to clinical informatics research as recalibrated by the introduction of the US Health Insurance Portability and Accountability Act (HIPAA) and similar regulations. METHODS: Randomized selection of clinical narratives from complete admissions written by diverse providers, reviewed using a two-tiered rater system and simple automated regular expression tools. For manual review, two independent reviewers used simple search and replace algorithms and visual scanning to find PHI as defined by HIPAA, followed by an independent second review to detect any missed PHI. Simple automated review was also performed for the "easy" PHI that are number- or date-based. RESULTS: From 262 notes, 2074 PHI, or 7.9 +/- 6.1 per note, were found. The average recall (or sensitivity) was 95.9% while precision was 99.6% for single reviewers. Agreement between individual reviewers was strong (ICC = 0.99), although some asymmetry in errors was seen between reviewers (p = 0.001). The automated technique had better recall (98.5%) but worse precision (88.4%) for its subset of identifiers. Manually de-identifying a note took 87.3 +/- 61 seconds on average. CONCLUSIONS: Manual de-identification of free-text notes is tedious and time-consuming, but even simple PHI is difficult to automatically identify with the exactitude required under HIPAA.

Confidentiality↗

Concept-value pair extraction from semi-structured clinical narrative: a case study using echocardiogram reports.

The task of gathering detailed patient information from narrative text presents a significant barrier to clinical research. A prototype information extraction system was developed to identify concepts and their associated values from narrative echocardiogram reports. The system uses a Unified Medical Language System compatible architecture and takes advantage of canonical language use patterns to identify sentence templates with which concepts and their related values can be identified. The data extracted from this system will be used to enrich an existing database used by clinical researchers in a large university healthcare system to identify potential research candidates fulfilling clinical inclusion criteria. The system was developed and evaluated using ten clinical concepts. Concept-value pairs extracted by the system were compared with findings extracted manually by the author. The system was able to recall 78% [95%CI, 76-80%] of the relevant findings, with a precision of 99% [95%CI, 98-99%].

Echocardiography↗

Using patient data to retrieve health knowledge.

BACKGROUND: We sought to study a variety of terminologic approaches to the use of clinical data for searching on-line information resources. METHODS: We used a collection of narrative text and coded data to search a variety of text-based, concept-based, and concept-indexed resources. RESULTS: Automated retrievals using original terms could be accomplished and technically produced ample results. However, quality of the results varied with the resource. Terminology translations were difficult to accomplish and produced variable results. CONCLUSIONS: Current resources support automated retrieval; however, achieving quality results varies with the terms and the resources, with term translation productive only in select situations.

Abstracting and Indexing↗

ReportTutor - an intelligent tutoring system that uses a natural language interface.

ReportTutor is an extension to our work on Intelligent Tutoring Systems for visual diagnosis. ReportTutor combines a virtual microscope and a natural language interface to allow students to visually inspect a virtual slide as they type a diagnostic report on the case. The system monitors both actions in the virtual microscope interface as well as text created by the student in the reporting interface. It provides feedback about the correctness, completeness, and style of the report. ReportTutor uses MMTx with a custom data-source created with the NCI Metathesaurus. A separate ontology of cancer specific concepts is used to structure the domain knowledge needed for evaluation of the student's input including co-reference resolution. As part of the early evaluation of the system, we collected data from 4 pathology residents who typed in their reports without the tutoring aspects of the system, and compared responses to an expert dermatopathologist. We analyzed the resulting reports to (1) identify the error rates and distribution among student reports, (2) determine the performance of the system in identifying features within student reports, and (3) measure the accuracy of the system in distinguishing between correct and incorrect report elements.

Artificial Intelligence↗

Semi-automatic indexing of full text biomedical articles.

The main application of U.S. National Library of Medicine's Medical Text Indexer (MTI) is to provide indexing recommendations to the Library's indexing staff. The current input to MTI consists of the titles and abstracts of articles to be indexed. This study reports on an extension of MTI to the full text of articles appearing in online medical journals that are indexed for Medline. Using a collection of 17 journal issues containing 500 articles, we report on the effectiveness of the contribution of terms by the whole article and also by each section. We obtain the best results using a model consisting of the sections Results, Results and Discussion, and Conclusions together with the article's title and abstract, the captions of tables and figures, and sections that have no titles. The resulting model provides indexing significantly better (7.4%) than what is currently achieved using only titles and abstracts.

Abstracting and Indexing↗

A strategy for assigning new concepts in the MEDLINE database.

The MeSH indexing done in MEDLINE is engineered by humans. Humans define the MeSH concepts and human indexers assign MeSH terms to MEDLINE records. Methods have been designed in an attempt to assign MeSH terms to MEDLINE documents automatically with some success. Methods have also been designed to locate useful phrases as potential concepts for indexing. However, little work has been done on the problem of how one might automatically index with the concepts represented by such phrases. Here we examine this issue and present a method for such indexing.

Abstracting and Indexing↗

Towards semantic role labeling & IE in the medical literature.

INTRODUCTION: In this work, we introduce the concept of semantic role labeling to the medical domain. We report first results of porting and adapting an existing resource, Propbank, to the medical field. Propbank is an adjunct to Penn Treebank that provides semantic annotation of predicates and the roles played by their arguments. The main aim of this work is the applicability of the Propbank frame files to predicates typically encountered in the medical literature. METHODS: We analyzed a target corpus of 610,100 abstracts, which was selected by searching for publication type "case reports". From this target corpus, we randomly selected 10,000 sample abstracts to estimate the predicate distribution, and matched the predicates from this sample to the predicates in Propbank. RESULTS: Of the 1998 unique verbs in our sample, 76% were represented in Propbank. This included the 40 most frequent verbs, which represented 49% of all predicate instances in our sample and which matched the Propbank usage in a study of representative sentences. We propose extensions to Propbank that handle medical predicates, which are not adequately covered by Propbank. CONCLUSION: We believe that semantic role labeling using Propbank is a valid approach to capture predicate relations in the medical literature.

Abstracting and Indexing↗