PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Coping with the variability of medical terms.

OBJECTIVES: To cope with medical terms, which present a high variability of expression through a single natural language, in the sense that any term may be reformulated in hundred of different ways. METHODS: A typology of term variants is presented as a systematic approach in order to favour the implementation of an exhaustive solution. Then, an algorithm able to handle all variants is designed. RESULTS: Using MetaMap, single terms are analyzed with a success rate varying between 68 and 88 %; the algorithm presented in this paper improves this situation. CONCLUSIONS: This experience shows that a semantic driven method, based on a thesaurus, provides a satisfactory solution to the problem of variability of a single term. The presented typology is representative of most variants in a language.

Algorithms↗

WRAPIN: new generation health search engine using UMLS knowledge sources for MeSH term extraction from health documentation.

To realize the potential of the Internet as a source of valuable healthcare information, for the general public, patients or practitioners, it is imperative to establish a validation system based on standards of quality. The WRAPIN project (World-wide online Reliable Advice to Patients and Individuals) from the European Community has this ambitious goal. WRAPIN is a federating system for medical information with an editorial policy of intelligently sharing quality and professional information. The WRAPIN project has two main axes: the efficient and intelligent search of information and the assertion of the trustworthiness of content. This article presents the scientific challenges involved in extracting the knowledge from text-based information in order to better manage the knowledge and the rest of the retrieval proc-ess. Our innovative approach is to efficiently extract MeSH terms from the analyzed documents exploiting UMLS knowledge sources. A benefit has been measured when comparing extraction results. Even if the evaluation is made with a limited corpus, this research work proposes heuristics that can be validated to the whole biomedical domain, and possibly enhanced by the adjunction of other methods.

Abstracting and Indexing↗

Using a terminology server and consumer search phrases to help patients find physicians with particular expertise.

OBJECTIVES: To design and implement a real world application using a terminology server to assist patients and physicians who use common language search terms to find specialist physicians with a particular clinical expertise. METHOD: Terminology servers have been developed to help users encoding of information using complicated structured vocabulary during data entry tasks, such as recording clinical information. We describe a methodology using Personal Health Terminology trade mark and a SNOMED CT-based hierarchical concept server. RESULTS: Construction of a pilot mediated-search engine to assist users who use vernacular speech in querying data which is more technical than vernacular. CONCLUSION: This approach, which combines theoretical and practical requirements, provides a useful example of concept-based searching for physician referrals.

Abstracting and Indexing↗

Summarization of an online medical encyclopedia.

We explore a knowledge-rich (abstraction) approach to summarization and apply it to multiple documents from an online medical encyclopedia. A semantic processor functions as the source interpreter and produces a list of predications. A transformation stage then generalizes and condenses this list, ultimately generating a conceptual condensate for a given disorder topic. We provide a preliminary evaluation of the quality of the condensates produced for a sample of four disorders. The overall precision of the disorder conceptual condensates was 87%, and the compression ratio from the base list of predications to the final condensate was 98%. The conceptual condensate could be used as input to a text generator to produce a natural language summary for a given disorder topic.

Disease↗

A surgeon can operate or approach a complicated integral calculus instead of a renal one: implications for conceptual and lexical semantic disambiguation.

This paper approaches lexical semantic disambiguation and polysemy in the context of medical language understanding. These issues have been addressed as linguistic and cognitive requirements to build a knowledge extraction tool that turns natural language input into conceptual graphs. Previous results obtained with semantic analysis of medical terms in the domain of transplantation and organ failure prompt us to check the capabilities of our prototype to deal with ambiguities and polysemy. Starting from linguistic observation, we attempt to demonstrate how the respect of ambiguities when the co-text is not sufficient for disambiguation implies to introduce a void type in the concept type lattice. Then we show how the "creative use of words" (i.e. new senses in novel context) imposes to dynamically allocate type, categories and roles with the co-text.

Linguistics↗

Acquiring meaning for French medical terminology: contribution of morphosemantics.

Morphologically complex words, and particularly neoclassical compounds, form more than 60% of the neologisms in the biomedical field. Guessing their definitions and grouping them into semantic classes by means of lexical relations are thus two crucial improvements for handling these words, e.g., for information retrieval, indexing and text understanding applications. This paper describes a morphosemantic linguistic-based parser called DériF, currently developed in the framework of two projects, UMLF and VUMeF, and its application to French biomedical derived and compound words. It shows how the resulting morphologically tagged lexicon is enriched by semantic relations leading both to the synthesis of pseudo-definitions and to the constitution of classes of synonyms, hypo- and hypernyms.

Algorithms↗

A comparative study on concept representation between the UMLS and the clinical terms in Korean medical records.

The Unified Medical Language System (UMLS) is a rich source of knowledge in the biomedical domain. In this paper, we evaluated the coverage of UMLS as compared with Korean medical terms and identified differences in concept representation between two vocabulary sets. We measured the concept coverage by mapping clinical terms extracted from the discharge records of Seoul National University Hospital (SNUH) and the UMLS "Sign or Symptom" and "Disease or Syndrome" concepts. Thirty-five percent of the complaints in the SNUH were conceptually matched with the UMLS "Sign or Symptom" concept. Fifty-eight percent were found to be matched with the UMLS "Disease or Syndrome" concept rather than the "Sign or Symptom" concept. The remaining seven percent were not found in the UMLS concepts mapped above or those terms that used to special circumstances of a tertiary hospital in Korea. We then analyzed some of different expression patterns used by the two vocabulary sets and addressed issues to be taken into consideration.

Humans↗

Automatic generation of repeated patient information for tailoring clinical notes.

Dictating clear, readable, and accurate clinical notes can be a time-consuming task for physicians. Clinical notes often contain information concerning the patient's medical history and current medical condition which is propagated from one clinical note to all follow-up clinical notes for the same patient. In this paper, we present a system which, given a clinical note, automatically determines what information should be repeated, and then generates this information for the physician for a new clinical note. We use semantic patterns for capturing the rhetorical category of sentences, which we show to be useful for determining whether the sentence should be repeated. Our system is shown to perform better than a baseline metric based on precision/recall results. Such a system would allow clinical notes to be more complete, timely, and accurate.

Humans↗

Failure analysis of MetaMap Transfer (MMTx).

A pilot study was conducted to evaluate the performance of the MetaMap Transfer (MMTx), a tool that extracts terms from free text and suggests matches to concepts in the Unified Medical Language System' (UMLS'). Five participants, including a content domain expert and a UMLS Expert, manually extracted and mapped terms to UMLS concepts for two disease summary documents from NLM's consumer health site, Genetic Home Reference. The resulting adjudicated annotations were used as a gold standard. Differences in auto-mated term extraction and mapping between MMTx and MetaMap were noted. A failure analysis was conducted to categorize the types of terms not correctly mapped by MMTx. The most frequent type of failure (30%) resulted from missing inferential or world knowledge. Characteristics of each category are discussed. We distinguish between classes of failures that may be easily rectified, such as alternative retrieval strategies to extract exact matches, and ones that require additional research, such as coordinating conjunctions, co-reference resolution, and word sense disambiguation

Abstracting and Indexing↗

The medical appointment scheduler.

In order to enhance our understanding of how to best improve patient care, it was necessary to initially identify key questions that patients ask most frequently by recording actual calls made by home hemodialysis patients to a dialysis clinic over a three-month period. From an initial analysis of the recorded conversations, one of the most frequent reasons for patient calls was for scheduling concerns. Since the use of automated systems often enable a quicker response to patients' needs and alleviate time demands on the personnel that have to answer these calls, an automated scheduling system was designed and implemented to address this concern. The SCHEDULER is a mixed-initiative spoken dialogue system that allows a patient to schedule an appointment with a provider over the telephone. The system was implemented using SPEECHBUILDER, a utility developed by the Spoken Language System group at M.I.T. that helps automatically configure human language technology servers to create a new conversational system. We describe the system architecture, general considerations in the design, and a preliminary evaluation of the system.

Appointments and Schedules↗

Inter-document coreference resolution of abnormal findings in radiology documents.

In the clinical environment, it is often necessary to track the progression of a condition or various pertinent findings over time. Establishing automatic mechanisms for tracking pertinent findings can aid in the management of a condition as well as provide feedback for treatment outcomes assessment. This work focuses on the challenge of correlating observation of pertinent findings, specifically lung masses, across documents from serial computed tomography examinations for lung cancer patients. A probabilistic model is presented to characterize the likeliness of two observed findings from different documents referring to the same entity. A greedy algorithm is also presented that utilizes the probabilistic model to establish coreference links between findings. Results from a preliminary evaluation of this methodology show a precision of 72% and a recall of 63% for the described inter-document coreference resolution task.

Algorithms↗

General-purpose search techniques for genomic text.

Fast and accurate techniques for searching large genomic text collections are becoming increasingly important. While Information Retrieval is well-established for general-purpose text retrieval tasks, less is known about retrieval techniques for genomic text data. In this paper, we investigate and propose general-purpose search techniques for genomic text. In particular, we show that significant improvements can result from manual term expansion, where additional words are added to queries and documents. We also show that collection partitioning, where documents are included in or excluded from the search space, is highly effective for some tasks. We experiment with our techniques on four text collections and show, for example, that the collection partitioning scheme can improve effectiveness by almost 9.5% over a standard retrieval baseline. We conclude by recommending techniques that can be considered for most genomic search tasks.

Database Management Systems↗

A literature based method for identifying gene-disease connections.

We present a statistical method that can swiftly identify, from the literature, sets of genes known to be associated with given diseases. It offers a comprehensive way to treat alias symbols, a statistical method for computing the relevance of the gene to the query, and a novel way to disambiguate gene symbols from other abbreviations. The method is illustrated by finding genes related to breast cancer.

Abstracting and Indexing↗

Advances in text analytics for drug discovery.

The automated extraction of biological and chemical information has improved over the past year, with advances in access to content, entity extraction of genes, chemicals, kinetic data and relationships, and algorithms for generating and testing hypotheses. As the systems for reading and understanding scientific literature grow more powerful, so must the infrastructure in which to assemble information. Advances in infrastructure systems are discussed in this review. Research efforts have flourished as a result of text analytics competitions that attract participants from various disciplines, from computer science to bioinformatics.

Animals↗

Large-scale extraction of gene regulation for model organisms in an ontological context.

This paper presents an approach using syntactosemantic rules for the extraction of relational information from biomedical abstracts. The results show that by overcoming the hurdle of technical terminology, high precision results can be achieved. From abstracts related to baker's yeast, we manage to extract a regulatory network comprised of 441 pairwise relations from 58,664 abstracts with an accuracy of 83 - 90%. To achieve this, we made use of a resource of gene/protein names considerably larger than those used in most other biology related information extraction approaches. This list of names was included in the lexicon of our retrained partof- speech tagger for use on molecular biology abstracts. For the domain in question an accuracy of 93.6 - 97.7% was attained on Part-of-speech-tags. The method can be easily adapted to other organisms than yeast, allowing us to extract many more biologically relevant relations. The main reason for the comparable precision rates is the ontological model that was built beforehand and served as a guiding force for the manual coding of the syntactosemantic rules.

Abstracting and Indexing↗

Refinement of an automatic method for indexing medical literature--a preliminary study.

OBJECTIVES: to rank according to their significance MeSH terms automatically extracted from Internet sites in the framework of a French project, VUMeF, a contribution to the NLM' UMLS project. MATERIAL AND METHODS: scores are affected to key-words of a given document on the basis of the Semantic Network of the UMLS and frequencies of co-occurring major terms in the Medline literature. If N is the number of major terms of a document, and n is the number of major terms retrieved in the N first terms ranked in descending order according to their scores, the measure of the achievement of the method is n/N. RESULTS: a set of 1444 randomized documents have been extracted from Medline. For each document we computed the retrieved major terms among the first N terms with two methods: a statistical method using only frequencies given by co-occurrences, and our method that uses furthermore the UMLS semantic network. In 34% of cases corresponding to documents indexed by about 16 key-words, about 3 major terms among them, our method produces a better precision (7%) than the statistical method. DISCUSSION: the rough calculation of the proportion of retrieved major terms should be enhanced by the use of a probability law allowing to enlarge the list of terms to select taking into account both the number of major terms and the total number of key-words used to index each document.

Humans↗

GALEN Based Formal Representation of ICD10.

The authors present a formal representation of ICD10 based on GALEN CRM. The goal of the work is to create a coding support tool for coding clinical diagnoses to ICD10. The formal representation of the first two chapters of ICD10 has been almost completed. The paper presents the main aspects of the modelling, and the experienced problems. The constructed ontology has been converted to OWL, and a test system has been implemented in Prolog to verify the feasibility of the approach. The system successfully identified diseases in medical records from gastrointestinal oncology. The classifier module is still under development.

Clinical Coding↗