PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A continuous-speech interface to a decision support system: II. An evaluation using a Wizard-of-Oz experimental paradigm.

OBJECTIVE: Evaluate the performance of a continuous-speech interface to a decision support system. DESIGN: The authors performed a prospective evaluation of a speech interface that matches unconstrained utterances of physicians with controlled-vocabulary terms from Quick Medical Reference (QMR). The performance of the speech interface was assessed in two stages: in the real-time experiment, physician subjects viewed audiovisual stimuli intended to evoke clinical findings, spoke a description of each finding into the speech interface, and then chose from a list generated by the interface the QMR term that most closely matched the finding. Subjects believed that the speech recognizer decoded their utterances; in reality, a hidden experimenter typed utterances into the interface (Wizard-of-Oz experimental design). Later, the authors replayed the same utterances through the speech recognizer and measured how accurately utterances matched with appropriate QMR terms using the results of the real-time experiment as the "gold standard." MEASUREMENTS: The authors measured how accurately the speech-recognition system converted input utterances to text strings (recognition accuracy) and how accurately the speech interface matched input utterances to appropriate QMR terms (semantic accuracy). RESULTS: Overall recognition accuracy was less than 50%. However, using language-processing techniques that match keywords in recognized utterances to keywords in QMR terms, the semantic accuracy of the system was 81%. CONCLUSIONS: Reasonable semantic accuracy was attained when language-processing techniques were used to accommodate for speech misrecognition. In addition, the Wizard-of-Oz experimental design offered many advantages for this evaluation. The authors believe that this technique may be useful to future evaluators of speech-input systems.

Adolescent↗

UMLS language and vocabulary tools.

A variety of resources developed for use with the Unified Medical Language System are presented. These resources include the UMLS Knowledge Source Server, the SPECIALIST lexicon, a set of lexical tools that work with the SPECIALIST lexicon, and a variety of other NLP document processing tools. These tools manage lexical variation, tokenize and parse text strings, suggest spelling variants, and provide text-to-concept mapping capabilities. The UMLS Knowledge Source Server is available under a license agreement. The other tools are freely downloadable.

Abstracting and Indexing↗

A method for vocabulary development and visualization based on medical language processing and XML.

A comprehensive controlled clinical vocabulary is critical to the effectiveness of many automated clinical systems. Vocabulary development and maintenance is an important aspect of a vocabulary, and should be linked to terms physicians actually use. This paper presents a method to help vocabulary builders capture, visualize, and analyze both compositional and quantitative information related to terms physicians use. The method includes several components: an MLP system, a corpus of relevant reports and a visualization tool based on XML and JAVA.

Humans↗

The use of concept maps during knowledge elicitation in ontology development processes--the nutrigenomics use case.

BACKGROUND: Incorporation of ontologies into annotations has enabled 'semantic integration' of complex data, making explicit the knowledge within a certain field. One of the major bottlenecks in developing bio-ontologies is the lack of a unified methodology. Different methodologies have been proposed for different scenarios, but there is no agreed-upon standard methodology for building ontologies. The involvement of geographically distributed domain experts, the need for domain experts to lead the design process, the application of the ontologies and the life cycles of bio-ontologies are amongst the features not considered by previously proposed methodologies. RESULTS: Here, we present a methodology for developing ontologies within the biological domain. We describe our scenario, competency questions, results and milestones for each methodological stage. We introduce the use of concept maps during knowledge acquisition phases as a feasible transition between domain expert and knowledge engineer. CONCLUSION: The contributions of this paper are the thorough description of the steps we suggest when building an ontology, example use of concept maps, consideration of applicability to the development of lower-level ontologies and application to decentralised environments. We have found that within our scenario conceptual maps played an important role in the development process.

Algorithms↗

RelEx--relation extraction using dependency parse trees.

MOTIVATION: The discovery of regulatory pathways, signal cascades, metabolic processes or disease models requires knowledge on individual relations like e.g. physical or regulatory interactions between genes and proteins. Most interactions mentioned in the free text of biomedical publications are not yet contained in structured databases. RESULTS: We developed RelEx, an approach for relation extraction from free text. It is based on natural language preprocessing producing dependency parse trees and applying a small number of simple rules to these trees. We applied RelEx on a comprehensive set of one million MEDLINE abstracts dealing with gene and protein relations and extracted approximately 150,000 relations with an estimated performance of both 80% precision and 80% recall. AVAILABILITY: The used natural language preprocessing tools are free for use for academic research. Test sets and relation term lists are available from our website (http://www.bio.ifi.lmu.de/publications/RelEx/).

Algorithms↗

A temporal constraint structure for extracting temporal information from clinical narrative.

INTRODUCTION: Time is an essential element in medical data and knowledge which is intrinsically connected with medical reasoning tasks. Many temporal reasoning mechanisms use constraint-based approaches. Our previous research demonstrates that electronic discharge summaries can be modeled as a simple temporal problem (STP). OBJECTIVE: To categorize temporal expressions in clinical narrative text and to propose and evaluate a temporal constraint structure designed to model this temporal information and to support the implementation of higher-level temporal reasoning. METHODS: A corpus of 200 random discharge summaries across 18 years was applied in a grounded approach to construct a representation structure. Then, a subset of 100 discharge summaries was used to tally the frequency of each identified time category and the percentage of temporal expressions modeled by the structure. Fifty random expressions were used to assess inter-coder agreement. RESULTS: Six main categories of temporal expressions were identified. The constructed temporal constraint structure models time over which an event occurs by constraining its starting time and ending time. It includes a set of fields for the endpoint(s) of an event, anchor information, qualitative and metric temporal relations, and vagueness. In 100 discharge summaries, 1961 of 2022 (97%) identified temporal expressions were effectively modeled using the temporal constraint structure. Inter-coder evaluation of 50 expressions yielded exact match in 90%, partial match with trivial differences in 8%, partial match with large differences in 2%, and total mismatch in 0%. CONCLUSION: The proposed temporal constraint structure embodies a sufficient and successful implementation method to encode the diversity of temporal information in discharge summaries. Placing data within the structure provides a foundational representation upon which further reasoning, including the addition of domain knowledge and other post-processing to implement an STP, can be accomplished.

Database Management Systems↗

BarleyExpress: a web-based submission tool for enriched microarray database annotations.

UNLABELLED: BarleyExpress is a web-based microarray experiment data submission tool for BarleyBase, a public data resource of Affymetrix GeneChip data for plants. BarleyExpress uses the Plant Ontology vocabularies and enhances the MIAME guidelines to standardize the annotation of microarray gene expression experiments. In addition, BarleyExpress provides explicit support for factorial experiment design and template loading methods to ease the submission process for large experiments. AVAILABILITY: http://barleybase.org SUPPLEMENTARY INFORMATION: BarleyExpress Users Manual.

Database Management Systems↗

Fast parsers for Entrez Gene.

NCBI completed the transition of its main genome annotation database from Locuslink to Entrez Gene in Spring 2005. However, to this date few parsers exist for the Entrez Gene annotation file. Owing to the widespread use of Locuslink and the popularity of Perl programming language in bioinformatics, a publicly available high performance Entrez Gene parser in Perl is urgently needed. We present four such parsers that were developed using several parsing approaches (Parse::RecDescent, Parse::Yapp, Perl-byacc and Perl 5 regular expressions) and provide the first in-depth comparison of these sophisticated Perl tools. Our fastest parser processes the entire human Entrez Gene annotation file in under 12 min on one Intel Xeon 2.4 GHz CPU and can be of help to the bioinformatics community during and after the transition from Locuslink to Entrez Gene.

Algorithms↗

The usability axiom of medical information systems.

INTRODUCTION: In this article we begin by connecting the concept of simplicity of user interfaces of information systems with that of usability, and the concept of complexity of the problem-solving in information systems with the concept of usefulness. We continue by stating "the usability axiom" of medical information technology: information systems must be, at the same time, usable and useful. We then try to show why, given existing technology, the axiom is a paradox and we continue with analysing and reformulating it several times, from more fundamental information processing perspectives. DISCUSSION: We underline the importance of the concept of representation and demonstrate the need for context-dependent representations. By means of thought experiments and examples, we advocate the need for context-dependent information processing and argue for the relevance of algorithmic information theory and case-based reasoning in this context. Further, we introduce the notion of concept spaces and offer a pragmatic perspective on context-dependent representations. We conclude that the efficient management of concept spaces may help with the solution to the medical information technology paradox. Finally, we propose a view of informatics centred on the concepts of context-dependent information processing and management of concept spaces that aligns well with existing knowledge centric definitions of informatics in general and medical informatics in particular. In effect, our view extends M. Musen's proposal and proposes a definition of Medical Informatics as context-dependent medical information processing. SUMMARY: The axiom that medical information systems must be, at the same time, useful and usable, is a paradox and its investigation by means of examples and thought experiments leads to the recognition of the crucial importance of context-dependent information processing. On the premise that context-dependent information processing equates to knowledge processing, this view defines Medical Informatics as a context-dependent medical information processing which aligns well with existing knowledge centric definitions of our field.

Algorithms↗

Implementation and evaluation of a negation tagger in a pipeline-based system for information extract from pathology reports.

We have developed a pipeline-based system for automated annotation of Surgical Pathology Reports with UMLS terms that builds on GATE--an open-source architecture for language engineering. The system includes a module for detecting and annotating negated concepts, which implements the NegEx algorithm--an algorithm originally described for use in discharge summaries and radiology reports. We describe the implementation of the system, and early evaluation of the Negation Tagger. Our results are encouraging. In the key Final Diagnosis section, with almost no modification of the algorithm or phrase lists, the system performs with precision of 0.84 and recall of 0.80 against a gold-standard corpus of negation annotations, created by modified Delphi technique by a panel of pathologists. Further work will focus on refining the Negation Tagger and UMLS Tagger and adding additional processing resources for annotating free-text pathology reports.

Algorithms↗

Hospitexte: towards a document-based hypertextual electronic medical record.

The patient record is a repository for knowledge about a patient. Work in Artificial Intelligence and knowledge representation has evidenced the intrinsic difficulty of formalizing knowledge for computer processing. It is therefore not a surprise that most attempts at computerizing the patient record have only had a limited degree of success or applicability. We claim that this is due to the fact that medicine is an empirical domain, and thus fundamentally resists formalization. Therefore, the only way medical knowledge can be fully expressed is through natural languages which is indeed what clinicians actually use. We proposed and designed an electronic medical record which adheres to this hypothesis and where structured documents play a prominent role.

Artificial Intelligence↗

Molecular decomposition of complex clinical phenotypes using biologically structured analysis of microarray data.

MOTIVATION: Today, the characterization of clinical phenotypes by gene-expression patterns is widely used in clinical research. If the investigated phenotype is complex from the molecular point of view, new challenges arise and these have not been addressed systematically. For instance, the same clinical phenotype can be caused by various molecular disorders, such that one observes different characteristic expression patterns in different patients. RESULTS: In this paper we describe a novel algorithm called Structured Analysis of Microarrays (StAM), which accounts for molecular heterogeneity of complex clinical phenotypes. Our algorithm goes beyond established methodology in several aspects: in addition to the expression data, it exploits functional annotations from the Gene Ontology database to build biologically focussed classifiers. These are used to uncover potential molecular disease subentities and associate them to biological processes without compromising overall prediction accuracy. AVAILABILITY: Bioconductor compliant R package SUPPLEMENTARY INFORMATION: Complete analyses are available at http://compdiag.molgen.mpg.de/supplements/lottaz05.

Biomarkers, Tumor↗

The ANTHEM representation formalism for the alphabetic index of ICD.

This paper describes a formalism in which the knowledge of the Alphabetic Index of ICD is expressed unambiguously in the ANTHEM prototype. The Alphabetic Index may be viewed as a collection of hypotheses on which ICD-codes should be tried first for a given diagnostic expression. The hypotheses generation can be based upon characteristics of the semantic representation of the diagnostic expression. The formalism is described using an "operational" semantics by referring to the processes that have to operate on the expressions in the knowledge base. A close integration between a sound linguistic model of diagnostic expressions and the formalism itself is realized.

Diagnosis↗

Biomedical ontologies: what part-of is and isn't.

Mereological relations such as part-of and its inverse has-part are fundamental to the description of the structure of living organisms. Whereas classical mereology focuses on individual entities, mereological relations in biomedical ontologies are generally asserted between classes of individuals. In general, this practice leaves some basic issues unanswered: type constraints of mereological relations, e.g., concerning artifacts and biological entities, the relation between parthood and time, inferred parts and wholes as well as a delimitation of parthood against spatial inclusion. Furthermore, mereological relations can be asserted not only between physical objects but also between biological processes and medical procedures. We analyze these ambiguities and make suggestions for a standardization of mereological relations in biomedical ontologies.

Anatomy↗

Conceptual search in electronic patient record.

Search by content in a large corpus of free texts in the medical domain is, today, only partially solved. The so-called GREP approach (Get Regular Expression and Print), based on highly efficient string matching techniques, is subject to inherent limitations, especially its inability to recognize domain specific knowledge. Such methods oblige the user to formulate his or her query in a logical Boolean style; if this constraint is not fulfilled, the results are poor. The authors present an enhancement to string matching search by the addition of a light conceptual model behind the word lexicon. The new system accepts any sentence as a query and radically improves the quality of results. Efficiency regarding execution time is obtained at the expense of implementing advanced indexing algorithms in a pre-processing phase. The method is described and commented and a brief account of the results illustrates this paper.

Artificial Intelligence↗

Identifying respiratory findings in emergency department reports for biosurveillance using MetaMap.

Clinical conditions described in patients' dictated reports are necessary for automated detection of patients with respiratory illnesses such as inhalational anthrax and pneumonia. We applied MetaMap to emergency department reports to extract a set of 71 clinical conditions relevant to detection of a lower respiratory outbreak. We indexed UMLS terms in emergency department reports with MetaMap, filtered the indexed output with a specialized lexicon of UMLS terms for the domain, and mapped the clinical conditions of interest to concepts in the lexicon. We compared MetaMap's ability to accurately identify the conditions against a physician's manual annotations and evaluated incorrectly indexed features to determine what additional processing is necessary. MetaMap identified the clinical conditions with a recall of 0.72 and a precision of 0.56. Necessary processing beyond MetaMap's indexing includes finding validation, temporal discrimination, anatomic location discrimination, finding-disease discrimination, and contextual inference. Successful identification of clinical conditions in an emergency department report with MetaMap requires processing techniques specific to the clinical question of interest.

Abstracting and Indexing↗

A cooperative methodology to build conceptual models in medicine.

We designed a methodology to perform distribute activities on conceptual modelling among cooperating centers. Our methodology assigns responsibilities and tasks and regulates interactions preserving coherence; it passes through the construction of unambiguous paraphrases to make explicit the context within the original sources, and through their compositional representation in an intermediate language. The process is intrinsically iterative, with continuous feedbacks and refinements, alternating analytic view on details and synthetic view on regularities and structures. Our methodology is based on requirements and experience made in the first GALEN project, and was applied in the GALEN-IN-USE project to coordinate modelling activities of three teams of surgeons in Rome with activities of other partners, during the production of an extensive model of surgical procedures.

Humans↗

Knowledge management prerequisites for building an information society in healthcare.

The European Research Area requires either technological development or information literacy of health professionals. This information literacy shall be understood much deeper and broader than a basic preparation to use ICT tools in everyday life only. The author's first aim is to present the literature review and analysis of different definition of the "information" concept in Polish and foreign sources for health sciences, to emphasize a problem fundamental for an information society development, i.e. lack of adequate "information" understanding. Health professionals' information literacy shall also build an awareness of conceptual differences among numerous classifications, thesauruses, and information-retrieval languages, which result in different information received in a retrieval process. This problem can be of crucial effect for either health research or practice. Understanding the problem shall mobilize the researchers, classifiers, and indexers to co-ordinate efforts aimed in organizing "a translator" covering the most popular classifications' and thesauruses' concepts, to make an international research co-operation easier, relevant, and safe for the patients.

Artificial Intelligence↗