PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Why discourse structures in medical reports matter for the validity of automatically generated text knowledge bases.

The automatic analysis of medical full-texts currently suffers from neglecting text coherence phenomena such as reference relations between discourse units. This has unwarranted effects on the description adequacy of medical knowledge bases automatically generated from texts. The resulting representation bias can be characterized in terms of artificially fragmented, incomplete and invalid knowledge structures. We discuss three types of textual phenomena (pronominal and nominal anaphora, as well as textual ellipsis) and outline basic methodologies how to deal with them.

Artificial Intelligence↗

The role of compositionality in standardized problem list generation.

Compositionality is the ability of a Vocabulary System to record non-atomic strings. In this manuscript we define the types of composition, which can occur. We will then propose methods for both server based and client-based composition. We will differentiate the terms Pre-Coordination, Post-Coordination, and User-Directed Coordination. A simple grammar for the recording of terms with concept level identification will be presented, with examples from the Unified Medical Language System's (UMLS) Metathesaurus. We present an implementation of a Window's NT based client application and a remote Internet Based Vocabulary Server, which makes use of this method of compositionality. Finally we will suggest a research agenda which we believe is necessary to move forward toward a more complete understanding of compositionality. This work has the promise of paving the way toward a robust and complete Problem List Entry Tool.

Humans↗

Comparing expert systems for identifying chest x-ray reports that support pneumonia.

We compare the performance of four computerized methods in identifying chest x-ray reports that support acute bacterial pneumonia. Two of the computerized techniques are constructed from expert knowledge, and two learn rules and structure from data. The two machine learning systems perform as well as the expert constructed systems. All of the computerized techniques perform better than a baseline keyword search and a lay person, and perform as well as a physician. We conclude that machine learning can be used to identify chest x-ray reports that support pneumonia.

Algorithms↗

Classification algorithms applied to narrative reports.

Narrative text reports represent a significant source of clinical data. However, the information stored in these reports is inaccessible to many automated decision support systems. Data mining techniques can assist in extracting information from narrative data. Multiple classification methods, such as rule generation, decision trees, Bayesian classifiers, and information retrieval were used to classify a set of 200 chest X-ray reports according to 6 clinical conditions indicated. A general-purpose natural language processor was used to convert the narrative text into a coded form that could be used by the classification algorithms. Significant differences in performance were found between algorithms. The best performing algorithm applied to the processor output was significantly better than information retrieval applied to raw text. Predictor variables from the coded processor output were limited to avoid overfitting. Methods that limited by domain knowledge performed significantly better than those that limited by conditional probabilities of the variables in the training set. Algorithms were also shown to be dependent on training set size.

Algorithms↗

Extracting noun phrases for all of MEDLINE.

A natural language parser that could extract noun phrases for all medical texts would be of great utility in analyzing content for information retrieval. We discuss the extraction of noun phrases from MEDLINE, using a general parser not tuned specifically for any medical domain. The noun phrase extractor is made up of three modules: tokenization; part-of-speech tagging; noun phrase identification. Using our program, we extracted noun phrases from the entire MEDLINE collection, encompassing 9.3 million abstracts. Over 270 million noun phrases were generated, of which 45 million were unique. The quality of these phrases was evaluated by examining all phrases from a sample collection of abstracts. The precision and recall of the phrases from our general parser compared favorably with those from three other parsers we had previously evaluated. We are continuing to improve our parser and evaluate our claim that a generic parser can effectively extract all the different phrases across the entire medical literature.

Linguistics↗

A statistical natural language processor for medical reports.

Statistical natural language processors have been the focus of much research during the past decade. The main advantage of such an approach over grammatical rule-based approaches is its scalability to new domains. We present a statistical NLP for the domain of radiology and report on methods of knowledge acquisition, parsing, semantic interpretation, and evaluation. Preliminary performance data are given. A discussion of the perceived benefit, limitations and future work is presented.

Algorithms↗

Synapses/SynEx goes XML.

This paper describes the first approach to use the Extensible Markup Language (XML) as a data interchange/delivery format in the Synapses and SynEx healthcare environment. It will ease the semantic mapping between healthcare systems (server-server, server-client) and make them more interoperable.

Computer Systems↗

Requirements for speech recognition to support medical documentation.

Recent advances in the development of automated speech recognition (ASR) have made routine applications for medical documentation possible. To achieve this, ASR has to be optimally integrated into the specific documentation scenario. The classification presented in this paper allows the definition of specification requirements. For two different documentation scenarios the appropriate product selection has been done according to this classification. Two evaluation studies are presented, addressing the usefulness of applying automated speech recognition.

Data Collection↗

Clinical terminology: why is it so hard?

Despite years of work, no re-usable clinical terminology has yet been demonstrated in widespread use. This paper puts forward ten reasons why developing such terminologies is hard. All stem from underestimating the change entailed in using terminology in software for 'patient centred' systems rather than for its traditional functions of statistical and financial reporting. Firstly, the increase in scale and complexity are enormous. Secondly, the resulting scale exceeds what can be managed manually with the rigour required by software, but building appropriate rigorous representations on the necessary scale is, in itself, a hard problem. Thirdly, 'clinical pragmatics'--practical data entry, presentation and retrieval for clinical tasks--must be taken into account, so that the intrinsic differences between the needs of users and the needs of software are addressed. This implies that validation of clinical terminologies must include validation in use as implemented in software.

Medical Informatics↗

GALEN-IN-USE: application in Greek and influences on education.

GALEN-IN-USE is a European project that aims to promote greater European harmonization and to overcome the problems encountered in using traditional coding and classification systems. This paper presents the work done by the Greek Centre of Medical Informatics and Terminology, as a collaborating centre of GALEN-IN-USE(GIU), in order to apply GIU's tools to Greek Health Care System as well as the affect of this application in education.

Classification↗

The NLM Indexing Initiative.

The objective of NLM's Indexing Initiative (IND) is to investigate methods whereby automated indexing methods partially or completely substitute for current indexing practices. The project will be considered a success if methods can be designed and implemented that result in retrieval performance that is equal to or better than the retrieval performance of systems based principally on humanly assigned index terms. We describe the current state of the project and discuss our plans for the future.

Abstracting and Indexing↗

Using UMLS semantics for classification purposes.

The Unified Medical Language System (UMLS) contains semantic information about terms from various sources; each concept can be understood and located by its relationships to other concepts. We describe a method in which the semantic relationships between UMLS concepts are exploited for the purpose of classification. This method combines three existing components: 1) Mapping terms to UMLS concepts; 2) Restricting UMLS concepts to MeSH; and 3) Mapping MeSH terms to disease categories. When applied to the automatic classification of condition terms into broad disease categories in the Clinical Trials database, this method assigned relevant categories to 92% of the 1823 condition terms encountered. 135 (7%) failed to be classified and 14 (.77%) were misclassified. The limits of this method are discussed, as well as the reuse of existing components, and the tuning required to achieve automatic classification.

Algorithms↗

Contribution of a speech recognition system to a computerized pneumonia guideline in the emergency department.

OBJECTIVE: Evaluate the effect of a radiology speech recognition system on a real-time computerized guideline in the emergency department. METHODS: We collected all chest x-ray reports (n = 727) generated for patients in the emergency department during a six-week period. We divided the concurrently generated reports into those generated with speech recognition and those generated by traditional dictation. We compared the two sets of reports for availability during the patient's emergency department encounter and for readability. RESULTS: Reports generated by speech recognition were available seven times more often during the patients' encounters than reports generated by traditional dictation. Using speech recognition reduced the turnover time of reports from 12 hours 33 minutes to 2 hours 13 minutes. Readability scores were identical for both kinds of reports. CONCLUSION: Using speech recognition to generate chest x-ray reports reduces turnover time so reports are available while patients are in the emergency department.

Decision Making, Computer-Assisted↗

MedSynDiKATe--design considerations for an ontology-based medical text understanding system.

MedSynDiKATe is a natural language processor for automatically acquiring knowledge from medical finding reports. The content of these documents is transferred to formal representation structures which constitute a corresponding text knowledge base. The general system architecture we present integrates requirements from the analysis of single sentences, as well as those of referentially linked sentences forming cohesive texts. The strong demands MedSynDiKATe poses to the availability of expressive knowledge sources are accounted for by two alternative approaches to (semi)automatic ontology engineering.

Evaluation Studies as Topic↗

Ontology acquisition from on-line knowledge sources.

Electronic knowledge representation is becoming more and more pervasive both in the form of formal ontologies and less formal reference vocabularies, such as UMLS. The developers of clinical knowledge bases need to reuse these resources. Such reuse requires a new generation of tools for ontology development and management. Medical experts with little or no computer science experience need tools that will enable them to develop knowledge bases and provide capabilities for directly importing knowledge not only from formal knowledge bases but also from reference terminologies. The portions of knowledge bases that are imported from disparate resources then need to be merged or aligned to one another in order to link corresponding terms, to remove redundancies, to resolve logical conflicts. We discuss the requirements for ontology-management tools that will enable interoperability of disparate knowledge sources. Our group is developing a suite of tools for knowledge-base management based on the Protégé-2000 environment for ontology development and knowledge acquisition. We describe one such tool in detail here: an application for incorporating information from remote knowledge sources such as UMLS into a Protégé knowledge base.

Artificial Intelligence↗

Argument identification for arterial branching predications asserted in cardiac catheterization reports.

The language describing coronary vasculature provides a suitable paradigm for research in semantic interpretation of anatomical text. As a pilot project we investigate the possibility of highly accurate retrieval of arterial branching relationships asserted in cardiac catheterization reports. Our methodology relies on the cooperation of underspecified linguistic analysis and structured domain knowledge. The satisfactory results of formal evaluation on both a training and testing set support the promise of this approach.

Cardiac Catheterization↗