PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Having our cake and eating it too: how the GALEN Intermediate Representation reconciles internal complexity with users' requirements for appropriateness and simplicity.

Clinical terminologies are complex objects, getting more complex as the requirements on them grow, and as more complex technologies are used in their construction. But to the clinical end-user, functionality and utility is important, not inherent complexity--the simpler a clinical terminology can be for the end-user, the better. To reconcile these contradictory requirements, the GALEN Programme has developed an Intermediate Representation that allows the OpenGALEN Clinical Terminology to retain a high degree of internal complexity, whilst allowing it to be efficiently maintained, and easily used. This paper describes the elements of the Intermediate Representation, how it works, and some experience of its use.

Abstracting and Indexing↗

Medical text representations for inductive learning.

Inductive learning algorithms have been proposed as methods for classifying medical text reports. Many of these proposed techniques differ in the way the text is represented for use by the learning algorithms. Slight differences can occur between representations that may be chosen arbitrarily, but such differences can significantly affect classification algorithm performance. We examined 8 different data representation techniques used for medical text, and evaluated their use with standard machine learning algorithms. We measured the loss of classification-relevant information due to each representation. Representations that captured status information explicitly resulted in significantly better performance. Algorithm performance was dependent on subtle differences in data representation.

Algorithms↗

Tagging medical texts: a rule-based experiment.

In this paper we describe the construction of a part-of-speech tagger for medical document retrieval purposes, therefore we have designed a specific architecture called minimal commitment. The system uses local grammatical rules for conducting the disambiguation task. Four evaluations are conducted, with and without taking unknown words into account. In between each evaluation the modules (lexicon, guesser, rules) of the system are incrementally improved.

Disease↗

Good isn't enough.

Explore the source record for details and available documents.

Cost-Benefit Analysis↗

The connection between terms used in medical records and coding system: a study on Swedish primary health care data.

Implementation of problem lists and their relation to standardized coding systems have been approached and analysed in different ways. Most evaluations concern quantitative aspects such as content coverage in a specific domain. In order to reveal the qualitative aspects of diagnostic coding, medical record texts from primary health care encounters were compared with terms from a coding system that was used for describing them statistically. The records were coded by six general practitioners, and in some cases, an applied diagnostic term was found within the text, while other record text-coding system relationships were categorized as synonyms, alternative terms, and interpretations. Thus, the categories roughly corresponded to a measure of semantic distance between the terms in the record text and the rubrics of the coding system, and there was a correlation between semantic distance and inter-rater agreement. The subcategories of this scheme corresponded fairly well to recently published desiderata for clinical terminology servers, including functionality such as word normalization and spelling correction. However, not all problems could have been automatically coded by means of lexical methods, which can be partly explained by the fact that diagnostic coding also relies on clinical knowledge. In addition, proper automation relies on context representation within the records.

Abstracting and Indexing↗

Discovery of association rules in medical data.

Data mining is a technique for discovering useful information from large databases. This technique is currently being profitably used by a number of industries. A common approach for information discovery is to identify association rules which reveal relationships among different items. In this paper, we use this approach to analyse a large database containing medical-record data. Our aim is to obtain association rules indicating relationships between procedures performed on a patient and the reported diagnoses. Random sampling was used to obtain these association rules. After reviewing the basic concepts associated with data mining, we discuss our approach for identifying association rules and report on the rules generated.

Algorithms↗

Semantic features of an enterprise interface terminology for SNOMED RT.

OBJECTIVE: To evaluate the utility of SNOMED RT in support of a natural language interface for encoding of clinical assessments. METHOD: Using a random sample of clinical terms from the UNMC Lexicon, I mapped the terminology into canonical data entries using SNOMED RT. Working from the source term language, I evaluated lexical mapping to the SNOMED term set, and the function of the SNOMED RT semantic network in support of a language-based clinical coding interface. RESULTS: Ambiguity in the source terms was low at 0.3%. Lexical (language-based) mapping could account for only 48.8% of meaning from the source terms. The RT semantic network accounted for 39.5% of meaning, and supplementing the lexical map this led to 80.2% capture of source content. Error rates in the segment of RT which I reviewed were low at 0.6%. 97.6% of source content could be accurately captured in SNOMED RT. CONCLUSION: SNOMED RT supported an accurate and reliable representation of clinical assessment data in this sample. The semantic network of RT substantially enhanced the encoding of concepts relative to lexical mapping. However these data suggest that natural language encoding with SNOMED RT in an enterprise environment is unlikely at this time.

Natural Language Processing↗

An approach to guideline implementation with GEM.

Implementation of practice guidelines refers to the creation of strategies and systems to operationalize the knowledge and recommendations set forth by guideline developers. We describe an approach to guideline implementation that makes direct use of the guideline document as a knowledge base. The Guideline Elements Model (GEM) provides an XML-based guideline document model that facilitates implementation of guidelines. Knowledge extraction using GEM requires document markup rather than programming and can promote authenticity and consistent knowledge encoding. Knowledge customization for the local enterprise requires addition of meta-information to pertinent components of the GEM hierarchy in a design database. GEM provides an audit trail to track local adaptation. Knowledge integration with patient data can be promoted using information management services. A design goal is to devise a system that can be applied by local clinical domain experts, quality assurance experts, and information systems programmers without requiring trained informaticians and knowledge engineers to serve as intermediaries

Artificial Intelligence↗

Intelligent system for topic survey in MEDLINE by keyword recommendation and learning text characteristics.

We have implemented a system for assisting experts in selecting MEDLINE records for database construction purposes. This system has two specific features: The first is a learning mechanism which extracts characteristics in the abstracts of MEDLINE records of interest as patterns. These patterns reflect selection decisions by experts and are used for screening the records. The second is a keyword recommendation system which assists and supplements experts' knowledge in unexpected cases. Combined with a conventional keyword-based information retrieval system, this system may provide an efficient and comfortable environment for MEDLINE record selection by experts. Some computational experiments are provided to prove that this idea is useful.

Artificial Intelligence↗

The contribution of morphological knowledge to French MeSH mapping for information retrieval.

MeSH-indexed Internet health directories must provide a mapping from natural language queries to MeSH terms so that both health professionals and the general public can query their contents. We describe here the design of lexical knowledge bases for mapping French expressions to MeSH terms, and the initial evaluation of their contribution to Doc'CISMeF, the search tool of a MeSH-indexed directory of French-language medical Internet resources. The observed trend is in favor of the use of morphological knowledge as a moderate (approximately 5%) but effective factor for improving query to term mapping capabilities.

Algorithms↗

A light knowledge model for linguistic applications.

Content extraction from medical texts is achievable today by linguistic applications, in so far as sufficient domain knowledge is available. Such knowledge represents a model of the domain and is hard to collect with sufficient depth and good coverage, despite numerous attempts. To leverage this task is a priority in order to benefit from the awaited linguistic tools. The light model is designed with this goal in mind. Syntactic and lexical information are generally available with large lexicons. A domain model should add the necessary semantic information. The authors have designed a light knowledge model for the collection of semantic information on the basis of the recognized syntactical and lexical attributes. It has been tailored for the acquisition of enough semantic information in order to retrieve terms of a controlled vocabulary from free texts, as for example, to retrieve Mesh terms from patient records.

Information Storage and Retrieval↗

Speech recognition systems. Are they up to the task?

Clinical speech recognition systems record speech input from a physician and translate the acoustic data into text output that forms the basis of medical reports. These reports can then be edited by physicians either during or after dictation. Speech recognition systems have received a lot of attention as potential time-savers. But when we looked at four of these systems, we found their accuracy to be low enough that it will take most physicians 20% to 30% more time to produce a report than is required using traditional transcription methods. Even so, these systems offer enough benefits that you should at least investigate their use. For one thing, they're considerably less expensive than using a transcription service. For another, despite the additional physician time required, a report can be turned around far more quickly than with transcription--a matter of hours rather than days. With proper planning, most facilities should be able to implement these systems successfully.

Equipment Design↗

The German specialist lexicon.

The German language and in particular biomedical terms exhibit a rich and productive morphology. Beyond inflection and comparison forms frequently spelling variants, German - Greek/Latin synonyms and nominal compounds exist. For the English language, the SPECIALIST LEXICON, part of the UMLS project, covers a broad range of biomedical terms. In this paper we describe the database model and the functionality of the GERMAN SPECIALIST LEXICON, an ongoing project to develop a lexical resource for German-language medical terminology. Similar to the SPECIALIST LEXICON it is accompanied by tools for the recognition and generation of lexical variants, as well as by databases linking synonymous words, spelling variants, phrases and abbreviations.

Databases as Topic↗

Text generation in clinical medicine--a review.

OBJECTIVE: This article aims at an analysis of ways of producing documents (such as findings or referral letters) in clinical medicine. Special emphasis is given to the question of whether the field of "Natural Language Generation" (NLG) can provide new approaches to ameliorate the current situation. METHODS: In order to assess the currently used techniques in text production, an analysis of commercially available systems was performed in addition to an extensive review of the literature. The sketch of current NLG approaches is also based on a literature review. To estimate the applicability of several techniques to clinical documents, a typology of documents in clinical medicine was developed, based on rhetorical structure theory, speech act theory and certain recurrent linguistic phenomena exposed in the said documents. RESULTS: Current ways of producing text for documents in medicine are less than optimal in several respects. The field of NLG draws on the idea of generating text from a conceptual representation of not only certain facts, but also knowledge about how to express them via (written) language. Unfortunately, NLG does not yet offer "ready-to-run" solutions for the automatic production of most of the document types in the given typology. It seems, however, highly plausible that the demands of medical informatics for these kinds of systems will be satisfiable as NLG matures. CONCLUSIONS: NLG offers a promising way of generating text for clinical documents, a problem of enormous economical importance. The medical informatics community should therefore commit itself to the idea of NLG in medicine.

Communication↗

Using cross-lingual information to cope with underspecification in formal ontologies.

Description logics and other formal devices are frequently used as means for preventing or detecting mistakes in ontologies. Some of these devices are also capable of inferring the existence of inter-concept relationships that have not been explicity entered into an ontology. A prerequisite, however, is that this information can be derived from those formal definitions of concepts and relationships which are included within the ontology. In this paper, we present a novel algorithm that is able to suggest relationships among existing concepts in a formal ontology that are not derivable from such formal definitions. The algorithm exploits cross-lingual information that is implicity present in the collection of terms used in various languages to denote the concepts and relationships at tissue. By using a specific experimental design, we are able to quantify the impact of cross-lingual information in coping with underspecification in formal ontologies.

Algorithms↗