PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Application of a Medical Text Indexer to an online dermatology atlas.

Clinical dermatology cases are presented as images and semi-structured text describing skin lesions and their relationships to disease. Metadata assignment to such cases is hampered by lack of a standardized dermatology vocabulary and facilitated methods for indexing legacy collections. In this pilot study descriptive clinical text from Dermatlas, a Web-based repository of dermatology cases, was indexed to Medical Subject Heading (MeSH) terms using the National Library of Medicine's Medical Text Indexer (MTI). The MTI is an automated text processing system that derives ranked lists of MeSH terms to describe the content of medical journal citations using knowledge from the Unified Medical Language System (UMLS) and from MEDLINE. For a representative, random sample of 50 Dermatlas cases, the MTI frequently derived MeSH indexing terms that matched expert-assigned terms for Diagnoses (88%), Lesion Types (72%), and Patient Characteristics (Gender and Age Groups, 62% and 84% respectively). This pilot demonstrates the potential for extending the MTI to automate indexing of clinical case presentations and for using MeSH to describe aspects of clinical dermatology.

Abstracting and Indexing↗

A Dutch medical language processor.

This paper describes the current state of a medical language processor for Dutch. The goal is to implement a language specific front-end compatible with some existing applications that aim at the intelligent extraction and processing of information from patient discharge summaries. A complete chain for processing and understanding Dutch medical documents will be the ultimate result. The text focuses mainly on the language specific aspects of the language processing chain. Evaluation results of the already functioning components are given as well as an outline for future developments and enhancements. A short theoretical background is provided (cf. also [1-3]: Rossi Mori et al., Proc. SCAMC 90, 1990, pp. 185-189; Wingert, in: Informatics and Medicine, an advanced course, Springer-Verlag. 1977, pp. 579-646; Wingert, Proc. MEDINFO 80, 1980, pp. 1321-1331) before the description of each component in order to familiarise the non-experienced reader with the basic notions of computational linguistics.

Artificial Intelligence↗

Generating models of surgical procedures using UMLS concepts and multiple sequence alignment.

Surgical procedures can be viewed as a process composed of a sequence of steps performed on, by, or with the patient's anatomy. This sequence is typically the pattern followed by surgeons when generating surgical report narratives for documenting surgical procedures. This paper describes a methodology for semi-automatically deriving a model of conducted surgeries, utilizing a sequence of derived Unified Medical Language System (UMLS) concepts for representing surgical procedures. A multiple sequence alignment was computed from a collection of such sequences and was used for generating the model. These models have the potential of being useful in a variety of informatics applications such as information retrieval and automatic document generation.

Abstracting and Indexing↗

Using hit curves to compare search algorithm performance.

Databases continue to grow but the metrics available to evaluate information retrieval systems have not changed. Large collections such as MEDLINE and the World Wide Web contain many relevant documents for common queries. Ranking is therefore increasingly important and successful information retrieval systems, such as Google, have emphasized ranking. However, existing evaluation metrics such as precision and recall, do not directly account for ranking. This paper describes a novel way of measuring information retrieval performance using weighted hit curves adapted from the field of statistical detection to reflect multiple desirable characteristics such as relevance, importance, and methodologic quality. In statistical detection, hit curves have been proposed to represent occurrence of interesting events during a detection process. Similarly, hit curves can be used to study the position of relevant documents within large result sets. We describe hit curves in light of a formal model of information retrieval, show how hit curves represent system performance including ranking, and define ways to statistically compare performance of multiple systems using hit curves. We provide example scenarios where traditional measures are less suitable than hit curves and conclude that hit curves may be useful for evaluating retrieval from large collections where ranking performance is crucial.

Algorithms↗

Puya: a method of attracting attention to relevant physical findings.

Puya is a method that compares the physical exam in an electronic clinical note with a set of stereotypical physical exam sentences that have been previously classified as "normal". The note is then displayed in a web browser with normal findings clearly delineated. The list of stereotypical sentences comes from a set of physical findings found within extensive electronic medical record. This list is then screened to select only those that represent "normal" findings, a process that yields 96% total agreement among 4 clinicians surveyed. This final list of stereotypical "normal" sentences accounts for 64% of the clinical narrative text. Sentences in the clinical note that do not match sentences in the "normal" list are assumed to be "abnormal". Puya screened 98 clinical notes consisting of 610 individual sentences. Puya achieved a sensitivity of 100%, a specificity of 63%, a positive predictive value of 44% and a negative predictive value of 100%. This leads to an application that reduces informational noise.

Computer Communication Networks↗

Towards a unified medical lexicon for French.

Medical Informatics has a constant need for basic Medical Language Processing tasks, e.g., for coding into controlled vocabularies, free text indexing and information retrieval. Most of these tasks involve term matching and rely on lexical resources: lists of words with attached information, including inflected forms and derived words, etc. Such resources are publicly available for the English language with the UMLS Specialist Lexicon, but not in other languages. For the French language, several teams have worked on the subject and built local lexical resources. The goal of the present work is to pool and unify these resources and to add extensively to them by exploiting medical terminologies and corpora, resulting in a unified medical lexicon for French (UMLF). This paper exposes the issues raised by such an objective, describes the methods on which the project relies and illustrates them with experimental results.

Algorithms↗

Coh-metrix: analysis of text on cohesion and language.

Advances in computational linguistics and discourse processing have made it possible to automate many language- and text-processing mechanisms. We have developed a computer tool called Coh-Metrix, which analyzes texts on over 200 measures of cohesion, language, and readability. Its modules use lexicons, part-of-speech classifiers, syntactic parsers, templates, corpora, latent semantic analysis, and other components that are widely used in computational linguistics. After the user enters an English text, CohMetrix returns measures requested by the user. In addition, a facility allows the user to store the results of these analyses in data files (such as Text, Excel, and SPSS). Standard text readability formulas scale texts on difficulty by relying on word length and sentence length, whereas Coh-Metrix is sensitive to cohesion relations, world knowledge, and language and discourse characteristics.

Comprehension↗

GMD@CSB.DB: the Golm Metabolome Database.

UNLABELLED: Metabolomics, in particular gas chromatography-mass spectrometry (GC-MS) based metabolite profiling of biological extracts, is rapidly becoming one of the cornerstones of functional genomics and systems biology. Metabolite profiling has profound applications in discovering the mode of action of drugs or herbicides, and in unravelling the effect of altered gene expression on metabolism and organism performance in biotechnological applications. As such the technology needs to be available to many laboratories. For this, an open exchange of information is required, like that already achieved for transcript and protein data. One of the key-steps in metabolite profiling is the unambiguous identification of metabolites in highly complex metabolite preparations from biological samples. Collections of mass spectra, which comprise frequently observed metabolites of either known or unknown exact chemical structure, represent the most effective means to pool the identification efforts currently performed in many laboratories around the world. Here we present GMD, The Golm Metabolome Database, an open access metabolome database, which should enable these processes. GMD provides public access to custom mass spectral libraries, metabolite profiling experiments as well as additional information and tools, e.g. with regard to methods, spectral information or compounds. The main goal will be the representation of an exchange platform for experimental research activities and bioinformatics to develop and improve metabolomics by multidisciplinary cooperation. AVAILABILITY: http://csbdb.mpimp-golm.mpg.de/gmd.html CONTACT: Steinhauser@mpimp-golm.mpg.de SUPPLEMENTARY INFORMATION: http://csbdb.mpimp-golm.mpg.de/

Database Management Systems↗

Monitoring free-text data using medical language processing.

In this paper, we describe a software system for automated monitoring of free-text data in a medical information system that we call RadTRAC (Radiology Text Report Analyzer and Classifier). RadTRAC uses a medical language processing tool and rules derived from statistical analysis of a database to process free-text chest X-ray (CXR) reports and identify reports that describe new or expanding neoplasms for the purpose of monitoring the follow-up of these patients. To evaluate the RadTRAC system, we examined a set of 470 consecutive radiology reports at the Veterans Administration Medical Center, Palo Alto, CA. We compared RadTRAC classification of CXR reports with retrospective expert classification of the reports and with clinical classification from CXR films as recorded in a logbook while the films were being read. The RadTRAC system had a sensitivity of 90% and a specificity of 82% using the logbook as the gold standard. This was similar to the performance of expert radiologists (sensitivity, 92%; specificity, 90%). We then reviewed the charts, appointment schedule, and subsequent X-ray reports of cases either in the logbook or that were identified by RadTRAC as needing follow-up. Two cases in the logbook could have potentially benefited from an automatic monitoring system to ensure follow-up. RadTRAC identified six confirmed new tumors or new metastatic lesions that were not in the logbook. Six other cases were identified by the RadTRAC system with suspicious X-ray findings that had either no follow-up or no further mention of the X-ray lesion in medical records. This suggests that a reminder system based on the RadTRAC technology would be potentially useful.

Abstracting and Indexing↗

Information extraction from biomedical text.

Information extraction is the process of scanning text for information relevant to some interest, including extracting entities, relations, and events. It requires deeper analysis than key word searches, but its aims fall short of the very hard and long-term problem of full text understanding. Information extraction represents a midpoint on this spectrum, where the aim is to capture structured information without sacrificing feasibility. One of the key ideas in this technology is to separate processing into several stages, in cascaded finite-state transducers. The earlier stages recognize smaller linguistic objects and work in a largely domain-independent fashion. The later stages take these linguistic objects as input and find domain-dependent patterns among them. There are now initial efforts to apply this technology to biomedical text. In other domains, the technology plateaued at about 60% recall and precision. Even if applications to biomedical text do no better than this, they could still prove to be of immense help to curatorial activities.

Abstracting and Indexing↗

The CAP cancer protocols--a case study of caCORE based data standards implementation to integrate with the Cancer Biomedical Informatics Grid.

BACKGROUND: The Cancer Biomedical Informatics Grid (caBIG) is a network of individuals and institutions, creating a world wide web of cancer research. An important aspect of this informatics effort is the development of consistent practices for data standards development, using a multi-tier approach that facilitates semantic interoperability of systems. The semantic tiers include (1) information models, (2) common data elements, and (3) controlled terminologies and ontologies. The College of American Pathologists (CAP) cancer protocols and checklists are an important reporting standard in pathology, for which no complete electronic data standard is currently available. METHODS: In this manuscript, we provide a case study of Cancer Common Ontologic Representation Environment (caCORE) data standard implementation of the CAP cancer protocols and checklists model--an existing and complex paper based standard. We illustrate the basic principles, goals and methodology for developing caBIG models. RESULTS: Using this example, we describe the process required to develop the model, the technologies and data standards on which the process and models are based, and the results of the modeling effort. We address difficulties we encountered and modifications to caCORE that will address these problems. In addition, we describe four ongoing development projects that will use the emerging CAP data standards to achieve integration of tissue banking and laboratory information systems. CONCLUSION: The CAP cancer checklists can be used as the basis for an electronic data standard in pathology using the caBIG semantic modeling methodology.

Clinical Protocols↗

Assessing semantic similarity measures for the characterization of human regulatory pathways.

MOTIVATION: Pathway modeling requires the integration of multiple data including prior knowledge. In this study, we quantitatively assess the application of Gene Ontology (GO)-derived similarity measures for the characterization of direct and indirect interactions within human regulatory pathways. The characterization would help the integration of prior pathway knowledge for the modeling. RESULTS: Our analysis indicates information content-based measures outperform graph structure-based measures for stratifying protein interactions. Measures in terms of GO biological process and molecular function annotations can be used alone or together for the validation of protein interactions involved in the pathways. However, GO cellular component-derived measures may not have the ability to separate true positives from noise. Furthermore, we demonstrate that the functional similarity of proteins within known regulatory pathways decays rapidly as the path length between two proteins increases. Several logistic regression models are built to estimate the confidence of both direct and indirect interactions within a pathway, which may be used to score putative pathways inferred from a scaffold of molecular interactions.

Databases, Protein↗

Terminology-driven mining of biomedical literature.

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective literature mining techniques that can help biologists to gather and make use of the knowledge encoded in text documents. Although the knowledge is organized around sets of domain-specific terms, few literature mining systems incorporate deep and dynamic terminology processing. RESULTS: In this paper, we present an overview of an integrated framework for terminology-driven mining from biomedical literature. The framework integrates the following components: automatic term recognition, term variation handling, acronym acquisition, automatic discovery of term similarities and term clustering. The term variant recognition is incorporated into terminology recognition process by taking into account orthographical, morphological, syntactic, lexico-semantic and pragmatic term variations. In particular, we address acronyms as a common way of introducing term variants in biomedical papers. Term clustering is based on the automatic discovery of term similarities. We use a hybrid similarity measure, where terms are compared by using both internal and external evidence. The measure combines lexical, syntactical and contextual similarity. Experiments on terminology recognition and clustering performed on a corpus of MEDLINE abstracts recorded the precision of 98 and 71% respectively. AVAILABILITY: software for the terminology management is available upon request.

Abbreviations as Topic↗

Medical-concept models and medical records: an approach based on GALEN and PEN&PAD.

OBJECTIVES: To investigate the issues raised in applying a preliminary version of the GALEN compositional concept reference (CORE) model to a series of radiographic reports, and to demonstrate that the same underlying concept model could be used in conjunction with both a detailed, fine-grained model of medical records based on that used in the PEN&PAD project and with other more conventional medical-record models. DESIGN: Following analysis and representation of concepts from a set of reports, a single report was taken as a "case study." This report was analyzed in detail in its entirety and represented using each of the medical-record models. RESULTS: The reports were successfully represented within the limits of the study, but a number of significant issues were raised. CONCLUSION: The compositional approach plus the PEN&PAD medical-record model allowed detailed information in the radiographic report to be represented, including information about the inferences and the clinical process. The resulting representation was large, and more compact representations may be necessary for some systems. Alternative encapsulations of the information as might be used in such systems were successfully prepared. The compositional approach avoided many issues that often cause controversy in the design of traditional coding and classification systems, but it raised other issues, including the handling of ambiguity and underspecification, linkage to information not explicitly present in the report, and questions concerning the focus of individual concepts. All work is preliminary and definitive conclusions await further studies and systematic evaluation.

Decision Making, Computer-Assisted↗

Notations for high efficiency data presentation in mammography.

As a result of improvements in Medical Language Processing, the availability of categorical information (such as diagnoses or radiology findings) is increasing rapidly. This increased availability has created a need for more efficient methods for computer presentation. One method for developing such presentations would be to adapt the hand-written notation systems already used in paper-based records. We have characterized one such notation system, the Mammography Notation Sublanguage(MNS). The MNS is a true medical sublanguage, with a definable lexicon and syntax. Compared with text reports, it represents a 37-fold size compression. A single "base", sublanguage pattern is identified for possible computer presentation of mammography findings. The issues involved in using such sublanguages for data presentation are discussed.

Language↗

Using medical language processing to support real-time evaluation of pneumonia guidelines.

OBJECTIVE: To evaluate if a medical language processing (MLP) system is able to support real-time computerization of community-acquired pneumonia (CAP) guidelines. METHODS: Prospective validation study in the emergency department of a tertiary care facility. All the chest x-ray reports available in real-time for an admission decision during a five-week period were included. The MLP system was compared to a physician for the automatic selection of eligible patients and on the extraction of radiographic findings required by five different CAP guidelines. The gold standard comprised of three independent physicians and reliability measures were calculated. The outcome measures were the area under the receiver operated characteristic curve (AUC) for selecting eligible patients, sensitivity, positive predictive value (PPV), and specificity for the extraction of radiographic findings. RESULTS: During the five-week period, 243 reports were available in real-time. The AUCs on selecting eligible CAP patients were 89.7% (CI: 84.2%, 93.7%) for the MLP system, and 93.3% (CI: 83.9%, 97.8%) for the physician. The average sensitivity, PPV, and specificity for radiographic findings that assessed localization and extension of CAP were respectively: 94%, 87%, 96% (physician); and 34%, 90%, 95% (MLP system). Both, the MLP system and the physician had average sensitivity, PPV, and specificity of 97%, 97%, and 99%, respectively, when localization was not an issue. Reliability measures for the gold standard were above 70%. CONCLUSION: The MLP system was able to support real-time computerization of guidelines by selecting eligible patients and extracting radiographic findings that do not assess localization and extension of CAP.

Community-Acquired Infections↗

The next generation of literature analysis: integration of genomic analysis into text mining.

Text-mining systems are indispensable tools to reduce the increasing flux of information in scientific literature to topics pertinent to a particular interest in focus. Most of the scientific literature is published as unstructured free text, complicating the development of data processing tools, which rely on structured information. To overcome the problems of free text analysis, structured, hand-curated information derived from literature is integrated in text-mining systems to improve precision and recall. In this paper several text-mining approaches are reviewed and the next step in development of text-mining systems, which is based on a concept of multiple lines of evidence, is described: results from literature analysis are combined with evidence from experiments and genome analysis to improve the accuracy of results and to generate additional knowledge beyond what is known solely from literature.

Abstracting and Indexing↗

Beyond the clause: extraction of phosphorylation information from medline abstracts.

MOTIVATION: Phosphorylation is an important biochemical reaction that plays a critical role in signal transduction pathways and cell-cycle processes. A text mining system to extract the phosphorylation relation from the literature is reported. The focus of this paper is on the new methods developed and implemented to connect and merge pieces of information about phosphorylation mentioned in different sentences in the text. The effectiveness and accuracy of the system as a whole as well as that of the methods for extraction beyond a clause/sentence is evaluated using an independently annotated dataset, the Phospho.ELM database. The new methods developed to merge pieces of information from different sentences are shown to be effective in significantly raising the recall without much difference in precision.

Artificial Intelligence↗