PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The LBI-method for automated indexing of diagnoses by using SNOMED. Part 2. Evaluation.

We present a simple, formal, lexicon-based method for automated indexing of diagnoses based on the Systematized Nomenclature of Medicine (SNOMED), called LBI-method. Part 1 gave an introduction to the LBI-method and presented its realisation as application system SALBIDH. Part 2 presents the design and the results of an evaluation study to judge the quality of the LBI-method. In this evaluation study the quality of automated indexing as well as the quality of the retrieval of patient data by using automated indexed diagnoses was examined. The results show that the retrieval based on SNOMED indices is at least as good as the retrieval based on ICD classes despite a lot of indexing errors. From this we gather that our system is not yet good enough for immediate routine use but that an appropriate indexing quality and, as a result, a higher retrieval quality can be achieved after few improvements of the LBI-method, especially after revision of the lexicons.

Abstracting and Indexing↗

Content-based retrieval of historical Ottoman documents stored as textual images.

There is an accelerating demand to access the visual content of documents stored in historical and cultural archives. Availability of electronic imaging tools and effective image processing techniques makes it feasible to process the multimedia data in large databases. In this paper, a framework for content-based retrieval of historical documents in the Ottoman Empire archives is presented. The documents are stored as textual images, which are compressed by constructing a library of symbols occurring in a document, and the symbols in the original image are then replaced with pointers into the codebook to obtain a compressed representation of the image. The features in wavelet and spatial domain based on angular and distance span of shapes are used to extract the symbols. In order to make content-based retrieval in historical archives, a query is specified as a rectangular region in an input image and the same symbol-extraction process is applied to the query region. The queries are processed on the codebook of documents and the query images are identified in the resulting documents using the pointers in textual images. The querying process does not require decompression of images. The new content-based retrieval framework is also applicable to many other document archives using different scripts.

Abstracting and Indexing↗

Online recognition of Chinese characters: the state-of-the-art.

Online handwriting recognition is gaining renewed interest owing to the increase of pen computing applications and new pen input devices. The recognition of Chinese characters is different from western handwriting recognition and poses a special challenge. To provide an overview of the technical status and inspire future research, this paper reviews the advances in online Chinese character recognition (OLCCR), with emphasis on the research works from the 1990s. Compared to the research in the 1980s, the research efforts in the 1990s aimed to further relax the constraints of handwriting, namely, the adherence to standard stroke orders and stroke numbers and the restriction of recognition to isolated characters only. The target of recognition has shifted from regular script to fluent script in order to better meet the requirements of practical applications. The research works are reviewed in terms of pattern representation, character classification, learning/adaptation, and contextual processing. We compare important results and discuss possible directions of future research.

Algorithms↗

Font adaptive word indexing of modern printed documents.

We propose an approach for the word-level indexing of modern printed documents which are difficult to recognize using current OCR engines. By means of word-level indexing, it is possible to retrieve the position of words in a document, enabling queries involving proximity of terms. Web search engines implement this kind of indexing, allowing users to retrieve Web pages on the basis of their textual content. Nowadays, digital libraries hold collections of digitized documents that can be retrieved either by browsing the document images or relying on appropriate metadata assembled by domain experts. Word indexing tools would therefore increase the access to these collections. The proposed system is designed to index homogeneous document collections by automatically adapting to different languages and font styles without relying on OCR engines for character recognition. The approach is based on three main ideas: the use of Self Organizing Maps (SOM) to perform unsupervised character clustering, the definition of one suitable vector-based word representation whose size depends on the word aspect-ratio, and the run-time alignment of the query word with indexed words to deal with broken and touching characters. The most appropriate applications are for processing modern printed documents (17th to 19th centuries) where current OCR engines are less accurate. Our experimental analysis addresses six data sets containing documents ranging from books of the 17th century to contemporary journals.

Abstracting and Indexing↗

Automatic SNOMED classification--a corpus-based method.

This paper presents a method of automatic classification of clinical narrative through text comparison. A diagnosis report can be classified by searching archive texts that show a high textual similarity, and the 'nearest neighbor classifies the case. This paper describes the method's theoretical background and gives implementation details. Large scale simulation experiments were run with a wide range of histology reports. Results showed that for 80-84% of the trials, relevant classification lines were included among the first five alternatives. In 5% of the cases, retrieval was unsuccessful due to the absence of relevant archive reports. From the results it is concluded that the method is a versatile approach for finding potentially good classifications.

Algorithms↗

Automatic processing of spoken dialogue in the home hemodialysis domain.

Spoken medical dialogue is a valuable source of information, and it forms a foundation for diagnosis, prevention and therapeutic management. However, understanding even a perfect transcript of spoken dialogue is challenging for humans because of the lack of structure and the verbosity of dialogues. This work presents a first step towards automatic analysis of spoken medical dialogue. The backbone of our approach is an abstraction of a dialogue into a sequence of semantic categories. This abstraction uncovers structure in informal, verbose conversation between a caregiver and a patient, thereby facilitating automatic processing of dialogue content. Our method induces this structure based on a range of linguistic and contextual features that are integrated in a supervised machine-learning framework. Our model has a classification accuracy of 73%, compared to 33% achieved by a majority baseline (p<0.01). This work demonstrates the feasibility of automatically processing spoken medical dialogue.

Algorithms↗

Neural correlates for the acquisition of natural language syntax.

Some types of simple and logically possible syntactic rule never occur in human language grammars, leading to a distinction between grammatical and nongrammatical syntactic rules. Comparison of the neuroanatomical correlates underlying the acquisition of grammatical and nongrammatical rules can provide relevant evidence on the neural processes dedicated to language acquisition in a given developmental stage. Until present no direct evidence on the neural mechanisms subserving language acquisition at any developmental stage has been supplied. We used fMRI in investigating the acquisition of grammatical and nongrammatical rules in the specified sense in 14 healthy adults. Grammatical rules compared with nongrammatical rules specifically activated a left hemispheric network including Broca's area, as shown by direct comparisons between the two rule types. The selective role of Broca's area was further confirmed by time x condition interactions and by proficiency effects, in that higher proficiency in grammatical rule usage, but not in usage of nongrammatical rules, led to higher levels of activation in this area. These findings provide evidence for the neural mechanisms underlying language acquisition in adults.

Adult↗

Semantic enrichment for medical ontologies.

The Unified Medical Language System (UMLS) contains two separate but interconnected knowledge structures, the Semantic Network (upper level) and the Metathesaurus (lower level). In this paper, we have attempted to work out better how the use of such a two-level structure in the medical field has led to notable advances in terminologies and ontologies. However, most ontologies and terminologies do not have such a two-level structure. Therefore, we present a method, called semantic enrichment, which generates a two-level ontology from a given one-level terminology and an auxiliary two-level ontology. During semantic enrichment, concepts of the one-level terminology are assigned to semantic types, which are the building blocks of the upper level of the auxiliary two-level ontology. The result of this process is the desired new two-level ontology. We discuss semantic enrichment of two example terminologies and how we approach the implementation of semantic enrichment in the medical domain. This implementation performs a major part of the semantic enrichment process with the medical terminologies, with difficult cases left to a human expert.

Artificial Intelligence↗

Impact of voice- and knowledge-enabled clinical reporting--US example.

This study shows qualitative and quantitative estimates of the national and the clinic level impact of utilizing voice and knowledge enabled clinical reporting systems. Using common sense estimation methodology, we show that the delivery of health care can experience a dramatic improvement in four areas as a result of the broad use of voice and knowledge enabled clinical reporting: (1) Process Quality as measured by cost savings, (2) Organizational Quality as measured by compliance, (3) Clinical Quality as measured by clinical outcomes and (4) Service Quality as measured by patient satisfaction. If only 15 percent of US physicians replaced transcription with modem clinical reporting voice-based methodology, about one half billion dollars could be saved. $6.7 Billion could be saved annually if all medical reporting currently transcribed was handled with voice-and knowledge-enabled dictation and reporting systems.

Artificial Intelligence↗

Automated analysis of medical text. I. Clue gathering.

Clinical practice of medicine is highly information-intensive. At the bedside, past experience is the primary justification of reasoning and decisions. This past medical experience is an amalgamation of textbook information and personal experience. During the last 2-3 decades, both of these major sources of clinical information have appeared less and less effective. The pace of progress, resulting in better diagnostic tools and new therapies, has undermined our personal experience, and for the same reason, the time lapse between drafting the manuscripts and distributing the textbooks has become a growing problem. Emphasis has shifted from textbooks to scientific journals with shorter publishing delays, and the role of daily newspapers and television programs seems to be growing. The traditional ways of gathering clinical knowledge and experience seem to fail more and more. In addition to textbooks and scientific journals, current clinical experience is described in millions of patient records, stored in hospitals and ambulatory care offices. However, we have no easy access to patient charts, and we are lacking a method for cost-effective merging of clinical case histories to make them suitable for much-needed statistical inferences. Computers could make a major contribution in this area, but first we must bridge the gap between the narrative text in the medical record and computer technology. Recently, much encouraging progress has been made in automated medical text processing, the topic of this paper.

Abstracting and Indexing↗

Adaptation in statistical pattern recognition using tangent vectors.

We integrate the tangent method into a statistical framework for classification analytically and practically. The resulting consistent framework for adaptation allows us to efficiently estimate the tangent vectors representing the variability. The framework improves classification results on two real-world pattern recognition tasks from the domains handwritten character recognition and automatic speech recognition.

Algorithms↗

A scale space approach for automatically segmenting words from historical handwritten documents.

Many libraries, museums, and other organizations contain large collections of handwritten historical documents, for example, the papers of early presidents like George Washington at the Library of Congress. The first step in providing recognition/ retrieval tools is to automatically segment handwritten pages into words. State of the art segmentation techniques like the gap metrics algorithm have been mostly developed and tested on highly constrained documents like bank checks and postal addresses. There has been little work on full handwritten pages and this work has usually involved testing on clean artificial documents created for the purpose of research. Historical manuscript images, on the other hand, contain a great deal of noise and are much more challenging. Here, a novel scale space algorithm for automatically segmenting handwritten (historical) documents into words is described. First, the page is cleaned to remove margins. This is followed by a gray-level projection profile algorithm for finding lines in images. Each line image is then filtered with an anisotropic Laplacian at several scales. This procedure produces blobs which correspond to portions of characters at small scales and to words at larger scales. Crucial to the algorithm is scale selection, that is, finding the optimum scale at which blobs correspond to words. This is done by finding the maximum over scale of the extent or area of the blobs. This scale maximum is estimated using three different approaches. The blobs recovered at the optimum scale are then bounded with a rectangular box to recover the words. A postprocessing filtering step is performed to eliminate boxes of unusual size which are unlikely to correspond to words. The approach is tested on a number of different data sets and it is shown that, on 100 sampled documents from the George Washington corpus of handwritten document images, a total error rate of 17 percent is observed. The technique outperforms a state-of-the-art gap metrics word-segmentation algorithm on this collection.

Abstracting and Indexing↗

Automating tissue bank annotation from pathology reports - comparison to a gold standard expert annotation set.

Surgical pathology specimens are an important resource for medical research, particularly for cancer research. Although research studies would benefit from information derived from the surgical pathology reports, access to this information is limited by use of unstructured free-text in the reports. We have previously described a pipeline-based system for automated annotation of surgical pathology reports with UMLS concepts, which has been used to code over 450,000 surgical pathology reports at our institution. In addition to coding UMLS terms, it annotates values of several key variables, such as TNM stage and cancer grade. The object of this study was to evaluate the potential and limitations of automated extraction of these variables, by measuring the performance of the system against a true gold standard - manually encoded data entered by expert tissue annotators. We categorized and analyzed errors to determine the potential and limitations of information extraction from pathology reports for the purpose of automated biospecimen annotation.

Abstracting and Indexing↗

Medical language processing with SGML display.

The paper demonstrates several ways that medical language processing can be combined with emerging display technologies to facilitate the extraction of data from free-text patient documents. The techniques allow rapid review via highlighting of the results of processing. Coupling of text markup with further procedures is envisioned.

Asthma↗

Acquisition of lexical resources from SNOMED for medical language processing.

Medical language processing depends on large-coverage, fine-grained specialized lexicons. The vast majority of existing electronic lexicons concern the English language; for other languages such as French, resources are scarce. In contrast, large medical thesauri exist in numerous languages, including French. Our goal was to study what kind of linguistic information could be extracted from thesauri into a lexicon, in which places human intervention is necessary, and what kind of issues arise in this process. We designed in this purpose a method to build a semantic lexicon from a subset of the SNOMED axes in their French translation.

France↗

Getting to the (c)ore of knowledge: mining biomedical literature.

Literature mining is the process of extracting and combining facts from scientific publications. In recent years, many computer programs have been designed to extract various molecular biology findings from Medline abstracts or full-text articles. The present article describes the range of text mining techniques that have been applied to scientific documents. It divides 'automated reading' into four general subtasks: text categorization, named entity tagging, fact extraction, and collection-wide analysis. Literature mining offers powerful methods to support knowledge discovery and the construction of topic maps and ontologies. An overview is given of recent developments in medical language processing. Special attention is given to the domain particularities of molecular biology, and the emerging synergy between literature mining and molecular databases accessible through Internet.

Abstracting and Indexing↗