PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Assessing the consistency of a biomedical terminology through lexical knowledge.

OBJECTIVE: We investigate the use of adjectival modification as a way of assessing the systematic use of linguistic phenomena to represent similar lexical or semantic features in the constituent terms of a vocabulary. METHODS: Terms consisting of one or more adjectival modifiers followed by a head noun are selected from disease and procedure terms in SNOMED. Frequently co-occurring adjectival modifiers are systematically combined with the contexts (i.e., terms minus modifier) of each modifier. The existence of these combinations is checked in both SNOMED and the entire UMLS Metathesaurus; the term corresponding to the context alone is similarly checked. Relationships among terms sharing a context and between each of these terms and their context are studied. RESULTS: Four pairs of modifiers were studied: (acute, chronic), (unilateral, bilateral), (primary, secondary), and (acquired, congenital). The numbers of contexts studied for each pair ranged from 73 to 974. The percentage of contexts associated with both modifiers ranged from 5 to 50% in SNOMED and from 10 to 60% in UMLS. The presence of the context term varied from 31 to 64% in SNOMED and from 43 to 79% in UMLS. Finally, 172 occurrences (9%) of synonymy between a modified term and the context term were found in SNOMED. One hundred and forty-five such occurrences (8%) were found in the entire Metathesaurus.

Dictionaries as Topic↗

Protein names and how to find them.

A prerequisite for all higher level information extraction tasks is the identification of unknown names in text. Today, when large corpora can consist of billions of words, it is of utmost importance to develop accurate techniques for the automatic detection, extraction and categorization of named entities in these corpora. Although named entity recognition might be regarded a solved problem in some domains, it still poses a significant challenge in others. In this work we focus on one of the more difficult tasks, the identification of protein names in text. This task presents several interesting difficulties because of the named entities variant structural characteristics, their sometimes unclear status as names, the lack of common standards and fixed nomenclatures, and the specifics of the texts in the molecular biology domain in which they appear. We describe how we approached these and other difficulties in the implementation of Yapex, a system for the automatic identification of protein names in text. We also evaluate Yapex under four different notions of correctness and compare its performance to that of another publicly available system for protein name recognition.

Dictionaries as Topic↗

MEDSYNDIKATE--a natural language system for the extraction of medical information from findings reports.

MEDSYNDIKATE is a natural language processor, which automatically acquires medical information from findings reports. In the course of text analysis their contents is transferred to conceptual representation structures, which constitute a corresponding text knowledge base. MEDSYNDIKATE is particularly adapted to deal properly with text structures, such as various forms of anaphoric reference relations spanning several sentences. The strong demands MEDSYNDIKATE poses on the availability of expressive knowledge sources are accounted for by two alternative approaches to acquire medical domain knowledge (semi)automatically. We also present data for the information extraction performance of MEDSYNDIKATE in terms of the semantic interpretation of three major syntactic patterns in medical documents.

Confidence Intervals↗

Learning anchor verbs for biological interaction patterns from published text articles.

Much of knowledge modeling in the molecular biology domain involves interactions between proteins, genes, various forms of RNA, small molecules, etc. Interactions between these substances are typically extracted and codified manually, increasing the cost and time for modeling and substantially limiting the coverage of the resulting knowledge base. In this paper, we describe an automatic system that learns from text interaction verbs; these verbs can then form the core of automatically retrieved patterns which model classes of biological interactions. We investigate text features relating verbs with genes and proteins, and apply statistical tests and a logistic regression statistical model to determine whether a given verb belongs to the class of interaction verbs. Our system, AVAD, achieves over 87% precision and 82% recall when tested on an 11 million word corpus of journal articles. In addition, we compare the automatically obtained results with a manually constructed database of interaction verbs and show that the automatic approach can significantly enrich the manual list by detecting rarer interaction verbs that were omitted from the database.

Artificial Intelligence↗

Terminology-driven literature mining and knowledge acquisition in biomedicine.

In this paper we describe Tagged Information Management System (TIMS), an integrated knowledge management system for the domain of molecular biology and biomedicine, in which terminology-driven literature mining, knowledge acquisition (KA), knowledge integration (KI), and XML-based knowledge retrieval are combined using tag information management and ontology inference. The system integrates automatic terminology acquisition, term variation management, hierarchical term clustering, tag-based information extraction (IE), and ontology-based query expansion. TIMS supports introducing and combining different types of tags (linguistic and domain-specific, manual and automatic). Tag-based interval operations and a query language are introduced in order to facilitate KA and retrieval from XML documents. Through KA examples, we illustrate the way in which literature mining techniques can be utilised for knowledge discovery from documents.

Artificial Intelligence↗

Restoring accents in unknown biomedical words: application to the French MeSH thesaurus.

In languages with diacritic marks, such as French, there remain instances of textual or terminological resources that are available in electronic form without diacritic marks, which hinders their use in natural language interfaces. In a specialized domain such as medicine, it is often the case that some words are not found in the available electronic lexicons. The issue of accenting unknown words then arises: it is the theme of this work. We propose two internal methods for accenting unknown words, which both learn on a reference set of accented words the contexts of occurrence of the various accented forms of a given letter. One method is adapted from part-of-speech tagging, the other is based on finite state transducers. We show experimental results for letter e on the French version of the Medical Subject Headings thesaurus. With the best training set, the tagging method obtains a precision-recall breakeven point of 84.2+/-4.4% and the transducer method 83.8+/-4.5% (with a baseline at 64%) for the unknown words that contain this letter. A consensus combination of both increases precision to 92.0+/-3.7% with a recall of 75%. We perform an error analysis and discuss further steps that might help improve over the current performance.

Algorithms↗

Evaluating and reducing the effect of data corruption when applying bag of words approaches to medical records.

Unlike journal corpora, which are supposed to be carefully reviewed before being published, the quality of documents in a patient record are often corrupted by mispelled words and conventional graphies or abbreviations. After a survey of the domain, the paper focuses on evaluating the effect of such corruption on an information retrieval (IR) engine. The IR system uses a classical bag of words approach, with stems as representation items and term frequency-inverse document frequency (tf-idf) as weighting schema; we pay special attention to the normalization factor. First results shows that even low corruption levels (3%) do affect retrieval effectiveness (4-7%), whereas higher corruption levels can affect retrieval effectiveness by 25%. Then, we show that the use of an improved automatic spelling correction system, applied on the corrupted collection, can almost restore the retrieval effectiveness of the engine.

Forecasting↗

Medical narratives in electronic medical records.

In this article, we describe the state of the art and directions of current development and research with respect to the inclusion of medical narratives in electronic medical-record systems. We used information about 20 electronic medical-record systems as presented in the literature. We divided these systems into 'classical' systems that matured before 1990 and are now used in a broad range of medical domains, and 'experimental' systems, more recently developed and, in general, more innovative. In the literature, three major challenges were addressed: facilitation of direct data entry, achieving unambiguous understandability of data, and improvement of data presentation. Promising approaches to tackle the first and second challenge are the use of dynamic data-entry forms that anticipate sensible input, and free-text data entry followed by natural-language interpretation. Both these approaches require a highly expressive medical terminology. How to facilitate the access to medical narratives has not been studied much. We found facilitating examples of presenting this information as fluent prose, of optimising the screen design with fixed position cues, and of imposing medical narratives with a structure of indexable paragraphs that can be used in flowsheets. We conclude that further study is needed to develop an optimal searching structure for medical narratives.

Database Management Systems↗

Computer assisted medical diagnosis using the Web.

The ADM (Aide au diagnostic Medical) project was started 15 years ago and was the first telematic project for physicians in France using the MINITEL terminal. The knowledge base contains information on more than 10000 diseases from all pathological fields, using more than 100000 signs or symptoms. The ADM system has two main functionalities for physicians: consultation of diseases descriptions and list of diseases containing one or more symptoms. The ADM knowledge base is supported by a relational database management system (DBMS ORACLE) and we developed a Web interface using the Perl language to produce HTML pages for the web server. We will describe our experience on redesigning a large existing medical knowledge base for diffusion on the web Internet.

Artificial Intelligence↗

Information retrieval: an overview of system characteristics.

The paper gives an overview of characteristics of information retrieval (IR) systems. The characteristics are identified from the descriptions of 23 IR systems. Four IR models are discussed: the Boolean model, the vector model, the probabilistic model and the connectionistic model. Twelve other characteristics of IR models are identified: search intermediary, domain knowledge, relevance feedback, natural language interface, graphical query language, conceptual queries, full-text IR, field searching, fuzzy queries, hypertext integration, machine learning, and ranked output. Finally, the relevance of IR systems for the World Wide Web is established.

Algorithms↗

Practical development of re-usable terminologies: GALEN-IN-USE and the GALEN Organisation.

Medical terminology is now playing a key role in medical software. This requires new techniques with which many clinical users, classification experts and applications developers are unfamiliar. There is a conflict in that the more re-usable techniques for terminology needed to support sharing of information among many different applications are more difficult to use for any one application. A layered approach to re-use is described which combines techniques from first generation systems and relatively easily understood second generation systems with the formal rigour of third generation systems to resolve this conflict. The methodology also provides a potentially rigorous approach to defining the relationship between terminology and structure in the electronic healthcare record architecture. It provides a natural migration pathway from existing systems to powerful re-usable multilingual terminologies.

Expert Systems↗

From a time standard for medical informatics to a controlled language for health.

CEN ENV 12381 is a European Prestandard focusing on formal representation and explicit reference of temporal information in healthcare informatics and telematics. One of its merits is not just the possibility to represent natural language expressions containing time-related information in a structured way, but also to give some mechanisms on how clinical language itself can be used to convey meaning unambiguously. As such, CEN ENV 12381 introduces the notion of 'controlled language use' in the domain of healthcare. In this paper the principles behind controlled language design and use are explained. Through a detailed study of the inconsistencies and ambiguities that arise when interpreting Snomed procedure terms in the framework of the Galen-In-Use project, it is shown that most of them can be explained as a violation of sound term-formation principles. A proposal is made to develop a controlled language for health and to use it in subsequent versions of coding and classification systems. It is expected that such an endeavour will lead to a more effective application of linguistic engineering in areas such as automatic knowledge acquisition, automatic translation, and terminology validation in the domain of healthcare informatics.

Artificial Intelligence↗

Exploiting the terminological approach from CEN/TC251 and GALEN to support semantic interoperability of healthcare record systems.

We apply the principles included in two CEN standards (ENV 12265, ENV 12264) to the analysis of the semantic structure of health record systems, to support their semantic interoperability. This result was made possible by dramatic methodological progress in the field of terminological systems--due to a worldwide evolution towards a new generation--and by the experience we acquired in the GALEN-IN-USE (formerly GALEN) project. The meaning behind names, content and context of record items and record item complexes can be considered as a 'semantic continuum'. This continuum is made explicit, by building a suitable paraphrase in a controlled language. We can then apply the principles we previously elaborated for the second generation of terminological system. Methodology and tools for generating a controlled language and a second-generation terminological system were developed and successfully used in the GALEN-IN-USE project and promising experiments were performed on elements of record structure listed in LOINC and SDM. In this way, the semantic structures of different record systems can be expressed by the resulting common formalism and thus, information units can be faithfully exchanged among different structures.

Italy↗

Discourse structures in medical reports--watch out! The generation of referentially coherent and valid text knowledge bases in the MEDSYNDIKATE system.

The automatic analysis of medical narratives currently suffers from neglecting text structure phenomena such as referential relations between discourse units. This has unwarranted effects on the descriptional adequacy of medical knowledge bases automatically generated from texts. The resulting representation bias can be characterized in terms of incomplete, artificially fragmented and referentially invalid knowledge structures. We focus here on four basic types of textual reference relations, viz. pronominal and nominal anaphora, textual ellipsis and metonymy and show how to deal with them in an adequate text parsing device. Since the types of reference relations we discuss show an increasing dependence on conceptual background knowledge, we stress the need for formally grounded, expressive conceptual representation systems for medical knowledge. Our suggestions are based on experience with MEDSYNDIKATE, a medical text knowledge acquisition system designed to properly deal with various sorts of discourse structure phenomena.

Artificial Intelligence↗

From syntactic-semantic tagging to knowledge discovery in medical texts.

In the GALEN project, the syntactic-semantic tagger MultiTALE is upgraded to extract knowledge from natural language surgical procedure expressions. In this paper, we describe the methodology applied and show that out of a randomly selected sample of such expressions coming from the procedure axis of Snomed International, 81% could be analysed correctly. The problems encountered fall in three different categories: unusual grammatical configurations within the Snomed terms, insufficient domain knowledge and different categorisation of concepts and semantic links in the domain and linguistic models used. It is concluded that the Multi-TALE system can be used to attach meaning to words that not have been encountered previously, but that an interface ontology mediating between domain models and linguistic models is needed to arrive at a higher level of independence from both particular languages and from particular domains.

Forecasting↗

Natural language generation of surgical procedures.

A number of compositional Medical Concept Representation systems are being developed. Although these provide for a detailed conceptual representation of the underlying information, they have to be translated back to natural language for used by end-users and applications. The GALEN programme has been developing one such representation and we report here on a tool developed to generate natural language phrases from the GALEN conceptual representations. This tool can be adapted to different source modelling schemes and to different destination languages or sublanguages of a domain. It is based on a multilingual approach to natural language generation, realised through a clean separation of the domain model from the linguistic model and their link by well defined structures. Specific knowledge structures and operations have been developed for bridging between the modelling 'style' of the conceptual representation and natural language. Using the example of the scheme developed for modelling surgical operative procedures within the GALEN-IN-USE project, we show how the generator is adapted to such a scheme. The basic characteristics of the surgical procedures scheme are presented together with the basic principles of the generation tool. Using worked examples, we discuss the transformation operations which change the initial source representation into a form which can more directly be translated to a given natural language. In particular, the linguistic knowledge which has to be introduced--such as definitions of concepts and relationships is described. We explain the overall generator strategy and how particular transformation operations are triggered by language-dependent and conceptual parameters. Results are shown for generated French phrases corresponding to surgical procedures from the urology domain.

Linguistics↗

The structure of science information.

The organization of information within science can be investigated in a principled way through analysis of science language. The restricted use of language in science enables description of the informational structure of science and of particular subfields, with strong similarities to structures in mathematics and programming languages. This result rests on decades of research into the relation between form and content in language, based on an information-theoretic approach to the structure of information. Examples are provided from immunology and the social sciences. Practical applications include storage of science information in databases, indexing the literature, and identification and resolution of controversy.

Abstracting and Indexing↗

Information extraction for enhanced access to disease outbreak reports.

Document search is generally based on individual terms in the document. However, for collections within limited domains it is possible to provide more powerful access tools. This paper describes a system designed for collections of reports of infectious disease outbreaks. The system, Proteus-BIO, automatically creates a table of outbreaks, with each table entry linked to the document describing that outbreak; this makes it possible to use database operations such as selection and sorting to find relevant documents. Proteus-BIO consists of a Web crawler which gathers relevant documents; an information extraction engine which converts the individual outbreak events to a tabular database; and a database browser which provides access to the events and, through them, to the documents. The information extraction engine uses sets of patterns and word classes to extract the information about each event. Preparing these patterns and word classes has been a time-consuming manual operation in the past, but automated discovery tools now make this task significantly easier. A small study comparing the effectiveness of the tabular index with conventional Web search tools demonstrated that users can find substantially more documents in a given time period with Proteus-BIO.

Abstracting and Indexing↗