PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Using argumentation to retrieve articles with similar citations: an inquiry into improving related articles search in the MEDLINE digital library.

The aim of this study is to investigate the relationships between citations and the scientific argumentation found abstracts. We design a related article search task and observe how the argumentation can affect the search results. We extracted citation lists from a set of 3200 full-text papers originating from a narrow domain. In parallel, we recovered the corresponding MEDLINE records for analysis of the argumentative moves. Our argumentative model is founded on four classes: PURPOSE, METHODS, RESULTS and CONCLUSION. A Bayesian classifier trained on explicitly structured MEDLINE abstracts generates these argumentative categories. The categories are used to generate four different argumentative indexes. A fifth index contains the complete abstract, together with the title and the list of Medical Subject Headings (MeSH) terms. To appraise the relationship of the moves to the citations, the citation lists were used as the criteria for determining relatedness of articles, establishing a benchmark; it means that two articles are considered as "related" if they share a significant set of co-citations. Our results show that the average precision of queries with the PURPOSE and CONCLUSION features is the highest, while the precision of the RESULTS and METHODS features was relatively low. A linear weighting combination of the moves is proposed, which significantly improves retrieval of related articles.

Abstracting and Indexing↗

Evaluation of two dependency parsers on biomedical corpus targeted at protein-protein interactions.

We present an evaluation of Link Grammar and Connexor Machinese Syntax, two major broad-coverage dependency parsers, on a custom hand-annotated corpus consisting of sentences regarding protein-protein interactions. In the evaluation, we apply the notion of an interaction subgraph, which is the subgraph of a dependency graph expressing a protein-protein interaction. We measure the performance of the parsers for recovery of individual dependencies, fully correct parses, and interaction subgraphs. For Link Grammar, an open system that can be inspected in detail, we further perform a comprehensive failure analysis, report specific causes of error, and suggest potential modifications to the grammar. We find that both parsers perform worse on biomedical English than previously reported on general English. While Connexor Machinese Syntax significantly outperforms Link Grammar, the failure analysis suggests specific ways in which the latter could be modified for better performance in the domain.

Abstracting and Indexing↗

A hybrid method for relation extraction from biomedical literature.

PURPOSE: Over recent years, there has been a growing interest in extracting entities and relations from biomedical literature. There are a vast number of systems and approaches being proposed to extract biological relations, but none of them achieves satisfactory results. These methodologies are either parsing-based or pattern-based, which are not competent to handle the grammatical complexities of biomedical texts, or too complicated to be adapted. It is well known that appositive, coordinative propositions and such grammatical structures are extremely common in biomedical texts, particularly in full texts. However, these problems are still untouched for most of researchers. METHODS: In this paper, we have proposed a new approach, which is hybrid with both shallow parsing and pattern matching, to extract relations between proteins from scientific papers of biomedical themes. In the method, appositive and coordinative structures are interpreted based on the shallow parsing analysis, with both syntactic and semantic constraints. Then long sentences are splitted into sub-ones, from which relations are extracted by a greedy pattern matching algorithm, along with automatically generated patterns. RESULTS: Our approach is experimented to extract protein-protein interactions from full biomedical texts, and has achieved an average F-score of 80% on individual verbs, and 66% on all verbs. With the help of shallow parsing analysis, pattern matching is improved remarkably. Compared with the traditional pattern matching algorithm, our approach achieves about 7% improvement of both precision and F-score. In contrast to other systems, our approach achieves performance comparable to the best. A demo system has been available at http://spies.cs.tsinghua.edu.cn.

Abstracting and Indexing↗

Zone analysis in biology articles as a basis for information extraction.

In the field of biomedicine, an overwhelming amount of experimental data has become available as a result of the high throughput of research in this domain. The amount of results reported has now grown beyond the limits of what can be managed by manual means. This makes it increasingly difficult for the researchers in this area to keep up with the latest developments. Information extraction (IE) in the biological domain aims to provide an effective automatic means to dynamically manage the information contained in archived journal articles and abstract collections and thus help researchers in their work. However, while considerable advances have been made in certain areas of IE, pinpointing and organizing factual information (such as experimental results) remains a challenge. In this paper we propose tackling this task by incorporating into IE information about rhetorical zones, i.e. classification of spans of text in terms of argumentation and intellectual attribution. As the first step towards this goal, we introduce a scheme for annotating biological texts for rhetorical zones and provide a qualitative and quantitative analysis of the data annotated according to this scheme. We also discuss our preliminary research on automatic zone analysis, and its incorporation into our IE framework.

Abstracting and Indexing↗

Semantic representation of consumer questions and physician answers.

The aim of this study was to identify the underlying semantics of health consumers' questions and physicians' answers in order to analyze the semantic patterns within these texts. We manually identified semantic relationships within question-answer pairs from Ask-the-Doctor Web sites. Identification of the semantic relationship instances within the texts was based on the relationship classes and structure of the Unified Medical Language System (UMLS) Semantic Network. We calculated the frequency of occurrence of each semantic relationship class, and conceptual graphs were generated, joining concepts together through the semantic relationships identified. We then analyzed whether representations of physician's answers exactly matched the form of the question representations. Lastly, we examined characteristics of the answer conceptual graphs. We identified 97 semantic relationship instances in the questions and 334 instances in the answers. The most frequently identified semantic relationship in both questions and answers was brings_about (causal). We found that the semantic relationship propositions identified in answers that most frequently contain a concept also expressed in the question were: brings_about, isa, co_occurs_with, diagnoses, and treats. Using extracted semantic relationships from real-life questions and answers can produce a valuable analysis of the characteristics of these texts. This can lead to clues for creating semantic-based retrieval techniques that guide users to further information. For example, we determined that both consumers and physicians often express causative relationships and these play a key role in leading to further related concepts.

Humans↗

Qualitative assessment of the International Classification of Functioning, Disability, and Health with respect to the desiderata for controlled medical vocabularies.

BACKGROUND: The International Classification of Functioning, Disability, and Health (ICF), a classification system published in 2001 by the World Health Organization (WHO), provides a common language and framework for describing functional status information (FSI) in health records. METHODS: Informed by ongoing research in coding FSI in patient records, this paper qualitatively assesses the ICF framework with respect to the desiderata for controlled medical vocabularies, an enumerated a list of desirable qualities for controlled medical vocabularies proposed by Cimino [J.J. Cimino, Desiderata for controlled medical vocabularies in the twenty-first century, Meth. Inform. Med. 37 (1998) 394-403]. RESULTS: The ICF satisfies 5 of the 12 desiderata. Five points were not satisfied and two points could not be evaluated. CONCLUSION: The ICF is a rich source of relevant terms, concepts, and relationships, but it was not developed in consideration of requirements for formal terminologies. Therefore, it could serve as a base from which to develop a formal terminology of functioning and disability. This assessment is a key next step in the development of the ICF as a sensitive, universal measure of functional status.

Decision Support Techniques↗

Customizing clinical narratives for the electronic medical record interface using cognitive methods.

OBJECTIVE: As healthcare practice transitions from paper-based to computer-based records, there is increasing need to determine an effective electronic format for clinical narratives. Our research focuses on utilizing a cognitive science methodology to guide the conversion of medical texts to a more structured, user-customized presentation in the electronic medical record (EMR). DESIGN: We studied the use of discharge summaries by psychiatrists with varying expertise-experts, intermediates, and novices. Experts were given two hypothetical emergency care scenarios with narrative discharge summaries and asked to verbalize their clinical assessment. Based on the results, the narratives were presented in a more structured form. Intermediate and novice subjects received a narrative and a structured discharge summary, and were asked to verbalize their assessments of each. MEASUREMENTS: A qualitative comparison of the interview transcripts of all subjects was done by analysis of recall and inference made with respect to level of expertise. RESULTS: For intermediate and novice subjects, recall was greater with the structured form than with the narrative. Novices were also able to make more inferences (not always accurate) from the structured form than with the narrative. Errors occurred in assessments using the narrative form but not the structured form. CONCLUSIONS: Our cognitive methods to study discharge summary use enabled us to extract a conceptual representation of clinical narratives from end-users. This method allowed us to identify clinically relevant information that can be used to structure medical text for the EMR and potentially improve recall and reduce errors.

Cognitive Science↗

Amplification of Terminologia anatomica by French language terms using Latin terms matching algorithm: a prototype for other language.

OBJECTIVE: Terminologia anatomica is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subject to the addition of terms from others languages. On the other hand, Nomina anatomica, the previous standard, has been widely translated. Aim of this work was to append foreign terms to Terminologia by using similarity-matching algorithm between its Latin terms and those from Nomina. METHODS: A semi-automatic matching of Latin terms from Terminologia with those of Nomina was performed using a string-to-string distance algorithm and manual assessment. We used a French-Latin version of Nomina together with Terminologia and we suggested French terms for Terminologia. Coverage was evaluated by the number of exact and approximate matches. A target of 78% was set due to the higher number of terms in Terminologia compared to Nomina. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5982 terms (76.5%) of Terminologia. Our results indicated that more than 75% of the terms from Terminologia came from Nomina, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 78% target. The method is based only on Latin terms and can be used for other languages. We consider this work as a starting point for adding terms to other knowledge sources, such as the foundational model of anatomy or the Unified Medical Language System (UMLS).

Algorithms↗

Defining and relating biomedical terms: towards a cross-language morphosemantics-based system.

This paper addresses the issue of how semantic information can be automatically assigned to compound terms, i.e. both a definition and a set of semantic relations. This is particularly crucial when elaborating multilingual databases and when developing cross-language information retrieval systems. The paper shows how morphosemantics can contribute in the constitution of multilingual lexical networks in biomedical corpora. It presents a system capable of labelling terms with morphologically related words, i.e. providing them with a definition, and grouping them according to synonymy, hyponymy and proximity relations. The approach requires the interaction of three techniques: (1) a language-specific morphosemantic parser, (2) a multilingual table defining basic relations between word roots and (3) a set of language-independent rules to draw up the list of related terms. This approach has been fully implemented for French, on an about 29,000 terms biomedical lexicon, resulting to more than 3000 lexical families. A validation of the results against a manually annotated file by experts of the domain is presented, followed by a discussion of our method.

France↗

Using argumentation to extract key sentences from biomedical abstracts.

PROBLEM: key word assignment has been largely used in MEDLINE to provide an indicative "gist" of the content of articles and to help retrieving biomedical articles. Abstracts are also used for this purpose. However with usually more than 300 words, MEDLINE abstracts can still be regarded as long documents; therefore we design a system to select a unique key sentence. This key sentence must be indicative of the article's content and we assume that abstract's conclusions are good candidates. We design and assess the performance of an automatic key sentence selector, which classifies sentences into four argumentative moves: PURPOSE, METHODS, RESULTS and CONCLUSION METHODS: we rely on Bayesian classifiers trained on automatically acquired data. Features representation, selection and weighting are reported and classification effectiveness is evaluated on the four classes using confusion matrices. We also explore the use of simple heuristics to take the position of sentences into account. Recall, precision and F-scores are computed for the CONCLUSION class. For the CONCLUSION class, the F-score reaches 84%. Automatic argumentative classification using Bayesian learners is feasible on MEDLINE abstracts and should help user navigation in such repositories.

Abstracting and Indexing↗

GALEN based formal representation of ICD10.

OBJECTIVES: The main objective is to create a knowledge-intensive coding support tool for the International Classification of Diseases (ICD10), which is based on formal representation of ICD10 categories. Beyond this task the resulting ontology could be reused in various ways. Decidability is an important issue for computer-assisted coding; consequently the ontology should be represented in description logic. METHODS: The meaning of the ICD10 categories is represented using the GALEN Core Reference Model. Due to the deficiencies of its representation language (GRAIL) the ontology is transformed to the quasi-standard OWL. A test system which extracts disease concepts and classifies them to ICD10 categories has been implemented in Prolog to verify the feasibility of the approach. RESULTS: The formal representation of the first two chapters of ICD10 (infectious diseases and neoplasms) has been almost completed. The constructed ontology has been converted to OWL DL. The test system successfully identified diseases in medical records from gastrointestinal oncology (84% recall, however precision is only 45%). The classifier module is still under development. Due to the experiences gained during the modelling, in the future work FMA is going to be used as anatomical reference ontology.

Abstracting and Indexing↗

Consistency across the hierarchies of the UMLS Semantic Network and Metathesaurus.

OBJECTIVE: To develop and test a method for automatically detecting inconsistencies between the parent-child is-a relationships in the Metathesaurus and the ancestor-descendant relationships in the Semantic Network of the Unified Medical Language System (UMLS). METHODS: We exploited the fact that each Metathesaurus concept is assigned one or more semantic types from the UMLS Semantic Network and that the semantic types are arranged in a hierarchy. We compared the semantic types of each pair of parent and child concepts to determine if the types "explained" the Metathesaurus is-a relationships. We considered cases where the semantic type of the parent was neither the same as, nor an ancestor of, the semantic type of the child to be "unexplained." We applied this method to the January 2002 release of the UMLS and examined the unexplained cases we discovered to determine their causes. RESULTS: We found that 17022 (24.3%) of the parent-child is-a relationships in the UMLS Metathesaurus could not be explained based on the semantic types of the concepts. Causes for these discrepancies included cases where the parent or child was missing a semantic type, cases where the semantic type of the child was too general or the semantic type of the parent was too specific, cases where the parent-child relationship was incorrect, and cases where an ancestor-descendant relationship should be added to the UMLS Semantic network. In many cases, the specific cause of the discrepancy cannot be resolved without authoritative judgment by the UMLS developers. CONCLUSIONS: Our method successfully detects inconsistencies between the hierarchies of the UMLS Metathesaurus and Semantic Network. We believe that our method should be added to the set of tools that the UMLS developers use to maintain and audit the UMLS knowledge sources.

Abstracting and Indexing↗

Exploring semantic groups through visual approaches.

Objectives. We investigate several visual approaches for exploring semantic groups, a grouping of semantic types from the Unified Medical Language System (UMLS) semantic network. We are particularly interested in the semantic coherence of the groups, and we use the semantic relationships as important indicators of that coherence. Methods. First, we create a radial representation of the number of relationships among the groups, generating a profile for each semantic group. Second, we show that, in our partition, the relationships are organized around a limited number of pivot groups and that partitions created at random do not exhibit this property. Finally, we use correspondence analysis to visualize groupings resulting from the association between semantic types and the relationships. Results. The three approaches provide different views on the semantic groups and help detect potential inconsistencies. They make outliers immediately apparent, and, thus, serve as a tool for auditing and validating both the semantic network and the semantic groups.

Abstracting and Indexing↗

OQAFMA Querying agent for the Foundational Model of Anatomy: a prototype for providing flexible and efficient access to large semantic networks.

The development of large semantic networks, such as the UMLS, which are intended to support a variety of applications, requires a flexible and efficient query interface for the extraction of information. Using one of the source vocabularies of UMLS as a test bed, we have developed such a prototype query interface. We first identify common classes of queries needed by applications that access these semantic networks. Next, we survey StruQL, an existing query language that we adopted, which supports all of these classes of queries. We then describe the OQAFMA Querying Agent for the Foundational Model of Anatomy (OQAFMA), which provides an efficient implementation of a subset of StruQL by pre-computing a variety of indices. We describe how OQAFMA leverages database optimization by converting StruQL queries to SQL. We evaluate the flexibility and efficiency of our implementation using English queries written by anatomists. This evaluation verifies that OQAFMA provides flexible, efficient access to one such large semantic network, the Foundational Model of Anatomy, and suggests that OQAFMA could be an efficient query interface to other large biomedical knowledge bases, such as the Unified Medical Language System.

Abstracting and Indexing↗

A reference ontology for biomedical informatics: the Foundational Model of Anatomy.

The Foundational Model of Anatomy (FMA), initially developed as an enhancement of the anatomical content of UMLS, is a domain ontology of the concepts and relationships that pertain to the structural organization of the human body. It encompasses the material objects from the molecular to the macroscopic levels that constitute the body and associates with them non-material entities (spaces, surfaces, lines, and points) required for describing structural relationships. The disciplined modeling approach employed for the development of the FMA relies on a set of declared principles, high level schemes, Aristotelian definitions and a frame-based authoring environment. We propose the FMA as a reference ontology in biomedical informatics for correlating different views of anatomy, aligning existing and emerging ontologies in bioinformatics ontologies and providing a structure-based template for representing biological functions.

Abstracting and Indexing↗

Towards the development of a conceptual distance metric for the UMLS.

The objective of this work is to investigate the feasibility of conceptual similarity metrics in the framework of the Unified Medical Language System (UMLS). We have investigated an approach based on the minimum number of parent links between concepts, and evaluated its performance relative to human expert estimates on three sets of concepts for three terminologies within the UMLS (i.e., MeSH, ICD9CM, and SNOMED). The resulting quantitative metric enables computer-based applications that use decision thresholds and approximate matching criteria. The proposed conceptual matching supports problem solving and inferencing (using high-level, generic concepts) based on readily available data (typically represented as low-level, specific concepts). Through the identification of semantically similar concepts, conceptual matching also enables reasoning in the absence of exact, or even approximate, lexical matching. Finally, conceptual matching is relevant for terminology development and maintenance, machine learning research, decision support system development, and data mining research in biomedical informatics and other fields.

Algorithms↗

Fever detection from free-text clinical records for biosurveillance.

Automatic detection of cases of febrile illness may have potential for early detection of outbreaks of infectious disease either by identification of anomalous numbers of febrile illness or in concert with other information in diagnosing specific syndromes, such as febrile respiratory syndrome. At most institutions, febrile information is contained only in free-text clinical records. We compared the sensitivity and specificity of three fever detection algorithms for detecting fever from free-text. Keyword CC and CoCo classified patients based on triage chief complaints; Keyword HP classified patients based on dictated emergency department reports. Keyword HP was the most sensitive (sensitivity 0.98, specificity 0.89), and Keyword CC was the most specific (sensitivity 0.61, specificity 1.0). Because chief complaints are available sooner than emergency department reports, we suggest a combined application that classifies patients based on their chief complaint followed by classification based on their emergency department report, once the report becomes available.

Algorithms↗