PubMed Health⌕ Search

Biomedical subjects

Thomas C Rindflesch

Publications and source records attributed to Thomas C Rindflesch.

18 recordsLinked to original sources

Argument-predicate distance as a filter for enhancing precision in extracting predications on the genetic etiology of disease.

BACKGROUND: Genomic functional information is valuable for biomedical research. However, such information frequently needs to be extracted from the scientific literature and structured in order to be exploited by automatic systems. Natural language processing is increasingly used for this purpose although it inherently involves errors. A postprocessing strategy that selects relations most likely to be correct is proposed and evaluated on the output of SemGen, a system that extracts semantic predications on the etiology of genetic diseases. Based on the number of intervening phrases between an argument and its predicate, we defined a heuristic strategy to filter the extracted semantic relations according to their likelihood of being correct. We also applied this strategy to relations identified with co-occurrence processing. Finally, we exploited postprocessed SemGen predications to investigate the genetic basis of Parkinson's disease. RESULTS: The filtering procedure for increased precision is based on the intuition that arguments which occur close to their predicate are easier to identify than those at a distance. For example, if gene-gene relations are filtered for arguments at a distance of 1 phrase from the predicate, precision increases from 41.95% (baseline) to 70.75%. Since this proximity filtering is based on syntactic structure, applying it to the results of co-occurrence processing is useful, but not as effective as when applied to the output of natural language processing. In an effort to exploit SemGen predications on the etiology of disease after increasing precision with postprocessing, a gene list was derived from extracted information enhanced with postprocessing filtering and was automatically annotated with GFINDer, a Web application that dynamically retrieves functional and phenotypic information from structured biomolecular resources. Two of the genes in this list are likely relevant to Parkinson's disease but are not associated with this disease in several important databases on genetic disorders. CONCLUSION: Information based on the proximity postprocessing method we suggest is of sufficient quality to be profitably used for subsequent applications aimed at uncovering new biomedical knowledge. Although proximity filtering is only marginally effective for enhancing the precision of relations extracted with co-occurrence processing, it is likely to benefit methods based, even partially, on syntactic structure, regardless of the relation.

Genetic Diseases, Inborn↗

Semantic representation of consumer questions and physician answers.

The aim of this study was to identify the underlying semantics of health consumers' questions and physicians' answers in order to analyze the semantic patterns within these texts. We manually identified semantic relationships within question-answer pairs from Ask-the-Doctor Web sites. Identification of the semantic relationship instances within the texts was based on the relationship classes and structure of the Unified Medical Language System (UMLS) Semantic Network. We calculated the frequency of occurrence of each semantic relationship class, and conceptual graphs were generated, joining concepts together through the semantic relationships identified. We then analyzed whether representations of physician's answers exactly matched the form of the question representations. Lastly, we examined characteristics of the answer conceptual graphs. We identified 97 semantic relationship instances in the questions and 334 instances in the answers. The most frequently identified semantic relationship in both questions and answers was brings_about (causal). We found that the semantic relationship propositions identified in answers that most frequently contain a concept also expressed in the question were: brings_about, isa, co_occurs_with, diagnoses, and treats. Using extracted semantic relationships from real-life questions and answers can produce a valuable analysis of the characteristics of these texts. This can lead to clues for creating semantic-based retrieval techniques that guide users to further information. For example, we determined that both consumers and physicians often express causative relationships and these play a key role in leading to further related concepts.

Humans↗

Effects of information and machine learning algorithms on word sense disambiguation with small datasets.

Current approaches to word sense disambiguation use (and often combine) various machine learning techniques. Most refer to characteristics of the ambiguity and its surrounding words and are based on thousands of examples. Unfortunately, developing large training sets is burdensome, and in response to this challenge, we investigate the use of symbolic knowledge for small datasets. A naïve Bayes classifier was trained for 15 words with 100 examples for each. Unified Medical Language System (UMLS) semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in nine experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. To investigate this large variance, we performed several follow-up evaluations, testing additional algorithms (decision tree and neural network), and gold standards (per expert), but the results did not significantly differ. However, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators. We conclude that neither algorithm nor individual human behavior cause these large differences, but that the structure of the UMLS Metathesaurus (used to represent senses of ambiguous words) contributes to inaccuracies in the gold standard, leading to varied performance of word sense disambiguation techniques.

Algorithms↗

Determining prominent subdomains in medicine.

We discuss an automated method for identifying prominent subdomains in medicine. The motivation is to enhance the results of natural language processing by focusing on sublanguages associated with medical specialties concerned with prevalent disorders. At the core of our approach is a statistical system for topical categorization of medical text. A method based on epidemiological evidence is compared to another that considers frequency of occurrence of Medline citations. We suggest the isolation of UMLS terminology peculiar to individual medical specialties as a way of enhancing natural language processing systems in the biomedical domain.

Abstracting and Indexing↗

Medical facts to support inferencing in natural language processing.

We report on the use of medical facts to support the enhancement of natural language processing of biomedical text. Inferencing in semantic interpretation depends on a fact repository as well as an ontology. We used statistical methods to construct a repository of drug-disorder co-occurrences from a large collection of clinical notes, and this resource is used to validate inferences automatically drawn during semantic interpretation of Medline citations about pharmacologic interventions for disease. We evaluated the results against a published reference standard for treatment of diseases.

Disease↗

Using symbolic knowledge in the UMLS to disambiguate words in small datasets with a naïve Bayes classifier.

Current approaches to word sense disambiguation use and combine various machine-learning techniques. Most refer to characteristics of the ambiguous word and surrounding words and are based on hundreds of examples. Unfortunately, developing large training sets is time-consuming. We investigate the use of symbolic knowledge to augment machine-learning techniques for small datasets. UMLS semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. A naïve Bayes classifier was trained for 15 words with 100 examples for each. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in eight experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. In a follow-up evaluation, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators.

Abstracting and Indexing↗

Identifying respiratory findings in emergency department reports for biosurveillance using MetaMap.

Clinical conditions described in patients' dictated reports are necessary for automated detection of patients with respiratory illnesses such as inhalational anthrax and pneumonia. We applied MetaMap to emergency department reports to extract a set of 71 clinical conditions relevant to detection of a lower respiratory outbreak. We indexed UMLS terms in emergency department reports with MetaMap, filtered the indexed output with a specialized lexicon of UMLS terms for the domain, and mapped the clinical conditions of interest to concepts in the lexicon. We compared MetaMap's ability to accurately identify the conditions against a physician's manual annotations and evaluated incorrectly indexed features to determine what additional processing is necessary. MetaMap identified the clinical conditions with a recall of 0.72 and a precision of 0.56. Necessary processing beyond MetaMap's indexing includes finding validation, temporal discrimination, anatomic location discrimination, finding-disease discrimination, and contextual inference. Successful identification of clinical conditions in an emergency department report with MetaMap requires processing techniques specific to the clinical question of interest.

Abstracting and Indexing↗

Summarization of an online medical encyclopedia.

We explore a knowledge-rich (abstraction) approach to summarization and apply it to multiple documents from an online medical encyclopedia. A semantic processor functions as the source interpreter and produces a list of predications. A transformation stage then generalizes and condenses this list, ultimately generating a conceptual condensate for a given disorder topic. We provide a preliminary evaluation of the quality of the condensates produced for a sample of four disorders. The overall precision of the disorder conceptual condensates was 87%, and the compression ratio from the base list of predications to the final condensate was 98%. The conceptual condensate could be used as input to a text generator to produce a natural language summary for a given disorder topic.

Disease↗

The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text.

Interpretation of semantic propositions in free-text documents such as MEDLINE citations would provide valuable support for biomedical applications, and several approaches to semantic interpretation are being pursued in the biomedical informatics community. In this paper, we describe a methodology for interpreting linguistic structures that encode hypernymic propositions, in which a more specific concept is in a taxonomic relationship with a more general concept. In order to effectively process these constructions, we exploit underspecified syntactic analysis and structured domain knowledge from the Unified Medical Language System (UMLS). After introducing the syntactic processing on which our system depends, we focus on the UMLS knowledge that supports interpretation of hypernymic propositions. We first use semantic groups from the Semantic Network to ensure that the two concepts involved are compatible; hierarchical information in the Metathesaurus then determines which concept is more general and which more specific. A preliminary evaluation of a sample based on the semantic group Chemicals and Drugs provides 83% precision. An error analysis was conducted and potential solutions to the problems encountered are presented. The research discussed here serves as a paradigm for investigating the interaction between domain knowledge and linguistic structure in natural language processing, and could also make a contribution to research on automatic processing of discourse structure. Additional implications of the system we present include its integration in advanced semantic interpretation processors for biomedical text and its use for information extraction in specific domains. The approach has the potential to support a range of applications, including information retrieval and ontology engineering.

Abstracting and Indexing↗

Integrating a hypernymic proposition interpreter into a semantic processor for biomedical texts.

Semantic processing provides the potential for producing high quality results in natural language processing (NLP) applications in the biomedical domain. In this paper, we address a specific semantic phenomenon, the hypernymic proposition, and concentrate on integrating the interpretation of such predications into a more general semantic processor in order to improve overall accuracy. A preliminary evaluation assesses the contribution of hypernymic propositions in providing more specific semantic predications and thus improving effectiveness in retrieving treatment propositions in MEDLINE abstracts. Finally, we discuss the generalization of this methodology to additional semantic propositions as well as other types of biomedical texts.

Abstracting and Indexing↗

Semantic relations asserting the etiology of genetic diseases.

Considerable research is being directed at extracting molecular biology information from text. Particularly challenging in this regard is to identify relations between entities, such as protein-protein interactions or molecular pathways. In this paper we present a natural language processing method for extracting causal relations between genetic phenomena and diseases. After presenting the results of preliminary evaluation, we suggest the use of a graphical display application for viewing the semantic predications produced by the system.

Computer Graphics↗

Interpreting hypernymic propositions in an online medical encyclopedia.

Interpretation of semantic propositions from bio-medical texts documents would provide valuable support to natural language processing (NLP) applications. We are developing a methodology to interpret a kind of semantic proposition, the hypernymic proposition, in MEDLINE abstracts. In this paper, we expanded the system to identify these structures in a different discourse domain: the Medical Encyclopedia from the National Library of Medi-cine's MEDLINEplus Website.

Encyclopedias as Topic↗

Toward (semi-)automatic generation of bio-medical ontologies.

The design and construction of domain specific ontologies and taxonomies requires allocation of huge resources in terms of cost and time. These efforts are human intensive and we need to explore ways of minimizing human involvement and other resources. In the biomedical domain, we seek to leverage resources such as the UMLS Metathesaurus and NLP-based applications such as MetaMap in conjunction with statistical clustering techniques, to (partially) automate the process. This is expected to be useful to the team involved in developing MeSH and other biomedical taxonomies to identify gaps in the existing taxonomies, and to be able to quickly bootstrap taxonomy generation for new research areas in biomedical informatics.

Algorithms↗

Assessing the consistency of a biomedical terminology through lexical knowledge.

OBJECTIVE: We investigate the use of adjectival modification as a way of assessing the systematic use of linguistic phenomena to represent similar lexical or semantic features in the constituent terms of a vocabulary. METHODS: Terms consisting of one or more adjectival modifiers followed by a head noun are selected from disease and procedure terms in SNOMED. Frequently co-occurring adjectival modifiers are systematically combined with the contexts (i.e., terms minus modifier) of each modifier. The existence of these combinations is checked in both SNOMED and the entire UMLS Metathesaurus; the term corresponding to the context alone is similarly checked. Relationships among terms sharing a context and between each of these terms and their context are studied. RESULTS: Four pairs of modifiers were studied: (acute, chronic), (unilateral, bilateral), (primary, secondary), and (acquired, congenital). The numbers of contexts studied for each pair ranged from 73 to 974. The percentage of contexts associated with both modifiers ranged from 5 to 50% in SNOMED and from 10 to 60% in UMLS. The presence of the context term varied from 31 to 64% in SNOMED and from 43 to 79% in UMLS. Finally, 172 occurrences (9%) of synonymy between a modified term and the context term were found in SNOMED. One hundred and forty-five such occurrences (8%) were found in the entire Metathesaurus.

Dictionaries as Topic↗

NLP-based information extraction for managing the molecular biology literature.

We present research aimed at devising a tool for using natural language processing to identify and extract biomedical information from text for the purpose of assisting researchers in molecular biology manage large amounts of information. A pilot project based on the molecular genetics of diabetes demonstrates our ability to explore the interaction of genomic phenomena and clinical findings. We suggest the cooperation of this extracted information with systems for clustering text and constructing labeled networks of data.

Diabetes Mellitus↗

Discovering protein similarity using natural language processing.

Extracting protein interaction relationships from textual repositories, such as MEDLINE, may prove useful in generating novel biological hypotheses. Using abstracts relevant to two known functionally related proteins, we modified an existing natural language processing tool to extract protein interaction terms. We were able to obtain functional information about two proteins, Amyloid Precursor Protein and Prion Protein, that have been implicated in the etiology of Alzheimer's Disease and Creutzfeldt-Jakob Disease, respectively.

Amyloid beta-Protein Precursor↗

Finding UMLS Metathesaurus concepts in MEDLINE.

The entire collection of 11.5 million MEDLINE abstracts was processed to extract 549 million noun phrases using a shallow syntactic parser. English language strings in the 2002 and 2001 releases of the UMLS Metathesaurus were then matched against these phrases using flexible matching techniques. 34% of the Metathesaurus names (occurring in 30% of the concepts) were found in the titles and abstracts of articles in the literature. The matching concepts are fairly evenly chemical and non-chemical in nature and span a wide spectrum of semantic types. This paper details the approach taken and the results of the analysis.

MEDLINE↗