PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Methods for automated concept mapping between medical databases.

The retrieval and exchange of information between medical databases is often impeded by the semantic heterogeneity of concepts contained within the databases. Manual identification of equivalent database elements consumes time and resources, and may often be the rate-limiting technological step in integrating disparate data sources. By employing semantic networks as an intermediary representation of the native databases, automated mapping algorithms can identify equivalent concepts in disparate databases. The algorithms take advantage of the conceptual "context" embodied within a semantic network to produce candidate concept mappings. The performance of automated concept mapping was evaluated by creating semantic network representations for two test laboratory databases. The mapping algorithms identified all equivalent concepts that were present in the databases, and did not leave any equivalent concepts unmapped. The utilization of conceptual context to perform automated concept mapping facilitates the identification of equivalent database concepts and may help decrease the work and costs associated with retrieval and integration of information from disparate databases.

Algorithms↗

GLIF3: a representation format for sharable computer-interpretable clinical practice guidelines.

The Guideline Interchange Format (GLIF) is a model for representation of sharable computer-interpretable guidelines. The current version of GLIF (GLIF3) is a substantial update and enhancement of the model since the previous version (GLIF2). GLIF3 enables encoding of a guideline at three levels: a conceptual flowchart, a computable specification that can be verified for logical consistency and completeness, and an implementable specification that is intended to be incorporated into particular institutional information systems. The representation has been tested on a wide variety of guidelines that are typical of the range of guidelines in clinical use. It builds upon GLIF2 by adding several constructs that enable interpretation of encoded guidelines in computer-based decision-support systems. GLIF3 leverages standards being developed in Health Level 7 in order to allow integration of guidelines with clinical information systems. The GLIF3 specification consists of an extensible object-oriented model and a structured syntax based on the resource description framework (RDF). Empirical validation of the ability to generate appropriate recommendations using GLIF3 has been tested by executing encoded guidelines against actual patient data. GLIF3 is accordingly ready for broader experimentation and prototype use by organizations that wish to evaluate its ability to capture the logic of clinical guidelines, to implement them in clinical systems, and thereby to provide integrated decision support to assist clinicians.

Artificial Intelligence↗

A framework for a distributed, hybrid, multiple-ontology clinical-guideline library, and automated guideline-support tools.

Clinical guidelines are a major tool in improving the quality of medical care. However, most guidelines are in free text, not in a formal, executable format, and are not easily accessible to clinicians at the point of care. We introduce a Web-based, modular, distributed architecture, the Digital Electronic Guideline Library (DeGeL), which facilitates gradual conversion of clinical guidelines from text to a formal representation in chosen target guideline ontology. The architecture supports guideline classification, semantic markup, context-sensitive search, browsing, run-time application, and retrospective quality assessment. The DeGeL hybrid meta-ontology includes elements common to all guideline ontologies, such as semantic classification and domain knowledge; it also includes four content-representation formats: free text, semi-structured text, semi-formal representation, and a formal representation. These formats support increasingly sophisticated computational tasks. The DeGeL tools for support of guideline-based care operate, at some level, on all guideline ontologies. We have demonstrated the feasibility of the architecture and the tools for several guideline ontologies, including Asbru and GEM.

Computer Communication Networks↗

Enhancing HMM-based biomedical named entity recognition by studying special phenomena.

The purpose of this research is to enhance an HMM-based named entity recognizer in the biomedical domain. First, we analyze the characteristics of biomedical named entities. Then, we propose a rich set of features, including orthographic, morphological, part-of-speech, and semantic trigger features. All these features are integrated via a Hidden Markov Model with back-off modeling. Furthermore, we propose a method for biomedical abbreviation recognition and two methods for cascaded named entity recognition. Evaluation on the GENIA V3.02 and V1.1 shows that our system achieves 66.5 and 62.5 F-measure, respectively, and outperforms the previous best published system by 8.1 F-measure on the same experimental setting. The major contribution of this paper lies in its rich feature set specially designed for biomedical domain and the effective methods for abbreviation and cascaded named entity recognition. To our best knowledge, our system is the first one that copes with the cascaded phenomena.

Abbreviations as Topic↗

Using name-internal and contextual features to classify biological terms.

There has been considerable work done recently in recognizing named entities in biomedical text. In this paper, we investigate the named entity classification task, an integral part of the named entity extraction task. We focus on the different sources of information that can be utilized for classification, and note the extent to which they are effective in classification. To classify a name, we consider features that appear within the name as well as nearby phrases. We also develop a new strategy based on the context of occurrence and show that they improve the performance of the classification system. We show how our work relates to previous works on named entity classification in the biological domain as well as to those in generic domains. The experiments were conducted on the GENIA corpus Ver. 3.0 developed at University of Tokyo. We achieve f value of 86 in 10-fold cross validation evaluation on this corpus.

Abstracting and Indexing↗

Comparison of character-level and part of speech features for name recognition in biomedical texts.

The immense volume of data which is now available from experiments in molecular biology has led to an explosion in reported results most of which are available only in unstructured text format. For this reason there has been great interest in the task of text mining to aid in fact extraction, document screening, citation analysis, and linkage with large gene and gene-product databases. In particular there has been an intensive investigation into the named entity (NE) task as a core technology in all of these tasks which has been driven by the availability of high volume training sets such as the GENIA v3.02 corpus. Despite such large training sets accuracy for biology NE has proven to be consistently far below the high levels of performance in the news domain where F scores above 90 are commonly reported which can be considered near to human performance. We argue that it is crucial that more rigorous analysis of the factors that contribute to the model's performance be applied to discover where the underlying limitations are and what our future research direction should be. Our investigation in this paper reports on variations of two widely used feature types, part of speech (POS) tags and character-level orthographic features, and makes a comparison of how these variations influence performance. We base our experiments on a proven state-of-the-art model, support vector machines using a high quality subset of 100 annotated MEDLINE abstracts. Experiments reveal that the best performing features are orthographic features with F score of 72.6. Although the Brill tagger trained in-domain on the GENIA v3.02p POS corpus gives the best overall performance of any POS tagger, at an F score of 68.6, this is still significantly below the orthographic features. In combination these two features types appear to interfere with each other and degrade performance slightly to an F score of 72.3.

Abbreviations as Topic↗

Prospective recruitment of patients with congestive heart failure using an ad-hoc binary classifier.

This paper addresses a very specific problem of identifying patients diagnosed with a specific condition for potential recruitment in a clinical trial or an epidemiological study. We present a simple machine learning method for identifying patients diagnosed with congestive heart failure and other related conditions by automatically classifying clinical notes dictated at Mayo Clinic. This method relies on an automatic classifier trained on comparable amounts of positive and negative samples of clinical notes previously categorized by human experts. The documents are represented as feature vectors, where features are a mix of demographic information as well as single words and concept mappings to MeSH and HICDA classification systems. We compare two simple and efficient classification algorithms (Naïve Bayes and Perceptron) and a baseline term spotting method with respect to their accuracy and recall on positive samples. Depending on the test set, we find that Naïve Bayes yields better recall on positive samples (95 vs. 86%) but worse accuracy than Perceptron (57 vs. 65%). Both algorithms perform better than the baseline with recall on positive samples of 71% and accuracy of 54%.

Artificial Intelligence↗

Inductive creation of an annotation schema for manually indexing clinical conditions from emergency department reports.

Evaluating automated indexing applications requires comparing automatically indexed terms against manual reference standard annotations. However, there are no standard guidelines for determining which words from a textual document to include in manual annotations, and the vague task can result in substantial variation among manual indexers. We applied grounded theory to emergency department reports to create an annotation schema representing syntactic and semantic variables that could be annotated when indexing clinical conditions. We describe the annotation schema, which includes variables representing medical concepts (e.g., symptom, demographics), linguistic form (e.g., noun, adjective), and modifier types (e.g., anatomic location, severity). We measured the schema's quality and found: (1) the schema was comprehensive enough to be applied to 20 unseen reports without changes to the schema; (2) agreement between author annotators applying the schema was high, with an F measure of 93%; and (3) the authors made complementary errors when applying the schema, demonstrating that the schema incorporates both linguistic and medical expertise.

Abstracting and Indexing↗

Towards new information resources for public health--from WordNet to MedicalWordNet.

In the last two decades, WordNet has evolved as the most comprehensive computational lexicon of general English. In this article, we discuss its potential for supporting the creation of an entirely new kind of information resource for public health, viz. MedicalWordNet. This resource is not to be conceived merely as a lexical extension of the original WordNet to medical terminology; indeed, there is already a considerable degree of overlap between WordNet and the vocabulary of medicine. Instead, we propose a new type of repository, consisting of three large collections of (1) medically relevant word forms, structured along the lines of the existing Princeton WordNet; (2) medically validated propositions, referred to here as medical facts, which will constitute what we shall call MedicalFactNet; and (3) propositions reflecting laypersons' medical beliefs, which will constitute what we shall call the MedicalBeliefNet. We introduce a methodology for setting up the MedicalWordNet. We then turn to the discussion of research challenges that have to be met to build this new type of information resource. We build a database of sentences relevant to the medical domain. The sentences are generated from WordNet via its relations as well as from medical statements broken down into elementary propositions. Two subcorpora of sentences are distinguished, MedicalBeliefNet and MedicalFactNet. The former is rated for assent by laypersons; the latter for correctness by medical experts. The sentence corpora will be valuable for a variety of applications in information retrieval as well as in research in linguistics and psychology with respect to the study of expert and non-expert beliefs and their linguistic expressions. Our work has to meet several considerable challenges. These include accounting for the distinction between medical experts and laypersons, the social issues of expert-layperson communication in different media, the linguistic aspects of encoding medical knowledge, and the reliability, volume, and emergence of medical knowledge. The work described here has been tested in a small pilot experiment and awaits large-scale implementation.

Computational Biology↗

Using statistical and knowledge-based approaches for literature-based discovery.

The explosive growth in biomedical literature has made it difficult for researchers to keep up with advancements, even in their own narrow specializations. While researchers formulate new hypotheses to test, it is very important for them to identify connections to their work from other parts of the literature. However, the current volume of information has become a great barrier for this task and new automated tools are needed to help researchers identify new knowledge that bridges gaps across distinct sections of the literature. In this paper, we present a literature-based discovery system called LitLinker that incorporates knowledge-based methodologies with a statistical method to mine the biomedical literature for new, potentially causal connections between biomedical terms. We demonstrate LitLinker's ability to capture novel and interesting connections between diseases and chemicals, drugs, genes, or molecular sequences from the published biomedical literature. We also evaluate LitLinker's performance by using the information retrieval metrics of precision and recall.

Abstracting and Indexing↗

Inter-patient distance metrics using SNOMED CT defining relationships.

BACKGROUND: Patient-based similarity metrics are important case-based reasoning tools which may assist with research and patient care applications. Ontology and information content principles may be potentially helpful tools for similarity metric development. METHODS: Patient cases from 1989 through 2003 from the Columbia University Medical Center data repository were converted to SNOMED CT concepts. Five metrics were implemented: (1) percent disagreement with data as an unstructured "bag of findings," (2) average links between concepts, (3) links weighted by information content with descendants, (4) links weighted by information content with term prevalence, and (5) path distance using descendants weighted by information content with descendants. Three physicians served as gold standard for 30 cases. RESULTS: Expert inter-rater reliability was 0.91, with rank correlations between 0.61 and 0.81, representing upper-bound performance. Expert performance compared to metrics resulted in correlations of 0.27, 0.29, 0.30, 0.30, and 0.30, respectively. Using SNOMED axis Clinical Findings alone increased correlation to 0.37. CONCLUSION: Ontology principles and information content provide useful information for similarity metrics but currently fall short of expert performance.

Algorithms↗

Using MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles.

Biomedical abbreviations and acronyms are widely used in biomedical literature. Since many of them represent important content in biomedical literature, information retrieval and extraction benefits from identifying the meanings of those terms. On the other hand, many abbreviations and acronyms are ambiguous, it would be important to map them to their full forms, which ultimately represent the meanings of the abbreviations. In this study, we present a semi-supervised method that applies MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles. We first automatically generated from the MEDLINE abstracts a dictionary of abbreviation-full pairs based on a rule-based system that maps abbreviations to full forms when full forms are defined in the abstracts. We then trained on the MEDLINE abstracts and predicted the full forms of abbreviations in full-text journal articles by applying supervised machine-learning algorithms in a semi-supervised fashion. We report up to 92% prediction precision and up to 91% coverage.

Artificial Intelligence↗

Enrichment of OBO ontologies.

This paper describes a frame-based integration of the three GO subontologies, the Chemical Entities of Biological Interest ontology, and the Cell Type Ontology in which relationships are modeled in a way that better captures the semantics between biological concepts represented by the terms, rather than between the terms themselves, than previous frame-based efforts. We also describe a methodology for creating suggested enriching assertions by identifying patterns in GO terms, mapping these patterns to new, specific relationships, and matching term substrings to concepts. Using this methodology, a predicted assertion was made for 62% of GO terms that matched one of 31 patterns, and 97% of these predicted assertions were assessed to be valid, resulting in an initial set of over 4000 assertions. Furthermore, this methodology programmatically integrates assertions into an ontology such that each assertion is fully consistent with respect to higher (i.e., more general) relevant class and slot levels.

Computational Biology↗

Evaluation of techniques for increasing recall in a dictionary approach to gene and protein name identification.

Gene and protein name identification in text requires a dictionary approach to relate synonyms to the same gene or protein, and to link names to external databases. However, existing dictionaries are incomplete. We investigate two complementary methods for automatic generation of a comprehensive dictionary: combination of information from existing gene and protein databases and rule-based generation of spelling variations. Both methods have been reported in literature before, but have hitherto not been combined and evaluated systematically. We combined gene and protein names from several existing databases of four different organisms. The combined dictionaries showed a substantial increase in recall on three different test sets, as compared to any single database. Application of 23 spelling variation rules to the combined dictionaries further increased recall. However, many rules appeared to have no effect and some appear to have a detrimental effect on precision.

Abstracting and Indexing↗

ASR for emotional speech: clarifying the issues and enhancing performance.

There are multiple reasons to expect that recognising the verbal content of emotional speech will be a difficult problem, and recognition rates reported in the literature are in fact low. Including information about prosody improves recognition rate for emotions simulated by actors, but its relevance to the freer patterns of spontaneous speech is unproven. This paper shows that recognition rate for spontaneous emotionally coloured speech can be improved by using a language model based on increased representation of emotional utterances. The models are derived by adapting an already existing corpus, the British National Corpus (BNC). An emotional lexicon is used to identify emotionally coloured words, and sentences containing these words are recombined with the BNC to form a corpus with a raised proportion of emotional material. Using a language model based on that technique improves recognition rate by about 20%.

Emotions↗

Syntactic-semantic tagging as a mediator between linguistic representations and formal models: an exercise in linking SNOMED to GALEN.

Natural language understanding applications are good candidates to solve the knowledge acquisition bottleneck when designing large scale concept systems. However, a necessary condition is that systems are built that transform sentences into a meaning representation that is independent of the subtleties of linguistic structure that nevertheless underly the way language works. The Cassandra II syntactic-semantic tagging system fulfills this goal partially. Within the GALEN-IN-USE project, it is used to transform linguistic representations of surgical procedure expressions into conceptual representations. In this paper, the proctology chapter of the SNOMED V3.1 procedure axis was used as a testbed to evaluate the usefulness of this approach. A quantitative and qualitative analysis of the data obtained is presented, showing that the Cassandra system can indeed complement the manual modelling efforts being conducted in the GALEN-IN-USE project. The different requirements related to linguistic modelling versus conceptual modelling can partly be accounted for by using an interface ontology, of which the fine tuning will however remain an important effort.

Artificial Intelligence↗

Text-based knowledge discovery: search and mining of life-sciences documents.

Text literature is playing an increasingly important role in biomedical discovery. The challenge is to manage the increasing volume, complexity and specialization of knowledge expressed in this literature. Although information retrieval or text searching is useful, it is not sufficient to find specific facts and relations. Information extraction methods are evolving to extract automatically specific, fine-grained terms corresponding to the names of entities referred to in the text, and the relationships that connect these terms. Information extraction is, in turn, a means to an end, and knowledge discovery methods are evolving for the discovery of still more-complex structures and connections among facts. These methods provide an interpretive context for understanding the meaning of biological data.

Biological Science Disciplines↗

Power of expression in the electronic patient record: structured data or narrative text?

This paper presents the authors' experience with the development and use of a document-centered electronic patient record (EPR) in a large teaching hospital. The development of the document-centered EPR began with the formulation of a set of critical hypotheses to facilitate both the continuation of the best medical practice and the implementation and use of the EPR. An alternate and more conventional approach - the data-centered EPR - is compared with the document-centered EPR. Various benefits and pitfalls are discussed. Finally, the choice was to offer both solutions in a tightly linked system. The need for an EPR which combines the document and data centered approaches is a reflection of the more general discussion of what the medical record will be in the future. All too often, the need for structured data conflicts with the need for free texts and the power of expression. It is not easy to evaluate the consequences of this initial decision. However, changing the foundations of the EPR after its implementation is difficult and expensive. Therefore, the selection of the correct orientation in a given hospital requires a broad-based discussion.

Hospitals, Teaching↗