PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Evaluating the UMLS as a source of lexical knowledge for medical language processing.

Medical language processing (MLP) systems rely on specialized lexicons in order to recognize, classify, and normalize medical terminology, and the performance of an MLP system is dependent on the coverage and quality of such lexicons. However, the acquisition of lexical knowledge is expensive and time-consuming. The UMLS is a comprehensive resource that can be used to acquire lexical knowledge needed for medical language processing. This paper describes methods that use these resources to automatically create lexical entries and generate two lexicons. The first lexicon was created primarily using the UMLS, whereas the second was created by supplementing the lexicon of an existing MLP system called MedLEE with entries based on the UMLS. We subsequently carried out a study, which is the primary focus of this paper, using MedLEE with each of the two lexicons and also the current MedLEE lexicon to measure performance. Overall accuracy, sensitivity, and specificity using the lexicon primarily based on the UMLS were.86,.60, and.96 respectively. Those measures using the MedLEE lexicon alone were.93,.81, and.93, which was significantly better except for specificity; performance using the supplemental lexicon was exactly the same as performance using solely the MedLEE lexicon.

Natural Language Processing↗

Analysis of biomedical text for chemical names: a comparison of three methods.

At the National Library of Medicine (NLM), a variety of biomedical vocabularies are found in data pertinent to its mission. In addition to standard medical terminology, there are specialized vocabularies including that of chemical nomenclature. Normal language tools including the lexically based ones used by the Unified Medical Language System (UMLS) to manipulate and normalize text do not work well on chemical nomenclature. In order to improve NLM's capabilities in chemical text processing, two approaches to the problem of recognizing chemical nomenclature were explored. The first approach was a lexical one and consisted of analyzing text for the presence of a fixed set of chemical segments. The approach was extended with general chemical patterns and also with terms from NLM's indexing vocabulary, MeSH, and the NLM SPECIALIST lexicon. The second approach applied Bayesian classification to n-grams of text via two different methods. The single lexical method and two statistical methods were tested against data from the 1999 UMLS Metathesaurus. One of the statistical methods had an overall classification accuracy of 97%.

Algorithms↗

Noise affects auditory and linguistic processing differently: an MEG study.

We investigated the influence of noise on brain responses to spoken sentences in MEG. Sixteen subjects had to listen to acoustically presented sentences and judge their syntactic correctness. Sentences were either presented on a silent background or with noise. Noise had differential effects on early auditory and syntactic processes. While noise affected early auditory processes only in the right hemisphere, noise had a general effect on syntactical processes. The evoked responses to syntactic violations compared with correct sentences, namely an early left anterior negativity, were significantly suppressed when noise was present The noise suppression effect, however, was not lateralized.

Acoustic Stimulation↗

TO-GO: a Java-based Gene Ontology navigation environment.

UNLABELLED: TO-GO is a Gene Ontology (GO) navigation tool, which is implemented as a Java application. After the initial data downloading, the GO term tree can be interactively navigated without further network transfer. Local annotation can be incorporated. It supports querying by GO terms or associated gene product information, displaying the result as a table or a sub-tree. The result from the search for a set of external database accessions includes the number of gene products associated with each node, inclusive of sub-nodes. Search results can be further processed by set operations and these set operations can be quite useful for expression profile data analysis. A copy/paste function is also implemented in order to facilitate data exchange between applications. AVAILABILITY: TO-GO is freely available at http://www.ngic.re.kr/togo/index.html CONTACT: ungsik@kribb.re.kr

Algorithms↗

Text mining and its potential applications in systems biology.

With biomedical literature increasing at a rate of several thousand papers per week, it is impossible to keep abreast of all developments; therefore, automated means to manage the information overload are required. Text mining techniques, which involve the processes of information retrieval, information extraction and data mining, provide a means of solving this. By adding meaning to text, these techniques produce a more structured analysis of textual knowledge than simple word searches, and can provide powerful tools for the production and analysis of systems biology models.

Artificial Intelligence↗

The MIPS mammalian protein-protein interaction database.

SUMMARY: The MIPS mammalian protein-protein interaction database (MPPI) is a new resource of high-quality experimental protein interaction data in mammals. The content is based on published experimental evidence that has been processed by human expert curators. We provide the full dataset for download and a flexible and powerful web interface for users with various requirements.

Animals↗

Evaluation of Meta-1 for a concept-based approach to the automated indexing and retrieval of bibliographic and full-text databases.

SAPHIRE is a concept-based approach to information retrieval in the biomedical domain. Indexing and retrieval are based on a concept-matching algorithm that processes free text to identify concepts and map them to their canonical form. This process requires a large vocabulary containing a breadth of medical concepts and a diversity of synonym forms, which is provided by the Meta-1 vocabulary from the Unified Medical Language System Project of the National Library of Medicine. This paper describes the use of Meta-1 in SAPHIRE and an evaluation of both entities in the context of an information retrieval study.

Abbreviations as Topic↗

Gene Ontology friendly biclustering of expression profiles.

The soundness of clustering in the analysis of gene expression profiles and gene function prediction is based on the hypothesis that genes with similar expression profiles may imply strong correlations with their functions in the biological activities. Gene Ontology (GO) has become a well accepted standard in organizing gene function categories. Different gene function categories in GO can have very sophisticated relationships, such as 'part of' and 'overlapping'. Until now, no clustering algorithm can generate gene clusters within which the relationships can naturally reflect those of gene function categories in the GO hierarchy. The failure in resembling the relationships may reduce the confidence of clustering in gene function prediction. In this paper, we present a new clustering technique, Smart Hierarchical Tendency Preserving clustering (SHTP-clustering), based on a bicluster model, Tendency Preserving cluster (TP-Cluster). By directly incorporating Gene Ontology information into the clustering process, the SHTP-clustering algorithm yields a TP-cluster tree within which any subtree can be well mapped to a part of the GO hierarchy. Our experiments on yeast cell cycle data demonstrate that this method is efficient and effective in generating the biological relevant TP-Clusters.

Algorithms↗

Identification of biological relationships from text documents using efficient computational methods.

The biological literature databases continue to grow rapidly with vital information that is important for conducting sound biomedical research and development. The current practices of manually searching for information and extracting pertinent knowledge are tedious, time-consuming tasks even for motivated biological researchers. Accurate and computationally efficient approaches in discovering relationships between biological objects from text documents are important for biologists to develop biological models. The term "object" refers to any biological entity such as a protein, gene, cell cycle, etc. and relationship refers to any dynamic action one object has on another, e.g. protein inhibiting another protein or one object belonging to another object such as, the cells composing an organ. This paper presents a novel approach to extract relationships between multiple biological objects that are present in a text document. The approach involves object identification, reference resolution, ontology and synonym discovery, and extracting object-object relationships. Hidden Markov Models (HMMs), dictionaries, and N-Gram models are used to set the framework to tackle the complex task of extracting object-object relationships. Experiments were carried out using a corpus of one thousand Medline abstracts. Intermediate results were obtained for the object identification process, synonym discovery, and finally the relationship extraction. For the thousand abstracts, 53 relationships were extracted of which 43 were correct, giving a specificity of 81 percent. These results are promising for multi-object identification and relationship finding from biological documents.

Algorithms↗

IDEM: a Web application of case-based reasoning in histopathology.

Different software engineering and artificial intelligence methods can be used to design Internet retrieval of prototypical medical images. We used the case-based reasoning (CBR) approach to provide an 'intelligent' access to a collection of illustrated medical cases through the Internet. This paper presents a Web interface for the CBR system IDEM (image and diagnosis from examples in medicine) in the domain of breast pathology. Thanks to the definition of a similarity measure between the descriptions of cases we propose a flexible querying of the case-base and a quantitative browsing among cases through similarity links. The resemblance rates provided by the system argue for the quality and the relevancy of the retrieved data. The flexibility of the querying process is robust to missing information and could be adapted to a daily practice. The CBR approach is a promising method for a clinical relevant and an efficient retrieval of reference images and diagnosis clues through Internet.

Artificial Intelligence↗

Data mining in bioinformatics using Weka.

UNLABELLED: The Weka machine learning workbench provides a general-purpose environment for automatic classification, regression, clustering and feature selection-common data mining problems in bioinformatics research. It contains an extensive collection of machine learning algorithms and data pre-processing methods complemented by graphical user interfaces for data exploration and the experimental comparison of different machine learning techniques on the same problem. Weka can process data given in the form of a single relational table. Its main objectives are to (a) assist users in extracting useful information from data and (b) enable them to easily identify a suitable algorithm for generating an accurate predictive model from it. AVAILABILITY: http://www.cs.waikato.ac.nz/ml/weka.

Algorithms↗

Techniques for optimization of queries on integrated biological resources.

Today, scientific data are inevitably digitized, stored in a wide variety of formats, and are accessible over the Internet. Scientific discovery increasingly involves accessing multiple heterogeneous data sources, integrating the results of complex queries, and applying further analysis and visualization applications in order to collect datasets of interest. Building a scientific integration platform to support these critical tasks requires accessing and manipulating data extracted from flat files or databases, documents retrieved from the Web, as well as data that are locally materialized in warehouses or generated by software. The lack of efficiency of existing approaches can significantly affect the process with lengthy delays while accessing critical resources or with the failure of the system to report any results. Some queries take so much time to be answered that their results are returned via email, making their integration with other results a tedious task. This paper presents several issues that need to be addressed to provide seamless and efficient integration of biomolecular data. Identified challenges include: capturing and representing various domain specific computational capabilities supported by a source including sequence or text search engines and traditional query processing; developing a methodology to acquire and represent semantic knowledge and metadata about source contents, overlap in source contents, and access costs; developing cost and semantics based decision support tools to select sources and capabilities, and to generate efficient query evaluation plans.

Algorithms↗

A multi-level text mining method to extract biological relationships.

Accurate and computationally efficient approaches in discovering relationships between biological objects from text documents are important for biologists to develop biological models. This paper presents a novel approach to extract relationships between multiple biological objects that are present in a text document. The approach involves object identification, reference resolution, ontology and synonym discovery, and extracting object-object relationships. Hidden Markov Models (HMMs), dictionaries, and N-Gram models are used to set the framework to tackle the complex task of extracting object-object relationships. Experiments were carried out using a corpus of one thousand Medline abstracts. Intermediate results were obtained for the object identification process, synonym discovery, and finally the relationship extraction. For a corpus of thousand abstracts, 53 relationships were extracted of which 43 were correct, giving a specificity of 81%. The approach is both adaptable and scalable to new problems as opposed to rule-based methods.

Abstracting and Indexing↗

Expert knowledge without the expert: integrated analysis of gene expression and literature to derive active functional contexts.

MOTIVATION: The interpretation of expression data without appropriate expert knowledge is difficult and usually limited to exploratory data analysis, such as clustering and detecting differentially regulated genes. However, comparing experimental results against manually compiled knowledge resources might limit or bias the perspective on the data. Thus, manual analysis by experts is required to obtain confident predictions about involved processes. RESULTS: We present an algorithm to simultaneously derive interpretations of expression measurements together with biological hypotheses from biomedical publications. It identifies active functional contexts ('concepts'), i.e. gene clusters that exhibit both a significant gene expression as well as a coherent literature profile. Manual intervention by an expert in specifying prior knowledge is not required. The approach scales to realistic applications and does not rely on controlled vocabularies or pathway resources. We validated our algorithm by analyzing a current juvenile arthritis dataset. A number of gene clusters and accompanying literature topics are identified as an interpretation of the data that coincide well with the phenotype and biological processes known to be involved in the disease. We demonstrate that generated clusters are both more sensitive and more specific than Gene Ontology categories detected on the same data. The method allows for in-depth investigation of subsets of genes, the associated literature topics and publications. AVAILABILITY: Supplementary data on clusters is available upon request.

Expert Systems↗

Designing metaschemas for the UMLS enriched semantic network.

The enriched semantic network (ESN) has previously been presented as an enhancement of the semantic network (SN) of the UMLS. The ESN's hierarchy is a DAG (Directed Acyclic Graph) structure allowing for multiple parents. The ESN is thus more complex than the SN and can be more difficult to view and comprehend. We have previously introduced the notion of a metaschema for the SN as a compact abstraction to support SN comprehension. We extend the definition of metaschema to make it applicable to a DAG classification hierarchy, such as the one exhibited by the ESN. We specify the requirements for and describe the general process of deriving such a metaschema. We derive two particular metaschemas of the ESN based on a pair of partitions. These two metaschemas and their underlying partitions are compared. Both metaschemas serve as compact representations of the ESN, allowing for convenient viewing of its hierarchy and easier comprehension.

Abstracting and Indexing↗

MediClass: A system for detecting and classifying encounter-based clinical events in any electronic medical record.

MediClass is a knowledge-based system that processes both free-text and coded data to automatically detect clinical events in electronic medical records (EMRs). This technology aims to optimize both clinical practice and process control by automatically coding EMR contents regardless of data input method (e.g., dictation, structured templates, typed narrative). We report on the design goals, implemented functionality, generalizability, and current status of the system. MediClass could aid both clinical operations and health services research through enhancing care quality assessment, disease surveillance, and adverse event detection.

Artificial Intelligence↗

Text mining of DNA sequence homology searches.

Primary tasks in analysis and annotation of expressed sequence tag (EST) datasets are to identify similarity among sequences by unsupervised clustering and assign putative function based on BLAST homology searches. We investigated the usefulness of text mining as a simple approach for further higher-level clustering of EST datasets using IBM Intelligent Miner for Text v2.3 tools. Agglomerative and k-means clustering tools were used to cluster BLASTx homology search documents from two onion EST datasets and optimised by pre-processing and pruning. Subjective evaluation confirmed that these tools provided biologically useful and complementary views of the two libraries, provided new insights into their composition and revealed clusters previously identified by human experts. We compared BLASTx textual clusters for two gene families with their DNA sequence-based clusters and confirmed that these shared similar morphology.

Abstracting and Indexing↗

Automating concept identification in the electronic medical record: an experiment in extracting dosage information.

We discuss the development and evaluation of an automated procedure for extracting drug-dosage information from clinical narratives. The process was developed rapidly using existing technology and resources, including categories of terms from UMLS96. Evaluations over a large training and smaller test set of medical records demonstrate an approximately 80% rate of exact and partial matches' on target phrases, with few false positives and a modest rate of false negatives. The results suggest a strategy for automating general concept identification in electronic medical records.

Classification↗