PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Textpresso: an ontology-based information retrieval and extraction system for biological literature.

We have developed Textpresso, a new text-mining system for scientific literature whose capabilities go far beyond those of a simple keyword search engine. Textpresso's two major elements are a collection of the full text of scientific articles split into individual sentences, and the implementation of categories of terms for which a database of articles and individual sentences can be searched. The categories are classes of biological concepts (e.g., gene, allele, cell or cell group, phenotype, etc.) and classes that relate two objects (e.g., association, regulation, etc.) or describe one (e.g., biological process, etc.). Together they form a catalog of types of objects and concepts called an ontology. After this ontology is populated with terms, the whole corpus of articles and abstracts is marked up to identify terms of these categories. The current ontology comprises 33 categories of terms. A search engine enables the user to search for one or a combination of these tags and/or keywords within a sentence or document, and as the ontology allows word meaning to be queried, it is possible to formulate semantic queries. Full text access increases recall of biological data types from 45% to 95%. Extraction of particular biological facts, such as gene-gene interactions, can be accelerated significantly by ontologies, with Textpresso automatically performing nearly as well as expert curators to identify sentences; in searches for two uniquely named genes and an interaction term, the ontology confers a 3-fold increase of search efficiency. Textpresso currently focuses on Caenorhabditis elegans literature, with 3,800 full text articles and 16,000 abstracts. The lexicon of the ontology contains 14,500 entries, each of which includes all versions of a specific word or phrase, and it includes all categories of the Gene Ontology database. Textpresso is a useful curation tool, as well as search engine for researchers, and can readily be extended to other organism-specific corpora of text. Textpresso can be accessed at http://www.textpresso.org or via WormBase at http://www.wormbase.org.

Abstracting and Indexing↗

GO-Diff: mining functional differentiation between EST-based transcriptomes.

BACKGROUND: Large-scale sequencing efforts produced millions of Expressed Sequence Tags (ESTs) collectively representing differentiated biochemical and functional states. Analysis of these EST libraries reveals differential gene expressions, and therefore EST data sets constitute valuable resources for comparative transcriptomics. To translate differentially expressed genes into a better understanding of the underlying biological phenomena, existing microarray analysis approaches usually involve the integration of gene expression with Gene Ontology (GO) databases to derive comparable functional profiles. However, methods are not available yet to process EST-derived transcription maps to enable GO-based global functional profiling for comparative transcriptomics in a high throughput manner. RESULTS: Here we present GO-Diff, a GO-based functional profiling approach towards high throughput EST-based gene expression analysis and comparative transcriptomics. Utilizing holistic gene expression information, the software converts EST frequencies into EST Coverage Ratios of GO Terms. The ratios are then tested for statistical significances to uncover differentially represented GO terms between the compared transcriptomes, and functional differences are thus inferred. We demonstrated the validity and the utility of this software by identifying differentially represented GO terms in three application cases: intra-species comparison; meta-analysis to test a specific hypothesis; inter-species comparison. GO-Diff findings were consistent with previous knowledge and provided new clues for further discoveries. A comprehensive test on the GO-Diff results using series of comparisons between EST libraries of human and mouse tissues showed acceptable levels of consistency: 61% for human-human; 69% for mouse-mouse; 47% for human-mouse. CONCLUSION: GO-Diff is the first software integrating EST profiles with GO knowledge databases to mine functional differentiation between biological systems, e.g. tissues of the same species or the same tissue cross species. With rapid accumulation of EST resources in the public domain and expanding sequencing effort in individual laboratories, GO-Diff is useful as a screening tool before undertaking serious expression studies.

Animals↗

GOurmet: a tool for quantitative comparison and visualization of gene expression profiles based on gene ontology (GO) distributions.

BACKGROUND: The ever-expanding population of gene expression profiles (EPs) from specified cells and tissues under a variety of experimental conditions is an important but difficult resource for investigators to utilize effectively. Software tools have been recently developed to use the distribution of gene ontology (GO) terms associated with the genes in an EP to identify specific biological functions or processes that are over- or under-represented in that EP relative to other EPs. Additionally, it is possible to use the distribution of GO terms inherent to each EP to relate that EP as a whole to other EPs. Because GO term annotation is organized in a tree-like cascade of variable granularity, this approach allows the user to relate (e.g., by hierarchical clustering) EPs of varying length and from different platforms (e.g., GeneChip, SAGE, EST library). RESULTS: Here we present GOurmet, a software package that calculates the distribution of GO terms represented by the genes in an individual expression profile (EP), clusters multiple EPs based on these integrated GO term distributions, and provides users several tools to visualize and compare EPs. GOurmet is particularly useful in meta-analysis to examine EPs of specified cell types (e.g., tissue-specific stem cells) that are obtained through different experimental procedures. GOurmet also introduces a new tool, the Targetoid plot, which allows users to dynamically render the multi-dimensional relationships among individual elements in any clustering analysis. The Targetoid plotting tool allows users to select any element as the center of the plot, and the program will then represent all other elements in the cluster as a function of similarity to the selected central element. CONCLUSION: GOurmet is a user-friendly, GUI-based software package that greatly facilitates analysis of results generated by multiple EPs. The clustering analysis features a dynamic targetoid plot that is generalizable for use with any clustering application.

Artificial Intelligence↗

Combining evidence, biomedical literature and statistical dependence: new insights for functional annotation of gene sets.

BACKGROUND: Large-scale genomic studies based on transcriptome technologies provide clusters of genes that need to be functionally annotated. The Gene Ontology (GO) implements a controlled vocabulary organised into three hierarchies: cellular components, molecular functions and biological processes. This terminology allows a coherent and consistent description of the knowledge about gene functions. The GO terms related to genes come primarily from semi-automatic annotations made by trained biologists (annotation based on evidence) or text-mining of the published scientific literature (literature profiling). RESULTS: We report an original functional annotation method based on a combination of evidence and literature that overcomes the weaknesses and the limitations of each approach. It relies on the Gene Ontology Annotation database (GOA Human) and the PubGene biomedical literature index. We support these annotations with statistically associated GO terms and retrieve associative relations across the three GO hierarchies to emphasise the major pathways involved by a gene cluster. Both annotation methods and associative relations were quantitatively evaluated with a reference set of 7397 genes and a multi-cluster study of 14 clusters. We also validated the biological appropriateness of our hybrid method with the annotation of a single gene (cdc2) and that of a down-regulated cluster of 37 genes identified by a transcriptome study of an in vitro enterocyte differentiation model (CaCo-2 cells). CONCLUSION: The combination of both approaches is more informative than either separate approach: literature mining can enrich an annotation based only on evidence. Text-mining of the literature can also find valuable associated MEDLINE references that confirm the relevance of the annotation. Eventually, GO terms networks can be built with associative relations in order to highlight cooperative and competitive pathways and their connected molecular functions.

Algorithms↗

Various criteria in the evaluation of biomedical named entity recognition.

BACKGROUND: Text mining in the biomedical domain is receiving increasing attention. A key component of this process is named entity recognition (NER). Generally speaking, two annotated corpora, GENIA and GENETAG, are most frequently used for training and testing biomedical named entity recognition (Bio-NER) systems. JNLPBA and BioCreAtIvE are two major Bio-NER tasks using these corpora. Both tasks take different approaches to corpus annotation and use different matching criteria to evaluate system performance. This paper details these differences and describes alternative criteria. We then examine the impact of different criteria and annotation schemes on system performance by retesting systems participated in the above two tasks. RESULTS: To analyze the difference between JNLPBA's and BioCreAtIvE's evaluation, we conduct Experiment 1 to evaluate the top four JNLPBA systems using BioCreAtIvE's classification scheme. We then compare them with the top four BioCreAtIvE systems. Among them, three systems participated in both tasks, and each has an F-score lower on JNLPBA than on BioCreAtIvE. In Experiment 2, we apply hypothesis testing and correlation coefficient to find alternatives to BioCreAtIvE's evaluation scheme. It shows that right-match and left-match criteria have no significant difference with BioCreAtIvE. In Experiment 3, we propose a customized relaxed-match criterion that uses right match and merges JNLPBA's five NE classes into two, which achieves an F-score of 81.5%. In Experiment 4, we evaluate a range of five matching criteria from loose to strict on the top JNLPBA system and examine the percentage of false negatives. Our experiment gives the relative change in precision, recall and F-score as matching criteria are relaxed. CONCLUSION: In many applications, biomedical NEs could have several acceptable tags, which might just differ in their left or right boundaries. However, most corpora annotate only one of them. In our experiment, we found that right match and left match can be appropriate alternatives to JNLPBA and BioCreAtIvE's matching criteria. In addition, our relaxed-match criterion demonstrates that users can define their own relaxed criteria that correspond more realistically to their application requirements.

Algorithms↗

Automated linking of free-text complaints to reason-for-visit categories and International Classification of Diseases diagnoses in emergency department patient record databases.

STUDY OBJECTIVE: The use of the International Classification of Diseases system to describe emergency department (ED) case mix has disadvantages. We therefore developed computer algorithms that recognize a combination of words, word fragments, and word patterns to link free-text complaint fields to 20 reason-for-visit categories. We examine the feasibility and reliability of applying these reason-for-visit categories to ED patient-visit databases. METHODS: We analyzed a database (containing complaints and International Classification of Diseases diagnoses for 1 year's visits to a single ED) using a 3-step process (create initial terms, maximize sensitivity, maximize specificity) to define inclusion and exclusion terms for 20 reason-for-visit categories. To assess the reliability of the reason-for-visit assignment algorithm, we repeated the final 2 steps on a second database, composed of visits sampled from 21 EDs. For each database, we determined the prevalence of complaints that link to each reason-for-visit category and the distributions of International Classification of Diseases, Ninth Revision diagnoses that resulted for all patients and patients stratified by age. RESULTS: The 20 reason-for-visit categories capture 77% of all patients in database 1 (mean age 33.5 years) and 67% of all patients in database 2 (mean age 38.9 years). The percentage of visits captured by the 20 reason-for-visit categories, by age range, for databases 1 and 2 are (respectively) 0 to 2 years (84% and 76%), 3 to 10 years (82% and 74%), 11 to 65 years (76% and 68%), and 66 years or older (69% and 60%). The proportions of all complaints that link to each reason-for-visit category are largely similar between databases. Every complaint field that is linked to each reason-for-visit category includes at least 1 term that relates it to the category title, and the most frequently assigned diagnoses in each reason-for-visit category are those that one would expect to be associated with the reason-for-visit category complaints. CONCLUSION: The method by which free-text complaint fields are parsed into reason-for-visit categories is feasible and reasonably reliable; the finalized database 1 reason-for-visit category inclusion/exclusion terms lists required only modest changes to work well in database 2. The reason-for-visit categories used here are broadly defined to maximize the proportion of visits that they capture; more narrowly defined reason-for-visit categories will require more extensive revision of their inclusion/exclusion terms lists when used in different databases. A prospective, reason-for-visit-based ED classification system could have several useful applications (including syndromic surveillance), although content validity analysis will be necessary to investigate this hypothesis.

Adolescent↗

An algorithm to derive a numerical daily dose from unstructured text dosage instructions.

PURPOSE: The General Practice Research Database (GPRD) is a database of longitudinal patient records from general practices in the United Kingdom. It is an important data source for pharmacoepidemiology studies, but until now it has been tedious to calculate the daily dose and duration of exposure to drugs prescribed. This is because general practitioners routinely record dosage instructions as free text rather than in a structured way. The objective was to develop and assess the validity of an automated algorithm to derive the daily dose from text dosage instructions. METHODS: A computer program was developed to derive numerical information from unstructured text dosage instructions. It was tested on dosage texts from a random sample of one million prescription entries. A random sample of 1,000 of these converted texts were manually checked for their accuracy. RESULTS: Out of the sample of one million prescription entries, 74.5% had text containing the daily dose, 14.5% had text but did not include a quantitative daily dose statement and 11.0% had no text entered. Of the 1000 texts which were checked manually, 767 stated the daily dose. The program interpreted 758 (98.8%) of these correctly, produced errors in four cases and failed to extract the dose from five texts. CONCLUSIONS: An automated algorithm has been developed which can accurately extract the daily dose from almost 99% of general practitioners' text dosage instructions. It increases the utility of GPRD and other prescription data sources by enabling researchers to estimate the duration of drug exposure more efficiently.

Adverse Drug Reaction Reporting Systems↗

PhosphaBase: an ontology-driven database resource for protein phosphatases.

PhosphaBase is an ontology-driven database resource containing information on the protein phosphatase family. It is the first public resource dedicated to protein phosphatases, which are enzymes that perform dephosphorylation reactions. In conjunction with the phosphorylation action of protein kinases, phosphatases are involved in important control and communication mechanisms in the cell. They have also been implicated in many human diseases, including diabetes and obesity, cancers, and neurodegenerative conditions. PhosphaBase aims to centralize the growing base of knowledge in the phosphatase research domain. The resource is built around a formal, domain-specific DAML+OIL ontology, and the data are collected from heterogeneous biological sources using Gene Ontology terms as a means of data extraction. The overall ontology-driven architecture provides a robust structure with distinct advantages for sustainability and provides the potential for the development of diagnostic tools, as well as a data repository.

Animals↗

A comparison of classification algorithms to automatically identify chest X-ray reports that support pneumonia.

We compared the performance of expert-crafted rules, a Bayesian network, and a decision tree at automatically identifying chest X-ray reports that support acute bacterial pneumonia. We randomly selected 292 chest X-ray reports, 75 (25%) of which were from patients with a hospital discharge diagnosis of bacterial pneumonia. The reports were encoded by our natural language processor and then manually corrected for mistakes. The encoded observations were analyzed by three expert systems to determine whether the reports supported pneumonia. The reference standard for radiologic support of pneumonia was the majority vote of three physicians. We compared (a) the performance of the expert systems against each other and (b) the performance of the expert systems against that of four physicians who were not part of the gold standard. Output from the expert systems and the physicians was transformed so that comparisons could be made with both binary and probabilistic output. Metrics of comparison for binary output were sensitivity (sens), precision (prec), and specificity (spec). The metric of comparison for probabilistic output was the area under the receiver operator characteristic (ROC) curve. We used McNemar's test to determine statistical significance for binary output and univariate z-tests for probabilistic output. Measures of performance of the expert systems for binary (probabilistic) output were as follows: Rules--sens, 0.92; prec, 0.80; spec, 0.86 (Az, 0.960); Bayesian network--sens, 0.90; prec, 0.72; spec, 0.78 (Az, 0.945); decision tree--sens, 0.86; prec, 0.85; spec, 0.91 (Az, 0.940). Comparisons of the expert systems against each other using binary output showed a significant difference between the rules and the Bayesian network and between the decision tree and the Bayesian network. Comparisons of expert systems using probabilistic output showed no significant differences. Comparisons of binary output against physicians showed differences between the Bayesian network and two physicians. Comparisons of probabilistic output against physicians showed a difference between the decision tree and one physician. The expert systems performed similarly for the probabilistic output but differed in measures of sensitivity, precision, and specificity produced by the binary output. All three expert systems performed similarly to physicians.

Acute Disease↗

Selective automated indexing of findings and diagnoses in radiology reports.

The recent improvements in capabilities of desktop computers and communications networks give impetus for the development of clinical image repositories that can be used for patient care and medical education. A challenge in the use of these systems is the accurate indexing of images for retrieval performance acceptable to users. This paper describes a series of experiments aiming to adapt the SAPHIRE system, which matches text to concepts in the UMLS Metathesaurus, for the automated indexing of image reports. A series of enhancements to the baseline system resulted in a recall of 63% but a precision of only 30% in detecting concepts. At this level of performance, such a system might be problematic for users in a purely automated indexing environment. However, if the ability to retrieve images in repositories based on content in their reports is desired by clinical users, and no other current systems offer this functionality, then follow-up research questions include whether these imperfect results would be useful in a completely or partially automated indexing environment and/or whether other approaches can improve upon them.

Abstracting and Indexing↗

A simple algorithm for identifying negated findings and diseases in discharge summaries.

Narrative reports in medical records contain a wealth of information that may augment structured data for managing patient information and predicting trends in diseases. Pertinent negatives are evident in text but are not usually indexed in structured databases. The objective of the study reported here was to test a simple algorithm for determining whether a finding or disease mentioned within narrative medical reports is present or absent. We developed a simple regular expression algorithm called NegEx that implements several phrases indicating negation, filters out sentences containing phrases that falsely appear to be negation phrases, and limits the scope of the negation phrases. We compared NegEx against a baseline algorithm that has a limited set of negation phrases and a simpler notion of scope. In a test of 1235 findings and diseases in 1000 sentences taken from discharge summaries indexed by physicians, NegEx had a specificity of 94.5% (versus 85.3% for the baseline), a positive predictive value of 84.5% (versus 68.4% for the baseline) while maintaining a reasonable sensitivity of 77.8% (versus 88.3% for the baseline). We conclude that with little implementation effort a simple regular expression algorithm for determining whether a finding or disease is absent can identify a large portion of the pertinent negatives from discharge summaries.

Algorithms↗

TransMiner: mining transitive associations among biological objects from text.

Associations among biological objects such as genes, proteins, and drugs can be discovered automatically from the scientific literature. TransMiner is a system for finding associations among objects by mining the Medline database of the scientific literature. The direct associations among the objects are discovered based on the principle of co-occurrence in the form of an association graph. The principle of transitive closure is applied to the association graph to find potential transitive associations. The potential transitive associations that are indeed direct are discovered by iterative retrieval and mining of the Medline documents. Those associations that are not found explicitly in the entire Medline database are transitive associations and are the candidates for hypothesis generation. The transitive associations were ranked based on the sum of weight of terms that co-occur with both the objects. The direct and transitive associations are visualized using a graph visualization applet. TransMiner was tested by finding associations among 56 breast cancer genes and among 24 objects in the calpain signal transduction pathway. TransMiner was also used to rediscover associations between magnesium and migraine.

Abstracting and Indexing↗

The frequencies of disease names with the natural language used in the hospital information system.

The statistical behavior of disease names referred by physicians with the natural language in a large hospital information system is little known despite the theoretical and practical interest. To address this issue, we reviewed and investigated the usage-frequencies of 18,274 disease names, 10,288 for outpatient care and 7986 for inpatient care, referred from October 1983 to June 1992 with the notation of the natural language in Japanese by the use of the registration-retrieval system of disease names at Fukui Medical School, Japan. Consequently, we found that the investigated distributions did not conform to the Poisson distribution, but conformed well to the Polya-Eggenberger distribution in both cases of outpatient and inpatient care. It implies that the disease names with the natural language are possibly referred by physicians with some interrelations.

Disease↗

The implementation of speech recognition in an electronic radiology practice.

For both efficiency and economic reasons, our practice (200,000 examinations) has converted all remote dictation to speech recognition transcription (PowerScribe, L & H, Burlington, MA). The design criteria included complete automation to the existing radiology information system (RIS), with full RIS capabilities immediately available following dictation. All dictations for computed tomography, magnetic resonance imaging, ultrasound, and nuclear medicine were converted from remote transcription to speech recognition over a 2-week period (following a 4-week installation phase and 8 days of training). The average turnaround time for these reports decreased from approximately 2 hours to less than 1 minute. Reports are then sent to the institutional Electronic Medical Record and are available throughout all facilities in a nominal 2 minutes. Speech recognition rates were surprisingly high, although certain phrases caused consistent difficulties and certain staff required retraining. This presents our analysis of both successful and problematic areas during our design and implementation, as well as statistical performance analyses.

Efficiency↗

[Experiences with a current speech recognition system in creating cardiology reports].

Development of speech recognition software is at a stage where you can use it effectively for creating cardiological reports or at least parts of it. We were very successful in using the Dragon Naturally Speaking system for our reports and we don't need a secretary for our writings any longer. Very important for effective work is a fast and good PC hardware, especially a good sound system and a sufficient amount of internal memory, also important is patience of the user, because a longer training phase is required. I recommend for own experiences to use the standard version which is cheap and includes all necessary features. Following some fundamental rules anyone can be successful with speech recognition. Development of speech recognition is increasing rapidly so that everyone of us will get in contact with it sooner or later, it would be better to be prepared for it now.

Cardiology↗