PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

THEA: ontology-driven analysis of microarray data.

MOTIVATION: Microarray technology makes it possible to measure thousands of variables and to compare their values under hundreds of conditions. Once microarray data are quantified, normalized and classified, the analysis phase is essentially a manual and subjective task based on visual inspection of classes in the light of the vast amount of information available. Currently, data interpretation clearly constitutes the bottleneck of such analyses and there is an obvious need for tools able to fill the gap between data processed with mathematical methods and existing biological knowledge. RESULTS: THEA (Tools for High-throughput Experiments Analysis) is an integrated information processing system allowing convenient handling of data. It allows to automatically annotate data issued from classification systems with selected biological information coming from a knowledge base and to either manually search and browse through these annotations or automatically generate meaningful generalizations according to statistical criteria (data mining). AVAILABILITY: The software is available on the website http://thea.unice.fr/

Abstracting and Indexing↗

On the retranslation process in Zadeh's paradigm of computing with words.

We discuss Zadeh's paradigm of computing with words and indicate the three important stages. We focus on the retranslation process, selecting a term from our prescribed vocabulary to express information represented using fuzzy sets. A number of criteria of concern in this retranslation process are introduced. Some of these criteria can be seen to correspond to a desire to accurately reflect the given information. Other criteria may correspond to a desire, on the part of the provider of the information, to give a particular perception or "spin." These types of criteria can be of particular importance in many types of information warfare. We discuss some methods for combining these criteria to evaluate potential retranslations.

Algorithms↗

DIG--a system for gene annotation and functional discovery.

SUMMARY: We describe a database and information discovery system named DIG (Duke Integrated Genomics) designed to facilitate the process of gene annotation and the discovery of functional context. The DIG system collects and organizes gene annotation and functional information, and includes tools that support an understanding of genes in a functional context by providing a framework for integrating and visualizing gene expression, protein interaction and literature-based interaction networks.

Chromosome Mapping↗

Conceptual approach for the design of radiology reporting interfaces: the talking template.

Within the coming decade, traditional dictation supported by human transcription for radiology reports will be replaced by one or more computerized methods. This paper discusses the cognitive and process efficiency problems arising from currently available technology including speech recognition and menu-driven interfaces. A specific concept for interaction with the reporting interface is proposed. This is called the "talking template" and departs from other designs by providing for all interactions to be mediated through audible prompts and microphone controls. The radiologist can recapture efficiency and cognitive focus by dictating while viewing images without the "look away" problem inherent in other interfaces.

Cognition↗

Aggregating automatically extracted regulatory pathway relations.

Automatic tools to extract information from biomedical texts are needed to help researchers leverage the vast and increasing body of biomedical literature. While several biomedical relation extraction systems have been created and tested, little work has been done to meaningfully organize the extracted relations. Organizational processes should consolidate multiple references to the same objects over various levels of granularity, connect those references to other resources, and capture contextual information. We propose a feature decomposition approach to relation aggregation to support a five-level aggregation framework. Our BioAggregate tagger uses this approach to identify key features in extracted relation name strings. We show encouraging feature assignment accuracy and report substantial consolidation in a network of extracted relations.

Artificial Intelligence↗

Using regular expressions to abstract blood pressure and treatment intensification information from the text of physician notes.

This case study examined the utility of regular expressions to identify clinical data relevant to the epidemiology of treatment of hypertension. We designed a software tool that employed regular expressions to identify and extract instances of documented blood pressure values and anti-hypertensive treatment intensification from the text of physician notes. We determined sensitivity, specificity and precision of identification of blood pressure values and anti-hypertensive treatment intensification using a gold standard of manual abstraction of 600 notes by two independent reviewers. The software processed 370 Mb of text per hour, and identified elevated blood pressure documented in free text physician notes with sensitivity and specificity of 98%, and precision of 93.2%. Anti-hypertensive treatment intensification was identified with sensitivity 83.8%, specificity of 95.0%, and precision of 85.9%. Regular expressions can be an effective method for focused information extraction tasks related to high-priority disease areas such as hypertension.

Blood Pressure↗

A free-text processing system to capture physical findings: Canonical Phrase Identification System (CAPIS).

The task of gathering detailed patient information from free-text medical records presents a significant barrier to clinical research. In this paper, we describe a prototype system for extracting physical examination findings from dictated admission summaries. Our computer program applies a concept-based free-text processing algorithm that identifies user-selected target physical examination findings. We are using the extraction system to enrich an existing clinical database. The system was evaluated by comparing the physical examination findings extracted by our computer program with findings extracted by an independent investigator. Our prototype system was able to recall 92 percent (sensitivity) of the relevant physical findings, with a precision of 96 percent (positive predictive value).

Academic Medical Centers↗

TEXTINFO: a tool for automatic determination of patient clinical profiles using text analysis.

The clinical data contained in narrative patient documents is made available via grammatical and semantic processing. Retrievals from the resulting relational database tables are matched against a set of clinical descriptors to obtain clinical profiles of the patients in terms of the descriptors present in the documents. Discharge summaries of 57 Dept. of Digestive Surgery patients were processed in this manner. Factor analysis and discriminant analysis procedures were then applied, showing the profiles to be useful for diagnosis definitions (by establishing relations between diagnoses and clinical findings), for diagnosis assessment (by viewing the match between a definition and observed events recorded in a patient text), and potentially for outcome evaluation based on the classification abilities of clinical signs.

Databases, Factual↗

Applying hybrid algorithms for text matching to automated biomedical vocabulary mapping.

Several biomedical vocabularies are often used by clinical applications due to their different domain(s) of coverage, intended use, etc. Mapping them to a reference terminology is essential for inter-systems interoperability. Manual vocabulary mapping is labor-intensive and allows room for inconsistencies. It requires manual searching for synonyms, abbreviation expansions, variations, etc., placing additional burden on the mappers. Furthermore, local vocabularies may use non-standard words and abbreviations, posing additional problems. However, much of this process can be automated to provide decision-support, allowing mappers to focus on steps that absolutely need their expertise. We developed hybrid algorithms comprising of rules, permutations, sequence alignment and cost algorithms that utilize the UMLS SPECIALIST Lexicon, a custom knowledgebase and a search engine to automatically find probable matches, allowing mappers to select the best match from this list. We discuss the techniques, results from assisting to map a local codeset, and scope for generalizability.

Algorithms↗

GOChase: correcting errors from Gene Ontology-based annotations for gene products.

SUMMARY: The Gene Ontology (GO) is a controlled biological vocabulary that provides three structured networks of terms to describe biological processes, cellular components and molecular functions. Many databases of gene products are annotated using the GO vocabularies. We found that some GO-updating operations are not easily traceable by the current biological databases and GO browsers. Consequently, numerous annotation errors arise and are propagated throughout biological databases and GO-based high-level analyses. GOChase is a set of web-based utilities to detect and correct the errors in GO-based annotations.

Database Management Systems↗

Interpreting procedures from descriptive guidelines.

Errors in clinical practice guidelines may translate into errors in real-world clinical practice. The best way to eliminate these errors is to understand how they are generated, thus enabling the future development of methods to catch errors made in creating the guideline before publication. We examined the process by which a medical expert from the American College of Physicians (ACP) created clinical algorithms from narrative guidelines, as a case study. We studied this process by looking at intermediate versions produced during the algorithm creation. We identified and analyzed errors that were generated at each stage, categorized them using Knuth's classification scheme, and studied patterns of errors that were made over the set of algorithm versions that were created. We then assessed possible explanations for the sources of these errors and provided recommendations for reducing the number of errors, based on cognitive theory and on experience drawn from software engineering methodologies.

Artificial Intelligence↗

Fast, cheap and out of control: a zero curation model for ontology development.

During two days at a conference focused on circulatory and respiratory health, 68 volunteers untrained in knowledge engineering participated in an experimental knowledge capture exercise. These volunteers created a shared vocabulary of 661 terms, linking these terms to each other and to a pre-existing upper ontology by adding 245 hyponym relationships and 340 synonym relationships. While ontology-building has proved to be an expensive and labor-intensive process using most existing methodologies, the rudimentary ontology constructed in this study was composed in only two days at a cost of only 3 t-shirts, 4 coffee mugs, and one chocolate moose. The protocol used to create and evaluate this ontology involved a targeted, web-based interface. The design and implementation of this protocol is discussed along with quantitative and qualitative assessments of the constructed ontology.

Artificial Intelligence↗

Medical language processing applied to extract clinical information from Dutch medical documents.

In this paper, we want to show how an existing morpho-syntactic analyser for Dutch (Dutch Medical Language Processor--DMLP) has been extended in order to produce output that is compatible with the language independent modules of the LSP-MLP system (Linguistic String Project--Medical Language Processor) of the New York University. The former can focus on idiosyncrasies for Dutch and take advantage of the language independent developments of the latter. This general strategy will be illustrated by a practical application, namely the extraction of clinical information from Dutch patient discharge summaries. Such an application can be of use for education, research and quality control purposes in a hospital environment.

Humans↗

Recognizing names in biomedical texts using mutual information independence model and SVM plus sigmoid.

In this paper, we present a biomedical name recognition system, called PowerBioNE. In order to deal with the special phenomena in the biomedical domain, various evidential features are proposed and integrated through a mutual information independence model (MIIM). In addition, a support vector machine (SVM) plus sigmoid is proposed to resolve the data sparseness problem in the MIIM. In this way, the data sparseness problem in MIIM-based biomedical name recognition can be resolved effectively and a biomedical name recognition system with better performance and better portability can be achieved. Finally, we present two post-processing modules to deal with the nested entity name and abbreviation phenomena in the biomedical domain to further improve the performance. Evaluation shows that our system achieves F-measures of 69.1 and 71.2 on the 23 classes of GENIA V1.1 and V3.0, respectively. In particular, our system achieves an F-measure of 77.8 on the "protein" class of GENIA V3.0. It also shows that our system outperforms the best-reported system on GENIA V1.1 and V3.0.

Abstracting and Indexing↗

Cognitive models of clinical reasoning and conceptual representation.

This paper presents an approach to conceptual representation, informed by theories and methods from cognitive psychology. Our investigation of clinical case comprehension and reasoning from textual information has shifted from instantiation models in which text processing is carried out through schema fitting to more dynamic models that account for how schemata are constructed by a process of construction and integration of meaning, which depends on specific situations. We give an example involving doctor-patient dialogue to illustrate this point. Nonetheless, our main approach has been propositionally-based. As we conduct research into more specific aspects of medical understanding, such as understanding of physiological systems, we have included alternative approaches, such as qualitative functional graphs. We present examples of their use in our research. These representational formalisms allow us better to capture reasoning and understanding in dynamic systems.

Artificial Intelligence↗

Status of text-mining techniques applied to biomedical text.

Scientific progress is increasingly based on knowledge and information. Knowledge is now recognized as the driver of productivity and economic growth, leading to a new focus on the role of information in the decision-making process. Most scientific knowledge is registered in publications and other unstructured representations that make it difficult to use and to integrate the information with other sources (e.g. biological databases). Making a computer understand human language has proven to be a complex achievement, but there are techniques capable of detecting, distinguishing and extracting a limited number of different classes of facts. In the biomedical field, extracting information has specific problems: complex and ever-changing nomenclature (especially genes and proteins) and the limited representation of domain knowledge.

Abstracting and Indexing↗

Positional candidate gene selection from livestock EST databases using Gene Ontology.

MOTIVATION: The number of expressed sequence tags (ESTs) in GenBank has now surpassed 200,000 for cattle and 100,000 for swine. The Institute of Genome Research (TIGR) has organized these sequences into approximately 60,000 non-redundant consensus sequences (identified by TIGR Gene Indices) for cattle and 40,000 for swine. Anonymous ESTs are of limited value unless they are connected to function. Functional information is difficult to manage electronically because of heterogeneity of meaning and form among databases. The Gene Ontology (GO) Consortium has produced ontologies for gene function with consistent meaning and form across species. Linking livestock EST to gene function through similarity with sequences from other annotation-rich mammals could accelerate: (1) the discovery of positional candidate genes underlying a livestock quantitative trait locus (QTL) and (2) comparative mapping between livestock and other mammals (e.g. humans, mouse and rat). We initiated this investigation to determine if incorporation of the GO into the annotation process could accelerate livestock positional candidate gene discovery. RESULTS: We have associated livestock ESTs with GO nodes through sequence similarity to the NCBI Reference Sequences (RefSeq). Positional candidate genes are identified within minutes that otherwise required days. The schema described here accommodates queries that return GO nodes from terms familiar to biologists, such as gene name, alternate/alias symbol, and OMIM phenotype. AVAILABILITY: Scripts and schema are available on request from the authors.

Animals↗

Heuristic evaluation of paper-based Web pages: a simplified inspection usability methodology.

Online medical information, when presented to clinicians, must be well-organized and intuitive to use, so that the clinicians can conduct their daily work efficiently and without error. It is essential to actively seek to produce good user interfaces that are acceptable to the user. This paper describes the methodology used to develop a simplified heuristic evaluation (HE) suitable for the evaluation of screen shots of Web pages, the development of an HE instrument used to conduct the evaluation, and the results of the evaluation of the aforementioned screen shots. In addition, this paper presents examples of the process of categorizing problems identified by the HE and the technological solutions identified to resolve these problems. Four usability experts reviewed 18 paper-based screen shots and made a total of 108 comments. Each expert completed the task in about an hour. We were able to implement solutions to approximately 70% of the violations. Our study found that a heuristic evaluation using paper-based screen shots of a user interface was expeditious, inexpensive, and straightforward to implement.

Computer Graphics↗