PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

MQAF: a medical question-answering framework.

Traditional question-answering programs are difficult to write and require: question analysis, document identification, and text extraction. Thousands of new documents are created daily, making it difficult to determine which has useful data. MQAF facilitates the process by limiting the factors used in the process: a Lexicon (the UMLS), a medical term identifier (MetaMap), a Question Taxonomy, and a Medical Information website that is both evidence-based and kept up-to-date.

Information Storage and Retrieval↗

Development of a cross-thesaurus with Internet-based refinement supported by UMLS.

Combinatorial terminological systems are appearing, to solve the issues related to flexibility and precision of representation requested by modern healthcare information systems and in particular by messaging standards. The development of a robust system of descriptors ('cross-thesaurus') is a crucial activity in the production of combinatorial terminological systems. We developed a tool (I-BROWSE) to produce a cross-thesaurus by analysing existing terminological corpora. To facilitate the work of experts and to produce re-usable results, our application interacts via the Internet with the UMLS Knowledge Sources Server. We applied our tool on 6372 dissections on surgical procedures produced in the project GALEN-IN-USE, as a part of the internal Quality Assurance Program. Support from UMLS seems mostly promising about descriptors on , , and . Additional assistance can be given to domain experts on less frequent descriptors on pervasive modifiers. We plan to apply our tool also to production of terminological standards in CEN, as a part of a world-wide process of gradual convergence and transformation of coding systems into second-generation systems and terminological services.

Databases as Topic↗

From text to knowledge: a unifying document-centered view of analyzed medical language.

Although medical language processing (MLP) has achieved some success, the actual use and dissemination of data extracted from free text by MLP systems is still very limited. We claim that the adoption of an 'enriched-document' paradigm (or 'document-centered' view) can help to address this issue. We present this paradigm and explain how it can be implemented, then discuss its expected benefits both for end-users and MLP researchers.

Artificial Intelligence↗

Improving the human readability of Arden Syntax medical logic modules using a concept-oriented terminology and object-oriented programming expressions.

Medical logic modules are a procedural representation for sharing task-specific knowledge for decision support systems. Based on the premise that clinicians may perceive object-oriented expressions as easier to read than procedural rules in Arden Syntax-based medical logic modules, we developed a method for improving the readability of medical logic modules. Two approaches were applied: exploiting the concept-oriented features of the Medical Entities Dictionary and building an executable Java program to replace Arden Syntax procedural expressions. The usability evaluation showed that 66% of participants successfully mapped all Arden Syntax rules to Java methods. These findings suggest that these approaches can play an essential role in the creation of human readable medical logic modules and can potentially increase the number of clinical experts who are able to participate in the creation of medical logic modules. Although our approaches are broadly applicable, we specifically discuss the relevance to concept-oriented nursing terminologies and automated processing of task-specific nursing knowledge.

Attitude of Health Personnel↗

Joint learning of gene functions--a Bayesian network model approach.

In this paper, we develop a machine learning system for determining gene functions from heterogeneous data sources using a Weighted Naive Bayesian network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or Open Reading Frames (ORFs) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore, many functional links would be missing when only one or two sources of data are used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations, and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

A methodology for partitioning a vocabulary hierarchy into trees.

Controlled medical vocabularies are useful in application areas such as medical information systems and decision-support systems. However, such vocabularies are large and complex, and working with them can be daunting. It is important to provide a means for orienting vocabulary designers and users to the vocabulary's contents. We describe a methodology for partitioning a vocabulary based on an IS-A hierarchy into small meaningful pieces. The methodology uses our disciplined modeling framework to refine the IS-A hierarchy according to prescribed rules in a process carried out by a user in conjunction with the computer. The partitioning of the hierarchy implies a partitioning of the vocabulary. We demonstrate the methodology with respect to a complex sample of the MED, an existing medical vocabulary.

Models, Theoretical↗

Bio-medical entity extraction using support vector machines.

OBJECTIVE: Support vector machines (SVMs) have achieved state-of-the-art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named entity task to special domains with large-scale terminologies such as those in medicine and related disciplines. METHODS AND MATERIALS: The foundation for the model is a sample of text annotated by a domain expert according to an ontology of concepts, properties and relations. The model then learns to annotate unseen terms in new texts and contexts. The results can be used for a variety of intelligent language processing applications. We illustrate SVMs capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. RESULTS: Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance. CONCLUSION: Our experiments indicate a relationship between feature window size and the amount of training data and that a combination of surface words, orthographic features and head noun features achieve the best performance among the feature sets tested.

Algorithms↗

Semantic annotation for concept-based cross-language medical information retrieval.

We present a framework for concept-based cross-language information retrieval in the medical domain, which is under development in the MUCHMORE project. Our approach is based on using the Unified Medical Language System (UMLS) as the primary source of semantic data. Documents and queries are annotated with multiple layers of linguistic information. Linguistic processing includes part-of-speech tagging, morphological analysis, phrase recognition and the identification of medical terms and semantic relations between them. The paper describes experiments in monolingual and cross-language document retrieval, performed on a corpus of medical abstracts. Results show that linguistic processing, especially lemmatization and compound analysis for German, is a crucial step in achieving a good baseline performance. On the other hand, they show that semantic information, specifically the combined use of concepts and relations, increases the performance in monolingual and cross-language retrieval.

Humans↗

Distributed modules for text annotation and IE applied to the biomedical domain.

Biological databases contain facts from scientific literature that have been curated by hand to ensure high quality. Curation is time-consuming and can be supported by information extraction methods. We present a server software infrastructure which allows to easily plug in modules to identify biologically interesting pieces of text to be then presented in a web interface to the curator. There are modules which identify UniProt, UMLS and GO terminology, gene and protein names, mutations and protein-protein interactions. UniProt, UMLS and GO concepts are automatically linked to the original source. The module for mutations is based on syntax patterns and the one for protein-protein interactions relies on chunk parsing. All modules work as separate servers possibly distributed on different machines and can be combined into processing pipelines as necessary. Communication is based on XML annotated text streams, each server processing the XML elements it is designed for, and possibly adding more information in the form of XML annotation. The server and the underlying software are available to the public.

Abstracting and Indexing↗

Automatic identification of pneumonia related concepts on chest x-ray reports.

A medical language processing system called SymText, two other automated methods, and a lay person were compared against an internal medicine resident for their ability to identify pneumonia related concepts on chest x-ray reports. Sensitivity (recall), specificity, and positive predictive value (precision) are reported with respect to an independent panel of physicians. Overall the performance of SymText was similar to the physician and superior to the other methods. The automatic encoding of pneumonia concepts will support clinical research, decision making, computerized clinical protocols, and quality assurance in a radiology department.

Bayes Theorem↗

Integrating order and distance relationships from heterogeneous maps.

There is no automatic mechanism to integrate information between heterogeneous genome maps. Currently, integration is a difficult, manual process. We have developed a process for knowledge base design, and we use this to integrate order and distance relationships between genetic linkage, radiation hybrid, and physical maps. Until now, the only way to develop a persistent, knowledge-intensive application was to either develop a new knowledge base from scratch or coerce the application to fit an existing knowledge base. This was not from lack of interest by the knowledge base or database community, but merely from a lack of theoretical tools powerful enough to tackle the problem. We import formalisms from knowledge representation, natural language semantics, programming language research, and databases. These form a strong, theoretical foundation for knowledge base design upon which we have implemented the knowledge base design tool called WEAVE.

Artificial Intelligence↗

[CILAB--a PC-based laboratory speech processor for implementation and evaluation of new stimulation strategies for cochlear implants].

CILab is a computer-based versatile laboratory system for the implementation and evaluation of innovative stimulation strategies for cochlear implants. In contrast to existing laboratory systems the entire signal processing from the input signal to the creation of the data word for the implant is effected with the aid of a personal computer (PC). This permits rapid implementation of new stimulation strategies or psycho-acoustic tests. Real-time audio processing is also possible by using the CILab as a cochlear implant speech processor. The laboratory system has been employed with success for the evaluation of new strategies and numerous psycho-acoustic tests.

Acoustic Stimulation↗

Terminological systems: bridging the generation gap.

A rigorous formal description of the intended behaviour of a compositional terminology, a 'third generation' system, enables powerful semantic processing techniques to assist in the building of a large terminology. Use of an intermediate representation derived from such a formalism, but simplified to resemble a 'second generation' system, enables authors to work in an simpler and more familiar environment, avoiding many of the technical complications of the 'third generation' system.

Classification↗

Medical linguistics: automated indexing into SNOMED.

This paper reviews the state of the art in processing medical language data. The area is divided into the topics: (1) morphologic analysis, (2) syntactic analysis, (3) semantic analysis, and (4) pragmatics. Additional attention is given to medical nomenclatures and classifications as the bases of (automated) indexing procedures which are required whenever medical information is formalized. These topics are completed by an evaluation of related data structures and methods used to organize language-based medical knowledge.

Abstracting and Indexing↗

Data mining techniques to study the disulfide-bonding state in proteins: signal peptide is a strong descriptor.

In the eucaryotic cell, the formation of disulfide bonds takes place in general inside the endoplasmic reticulum which provides a unique folding environment. The DisulfideDB database gathers information about this biological process with structural, evolutionary and neighborhood information on cysteines in proteins. Mining this information with an association rule discovery program permits to extract some strong rules for the prediction of the disulfide-bonding state of cysteines.

Binding Sites↗

Rubrics to dissections to GRAIL to classifications.

This paper summarises the process in the GALEN-IN-USE project by which rubrics from traditional medical coding schemes are analysed into an intermediate, relatively informal conceptual representation which is then automatically translated into the GRAIL formalism and its Common Reference Model.

Europe↗

Automatic parsing of parental verbal input.

To evaluate theoretical proposals regarding the course of child language acquisition, researchers often need to rely on the processing of large numbers of syntactically parsed utterances, both from children and from their parents. Because it is so difficult to do this by hand, there are currently no parsed corpora of child language input data. To automate this process, we developed a system that combined the MOR tagger, a rule-based parser, and statistical disambiguation techniques. The resultant system obtained nearly 80% correct parses for the sentences spoken to children. To achieve this level, we had to construct a particular processing sequence that minimizes problems caused by the coverage/ambiguity tradeoff in parser design. These procedures are particularly appropriate for use with the CHILDES database, an international corpus of transcripts. The data and programs are now freely available over the Internet.

Adult↗

Utilizing weakly controlled vocabulary for sentence segmentation in biomedical literature.

Since biomedical texts contain a wide variety of domain specific terms, building a large dictionary to perform term matching is of great relevance. However, due to the existence of null boundary between adjacent terms, this matching is not a trivial problem. Moreover, it is known that generative words cannot be comprehensively included in a dictionary because their possible variations are infinite. In this study, we report our approach to dictionary building and term matching in biomedical texts. Large amount of terms with/without part-of-speech (POS) and/or category information were gathered, and a completion program generated approximately 1.36 million term variants to avoid stemming problems when matching terms. The dictionary was stored in a relational database management system (RDBMS) for quick lookup, and used by a matching program. Since the matching operation is not restricted to a substring surrounded by space characters, we can avoid the problem of null boundaries. This feature is also useful for generative words. Experimental results on GENIA corpus are promising: nearly half of the possible terms were correctly recognized as a meaningful segment, and most of the remaining half could be correctly recognized by some post-processing process, like chunking and further decomposition. It should be remarked that although we have not used term cost, connectivity cost, or syntactic information, reasonable segmentation and dictionary lookup were performed in most cases.

Abstracting and Indexing↗