PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Paraphrasing for condensation in journal abstracting.

When authors of empirical science articles write abstracts, they employ a wide variety of distinct linguistic operations which interact to condense and rephrase a subset of sentences from the source text. An on-going comparison of biological and biomedical journal articles with their author-written abstracts is providing a basis for a more linguistically detailed model of abstract derivation using syntactic representations of selected source sentences. The description makes use of rich dictionary information to formulate paraphrasing rules of differing degrees of generality, including some which are sublanguage-specific, and others which appear valid in several languages when formulated using "lexical functions" to express important semantic relationships between lexical items. Some paraphrase operations may use both lexical functions and rhetorical relations between sentences to reformulate larger chunks of text in a concise abstract sentence. The descriptive framework is computable and utilizes existing linguistic resources.

Abstracting and Indexing↗

Assessing explicit error reporting in the narrative electronic medical record using keyword searching.

BACKGROUND: Many types of medical errors occur in and outside of hospitals, some of which have very serious consequences and increase cost. Identifying errors is a critical step for managing and preventing them. In this study, we assessed the explicit reporting of medical errors in the electronic record. METHOD: We used five search terms "mistake," "error," "incorrect," "inadvertent," and "iatrogenic" to survey several sets of narrative reports including discharge summaries, sign-out notes, and outpatient notes from 1991 to 2000. We manually reviewed all the positive cases and identified them based on the reporting of physicians. RESULT: We identified 222 explicitly reported medical errors. The positive predictive value varied with different keywords. In general, the positive predictive value for each keyword was low, ranging from 3.4 to 24.4%. Therapeutic-related errors were the most common reported errors and these reported therapeutic-related errors were mainly medication errors. CONCLUSION: Keyword searches combined with manual review indicated some medical errors that were reported in medical records. It had a low sensitivity and a moderate positive predictive value, which varied by search term. Physicians were most likely to record errors in the Hospital Course and History of Present Illness sections of discharge summaries. The reported errors in medical records covered a broad range and were related to several types of care providers as well as non-health care professionals.

Abstracting and Indexing↗

Design and development of chemical ontologies for reaction representation.

This paper describes the development of chemical ontologies applied to the representation of organic chemical reactions. The ontologies are built using the methodology known as methontology. The hierarchically structured set of terms describing the subdomains, namely, organic reactions, organic compounds, and reagents, are constructed into individual ontologies. The ontologies consist of about 200 concepts and around 125 individuals. A set of binary relations is defined in order to integrate the ontologies with applications. The ontologies are implemented as an XML application with a set of vocabulary describing the domain knowledge. This paper also features an easy-to-use chemical ontological support system (COSS) intended to represent organic chemical reactions automatically. As a model application, the automatic representation of aliphatic nucleophilic substitution reactions is demonstrated using COSS. The paper also describes a keyword-based search system whose functionality is backed with COSS.

Algorithms↗

Current status of the evaluation of information retrieval.

This is the second in the series of the articles on an application of the systems analytic approach to evaluation of information retrieval (IR). In the previous article a historical overview of IR was presented and existing terminological problems associated with IR were identified and discussed. In the presented article the current status of IR evaluation is summarized, and different evaluation approaches are discussed. The Cranfield evaluation model and the most often used relevance-based measures of recall and precision are explained, and their problems are presented. Possible evaluation alternatives to the Cranfield model are discussed, and the case for a systems analytic approach to IR is summarized.

Abstracting and Indexing↗

[Speech recognition in clinical routine, a pilot trial at the Zurich University Hospital].

Two systems for continuous German speech recognition were evaluated in pilot installations in a Department of Medicine. Word recognition accuracies of 92-94% were achieved one month after implementation. For standardized text the performance increased to 97%. Speech recognition proved to be economical for physicians when producing short reports, whereas for comprehensive discharge summaries no significant time savings were observed. All reports could be made available much faster by using speech recognition. However, some particle difficulties may still limit the success of comprehensive installations in clinical settings. Adequate hardware, appropriate choice of areas of application, reasonably quiet working conditions and well motivated users are essential for successful implementations.

Hospital Information Systems↗

Creation and implications of a phenome-genome network.

Although gene and protein measurements are increasing in quantity and comprehensiveness, they do not characterize a sample's entire phenotype in an environmental or experimental context. Here we comprehensively consider associations between components of phenotype, genotype and environment to identify genes that may govern phenotype and responses to the environment. Context from the annotations of gene expression data sets in the Gene Expression Omnibus is represented using the Unified Medical Language System, a compendium of biomedical vocabularies with nearly 1-million concepts. After showing how data sets can be clustered by annotative concepts, we find a network of relations between phenotypic, disease, environmental and experimental contexts as well as genes with differential expression associated with these concepts. We identify novel genes related to concepts such as aging. Comprehensively identifying genes related to phenotype and environment is a step toward the Human Phenome Project.

Aging↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

Broca's area and the language instinct.

Language acquisition in humans relies on abilities like abstraction and use of syntactic rules, which are absent in other animals. The neural correlate of acquiring new linguistic competence was investigated with two functional magnetic resonance imaging (fMRI) studies. German native speakers learned a sample of 'real' grammatical rules of different languages (Italian or Japanese), which, although parametrically different, follow the universal principles of grammar (UG). Activity during this task was compared with that during a task that involved learning 'unreal' rules of language. 'Unreal' rules were obtained manipulating the original two languages; they used the same lexicon as Italian or Japanese, but were linguistically illegal, as they violated the principles of UG. Increase of activation over time in Broca's area was specific for 'real' language acquisition only, independent of the kind of language. Thus, in Broca's area, biological constraints and language experience interact to enable linguistic competence for a new language.

Adult↗

Breast cancer: patient information needs reflected in English and German web sites.

Individual belief and knowledge about cancer were shown to influence coping and compliance of patients. Supposing that the Internet information both has impact on patients and reflects patients' information needs, breast cancer web sites in English and German language were evaluated to assess the information quality and were compared with each other to identify intercultural differences. Search engines returned 10 616 hits related to breast cancer. Of these, 4590 relevant hits were analysed. In all, 1888 web pages belonged to 132 English-language web sites and 2702 to 65 German-language web sites. Results showed that palliative therapy (4.5 vs 16.7%; P=0.004), alternative medicine (18.2 vs 46.2%; P<0.001), and disease-related information (prognosis, cancer aftercare, self-help groups, and epidemiology) were significantly more often found on German-language web sites. Therapy-related information (including the side effects of therapy and new studies) was significantly more often given by English-language web sites: for example, details about surgery, chemotherapy, radiotherapy, hormone therapy, immune therapy, and stem cell transplantation. In conclusion, our results have implications for patient education by physicians and may help to improve patient support by tailoring information, considering the weak points in information provision by web sites and intercultural differences in patient needs.

Breast Neoplasms↗

Development of a speech-based dialogue system for report dictation and machine control in the endoscopic laboratory.

BACKGROUND AND STUDY AIMS: Reporting and machine control based on speech technology can enhance work efficiency in the gastrointestinal endoscopy laboratory. MATERIALS AND METHODS: The status and activation of endoscopy laboratory equipment were described as a multivariate parameter and function system. Speech recognition, text evaluation and action definition engines were installed. Special programs were developed for the grammatical analysis of command sentences, and a rule-based expert system for the definition of machine answers. A speech backup engine provides feedback to the user. Techniques were applied based on the "Hidden Markov" model of discrete word, user-independent speech recognition and on phoneme-based speech synthesis. Speech samples were collected from three male low-tone investigators. RESULTS: The dictation module and machine control modules were incorporated in a personal computer (PC) simulation program. Altogether 100 unidentified patient records were analyzed. The sentences were grouped according to keywords, which indicate the main topics of a gastrointestinal endoscopy report. They were: "endoscope", "esophagus", "cardia", "fundus", "corpus", "antrum", "pylorus", "bulbus", and "postbulbar section", in addition to the major pathological findings: "erosion", "ulceration", and "malignancy". "Biopsy" and "diagnosis" were also included. We implemented wireless speech communication control commands for equipment including an endoscopy unit, video, monitor, printer, and PC. The recognition rate was 95%. CONCLUSIONS: Speech technology may soon become an integrated part of our daily routine in the endoscopy laboratory. A central speech and laboratory computer could be the most efficient alternative to having separate speech recognition units in all items of equipment.

Artificial Intelligence↗

[Name-based identification of cases of Turkish origin in the childhood cancer registry in Mainz].

Until now few analyses of routine data relating to the health of migrants have been conducted in Germany. A major obstacle is that most data sources do not provide reliable information on the origin of migrants. While some sources contain the nationality of persons registered, this information does not allow one to identify migrants who have taken up German citizenship, i.e., a substantial part of second-generation migrants. In this paper we demonstrate how a computer-aided, name-based algorithm can be used to identify persons of Turkish origin in the German Childhood Cancer Registry in Mainz, Germany. The performance of the algorithm, as assessed against the gold standard of assessing names manually, was very good (sensitivity and specificity > or = 0.975). In total, we identified 1774 of the 37,259 cases in the registry as being of Turkish origin. The name algorithm proved to be a useful tool to identify Turkish migrants in routine data sources, thus avoiding potential bias due to changes in citizenship. This approach aims at improving migrant-sensitive health reporting and research in Germany. In future, additional information on migrant status should be obtained already during primary data collection so that health data for all migrant groups can be provided.

Algorithms↗

Knowledge discovery in biology and biotechnology texts: a review of techniques, evaluation strategies, and applications.

Arguably, the richest source of knowledge (as opposed to fact and data collections) about biology and biotechnology is captured in natural-language documents such as technical reports, conference proceedings and research articles. The automatic exploitation of this rich knowledge base for decision making, hypothesis management (generation and testing) and knowledge discovery constitutes a formidable challenge. Recently, a set of technologies collectively referred to as knowledge discovery in text (KDT) has been advocated as a promising approach to tackle this challenge. KDT comprises three main tasks: information retrieval, information extraction and text mining. These tasks are the focus of much recent scientific research and many algorithms have been developed and applied to documents and text in biology and biotechnology. This article introduces the basic concepts of KDT, provides an overview of some of these efforts in the field of bioscience and biotechnology, and presents a framework of commonly used techniques for evaluating KDT methods, tools and systems.

Algorithms↗

Learning and performance of able-bodied individuals using scanning systems with and without word prediction.

This study examines how the cognitive and perceptual loads introduced by a word prediction feature impact learning and performance. Two groups of able-bodied subjects transcribed text using two row-column scanning systems for 10 consecutive trials each. The two systems differed only in that one system had a word prediction feature. Subject groups differed in their order of system use. The results show that, under the conditions of this study, the word prediction system was not substantially more difficult to learn, but it did not yield a statistically significant improvement in text generation rate. This suggests that the cost of using this word prediction system balanced the benefit of the keystroke savings achieved by these subjects. The relationship between keystroke savings, cost in item selection rate, and improvement in text generation rate is explored in order to provide insight into this outcome.

Cognition↗