PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

The effect of a general lexicon in corpus-based identification of French-English medical word translations.

We present a method, based on the similarity of word distribution across languages, of finding 'new' words' translations in French-English comparable medical texts, starting from a partial bilingual medical lexicon. In this paper, we test the influence of adding general-language words to this initial lexicon. Our experimental results show that all test words are correctly translated within the top 25 candidates; and that the addition of general words to the lexicon helps to improve translation accuracy for medical words.

Algorithms↗

Matching controlled vocabulary words.

This study examines an enabling condition for natural languages access to medical knowledge resources (Medline, CISMeF) indexed with controlled vocabularies (e.g., the MeSH): is the vocabulary of user queries comparable with that of the index terms? The two vocabularies were compared in their original form, then under incrementally normalized forms, using character-based normalizations then linguistic normalizations. Only 16.7% of the user vocabulary, in its original form, is in the MeSH. Progressive normalizations increase this proportion to 65.5%. Besides, if the frequencies of occurrence of words are taken into account, 89.3% of user word occurrences can be matched to MeSH words. This shows the interest of taking into account further matching methods between queries and index terms than those presented here.

France↗

A prototype natural language interface to a large complex knowledge base, the Foundational Model of Anatomy.

We describe a constrained natural language interface to a large knowledge base, the Foundational Model of Anatomy (FMA). The interface, called GAPP, handles simple or nested questions that can be parsed to the form, subject-relation-object, where subject or object is unknown. With the aid of domain-specific dictionaries the parsed sentence is converted to queries in the StruQL graph-searching query language, then sent to a server we developed, called OQAFMA, that queries the FMA and returns output as XML. Preliminary evaluation shows that GAPP has the potential to be used in the evaluation of the FMA by domain experts in anatomy.

Anatomy↗

Medical problem and document model for natural language understanding.

We are developing tools to help maintain a complete, accurate and timely problem list within a general purpose Electronic Medical Record system. As a part of this project, we have designed a system to automatically retrieve medical problems from free-text documents. Here we describe an information model based on XML (eXtensible Markup Language) and compliant with the CDA (Clinical Document Architecture). This model is used to ease the exchange of clinical data between the Natural Language Understanding application that retrieves potential problems from narrative document, and the problem list management application.

Computer Simulation↗

A data-driven approach for extracting "the most specific term" for ontology development.

We present a data-driven approach to extract the "most specific" terms relevant to an ontology of functioning, disability and health. The algorithm is a combination of statistical and linguistic approaches. The statistical filter is based on the frequency of the content words in a given text string; the linguistic heuristic is an extension of existing algorithms but goes beyond noun phrases and is formulated as a "complete syntactic node". Thus, it can be applied to any syntactic node of interest in the particular domain. Two test sets were marked by three experts. Test set 1 is a well-constructed text from pain abstracts; test set 2 is actual medical reports. Results are reported as recall, precision, F-score and rate of valid terms in false positives. A limitation of the current research is the relatively small test set.

Algorithms↗

Creating knowledgebases to text-mine PUBMED articles using clustering techniques.

Knowledgebase-mediated text-mining approaches work best when processing the natural language of domain-specific text. To enhance the utility of our successfully tested program-NeuroText, and to extend its methodologies to other domains, we have designed clustering algorithms, which is the principal step in automatically creating a knowledgebase. Our algorithms are designed to improve the quality of clustering by parsing the test corpus to include semantic and syntactic parsing

Algorithms↗

Extracting diagnosis from Japanese radiological report.

This study is aimed at extracting diagnosis with positive or negative assertion from radiological report written in Japanese Natural Language. We get frequency of verb patterns that indicate pos/neg assertion, and extract a rule in order of the occurrence. We made customized dictionary of 36,152 terms relating to disease names or radiological findings, and tried to extract pairs of (pos/neg, disease and verb pattern ) by using rules according to the most frequent pattern from 1,524/5,000 CT reports (each report consists of 15.1 words on the average). We tried only a few rules so far, and continue to find other rules.

Diagnosis↗

Modeling interventions to improve access to public health information.

In a Robert Wood Johnson funded project, we established a model-based means for automatically analyzing and representing grey literature that reports on public health (PH) interventions. We summarize the development of an intervention model for public health documents and provide a project update on the implementation of natural language technology to improve access to difficult to find public health information.

Algorithms↗

Automatic learning of the morphology of medical language using information compression.

Conversion of free-text strings in a natural language to a standard representation (codes) is an important reoccurring problem in biomedical informatics. Determining the content of a string involves identifying its meaningful constituents (morphemes). One current method of identifying these constituents is to look them up in a preexisting table (lexicon). Manual construction of lexicons and grammars in complex domains such as biomedicine is extremely laborious. As an alternative to the lexico-grammatical approach, we introduce a segmentation algorithm that automatically learns lexical and structural preferences from corpora via information compression. The method is based on the Minimum Description Length (MDL) principle from classic information theory.

Algorithms↗

Implementation and evaluation of an virtual intelligent agent.

Reference librarians at the National Library of Medicine answer over 100,000 client questions annually. Although many answers are the NLM Web, clients may have difficulty finding them. Using intelligent virtual agent software, the NLM launched Cosmo, the Customer Service Owl (http://wwwns.nlm.nih.gov/). Cosmo uses natural language pattern matching to answer common questions. Early evaluation shows Cosmo can answer 25% of questions asked, but some users mistake the service for chat reference or a search engine.

Artificial Intelligence↗

Generating medical logic modules for clinical trial eligibility criteria.

Clinical trials are an important part of modern medical research, however the effort required to find candidates for participation in such trial is significant. With the increasing prevalence of electronic medical records, automated or semi-automated solutions become feasible. We present an semi-automated approach for determining clinical trial eligibility based on information available in an electronic medical record.

Clinical Trials as Topic↗

A protocol for the update of references to scientific literature in biological databases.

Entries in biological databases are usually linked to scientific references. To generate those links and to keep them up-to-date, database maintainers have to continuously scan the scientific literature to select references that are relevant for each single database entry. The continuous growth of both the corpus of scientific literature and the size of biological databases makes this task very hard. We present a protocol intended to assist the updating of an existing set of literature (abstract) links from a single database entry with new references. It consists of taking the set of MEDLINE neighbour references of the existing linked abstracts and evaluating their relevance according to the existing set of abstracts. To test the applicability of the algorithm, we did a simple benchmark of the system using the references associated with the entries of a protein domain database. Human experts found the references that the algorithm scored highly were more relevant to the database entry than those scored lowly, suggesting that the algorithm was useful.

Abstracting and Indexing↗

DNA sequence analysis linguistic tools: contrast vocabularies, compositional spectra and linguistic complexity.

This is a review of the methods based on counting oligomers in nucleotide and amino acid sequences. Such methods are analogous to the formal linguistic analysis of human texts. This review includes methods based on the calculation of observed occurrences (frequencies) of oligomers and their distribution, as well as those based on deviations between the observed and the expected occurrences (contrast words, genome signatures) in biological sequences. Both types of methods have a wide range of sensitivity and can identify homologous as well as functionally and taxonomically related sequences.

Algorithms↗

University of Wyoming, College of Engineering, Undergraduate Senior Design Project: the talking hand.

A glove was built using flex sensors to produce a voltage according to the amount of bend each finger produces when signing a letter of the alphabet. The glove detects and outputs in text the letter of the alphabet being signed as the wearer signs the different letters. The amount of bend causes a change in resistance, which in turn produces a specific voltage in accordance to the letter being signed. That voltage is then fed into a data acquisition card that runs into a personal computer. Through intensive programming and training of a special algorithm called a neural network; the input voltage to the data acquisition card will result in that letter being displayed in font on the monitor of the computer. The computer is then programmed to take the text that is displayed on the monitor and run it out of the PC into a store bought chipset that will convert the text to speech. Therefore, a person will be able to put on this glove, sign all twenty-six letters of the alphabet, see the letter they are currently signing output on the monitor, and hear it spoken in a pre-recorded voice.

Algorithms↗

Aligning words in French-English non-parallel medical texts: effect of term frequency distributions.

In this paper, we present a method for aligning words based on a statistical model of word distribution similarity. The basis underlying our method is that there is a correlation between the patterns of word co-occurrences in texts of different languages. Using automatically downloaded pages from different medical web sites and a combined bilingual lexicon of general and medical terms as language sources, a similarity score is assigned to each proposed translated pair of words, based on the distributional contexts of these two words. We vary several parameters of the method. Experimental results confirm a positive effect of frequency, show that medical words are better handled than less specialized words, and do not evidence a clear influence of context window size. Future directions for improvement include working with very large, part-of-speech tagged corpora.

Algorithms↗

The NLM Indexing Initiative's Medical Text Indexer.

The Medical Text Indexer (MTI) is a program for producing MeSH indexing recommendations. It is the major product of NLM's Indexing Initiative and has been used in both semi-automated and fully automated indexing environments at the Library since mid 2002. We report here on an experiment conducted with MEDLINE indexers to evaluate MTI's performance and to generate ideas for its improvement as a tool for user-assisted indexing. We also discuss some filtering techniques developed to improve MTI's accuracy for use primarily in automatically producing the indexing for several abstracts collections.

Abstracting and Indexing↗