PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Compositional and enumerative designs for medical language representation.

Medical language is in essence highly compositional, allowing complex information to be expressed from more elementary pieces. Embedding the expressive power of medical language into formal systems of representation is recognized in the medical informatics community as a key step towards sharing such information among medical record, decision support, and information retrieval systems. Accordingly, such representation requires managing both the expressiveness of the formalism and its computational tractability, while coping with the level of detail expected by clinical applications. These desiderata can be supported by enumerative as well as compositional approaches, as argued in this paper. These principles have been applied in recasting a frame-based system for general medical findings developed during the 1980s. The new system captures the precise meaning of a subset of over 1500 medical terms for general internal medicine identified from the Quick Medical Reference (QMR) lexicon. In order to evaluate the adequacy of this formal structure in reflecting the deep meaning of the QMR findings, a validation process was implemented. It consists of automatically rebuilding the semantic representation of the QMR findings by analyzing them through the RECIT natural language analyzer, whose semantic components have been adjusted to this frame-based model for the understanding task.

Internal Medicine↗

Analysis of medical texts based on a sound medical model.

Automatic understanding of natural language is a complex task due to the presence of ambiguities. In particular, semantic ambiguities which are often immediately and unconsciously solved by human beings, are raised when analyzing natural language sentences by computer. The latter has to know the implicit and contextual information in order to resolve these difficulties. Nowadays in medicine, a considerable effort is deployed to model semantic contents of the medical domain. Such a task is usually performed separately from linguistic considerations. The goal of this paper is to highlight the key issues of basing a medical language processing system on a sound semantic model. To illustrate the requirements and advantages of such a conceptual approach to the analysis process, the experiment conducted to adjust the RECIT analyzer to the GALEN model is shown.

Models, Theoretical↗

The LBI-method for automated indexing of diagnoses by using SNOMED. Part 2. Evaluation.

We present a simple, formal, lexicon-based method for automated indexing of diagnoses based on the Systematized Nomenclature of Medicine (SNOMED), called LBI-method. Part 1 gave an introduction to the LBI-method and presented its realisation as application system SALBIDH. Part 2 presents the design and the results of an evaluation study to judge the quality of the LBI-method. In this evaluation study the quality of automated indexing as well as the quality of the retrieval of patient data by using automated indexed diagnoses was examined. The results show that the retrieval based on SNOMED indices is at least as good as the retrieval based on ICD classes despite a lot of indexing errors. From this we gather that our system is not yet good enough for immediate routine use but that an appropriate indexing quality and, as a result, a higher retrieval quality can be achieved after few improvements of the LBI-method, especially after revision of the lexicons.

Abstracting and Indexing↗

Automatic SNOMED classification--a corpus-based method.

This paper presents a method of automatic classification of clinical narrative through text comparison. A diagnosis report can be classified by searching archive texts that show a high textual similarity, and the 'nearest neighbor classifies the case. This paper describes the method's theoretical background and gives implementation details. Large scale simulation experiments were run with a wide range of histology reports. Results showed that for 80-84% of the trials, relevant classification lines were included among the first five alternatives. In 5% of the cases, retrieval was unsuccessful due to the absence of relevant archive reports. From the results it is concluded that the method is a versatile approach for finding potentially good classifications.

Algorithms↗

Neural correlates for the acquisition of natural language syntax.

Some types of simple and logically possible syntactic rule never occur in human language grammars, leading to a distinction between grammatical and nongrammatical syntactic rules. Comparison of the neuroanatomical correlates underlying the acquisition of grammatical and nongrammatical rules can provide relevant evidence on the neural processes dedicated to language acquisition in a given developmental stage. Until present no direct evidence on the neural mechanisms subserving language acquisition at any developmental stage has been supplied. We used fMRI in investigating the acquisition of grammatical and nongrammatical rules in the specified sense in 14 healthy adults. Grammatical rules compared with nongrammatical rules specifically activated a left hemispheric network including Broca's area, as shown by direct comparisons between the two rule types. The selective role of Broca's area was further confirmed by time x condition interactions and by proficiency effects, in that higher proficiency in grammatical rule usage, but not in usage of nongrammatical rules, led to higher levels of activation in this area. These findings provide evidence for the neural mechanisms underlying language acquisition in adults.

Adult↗

Impact of voice- and knowledge-enabled clinical reporting--US example.

This study shows qualitative and quantitative estimates of the national and the clinic level impact of utilizing voice and knowledge enabled clinical reporting systems. Using common sense estimation methodology, we show that the delivery of health care can experience a dramatic improvement in four areas as a result of the broad use of voice and knowledge enabled clinical reporting: (1) Process Quality as measured by cost savings, (2) Organizational Quality as measured by compliance, (3) Clinical Quality as measured by clinical outcomes and (4) Service Quality as measured by patient satisfaction. If only 15 percent of US physicians replaced transcription with modem clinical reporting voice-based methodology, about one half billion dollars could be saved. $6.7 Billion could be saved annually if all medical reporting currently transcribed was handled with voice-and knowledge-enabled dictation and reporting systems.

Artificial Intelligence↗

Automated analysis of medical text. I. Clue gathering.

Clinical practice of medicine is highly information-intensive. At the bedside, past experience is the primary justification of reasoning and decisions. This past medical experience is an amalgamation of textbook information and personal experience. During the last 2-3 decades, both of these major sources of clinical information have appeared less and less effective. The pace of progress, resulting in better diagnostic tools and new therapies, has undermined our personal experience, and for the same reason, the time lapse between drafting the manuscripts and distributing the textbooks has become a growing problem. Emphasis has shifted from textbooks to scientific journals with shorter publishing delays, and the role of daily newspapers and television programs seems to be growing. The traditional ways of gathering clinical knowledge and experience seem to fail more and more. In addition to textbooks and scientific journals, current clinical experience is described in millions of patient records, stored in hospitals and ambulatory care offices. However, we have no easy access to patient charts, and we are lacking a method for cost-effective merging of clinical case histories to make them suitable for much-needed statistical inferences. Computers could make a major contribution in this area, but first we must bridge the gap between the narrative text in the medical record and computer technology. Recently, much encouraging progress has been made in automated medical text processing, the topic of this paper.

Abstracting and Indexing↗

Medical language processing with SGML display.

The paper demonstrates several ways that medical language processing can be combined with emerging display technologies to facilitate the extraction of data from free-text patient documents. The techniques allow rapid review via highlighting of the results of processing. Coupling of text markup with further procedures is envisioned.

Asthma↗

Acquisition of lexical resources from SNOMED for medical language processing.

Medical language processing depends on large-coverage, fine-grained specialized lexicons. The vast majority of existing electronic lexicons concern the English language; for other languages such as French, resources are scarce. In contrast, large medical thesauri exist in numerous languages, including French. Our goal was to study what kind of linguistic information could be extracted from thesauri into a lexicon, in which places human intervention is necessary, and what kind of issues arise in this process. We designed in this purpose a method to build a semantic lexicon from a subset of the SNOMED axes in their French translation.

France↗

Getting to the (c)ore of knowledge: mining biomedical literature.

Literature mining is the process of extracting and combining facts from scientific publications. In recent years, many computer programs have been designed to extract various molecular biology findings from Medline abstracts or full-text articles. The present article describes the range of text mining techniques that have been applied to scientific documents. It divides 'automated reading' into four general subtasks: text categorization, named entity tagging, fact extraction, and collection-wide analysis. Literature mining offers powerful methods to support knowledge discovery and the construction of topic maps and ontologies. An overview is given of recent developments in medical language processing. Special attention is given to the domain particularities of molecular biology, and the emerging synergy between literature mining and molecular databases accessible through Internet.

Abstracting and Indexing↗

Text structures in medical text processing: empirical evidence and a text understanding prototype.

We consider the role of textual structures in medical texts. In particular, we examine the impact the lacking recognition of text phenomena has on the validity of medical knowledge bases fed by a natural language understanding front-end. First, we review the results from an empirical study on a sample of medical texts considering, in various forms of local coherence phenomena (anaphora and textual ellipses). We then discuss the representation bias emerging in the text knowledge base that is likely to occur when these phenomena are not dealt with--mainly the emergence of referentially incoherent and invalid representations. We then turn to a medical text understanding system designed to account for local text coherence.

Hospital Information Systems↗

Can computer autoacquisition of medical information meet the needs of the future? A feasibility study in direct computation of the fine grained electronic medical record.

The project describes feasibility testing of a two-year clinical deployment of an electronic record keeping system for primary care medicine that allowed financial medical management and clinical disease study without the encumbrance of human encoding. The software used an expert system for acquisition of historical information and automatic database encoding of each independent fact. The historical acquisition system was combined with a screen-based physician data entry system to create a fine-grained medical record. Fine-grained data allowed direct computer processing to mimic the ends that presently require human encoding--gatekeeping, disease characterization and remote disease surveillance. The project demonstrated the possibility of real time gatekeeping through direct analysis of data. Detection and characterization of disease states using statistical methods within the database was possible, however, limited in this study because of the large numbers of patient interviews required. The possibilities for remote disease monitoring and clinical studies are also discussed.

Chi-Square Distribution↗

The impact of tokenizer selection in genomic language models.

MOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130.

Natural Language Processing↗

Unlimited capacity and processibility of sequence information: prerequisites for a system of biological chains.

Both genetic chains in a cell and phonological chains in a human phonological working memory simultaneously store and process sequence information. Unlimited capacity and processibility of sequence information are two prerequisites for such a system of biological chains. It is demonstrated that information chains (I-chains) and conformation chains (C-chains) satisfy these two prerequisites. Namely, in both kinds of chains constant efficiency and precision of intra- and inter-sequence interactions are guaranteed irrespective of the chain length. Nucleic acids and proteins are I-chains and C-chains of genetic chains, respectively. A 'molecular' model of a phonological chain is formulated based on the properties of phonological working memory. It is proposed that prose and verse are I-chains and C-chains of phonological chains, respectively. The correspondence between a system of genetic chains and a system of phonological chains is explored in detail. A critical difference between systems of biological chains and artificial information-processing systems is attributed to the existence of C-chains.

Humans↗

Medical language processing: applications to patient data representation and automatic encoding.

A linguistic approach is presented to develop a representation of patient data. Semantic categories developed for computer processing of narrative clinical reports are shown to be similar to the Medical Concepts used manually to extract data from narrative in Exercises of the Computer-based Patient Record Institute. Clinical statement types composed of these categories are used in the Linguistic String Project (LSP) medical language processing (MLP) system to convert narrative information into relational database tables of patient information. A procedure for mapping the output of the LSP MLP system into SNOMED International codes was developed. Preliminary results and further requirements are discussed.

Abstracting and Indexing↗

Description generation of abnormal densities found in radiographs.

In this paper we present a system for describing renal stones found in radiographs. The system generates descriptions that adhere to those generated by radiologists. The descriptions are formulated by discovering the spatial relationships that exist between the major organs and the renal stones. The system consists of three major components. The first is the image processing component which is responsible for locating the stone. The second component is the inference network minimization component which determines which spatial relationships, of all those that exist between the stone and the organs, is the most descriptive. The third component is the natural language generation component which is responsible for translating the spatial relationships into appropriate medical terminology. We will illustrate all these components on several examples.

Algorithms↗