PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A reduced ambiguity lexical system.

Natural human languages have proven to be sub-optimal in artificial intelligence applications because of their tendency to inexact representation of meaning. The author has devised a technique for converting human language to and from a compact byte-coded intermediate representation, which is processed more easily by computer systems. A specialized lexical engine based on IEEE Standard 1275-1994 was created to embed redundant information invisibly within the byte-coded text stream, to enable use of a variety of alphabets, grammars, and pronunciation rules (including slang and regional dialects). Very large vocabularies in a variety of human languages are supported. These lexical tools are designed to facilitate speech recognition and speech synthesis subsystems, universal translators and machine intelligence systems.

Algorithms↗

Text structures in medical text processing: empirical evidence and a text understanding prototype.

We consider the role of textual structures in medical texts. In particular, we examine the impact the lacking recognition of text phenomena has on the validity of medical knowledge bases fed by a natural language understanding front-end. First, we review the results from an empirical study on a sample of medical texts considering, in various forms of local coherence phenomena (anaphora and textual ellipses). We then discuss the representation bias emerging in the text knowledge base that is likely to occur when these phenomena are not dealt with--mainly the emergence of referentially incoherent and invalid representations. We then turn to a medical text understanding system designed to account for local text coherence.

Hospital Information Systems↗

A Bayesian network coding scheme for annotating biomedical information presented to genetic counseling clients.

We developed a Bayesian network coding scheme for annotating biomedical content in layperson-oriented clinical genetics documents. The coding scheme supports the representation of probabilistic and causal relationships among concepts in this domain, at a high enough level of abstraction to capture commonalities among genetic processes and their relationship to health. We are using the coding scheme to annotate a corpus of genetic counseling patient letters as part of the requirements analysis and knowledge acquisition phase of a natural language generation project. This paper describes the coding scheme and presents an evaluation of intercoder reliability for its tag set. In addition to giving examples of use of the coding scheme for analysis of discourse and linguistic features in this genre, we suggest other uses for it in analysis of layperson-oriented text and dialogue in medical communication.

Artificial Intelligence↗

Can computer autoacquisition of medical information meet the needs of the future? A feasibility study in direct computation of the fine grained electronic medical record.

The project describes feasibility testing of a two-year clinical deployment of an electronic record keeping system for primary care medicine that allowed financial medical management and clinical disease study without the encumbrance of human encoding. The software used an expert system for acquisition of historical information and automatic database encoding of each independent fact. The historical acquisition system was combined with a screen-based physician data entry system to create a fine-grained medical record. Fine-grained data allowed direct computer processing to mimic the ends that presently require human encoding--gatekeeping, disease characterization and remote disease surveillance. The project demonstrated the possibility of real time gatekeeping through direct analysis of data. Detection and characterization of disease states using statistical methods within the database was possible, however, limited in this study because of the large numbers of patient interviews required. The possibilities for remote disease monitoring and clinical studies are also discussed.

Chi-Square Distribution↗

The impact of tokenizer selection in genomic language models.

MOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130.

Natural Language Processing↗

Unlimited capacity and processibility of sequence information: prerequisites for a system of biological chains.

Both genetic chains in a cell and phonological chains in a human phonological working memory simultaneously store and process sequence information. Unlimited capacity and processibility of sequence information are two prerequisites for such a system of biological chains. It is demonstrated that information chains (I-chains) and conformation chains (C-chains) satisfy these two prerequisites. Namely, in both kinds of chains constant efficiency and precision of intra- and inter-sequence interactions are guaranteed irrespective of the chain length. Nucleic acids and proteins are I-chains and C-chains of genetic chains, respectively. A 'molecular' model of a phonological chain is formulated based on the properties of phonological working memory. It is proposed that prose and verse are I-chains and C-chains of phonological chains, respectively. The correspondence between a system of genetic chains and a system of phonological chains is explored in detail. A critical difference between systems of biological chains and artificial information-processing systems is attributed to the existence of C-chains.

Humans↗

Repair-FunMap: a functional database of proteins of the DNA repair systems.

UNLABELLED: Repair-FunMap is a functional database of the DNA repair systems. This database contains not only the proteins directly involved in DNA repair, but also the proteins that interact with the DNA repair proteins. A protein interaction network associated with the human DNA repair processes was established according to the functional relationship between proteins in the database. This network represents the current knowledge on the intrinsic signaling pathways related to DNA repair. The Repair-FunMap could become an essential resource center for cancer research, providing clues to understanding the inter-relationship between proteins in the network, and to building scientific models of the DNA repair processes. AVAILABILITY: http://astro.temple.edu/~feng/Servers/BioinformaticServers.htm

Abstracting and Indexing↗

Which gene did you mean?

Computational Biology needs computer-readable information records. Increasingly, meta-analysed and pre-digested information is being used in the follow up of high throughput experiments and other investigations that yield massive data sets. Semantic enrichment of plain text is crucial for computer aided analysis. In general people will think about semantic tagging as just another form of text mining, and that term has quite a negative connotation in the minds of some biologists who have been disappointed by classical approaches of text mining. Efforts so far have tried to develop tools and technologies that retrospectively extract the correct information from text, which is usually full of ambiguities. Although remarkable results have been obtained in experimental circumstances, the wide spread use of information mining tools is lagging behind earlier expectations. This commentary proposes to make semantic tagging an integral process to electronic publishing.

Abstracting and Indexing↗

A graph-theoretic modeling on GO space for biological interpretation of gene clusters.

MOTIVATION: With the advent of DNA microarray technologies, the parallel quantification of genome-wide transcriptions has been a great opportunity to systematically understand the complicated biological phenomena. Amidst the enthusiastic investigations into the intricate gene expression data, clustering methods have been the useful tools to uncover the meaningful patterns hidden in those data. The mathematical techniques, however, entirely based on the numerical expression data, do not show biologically relevant information on the clustering results. RESULTS: We present a novel methodology for biological interpretation of gene clusters. Our graph theoretic algorithm extracts common biological attributes of the genes within a cluster or a group of interest through the modified structure of gene ontology (GO) called GO tree. After genes are annotated with GO terms, the hierarchical nature of GO terms is used to find the representative biological meanings of the gene clusters. In addition, the biological significance of gene clusters can be assessed quantitatively by defining a distance function on the GO tree. Our approach has a complementary meaning to many statistical clustering techniques; we can see clustering problems from a different viewpoint by use of biological ontology. We applied this algorithm to the well-known data set and successfully obtained the biological features of the gene clusters with the quantitative biological assessment of clustering quality through GO Biological Process.

Algorithms↗

Medical language processing: applications to patient data representation and automatic encoding.

A linguistic approach is presented to develop a representation of patient data. Semantic categories developed for computer processing of narrative clinical reports are shown to be similar to the Medical Concepts used manually to extract data from narrative in Exercises of the Computer-based Patient Record Institute. Clinical statement types composed of these categories are used in the Linguistic String Project (LSP) medical language processing (MLP) system to convert narrative information into relational database tables of patient information. A procedure for mapping the output of the LSP MLP system into SNOMED International codes was developed. Preliminary results and further requirements are discussed.

Abstracting and Indexing↗

Description generation of abnormal densities found in radiographs.

In this paper we present a system for describing renal stones found in radiographs. The system generates descriptions that adhere to those generated by radiologists. The descriptions are formulated by discovering the spatial relationships that exist between the major organs and the renal stones. The system consists of three major components. The first is the image processing component which is responsible for locating the stone. The second component is the inference network minimization component which determines which spatial relationships, of all those that exist between the stone and the organs, is the most descriptive. The third component is the natural language generation component which is responsible for translating the spatial relationships into appropriate medical terminology. We will illustrate all these components on several examples.

Algorithms↗

Modeling a description logic vocabulary for cancer research.

The National Cancer Institute has developed the NCI Thesaurus, a biomedical vocabulary for cancer research, covering terminology across a wide range of cancer research domains. A major design goal of the NCI Thesaurus is to facilitate translational research. We describe: the features of Ontylog, a description logic used to build NCI Thesaurus; our methodology for enhancing the terminology through collaboration between ontologists and domain experts, and for addressing certain real world challenges arising in modeling the Thesaurus; and finally, we describe the conversion of NCI Thesaurus from Ontylog into Web Ontology Language Lite. Ontylog has proven well suited for constructing big biomedical vocabularies. We have capitalized on the Ontylog constructs Kind and Role in the collaboration process described in this paper to facilitate communication between ontologists and domain experts. The artifacts and processes developed by NCI for collaboration may be useful in other biomedical terminology development efforts.

Animals↗

A virtual university Web system for a medical school.

This paper describes a Virtual Medical University Web Server. This project started in 1994 by the development of the French Radiology Server. The main objective of our Medical Virtual University is to offer not only an initial training (for students) but also the Continuing Professional Education (for practitioners). Our system is based on electronic textbooks, clinical cases (around 4000) and a medical knowledge base called A.D.M. ("Aide au Diagnostic Medical"). We have indexed all electronic textbooks and clinical cases according to the ADM base in order to facilitate the navigation on the system. This system base is supported by a relational database management system. The Virtual Medical University, available on the Web Internet, is presently in the process of external evaluations.

Computer-Assisted Instruction↗

Genome wide prediction of protein function via a generic knowledge discovery approach based on evidence integration.

BACKGROUND: The automation of many common molecular biology techniques has resulted in the accumulation of vast quantities of experimental data. One of the major challenges now facing researchers is how to process this data to yield useful information about a biological system (e.g. knowledge of genes and their products, and the biological roles of proteins, their molecular functions, localizations and interaction networks). We present a technique called Global Mapping of Unknown Proteins (GMUP) which uses the Gene Ontology Index to relate diverse sources of experimental data by creation of an abstraction layer of evidence data. This abstraction layer is used as input to a neural network which, once trained, can be used to predict function from the evidence data of unannotated proteins. The method allows us to include almost any experimental data set related to protein function, which incorporates the Gene Ontology, to our evidence data in order to seek relationships between the different sets. RESULTS: We have demonstrated the capabilities of this method in two ways. We first collected various experimental datasets associated with yeast (Saccharomyces cerevisiae) and applied the technique to a set of previously annotated open reading frames (ORFs). These ORFs were divided into training and test sets and were used to examine the accuracy of the predictions made by our method. Then we applied GMUP to previously un-annotated ORFs and made 1980, 836 and 1969 predictions corresponding to the GO Biological Process, Molecular Function and Cellular Component sub-categories respectively. We found that GMUP was particularly successful at predicting ORFs with functions associated with the ribonucleoprotein complex, protein metabolism and transportation. CONCLUSION: This study presents a global and generic gene knowledge discovery approach based on evidence integration of various genome-scale data. It can be used to provide insight as to how certain biological processes are implemented by interaction and coordination of proteins, which may serve as a guide for future analysis. New data can be readily incorporated as it becomes available to provide more reliable predictions or further insights into processes and interactions.

Algorithms↗

A literature-based similarity metric for biological processes.

BACKGROUND: Recent analyses in systems biology pursue the discovery of functional modules within the cell. Recognition of such modules requires the integrative analysis of genome-wide experimental data together with available functional schemes. In this line, methods to bridge the gap between the abstract definitions of cellular processes in current schemes and the interlinked nature of biological networks are required. RESULTS: This work explores the use of the scientific literature to establish potential relationships among cellular processes. To this end we have used a document based similarity method to compute pair-wise similarities of the biological processes described in the Gene Ontology (GO). The method has been applied to the biological processes annotated for the Saccharomyces cerevisiae genome. We compared our results with similarities obtained with two ontology-based metrics, as well as with gene product annotation relationships. We show that the literature-based metric conserves most direct ontological relationships, while reveals biologically sounded similarities that are not obtained using ontology-based metrics and/or genome annotation. CONCLUSION: The scientific literature is a valuable source of information from which to compute similarities among biological processes. The associations discovered by literature analysis are a valuable complement to those encoded in existing functional schemes, and those that arise by genome annotation. These similarities can be used to conveniently map the interlinked structure of cellular processes in a particular organism.

Databases, Bibliographic↗

ADM-INDEX: an automated system for indexing and retrieval of medical texts.

ADM-INDEX is a system for indexing and retrieval of Patients Discharge Summaries (PDSs) by using linguistic methods (morphologic, syntaxic and semantic processing). The ADM-INDEX knowledge base is a restructuring of a diagnostic aid knowledge base (ADM) in order to allow the linguistic analysis of medical texts. The ADM system is a comprehensive medical knowledge base which has been developed since 1972 at the University Hospital of Rennes and which has been the first professional videotex medical diagnostic aid in France. After linguistic analysis, ADM-INDEX build the index table with thesaurus wording, medical words, concepts and phrases, unknown words contained in each PDS. The benefit of using those different elements is to improve information retrieval. Although our system is constructed with the ADM dictionary, it can be easily applied to other medical nomenclature or thesaurus. In this paper, we present on the one hand the ADM-INDEX knowledge base which is constituted by rules, a dictionary and a thesaurus, and on the other hand, the process of indexing and retrieval information.

Abstracting and Indexing↗

Morse code recognition system with fuzzy algorithm for disabled persons.

It is generally known that Morse code is an efficient input method for one or two switches and it is made from long and short sounds separated by silence between the sounds. The long-to-short ratio in the definition is always 3 to 1, but the long-to-short ratio variation for a disabled person is so large that it is difficult to recognize. In the last few years, several Morse code recognition methods have been successfully built on the LMS adaptive algorithms and neural network algorithm. But LMS-related adaptive algorithms need mass computation to infer the characteristic of the controller; also the neural network must learn first, by inputting some data before it is used to recognize the Morse code sequence. In this study, two fuzzy algorithms are used to recognize the unstable Morse code sequences and the result demonstrates a significant improvement of recognition for real time signal processing in a single-chip microprocessor.

Adolescent↗

The evolution of tools and processes for data mapping.

The 3M Healthcare Data Dictionary (HDD) group provides vocabulary mapping services for the Department of Defense (DoD). The process to manage changes in these mappings, the so-called "delta process," is complex and requires the exchange and management of large amounts of data. To aid in this process, the 3M HDD group has created and refined a vocabulary Mapping Environment (ME).

Databases as Topic↗