PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Computer assisted information resources navigation.

In this paper, the design and development of Computer Assisted Information Resources Navigation (CAIRN) is discussed. CAIRN system is a medical information retrieval system that allows physicians and students to store full text medical information from any resource, organize and retrieve it. The most important feature of CAIRN is its capability to assist the user, physician, student etc. in selecting documents against a submitted query in Natural Language. The retrieved documents are presented in decreasing order according to their similarity to the submitted query. The nearest neighbour method is used. An alternative similarity measure based on a new calculation of the length of documents is proposed and some experimentation with it is discussed.

Database Management Systems↗

A mathematical specification of the New Deal on junior doctors' hours.

OBJECTIVES: Our objective is to make the New Deal on junior doctors' hours sufficiently precise that the definitions may be used as a basis for computer software that checks the compliance of rotas or that automatically generates compliant rotas. METHODS: We formalize the clauses of the New Deal, as relevant to 'full shifts', using the Z specification language. RESULTS: The mathematical definitions are simple and concise. CONCLUSIONS: Mathematical specification is a useful way to express constraints on rotas unambiguously.

Humans↗

Automatic construction of gene relation networks using text mining and gene expression data.

Microarray gene expression analysis is a powerful high-throughput technique that enables researchers to monitor the expression of thousands of genes simultaneously. Using this methodology huge amounts of data are produced which have to be analysed. Clustering algorithms are used to group genes together based on a predefined distance measure. However, clustering algorithms do not necessarily group the genes in a biological meaningful way. Additional information is needed to improve the identification of disease relevant genes. The primary objective of our project is to support the analysis of microarray gene expression data by construction of gene relation networks (GRNs). Required information can not be found in a structured representation like a database. In contrast, a large number of relations are described in biomedical literature. The main outcome of this project is the implementation of a software system that provides clinicians and researchers with a tool that supports the analysis of microarray gene expression data by mapping known relationships from the biomedical literature to local gene expression experiments.

Abbreviations as Topic↗

Implicit learning in children and adults with Williams syndrome.

In comparison to explicit learning, implicit learning is hypothesized to be a phylogenetically older form of learning that is important in early developmental processes (e.g., natural language acquisition, socialization)and relatively impervious to individual differences in age and IQ. We examined implicit learning in a group of children and adults (9.49 years of age)with Williams syndrome (WS)and in a comparison group of typically developing individuals matched for chronological age. Participants were tested in an artificial-grammar learning paradigm and in a rotor-pursuit task. For both groups, implicit learning was largely independent of age. Both groups showed evidence of implicit learning but the comparison group outperformed the WS group on both tasks. Performance advantages for the comparison group were no longer significant when group differences in working memory or nonverbal intelligence were held constant.

Adolescent↗

Capturing behaviour for the use of avatars in virtual environments.

Avatars, representations of people in virtual environments, are subject to human control. However, for most applications, it is impractical for a person to directly control each joint in a complex avatar. Rather, people must be allowed to specify complex behaviours with simple instructions and the avatar permitted to select the correct movements in sequence to execute the instruction. This requires a variety of technologies that are currently available. Human behaviour must be captured and stored it so that it can be retrieved at a later time for use by the avatar. This has been done successfully with a variety of haptic interfaces, with visual observation of human head movements, and with verbal behaviour in natural language applications. The behaviour must be broken into atomic actions that can be sequenced with a regular grammar, and an appropriate grammar developed. Finally, a user interface must be developed so that a person can deliver instructions to the avatar.

Behavior↗

Semantic search among heterogeneous biological databases based on gene ontology.

Semantic search is a key issue in integration of heterogeneous biological databases. In this paper, we present a methodology for implementing semantic search in BioDW, an integrated biological data warehouse. Two tables are presented: the DB2GO table to correlate Gene Ontology (GO) annotated entries from BioDW data sources with GO, and the semantic similarity table to record similarity scores derived from any pair of GO terms. Based on the two tables, multifarious ways for semantic search are provided and the corresponding entries in heterogeneous biological databases in semantic terms can be expediently searched.

Database Management Systems↗

A survey of current work in biomedical text mining.

The volume of published biomedical research, and therefore the underlying biomedical knowledge base, is expanding at an increasing rate. Among the tools that can aid researchers in coping with this information overload are text mining and knowledge extraction. Significant progress has been made in applying text mining to named entity recognition, text classification, terminology extraction, relationship extraction and hypothesis generation. Several research groups are constructing integrated flexible text-mining systems intended for multiple uses. The major challenge of biomedical text mining over the next 5-10 years is to make these systems useful to biomedical researchers. This will require enhanced access to full text, better understanding of the feature space of biomedical literature, better methods for measuring the usefulness of systems to users, and continued cooperation with the biomedical research community to ensure that their needs are addressed.

Abstracting and Indexing↗

Hairpins in bookstacks: information retrieval from biomedical text.

Current advances in high-throughput biology are accompanied by a tremendous increase in the number of related publications. Much biomedical information is reported in the vast amount of literature. The ability to rapidly and effectively survey the literature is necessary for both the design and the interpretation of large-scale experiments, and for curation of structured biomedical knowledge in public databases. Given the millions of published documents, the field of information retrieval, which is concerned with the automatic identification of relevant documents from large text collections, has much to offer. This paper introduces the basics of information retrieval, discusses its applications in biomedicine, and presents traditional and non-traditional ways in which it can be used.

Abstracting and Indexing↗

Text mining and ontologies in biomedicine: making sense of raw text.

The volume of biomedical literature is increasing at such a rate that it is becoming difficult to locate, retrieve and manage the reported information without text mining, which aims to automatically distill information, extract facts, discover implicit links and generate hypotheses relevant to user needs. Ontologies, as conceptual models, provide the necessary framework for semantic representation of textual information. The principal link between text and an ontology is terminology, which maps terms to domain-specific concepts. This paper summarises different approaches in which ontologies have been used for text-mining applications in biomedicine.

Abstracting and Indexing↗

Information retrieval and knowledge discovery utilising a biomedical Semantic Web.

Although various ontologies and knowledge sources have been developed in recent years to facilitate biomedical research, it is difficult to assimilate information from multiple knowledge sources. To enable researchers to easily gain understanding of a biomedical concept, a biomedical Semantic Web that seamlessly integrates knowledge from biomedical ontologies, publications and patents would be very helpful. In this paper, current research efforts in representing biomedical knowledge in Semantic Web languages are surveyed. Techniques are presented for information retrieval and knowledge discovery from the Semantic Web that extend traditional keyword search and database querying techniques. Finally, some of the challenges that have to be addressed to make the vision of a biomedical Semantic Web a reality are discussed.

Abstracting and Indexing↗

Extraction of biological interaction networks from scientific literature.

Biology can be regarded as a science of networks: interactions between various biological entities (eg genes, proteins, metabolites) on different levels (eg gene regulation, cell signalling) can be represented as graphs and, thus, analysis of such networks might shed new light on the function of biological systems. Such biological networks can be obtained from different sources. The extraction of networks from text is an important technique that requires the integration of several different computational disciplines. This paper summarises the most important steps in network extraction and reviews common approaches and solutions for the extraction of biological networks from scientific literature.

Abstracting and Indexing↗

Online tools to support literature-based discovery in the life sciences.

In biomedical research, the amount of experimental data and published scientific information is overwhelming and ever increasing, which may inhibit rather than stimulate scientific progress. Not only are text-mining and information extraction tools needed to render the biomedical literature accessible but the results of these tools can also assist researchers in the formulation and evaluation of novel hypotheses. This requires an additional set of technological approaches that are defined here as literature-based discovery (LBD) tools. Recently, several LBD tools have been developed for this purpose and a few well-motivated, specific and directly testable hypotheses have been published, some of which have even been validated experimentally. This paper presents an overview of recent LBD research and discusses methodology, results and online tools that are available to the scientific community.

Abstracting and Indexing↗

Get ready to GO! A biologist's guide to the Gene Ontology.

The Gene Ontology (GO) project provides a controlled vocabulary to facilitate high-quality functional gene annotation for all species. Genes in biological databases are linked to GO terms, allowing biologists to ask questions about gene function in a manner independent of species. This tutorial provides an introduction for biologists to the GO resources and covers three of the most common methods of querying GO: by individual gene, by gene function and by using a list of genes. [For the sake of brevity, the term 'gene' is used throughout this paper to refer to genes and their products (proteins and RNAs). GO annotations are always based on the characteristics of gene products, even though it may be the gene that is cited in the annotation.].

Abstracting and Indexing↗

Evaluation of biomedical text-mining systems: lessons learned from information retrieval.

Biomedical text-mining systems have great promise for improving the efficiency and productivity of biomedical researchers. However, such systems are still not in routine use. One impediment to their development is the lack of systematic and rigorous evaluation, comparable to the approaches developed for information retrieval systems. The developers of text-mining systems need to improve both test collections for system-oriented evaluation and undertake user-oriented evaluations to determine the most effective use of their systems for their intended audience.

Algorithms↗

What makes a gene name? Named entity recognition in the biomedical literature.

The recognition of biomedical concepts in natural text (named entity recognition, NER) is a key technology for automatic or semi-automatic analysis of textual resources. Precise NER tools are a prerequisite for many applications working on text, such as information retrieval, information extraction or document classification. Over the past years, the problem has achieved considerable attention in the bioinformatics community and experience has shown that NER in the life sciences is a rather difficult problem. Several systems and algorithms have been devised and implemented. In this paper, the problems and resources in NER research are described, the principal algorithms underlying most systems sketched, and the current state-of-the-art in the field surveyed.

Algorithms↗

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models↗

Bio-ontologies: current trends and future directions.

In recent years, as a knowledge-based discipline, bioinformatics has been made more computationally amenable. After its beginnings as a technology advocated by computer scientists to overcome problems of heterogeneity, ontology has been taken up by biologists themselves as a means to consistently annotate features from genotype to phenotype. In medical informatics, artifacts called ontologies have been used for a longer period of time to produce controlled lexicons for coding schemes. In this article, we review the current position in ontologies and how they have become institutionalized within biomedicine. As the field has matured, the much older philosophical aspects of ontology have come into play. With this and the institutionalization of ontology has come greater formality. We review this trend and what benefits it might bring to ontologies and their use within biomedicine.

Animals↗

Disambiguating proteins, genes, and RNA in text: a machine learning approach.

We present an automated system for assigning protein, gene, or mRNA class labels to biological terms in free text. Three machine learning algorithms and several extended ways for defining contextual features for disambiguation are examined, and a fully unsupervised manner for obtaining training examples is proposed. We train and evaluate our system over a collection of 9 million words of molecular biology journal articles, obtaining accuracy rates up to 85%.

Algorithms↗