PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design↗

Interactive software for generation and visualization of structured findings in radiology reports.

OBJECTIVE: Our objectives were to develop a user-friendly graphic interface for a module that integrates traditional radiology reporting, natural language processing, and editing capabilities; to facilitate the structuring of radiology reports as part of routine clinical practice; to use a commercial speech recognition module for online transcription; and to implement the module in a hardware-independent environment. CONCLUSION: After implementation, the module was tested with 150 chest radiology reports by two radiologists and assessed for ease of use and accuracy. Overall, accuracy was close to 90% and user satisfaction was high. When radiology reports are structured as a part of routine clinical practice, it is possible to accomplish intelligent indexing and retrieval to facilitate teaching and research.

Medical Records Systems, Computerized↗

MANULEX: a grade-level lexical database from French elementary school readers.

This article presents MANULEX, a Web-accessible database that provides grade-level word frequency lists of nonlemmatized and lemmatized words (48,886 and 23,812 entries, respectively) computed from the 1.9 million words taken from 54 French elementary school readers. Word frequencies are provided for four levels: first grade (G1), second grade (G2), third to fifth grades (G3-5), and all grades (G1-5). The frequencies were computed following the methods described by Carroll, Davies, and Richman (1971) and Zeno, Ivenz, Millard, and Duvvuri (1995), with four statistics at each level (F, overall word frequency; D, index of dispersion across the selected readers; U, estimated frequency per million words; and SFI, standard frequency index). The database also provides the number of letters in the word and syntactic category information. MANULEX is intended to be a useful tool for studying language development through the selection of stimuli based on precise frequency norms. Researchers in artificial intelligence can also use it as a source of information on natural language processing to simulate written language acquisition in children. Finally, it may serve an educational purpose by providing basic vocabulary lists.

Adolescent↗

Developing NLP Tools for Genome Informatics: An Information Extraction Perspective.

Huge quantities of on-line medical texts such as Medline are available, and we would hope to extract useful information from these resources, as much as possible, hopefully in an automatic way, with the aid of computer technologies. Especially, recent advances in Natural Language Processing (NLP) techniques raise new challenges and opportunities for tackling genome-related on-line text; combining NLP techniques with genome informatics extends beyond the traditional realms of either technology to a variety of emerging applications. In this paper, we explain some of our current efforts for developing various NLP-based tools for tackling genome-related on-line documents for information extraction task.

Journal Article↗

Paragraph-oriented structure for narratives in medical documentation.

The authors present a 6 years experiment using a document- centered electronic patient record, based on a central document repository. The document management system is paragraph oriented and all documents are built automatically before editing using predefined ordered sets of para-graphs. Paragraphs can be preloaded with templates, text or images. Once edited, signed and printed, documents are again decomposed in paragraphs and permanently stored. This system, though the compositional aspect of paragraphs is limited and their semantic content wide, offers numerous advantages. The typology is easy to build and to maintain, it has been implemented widely in our hospitals without need for any natural language processing techniques and is used daily within commercially available text editors. The actual state of the system is discussed, emphasizing the structure of the documents, the various attributes and properties that have been needed in order to meet user's needs.

Documentation↗

A humanist's legacy in medical informatics: visions and accomplishments of Professor Jean-Raoul Scherrer.

OBJECTIVE: To report about the work of Prof. Jean-Raoul Scherrer, and show how his humanist vision, his medical skills and his scientific background have enabled and shaped the development of medical informatics over the last 30 years. RESULTS: Starting with the mainframe-based patient-centered hospital information system DIOGENE in the 70s, Prof. Scherrer developed, implemented and evolved innovative concepts of man-machine interfaces, distributed and federated environments, leading the way with information systems that obstinately focused on the support of care providers and patients. Through a rigorous design of terminologies and ontologies, the DIOGENE data would then serve as a basis for the development of clinical research, data mining, and lead to innovative natural language processing techniques. In parallel, Prof. Scherrer supported the development of medical image management, ranging from a distributed picture archiving and communication systems (PACS) to molecular imaging of protein electrophoreses. Recognizing the need for improving the quality and trustworthiness of medical information on the Web, Prof. Scherrer created the Health-On-the-Net (HON) foundation. CONCLUSIONS: These achievements, made possible thanks to his visionary mind, deep humanism, creativity, generosity and determination, have made of Prof. Scherrer a true pioneer and leader of the human-centered, patient-oriented application of information technology for improving healthcare.

History, 20th Century↗

Controlling the vocabulary for anatomy.

When confronted with the representation of human anatomy, natural language processing (NLP) system designers are facing an unsolved and frequent problem: the lack of a suitable global reference. The available sources in electronic format are numerous, but none fits adequately all the constraints and needs of language analysis. These sources are usually incomplete, difficult to use or tailored to specific needs. The anatomist's or ontologist's view does not necessarily match that of the linguist. The purpose of this paper is to review most recognized sources of knowledge in anatomy usable for linguistic analysis. Their potential and limits are emphasized according to this point of view. Focus is given on the role of the consensus work of the International Federation of Associations of Anatomists (IFAA) giving the Terminologia Anatomica.

Anatomy↗

A study of abbreviations in MEDLINE abstracts.

Abbreviations are widely used in writing, and the understanding of abbreviations is important for natural language processing applications. Abbreviations are not always defined in a document and they are highly ambiguous. A knowledge base that consists of abbreviations with their associated senses and a method to resolve the ambiguities are needed. In this paper, we studied the UMLS coverage, textual variants of senses, and the ambiguity of abbreviations in MEDLINE abstracts. We restricted our study to three-letter abbreviations which were defined using parenthetical expressions. When grouping similar expansions together and representing senses using groups, we found that after ignoring senses where the total number of occurrences within the corresponding group was less than 100, 82.8% of the senses matched the UMLS, covered over 93% of occurrences that were considered, and had an average of 7.74 expansions for each sense. Abbreviations are highly ambiguous: 81.2% of the abbreviations were ambiguous, and had an average of 16.6 senses. However, after ignoring senses with occurrences of less than 5, 64.6% of the abbreviations were ambiguous, and had an average of 4.91 senses.

Abbreviations as Topic↗

A successful technique for removing names in pathology reports using an augmented search and replace method.

The ability to access large amounts of de-identified clinical data would facilitate epidemiologic and retrospective research. Previously described de-identification methods require knowledge of natural language processing or have not been made available to the public. We take advantage of the fact that the vast majority of proper names in pathology reports occur in pairs. In rare cases where one proper name is by itself, it is preceded or followed by an affix that identifies it as a proper name (Mrs., Dr., PhD). We created a tool based on this observation using substitution methods that was easy to implement and was largely based on publicly available data sources. We compiled a Clinical and Common Usage Word (CCUW) list as well as a fairly comprehensive proper name list. Despite the large overlap between these two lists, we were able to refine our methods to achieve accuracy similar to previous attempts at de-identification. Our method found 98.7% of 231 proper names in the narrative sections of pathology reports. Three single proper names were missed out of 1001 pathology reports (0.3%, no first name/last name pairs). It is unlikely that identification could be implied from this information. We will continue to refine our methods, specifically working to improve the quality of our CCUW and proper name lists to obtain higher levels of accuracy.

Algorithms↗

A probabilistic information retrieval approach to medical annotation in SWISS-PROT.

The goal of medical annotation of human proteins in Swiss-Prot is to add features specifically intended for researchers working on genetic diseases and polymorphisms. For this purpose, it is necessary to search through a vast number of publications containing relevant information. Promising results have been obtained by applying natural language processing and machine learning techniques to solve this problem. By using the Probabilistic Latent Categorizer on representative query sets, 69% recall and 59% precision was achieved for relevant documents. This classifier also rejected irrelevant abstracts with more than 96% precision. Better linguistic pre-processing of source documents can further improve such computer approach.

Databases, Protein↗

IndexFinder: a method of extracting key concepts from clinical texts for indexing.

Extracting key concepts from clinical texts for indexing is an important task in implementing a medical digital library. Several methods are proposed for mapping free text into standard terms defined by the Unified Medical Language System (UMLS). For example, natural language processing techniques are used to map identified noun phrases into concepts. They are, however, not appropriate for real time applications. Therefore, in this paper, we present a new algorithm for generating all valid UMLS concepts by permuting the set of words in the input text and then filtering out the irrelevant concepts via syntactic and semantic filtering. We have implemented the algorithm as a web-based service that provides a search interface for researchers and computer programs. Our preliminary experiment shows that the algorithm is effective at discovering relevant UMLS concepts while achieving a throughput of 43K bytes of text per second. The tool can extract key concepts from clinical texts for indexing.

Abstracting and Indexing↗

A native XML database design for clinical document research.

Health-care institutions are gaining an increasing interest in exploiting the data that are gathered through electronic medical records. Narrative data, generated by transcription or direct entry, represents a far greater challenge for analytic tasks. Moreover, a small number of institutions are beginning to explore deeper structuring of narrative data using natural language processing (NLP). The data produced by NLP systems has a complex, nested structure. Current electronic medical records do not have the ability to store and retrieve data of this complexity in a suitable way.

Database Management Systems↗

Formative evaluation to guide early deployment of an online content management tool for medical curriculum.

KM is a Web-accessible, comprehensive database that organizes course materials (at the level of full lectures, not just outlines or syllabi) from the Vanderbilt School of Medicine curriculum. KM uses natural language processing techniques to analyze educational documents for biomedical concepts. Lecture handouts and Microsoft PowerPoint presentations are indexed and available online for students, faculty and administrators to search for individual or interrelated concepts across the medical school curriculum.

Anatomy↗

Representation of relationships between data in healthcare documents.

In healthcare data are to a large extent represented in narrative textual documents. The analysis of narrative text has not been solved successfully up to now. The results of natural language processing or fully indexing systems have not been completely satisfying. It is in particular difficult to detect and describe reliably relationships between data items in narrative text. The eXtensible Markup Language XML has opened new and very promising perspectives in this environment which could improve the storage, processing and analysis of narrative textual documents in particular in healthcare.

Germany↗

Electronic Patient Information -- Pioneers and MuchMore. A vision, lessons learned, and challenges.

OBJECTIVES: This paper must fulfill three different tasks: First, to introduce the topic "Electronic Patient Information -- Pioneers and MuchMore", second, to introduce the invited authors of the symposium, and third, to serve as the author's academic farewell lecture as professor emeritus. RESULTS: The electronic patient record, with all its different kinds of patient information, can be structured in many ways. Here, an historical approach is presented with a primary focus on the development of an information system for in- and outpatients in Germany, especially in Frankfurt, but also in comparison with US systems. The "Stone Age" and "Bronze Age" of patient-related computer applications started with expensive and insufficient hardware, but some years later, the first systems for patient documentation, text generation, and data acquisition could be implemented. The "iron age" and "golden age" yielded until the mid-1970s, e.g. in Oakland, Boston, Salt Lake City, and Frankfurt, quite successful Hospital Information Systems with some special emphasis on natural language processing. The following dark years were filled primarily with administrative systems, but beginning in the early 1980s, an era of enlightenment started, e.g. with rather inexpensive and easy to use PC application, broadly distributed MUMPS systems, and improved thesaurus-based text analysis. Especially in modern times, the medical text processing and classifying has been extended and successfully applied. CONCLUSIONS: Somewhat in contrast to other approaches, in the future the use of medical linguistics for the development of a successful electronic patient record should be better supported. Electronic patient information should be available wherever and whenever needed. For this, intelligent and automated reporting and controlled data exchange is necessary. The computer should do all classification, coding, and administrative work, and the physician should get all relevant information necessary for decision making.

Diffusion of Innovation↗

Decision support, knowledge representation and management: A broad methodological spectrum. Findings from the Decision Support, Knowledge Representation and Management.

OBJECTIVES: To summarize current excellent research in the field of decision support, knowledge management and representation. METHODS: Synopsis of the articles selected for the IMIA Yearbook 2006. RESULTS: Decision Support, Knowledge Representation and Management covers a broad spectrum of methods applied to a variety of medical problems and domains. Some particularly interesting and current topics were picked for the IMIA Yearbook 2006: the importance of ontologies for systematic system engineering of decision support systems, syndromic surveillance based on natural language processing, the evaluation of large semantic networks, and a comprehensive ontology for a randomised controlled trial database to support evidence-based practise. CONCLUSIONS: The best paper selection shows that methods for decision support, knowledge representation and management can decisively contribute to the solution of many different medical problems, but also that there is still a lot of exiting research to be done.

Awards and Prizes↗

Strategies for health information retrieval.

BACKGROUND: The amount of health data accessible on the Web is increasing and Internet has become a major source of health information. Many tools and search engines are available but medical information retrieval remains difficult for both the health professional and the patients. OBJECTIVE: In this paper we describe heuristics that aim at matching as much as possible queries with the content of the documents in the context of the CISMeF catalogue (Catalogue and Index of Health Resources in French) and its Doc'CISMeF search tool. The queries are represented by terms and the content of the documents is indexed by a terminology based on the MeSH thesaurus. RESULTS: Several operations are performed to match the terms of the terminology: natural language processing techniques on multi-words queries, phonemisation, spelling correction, plain text search with adjacency etc.. Each one is tested to evaluate its contribution in matching the terminology and the indexed documents. CONCLUSION: The implemented heuristics contribute significantly with good results in maximising as much as possible the recall of the Doc'CISMeF search tool.

France↗