PubMed HealthSearch

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Suffix interference in the recall of linguistically coherent speech.

Four experiments are presented that address the stimulus suffix effect for linguistically coherent spoken materials. In Experiment 1, definitions of low-frequency words were presented for online written recall. Each definition was followed by a nonword speech suffix presented in the same voice as the definition, the same nonword presented in a different voice, or a tone. The results yielded a significant reduction in the recall of the terminal words of the definitions in the speech suffix conditions compared with the tone control. This general pattern was replicated in Experiment 2, in which subjects did not begin their recall until the suffix item or tone was presented, although the magnitude of the suffix effect was reduced in this experiment. In Experiment 3, sentences that were part of a cohesive story were presented for on-line recall. Here, the suffix effect was considerably reduced compared with the suffix effect found with the definitions presented in Experiments 1 and 2. This pattern was replicated in Experiment 4, in which subjects did not begin their recall of the story sentences until the speech suffix or tone was presented. Overall, the results suggest that auditory memory interference can take place for linguistically coherent speech, although the magnitude of the interference decreases as one increases the level of linguistic structure in the to-be-recalled materials. Implications of the present results for current models of natural language processing are discussed.

Adult

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article

Artificial Intelligence for Colorectal Surgeons-Part II: Research Applications, Challenges in Adoption, and Practical Resources.

BACKGROUND: This is part II of a 2-part series examining artificial intelligence in colorectal surgery. Part I established foundational concepts and clinical applications. Implementation, however, requires understanding research methodologies, available resources, and the specific challenges currently limiting widespread adoption. These topics are the focus of part II. OBJECTIVE: To examine artificial intelligence's transformation of surgical research, provide practical implementation resources, address adoption challenges, and explore future directions in colorectal surgery. METHODS: Comprehensive literature review focusing on artificial intelligence research methodology, implementation barriers, educational resources, and emerging technologies relevant to colorectal surgeons. RESULTS: Artificial intelligence streamlines clinical trial design through predictive modeling and natural language processing, reducing enrollment challenges that contribute to failed or inadequate trial accrual. Machine learning enables heterogeneity analysis within clinical trials, identifying treatment-responsive subgroups. Foundation models unlock analysis of unstructured electronic health record data at scale. Professional societies and universities offer specialized artificial intelligence education programs, with open-access data sets facilitating research participation. However, implementation faces multifaceted challenges: technical infrastructure demands, with real-time processing requiring dedicated graphics processing unit clusters; regulatory frameworks struggling with continuously evolving algorithms; undefined liability distribution for artificial intelligence-assisted decisions; algorithmic bias risking health care disparities; and the "black box" problem limiting clinical trust. Economic barriers include substantial initial costs without clear reimbursement pathways. Future directions include multimodal artificial intelligence integrating imaging, genomics, and histopathology; cognitive robotic systems with real-time decision support; digital twin technology for patient-specific surgical simulation; and global surgical artificial intelligence networks enabling distributed learning across institutions. CONCLUSIONS: Although artificial intelligence offers transformative potential for colorectal surgery research and practice, successful implementation requires addressing technical, regulatory, ethical, and economic challenges. The surgeon's evolving role demands both traditional expertise and computational fluency. Future advances in multimodal integration, autonomous systems, and global collaboration will fundamentally reshape surgical practice but will require thoughtful implementation prioritizing patient benefit and clinical value.

Humans

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article

Development of a questionnaire for detecting potential adverse drug reactions.

OBJECTIVE: To develop a comprehensive list of symptoms categorized by body system as part of a questionnaire for detecting potential adverse drug reactions. DATA SOURCES: A preliminary list of symptoms in lay terminology was extracted from the "Side Effects" section of all drug monographs contained in the United States Pharmacopeia Dispensing Information (USP DI) computerized database (Volume II, Advice for the Patient) using natural language processing software. The list was sorted alphabetically and duplicate terms were eliminated. Symptoms were then categorized by body system or anatomic region. A preferred term for each symptom was selected when multiple synonyms and related words were listed. Finally, all of the symptom terms were incorporated into a thesaurus from which the questionnaire was derived. RESULTS: The questionnaire will be used as part of a computer-assisted interview, developed to solicit information from patients regarding their medication regimens and to systematically query them regarding the presence of salient symptoms or complaints. The computer system will eventually interface with the USP DI database to identify drugs from a patient's regimen that may be associated with adverse symptoms. The symptom thesaurus will provide the link to the USP DI database. Preliminary experience with the questionnaire in a limited number of patients has been encouraging. CONCLUSIONS: The questionnaire can assist clinicians in identifying drug-related symptoms including unreported adverse clinical effects of newly marketed or investigational therapeutic agents. When the questionnaire is computerized and linked to a comprehensive database, it can be more widely used to alert healthcare providers of potential adverse drug reactions that may otherwise go undetected.

Adverse Drug Reaction Reporting Systems

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

Patient data acquisition.

Patient data are acquired in three ways: by direct interrogation; physiological measurements; and analysis of specimens, signals, and images. When computers are used to acquire information from the patient directly, difficulties arise from the lack of standardized patient medical history, the complexities of natural language processing, and the problems of man/machine communication (patient with computer terminal, and physician with computer-generated history). A great variety of data input devices have been used for the acquisition of the patient medical history. Most have been extensively used in multiphasic health testing programs, and this experience is freely drawn upon in this paper.

Computers

A Medical Text Analysis System for German--syntax analysis.

Much information about patients is stored in free text. Hence, the computerized processing of medical language data has been a well-known goal of medical informatics resulting in different paradigms. In Göttingen, a Medical Text Analysis System for German (abbr. MediTAS) has been under development for some time, trying to combine and to extend these paradigms. This article concentrates on the automated syntax analysis of German medical utterances. The investigated text material consists of 8,790 distinct utterances extracted from the summary sections of about 18,400 cytopathological findings reports. The parsing is based upon a new approach called Left-Associative Grammar (LAG) developed by Hausser. By extending considerably the LAG approach, most of the grammatical constructions occurring in the text material could be covered.

Electronic Data Processing

A free-text processing system to capture physical findings: Canonical Phrase Identification System (CAPIS).

The task of gathering detailed patient information from free-text medical records presents a significant barrier to clinical research. In this paper, we describe a prototype system for extracting physical examination findings from dictated admission summaries. Our computer program applies a concept-based free-text processing algorithm that identifies user-selected target physical examination findings. We are using the extraction system to enrich an existing clinical database. The system was evaluated by comparing the physical examination findings extracted by our computer program with findings extracted by an independent investigator. Our prototype system was able to recall 92 percent (sensitivity) of the relevant physical findings, with a precision of 96 percent (positive predictive value).

Algorithms

Automated analysis of medical text. I. Clue gathering.

Clinical practice of medicine is highly information-intensive. At the bedside, past experience is the primary justification of reasoning and decisions. This past medical experience is an amalgamation of textbook information and personal experience. During the last 2-3 decades, both of these major sources of clinical information have appeared less and less effective. The pace of progress, resulting in better diagnostic tools and new therapies, has undermined our personal experience, and for the same reason, the time lapse between drafting the manuscripts and distributing the textbooks has become a growing problem. Emphasis has shifted from textbooks to scientific journals with shorter publishing delays, and the role of daily newspapers and television programs seems to be growing. The traditional ways of gathering clinical knowledge and experience seem to fail more and more. In addition to textbooks and scientific journals, current clinical experience is described in millions of patient records, stored in hospitals and ambulatory care offices. However, we have no easy access to patient charts, and we are lacking a method for cost-effective merging of clinical case histories to make them suitable for much-needed statistical inferences. Computers could make a major contribution in this area, but first we must bridge the gap between the narrative text in the medical record and computer technology. Recently, much encouraging progress has been made in automated medical text processing, the topic of this paper.

Abstracting and Indexing

The impact of tokenizer selection in genomic language models.

MOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130.

Natural Language Processing

Evaluation of Meta-1 for a concept-based approach to the automated indexing and retrieval of bibliographic and full-text databases.

SAPHIRE is a concept-based approach to information retrieval in the biomedical domain. Indexing and retrieval are based on a concept-matching algorithm that processes free text to identify concepts and map them to their canonical form. This process requires a large vocabulary containing a breadth of medical concepts and a diversity of synonym forms, which is provided by the Meta-1 vocabulary from the Unified Medical Language System Project of the National Library of Medicine. This paper describes the use of Meta-1 in SAPHIRE and an evaluation of both entities in the context of an information retrieval study.

Abbreviations as Topic

A free-text processing system to capture physical findings: Canonical Phrase Identification System (CAPIS).

The task of gathering detailed patient information from free-text medical records presents a significant barrier to clinical research. In this paper, we describe a prototype system for extracting physical examination findings from dictated admission summaries. Our computer program applies a concept-based free-text processing algorithm that identifies user-selected target physical examination findings. We are using the extraction system to enrich an existing clinical database. The system was evaluated by comparing the physical examination findings extracted by our computer program with findings extracted by an independent investigator. Our prototype system was able to recall 92 percent (sensitivity) of the relevant physical findings, with a precision of 96 percent (positive predictive value).

Academic Medical Centers

TEXTINFO: a tool for automatic determination of patient clinical profiles using text analysis.

The clinical data contained in narrative patient documents is made available via grammatical and semantic processing. Retrievals from the resulting relational database tables are matched against a set of clinical descriptors to obtain clinical profiles of the patients in terms of the descriptors present in the documents. Discharge summaries of 57 Dept. of Digestive Surgery patients were processed in this manner. Factor analysis and discriminant analysis procedures were then applied, showing the profiles to be useful for diagnosis definitions (by establishing relations between diagnoses and clinical findings), for diagnosis assessment (by viewing the match between a definition and observed events recorded in a patient text), and potentially for outcome evaluation based on the classification abilities of clinical signs.

Databases, Factual

Computerized extraction of coded findings from free-text radiologic reports. Work in progress.

A computerized data acquisition tool, the special purpose radiology understanding system (SPRUS), has been implemented as a module in the Health Evaluation through Logical Processing Hospital Information System. This tool uses semantic information from a diagnostic expert system to parse free-text radiology reports and to extract and encode both the findings and the radiologists' interpretations. These coded findings and interpretations are then stored in a clinical data base. The system recognizes both radiologic findings and diagnostic interpretations. Initial tests showed a true-positive rate of 87% for radiographic findings and a bad data rate of 5%. Diagnostic interpretations are recognized at a rate of 95% with a bad data rate of 6%. Testing suggests that these rates can be improved through enhancements to the system's thesaurus and the computerized medical knowledge that drives it. This system holds promise as a tool to obtain coded radiologic data for research, medical audit, and patient care.

Artificial Intelligence