PubMed HealthSearch

SEARCH · PubMed Health

Results for “large language model (LLM)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence

Large Language Model and Knowledge Graph-Driven AJCC Staging of Prostate Cancer Using Pathology Reports.

Background/Objectives: To develop an automated American Joint Committee on Cancer (AJCC) staging system for radical prostatectomy pathology reports using large language model-based information extraction and knowledge graph validation. Methods: Pathology reports from 152 radical prostatectomy patients were used. Five additional parameters (Prostate-specific antigen (PSA) level, metastasis stage (M-stage), extraprostatic extension, seminal vesicle invasion, and perineural invasion) were extracted using GPT-4.1 with zero-shot prompting. A knowledge graph was constructed to model pathological relationships and implement rule-based AJCC staging with consistency validation. Information extraction performance was evaluated using a local open-source large language model (LLM) (Mistral-Small-3.2-24B-Instruct) across 16 parameters. The LLM-extracted information was integrated into the knowledge graph for automated AJCC staging classification and data consistency validation. The developed system was further validated using pathology reports from 88 radical prostatectomy patients in The Cancer Genome Atlas (TCGA) dataset. Results: Information extraction achieved an accuracy of 0.973 and an F1-score of 0.986 on the internal dataset, and 0.938 and 0.968, respectively, on external validation. AJCC staging classification showed macro-averaged F1-scores of 0.930 and 0.833 for the internal and external datasets, respectively. Knowledge graph-based validation detected data inconsistencies in 5 of 150 cases (3.3%). Conclusions: This study demonstrates the feasibility of automated AJCC staging through the integration of large language model information extraction and knowledge graph-based validation. The resulting system enables privacy-protected clinical decision support for cancer staging applications with extensibility to broader oncologic domains.

artificial intelligence

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.

OBJECTIVE: Evaluate the utility of the large language model (LLM), ChatGPT, for the analysis of operative notes and the generation of Current Procedural Terminology (CPT) codes in comparison to human clinical coders. STUDY DESIGN: CPT billing codes assigned by ChatGPT were compared to existing billing data. Otology practice within a tertiary academic center. METHODS: About 191 operative notes from a single surgeon (9/2022-10/2023) were analyzed. ChatGPT-3.5 and 4 models were prompted for CPT codes based on operative notes. Assessment included determining exact and partial match rates, sensitivity and specificity for targeted procedures, and work Relative Value Units (wRVU) differences between ChatGPT-generated and human-assigned codes. RESULTS: ChatGPT-3.5 achieved exact matches in 22% of cases and partial matches in 32%, while ChatGPT-4 achieved 14% exact and 33% partial matches. When cochlear implantation (CI) was excluded, performance dropped significantly. For CI, ChatGPT-3.5 demonstrated a sensitivity of 94% and specificity of 90%, while ChatGPT-4 showed a sensitivity of 96% and specificity of 92%. In contrast, performance on cartilage grafting was poor, with sensitivities of 4.2% for ChatGPT-3.5 and 0% for ChatGPT-4. ChatGPT-3.5 and 4 showed moderate CPT code matching accuracy among themselves, with slight agreement to human coders. Both models tended to underbill for wRVUs compared to human coders, with significant differences in the values generated. CONCLUSION: This study assessed ChatGPT's effectiveness in automating CPT code assignment for otologic surgeries. While the models achieved high sensitivity values for assigning codes related to cochlear implantation, both models struggled with complex cases, failed to apply modifiers, and often assigned fewer wRVUs. The findings highlight ChatGPT's potential in medical billing but indicate a need for further refinement.

Humans

Multimodal alignment improves generalizability of genomic biomarker prediction in computational pathology.

Computational pathology models that use digitized histopathology whole-slide images have the potential to become a cost-effective and scalable alternative to molecular assays for the prediction of genomic biomarkers, a key task in precision oncology. However, as new genomic biomarkers are discovered or quantified, large, labeled datasets must be prospectively collected to train new models. To address this challenge, we developed multimodal alignment for biomarker learning and generalization (MARBLE), a multimodal contrastive pretraining strategy that integrates structured biomarker knowledge into representation learning of histopathology images. MARBLE aligns histopathology-derived representations with representations of genomic biomarkers generated by a large language model (LLM) and a protein language model (PLM). This biologically informed alignment enables data-efficient generalization to novel, out-of-distribution biomarkers. Using the MSK-IMPACT cohort of over 40,000 patients across multiple biomarker panel versions, we design experiments grounded in real-world data to demonstrate the value of our proposed approach.

CP: computational biology

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). &#xa9; 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models

The Use of ChatGPT to Assist in Diagnosing Glaucoma Based on Clinical Case Reports.

INTRODUCTION: The purpose of this study was to evaluate the capabilities of large language models such as Chat Generative Pretrained Transformer (ChatGPT) to diagnose glaucoma based on specific clinical case descriptions with comparison to the performance of senior ophthalmology resident trainees. METHODS: We selected 11 cases with primary and secondary glaucoma from a publicly accessible online database of case reports. A total of four cases had primary glaucoma including open-angle, juvenile, normal-tension, and angle-closure glaucoma, while seven cases had secondary glaucoma including pseudo-exfoliation, pigment dispersion glaucoma, glaucomatocyclitic crisis, aphakic, neovascular, aqueous misdirection, and inflammatory glaucoma. We input the text of each case detail into ChatGPT and asked for provisional and differential diagnoses. We then presented the details of 11 cases to three senior ophthalmology residents and recorded their provisional and differential diagnoses. We finally evaluated the responses based on the correct diagnoses and evaluated agreements. RESULTS: The provisional diagnosis based on ChatGPT was correct in eight out of 11 (72.7%) cases and three ophthalmology residents were correct in six (54.5%), eight (72.7%), and eight (72.7%) cases, respectively. The agreement between ChatGPT and the first, second, and third ophthalmology residents were 9, 7, and 7, respectively. CONCLUSIONS: The accuracy of ChatGPT in diagnosing patients with primary and secondary glaucoma, using specific case examples, was similar or better than senior ophthalmology residents. With further development, ChatGPT may have the potential to be used in clinical care settings, such as primary care offices, for triaging and in eye care clinical practices to provide objective and quick diagnoses of patients with glaucoma.

Artificial intelligence (AI)

AI-driven CRISPR screening: optimizing gene editing through automation and intelligent decision support.

BACKGROUND: CRISPR-based genetic screening has become a central methodology in functional genomics, enabling systematic interrogation of gene function, genetic interactions and context-dependent vulnerabilities at scale. However, the rapid expansion of screening modalities-including multi-condition designs, combinatorial perturbations, in vivo applications and single-cell readouts-has exposed fundamental limitations of heuristic-driven experimental design and post hoc statistical analysis. MAIN BODY: This Review synthesizes how artificial intelligence is reshaping CRISPR screening by introducing predictive, adaptive and system-level intelligence across the experimental lifecycle. We organize recent advances into two tightly coupled modules. First, machine learning and deep learning (ML/DL) methods optimize experimental design by learning context-dependent perturbation behavior, anticipating confounding effects and enabling iterative, information-efficient screening strategies. Second, large language model-agent (LLM-agent) systems complement these advances by externalizing scientific reasoning, integrating biological knowledge at scale and coordinating analysis and decision-making in human-in-the-loop workflows. CONCLUSIONS: Together, ML/DL and LLM-agent approaches reframe CRISPR screening from a static analytical pipeline into an intelligent experimental system, with important implications for robustness, scalability and biological discovery.

Artificial Intelligence

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning

AI-generated familiarity estimates are a useful new source of information about word knowledge in Simplified Chinese.

This study evaluated the usefulness of AI-generated estimates of word familiarity for predicting word difficulty in Simplified Chinese, building on previous research in alphabetic languages. We found that familiarity estimates produced using large language models (LLMs) showed moderate-to-strong correlations with human familiarity ratings. These LLM estimates were the most effective predictors of both word naming and lexical decision times, surpassing traditional metrics such as word frequency and human familiarity ratings, while the latter still provided modest, non-overlapping variance. GPT-4o with English instructions produced superior results compared to the Chinese-centered models currently available. The results imply that LLM familiarity estimates are a valuable resource for Chinese psycholinguistics, supporting work across experimental design, modeling, and norming. We release familiarity estimates for 27,624 words for unrestricted research and educational use.

Humans

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models

Searching the druggable genome using large language models.

SUMMARY: The druggable genome encompasses the genes that are known or predicted to interact with drugs. The Drug-Gene Interaction Database (DGIdb) provides an integrated resource for discovering and contextualizing these interactions, supporting a broad range of research and clinical applications. DGIdb is currently accessed through structured web interfaces and API calls, requiring users to translate natural-language questions into database-specific query patterns. To allow for the use of DGIdb through natural language, we developed the DGIdb Model Context Protocol (MCP) server, which allows large language models (LLMs) access to up-to-date information through the DGIdb API. We demonstrate that the MCP server improves an LLM's ability to answer questions requiring accurate, up-to-date biomedical knowledge drawn from structured external resources. AVAILABILITY AND IMPLEMENTATION: The DGIdb MCP server is detailed at https://github.com/dgidb/dgidb-mcp-server and includes instructions for accessing the server through the Claude desktop app.

Large Language Models