PubMed HealthSearch

SEARCH · PubMed Health

Results for “large language model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Knowledge-guided contextual gene set analysis with large language models.

MOTIVATION: Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. RESULTS: We introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis. AVAILABILITY AND IMPLEMENTATION: The demo website is publicly available at https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/cGSA/, while the data and code can be accessed at https://github.com/ncbi-nlp/cGSA.

Large Language Models

Knowledge-guided Contextual Gene Set Analysis Using Large Language Models.

Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. To address this limitation, we introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 manually curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis generation.

Journal Article

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Artificial intelligence agents and agentic artificial intelligence applied to precision medicine.

Precision medicine seeks to individualise care by integrating multimodal biomedical data, yet most deployed clinical artificial intelligence (AI) remains assistive, providing predictions without managing workflows or adapting autonomously. Agentic AI, built on large language models (LLMs), has emerged as a paradigm characterised by autonomy, goal-directed reasoning, memory, planning and tool use. This review synthesises evidence on agentic AI and LLMs applied to precision medicine, encompassing drug discovery, genomics, oncology, rare disease diagnostics and clinical pharmacology. This review also examines architectural components, recent validation milestones and emerging challenges, including hallucination, sociodemographic bias and evolving regulatory frameworks across the FDA, the EU AI Act and the WHO.

agentic AI

AI-generated familiarity estimates are a useful new source of information about word knowledge in Simplified Chinese.

This study evaluated the usefulness of AI-generated estimates of word familiarity for predicting word difficulty in Simplified Chinese, building on previous research in alphabetic languages. We found that familiarity estimates produced using large language models (LLMs) showed moderate-to-strong correlations with human familiarity ratings. These LLM estimates were the most effective predictors of both word naming and lexical decision times, surpassing traditional metrics such as word frequency and human familiarity ratings, while the latter still provided modest, non-overlapping variance. GPT-4o with English instructions produced superior results compared to the Chinese-centered models currently available. The results imply that LLM familiarity estimates are a valuable resource for Chinese psycholinguistics, supporting work across experimental design, modeling, and norming. We release familiarity estimates for 27,624 words for unrestricted research and educational use.

Humans

Systematic contextual biases in SegmentNT potentially relevant to other nucleotide transformer models.

Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT's raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT's prediction probabilities, revealing an intrinsic bias potentially linked to the model's training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.

Nucleotides

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning

Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.

OBJECTIVE: Evaluate the utility of the large language model (LLM), ChatGPT, for the analysis of operative notes and the generation of Current Procedural Terminology (CPT) codes in comparison to human clinical coders. STUDY DESIGN: CPT billing codes assigned by ChatGPT were compared to existing billing data. Otology practice within a tertiary academic center. METHODS: About 191 operative notes from a single surgeon (9/2022-10/2023) were analyzed. ChatGPT-3.5 and 4 models were prompted for CPT codes based on operative notes. Assessment included determining exact and partial match rates, sensitivity and specificity for targeted procedures, and work Relative Value Units (wRVU) differences between ChatGPT-generated and human-assigned codes. RESULTS: ChatGPT-3.5 achieved exact matches in 22% of cases and partial matches in 32%, while ChatGPT-4 achieved 14% exact and 33% partial matches. When cochlear implantation (CI) was excluded, performance dropped significantly. For CI, ChatGPT-3.5 demonstrated a sensitivity of 94% and specificity of 90%, while ChatGPT-4 showed a sensitivity of 96% and specificity of 92%. In contrast, performance on cartilage grafting was poor, with sensitivities of 4.2% for ChatGPT-3.5 and 0% for ChatGPT-4. ChatGPT-3.5 and 4 showed moderate CPT code matching accuracy among themselves, with slight agreement to human coders. Both models tended to underbill for wRVUs compared to human coders, with significant differences in the values generated. CONCLUSION: This study assessed ChatGPT's effectiveness in automating CPT code assignment for otologic surgeries. While the models achieved high sensitivity values for assigning codes related to cochlear implantation, both models struggled with complex cases, failed to apply modifiers, and often assigned fewer wRVUs. The findings highlight ChatGPT's potential in medical billing but indicate a need for further refinement.

Humans

Decoding gene regulation in plant genomes with artificial intelligence.

One of the central goals of plant functional genomics is to uncover regulatory mechanisms that shape agriculturally important traits to inform crop improvement. Recent advances in machine learning (ML) and artificial intelligence (AI), especially Large Language Models (LLMs), have greatly transformed our ability to derive regulatory information from complex genomics data. This review starts with a brief introduction of recent advances in AI and ML. We then present a plant-focused synthesis of emerging applications of AI- and LLM tools to: (i) predict epigenomic features, regulatory DNA elements, and gene expressions; (ii) infer gene regulatory network; and (iii) estimate post-transcriptional regulation.

Artificial intelligence

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

The Use of ChatGPT to Assist in Diagnosing Glaucoma Based on Clinical Case Reports.

INTRODUCTION: The purpose of this study was to evaluate the capabilities of large language models such as Chat Generative Pretrained Transformer (ChatGPT) to diagnose glaucoma based on specific clinical case descriptions with comparison to the performance of senior ophthalmology resident trainees. METHODS: We selected 11 cases with primary and secondary glaucoma from a publicly accessible online database of case reports. A total of four cases had primary glaucoma including open-angle, juvenile, normal-tension, and angle-closure glaucoma, while seven cases had secondary glaucoma including pseudo-exfoliation, pigment dispersion glaucoma, glaucomatocyclitic crisis, aphakic, neovascular, aqueous misdirection, and inflammatory glaucoma. We input the text of each case detail into ChatGPT and asked for provisional and differential diagnoses. We then presented the details of 11 cases to three senior ophthalmology residents and recorded their provisional and differential diagnoses. We finally evaluated the responses based on the correct diagnoses and evaluated agreements. RESULTS: The provisional diagnosis based on ChatGPT was correct in eight out of 11 (72.7%) cases and three ophthalmology residents were correct in six (54.5%), eight (72.7%), and eight (72.7%) cases, respectively. The agreement between ChatGPT and the first, second, and third ophthalmology residents were 9, 7, and 7, respectively. CONCLUSIONS: The accuracy of ChatGPT in diagnosing patients with primary and secondary glaucoma, using specific case examples, was similar or better than senior ophthalmology residents. With further development, ChatGPT may have the potential to be used in clinical care settings, such as primary care offices, for triaging and in eye care clinical practices to provide objective and quick diagnoses of patients with glaucoma.

Artificial intelligence (AI)

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial.

BACKGROUND: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin. METHODS: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0-10 points), with a prespecified non-inferiority margin of -0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment. RESULTS: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was -0.335 points (95% CI, -1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of -0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group. CONCLUSIONS: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.

Humans

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article