PubMed Health⌕ Search

PubMed · 36993670

Benchmarking large language models for genomic knowledge with GeneTuring.

Abstract

Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinyi Shang, Xu Liao, Zhicheng Ji, Wenpin Hou. 2025-09-12. Benchmarking large language models for genomic knowledge with GeneTuring.. https://doi.org/10.1101/2023.03.11.532238

Cite the original work for its findings. Save a collection to share your selection of sources.