PubMed · 36993670
Benchmarking large language models for genomic knowledge with GeneTuring.
Abstract
Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinyi Shang, Xu Liao, Zhicheng Ji, Wenpin Hou. 2025-09-12. Benchmarking large language models for genomic knowledge with GeneTuring.. https://doi.org/10.1101/2023.03.11.532238
Cite the original work for its findings. Save a collection to share your selection of sources.