PubMed HealthSearch

SEARCH · PubMed Health

Results for “ChatGPT”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

6 recordsLinked to original sources

Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.

OBJECTIVE: Evaluate the utility of the large language model (LLM), ChatGPT, for the analysis of operative notes and the generation of Current Procedural Terminology (CPT) codes in comparison to human clinical coders. STUDY DESIGN: CPT billing codes assigned by ChatGPT were compared to existing billing data. Otology practice within a tertiary academic center. METHODS: About 191 operative notes from a single surgeon (9/2022-10/2023) were analyzed. ChatGPT-3.5 and 4 models were prompted for CPT codes based on operative notes. Assessment included determining exact and partial match rates, sensitivity and specificity for targeted procedures, and work Relative Value Units (wRVU) differences between ChatGPT-generated and human-assigned codes. RESULTS: ChatGPT-3.5 achieved exact matches in 22% of cases and partial matches in 32%, while ChatGPT-4 achieved 14% exact and 33% partial matches. When cochlear implantation (CI) was excluded, performance dropped significantly. For CI, ChatGPT-3.5 demonstrated a sensitivity of 94% and specificity of 90%, while ChatGPT-4 showed a sensitivity of 96% and specificity of 92%. In contrast, performance on cartilage grafting was poor, with sensitivities of 4.2% for ChatGPT-3.5 and 0% for ChatGPT-4. ChatGPT-3.5 and 4 showed moderate CPT code matching accuracy among themselves, with slight agreement to human coders. Both models tended to underbill for wRVUs compared to human coders, with significant differences in the values generated. CONCLUSION: This study assessed ChatGPT's effectiveness in automating CPT code assignment for otologic surgeries. While the models achieved high sensitivity values for assigning codes related to cochlear implantation, both models struggled with complex cases, failed to apply modifiers, and often assigned fewer wRVUs. The findings highlight ChatGPT's potential in medical billing but indicate a need for further refinement.

Humans

The Use of ChatGPT to Assist in Diagnosing Glaucoma Based on Clinical Case Reports.

INTRODUCTION: The purpose of this study was to evaluate the capabilities of large language models such as Chat Generative Pretrained Transformer (ChatGPT) to diagnose glaucoma based on specific clinical case descriptions with comparison to the performance of senior ophthalmology resident trainees. METHODS: We selected 11 cases with primary and secondary glaucoma from a publicly accessible online database of case reports. A total of four cases had primary glaucoma including open-angle, juvenile, normal-tension, and angle-closure glaucoma, while seven cases had secondary glaucoma including pseudo-exfoliation, pigment dispersion glaucoma, glaucomatocyclitic crisis, aphakic, neovascular, aqueous misdirection, and inflammatory glaucoma. We input the text of each case detail into ChatGPT and asked for provisional and differential diagnoses. We then presented the details of 11 cases to three senior ophthalmology residents and recorded their provisional and differential diagnoses. We finally evaluated the responses based on the correct diagnoses and evaluated agreements. RESULTS: The provisional diagnosis based on ChatGPT was correct in eight out of 11 (72.7%) cases and three ophthalmology residents were correct in six (54.5%), eight (72.7%), and eight (72.7%) cases, respectively. The agreement between ChatGPT and the first, second, and third ophthalmology residents were 9, 7, and 7, respectively. CONCLUSIONS: The accuracy of ChatGPT in diagnosing patients with primary and secondary glaucoma, using specific case examples, was similar or better than senior ophthalmology residents. With further development, ChatGPT may have the potential to be used in clinical care settings, such as primary care offices, for triaging and in eye care clinical practices to provide objective and quick diagnoses of patients with glaucoma.

Artificial intelligence (AI)

AI Health message intervention: The role of message customization and message source in breast cancer screening among women of color.

OBJECTIVES: To examine the effectiveness of breast cancer screening messages with varying levels of customization (generic, targeted, and tailored) and to compare AI-generated versus human-generated messages. METHODS: A between-subjects experimental design with a control condition was employed. Message content followed a standardized structure and varied by level of customization: generic, targeted (demographic-based), and tailored (perceived susceptibility- and barrier-based). Messages were developed by either the authors or GenAI (ChatGPT-4o). A total of 391 participants recruited via Prolific were randomly assigned to five groups (generic, targeted-human, targeted-AI, tailored-human, and tailored-AI). Self-efficacy, behavioral intentions, attitudes, and message believability were measured using different scales. RESULTS: Customized (tailoring and targeting) health messages performed comparably to generic messages in shaping positive health outcomes. GenAI-generated messages also produced outcomes comparable to those of human-generated messages under standardized conditions. Significant negative indirect effects through message believability for the human-tailored condition was found relative to the generic condition. CONCLUSIONS: GenAI may be a useful tool for developing and customizing scalable health messages. Its effectiveness depends not only on customization but also on maintaining message quality, including readability, clarity, coherence, naturalness, and credibility. PRACTICAL IMPLICATIONS: GenAI may support health practitioners in developing customized and scalable breast cancer messages. However, professional review remains necessary to ensure that the message is culturally appropriate, responsive to patient concerns, and suitable for use alongside patient-provider communication.

Humans

Drug repurposing in status epilepticus.

The treatment of status epilepticus (SE) has changed little in the last 20 years, largely because of the high risks and costs of new drug development for SE. Moreover, SE poses specific challenges to drug development, such as patient diversity, logistical hurdles, and the need for acute treatment strategies that differ from chronic seizure prevention. This has reduced the appetite of industry to develop new drugs in this area. Drug repurposing is an attractive approach to address this unmet need. It offers significant advantages, including reduced development time, lower costs, and higher success rates, compared to novel drug development. Here I demonstrate how novel methods integrating biological knowledge and computational methods can be applied to drug repurposing in status epilepticus. Biological approaches focus on addressing mechanisms underlying drug resistance in SE (using for example ketamine, tacrolimus and safinamide) and longer-term consequences (using for example omaveloxolone, celecoxib and losartan). Additionally, artificial intelligence platforms, such as ChatGPT, can rapidly generate promising drug lists, while in silico methods can analyze gene expression changes to predict molecular targets. Combining AI and in silico approaches has identified several candidate drugs, including metformin, sirolimus and riluzole, for SE treatment. Despite the promise of repurposing, challenges remain, such as intellectual property issues and regulatory barriers. Nonetheless, drug repurposing presents a viable solution to the high costs and slow progress of traditional drug development for SE. This paper is based on a presentation made at the 9th London-Innsbruck Colloquium on Status Epilepticus and Acute Seizures, in April 2024.

Animals

Benchmarking large language models for genomic knowledge with GeneTuring.

Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.

Benchmark

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design