Skip to main content
StreamVetree
Sign in
AnesthesiaCase series / Retrospective2 min read · distilled by Vetree AI

Assessing the performance of large language models when used to determine ASA status of cats and dogs and generate anaesthetic protocols.

Okur S, Akçora Y, Okur DT, Şenocak MG, Bedir AG, Ersöz U, Kartal T, Kurt BK · The Veterinary record · 12 May 2026

Clinical bottom line

LLMs show promise in veterinary anaesthesia but require expert oversight before clinical use.

Summary

This retrospective study evaluated the performance of three large language models (ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Pro) in determining American Society of Anaesthesiologists (ASA) classifications and generating anaesthetic protocols for 225 feline and canine cases. ChatGPT-5 demonstrated superior performance with 53.3% accuracy in ASA classification, compared to ChatGPT-4o (46.7%) and Gemini 2.5 Pro (30.7%). The models performed better at classifying higher-risk patients (ASA 3-5) but frequently overestimated ASA status in healthy animals (ASA 1). ChatGPT-5 also generated the most clinically adequate anaesthetic protocols as evaluated by experienced veterinary anaesthesiologists using a standardized assessment scale. While LLMs show promise as decision-support tools in veterinary anaesthesiology, their accuracy remains suboptimal, and expert veterinary oversight is essential before clinical implementation. The study's single-centre design and limitation to small animals may restrict generalizability to other species and practice settings.

AnesthesiaSmall AnimalPharmacology

This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.