Evaluation of large language models for clinical sign-based oral assessment in dogs compared with veterinary practitioners.
Kocaman Y, Yanmaz LE, Okur S, Turgut F, Suzak Kocaman I, Aslan Canatan V, Golgeli Bedir A, Sahin O · Veterinary Journal · 24 February 2026
Current LLMs insufficient for independent oral diagnosis; novice vets outperform AI models.
This study evaluated the diagnostic accuracy of four Large Language Models (ChatGPT-5.1, ChatGPT-5.1 Thinking, Claude-Sonnet 4.5, and Gemini-Pro 3) for canine oral clinical sign assessment using 60 lateral oral photographs. Expert veterinary academicians established the reference standard for nine oral clinical signs. Novice veterinarians demonstrated superior diagnostic performance, achieving statistically significant agreement with experts in six clinical signs, including calculus and pigmentation. Among LLMs, Claude-Sonnet 4.5 showed the highest intra-model consistency (90% agreement, κ=0.75), while ChatGPT-5.1 models exhibited the strongest overall performance with weak-to-moderate agreement for several signs. LLMs outperformed novices only in traumatic lesion detection. All evaluators, including experts, showed chance-level performance for tooth fractures. The findings indicate that current LLMs lack sufficient diagnostic accuracy for independent clinical use in veterinary dentistry, though they show promise as supplementary tools with further development.
This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.