Skip to main content
StreamVetree
Sign in
OncologyCase series / Retrospective2 min read · distilled by Vetree AI

Comparative diagnostic performance of large language models and clinicians for splenic diseases in dogs and cats.

İlgün M, Eren E, Kibar B, Okur S, Aktaş MS, Eroğlu MS, Arslan T, Baykal B, Petek B, Gürel A, Özdemir D · Veterinary Journal · 10 July 2026

Clinical bottom line

LLMs provide diagnostic support for splenic diseases exceeding novice clinicians but below expert level.

Summary

This two-center retrospective diagnostic-accuracy study evaluated large language models (ChatGPT-5 and Gemini 1.5 Pro) against experienced and novice veterinary clinicians in diagnosing splenic diseases in 38 dogs and cats. Clinicians received standardized multimodal case packets including clinical, laboratory, ultrasonographic, and intraoperative macroscopic data, with histopathology as the reference standard. Expert clinicians achieved the highest exact diagnosis accuracy (92.1% and 89.5%), followed by ChatGPT-5 (76.3%) and Gemini 1.5 Pro (71.1%), while novices performed lowest (57.9% and 52.6%). ROC analysis demonstrated consistent performance gradients (Expert > LLM > Novice) with AUCs of 0.97, 0.86, and 0.71 respectively. Upper-category classification accuracy exceeded exact diagnosis across all groups. Adding imaging and macroscopic data improved performance for all assessors. The study demonstrates that modern LLMs can provide meaningful diagnostic triage support exceeding novice performance, though not yet equivalent to expert-level classification in zero-shot settings.

OncologySoft Tissue SurgerySmall AnimalRadiologyPathology

This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.