Performance of large language models versus clinicians and novices in veterinary theriogenology decision support.
Okur DT, Cengiz M, Küçükaslan İ, Peker C, Çiplak AY, Tohumcu V, Aydın Ş · Journal of the American Veterinary Medical Association · 13 February 2026
ChatGPT-5 Thinking performed comparably to expert clinicians in veterinary reproductive case management.
This study evaluated the clinical decision-support capabilities of two large language models (ChatGPT-5 and ChatGPT-5 Thinking) compared to experienced clinicians and novice veterinarians in veterinary theriogenology. Fifteen standardized obstetric and gynecologic scenarios were independently assessed by two expert clinicians, two novice veterinarians, and both LLMs under identical conditions. A blinded expert panel evaluated all responses using a standardized 5-point quality scoring system. Results demonstrated that ChatGPT-5 Thinking achieved the highest overall quality ratings, followed by ChatGPT-5, expert clinicians, and novice veterinarians. LLM-generated responses demonstrated greater consistency and completeness compared to human responses. Within the constraints of simulated scenarios, LLMs—particularly ChatGPT-5 Thinking—provided clinically appropriate guidance that exceeded novice performance and approached expert clinician standards. The authors conclude that LLMs show promise as adjunct decision-support tools in time-sensitive reproductive emergencies, potentially assisting both clinicians and trainees through rapid, structured, guideline-aligned recommendations. However, further validation in actual clinical practice settings is recommended before widespread implementation.
This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.