Skip to main content
StreamVetree
Sign in
OrthopedicsCase series / Retrospective2 min read · distilled by Vetree AI

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Okur S, Özkalipçi Ç, Baykal B, Kartal T, Şenocak MG, Bedir AG, Yanmaz LE · Journal of the American Veterinary Medical Association · 14 August 2026

Clinical bottom line

LLMs show variable and unreliable surgical triage performance for feline fractures; clinician oversight is essential.

Summary

This retrospective study evaluated the diagnostic performance of five large language models (LLMs) in determining appropriate surgical versus conservative management for feline metacarpal and metatarsal fractures across 73 clinical cases. Two board-certified veterinary orthopedic surgeons established the reference standard, classifying 49 cases (67.1%) as requiring surgical intervention and 24 (32.9%) as amenable to conservative management. Each LLM assessed anonymized case summaries via standardized zero-shot prompting, with performance measured by accuracy, sensitivity, specificity, and Cohen κ. ChatGPT demonstrated the strongest performance (accuracy 84.9%, sensitivity 79.6%, specificity 95.8%, κ = 0.69), approaching expert-level agreement. In contrast, Qwen and Claude Sonnet exhibited 0% sensitivity, failing to identify any surgical cases. A consistent systematic bias toward conservative management was identified across all models, resulting in high false-negative rates and clinically significant undertriage risk. This undertriage pattern is particularly concerning in orthopedic triage, where delayed surgical intervention can lead to malunion, chronic pain, and reduced limb function. The study highlights that LLM diagnostic performance is highly variable and that no model reliably replicated expert surgical decision-making. The authors conclude that while top-performing LLMs may serve as adjunctive preliminary triage tools, independent use for surgical decision-making is not recommended. Strict clinician oversight remains essential when incorporating LLMs into clinical orthopedic workflows.

OrthopedicsSmall AnimalEmergencyRadiology

This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.