Pilot study: external validation of commercial veterinary radiology artificial intelligence services shows deficiencies in interpretation of general practice-sourced canine abdominal radiographs.
Ma D, Faulkner JE, Stander N, Raisis A, Joslyn S · Journal of the American Veterinary Medical Association · 20 March 2026
Current commercial veterinary radiology AI platforms are unsuitable for clinical use due to high missed diagnosis rates.
This pilot study evaluated the diagnostic performance of six commercial veterinary radiology artificial intelligence platforms on canine abdominal radiographs with confirmed diagnoses. Fifty-three cases were submitted to the platforms, yielding 307 evaluations after rejections. Performance metrics were variable and generally suboptimal, with mean accuracy ranging from 70-90%, balanced accuracy 60-65%, and Matthews correlation coefficient -0.08 to 0.43. Sensitivity for radiographic finding classification was consistently low (28-78%), with F1 scores and positive predictive values between 25-54%, indicating frequent missed diagnoses. Notably, small intestinal obstruction, a critical finding requiring urgent intervention, demonstrated sensitivity of only 23-69% across platforms. While Matthews correlation coefficient performed better (0.16-0.45) as it was less affected by label misclassification, even the best-performing algorithm exhibited significant limitations. The authors concluded that none of the tested AI platforms are currently suitable for clinical use. The study emphasizes the need for larger-scale independent external validations and substantial performance improvements before veterinary AI platforms can be safely integrated into routine clinical practice for abdominal radiograph interpretation.
This summary was distilled by AI and may occasionally misinterpret data. Confirm critical details with the primary literature before clinical application.