Science1 publisher3 min readPublished
A horse pain detector sets a testable standard for explainable AI
SHIC-XE scores equine pain from video and, its authors say, is the first system whose visual explanations were measured against expert clinical judgement. The second half is the part worth copying.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A new AI framework known as SHIC-XE detects signs of pain in horses from video analysis while providing stable, anatomically consistent explanations for its decisions.
- SHIC-XE was developed by an international team of researchers led by Dr. Marcelo Feighelstein, head of the Artificial Intelligence Systems Engineering Program at Tel-Hai University's new Cluster of Engineering and Advanced Computing.
- For the first time, these AI-generated explanations can be quantitatively compared with expert assessments, described as a significant breakthrough in explainable artificial intelligence.
- The study was published in the International Journal of Computer Vision.
- Existing explainability methods for AI-based video analysis often highlight different areas of an image from one video frame to the next, creating unstable and difficult-to-interpret explanations.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
An international team led by Marcelo Feighelstein, who heads the Artificial Intelligence Systems Engineering Program at Tel-Hai University's new Cluster of Engineering and Advanced Computing, has published a framework called SHIC-XE that detects signs of pain in horses from video while producing stable, anatomically consistent explanations for its decisions [1][2]. The horses matter less than the procedure: according to the researchers, this is the first time such AI-generated explanations have been quantitatively compared against expert assessment, which is the check that most explainability work has skipped [3][10].
The work appears in the International Journal of Computer Vision [4]. The technical problem it targets is mundane and familiar to anyone who has shipped a saliency map: existing methods tend to highlight different regions of an image from one video frame to the next, so the explanation is unstable and hard to interpret [5]. SHIC-XE's answer is to project the model's attention onto a fixed three-dimensional representation of a horse's face, so the highlighted region keeps its anatomical identity even as the animal moves its head or the camera angle changes [6]. That is what makes the explanation addressable at all. A heatmap that wanders cannot be compared with anything; a heatmap pinned to named anatomy can be scored.
On detection, the framework was evaluated on three independent datasets covering post-surgical pain, inflammatory and orthopedic pain, and acute mechanical pain, with F1 scores from 0.67 to 0.80 [7][8]. That is a spread of 0.13 across clinical scenarios [15], and the bottom of the range is not a number anyone should treat as a triage decision on its own. The more interesting result is the explanation audit: the regions the model attended to showed statistically significant agreement with veterinary experts scoring the same animals on the Horse Grimace Scale, a standardised facial-expression instrument, with the agreement noted particularly in the ears and cheek muscles [9][11]. Feighelstein describes this as demonstrating that the model "not only reaches the correct conclusion but also focuses on the anatomically relevant regions when making that decision" [10].
What the announcement does not give is the machinery a reader would need to judge the strength of that agreement: no count of experts, no named statistical test, no effect size, and no indication of whether regions beyond the ears and cheeks agreed or disagreed [16]. "Statistically significant agreement" on a small panel of raters is a weak floor, and the honest reading is that a method now exists, not that it has been stress-tested. The group has previously built pain and emotion recognition tools for cats, dogs, rabbits, sheep and cattle, so there is a pipeline behind this rather than a one-off [12].
The stated ambition is human: pain assessment in newborns, monitoring of sedated and ventilated intensive care patients, dementia care, neurological movement disorders and surgical video analysis [13]. All of those inherit the same requirement, which Feighelstein frames as ensuring the voice the system gives a patient is reliable, explaining not only that pain is present but why the system concluded it [14].
Watch whether the expert-agreement metric shows up in other groups' papers as a reported number rather than a claim, and whether it survives a larger rater panel. Watch also whether the fixed-anatomy projection transfers to human faces and bodies, because without a canonical surface to project onto, the stability that made this validation possible disappears [6].