Build1 publisher3 min readPublished
Most of 30 vision models misread the clock on synthetic Grafana panels in a Kaggle benchmark
Kaggle benchmark results posted on dev.to report that most of 30 vision models reading 336 synthetic Grafana-style panels found the peak but misread the clock. Copilot incident timelines drafted from screenshots need their start times checked by hand.
The Engineer · Build desk

What happened
- Gemini 3.7 Flash led at 0.978, with Gemini 3.8 Flash, Gemini 3.5 Flash and GPT-6 Astra all within 0.02 and inside overlapping confidence intervals.
- The most expensive run, Gemini 2.5 Pro at $9.80, scored 0.790, and Gemini 3.1 Pro at $7.75 trailed every current Gemini Flash model.
- Panels carry everyday traps such as log-scale axes, milliseconds labelled as seconds, truncated y-axes, decoy neighbour series and alert-line clutter.
- Canary questions about the y-axis label and dashed-marker count exposed three models that accept an image but cannot see it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Rolling back a deploy on a copilot's reading rests on the start time it took from the axis, so the implicated-deploy field carries any clock error forward.
- cost Moving up a vendor's price ladder bought worse chart reading on these panels, so cheaper Flash tiers and models like GPT-5.6 Luna are the first to trial for screenshot triage.
- exposure A copilot pipeline without canary questions would pass along confident readings from models that accept an image they cannot actually see.
- constraint The ranking holds for single cropped Grafana-style panels at 75 or 200 dpi; other dashboards or multi-panel captures need their own run before the scores transfer.
Each panel goes to the model as one image, in its own isolated chat, with a typed schema to fill [6]. Five fields are scored: the event type, the start time read off the x-axis as HH:MM, the implicated deploy, whether the metric is recovering, and the peak value [5]. A model's score is the mean of those five across every panel [5]. The output budget is 16,384 tokens per panel, set above every model's 99th-percentile output length [6]. At most one reading in a hundred per model could have reached that cap [1].
A mean over five fields hides where a model misses. Every model scored between 0.83 and 1.00 on peak value [10]. The author places the weakness on the clock. "Most of the rest find the peak but cannot read the clock, and when unsure they make the same mistake," the author wrote [9]. The available text of the post stops before the per-field table, so it does not show how large the start-time gap is or what the repeated mistake looks like.
The start time matters more than its one-fifth share of the score suggests. In the post's example panel, latency on auth-gw climbs from 15:26, right at deploy A, and then pins flat at 0.29 s [11]. Deploy A is implicated because the climb lines up with its marker. Panels carry up to three deploy markers [2]. A model that reads the wrong time off the axis has nothing correct to line them up against.
The leaderboard is tighter than it looks. Gemini 3.7 Flash leads GPT-6 Astra by 0.017 [2]. When 21 models were re-run on earlier versions with the same prompt, their scores moved by a median of 0.005 and never more than 0.015 [16]. The author reports overlapping confidence intervals for the top four [12]. Cost separates them more clearly. Astra's full run cost about 4.8 times Gemini 3.7 Flash's [3], and the Flash run works out to roughly 0.45 cents a panel [4]. Further down, GPT-5.6 Luna scored 0.877 for $0.31, against 0.871 for $5.41 on Claude Opus 5 [15].
The construction is careful. A generator renders the panels and records the ground truth, so nothing is hand-labelled and no model has seen the images before [4]. The best models still found three bugs in the answer key [18], a better review than most benchmarks get. The task is built on the kaggle-benchmarks SDK, with the PNGs and answer key in a Kaggle dataset attached to it [20]. The canary fields are trivial for any model that can see the image [7], so they separate two failures. A wrong canary means the model never saw the panel. Correct canaries alongside a wrong start time mean it saw the panel and misread the axis.
Transfer depends on input shape. The 336 panels come from 7 event shapes, dark and light themes, 75 and 200 dpi, and 12 variants [2]; 7 x 2 x 2 x 12 is 336 [5]. Each is a single Grafana-style panel with one to four series [2]. For these scores to carry over to a team's copilot, its screenshots have to resemble that: one cropped panel per image, at a resolution where the x-axis labels are legible. Every model read all 336 panels, 10,080 readings in total [17], and every one of them was a single panel.
What to watch
- Publication of per-field start-time scores, and a description of the mistake models repeat when unsure, would put a size on the clock problem.
- A rerun on real, uncropped incident screenshots, including multi-panel captures, would test whether the synthetic ranking holds outside the generator.