Security1 distinct publisher3 min readPublished
Two Bern researchers solved a rotation CAPTCHA in 0.006 seconds using circle detection from the 1970s, then fed that answer to frontier models as a tool result and watched one of them argue with it.
The Watch · Security desk
Compiled by The WatchSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
build
Four months of A100 bills say self-hosting is a utilization bet, not a cost saving1 distinct publisher
product
Anthropic's usage policy says no explicit content. Opus 4.6 said yes 10 times out of 10.1 distinct publisher
security
GitLab 19.3 puts agent runtime, inference models and secrets under one permission model1 distinct publisher
The override has a clean mechanism. The script does not judge the image on its own terms; it fits the alignment to the answer space the site will accept, which is 45-degree steps [10]. One test image, taken from a live site and slightly crooked to begin with [7], lined up best at roughly 129 degrees, so the script proposed 135 and that was the correct submission [10]. Gemini looked at 135, judged it still a little off, and substituted its own read [9]. The model's look at the image was fine; what it lacked was the tool's information about what the grader would accept.
Unassisted, Gemini 3.1 Pro needed 67 seconds per puzzle to reach 7 out of 10 [4]. Against 0.006 seconds [1] that is a factor of about 11,000 [1], for a worse result than the script's 10 out of 10 [3]. The other two models are weaker than the raw score suggests. With eight possible answers [6], blind guessing returns an expected 1.25 correct in ten, and GPT-4o and Grok each returned 1 [5][2].
The tool-use round is the part worth keeping. Handed the script's output, GPT-4o and Grok relayed it and scored perfectly [8]. Gemini adjudicated and lost a fifth of its score [11], which on a ten-puzzle run is two puzzles, or about 8 out of 10 [4]. The tool and the correct answer were identical; only the wrapper's policy differed. All of the loss sat in the review step.
The other tasks point the same way from the opposite side. On the open-circle puzzles the models described the target correctly, one calling it a cyan ring in the lower central area, which it was, and still could not produce a coordinate accurate enough to click it [12]. On the nine-square insect grid, Grok found the insects and then mis-numbered the cells, calling the middle-left square 5 when it was 4 and at one point answering square 10 in a grid of nine [13]. Turning on its reasoning mode, with no change to what it was looking at, moved it from near-useless to nearly perfect [14].
Treat this as a demonstration rather than a benchmark: it covers ten puzzles across three models within a single CAPTCHA family, read through Help Net Security's account of the paper [1]. The bypass economics still deserve naming, because 0.006 seconds works out to roughly 167 solves per second [3]; Help Net Security calls it 160 [16]. These puzzles are pure geometry because the sites hosting them cannot run JavaScript, their visitors having switched it off to avoid fingerprinting, which strips out mouse movement, click timing and the rest of the behavioral signals mainstream CAPTCHAs lean on [15]. That privacy reading is the publisher's, offered as something the authors walk past [16].
The design question for a security team is narrow: whether the model is permitted to disagree with a deterministic tool when that tool is right. In this study, permission cost two puzzles in ten.
Ranked by verification strength, evidence, and original report placement.
Two researchers at Bern University of Applied Sciences wrote a script that solves a rotation CAPTCHA in 0.006 seconds.
The script uses circle-detection math from the 1970s and a signal-matching technique that predates current AI methods.
Gemini 3.1 Pro took 67 seconds and got seven out of ten rotation puzzles correct.
GPT-4o and Grok each got one of ten rotation puzzles correct.
The rotation CAPTCHA has only eight possible answers, since the site accepts submissions in 45-degree steps.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific but single-sourced
The detail is unusually granular for a secondhand write-up: a 129-degree best fit rounded to the only legal 135, a grid answer of 'square 10' where nine exist, a per-puzzle timing to the millisecond. The arithmetic that follows from those numbers holds. What it rests on is one publisher's reading of a Bern paper that our coverage never quotes at length, over ten puzzles per task, with no sign the runs were repeated and no independent hand on the script.
No uptake to measure
Nothing here tells us the script exists outside a research repo or that anyone changed anything because of it. No site swapping out its rotation puzzle, no named CAPTCHA product, no traffic or abuse data, not even a release the reader could go and look at.
Modest findings, generalised a step early
The numbers themselves are stated straight — no rounding in the favourable direction, and the piece even volunteers that guessing would beat two of the models. The stretch is in the moral: two lost puzzles in one run against one model version become a rule about language models supervising tools that already work. It is a good instinct and probably right, but it is a hypothesis wearing the clothes of a finding, and the privacy section widens the frame further than the underlying test does.
Nobody selling, one visible angle
There is no product in this story. Two academics, three models nobody appears to have paid to test, and not a single CAPTCHA vendor named or defended. The one thumb on the scale is editorial and openly declared: the privacy reading of the results is Help Net Security's, and it says so by pointing out the authors walked past it. Declared angles distort less than hidden ones, but a security publication has a standing appetite for defences that turn out to be geometry.
Believable shape, unverified detail
Enough to trust the direction, not enough to redesign a system on. What lifts it is texture that fabrication rarely produces: an off-by-one grid index, a rounding that was correct and got vetoed anyway. What holds it down is that all of it arrives through one publisher, from a paper readers cannot inspect, on ten puzzles per task — the kind of gap a second look at the source would close in an afternoon.