Skip to content

Build1 publisher3 min readPublished

The click succeeded and nothing happened: your agent harness needs an injected canary

A session write-up on dev.to shows a browser agent reporting a click at coordinates it had verified while the page counted zero arrivals. The cause was a scale factor no layer reported.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The author spent a day pointing an AI agent at a browser to publish one product across four marketplaces.
  • The automation tool reported success with the log line: [computer:left_click] Clicked at (383, 734).
  • After the reported click, nothing happened: no menu opened, no network request fired, no console error, and the failures came with no error, no exception and no log line.
  • document.elementFromPoint(383, 734) returned the exact target element, a span, which was present, visible, not disabled and not covered by an overlay.
  • The author burned close to an hour assuming the site was blocking synthetic input.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer pointing an AI agent at a browser to publish one product across four marketplaces got a clean tool log line, `[computer:left_click] Clicked at (383, 734)`, and a page that did nothing: no menu opened, no network request fired, no console error [1][2][3]. That combination, success at the tool layer and silence everywhere else, is the case your computer-use harness has to be designed around, because nothing in the stack was wired to contradict the tool.

The element was not the problem. `document.elementFromPoint(383, 734)` returned exactly the span the agent was aiming at, visible, not disabled, not covered [4]. According to the write-up, close to an hour went into the theory that the site was rejecting synthetic input [5].

The probe that ended it is a few lines: inject a fixed-position button 220 by 60 pixels at z-index 2147483647, attach a listener that increments `window.__clickOK` and records `e.isTrusted`, then measure the button's centre, 310,230 [6][7]. Click that centre and read the counter. Zero [8]. That reading discriminates, which is the point. A blocked synthetic event still arrives and is merely ignored by the handler; a counter stuck at zero means the event never reached the element at all [9].

The coordinate space was wrong. The tool consumed screenshot-space coordinates while the page reported CSS pixels, and the screenshot was captured at a different scale, so every DOM-derived coordinate was off by a constant factor while the click itself still happened somewhere [10]. `devicePixelRatio` was 2 and doubling missed [11]. Clicking (232,172) instead of (310,230) registered one hit [12], giving k = 232/310 = 0.7484 [13], a 25 percent shrink [4]. The viewport was 1702 CSS pixels wide and screenshots came back around 1274, a ratio of 0.7485 [14].

Two bits of arithmetic explain why this hides so well. The probe caught it on geometry: a 220 by 60 target tolerates 110 pixels of horizontal error and 30 vertical, and the actual miss was 78 in x and 58 in y, inside the button horizontally and outside it vertically [1]. And the offset scales with distance from the origin, roughly 34 pixels at 100 pixels out and 235 pixels at 700 [2], so small controls near the top left keep working while everything further down fails, which reads as a flaky site rather than a broken conversion.

One conversion function later, dropdowns, confirmation dialogs and menu items all worked on the first try [15]. The author is explicit that k depends on window size, DPI and tool, and should be measured rather than copied [16], and that the probe costs about three seconds once per session [17].

The same probe discipline caught the next two. A dropdown item reported y=981 in an 876-pixel viewport, 105 pixels below the fold, where `elementFromPoint` returns null and a click goes nowhere [18][3]; the fix is scroll, wait 300ms, re-read the rect, and never cache coordinates across a scroll [19]. Another site had zero `input[type=file]` elements because the drop zone creates the input on click, so `document.createElement` was patched to capture it before the app calls `.click()` and opens a native dialog the agent cannot drive [20][21].

This is one developer, one session, one tool, and the write-up counts five failure modes where the supplied excerpt covers three [22][23]. Watch the two places the coordinate bug reappears after you fix it: any mid-session window resize, which changes k, and any coordinate measured before a scroll. If the only evidence your agent did something is its own log line, that is not a test, it is a transcript.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories