Skip to content

Build1 publisher3 min readPublished

Webwright lifts GPT-5.4 from 33.5% to 60.1% on long web tasks by having it write browser code

Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Webwright's core has three parts: a Runner that tracks the task, a Model Endpoint for OpenAI, Anthropic and OpenRouter, and a terminal wired to Playwright on Chromium.
  • On Online-Mind2Web's 300 real-site tasks, GPT-5.4 with Webwright scored 86.7% and Claude Opus 4.7 scored 84.7%, though Opus led on the hard subset, 80.5% to 76.6%.
  • Every score is an AutoEval verdict from an LLM judge, not a unit test, and the Mind2Web headline figure used only 100 of the benchmark's 300 tasks.
  • Run as a Codex skill, one Microsoft example used about 3.3 million tokens, against 424,000 for the standalone harness.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Because the spend goes into building tools up front, a task that runs only once pays for the tooling and never gets the reuse back.
  • decision Whether to run Webwright inside Codex or on its own becomes a token-budget choice for identical work.
  • capability If smaller models can drive the saved tools, as the write-up says, a team can pay large-model prices once per workflow and run the repeats on cheaper models.
  • constraint Persisted scripts are code a team now owns, so a redesign of a target site turns into a script fix someone has to schedule.

A conventional web agent works inside the browser session. At each step the model gets the current page state and predicts one action, such as a click, a keystroke, a scroll, a DOM selection or a short tool call, inside a loop fixed in advance [4]. When it finishes, it leaves a one-time click sequence with nothing to rerun [5].

Webwright moves the durable state off the page, according to a dev.to write-up by Nokka that draws on Microsoft's blog and the repository README [18]. The model gets a terminal and writes its own browsing program [1]. It opens a browser, inspects it and throws it away as needed while the program develops, and the code and notes in a local workspace are what persist [7]. The loop has four steps:

1. Read the current state [10]. 2. Choose a command [10]. 3. Run it and look at what happened [10]. 4. Repeat until the model judges the task done, then confirm with a self-check step [10].

There is no multi-agent layer, graph engine, plugin system or hidden orchestration [8].

The researchers argue that tight harnesses helped when models were weak and became the bottleneck once models got good at writing and fixing code [6]. I think the design fits long tasks. A program can carry a plan across the 76.1 steps Webwright averaged on Odysseys [3]. A one-action loop has to choose its next move from whatever page is in front of it [4].

The 26.6-point gain [1] needs a qualification before it transfers. The baseline was GPT-5.4 steering by screen coordinates [2]. So the comparison sets written code against pixel clicking, and the code harness posted a success rate about 79% higher [2]. The reported results do not include a step-wise agent that selects DOM elements, so they cannot split the gap between the one-step loop and the coordinate interface.

The grader is the second qualification. For the gain to show up elsewhere, a team's tasks would have to resemble Odysseys' long ones, and its definition of done would have to match the LLM judge's [13].

Webwright spends its compute early. The write-up explains the per-task price as compute invested up front in sturdier tools that later runs reuse [14]. Claude Opus 4.7 costs about 2.6 times as much per task as GPT-5.4 [3]. The host changes the bill as well. As a Codex skill, the example used about 7.8 times the tokens of the standalone harness [4]. The write-up attributes that to context cached in the host's session [15].

The harness is small enough to review in one sitting. Microsoft's blog puts the core at about 1,000 lines: a 150-line Runner, a 550-line Model Endpoint and a 300-line Environment [16]. The README counts by file instead, listing a 450-line agent loop, a 570-line Playwright environment and a 150-line CLI [17]. That totals about 1,170 lines [5].

What to watch

  • Odysseys results for a step-wise agent that selects DOM elements, which would show how much of the 26.6 points comes from writing code.
  • Scores on all 300 Online-Mind2Web tasks graded by unit tests or human reviewers in place of the AutoEval judge.
  • Measured success rates for smaller models rerunning saved Webwright scripts on the same sites weeks later.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories