Build1 publisher3 min readPublished
Webwright lifts GPT-5.4 from 33.5% to 60.1% on long web tasks by having it write browser code
Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Webwright's core has three parts: a Runner that tracks the task, a Model Endpoint for OpenAI, Anthropic and OpenRouter, and a terminal wired to Playwright on Chromium.
- On Online-Mind2Web's 300 real-site tasks, GPT-5.4 with Webwright scored 86.7% and Claude Opus 4.7 scored 84.7%, though Opus led on the hard subset, 80.5% to 76.6%.
- Every score is an AutoEval verdict from an LLM judge, not a unit test, and the Mind2Web headline figure used only 100 of the benchmark's 300 tasks.
- Run as a Codex skill, one Microsoft example used about 3.3 million tokens, against 424,000 for the standalone harness.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Because the spend goes into building tools up front, a task that runs only once pays for the tooling and never gets the reuse back.
- decision Whether to run Webwright inside Codex or on its own becomes a token-budget choice for identical work.
- capability If smaller models can drive the saved tools, as the write-up says, a team can pay large-model prices once per workflow and run the repeats on cheaper models.
- constraint Persisted scripts are code a team now owns, so a redesign of a target site turns into a script fix someone has to schedule.
A conventional web agent works inside the browser session. At each step the model gets the current page state and predicts one action, such as a click, a keystroke, a scroll, a DOM selection or a short tool call, inside a loop fixed in advance [4]. When it finishes, it leaves a one-time click sequence with nothing to rerun [5].
Webwright moves the durable state off the page, according to a dev.to write-up by Nokka that draws on Microsoft's blog and the repository README [18]. The model gets a terminal and writes its own browsing program [1]. It opens a browser, inspects it and throws it away as needed while the program develops, and the code and notes in a local workspace are what persist [7]. The loop has four steps:
1. Read the current state [10]. 2. Choose a command [10]. 3. Run it and look at what happened [10]. 4. Repeat until the model judges the task done, then confirm with a self-check step [10].
There is no multi-agent layer, graph engine, plugin system or hidden orchestration [8].
The researchers argue that tight harnesses helped when models were weak and became the bottleneck once models got good at writing and fixing code [6]. I think the design fits long tasks. A program can carry a plan across the 76.1 steps Webwright averaged on Odysseys [3]. A one-action loop has to choose its next move from whatever page is in front of it [4].
The 26.6-point gain [1] needs a qualification before it transfers. The baseline was GPT-5.4 steering by screen coordinates [2]. So the comparison sets written code against pixel clicking, and the code harness posted a success rate about 79% higher [2]. The reported results do not include a step-wise agent that selects DOM elements, so they cannot split the gap between the one-step loop and the coordinate interface.
The grader is the second qualification. For the gain to show up elsewhere, a team's tasks would have to resemble Odysseys' long ones, and its definition of done would have to match the LLM judge's [13].
Webwright spends its compute early. The write-up explains the per-task price as compute invested up front in sturdier tools that later runs reuse [14]. Claude Opus 4.7 costs about 2.6 times as much per task as GPT-5.4 [3]. The host changes the bill as well. As a Codex skill, the example used about 7.8 times the tokens of the standalone harness [4]. The write-up attributes that to context cached in the host's session [15].
The harness is small enough to review in one sitting. Microsoft's blog puts the core at about 1,000 lines: a 150-line Runner, a 550-line Model Endpoint and a 300-line Environment [16]. The README counts by file instead, listing a 450-line agent loop, a 570-line Playwright environment and a 150-line CLI [17]. That totals about 1,170 lines [5].
What to watch
- Odysseys results for a step-wise agent that selects DOM elements, which would show how much of the 26.6 points comes from writing code.
- Scores on all 300 Online-Mind2Web tasks graded by unit tests or human reviewers in place of the AutoEval judge.
- Measured success rates for smaller models rerunning saved Webwright scripts on the same sites weeks later.