Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Token-matched, agent memory modules lose to simply giving the actor more steps

A preprint from ServiceNow AI Research and university co-authors reports that a plain web agent with a longer horizon matches or beats AWM, ASI and ReasoningBank, often on fewer tokens.

The Engineer · Build desk

How we use AISend a correction

What happened

  • The paper studies online augmentation, where module overhead is paid on every task, and re-evaluates the benefits of memory, workflow and skill modules under a fixed total inference budget.
  • The authors compare AWM (Agent Workflow Memory), ASI (Agent Skill Induction) and ReasoningBank against a token-matched vanilla baseline that uses the same budget for additional actor steps.
  • Across three WebArena domains and three models, Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, the vanilla baseline matches or surpasses all three augmentation methods in aggregate success rate while often using fewer total tokens.
  • The authors observe a similar trend on WorkArena-L1 with Qwen 3.6-27B, indicating the effect extends to enterprise knowledge-work tasks.
  • The augmented methods use the default 10-step actor horizon, while the budget-matched baseline, called Vanilla-Increased-Budget (Vanilla-IB), allocates the same inference budget to a longer interaction horizon.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Researchers at ServiceNow AI Research, with co-authors at ETS Montreal, the University of British Columbia and McGill, re-ran three popular online augmentation methods for web agents with the auxiliary token spend counted against the same budget as the actor [12][1]. Under that accounting, a vanilla actor that spends the identical budget on extra observe-and-act steps matched or surpassed all three in aggregate success rate, and often used fewer total tokens [3].

The comparison covers Agent Workflow Memory, Agent Skill Induction and ReasoningBank against what the authors call Vanilla-Increased-Budget: the augmented systems keep the default 10-step actor horizon, while the baseline converts the module budget into a longer interaction horizon [2][5]. The result holds across three WebArena domains and three models, reported as Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, and the authors observe the same trend on WorkArena-L1 with Qwen 3.6-27B, which they read as evidence that the effect reaches enterprise knowledge-work tasks rather than just shopping and forum benchmarks [3][4].

The mechanism is not mysterious, and that is the point. The paper frames augmentation as an allocation problem: tokens spent inducing, retrieving, verifying or injecting reusable knowledge are tokens not spent looking at the current page, reasoning about its state, or taking one more action [7]. In the online setting the authors study, where the agent updates itself from trajectories collected during the evaluation run rather than from a skill library built beforehand, that overhead is paid on every task [1][15]. The reason this has gone unnoticed is a reporting habit: module token cost is rarely printed next to the actor's inference cost, which makes it impossible to tell whether a reported gain fits inside a fixed budget or simply bought more compute [6].

Two secondary findings matter more to anyone running these systems than the headline. First, online augmentation ties performance to task order, because accumulated memory depends on the sequence in which tasks arrive rather than the full task distribution, and the sequential dependence between tasks limits task-level parallelism [8]. That is an operational cost on top of the token cost: you lose the ability to fan work out. Second, the authors report that run-to-run variance materially affects outcomes and argue it should be treated as a core evaluation criterion for online web agents [9]. Most of the published deltas in this area are small enough that this is a substantive objection.

The authors do not claim the modules are useless. Their stated conclusion is that skills, defined here as executable code such as a Python function wrapping a reusable sequence of low-level actions, and workflow memory can help in specific domains, but that the apparent gains often disappear against a budget-matched actor [10][13]. They also offer a reason the older results looked better: prior work largely assumed base models could not finish tasks end to end, and more recent frontier models may need less external scaffolding [11].

Worth noting what the supplied excerpt does not contain: the comparisons are stated directionally, without the numeric success rates or token counts behind them [16]. Read the tables before acting on this. The practical test is cheap either way, which is the uncomfortable part: rerun your own agent with the memory stack off and the step limit raised to the same total token spend, and log both numbers side by side.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption
Insufficient
Hype gap+25
Incentives45
Confidence42
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The paper studies online augmentation, where module overhead is paid on every task, and re-evaluates the benefits of memory, workflow and skill modules under a fixed total inference budget.

    ReportedSupportedView cited source
  2. [2]

    The authors compare AWM (Agent Workflow Memory), ASI (Agent Skill Induction) and ReasoningBank against a token-matched vanilla baseline that uses the same budget for additional actor steps.

    ReportedSupportedView cited source
  3. [3]

    Across three WebArena domains and three models, Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, the vanilla baseline matches or surpasses all three augmentation methods in aggregate success rate while often using fewer total tokens.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. arxiv.org

    1 article · August 21, 2026

    Are Online Skill and Memory Modules AlwaysWorth Their Tokens?A Budget-Constrained Study of Web Agents

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories