Skip to content

Invest1 publisher3 min readPublished

Microsoft's log-reading CASD method cuts agent prompt tuning to about $1.60 a pass

Microsoft researchers say CASD writes better agent prompts from saved logs for about $1.60 a pass, over 22 times cheaper than iterative tuning. Outside one telecom test its accuracy edge over GEPA is thin, so price and setup carry the case.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Microsoft's log-reading CASD method cuts agent prompt tuning to about $1.60 a pass
Generated illustration

What happened

  • A coding agent reads the full corpus of saved agent trajectories in one pass, extracts recurring failure modes and representative episodes, and writes behavioural rules into a new system prompt.
  • Across the tested tasks, the optimized prompts averaged a gain of 16.6 percentage points over the unoptimized baseline, according to the paper.
  • The paper says CASD kept its lead even when rival methods were given extra validation data and unrestricted environment access.
  • GEPA, the main comparison, is described in the paper as widely adopted at Databricks, Shopify, Dropbox and OpenAI, among others.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost For agent teams, a larger saving than the compute bill is staff time: CASD drops the held-out validation set that iterative tuners need someone to build and keep current.
  • exposure Tuning without live environment access stops an optimizer from testing draft prompts on systems where the agent can act on customer accounts or production tools.
  • constraint Output quality is capped by the saved logs, so a team whose trajectories are narrow or unrepresentative can end up with rules that do not generalize.

On the paper's own ratio, a validation-gated iterative pass costs more than $35.20, because CASD's roughly $1.60 is put at over 22 times cheaper [1][2][1]. The direct saving is at least about $33.60 a pass [2]. For a single tuning job, $33.60 is a small sum. It becomes a budget line only for a team re-tuning many agents often, and the source does not say how often teams re-run optimization or how many passes an iterative run takes.

The $1.60 is also a marginal price. CASD works from saved trajectories and needs no iterative search or live interaction with the environment [3]. Those past episodes were paid for when the agent first ran them. A team that has not kept its agent logs has nothing to feed the method.

Against GEPA, the margins are uneven. CASD led by 9.3 points on ALFWorld, 0.8 on tau2-bench retail and 21.7 on tau2-bench telecom, and trailed by 9.4 on SpreadsheetBench-Verified, where GEPA scored 60.7 to CASD's 51.3 [3][11]. Averaged across the four, CASD's lead is 5.6 points [4]. Take telecom out and the average edge across the other three is about 0.2 points [5]. Telecom supplies 21.7 of the 22.4 net points CASD gained across the four tests [6].

The write-up puts the gains down to scope. Iterative optimizers typically look at slices of data at a time, and patterns that only appear across many episodes are harder to spot that way [19].

Microsoft is also arguing against its own earlier work. SkillOpt, a validation-gated iterative method that treats agent skills as trainable text parameters, came out of Microsoft in June 2026 [15]. The new paper's six authors, Sumit Gulwani among them, report CASD beating SkillOpt-style search across its metrics [14][6]. Cryptobriefing reads the result as Microsoft saying that, in many cases, its earlier approach is more work than necessary [20].

Three outcomes look plausible. Results that hold on production logs would leave iterative loops a role mainly on structured work like spreadsheets. A corpus of poor logs would move the effort from building validation sets to curating trajectories. An outside replication that narrows the telecom gap [7] would leave the GEPA comparison close to a tie on accuracy, with CASD ahead on price and setup.

I think the third is the likeliest. It still favours CASD for teams that already store large volumes of agent logs, because a 0.2-point average edge at a 22nd of the cost is a good deal [5][2]. The counter-case is the spreadsheet benchmark, where GEPA's 9.4-point lead shows an iterative loop still paying on structured tasks [11]. The view is wrong if an outside replication on production logs finds CASD trailing GEPA across most task types, not only spreadsheets.

What to watch

  • An independent replication of the CASD versus GEPA results, especially the 21.7-point telecom gap, on agent logs outside the paper's benchmarks.
  • Whether GEPA users named in the paper, such as Databricks or Shopify, publish comparisons run on their own production trajectories.
  • Cost figures for CASD measured on real trajectory corpora, including how large a log store the method needs to beat iterative tuning.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories