Build1 distinct publisher3 min readUpdated
Anthropic's Workbench replacement stores nothing, and OpenAI's saved Prompts and Evals platform closes on November 30. Whatever a team kept in a vendor Console needs a repo and an eval runner it owns.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
Stateless is the word doing the work. A Console that saves prompts, keeps version history and lets a team share both is a system of record whether or not anyone signed off on it as one, and Anthropic's replacement keeps none of that on its servers [3][2]. The prompt that shipped and the eval that justified shipping it used to sit in the same browser tab, next to some indication of who touched them last.
Measured from the August 18 swap, Anthropic's export window ran 14 days; OpenAI's runs 104 from that same date [16][17]. The short one is the harder one, because notice of the store closing arrived with the tool that had already replaced it [1][2].
What the exports carry is the part to check before booking migration time. On the Anthropic side, The New Stack reports that a single toggle turned the browser session into Python with the instructions inline, and that the file ran in a terminal without edits [11]. That moves prompt text, which was never the expensive artifact. Version history and eval results are the expensive artifacts, and the features that held them are gone [2], so the repo absorbs the cheap half and something you build or buy has to absorb the rest. The same review skipped testing those removed features on the OpenAI side, on the grounds that OpenAI is retiring them in November anyway [15]. The supplied text also breaks off mid-sentence right after noting that OpenAI had not exported the bot, so there is no recorded result for that half of the test [21].
The performance numbers are worth reading as metering, not as a benchmark. OpenAI's run took roughly 3.3 times as long as Anthropic's [18] and reported about 28 times the tokens [19], but OpenAI's page exposes reasoning and verbosity controls and shows the model's reasoning [22][12], so the two token counts are not counting the same thing. More telling: one side reported a cost and the other reported none at all [10][12]. Anthropic's figure works out to about $0.0000284 per token, near $28 per million billed tokens on that single request [20].
That is a metered scratchpad, and it is not on anyone's subscription plan. Both playgrounds bill from prepaid credits [13], which produced the most useful detail in the whole comparison: the reviewer could not switch off the default GPT-5.4-mini until he topped up his balance [14]. A prompt surface that gates model selection on a credit balance is not somewhere to keep anything you need at 3 a.m.
The New Stack's reading is that both vendors landed on the same position, that prompts belong in code rather than a web console [6]. The vendors did not need to argue it. OpenAI's Playground has existed since June 2020, two and a half years before ChatGPT [7], and is now called Chat [8]; the storage layer around it is what got cut. Read together, the two announcements say the Console is a place to try things, and the durable copy is your problem. The head-to-head answers which scratchpad is faster [9][10][12]. It does not answer which vendor leaves you a clean path out of its store.
Ranked by verification strength, evidence, and original report placement.
The New Stack's conclusion is that both companies reached the same position: prompts belong in your code, not in a web console.
In Anthropic's Playground the bot returned the correct answer on the first try using the default claude-sonnet-5, in 1.9 seconds, using 102 tokens, at a cost of $0.0029.
OpenAI's tool also returned a correct answer on the first try, in 6.2 seconds, with color-coded formatting and a view of the model's reasoning; the test used about 2.900k tokens and no cost was listed.
The supplied article text breaks off mid-sentence immediately after stating that OpenAI had not exported the reviewer's bot, so no outcome is recorded for the OpenAI export test.
On August 18, Anthropic replaced Workbench, the prompt-testing tool in its developer Console, with Playground.
Anthropic removed features including saved prompts, version history, evals and team sharing in the change from Workbench to Playground.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One hands-on account, no vendor documentation
Everything rests on a single trade-press review by one reviewer. The product-change facts (dated replacement, removed features, statelessness, export deadline, November 30 shutdown) are specific and internally consistent, but no vendor announcement, changelog or second outlet is cited. The comparative measurements are one run per tool on two different models, with an ambiguously written token figure and no repetition, and the reviewer explicitly skipped the removed-feature comparison. One of the two planned tests on OpenAI is left unresolved in the supplied text.
Shipped and dated vendor changes, no usage data
The underlying changes are not proposals: Anthropic's Playground is live in the Console with Workbench removed, a hard export deadline was set, and OpenAI has a dated shutdown for saved Prompts and Evals. That is real, mandatory-for-users deployment by two major vendors. What is absent is any measure of uptake or impact — no user counts, no migration data, no statements from affected teams, and no third-party tooling adoption evidence.
Headline verdict outruns a one-run test
The vendor-change facts are reported soberly, but the framing — 'the week-old tool beat the six-year incumbent' and 'round one goes to Anthropic' — is a decisive competitive verdict drawn from one run per tool on two different models, with the second OpenAI test unfinished and the deprecated features never compared. The narrower findings that do hold up (Anthropic's export produced runnable code, OpenAI's exported a transcript; Anthropic surfaced a clear max-token message) are genuine but narrow. Overstatement is moderate rather than severe because the load-bearing product facts are concrete and dated.
Trade-press review incentives, no disclosed vendor tie
The cluster contains no vendor-authored or sponsored material: the reviewer describes paying for his own credits ($18 already on Anthropic, $10 added to OpenAI) and no partnership, embargo or affiliate relationship is disclosed. The observable pull is editorial — a developer-tooling outlet's comparative review benefits from a clean winner, which matches the underdog-beats-incumbent headline framing and the 'round one goes to Anthropic' scoring. No further incentive facts are available in the supplied material, so this is scored on framing alone.
Product facts credible, comparison weak
Confidence is split. The dated product and deprecation facts are specific, mutually consistent and low-ambiguity, and would be easy to falsify if wrong, so they carry moderate-to-good confidence even from one outlet. The comparative performance verdict carries low confidence: n=1 per tool, mismatched models, an oddly formatted token count, unverifiable model names, no vendor corroboration, and a truncated second test. The blended figure reflects reliable 'what changed and by when' with unreliable 'which tool is better'.
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
product
Anthropic's bioweapon filters skipped 133 million contractor chats for eleven months1 distinct publisher
product
Rillet's $100M reads as proof mid-market ERP is rip-and-replace, mostly at the cheap end1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026