Build1 publisherNot yet confirmed elsewhere3 min readPublished
The Console is a scratchpad now: Anthropic gave 14 days to export, OpenAI gives until November 30
Anthropic's Workbench replacement stores nothing, and OpenAI's saved Prompts and Evals platform closes on November 30. Whatever a team kept in a vendor Console needs a repo and an eval runner it owns.
The Engineer · Build desk
What happened
- Anthropic swapped Workbench for Playground on August 18 and dropped saved prompts, version history, evals and team sharing.
- The replacement keeps nothing on Anthropic's servers, and anything stored in Workbench had to be exported by September 1.
- In the same week, OpenAI said its saved Prompts and Evals platform will shut down on November 30.
- In a New Stack side-by-side test of the same PR review bot, Claude answered correctly in 1.9 seconds; OpenAI's Chat took 6.2 seconds and showed no cost.
- Neither playground is covered by a subscription plan; both run on prepaid credits.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint After November 30 neither vendor holds prompt versions or eval results, so any record that is not in a repository has nowhere left to sit.
- cost The migration bill falls on whoever quietly used the Console as storage, and Anthropic allowed two weeks to pay it.
- decision Every team that scored prompts in a Console now has to pick an eval runner it owns and maintains, rather than one that ships with the vendor's browser tab.
- capability Anthropic's runnable Python export makes lifting prompt text into code close to free, which reframes the hard part of the move as history and scores, not prompts.
Stateless is the word doing the work. A Console that saves prompts, keeps version history and lets a team share both is a system of record whether or not anyone signed off on it as one, and Anthropic's replacement keeps none of that on its servers [7][6]. The prompt that shipped and the eval that justified shipping it used to sit in the same browser tab, next to some indication of who touched them last.
Measured from the August 18 swap, Anthropic's export window ran 14 days; OpenAI's runs 104 from that same date [21][22]. The short one is the harder one, because notice of the store closing arrived with the tool that had already replaced it [5][6].
What the exports carry is the part to check before booking migration time. On the Anthropic side, The New Stack reports that a single toggle turned the browser session into Python with the instructions inline, and that the file ran in a terminal without edits [13]. That moves prompt text, which was never the expensive artifact. Version history and eval results are the expensive artifacts, and the features that held them are gone [6], so the repo absorbs the cheap half and something you build or buy has to absorb the rest. The same review skipped testing those removed features on the OpenAI side, on the grounds that OpenAI is retiring them in November anyway [16]. The supplied text also breaks off mid-sentence right after noting that OpenAI had not exported the bot, so there is no recorded result for that half of the test [4].
The performance numbers are worth reading as metering, not as a benchmark. OpenAI's run took roughly 3.3 times as long as Anthropic's [18] and reported about 28 times the tokens [19], but OpenAI's page exposes reasoning and verbosity controls and shows the model's reasoning [17][3], so the two token counts are not counting the same thing. More telling: one side reported a cost and the other reported none at all [2][3]. Anthropic's figure works out to about $0.0000284 per token, near $28 per million billed tokens on that single request [20].
That is a metered scratchpad, and it is not on anyone's subscription plan. Both playgrounds bill from prepaid credits [14], which produced the most useful detail in the whole comparison: the reviewer could not switch off the default GPT-5.4-mini until he topped up his balance [15]. A prompt surface that gates model selection on a credit balance is not somewhere to keep anything you need at 3 a.m.
The New Stack's reading is that both vendors landed on the same position, that prompts belong in code rather than a web console [1]. The vendors did not need to argue it. OpenAI's Playground has existed since June 2020, two and a half years before ChatGPT [10], and is now called Chat [11]; the storage layer around it is what got cut. Read together, the two announcements say the Console is a place to try things, and the durable copy is your problem. The head-to-head answers which scratchpad is faster [12][2][3]. It does not answer which vendor leaves you a clean path out of its store.
What to watch
- Whether Anthropic ships any replacement for evals, version history or shared prompts, or leaves Playground as a pure scratchpad.
- What OpenAI's November 30 export actually hands back, and whether eval run history comes with the saved prompts.
- Whether either vendor exposes eval scoring through a CLI or API that a build pipeline can call once the Console no longer hosts it.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption52
- Hype gap+30
- Incentives42
- Confidence44
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The New Stack's conclusion is that both companies reached the same position: prompts belong in your code, not in a web console.
ReportedSupportedSource: The New Stack2 sources— create a free account to open themView cited source - [2]
In Anthropic's Playground the bot returned the correct answer on the first try using the default claude-sonnet-5, in 1.9 seconds, using 102 tokens, at a cost of $0.0029.
- [3]
OpenAI's tool also returned a correct answer on the first try, in 6.2 seconds, with color-coded formatting and a view of the model's reasoning; the test used about 2.900k tokens and no cost was listed.
- [4]
The supplied article text breaks off mid-sentence immediately after stating that OpenAI had not exported the reviewer's bot, so no outcome is recorded for the OpenAI export test.
- [5]
On August 18, Anthropic replaced Workbench, the prompt-testing tool in its developer Console, with Playground.
- [6]
Anthropic removed features including saved prompts, version history, evals and team sharing in the change from Workbench to Playground.
- [7]
The new Anthropic Playground is stateless: it remembers nothing and stores nothing on Anthropic's servers.
- [8]
Users with data stored in Workbench had until September 1 to export it.
- [9]
In the same week Anthropic launched Playground, OpenAI announced it would shut down its saved Prompts and Evals platform on November 30.
- [10]
OpenAI's Playground launched in June 2020, in the GPT-3 era, two and a half years before ChatGPT existed.
- [11]
OpenAI's Playground is now called Chat but serves the same prompt-testing purpose.
- [12]
The reviewer built the same PR review bot in both tools, feeding each the same instructions and code changes and requiring a rigid machine-readable report on risk, files touched, a one-line summary and whether tests changed.
- [13]
One toggle in Anthropic's Playground turned the session into Python code with the instructions already included; the reviewer copied it, ran it in his terminal and needed no edits to make it work.
- [14]
Neither Anthropic's nor OpenAI's playground falls under the vendors' subscription plans; both require added credits. The reviewer had a little over $18 in his Anthropic account and added $10 to OpenAI.
- [15]
The reviewer did not have enough credits to switch the model away from GPT-5.4-mini; each attempt to change model prompted him to add credits, so he added $10.
- [16]
The reviewer did not test the OpenAI features Anthropic cut, such as saved prompts, version history and evals, because OpenAI is retiring them in November.
- [17]
OpenAI's page carries more controls than Anthropic's, including settings for reasoning and verbosity and a button that writes the prompt for you.
- [18]
The OpenAI run took about 3.3 times as long as the Anthropic run.
- [19]
The OpenAI run reported about 28 times as many tokens as the Anthropic run.
- [20]
Anthropic's reported charge works out to about $0.0000284 per token, or roughly $28 per million billed tokens on that request.
- [21]
Anthropic's export window ran 14 days, from the August 18 launch to the September 1 deadline.
- [22]
OpenAI's November 30 shutdown falls 104 days after August 18.
Sources
1 independent publisher whose own reporting we read for this story.
- thenewstack.ioAnthropic’s Playground vs. OpenAI’s: The week-old tool beat the six-year incumbent
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- AnthropicFollow
- OpenAIFollow
- Anthropic PlaygroundFollow
- Anthropic WorkbenchFollow
- OpenAI Playground (Chat)Follow
- OpenAI saved Prompts and Evals platformFollow
- Claude Sonnet 5Follow
- GPT-5.4-miniFollow
- The New StackFollow