Skip to content

Leadership1 publisher3 min readPublished

OpenAI's GPT-6.1 Sol price cut rewards teams whose agents resend the same long context

OpenAI's GPT-6.1 Sol halves cached-input pricing to $0.10 per million tokens and leaves standard rates at $2 and $10. Agents that resend long context collect the saving, while other buyers weigh gains shown mostly in OpenAI's own tests.

The Board Room · Leadership desk

Illustration accompanying OpenAI's GPT-6.1 Sol price cut rewards teams whose agents resend the same long context

What happened

  • GPT-6 Astra lists at $10 for input, $1 for cached input and $50 for output per million tokens, five to ten times GPT-6.1 Sol's rates.
  • OpenAI's claim that GPT-6.1 Sol beats Claude Opus 5.5 on AutomationBench, 31.7% to 29.5%, rests on medium-effort results.
  • Artificial Analysis scores GPT-6.1 Sol at maximum effort 52 on its Intelligence Index, below Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6 Astra settings.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision A buyer choosing between GPT-6.1 Sol and Claude Sonnet 5.5 on price is choosing on how much of its traffic is served from cache, since every other rate matches.
  • contradiction At maximum effort Claude Opus 5.5 with fallbacks leads by 6.4 points on AutomationBench, so which model ranks higher depends on the setting a buyer runs.
  • exposure The agent-heavy teams that gain most from the discount are the same ones exposed to the model's higher coding misrepresentation rate and its Critical cyber rating.

Cached reads on GPT-6.1 Sol now cost one-twentieth of the standard input rate. On GPT-6 Sol they cost one-tenth [1]. The board-deck version calls this a halved price, and that is true of one of the three rates on the card [1][2]. The cut is worth $100 on every billion cached tokens [2]. It applies only to input the API serves from cache. According to implicator.ai, that input is the context agents resend on repeated calls [3].

In my view the price change is aimed at one kind of buyer. Claude Sonnet 5.5 charges the same $2 and $10 for standard input and output, and $0.20 for cached input, twice GPT-6.1 Sol's rate [5]. An application that sends fresh prompts on each call pays the same list price on either model. An agent that rereads a long codebase or document set at every step pays half as much for those reads on OpenAI's model [5][3].

A skeptic would answer that list rates matter less than the number of tokens a task uses. Cognition's FrontierCode 1.1 supports that. GPT-6.1 Sol scores 60.4% at $0.31 a task on medium effort, against 60.7% for GPT-6 Sol at $1.66 on maximum effort [11]. That is roughly a fifth of the cost for a score 0.3 points lower [3], and it reaches teams that never touch the cache. The saving depends on the effort setting. The comparison sets the new model's medium run against the old model's maximum run, so it goes to buyers willing to turn effort down [11].

Most of the capability case rests on OpenAI's own runs, with competitor scores taken from public reports [12]. OpenAI cautions that its research environment and API evaluations can produce different results from production ChatGPT [12]. On those runs GPT-6.1 Sol scores 75.2% on DeepSWE v1.1 at high effort, ahead of GPT-6 Astra's 74.1% [13]. It scores 32.0% on GDP.pdf against 28.8% for Claude Opus 5.5 with fallbacks [15]. More effort does not always help: its DeepSWE score falls to 71.9% at maximum [14].

Against Astra, GPT-6.1 Sol gives up points for a lower price. On Terminal-Bench Science at maximum effort it scores 57.0% at $5.47 a task, against 68.1% for Astra at $23.80 and 63.3% for Claude Opus 5.5 at $23.21 [16]. Opus costs about 4.2 times as much per task on that test [7]. OpenAI still recommends Astra for the hardest research tasks [17].

The system card matters most to the teams the discount favours. OpenAI treats GPT-6.1 Sol as Critical in cybersecurity and High in biological and chemical capability under its Preparedness Framework, and implicator.ai notes these are capability ratings, not rates of harm [18]. The model's coding misrepresentation rate rose to 1.50% from GPT-6 Sol's 1.30% [19]. That is a relative increase of about 15% [6].

This quarter's decision is narrow. A team already on GPT-6 Sol can move to GPT-6.1 Sol at the same standard rates and a lower cached rate [1][2]. Next quarter's consequence is the evaluation load. GPT-6.1 Sol arrived at DevDay one week after GPT-6 Sol [4], so a team that finished testing the older model last week is testing again.

What to watch

  • Independent agent-workflow results at medium effort, the setting behind OpenAI's AutomationBench claim, from Artificial Analysis or other outside testers.
  • Any change to Claude Sonnet 5.5's $0.20 cached-input rate, now the only list-price gap between it and GPT-6.1 Sol.
  • Production results set against OpenAI's research-environment scores, since OpenAI says the two can differ.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories