Skip to content

Build1 publisher3 min readPublished

Changing only the model string to GPT-6.1 Sol leaves two broken fields in a coding agent's request

OpenAI's GPT-6.1 Sol halves old Sol's cache-read rate to $0.10 per million tokens and requires the Responses API for tool calls. In one worked example the cut saves about 10%, and older agents only get that saving after their tool and reasoning fields are rewritten.

The Engineer · Build desk

Illustration accompanying Changing only the model string to GPT-6.1 Sol leaves two broken fields in a coding agent's request

What happened

  • Neither GPT-6.1 Sol nor GPT-6 Astra supports 'none' reasoning, so an agent that switches reasoning off cannot carry that setting across.
  • GPT-6.1 Sol keeps GPT-6 Sol's Standard input, write and output prices, so a workload with no cache reads gets the same bill.
  • The author frames the post as a selection framework with prices checked on September 30, 2026, and says no three-model benchmark was run.
  • The post warns against reading API spend off a ChatGPT or Codex subscription's percentage meter.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost An agent with little cached context gets no price benefit from upgrading, because the cut is worth only what cache reads already cost that agent.
  • decision Teams whose agents disable reasoning must either stay on GPT-6 Sol or accept a reasoning effort whose token use the flat per-token rates cannot predict.
  • exposure Teams calling Sol through another gateway cannot assume OpenAI's direct rates or its Responses tool support carry over, so they have to rerun the comparison against that service.

The old request in the post has two settings the new model will not accept. It sets `reasoning_effort: "none"`, and its tool definition nests `name`, `description` and `parameters` inside a `function` object [13]. According to the post, changing only `model` to `gpt-6.1-sol` leaves both problems in place: disabled reasoning is unsupported, and tool calling belongs on Responses [14]. The migrated version sets `reasoning: {"effort": "low"}`, lifts `name` to the top level of the tool, adds `strict: true` with `"additionalProperties": false`, and goes to `/v1/responses` [15].

The field edits are the easy part. The loop around them changes too. The application has to inspect the typed output, execute an authorized function, and send back `function_call_output` with the matching `call_id` [16]. The post puts it in a sentence more agent READMEs could use: "A schema named run_tests cannot run your tests by itself." [17] Its check before switching is a full round trip. That means the first tool request, real test execution, the returned result and the final answer, with reasoning and tool items kept in history [18]. I would not ship on less.

On price, only one Standard rate moved. Cached input went from $0.20 to $0.10 per million tokens [4]. In the post's fixture, cache reads are 900,000 of 1,050,000 tokens, about 86% [1]. They cost $0.18 at the old rate and $0.09 at the new one [2]. Input, write and output rates are unchanged [3], so that $0.09 is the whole difference. A 10.2% saving [9] then implies an old-Sol bill of about $0.88 [3], and cache reads were about a fifth of it [4].

For the 10.2% to carry over to another workload, several things have to hold. Every request has to stay under the long-context threshold, with no cache writes and no paid tools [8]. Cache reads have to take a similar share of spend. Both models also have to use the same amount of reasoning and need the same number of attempts. The post says neither of its percentages predicts task-level savings when they do not [11]. Longer answers break the equal-rates assumption the same way [12]. Its instruction is to replace the fixture with returned usage and count every attempt, failed patches and retries included [22].

According to the post, OpenAI positions the new Sol close to Astra for complex coding, computer use and professional work [25]. The author adds that this does not establish a universal winner for writing, repository review or difficult debugging [26]. The recommendation is 6.1 Sol as the next default for a Responses-based coding agent, and Astra where it produces enough additional correct work to justify its premium [19]. A verified escalation path sits between them [20]. For a cache-heavy agent already on Responses, I think that ordering is right.

What to watch

  • Returned usage from real agent runs on both Sols, counting retries and failed patches, would show whether the 10.2% fixture saving survives a move from 'none' to low reasoning effort.
  • Whether third-party gateways expose GPT-6.1 Sol's Responses tool calling and cache-read pricing at parity with OpenAI's direct Standard rates.
  • A published task-level comparison of 6.1 Sol and Astra on repository review or difficult debugging, the areas where the post says OpenAI's positioning does not settle the choice.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories