Skip to content

Build1 publisher2 min readPublished

GPT-6.1 Sol closes 80% of the no-reasoning gap between GPT-6 Sol and Astra on 27 tasks

GPT-6.1 Sol closes 80% of the gap between GPT-6 Sol and GPT-6 Astra on 27 no-reasoning tasks, a test posted on LessWrong finds. For teams calling Astra with reasoning suppressed, one author's run is grounds to test Sol on their own tasks before deciding anything.

The Engineer · Build desk

Illustration accompanying GPT-6.1 Sol closes 80% of the no-reasoning gap between GPT-6 Sol and Astra on 27 tasks

What happened

  • Astra scored at least as well as GPT-6.1 Sol on 25 of the 27 tasks, and on the two where Sol came out ahead, the margin was one percentage point.
  • The 27 tasks drop nine where GPT-5.5 and Astra finished within 2 points of each other, including arithmetic and intuit_physical, plus one task that appeared to confuse the models.
  • Scores for Astra and GPT-5.5 were copied from an earlier post, Estimating GPT-6 Astra's no-CoT Time Horizon; the author did not re-run either model.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision The post does not cover price, so the case for moving no-reasoning calls from Astra to GPT-6.1 Sol depends on what each model costs under a team's own contract.
  • constraint Teams copying the setup inherit a prompt-level switch for reasoning on GPT-6.1 Sol, and the author's perfect compliance covers the 27 test tasks only until it is checked on their own traffic.
  • exposure Moving work from GPT-6 Sol to 6.1 Sol would weaken chain-of-thought monitoring if the author's scores hold, since 6.1 Sol's controllability sits near Astra's and it beats GPT-6 Sol at evading monitors.

The 80% is a ratio. It puts GPT-6 Sol at zero and Astra at 100, then reports where 6.1 Sol lands on mean accuracy across tasks from the Think Fast suite. The 95% interval is 72 to 87% [1][2]. To turn that into accuracy points you need the spread between the two reference models on your own workload. The per-task count is steadier than the mean. On 24 of the 27 tasks, 6.1 Sol sits nearer Astra than its own predecessor [3].

The time-horizon figures are much softer. GPT-6.1 Sol's estimate is 35 minutes, against 4.0 minutes for GPT-6 Sol [10]. That is a factor of 8.75 [1]. The 95% interval on the 6.1 Sol figure runs from 9.5 minutes to 23 hours [10], a spread of about 145 times [2]. The author calls the estimates highly uncertain because 6.1 Sol saturates many of the benchmarks, and relies on per-benchmark accuracy instead [9].

The call pattern matters as much as the task mix. GPT-6.1 Sol does not accept reasoning_effort=none. The author set reasoning_effort=low, added an immediate-recall system prompt following the method of an earlier post, and took one sample per question [5]. GPT-6 Sol ran at reasoning_effort=none with the other choices unchanged [6]. For the 80% to transfer, a production call has to look like the test call. That means low effort, a prompt that makes the model answer without reasoning first, a single attempt, and tasks that resemble the suite.

The explanation on offer is architectural. Twitter users spotted a registry path for gpt-6-astra-minor in the public playground configuration of Microsoft Azure. Others then guessed that Astra Minor had shipped under the 6.1 Sol name [11]. The first exhibit in an architecture argument is a line in someone else's cloud config. The author adds that 6.1 Sol came out seven days after GPT-6 Sol. On that timing, both were probably distilled from Astra, so distillation does not explain the gap [12]. "Combining these facts with 6.1 Sol's time horizons, the looped transformer hypothesis seems highly likely to me," the author wrote [13].

Take a shop that calls Astra with reasoning suppressed. There, I think the right move is to run 6.1 Sol through the existing eval at reasoning_effort=low with the recall prompt, and to decide on those numbers. If the looped hypothesis holds, the author expects OpenAI to keep deploying looped transformers "both at the frontier and below it" [16].

What to watch

  • A re-run of GPT-6 Astra in the same harness as the Sol models, replacing the scores borrowed from the earlier post.
  • Any OpenAI statement on whether GPT-6.1 Sol is Astra Minor or uses a looped architecture.
  • Support for reasoning_effort=none on GPT-6.1 Sol, letting teams suppress reasoning without a system prompt.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories