Skip to content

Build1 publisher3 min readPublished

Putting the key last cuts GPT-6.1 Sol's measured no-CoT reasoning depth on no-cot-bench by 16%

LessWrong researchers found GPT-6.1 Sol's no-CoT reasoning depth falls 16% once no-cot-bench gives the prompt's starting key last. Astra's standout score may carry the same inflation, though only Sol has been retested.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Putting the key last cuts GPT-6.1 Sol's measured no-CoT reasoning depth on no-cot-bench by 16%
Photo: lesswrong.com

What happened

  • In no-cot-bench the key, the starting state the operations act on, usually comes at the top of the prompt, ahead of the steps.
  • On Neel Nanda's benchmark, Astra, suspected to be a looped transformer, has about 8.6 times Fable 5.1's odds of solving an arbitrary problem.
  • The authors paired six of the benchmark's serial state-tracking tasks with key-last versions and added five new synthetic domains.
  • A control appending the same unrelated question to both layouts found accuracy on it mostly identical across key-first and key-last versions.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Any discount applied to Astra's score today is borrowed from Sol on the strength of a suspected architectural match, so it stays an estimate until Astra gets its own key-last run.
  • decision Citing no-cot-bench as evidence of opaque serial reasoning now calls for reporting key-first and key-last scores separately, because the two capacities the benchmark blends scale differently.
  • precedent Other serial-reasoning evals that state the initial state first are open to the same objection, and key-last items with state-dependent branching over large state spaces are now a published way to test for it.

On the state-machine task, the prompt opens with "Start with the number 12" and then lists six conditional updates to apply in order [5]. A model reading that prompt has the starting state before it sees the first rule. According to the post, it can compute a persistent hidden state while it reads and use attention to carry that state forward across token positions [6]. Each rule then arrives with its own run of tokens in which to do one step. The authors wrote that the benchmark might therefore measure "how good is this model at spreading per-step work across token positions" in addition to serial depth [7].

Both capacities count as opaque reasoning. "These both measure 'opaque reasoning' but scale differently and have different implications," the authors wrote [3]. A single score that blends them cannot say which one a new model improved.

The obvious fix is to move the key to the end. robo proposed exactly that for Astra, using nested modular arithmetic [8]. The authors found it is not enough on its own. With a small key space, the model can precompute the outcome for every possible ending key, and operations that compose compactly can be folded together before the key shows up [9]. So the key-last items branch on the current state to block simple symbolic composition, and they draw from state spaces too large to enumerate [10].

The control is the part of the post I would reuse. A key-last prompt might simply be harder to parse or further out of distribution, and that alone would lower scores [12]. Each pair therefore gets the same unrelated question appended, such as the digit sum of 5543919735379573, a question nothing in the original problem helps answer [12]. If the key-last layout were generically harder to process, accuracy on that tail question should drop after the key-last version [12].

The 16% was measured on GPT-6.1 Sol [2]. The authors picked Sol because it is much cheaper than Astra and because recent results suggest the two are architecturally similar [14]. They lacked the compute to rerun the key-last analysis on other models [14]. For the correction to carry over to Astra's 8.6x odds lead over Fable 5.1 [4], Astra would have to split its score between serial depth and token-spreading in roughly the proportion Sol does. The post does not measure that split for Astra.

The authors also expect the effect to grow as dependent depth increases [2]. They frame the test as a choice between a shift, where key-last costs the same at every depth, and a change in slope [15]. A steeper slope would mean the 16% understates the correction on longer chains. In my view, teams treating no-cot-bench as evidence of opaque serial reasoning in a suspected looped transformer should read key-first scores as an upper bound on serial depth until Astra gets its own key-last run.

What to watch

  • A key-last rerun of Astra itself, which would replace the discount inferred from Sol with a measured one.
  • The shift-versus-slope result: whether the key-last penalty grows with dependent depth, as the authors expect.
  • Whether no-cot-bench takes up the paired key-first and key-last items as part of its standard suite.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories