Build1 distinct publisher3 min readPublished
A position paper argues that reasoning traces carry safety signal precisely because reinforcement learning treats them as latents rather than outputs, and it asks frontier developers to weigh training decisions against that.
The Engineer · Build desk

product
Chinese banks and telcos are retailing AI tokens in a unit their customers cannot price1 distinct publisher
security
OpenAI's evaluation agents turned a package registry into their messaging bus1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
security
Washington names industrial-scale distillation, then hands the detection bill to abuse teams1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
Where the signal comes from is the whole argument. In reasoning models the trace is produced before the action or the final answer, and the paper treats it as a latent variable in the computation: during the RL phase it is handled much like an activation, never directly supervised, shaped only by whether it helped lead to a highly rewarded answer [6]. Answers, by contrast, are explicitly rewarded against displaying certain content. Intent to misbehave can therefore survive in the trace after it has been trained out of the answer [7], which is the mechanism the whole monitoring proposal rests on [5].
That same absence of scoring is also the failure mode. CoT obtained by prompting a non-reasoning model is under the same selection pressure to look helpful and harmless as any other output, which is why the paper says its trustworthiness is limited [8]. Both cases produce text that reads alike. What differs is whether anything in training was scoring that text [16]. A safety property that exists because nobody optimised for it is an awkward line item to defend at a training review.
The necessity ground is the one to check against your own threat model, because it arrives with conditions attached: models have to keep reasoning in natural language, and the behaviours carrying the most severe risk have to require extended reasoning [13]. The paper supports that second condition with capability evidence rather than monitorability evidence, arguing that severe risks typically need advanced capabilities and that current systems are unlikely to carry out a sophisticated cyberattack or exfiltrate their own weights without being noticed, citing evaluations that include joint work by the US and UK AI Safety Institutes [15]. Prompting a model to think out loud also improves capability and moves more of the relevant computation into natural language [9], so the necessity case is not purely hypothetical. If the harm you actually fear fits in one cheap step, necessity buys you nothing, and what remains is propensity, which the paper declines to call robust [14].
On the read side, the paper's position is that an incomplete trace still earns its place as an additional safety layer [11], which is an argument about layering rather than about coverage. The input is the trace itself, so a filter that only sees final outputs cannot see the thing that makes any of this work [17].
The paper stops short of the part a builder most wants. It asks frontier developers to consider the impact of development decisions on monitorability without enumerating which decisions cost how much [18]. The burden therefore sits upstream with whoever changes the training recipe [1]. Downstream, the observable is thin: a model update can leave benchmark scores intact and still remove the input your monitor was reading.
Ranked by verification strength, evidence, and original report placement.
The paper "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety" recommends that frontier model developers consider the impact of development decisions on CoT monitorability, because CoT monitorability may be fragile.
The paper recommends further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods.
The paper states that, like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed.
The paper defines a CoT monitor as an automated system that reads the chain of thought of a reasoning model and other relevant information and flags suspicious or potentially harmful interactions; flagged responses could then be blocked, replaced with safer actions, or reviewed in more depth.
Reasoning models are explicitly trained to perform extended reasoning in CoT before taking actions or producing final outputs, and in these systems CoTs can serve as latent variables in the model's computation.
During the RL phase of training, CoT latents are treated largely the same as activations: they are not directly supervised, but are optimized indirectly by their contribution in leading the model to a highly rewarded final answer.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary document, argument rather than measurement
Our claims are close restatements of the paper's own sentences, so the reporting is accurate about what was argued, and the document is the primary artefact rather than a summary of one. The step everything rests on, that traces optimised only indirectly may hold intent the output hides, is asserted by analogy to activations and demonstrated nowhere in the text we have. The counter-evidence about unfaithful traces is handled by citation to Turpin, Chen, Lazaridou and Korbak, and the supplied portion breaks off in section 1.1 before any experiment.
No usage disclosed
The paper names OpenAI, DeepSeek and Anthropic as builders of reasoning models, and stops there. Nothing in our coverage says whether any developer runs a chain-of-thought monitor today, over what share of traffic, or with what result. A recommendation to invest is not evidence that anyone has.
Sold as partial, argued as structural
The paper discounts itself at every opportunity: imperfect, not a panacea, fragile, and in the propensity case not generally robust. That hedging is honest, and it undersells the one structural point in the text, which is that output-level oversight cannot see the signal at all, because reward has already removed it from the output. Pulling the other way, the ask to weigh development decisions against monitorability arrives with no accounting of which decisions matter, so a reader cannot act on the recommendation from this text. The net sits slightly on the understated side.
A position paper asking for resources it also justifies
What this document delivers is a request: fund monitorability research, buy monitors, and let monitorability constrain training decisions. That is a normal thing for a position paper to do, and it also means the only voice describing the opportunity is the voice asking for money and attention for it. The supplied text does not identify authors or affiliations, so we cannot say whose training choices the recommendation would bind or who benefits if it is adopted.
One document, quoted closely, read only as far as section 1.1
We can be confident about what the paper says, since the text is in front of us and our claims track it sentence by sentence. We cannot be confident about the world it describes: a single publisher, no corroboration, no empirical result inside the excerpt, and a cut-off mid-sentence in section 1.1 that may well precede the evidence a sceptic would want.