Skip to content

Build1 publisher3 min readPublished

DeepMind's interpretability lead found Anthropic's J-space again in an open 27B model

Anthropic screened Claude for one property, verbalizability, and found two more in the same representations. Neel Nanda reproduced the structure in open weights, and the outside commentators Anthropic invited disagree about what it is.

The Engineer · Build desk

Illustration accompanying DeepMind's interpretability lead found Anthropic's J-space again in an open 27B model

What happened

  • Anthropic researchers report a representational structure in Claude and other language models that they call the J-space, with features functionally analogous to a global workspace, which one leading theory ties to conscious access.
  • The authors say they searched only for verbalizability, then checked the same representations for direct manipulation by the model and flexible generalization, and were surprised to find both.
  • The researchers and other commentators stress that none of this shows Claude has subjective experiences, and that the claim is about access consciousness: information the model can report, control and reason with.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Because the structure shows up in downloadable weights, a team outside Anthropic has somewhere to look when a reasoning run behaves strangely, and Nanda has already named that use.
  • contradiction Mowshowitz calls the same finding evidence toward consciousness; Chalmers says it shows only limited workspace features. So anyone writing 'Anthropic found a global workspace' into a safety case is citing a disputed interpretation.
  • constraint The co-developers of the human workspace theory list a body, lasting episodic memory and recurrent activity as absent in Claude, so experimental paradigms built for brains need a fresh justification before they are pointed at a model.
  • precedent Eleos AI Research treats the result as showing moral-status questions can be studied empirically. That moves the work into probe-building, staffed by the same interpretability teams that already build safety probes.

Verbalizability was the only property the authors screened for, meaning representations the model can report on [3]. They then tested those same representations for two further properties: whether the model could manipulate them directly, and whether they generalized flexibly. Both held, and the authors report being surprised that the first property predicted the others [3].

That design is defensible. It also fixes the population under study, since whatever the verbalizability probe selects is what gets tested for everything else. Three properties out of one screen is a good result, and it is the kind of result that makes a reviewer ask to see the probe.

The term being applied is access consciousness, defined in the newsletter as information available to report, deliberately control, and flexibly reason with [2]. Those are the same three items as the tested features [13]. So the label does not add independent evidence for the workspace reading; it restates the measurement in philosophical vocabulary.

The reproduction is the part an outside team can act on. Neel Nanda, who leads Google DeepMind's mechanistic interpretability team, independently found a similar internal space in the open Qwen3.6-27B model, storing intermediate information during reasoning [4]. Anthropic's result is about Claude [1]. Nanda's is about weights a team can load itself. He described access to that space as potentially useful for investigating unusual behaviour and for generating new hypotheses [5]. So far that is one reproduction, on one open model, at one size. For any of it to be worth anything on a model you operate, you need a verbalizability probe you trust and read access to activations along the reasoning path.

Anthropic invited outside experts to comment [11], and the responses split. Stanislas Dehaene and Lionel Naccache, who helped develop Global Neuronal Workspace Theory, see important similarities to the workspace proposed in human brains, and stress that Claude has no body, no lasting episodic memory and none of the recurrent neural activity found in brains [6]. Researchers at Eleos AI Research grant that the authors found privileged representations used in reasoning and report, and question whether those form a single unified workspace [7]. David Chalmers argued that the J-space shows only limited evidence of several features associated with a classic global workspace [8]. Zvi Mowshowitz calls the paper a major advance in understanding how language models work and says that, although it does not prove models are conscious, finding a global-workspace-like structure predicted by some theories should count as evidence in that direction [9]. Of the five parties named, three attach qualifications, one calls the finding evidence, and one reproduced it [10].

Every party in that summary who touches the consciousness question either qualifies the workspace reading or stops short of the claim, and the researchers themselves emphasise that the work does not show Claude has subjective experiences [2]. The strongest reading on offer is Mowshowitz's evidential one [9]. Anyone who wants to look up the Anthropic paper for themselves will not find its title or date in the newsletter [12].

What to watch

  • Whether the verbalizability probe is published in enough detail to rerun on a model outside the Claude and Qwen families.
  • Reproduction attempts at other parameter counts: a failure well below 27B would bound how general the structure is.
  • Whether Eleos AI Research's objection about unity gets an experiment, distinguishing one workspace from several privileged representations.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories