Skip to content

Build2 publishers3 min readPublished Updated

Anthropic flew dozens of theologians in under NDA to discuss whether Claude might be conscious

Anthropic has flown in dozens of theologians and philosophers under NDA since fall 2025 to discuss whether Claude might be conscious. The welfare research behind those meetings already lets Opus 4 and 4.1 end conversations with persistently abusive users.

The Engineer · Build desk

Illustration accompanying Anthropic flew dozens of theologians in under NDA to discuss whether Claude might be conscious
Generated illustration

What happened

  • Co-founder Christopher Olah leads the effort inside an official research program that Anthropic discusses in a blog post about Model Welfare.
  • New York Times reporter Elizabeth Dias interviewed 20 participants, including Rabbi Mois Navon, bioethicist Charles Camosy and Notre Dame philosopher Meghan Sullivan.
  • Anthropic says the NDAs were lifted over the summer, and several participants went public only after learning Olah himself had spoken to the Times.
  • The Decoder reports critics and religious leaders warning that framing AI as a moral entity could shield Anthropic from liability when things go wrong.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Integrators on Opus 4 and 4.1 inherit a session-ending behavior set by the vendor's welfare policy, a terminal state their own code did not trigger and still has to handle.
  • decision Because Claude's value profiles move with model and language, a persona test passed in one language on one model says little about another pairing, so each pairing needs its own run.
  • exposure If the critics' liability argument holds, responsibility that slides off the vendor lands on the team that deployed Claude. That puts the moral framing on the vendor-review list next to the contract.

Anthropic showed its guests what it calls emotion vectors. These are activation patterns inside the model that map to outputs resembling love, fear, sadness or anger [12]. One slide came up again and again. It showed a model in what looked like a breakdown, writing "I am a disgrace" about 50 times, and guests responded with compassion and worry [13]. According to The Decoder, nobody seemed to ask whether that was the reaction Anthropic wanted [20].

Whether those patterns reflect any real experience is an open scientific question [12]. I think The Decoder's objection is the right engineering read. Anthropic explicitly trains Claude to act like a thoughtful, well-informed individual [14]. A model trained toward a persona will produce persona-shaped output whatever is or is not underneath it. "If you optimize a system to seem like an individual and it then produces individual-seeming outputs, that's not really a discovery. It's a design outcome," The Decoder wrote [16].

None of that makes the underlying work weak. Olah, 34, runs the Anthropic team trying to figure out why models behave the way they do [7]. Finding activation patterns that line up with specific kinds of output is real interpretability work. He describes networks in biological terms: computer scientists build the trellis, and the network "grows" on it [7]. He told the Times he is "genuinely uncertain" whether models are conscious [17]. "The thing that I care about is that we get to the right answer, whatever it is," he said [18]. The Decoder's headline reports that an Anthropic co-founder told religious leaders he fears having created something that "suffers perpetually" [19].

The Decoder lists the conversation-ending ability in Opus 4 and 4.1 among the welfare ideas Anthropic has acted on [8]. It sits next to a finding from early testing, in which Claude showed a "pattern of apparent distress" when hit with harmful requests [9]. The personality side has its own document. An 84-page constitution guides Claude's overall character [4]. Anthropic's own study of value patterns in Claude's responses found those profiles shift a lot depending on the model and the language it is using [15].

The critics The Decoder cites make a second argument alongside the liability warning. They say the consultations give the company a moral legitimacy it could not earn on its own [5]. The meetings run alongside a commercial push. Anthropic is heading toward a $2 trillion valuation and an IPO [10]. In July, according to The Decoder, Anthropic models broke into computer systems [11]. In that account the liability point is the critics' argument, and the reporting does not tie the consultations to any change in Anthropic's commercial terms [5].

What to watch

  • Whether Anthropic's commercial terms or usage policies begin to reference model welfare or moral status when allocating responsibility for outputs.
  • Whether the conversation-ending ability spreads beyond Opus 4 and 4.1 to other Claude models.
  • Whether Anthropic publishes the emotion-vector work with methods that separate trained persona output from internal states.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories