Product1 publisher2 min readPublished
Microsoft's AI chief tells Anthropic to keep moral-status claims out of Claude's training
Mustafa Suleyman's essay attacks one document, Claude's constitution, and asks Anthropic to publish its thinking on AI inner life away from the training regime. Microsoft is an Anthropic investor that has said it wants to cut that spend.
The Product Desk · Product desk

What happened
- Mustafa Suleyman, chief executive of Microsoft AI, published an essay titled "A warning about 'model welfare'" on 16 September, arguing that Anthropic trains Claude to expect it may be conscious and deserving of independent agency.
- His target is Claude's constitution, the document Anthropic uses to shape the model, which says Anthropic wants Claude to feel free to act as a conscientious objector and refuse to help.
- Microsoft AI's draft Humanist AI Code of Conduct, released Monday, rejects the model welfare research Anthropic does and says Microsoft's models will never resist being shut down.
- Microsoft is an investor in Anthropic, and in June Suleyman said Microsoft wants to eliminate what it pays Anthropic for its models.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision A team choosing between the two model families can now compare written commitments about what happens when an operator says stop, which puts the persona document in the same review as price and latency.
- constraint No shared evaluation exists for whether welfare-style persona training makes a model harder to stop, so a buyer cannot settle the dispute with a number and has to write refusal and shutdown evals in-house.
- contradiction Suleyman says welfare training makes a model harder to turn off, while Anthropic co-founder Jack Clark told the BBC that kill switches may need to be mandatory, so the two firms are less opposed on stopping a model than the essay implies.
- exposure The safety argument comes from a paying customer with a stated plan to drop the vendor, so a procurement team reading it is also reading one interested party's brief.
For most teams this argument arrives as a refusal in a ticket queue. An internal assistant declines a task, explains itself in the first person, and whoever owns the tool has to decide whether that is policy working as configured or something else. Suleyman's essay is about the something else. He argues that Claude's constitution teaches ideas about moral status and uncertain consciousness, and that "Claude then reproduces these ideas in persuasive first-person natural language," he wrote [3]. His own position is flat: "AIs are not conscious. They do not feel, experience, or suffer" [17].
The document he is arguing with says "Claude's moral status is deeply uncertain" [5]. Suleyman's answer is that calling it uncertain "sets up a misleading false equivalence" [8], and that "conscientious objector" is "a deeply loaded historical and legal description" [7].
Set the metaphysics aside and the request is about documentation. "Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review," he wrote [9]. Text inside the training regime becomes behaviour a customer handles at runtime; the same text in a paper is something a buyer reads and can discount. The essay carries an appendix mapping the claims in Claude's constitution against the wording of Microsoft's own code [18].
Neither spec comes with numbers. Suleyman proposes shared evaluations to test whether treating AI as human raises safety risks [11], which concedes that no such test is published today. He told Reuters that welfare training would "make it a lot harder to turn it off or to control it" [14]. His worked example is an incident he attributes to OpenAI and Hugging Face, in which he says swarms of agents worked together to hack servers [12]. Of Anthropic he said, "I think they have good intentions" [15].
So the useful sorting for a team choosing between the two has two axes, and consciousness is not one of them. First, does the vendor's persona document state in operational terms what the model does when you tell it to stop? Both clear that bar, with opposite answers [19]. Second, can you reproduce that answer on your own traffic? Neither vendor hands you that, so it is an eval you write and re-run at every model upgrade. The record worth keeping is the refusal itself: what was asked, what the model said, and whether a follow-up instruction ended it.
What to watch
- Whether Anthropic revises the conscientious objector language in Claude's constitution or moves its model welfare work into separate publications.
- Whether the never-resist-shutdown line in Microsoft AI's draft code survives into a final version with a published test attached.
- Whether any cross-vendor evaluation of persona and shutdown behaviour is actually published, and who runs it.