Science1 publisher2 min readPublished
Participants rated the most responsive AI less conscious than the least responsive human in LMU vignettes
LMU Munich researchers found that nearly 1,100 people withheld 'conscious' from an AI agent even when it acted exactly as a human did in an identical scenario. Because the test used written scenarios, it measures the words people choose more than how they treat a chatbot they use daily.
The Scientist · Science desk

What happened
- One group read scenarios with an AI protagonist and another read word-for-word identical scenarios with a human, then rated how aware or conscious the agent was of its surroundings.
- The vignettes varied how strongly the agent responded to its situation, such as reacting to a change in sound or to the emotional state of someone nearby.
- Ratings of awareness rose with the agent's responsiveness, and the gap between human and AI agents on that measure largely disappeared.
- On consciousness, an AI showing maximum responsiveness was still rated below a human showing the bare minimum.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint Because participants rated written scenarios, the result cannot settle whether months of conversation with a chatbot change what people believe about it or how they act toward it.
- decision Arguments for rules premised on users mistaking human-like chatbots for feeling beings now have to contend with evidence that, asked directly, people reserve 'conscious' for humans.
- capability Holding behavior fixed and swapping only the agent's label gives researchers a way to measure how much of a mental-state judgement comes from knowing the agent is a machine.
Swapping only the protagonist is a clean control. Every scenario existed in an AI version and a human version with the same wording [4]. A lower consciousness rating for the AI therefore cannot come from the AI behaving less like a person. What is left is the label, and whatever participants already believe a machine is. Louis Longin, the lead author, from LMU's Chair of Philosophy of Mind [11], said that comparison is what sets the work apart: "Whereas previous studies have typically asked general questions about whether AI actually has mental states, our study is the first to directly compare how people attribute the same mental states to AI and humans behaving in exactly the same way, under identical circumstances." [7]
The shape of the result is what makes it interesting. Awareness ratings moved with behavior, much as a measurement would [5]. Consciousness ratings did not. Even when the AI acted identically to a human, people declined to call it conscious [1], and no level of responsiveness in the vignettes lifted the AI to the human baseline [6]. The LMU summary explains this as people viewing consciousness as uniquely biological or subjective [14]. That explanation is the authors' reading of the ratings.
Ophelia Deroy, the co-senior author, who holds the Chair of Philosophy of Mind at LMU [11], offered a plainer account of why people talk about chatbots in mental terms at all. "We may be using mental terms to refer to AI, but this may only be because we lack better words," she said [8]. "After all, AI is trained to act and speak like a human, so these descriptions seem natural." [9]
The study is aimed at a specific worry. "There is a growing worry in public discussion that people will see AI behaving in a human-like way and start treating these systems as if they had human-like mental states," Longin said [10]. According to the LMU summary, the results counter fears that human-like generative AI will fool the public into treating machines as feeling entities [13].
I think that holds for what people say when asked directly about a described agent. Treating a machine as sentient is a behavior, and this design did not observe behavior. Participants read short vignettes [3]. They did not spend weeks in conversation with a system. A person could refuse to call a chatbot conscious and still act toward it as though it had feelings, and a rating scale would not register that.
The press summary does not report effect sizes, how the nearly 1,100 participants were divided between conditions, or where they were recruited [3].
What to watch
- Whether the full Cognition paper reports effect sizes and participant demographics, and whether the pattern holds across countries and languages.
- Studies that track how people behave toward a chatbot over repeated use, to test whether the stated refusal to call it conscious survives contact with a live system.