Science1 publisher2 min readPublished
Specific suggestions help explain AI's emotional-support edge in two follow-up experiments
Manchester and Durham researchers had 390 people rate comfort messages without knowing who wrote them, and the AI text scored higher for anger and fear. Human messages carrying the same practical advice were judged just as supportive.
The Scientist · Science desk
What happened
- A team at the universities of Manchester and Durham ran five experiments comparing emotional-support messages written by people with messages produced by large language models.
- In one of them, 390 participants imagined an anger, sadness or fear scenario and read a single reply without being told its source, and they rated the AI replies more emotionally supportive for anger and for fear.
- For the sadness scenario, the difference between the AI and human messages in rated support was not reliable.
- Two further experiments pointed to what the team calls actionable support, meaning specific and realistic suggestions, as what made messages more comforting whether a model or a person wrote them.
- A 2025 US survey of 1,058 people aged 12 to 21 found 13% had sought advice from gen AI when feeling sad, angry or nervous, rising to 22% among those aged 18 to 21.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- exposure Assistants sold for drafting email and planning holidays are taking emotional disclosures from teenagers and young adults, so how a general-purpose tool handles a distressed user is a design question for teams that never scoped one.
- contradiction The same body of work supports two readings: blind readers score the AI text higher, and people still pick a human for emotional engagement when they know the difference and have to wait for it. Winning on rated message quality does not win the channel.
- decision Disclosure and perceived comfort pull against each other in the same product. A team that labels AI-written support messages should expect lower empathy and support ratings for text it has not changed.
- constraint Nothing here supports a claim about mental health outcomes. A vendor citing these results as evidence that its chatbot helps distressed users is going well past what the experiments measured.
The follow-up experiments asked what the AI messages contained. One tested whether explicitly acknowledging and validating the recipient's feelings explained the difference, and validation did not account for the greater emotional improvement associated with gen AI [6]. Two others tested specific and realistic suggestions. When human messages offered comparable help, they were judged just as supportive [8]. The team's own explanation for the gap is consistency: a model can produce a structured response on demand, usually with a concrete suggestion in it, while a person may be tired or unsure what to say [9].
The direction matches a review of 23 studies, in which people tended to rate gen AI messages as more empathic than human-written ones, an effect researchers call the "AI advantage" [5]. The 390 participants were not told the source, which separates the content of a message from its attribution. Attribution moves ratings by itself. Across nine further studies involving 6,282 people, the same gen AI responses were rated more empathic and supportive when described as human-written than when attributed to AI [12]. Those nine studies together involved about 16 times as many people as the blind test [18].
The improvements were not uniform across emotions. In the fear scenario, the gen AI messages increased calm more than the human messages and did not produce a greater reduction in fear itself [4]. More advice is not better either: a 2024 study found that excessive suggestions could be less effective at making people feel heard [10], and the same study found that the gen AI advantage in making people feel heard declined when recipients believed the message came from AI [11].
People still chose people. Participants preferred human interaction when seeking emotional engagement, even when choosing a person meant waiting longer [13]. The authors write that research has yet to establish why, and offer one possibility: a human response signals someone's willingness to spend time and emotional effort [17].
The outcome measures here are how supportive a message seemed and which emotions a participant reported after reading it, in situations they were asked to imagine [2] [4]. The couples study in the same programme is correlational, and in it the recipient's perception of a partner's efforts was more consistently associated with both partners' relationship ratings than the helper's own account [15]. The authors state that these studies identified associations and did not establish that particular forms of support caused the outcomes [16].
What to watch
- A study that recruits people in actual distress and measures how they are doing later, instead of asking participants to imagine anger, sadness or fear.
- Whether the actionable-support effect replicates when the suggestion offered is impractical or wrong for the recipient.
- Whether AI-disclosure rules get tested against the finding that an AI label lowers how supportive identical text feels.