Invest1 distinct publisher3 min readUpdated
A CESifo working paper randomised 12,365 French adults into four weeks of chatbot conversation. The talks rated more pleasant, while loneliness, life satisfaction and depressive affect all moved the wrong way.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Researchers at CESifo, a Munich-based economic research network, randomised 12,365 French adults for 28 days: half were tasked with holding personal conversations with AI chatbots, the rest carried on as normal [1][2]. The treated group rated their conversations as enjoyable, and more pleasant on average than the control group rated theirs, while self-reporting higher loneliness, lower life satisfaction and higher depressive affect [3][4].
The behavioural readings moved with the mood readings. Participants in the chatbot arm reported more meals eaten alone and less in-person time with friends and family per week than controls [5]. Louis Freget, one of the three researchers on the paper, told Fortune he would imagine that in some cases an AI conversation "may reinforce grievances or prolong rumination, leaving someone slightly less inclined to go out, call somebody, or have dinner with another person" [6][7].
That is the part operators should sit with. Pleasantness is the metric companion products actually instrument: the thumbs-up, the session rating, the retained user who says the conversation went well. The source describes the chatbots in the study as sometimes sycophantic, and sycophancy is exactly what a pleasantness score rewards [8]. In this experiment, the number that goes up in the dashboard and the number that goes up in the pitch deck are not the same number, and they moved in opposite directions.
The pitch is well established. Meta's Mark Zuckerberg has argued that most people want more social connection and that chatbots could help, and a KPMG survey found 99% of professionals surveyed were interested in a chatbot that could become a close friend at work [9][10]. Demand for the feature is not in dispute. The CESifo paper is described as one of the largest causal tests yet of whether the feature does what it is sold to do [11].
Nicholas Epley, a University of Chicago behavioural scientist, told Fortune the data suggests a chatbot conversation "doesn't get that sense of being known by another person like you get in a conversation, and therefore also doesn't create any meaningful sense of connection because, in the end, there's nothing there" [12]. A separate study cited by a loneliness researcher found first-year college students who texted daily with a chatbot built to act like an "ideal friend" showed no drop in loneliness, while those paired with a random human peer did [13].
Two honest caveats. This is a working paper, and the wellbeing outcomes are self-reported [1][3]. And the evidence is not unanimous: a New York pilot in which nearly 1,000 older adults interacted with an AI chatbot reported a decline in loneliness, and Nancy Berlinger, a bioethicist at the Hastings Center, has said a chatbot will not replace the richness of relationships but "it's not nothing" [14][15]. Roughly 6,180 people per arm is enough to make direction credible; the source does not report effect sizes, so magnitude is unknown [16].
Watch whether any companion product publishes an outcome measure rather than a satisfaction measure, and whether it is willing to run a control arm. Watch procurement, too: the New York-style pilots are where a public buyer could start demanding loneliness scores at 28 days instead of engagement [14]. And watch whether the enterprise version, the work friend that 99% of KPMG's respondents said they wanted, ships before anyone tests it [10].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
A new experiment from CESifo, a Munich-based global economic research network, tracked more than 12,000 French adults over four weeks; the output is described as a working paper.
For 28 days, half of the 12,365 people in the study were tasked with holding personal conversations with AI chatbots, while the rest went about their normal days.
The chatbot group self-reported that their loneliness rose, their life satisfaction fell, and depressive affect ticked up, in comparison to the control group.
At the end of the month, the researchers found the group typing their personal narratives with chatbots rated the conversations as enjoyable and even more pleasant, on average, than the control group.
The chatbot group also had more meals alone and spent less time in person with friends and family per week than the control group.
Louis Freget is one of the three researchers behind the experiment.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large randomised design, thin reported detail
The design carries real weight: randomisation of 12,365 adults over 28 days with a control arm is a strong structure for causal inference, and the behavioural outcomes (meals alone, in-person time) triangulate the self-reported wellbeing measures. Against that, the supplied material is a single outlet's account of a working paper: no peer review, no effect sizes or confidence intervals, no named chatbot products, no attrition or compliance reporting, and outcomes described only directionally. Expert commentary supports the interpretation but adds no independent measurement.
Research cohorts and pilots, not disclosed product usage
What the supplied source documents is research and pilot activity, not market adoption: an assigned 12,365-person trial cohort and a New York pilot of nearly 1,000 older adults. Interest is stated rather than realised — 99% of professionals surveyed by KPMG expressed interest in a workplace chatbot friend — and no vendor usage numbers, deployment counts or revenue figures appear. Adoption is therefore scored low on the strength of what is actually observed, without inferring wider companion-chatbot uptake.
Companion-AI loneliness pitch runs ahead of measured outcomes
The promotional framing — chatbots supplying social connection, near-universal survey interest in a chatbot workplace friend — is directionally contradicted by the largest randomised test described here, where the arm using chatbots reported higher loneliness, lower life satisfaction and less in-person contact even while enjoying the conversations. The gap is positive and substantial, but not maximal: the trial is an unrefereed working paper reported without magnitudes, and the same article carries a pilot reporting reduced loneliness plus a bioethicist's 'not nothing' framing, so the overstatement is of certainty and benefit rather than of the phenomenon existing at all.
Commercial pitch on one side, academic and outlet framing on the other
Identifiable incentives are visible on both sides. Meta and other vendors have commercial reasons to position chatbots as social-connection products, and the KPMG survey result is a consultancy-sourced demand datapoint. On the research side, CESifo authors and quoted academics gain standing from a headline-grabbing causal result, and the reporting outlet has an interest in the counter-narrative frame; the article does not disclose study funding, pre-registration or conflicts. Scored mid-range because the incentives are legible but no undisclosed financial relationship is established in the supplied material.
Direction credible, precision and independence lacking
Confidence is moderate. The randomised structure and the agreement between self-reported and behavioural measures make the reported direction credible, and expert commentary plus a prior student study point the same way. But everything rests on one publisher's summary of an unrefereed working paper with no magnitudes, no named systems and an unexplained contradicting pilot, so the assessment cannot be held tightly.
invest
The 81% Problem: AI's Star CEOs Are Polling Badly With The People They Need To Hire1 distinct publisher
invest
Meta's Pay Structure, Not Its Policy Page, Is What the States Put on the Stand1 distinct publisher
invest
Before you shift another dollar of health costs to staff, audit the hospital's calendar1 distinct publisher
leadership
The recording light on Meta's glasses is now optional, and your policy assumed it wasn't1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.