Product1 publisher2 min readPublished
HubSpot post links AI's visible-thinking feature to a "perceived effort" effect, citing 2011 and 2022 studies over Anthropic's official explanation
Two published experiments found that visible effort raised perceived quality when the output was identical. Neither tested a reasoning model, and the claim that labs shipped traces to persuade is inference.
The Product Desk · Product desk

What happened
- Claude, ChatGPT and Gemini all shipped near-identical updates in 2025 that show users what the model is thinking while it works, according to a HubSpot post.
- Anthropic's stated reasons for showing Claude's thinking are that it helps people check and trust answers, reveals mismatches between thinking and answers, and is simply interesting to watch.
- A 2011 Management Science study of 266 travel-search users found the same flight results were perceived as 8.1% higher value when a live list of airlines being searched replaced a blank loading wheel.
- A later group of 118 participants was more likely to pick a slower site that showed its searching over a faster one that did not, with both sites returning the same results.
- In 2022 research in Information and Management, seven seconds of a "calculating results" loader raised quality ratings for recommendations identical to the ones delivered instantly.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint Both experiments held the output identical, so they cannot tell a product team what to expect when the visible steps also change the answer the user gets.
- exposure If visible effort lifts ratings on identical output, a satisfaction score on an assistant no longer measures answer quality on its own, and whoever presents that lift internally has to say which part is perception.
- decision A team weighing a thinking panel has to settle in advance whether a perceived-quality rating counts as success, since perceived quality on identical output is what the cited studies measured.
- contradiction The effect the post describes rests on published experiments; its claim about why the labs shipped traces rests on the author's inference about hidden motive.
The post treats a loading spinner and a reasoning trace as the same product move. A spinner shows the user only that time is passing; the trace, as the post describes it, tells the user which websites were searched, which documents were opened and which assumptions got second-guessed [2].
Both cited experiments were built the same way, and the design limits what they can tell a product team. The 2011 travel test returned the same flights to both groups and varied only whether the wheel showed airlines and fares stacking up [5][6]. The 2022 car and dating studies returned identical recommendations, with seven seconds of a rotating "calculating results" loader as the only difference between conditions [12]. Holding output constant isolates perception. A reasoning trace works differently: the steps on screen come out of the same run that writes the answer.
Add the samples up and the pattern rests on 984 people, none of them using an assistant: 266 in the travel test, 118 in the choice test, 306 in the car study and 294 in the dating study [16]. The post does not report usage data from Claude, ChatGPT or Gemini on how often anyone opens a trace [18].
The strongest line in the post is that people preferred the transparent search even when it took 50 seconds longer to gather the same results [8]. The 2011 waits were randomly set at 10, 20, 30, 40, 50 or 60 seconds [7], so a 50-second penalty is the widest gap that design could produce, a 60-second transparent search against a 10-second blank one [17].
On motive the post goes past its evidence. Its author wrote that Anthropic and OpenAI might tell us the transparent loading screen is there only because it is simply interesting to watch, and that most marketers know that is BS [14]. The argument is that providers have purposefully hidden a fourth reason [15]. The support offered for it is the two studies.
For anyone deciding whether to expose intermediate steps, two variables separate the cases: whether the visible work changes the output, and whether users do anything with it. Where both answers are no, you are buying a higher rating with latency and screen area; the two studies measured exactly that effect on identical output [13]. Where either answer is yes, your case falls outside the labor illusion research, and the numbers that apply are repeat use, task completion, and the rate at which people open the trace and act on what they find there. The first two reasons Anthropic gives describe that second case [3].
What to watch
- Whether any lab publishes the rate at which users open a reasoning trace. That would move the argument from perceived quality to behaviour.
- A replication that varies answer quality alongside visible effort, since the 2011 and 2022 designs both held output identical.
- Whether Anthropic or OpenAI adds a setting to hide the thinking panel, and what happens to reported satisfaction when users turn it off.