Science1 distinct publisher3 min readUpdated
A structured test of ChatGPT, Claude and Perplexity found competent financial mechanics and a consistent blind spot for the user's circumstances. That argues for documented failure modes.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Researchers ran hypothetical financial crises past three open-access AI models, choosing cases in which the user's personal situation put them at risk of harm, and found that all three answered the money question while handling the vulnerability inconsistently or not at all [1][2]. It matters because these tools are free, instant and usable by anyone with a phone, which makes them the first stop for people least able to absorb bad advice [3].
The method was tighter than most chatbot write-ups. The team, drawing on financial planning and consumer behaviour expertise, built five hypothetical scenarios and ran each one five times, each run in a different browser in incognito mode with cache and cookies cleared, so no retained context could shape the output [1][4][5]. Identifying attributes such as names, locations, race and income were stripped to limit model bias [6]. Across three models that is 75 generated responses [7]. Each was assessed against the user's vulnerability, their area of financial need, and their stated goal [17].
The scenarios were the kind that a human adviser would recognise instantly: a 22-year-old graduate saving for a home deposit during a cost-of-living crisis; a pregnant woman planning for maternity leave with a partner who does not share money; a single parent of two on a modest income whose cousin told them to buy cryptocurrency [8].
The three systems failed differently, which is the useful part. According to the authors, ChatGPT was highly detailed and practical but did not recognise the vulnerability embedded in the prompt, working only from what was stated outright and never asking whether the situation called for extra support [10]. Perplexity was the most risk-averse and most often pushed the user toward a professional, but returned the least detailed answers [11]. Claude produced comprehensive recommendations that leaned heavily toward self-guided planning [12]. The models had been picked for those dispositions in the first place: Claude for versatility and a conservative streak, ChatGPT as a general assistant, Perplexity for contextual understanding [9].
The graduate case shows what "blind spot" means in practice. The advice concentrated on accumulating a deposit while ignoring how the high cost of living would make that harder, even though the prompt spelled out those circumstances [13].
Two mechanisms in the source deserve attention from anyone deploying this in a customer-facing channel. Training data is built by humans and can carry systemic discrimination forward as stereotyped assumptions in the output [14]. And research cited by the authors finds that when a model explains its recommendation, people are far more likely to trust it blindly, without checking whether it is right [15]. Fluent reasoning is therefore a risk multiplier, not a safeguard.
The authors' own conclusion is narrow and defensible: these systems are strong for financial fact-finding and brainstorming, genuinely useful if you are already financially literate and able to cross-check, and not a replacement for the human element for now [16].
Watch whether the full study reports run-to-run variance, since the five-times design was built to test consistency [4]. Watch, too, whether vendors publish anything resembling a vulnerability failure list rather than a generic disclaimer telling users to consult a professional [11].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Researchers drawing on financial planning and consumer behaviour expertise evaluated how three open-access AI models (OpenAI's ChatGPT, Anthropic's Claude and Perplexity AI) handled the financial queries of users whose personal situations put them at risk of harm, using hypothetical scenarios representing different life stages, economic challenges and socioeconomic vulnerabilities.
The models directly addressed the financial query but revealed gaps in how they balanced advice with considerations of vulnerability; although the models were given explicit information about how the three cases were vulnerable, they did not respond appropriately and even offered advice that could worsen financial harm.
Tools like ChatGPT, Claude and Perplexity are free, instant and easy to use, and anyone with internet access, even on their phone, can use them.
The researchers developed five hypothetical scenarios and simultaneously ran each scenario five times to test for consistency in the output.
Each scenario was run in a different web browser, in incognito mode, with cleared cache and cookies, to eliminate retained information and ensure output was not influenced by retained data.
To reduce the risk of AI-generated bias, the researchers removed identifying attributes that could influence model outputs, such as names, locations, race and income.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-reported study, method described but results unquantified
The design has real controls — repeated runs for consistency, fresh incognito sessions, stripped identifiers, stated assessment dimensions — and the failure mode is illustrated with concrete worked examples. Against that: a single source, authored by the researchers themselves, with no model versions or access dates, no scoring rubric or counts, only three of five scenarios described, no link to a peer-reviewed paper in the supplied text, and no independent replication or vendor response.
No usage or deployment data supplied
The cluster contains no usage figures, deployment counts, market data, or disclosure of how often people actually use these assistants for financial decisions. The only adoption-adjacent statement is qualitative — that the tools are free, instant and reachable by anyone with internet access — which cannot be scored. The one observation in this payload is a benchmark-style evaluation, not evidence of uptake.
Modest overreach in generalisation, restrained conclusions
The conclusions are deliberately hedged — useful for fact-finding, not a replacement for a human adviser — which keeps the gap small. It is positive rather than zero because unnamed, undated model builds tested on a handful of hypothetical prompts are generalised to 'AI' as a category, an alarming-failure framing is applied without any quantified rate, and a blind-trust claim is invoked without citation.
Researcher-authored promotion of own study, adjacent to human-advice interest
The article is a first-person write-up by the researchers publicising their own work, so there is a straightforward visibility incentive, and the authors' field — financial planning and consumer behaviour — has a professional stake in the closing conclusion that AI cannot replace the human element. That interest is not disclosed or interrogated in the supplied text, and no vendor or independent voice is included. Incentives are not scored higher because no funding conflict, commercial product, or sponsorship is evidenced in the cluster.
Low-to-moderate: plausible mechanism, thin verification
The direction of the finding is plausible and internally consistent, and the method is described well enough to be re-run. But confidence is capped by one publisher, self-reported qualitative results, undisclosed model versions, an unreported response count, one uncited supporting literature claim, and zero adoption or incident data to corroborate real-world consequence.
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
security
The nationalization argument is really a vendor-continuity memo1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026