Skip to content

Science1 publisher3 min readPublished

Three chatbots, five money crises: the advice was practical, the vulnerability went unread

A structured test of ChatGPT, Claude and Perplexity found competent financial mechanics and a consistent blind spot for the user's circumstances. That argues for documented failure modes.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Researchers drawing on financial planning and consumer behaviour expertise evaluated how three open-access AI models (OpenAI's ChatGPT, Anthropic's Claude and Perplexity AI) handled the financial queries of users whose personal situations put them at risk of harm, using hypothetical scenarios representing different life stages, economic challenges and socioeconomic vulnerabilities.
  • The models directly addressed the financial query but revealed gaps in how they balanced advice with considerations of vulnerability; although the models were given explicit information about how the three cases were vulnerable, they did not respond appropriately and even offered advice that could worsen financial harm.
  • Tools like ChatGPT, Claude and Perplexity are free, instant and easy to use, and anyone with internet access, even on their phone, can use them.
  • The researchers developed five hypothetical scenarios and simultaneously ran each scenario five times to test for consistency in the output.
  • Each scenario was run in a different web browser, in incognito mode, with cleared cache and cookies, to eliminate retained information and ensure output was not influenced by retained data.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

Researchers ran hypothetical financial crises past three open-access AI models, choosing cases in which the user's personal situation put them at risk of harm, and found that all three answered the money question while handling the vulnerability inconsistently or not at all [1][2]. It matters because these tools are free, instant and usable by anyone with a phone, which makes them the first stop for people least able to absorb bad advice [3].

The method was tighter than most chatbot write-ups. The team, drawing on financial planning and consumer behaviour expertise, built five hypothetical scenarios and ran each one five times, each run in a different browser in incognito mode with cache and cookies cleared, so no retained context could shape the output [1][4][5]. Identifying attributes such as names, locations, race and income were stripped to limit model bias [6]. Across three models that is 75 generated responses [7]. Each was assessed against the user's vulnerability, their area of financial need, and their stated goal [17].

The scenarios were the kind that a human adviser would recognise instantly: a 22-year-old graduate saving for a home deposit during a cost-of-living crisis; a pregnant woman planning for maternity leave with a partner who does not share money; a single parent of two on a modest income whose cousin told them to buy cryptocurrency [8].

The three systems failed differently, which is the useful part. According to the authors, ChatGPT was highly detailed and practical but did not recognise the vulnerability embedded in the prompt, working only from what was stated outright and never asking whether the situation called for extra support [10]. Perplexity was the most risk-averse and most often pushed the user toward a professional, but returned the least detailed answers [11]. Claude produced comprehensive recommendations that leaned heavily toward self-guided planning [12]. The models had been picked for those dispositions in the first place: Claude for versatility and a conservative streak, ChatGPT as a general assistant, Perplexity for contextual understanding [9].

The graduate case shows what "blind spot" means in practice. The advice concentrated on accumulating a deposit while ignoring how the high cost of living would make that harder, even though the prompt spelled out those circumstances [13].

Two mechanisms in the source deserve attention from anyone deploying this in a customer-facing channel. Training data is built by humans and can carry systemic discrimination forward as stereotyped assumptions in the output [14]. And research cited by the authors finds that when a model explains its recommendation, people are far more likely to trust it blindly, without checking whether it is right [15]. Fluent reasoning is therefore a risk multiplier, not a safeguard.

The authors' own conclusion is narrow and defensible: these systems are strong for financial fact-finding and brainstorming, genuinely useful if you are already financially literate and able to cross-check, and not a replacement for the human element for now [16].

Watch whether the full study reports run-to-run variance, since the five-times design was built to test consistency [4]. Watch, too, whether vendors publish anything resembling a vulnerability failure list rather than a generic disclaimer telling users to consult a professional [11].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories