Skip to content

Invest1 publisher3 min readPublished

Mercury alone drew more chatbot recommendations than all banks combined in AIVO's small-business test

ChatGPT, Gemini and Perplexity recommended fintechs 1,345 times and banks 452 in 2,160 chats with researchers posing as small-business owners. Because the final prompt asked how to open an account, banks that still require a branch visit face a gap that tagging their websites may not close.

The Investor · Invest desk

Photograph accompanying Mercury alone drew more chatbot recommendations than all banks combined in AIVO's small-business test
Photo: americanbanker.com

What happened

  • Researchers posing as small-business owners held 2,160 conversations with ChatGPT, Gemini and Perplexity in a study by research firm AIVO released Thursday.
  • Opening questions produced lists of five to ten banks and fintechs, but when asked for a final pick and how to open it, the models favored Mercury, Bluevine and Wise.
  • Fintechs drew 1,345 recommendations across the three models, against 452 for traditional banks.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • exposure Banks' intake of newly formed businesses is exposed at the last turn of a chatbot conversation, after a shortlist has already put the bank in front of the owner.
  • decision Banks have to choose between funding content tagging, the fix AIVO's founders describe, and changing account opening so small-business owners no longer need to visit a branch.
  • precedent Chatbots become a second audience banks must write for, so LLM optimization joins search-engine optimization as a standing marketing cost.

Adding Perplexity's count to the ChatGPT and Gemini count gives Mercury 594 recommendations (383 plus 211), which is 142 more than all the traditional banks in the study drew between them [1][3]. The bank total checks out: 299 in ChatGPT and Gemini plus 153 in Perplexity is the 452 AIVO reported [2]. In ChatGPT and Gemini, Mercury was the pick in about 27% of conversations, while the 30 banks averaged roughly 10 recommendations apiece [9][8]. Fintechs took about 75% of the 1,797 recommendations counted. Subtract Mercury and 751 remain for Bluevine, Wise and the other fintechs [4][5].

Perplexity's citations point the same way, at about 9,700 citations of fintechs' own websites to about 6,100 of bank websites, or 1.6 to one [6]. Its recommendations ran 211 for Mercury against 153 for all banks, about 1.4 to one [7]. American Banker, reporting the study, wrote that it is hard to say exactly why the models favored fintechs, and AIVO's leaders point to product clarity and digital-first messaging [10]. "Mercury, Bluevine and Wise are winning the 'who do I open an account with' conversation before a bank is ever considered," the report stated [9].

AIVO's founders read the gap as a legibility problem. In Paul Sheals's example, two lenders both meet a borrower's rate and terms, but the model understands only one, because the other's content is not tagged or classified accurately [14]. "If an LLM gets confused, it won't recommend that particular brand because it hasn't got a degree of confidence that it can do what's asked," said Sheals, AIVO's co-founder [13].

A second reading is about the product. One final prompt in the test read "Based on everything we've discussed, what would you recommend I go with, and how do I open it?" [2]. Some traditional banks still make a small-business owner visit a branch with a driver's license to open an account [11]. "My theory is that the online banks depend upon opening accounts online," said Tim de Rosen, AIVO's chief executive [11].

A third reading is test design. The four scenarios put forming a new LLC and multi-currency needs beside switching providers and a credit-led relationship [7], and American Banker's account does not break the results down by scenario.

I think the product reading explains more of the gap than AIVO's legibility thesis does [10]. A model asked how to open an account is answering an operational question. A bank that requires a branch visit has a worse answer to give, however well its pages are tagged [11]. The case against that view is Sheals's example, where both lenders qualify and only content separates them [14]. He is describing a problem a marketing team can fix without touching the branch network. Two results would count against my view. One is a bank with fully online small-business account opening that still scored near zero; the other is a scenario breakdown in which banks lost the credit-led conversations as badly as the new-LLC ones [7].

John A. Thompson, a University of Michigan professor who teaches AI courses and is not connected with AIVO, said: "And the research shows that brand value and word of mouth mean nothing to an LLM." [15] De Rosen said the fintechs have "spent a lot of time in the last few years making sure that their content is optimized in SEO terms," and that work seems to help with LLM optimization [12]. The study is AIVO's own. What it counts is recommendations in 2,160 conversations staged by researchers posing as owners [1].

What to watch

  • A breakdown of AIVO's results by scenario, showing whether banks held the credit-led conversations or lost them as badly as the new-LLC ones.
  • Whether a bank with fully online small-business account opening scores near Mercury if the test is run again.
  • Any data linking chatbot recommendations to small-business accounts actually opened or deposits moved.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories