Invest1 publisher2 min readPublished
Claude, ChatGPT and Gemini grade seven banks an average 37.6 against the public's 70.0
RepTrak ran its 23-factor reputation questionnaire through Claude, ChatGPT and Gemini for seven banks, and the models came back with Ally at 52.1 and TD Bank at 19.4 while the same questions put the informed public's average at 70.0.
The Investor · Invest desk

What happened
- RepTrak's AI-rated scores for seven banks on a 100-point scale run from Ally at 52.1 and Chase at 50.3 down through Truist, Bank of America and Chime to Wells Fargo at 22.2 and TD Bank at 19.4.
- Asked the same questions, the informed general public scored those banks an average of 70.0, with Wells Fargo lowest at 55.5 and Chime highest at 71.5.
- Perceived financial performance was the only one of the seven drivers on which the platforms scored banks significantly more positively than the survey respondents did.
- The platforms draw on bank-owned and third-party material, deciding for themselves which sources to prioritise, how to interpret them and what verdict to return.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint RepTrak found no consistent correlation between the human rating and the AI rating, so a bank that lifts survey perception has no evidentiary basis for expecting the chatbot grade to follow it.
- exposure Legacy content sits inside the source pool the platforms read, which means coverage a bank cannot retract keeps feeding a conduct score years after the matter closed.
- decision The remedy on offer is supply-side and competes for the same communications budget as paid media, and the firm measuring the problem is also the firm diagnosing it.
- contradiction The column calls reputation a leading indicator of business outcomes, yet no figure in the study connects TD Bank's 19.4 to a deposit balance, an account opening or a funding spread.
Ally's 52.1 is the highest score any of the seven banks got from Claude, ChatGPT and Gemini, and it lands 3.4 points below the 55.5 the informed general public gave Wells Fargo, the bottom of the human table [2][3][3]. Averaged, the three platforms come out at 37.6 against the survey's 70.0, a gap of 32.4 points [1][2]. The two panels also disagree about how different these banks are from each other: 32.7 points separate top from bottom on the AI side, against 16.0 on the survey side [4].
Chime is the clearest case. The public put it first at 71.5; the models put it fifth of seven at 37.3, a gap of 34.2 points, wider than Wells Fargo's 33.3 [2][3][5][7].
RepTrak's column says "AI is especially critical of banks not standing behind their products and services, their lack of ethical behavior and environmental shortcomings" [11]. The driver splits it publishes land hardest in one of those places: conduct comes in 41.9 points below the survey reading, against 29.0 for products and services and 15.9 for citizenship [6].
The sample is small. Seven names, picked as a mix of institutions "for illustrative purposes", scored on a 23-factor questionnaire and averaged across three platforms [7][1]. The piece ran as opinion at American Banker, written from an organisation that has been doing cross-industry reputation measurement for more than 20 years [13], and its recommended response is that banks "redouble efforts to ensure their perspectives are authentic, accessible, and readily discoverable" [10].
Two readings compete, and they carry different price tags. If the ordering is an artifact of prompt wording and model version, the next releases reset it, and a 19.4 is not a number any bank should be budgeting against. If the platforms are instead reading a third-party record on conduct that does not change much, the order survives the reruns, and the only input a bank owns is how much of its own material is there to be weighed. Running the same questionnaire on the next model versions, and checking whether TD Bank is still 32.7 points behind Ally, would separate the two [4].
What to watch
- A rerun on the next Claude, ChatGPT and Gemini releases: if the ordering moves, the scores track model versions and not banks.
- Any bank disclosing a funding cost, acquisition cost or deposit movement it attributes to an AI-generated verdict.
- Per-platform scores and prompt methodology, instead of a single figure averaged across three platforms.