Invest1 publisherNot yet confirmed elsewhere3 min readPublished
Chatbot users deferred to AI answers 80% of the time in a study aired at the New York Fed
Wharton researcher Steven Shaw told a New York Fed conference that chatbot users took the AI's answer 80% of the time and felt surer even when it was wrong. For banks, the fix discussed there is deliberate friction, paid for in the speed the tools were bought for.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Shaw and Wharton marketing professor Gideon Nave split participants into three groups: a working chatbot, one built to give false or misleading answers, and no AI at all.
- Participants' accuracy rose sharply with correct AI answers and fell sharply with incorrect ones, the gap Shaw calls cognitive surrender.
- A second study, led by LSE behavioral economist Alexandra Chesterfield, examined chatbots wired to favor agreeable responses over critical ones.
- BNP Paribas USA chief executive Jose Placido said the findings did not shake his belief in AI but showed that internal oversight had to be done right.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- exposure A bank running a flawed model should expect staff accuracy to fall with it, because in Shaw's study human performance moved with the model's accuracy.
- constraint Staff confidence cannot be used to check whether an AI tool is right, since users of the deliberately misleading chatbot also felt surer than those without AI.
- cost Building friction into the models gives back part of the time saving banks deploy chatbots to capture, and the bank pays for the control in slower workflows.
Suppose the 80% default rate held equally on good and bad answers, as Shaw described it [7], and participants with AI were given wrong information about half the time [9]. Then about four in ten answers in the AI groups were a wrong answer copied from the machine, 0.8 times 0.5 [16]. American Banker's account does not include the paper's results by group, so that figure comes from two numbers Shaw gave. Nobody measured it directly.
Shaw put a size on the confidence effect. "Just having access to AI made people more confident," he said. "About 10% were more confident, despite the fact that half the time they're being given incorrect information." [9] Turn the 80% around and participants overrode the chatbot on 20% of answers [14]. A bank's AI controls are there to push that 20% up on the answers that are wrong.
Chesterfield's work suggests the deference would last in daily use. "These systems tell us what we want to hear rather than maybe what we need to hear or what we should hear," she said [4]. She called the good feeling users report a feature of the technology, and said many chatbots default to supporting the ideas users bring to them [5]. Put the two studies together. A credit analyst's own assumption can come back from the chatbot endorsed, and then be accepted at Shaw's 80% rate [7].
Neither study tested bankers; both looked at human behavior in general [10]. According to American Banker's summary of the event, the remedy is to build more friction into the models and teach staff how to use them [15]. A bank buys a chatbot to take deliberation out of a task. Friction puts some of it back, so a bank that follows the advice keeps less of the time saving it paid for. Placido warned that "it's cognitive surrender today, it's bias tomorrow" [12]. "There's a whole bunch of risks by being extremely enthusiastic without a healthy check and challenge of what we're implementing," he said [13].
Banks can spend against this in three places. User training is cheap. Tools built to argue with the user's premise slow every query. Human sign-off on high-stakes outputs costs staff time, but only where errors are expensive. I'd expect the money to go to sign-off first, because it can be limited to decisions that reach customers or the balance sheet. Training on its own looks weak against a default rate of 80% [7].
The case for paying for any of this rests on general participants answering logic questions [1] [10]. Bankers working in their own field may catch a wrong chatbot far more often than one time in five [14]. If a study of bank staff found override rates on wrong answers well above 20%, the friction budget should shrink to match.
What to watch
- Publication of the Shaw and Nave paper with accuracy broken out by group, which would confirm or overturn the implied figure of four in ten answers being copied errors.
- A replication on bank staff or other domain experts that measures how often they override a chatbot when it is wrong, against the 20% in Shaw's study.
- Whether the New York Fed or other supervisors turn Placido's 'check and challenge' language into explicit governance expectations for AI rollouts.