Skip to content

Product1 publisherNot yet confirmed elsewhere3 min readPublished

After Arup, a face on a video call is not a credential

Fifteen transfers worth about $25m went out because the people on the call looked and sounded right. Detection accuracy is near chance, so the surviving control is procedural.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • In January 2024, an employee at professional services firm Arup joined a video call with someone they believed included the company's CFO; the call resulted in 15 wire transfers to third-party accounts totalling about $25m.
  • Every participant on the other side of the Arup video call was an AI-generated clone cobbled together from public appearances and earnings calls of Arup executives.
  • The Arup transfers averaged roughly $1.7m each.
  • Deepak Gupta, technical CEO at GrackerAI, told ZDNET when discussing the Arup incident: "Seeing and hearing someone is no longer proof they are real." and "Any protocol that relies on 'I recognized their face and voice' is now broken."
  • Deepak Gupta holds the title technical CEO at GrackerAI.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

In January 2024, an employee at professional services firm Arup joined a video call they believed included the company's CFO, and the call resulted in 15 wire transfers to third-party accounts totalling about $25m [1]. Every participant on the other side of that call was an AI-generated clone assembled from public appearances and earnings calls of Arup executives [2].

That works out to roughly $1.7m per transfer [7]. The detail worth sitting with is not the total but the count: the control did not fail once, it failed fifteen times in sequence, because each time the employee re-verified by looking at a screen. "Seeing and hearing someone is no longer proof they are real," Deepak Gupta, technical CEO at GrackerAI, told ZDNET about the incident. "Any protocol that relies on 'I recognized their face and voice' is now broken" [3][4].

Live voice and video calls have been the default standard for identity verification through most of corporate history, used to authenticate high-value transactions and sensitive legal and medical exchanges [8]. The tells staff were taught to listen for were background noise, synthetic voice modulation, and the absence of breathing sounds [9]. The measured value of that training is thin. A University College London study found listeners identified deepfakes correctly only about 73% of the time, and training with deepfake samples improved accuracy by 3.84 percentage points [5], which is a trained accuracy of about 77% and a miss on roughly one attempt in four [10]. Researchers at the University of Duisburg-Essen and Indiana University then aggregated 56 similar studies and found human detection rates closer to chance [14]. James Scobey, CTO at S2i2, told ZDNET that tells retain marginal value as a supporting signal but no longer work as a control, because they train employees to rely on the exact perception the attacker is engineering [11].

Buying a detector does not resolve this either. A joint information sheet from the NSA, FBI and CISA ruled out traditional automated detection that visualises evidence of manipulation in voice or video, because those methods assume statistically significant traces of manipulation exist to be found [6].

The threat is also not confined to a single fraudulent call. Gupta pointed to a 2024 incident in which security training firm KnowBe4 hired an attacker whom interviews and screening did not flag, and issued a company workstation before discovering the hire was a North Korean operative using the device to upload malware [12]. Gupta called it "almost poetic" that the victim was a company whose business is teaching people to spot social engineering [13], and said North Korean operators run this at scale, sending thousands of fake workers a year [15].

Which leaves the unglamorous answer. Security experts including Gupta and Scobey argue the defence lies in low-tech protocols once dismissed as too simple for large-scale business operations, and ZDNET framed the whole piece as a solution from the past [16][17]. The reading that follows from the evidence above is procedural: verification has to leave the channel the attacker controls. A callback to a number pulled from the directory rather than from the meeting invite, and a shared phrase agreed in advance, are cheap and do not depend on an employee out-detecting a model.

Watch whether these procedures get written into payment authorisation itself, with a hard rule that no transfer is released on the strength of a live call. Watch whether hiring and device issuance get the same treatment, since the KnowBe4 case shows the failure is not only in finance [12]. And watch whether firms stop funding detection training whose measured lift is under four points [5].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories