Product1 distinct publisher3 min readUpdated
Fifteen transfers worth about $25m went out because the people on the call looked and sounded right. Detection accuracy is near chance, so the surviving control is procedural.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
In January 2024, an employee at professional services firm Arup joined a video call they believed included the company's CFO, and the call resulted in 15 wire transfers to third-party accounts totalling about $25m [1]. Every participant on the other side of that call was an AI-generated clone assembled from public appearances and earnings calls of Arup executives [2].
That works out to roughly $1.7m per transfer [3]. The detail worth sitting with is not the total but the count: the control did not fail once, it failed fifteen times in sequence, because each time the employee re-verified by looking at a screen. "Seeing and hearing someone is no longer proof they are real," Deepak Gupta, technical CEO at GrackerAI, told ZDNET about the incident. "Any protocol that relies on 'I recognized their face and voice' is now broken" [4][5].
Live voice and video calls have been the default standard for identity verification through most of corporate history, used to authenticate high-value transactions and sensitive legal and medical exchanges [6]. The tells staff were taught to listen for were background noise, synthetic voice modulation, and the absence of breathing sounds [7]. The measured value of that training is thin. A University College London study found listeners identified deepfakes correctly only about 73% of the time, and training with deepfake samples improved accuracy by 3.84 percentage points [8], which is a trained accuracy of about 77% and a miss on roughly one attempt in four [9]. Researchers at the University of Duisburg-Essen and Indiana University then aggregated 56 similar studies and found human detection rates closer to chance [10]. James Scobey, CTO at S2i2, told ZDNET that tells retain marginal value as a supporting signal but no longer work as a control, because they train employees to rely on the exact perception the attacker is engineering [11].
Buying a detector does not resolve this either. A joint information sheet from the NSA, FBI and CISA ruled out traditional automated detection that visualises evidence of manipulation in voice or video, because those methods assume statistically significant traces of manipulation exist to be found [12].
The threat is also not confined to a single fraudulent call. Gupta pointed to a 2024 incident in which security training firm KnowBe4 hired an attacker whom interviews and screening did not flag, and issued a company workstation before discovering the hire was a North Korean operative using the device to upload malware [13]. Gupta called it "almost poetic" that the victim was a company whose business is teaching people to spot social engineering [14], and said North Korean operators run this at scale, sending thousands of fake workers a year [15].
Which leaves the unglamorous answer. Security experts including Gupta and Scobey argue the defence lies in low-tech protocols once dismissed as too simple for large-scale business operations, and ZDNET framed the whole piece as a solution from the past [16][17]. The reading that follows from the evidence above is procedural: verification has to leave the channel the attacker controls. A callback to a number pulled from the directory rather than from the meeting invite, and a shared phrase agreed in advance, are cheap and do not depend on an employee out-detecting a model.
Watch whether these procedures get written into payment authorisation itself, with a hard rule that no transfer is released on the strength of a live call. Watch whether hiring and device issuance get the same treatment, since the KnowBe4 case shows the failure is not only in finance [13]. And watch whether firms stop funding detection training whose measured lift is under four points [8].
Ranked by verification strength, evidence, and original report placement.
In January 2024, an employee at professional services firm Arup joined a video call with someone they believed included the company's CFO; the call resulted in 15 wire transfers to third-party accounts totalling about $25m.
Every participant on the other side of the Arup video call was an AI-generated clone cobbled together from public appearances and earnings calls of Arup executives.
Deepak Gupta, technical CEO at GrackerAI, told ZDNET when discussing the Arup incident: "Seeing and hearing someone is no longer proof they are real." and "Any protocol that relies on 'I recognized their face and voice' is now broken."
Deepak Gupta holds the title technical CEO at GrackerAI.
A University College London study found listeners could accurately identify deepfakes only about 73% of the time, and with training and familiarisation with deepfake samples the accuracy rate improved by just 3.84%.
A joint information sheet released by the NSA, FBI and CISA ruled out traditional automated detection protocols that visualise evidence of manipulation in voice or video content, because these methods assume statistically significant traces of manipulation can be found and recorded, which is no longer necessarily the case.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Concrete incidents, uncited research
The two anchor incidents (Arup, KnowBe4) are specific, named and internally consistent, and the arithmetic on 15 transfers totalling ~$25m holds. Against that, every research and government reference is second-hand: no study titles, authors, dates or links, no identifier for the NSA/FBI/CISA information sheet, and a direct numerical tension between the ~73% UCL accuracy figure and the near-chance aggregation. Everything rests on one outlet and two vendor interviews, so the load-bearing detection claims cannot be independently checked from the supplied material.
Guidance exists, uptake unmeasured
There is real institutional movement to point at: reported CISA guidance naming FIDO2 and PIV hardware credentials as the MFA gold standard, recurring FBI advice on shared verbal passphrases, and a joint NSA/FBI/CISA position against trace-based detection. But adoption of the prescribed controls is described only as 'more fraud consultants now recommend' — no deployment counts, no enterprise rollout examples, no before/after fraud outcomes. Two documented incidents demonstrate attacker adoption of the technique, not defender adoption of the fix.
Prescription outruns its evidence
The incident reporting is not overstated — if anything the Arup facts are underplayed relative to the headline. The overstatement sits in the framing: a headline promising that a 'low-tech solution from the past may be your best defense' is supported only by two vendor executives' opinions plus general agency guidance, with no efficacy data, no failure modes and no cost. The article also reaches for the strongest available detection statistic ('closer to chance') while leaving it unreconciled with its own 73% figure, which inflates the rhetorical case.
Vendor-sourced, undisclosed interest
Both named experts sell into the problem they describe: Gupta is technical CEO at GrackerAI and Scobey is CTO at B2B cybersecurity firm S2i2, and the article carries no conflict disclosure while their commentary drives directly to buy-side conclusions about controls and credentials. The KnowBe4 anecdote is also supplied by a competitor-adjacent vendor executive. Layered on top is trade-publication incentive: a prescriptive 'old-fashioned fix' headline is engagement-optimised framing for a security-buyer audience.
One outlet, mixed verifiability
Confidence is capped by single-publisher, single-source coverage with no corroboration available in the cluster. The incident facts and the direct quotes are high-confidence within that limit; the research statistics, the agency documents and the scale claim about North Korean operations are not traceable to primaries, and one key statistic is internally contested. The truncated body also cuts off part of the recommended control's explanation.
Follow any of these and your For You feed starts watching them — no settings page required.
security
Gunra Goes Franchise: Conti's Leaked Code Now Ships With a Builder and an Affiliate Panel2 distinct publishers
product
White House lets vetted firms hack back and leaves liability blank for 60 days1 distinct publisher
build
AI-written snap7 scripts move the scarce resource in OT attacks from skill to exposure2 distinct publishers
invest
Washington licenses private hacking, and hands the contractor the liability1 distinct publisher