Science1 publisher2 min readPublished
AI-written phishing fooled people 28% of the time in a BYU test, against 21% for human-written
In the same experiments, participants identified the author correctly 52% of the time. The strongest single lever the BYU team reports is a mention of a co-worker, worth 2.3 times the clicks of a generic sender.
The Scientist · Science desk

What happened
- BYU researchers examining AI-generated deception across text messages and emails found AI-written messages fooled people 28% of the time, against 21% for human-authored messages.
- Messages that mentioned a co-worker were 2.3 times as likely to be clicked as messages that appeared to come from a generic organization.
- The team reports that company websites, LinkedIn profiles, social media accounts and organizational directories supply enough material to build a highly persuasive message.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint Content-based detection is weak here: a training program that scores employees on spotting clumsy phrasing is testing a discrimination these participants performed at 52%.
- decision Security teams weighing another recognition module against a mandatory second-channel check for anything unexpected now have a number that favors the check.
- exposure Staff whose employer and job title are public sit at the study's most persuasive configuration, and public information is all an attacker needs to get there.
- cost If Hansen is right that AI cuts the effort of personalizing a lure, defenders pay in volume, and per-message success rates will not tell them how much.
Seven percentage points separate the two arms: 28% for the AI-written messages, 21% for the human-written ones [1][1]. In relative terms that is about a third more clicks per message [2]. It also leaves roughly seven in ten AI messages ignored [4]. Across the comparisons, the researchers report that AI matched or beat human authors at producing a click 80% of the time [2].
For anyone building awareness training, the identification result matters more than the click rate. Asked whether a person or a model had written a given message, participants were right 52% of the time [4]. On a two-way judgement, chance is 50%, so they were two points above guessing [3]. That was a question about authorship, not about whether the message was a scam. Still, a program that trains staff to judge a message by how it reads is asking for a discrimination these participants made at 52%.
What moved the click rate was the detail about the target. A message naming a co-worker was 2.3 times as likely to be clicked as one appearing to come from a generic organization [3], and messages carrying job-related details, a workplace or an occupation, were the most convincing the team tested [6]. The inputs are ordinary: company websites, LinkedIn profiles, social media accounts, staff directories [5].
A click in a study is a click. Whether it becomes a compromise depends on what sits behind the link, and the published account measures clicks, across text messages and emails [13]. It gives no participant counts and no confidence interval for the seven-point gap [15], so treat the size of that gap as provisional and the direction as the finding. The same account carries the figure that 90% of corporate hacks originate from spear phishing, unattributed [8].
The economic claim sits apart from the click numbers. "AI can reduce the time and effort required to create personalized spear-phishing messages, thus making them more effective and more common," said Derek Hansen, the BYU cybersecurity professor who led the work [9][7]. A per-message success comparison cannot test that; what could is time per message, or messages per hour.
Jerson Francia, a BYU cybersecurity doctoral student and co-author, points the defense at the channel instead of the text, recommending verification of unexpected messages through a separate, trusted route before anyone responds or clicks [7][12]. "We need to more thoroughly verify the authenticity of messages through other means, rather than relying solely on the content of the message itself," Francia said [11].
What to watch
- The peer-reviewed papers, with participant counts and a confidence interval for the seven-point click-rate gap.
- A field test in a live corporate mail environment that follows clicks through to compromise rates.
- Any measurement of time or cost per personalized message. Hansen's volume claim turns on that, and this comparison does not test it.