Security1 publisher2 min readPublished
Surfshark's agreeable bot personas evaded detection 62% of the time
The company showed 1,722 people a mix of human and AI-written comments in a simulated feed. The accounts they caught most reliably were the hostile ones that awareness training already teaches staff to distrust.
The Watch · Security desk

What happened
- Negative, aggressive personas were flagged 50.2% of the time, while positive, agreeable ones were flagged 38%, a 12-point gap in favour of the polite account.
- Bots that leaned on emojis were caught over 60% of the time; bots that kept their language plain were caught 35% of the time.
- Detection peaked among the youngest participants and was lowest among over-50s, who were also the most likely to flag a real person as a bot.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- constraint Training built around hostile or obviously odd accounts is aimed at the persona participants already caught half the time; the accounts that agree with the target slip straight past it.
- exposure Any process whose final check is a person judging whether an account is real, including user reporting, fails against the polite persona close to two times in three.
- capability An operator can cut the flagging rate against its accounts by writing plainly and staying agreeable, with no change to infrastructure or volume.
- contradiction Surfshark's security conclusion runs from a missed label to clicked links and private chats, while the measurement stops at whether a participant labelled a comment correctly.
Converted into miss rates: positive, agreeable personas went unflagged 62% of the time against 49.8% for the hostile ones, and the miss rate across the whole test was 60% [1][2][3].
User reporting inherits those numbers, and so does the awareness training that tells staff to watch for accounts behaving strangely. Training describes the persona participants caught half the time [4].
On the attacker's side the adjustment is free. Surfshark said an account that stops overusing emojis becomes twice as hard to identify [7]. On its own published figures the detection rate falls by a factor of about 1.7 and the unflagged share rises by about 1.6 [4].
"Often, AI-powered accounts used in sophisticated manipulation campaigns are agreeable, logical, or simply unremarkable. They can support real users' opinions and just inflate the number of comments in the discussion to make a minority view look like everyone feels the same way. Seldom do people report such accounts," said Luis Costa, Research and Insights Lead at Surfshark [15].
The error ran both ways on the harder subject. On women's rights participants missed more bots than on a light topic such as pineapple on pizza, and they also wrongly accused more real humans [8]. Immigration reordered the result: negative bots were still the easiest to catch there, but positive bots were harder to spot than in any of the other three topics [9]. Neutral bots sat near the bottom for detection across most topics [10].
Threads users had the highest detection rate, on what Surfshark notes is a small sample, with X users close behind and TikTok and Facebook users well behind both [12]. Participants who use social media almost all the time caught close to half the bots they saw, while people who do not use it at all caught roughly a third [13].
The test measured whether participants could separate human comments from AI-generated ones in a simulated social media setting [1]. Surfshark's researchers argue the agreeable, logical bot is a bigger privacy and security concern than the confrontational one precisely because it does not set off suspicion [14]. The step from a missed label to a private chat, a handed-over detail or a clicked link is their argument from the detection gap, and the study is Surfshark's own [1].
What to watch
- Publication of the model and prompts behind the AI-generated comments would show whether the 38% detection figure travels to newer generators.
- Per-platform sample sizes from Surfshark: it flags the Threads result as small-sample, and X sits close behind it.
- A replication that measures a downstream action, such as a link click or a move to private chat, instead of a label.