Science1 distinct publisher2 min readUpdated
A European Commission analysis found women candidates drew 71% of the worst abuse on X, and a different kind of it. Divide by reply volume and the gap stops being significant.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Perspective API returns one number per message, from zero to one, and what that number estimates is the chance a reader would find the comment rude enough to leave the discussion [3]. That is a serviceable instrument for asking whether a thread has gone bad. It is a weak one for the question the Joint Research Centre team ended up asking, because a reply bundling a sexist insult, a sexual threat, a reference to mental illness and a demand that the candidate quit politics sits on the same single axis as a reply calling a man incompetent [13][12].
The count that survived contact with human eyes is the interesting one. Of the roughly 2,700 highest-confidence toxic replies the team read by hand, 71% went to women [8][9]: about 1,900 replies against about 780 for the men, a little over two to one [1][2]. Part of that follows from attention, since the women in the sample posted less and were replied to more [7], and the adjustment that makes the gap vanish is a division, toxic replies over total replies [6]. Volume can explain a count. It cannot explain a vocabulary. Receiving more replies does not convert generic hostility into sexually explicit remarks and instructions to vacate political space [10][11], which is the material the authors connect to the literature on semiotic violence, harm done through language that works to delegitimise women as political actors [15].
The intensity result is narrower than it first reads. The team reports that 62% of toxic replies to women carried two or more abusive elements, and that 52% of those to men carried only one [14]. Read as complements, men sit at 48% with two or more, a gap of 14 percentage points [3], real but well short of what the 71-to-29 targeting split implies. The summary does not publish the men's two-or-more figure directly, which is the kind of asymmetry that makes a replication attempt guess.
Where this bites is reporting. Prevalence, the share of replies crossing a threshold, is the metric that scales to millions of messages, and it is the metric that came out flat here. Composition, what the abuse actually says, is what separated the sexes [10][12], and composition required hand-reading roughly 0.08% of the corpus [4]. A transparency regime built on a rate and a threshold will collect numbers that are accurate and almost empty of this finding. Concentration makes it worse: the JRC reports that women from left-leaning groups drew the higher volumes, matching what studies of Brazilian elections found [16], so an average across a national ballot spreads the load across candidates who were not carrying it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
A team at the European Commission's Joint Research Centre analysed more than 3.2 million responses on X to 88 candidates for the 2024 European Parliament elections.
The 88 candidates studied were from France, Ireland, Italy, Portugal and Spain.
The study used the Perspective API toxicity detection algorithm, whose score ranges from zero to one and estimates the likelihood a reader would perceive text as toxic, with toxicity defined as a rude, disrespectful or unreasonable comment likely to make you leave a discussion.
After accounting for differences in the overall number of replies per post, the gender difference in the proportion of toxic replies was no longer statistically significant.
Women candidates posted less than men and yet were addressed more; their accounts attracted more attention overall, toxic and benign alike.
Of those clearly toxic replies, 71% were directed at women candidates.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed method, single non-independent account
The quantitative spine is specific and internally checkable - 3.2 million replies, 88 candidates, five countries, a named classifier, ~2,700 hand-coded replies, 71%, 62%/52% - and the authors report a null result against their own headline, which is a marker of candour. But everything comes from one item written by the study's own team, with no linked paper, dataset, codebook, threshold value, effect sizes, or inter-coder reliability, and no independent replication in the cluster. The strongest content finding also rests on a severity-filtered ~0.08% slice, which caps how far the evidence can carry.
One research deployment plus regulatory pressure; no platform uptake shown
There is a concrete, disclosed deployment: an EU institutional research team ran a production toxicity classifier over millions of replies and built an additional severity metric on top. The article also situates this against live DSA obligations and a new EU directive, which is real regulatory demand. What is absent is any evidence that platforms, election authorities, or vendors have adopted severity- or content-aware gendered-abuse measurement - the article instead argues they should, which is advocacy rather than adoption.
Mildly overstated by subsample generalisation
Modestly positive rather than high. The article does the honest thing on the prevalence question, reporting that the normalised gender difference was not statistically significant, and its qualitative conclusion aligns with cited work from Sweden, New Zealand and the Netherlands. The overstatement is narrower: the 71% figure and the content contrast come from a top-severity slice of about 0.08% of replies, are not normalised by candidate gender counts or reply volume, and are then generalised into 'a different phenomenon' with no flag on that selection effect. The ideology finding is asserted without any figures at all.
Authors are EU institutional researchers advocating platform obligations
The piece is written by the researchers themselves, based at the European Commission's Joint Research Centre, and closes by endorsing EU regulatory instruments - the DSA and the new violence-against-women directive - while arguing platforms must go further on moderation and detection investment. That is a coherent institutional interest in findings that justify platform obligations, and no conflict-of-interest or funding disclosure accompanies it. This is a structural read of position, not an allegation of distortion; the reporting of an unfavourable null result runs against the incentive.
Single-source, self-reported, method partly unverifiable
Confidence is limited chiefly by source structure rather than by internal contradiction: one publisher, one item, written by the study team, with no paper, data, thresholds, reliability statistics, or platform comment available to check. The core numbers are consistent and the arithmetic derivations follow cleanly from them, so the direction of the finding is reasonably firm; the precise magnitudes and their generality are not.
product
France's under-15 ban failed on the age check, not the age limit1 distinct publisher
invest
France's under-15 ban dies on the arithmetic of age gates: to check minors, you check everyone1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
leadership
Daycare Does Not Break Children's Brains, And It Does Not Fix Economies Either1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 21, 2026