Product2 publishers3 min readPublished
Claude Opus 5.5 writes 'this matters' 116 times as often as human writers, Graphite finds
Graphite counted 13,000 phrases that AI models use at least twice as often as human writers, with a different set for every model version. Editors cleaning AI drafts need a phrase list tied to the model that wrote them, rebuilt at each release.
The Product Desk · Product desk

What happened
- Opus 5.5 writes 'why X matters' 92 times as often as the human-written samples Graphite compared it against.
- OpenAI's Astra uses corrective framing such as 'not simply X' and 'rather than relying on X' more than 100 times as often as human writers, according to Graphite.
- Labs have cut em dashes: Astra now uses them 88% less than human writers, and Gemini 3.1 Pro has almost dropped them entirely.
- Graphite used 10,000 articles published before ChatGPT's release as its human baseline and had each AI model rewrite them from summaries.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision A phrase checklist holds only for the model version it was measured on, so teams publishing AI-assisted copy need a record of which version drafted each piece.
- cost Whoever owns the house style pays again at each model release, since Graphite finds old tells disappearing while new ones keep the total roughly level.
- constraint Opus 5.5 overstates and Astra hedges, so one shared edit pass for AI copy cannot correct both models' drafts in the same direction.
Give Claude Opus 5.5 and OpenAI's Astra the same brief and, going by Graphite's counts, the editor gets back two different cleanup jobs. The Opus draft qualifies its claims less often and is far more likely to use superlatives [9]. It has dropped "it's not X, it's Y" and now writes that something "is more than an X, it's a Y" [10]. The Astra draft hedges, saying an action "may provide" or "can provide" a benefit, and finds "another dimension" in whatever it covers [11].
The labs describe their writing upgrades in general terms. Anthropic said Opus 5.5 "communicates more naturally than prior models" [20]. OpenAI told users of the GPT-6 versions of Sol and Luna to expect more clarity, less jargon and fewer odd turns of phrase [21]. A team reading those release notes could fairly assume its existing checks for AI copy still hold. Graphite measured something narrower. The famous tell did go: Opus 5.5 uses em dashes 99% less often than Opus 5 [14]. The total number of tells held mostly steady anyway, according to Graphite [24]. "They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own," Greg Druck, Graphite's chief AI officer, told TechCrunch [18].
The labs are also moving in different directions. "It turns out that Claude models are actually getting closer to the human word distribution over time," Druck said. "And for the GPT models, it's getting further away." [17] He doubts any lab can finish the job. "These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through," he said [19].
Treat the exact multiples as approximate. Graphite is a marketing company, and Gizmodo advised taking its study "with a grain of salt" [1]. Gizmodo puts "dependable" at 26 times the human rate and calls "this matters" the biggest single tell [6][8]. TechCrunch has "dependable" at 23 times and names it Opus 5.5's biggest tell [7]. Astra's 275x figure for "does not establish" is measured against Claude's output, not against human writing [13]. Every ratio comes from the one rewriting task in Graphite's design [16]. Gizmodo describes an AI-detector business "somehow worth millions of dollars" [23]; neither report says how those detectors score against the newer phrases.
At a 2x threshold the list runs to 13,000 entries [2], far more than an editor can check by hand. The usable part is the short head of the list for one model version. For Opus 5.5, I'd check "this matters," "why X matters," "looking ahead the" and "is more than an X, it's a Y" [3][4][5][10]. For Astra, the list starts with "not simply," "rather than relying on" and "may provide" [12][11]. I'd keep a list like that for each model in use, knowing it expires at the next release [18] and rests on one firm's single-task study [1][16].
Sort each AI-assisted draft on two things: whether the editor can name the model version that wrote it, and whether the phrase list was built on that version. Two yeses mean the list applies. A named model with an older list catches the previous version's habits, the way an em dash check now finds little in Opus 5.5 output [14]. When nobody can name the model, a per-model list has nothing to match. The edit then falls back on the habit Graphite found across models, contrast-heavy constructions [22], plus a read for whether claims are overstated or over-hedged [9].
What to watch
- Whether Graphite publishes its per-model ratios in full, settling the 23x versus 26x gap on 'dependable'.
- The next Opus and Astra releases, and whether 'this matters' and corrective framing fall away the way em dashes did.
- Any published test of commercial AI detectors against the newer per-model phrases.