Skip to content

Build1 publisher3 min readPublished

Claude Opus 5 cuts Russian text into nearly three times as many tokens as English

Habr user donseo's test of 11 tokenizers found Claude Opus 5 turns Russian into 2.96 times the tokens of the same English text. Anthropic bills $5 per million input tokens in either language, so Russian on Opus 5 costs nearly three times as much to send.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Claude Opus 5 cuts Russian text into nearly three times as many tokens as English
Generated illustration

What happened

  • OpenAI's o200k tokenizer, used by GPT-5.x, needed 1.19 times the English token count for Russian, down from 2.09 times on GPT-4's cl100k.
  • YandexGPT 5 Lite and GigaChat 3 were the only two of the 11 tokenizers that needed fewer tokens for Russian than for English, at 0.91x and 0.96x.
  • Anthropic says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text in any language.
  • Petrov et al. measured the GPT-4 tokenizer at about 3x for Arabic and 2.6x for Bulgarian at NeurIPS 2023, with some languages reaching 15x.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Russian replies on Opus 5 pay the multiple at the $25 output rate, roughly $74 per million English-equivalent tokens against $25 for the English original.
  • constraint At 2.96x the context window holds about a third as much Russian as English, and per-minute token caps stop a Russian team after less work.
  • contradiction Anthropic's guide of about 4 characters per token would count Opus 5 Russian, measured at 1.03 characters per token, at roughly a quarter of its real token count.

The premium is set before the model trains. A tokenizer builds a fixed vocabulary from a corpus, and sequences that are frequent in that corpus get tokens of their own [13]. Common English words come out as one token each [14]. A Cyrillic word breaks into fragments, and in GPT-4's cl100k_base some Cyrillic letters take two tokens apiece [14]. Petrov et al. wrote that "tokenizers are heavily influenced by the biases of the corpus source" [15]. They also wrote that "the unequal treatment of languages arises at the tokenization stage, well before the language model sees any data at all." [16]

Opus 5 starts behind in English. According to the Habr test, Opus 5 English averaged 3.25 characters per token, against 5.07 for o200k, the tokenizer behind GPT-5.x [10]. Opus 5 Russian came in at 1.03, close to one token per letter [10]. On the same English text, Opus 5 needs about 1.56 times the tokens o200k does [1]. Multiply by Opus 5's 2.96x Russian ratio and divide by o200k's 1.19x, and Russian on Opus 5 comes to about 3.88 times the tokens of Russian on o200k [2]. The write-up rounds the same comparison to 3.9x against GPT-5.5 [12]. It notes that reposts present that figure as Russian against English, which overstates the language gap and hides the model gap [12].

The Habr table describes four text pairs. Their Russian side totals about 2,500 characters across a technical article, a support chat, a business notice and commented Python [3]. The write-up says prose is where the premium bites hardest [24]. On Opus 5 the technical article went from 327 English tokens to 872 Russian, about 2.67x [9][3]. The Python sample went from 231 to 424, about 1.84x, the smallest gap of the four [11][4]. Both sit below the 2.96x aggregate [6]. Whichever way the aggregate was weighted, the support chat or the business notice must have run above it [5]. A code-heavy workload would see less than the headline ratio, and a support desk could see more [4][5].

The method carries over more cleanly than the numbers. The open tokenizers were run on a local machine with tiktoken and Hugging Face [4]. Closed models had to go through their APIs: the script submitted each text together with a brief fixed prompt, submitted the prompt by itself, and subtracted one reported token count from the other [4]. That needs nothing beyond the usage count the vendor already returns, so it can run against a team's own text in each language it serves [4]. The code and the corpus are on GitHub under MIT [5].

In my view a published ratio, this one included, is a first estimate drawn from someone else's text. A per-language budget needs counts taken from the team's own traffic. Petrov et al. put the money consequence in one sentence: "the tokenization premiums discussed in Section 4 directly map to cost premiums." [23]

What to watch

  • Whether Anthropic's next tokenizer revision lowers Opus 5's 2.96x Russian ratio the way o200k cut OpenAI's from 2.09x to 1.19x.
  • Reruns of donseo's public script on languages Petrov et al. found costlier than Russian, such as Arabic, against current tokenizers.
  • Whether Anthropic revises its 4-characters-per-token guidance for models on the post-4.7 tokenizer.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories