Skip to content

Build1 publisher3 min readPublished

The Tokenizer Is Your Real Price List, Not the Per-Million Rate Card

Two models can quote identical per-token prices and still bill differently for the same string. The split is decided by merge tables you did not train and cannot assume.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying The Tokenizer Is Your Real Price List, Not the Per-Million Rate Card
Photo: aclanthology.org

What happened

  • If you build with LLMs, you pay by the token, not by the word and not by the character.
  • A rough rule of thumb for English: 1 token is about 4 characters, or about 0.75 words, so ~100 tokens is about 75 words.
  • A page of English prose (about 500 words) is roughly 650-750 tokens.
  • The 500-word page figure implies roughly 1.3 to 1.5 tokens per English word.
  • Nearly every major model today, including GPT, Claude, Gemini, Llama and Mistral, uses some flavor of Byte-Pair Encoding (BPE) or a close relative.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

The unit you are billed in is not the unit you write in. A dev.to explainer on tokenizers makes the consequence explicit: LLM APIs charge per token, not per word or character [1], and the same short string may cost four tokens on GPT and three on Claude [14].

That is a 33 percent gap on one string [15], created by nothing except which merge table ran. Nearly every current major model, including GPT, Claude, Gemini, Llama and Mistral, uses Byte-Pair Encoding or a close relative [5]: a compression trick from 1994 that was adapted for language models in 2016 [6]. BPE counts adjacent pairs of units, merges the most frequent pair, repeats until it hits a target vocabulary size, and records each merge in order so the tokenizer can replay them deterministically at inference [7]. The important part is what that implies about your bill: training frequency decides the splits. Common words collapse to one token, rare ones fragment, so "the" is a single token while "tokenization" may be two or three [8].

Vocabulary size is the other lever. A bigger vocabulary means fewer tokens per sentence but a larger embedding table and more memory [10]. GPT-2 learned roughly 50,000 merges; o200k_base uses roughly 200,000, which the post credits as a large part of why newer models are more token-efficient per word [11]. Each provider trained its own tokenizer on its own data, with its own vocabulary size, merge tables, and rules for whitespace and non-Latin scripts [16].

This is where non-English workloads stop being a rounding error. Modern tokenizers start from the 256 raw byte values, so nothing is ever an unknown word, but a character outside the well-merged region falls back to several byte-level tokens [9]. The post reports that a Spanish chatbot costs more than the English one [23]; it gives no multiplier, and you should not invent one. The same fragmentation hits content you may think of as English: long runs of digits break into several tokens, and JSON spends tokens on braces, quotes and keys [13]. Measured counts from tiktoken put "Hello, world!" at 4 tokens, "tokenization" at 2, and a nine-word pangram at 10 [12].

So the rule of thumb, 1 token to about 4 characters or 0.75 words, with a 500-word page landing around 650 to 750 tokens [2][3], is an English average, roughly 1.3 to 1.5 tokens per word [4]. It is a sanity check, not an estimate. The estimate has to be measured: OpenAI's tiktoken is open source and gives exact local counts, with cl100k_base at about 100k vocabulary on older models and o200k_base at about 200k on newer ones [17]. Anthropic does not publish its tokenizer, and directs you to POST /v1/messages/count_tokens, which takes the same shape as a real request, returns the input token total, and is free to call [18]. Google exposes countTokens [19]. Counts do not transfer between models [20].

Two things to watch. The post asserts that newer Claude models use a tokenizer that can produce about 30 percent more tokens, though the passage is truncated and never names the baseline [21]; if that holds, an identical per-token rate is a 30 percent higher input bill [22]. And run your own prompt corpus, in every language you ship, through each provider's counting endpoint before you commit to a feature price [25].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories