Skip to content

Build1 publisher3 min readPublished

Thai sentences cost up to 3.6 times as many Bedrock tokens as English in one architect's dated tests

Amazon Bedrock bills a Thai customer sentence at 2.4 to 3.6 times the tokens of its English version, by an Iglu architect's dated count. Each figure belongs to one account, one model and one date, so teams sizing agents for Thai users have to rerun the scripts in their own accounts.

The Engineer · Build desk

Photograph accompanying Thai sentences cost up to 3.6 times as many Bedrock tokens as English in one architect's dated tests
Photo: aboutamazon.com

What happened

  • In the four weeks of preparation, Bedrock's Thailand Region model list grew from 21 to 31, two new Claude models appeared and a cited documentation page changed its numbers.
  • On 2026-09-30, Claude Sonnet 5.5 and Opus 5.5 counted the Thai test sentence at 91 tokens against 39 for the English one, a ratio of 2.3.
  • The author found no published token figure for Thai from either AWS or Anthropic, and measured one instead.
  • The post traces three production surprises, a model that forgets, a bill estimated in English and throttling well below expected traffic, back to a single model call.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A budget priced from English test prompts undercounts Thai traffic even at the correct per-token rate, and the shortfall appears only after real customers start writing in Thai.
  • decision Teams moving to Sonnet 5.5 or Opus 5.5 have to recount even unchanged prompts, because the per-message count shifted by two tokens with no documented reason.
  • constraint An agent pays any per-call language overhead again at every step of its loop, so a Thai-facing agent's token budget grows with the number of steps it takes.

The test needs one script and no framework. It uses plain boto3 and the Converse API [11]. One customer sentence, written in four languages, goes to each model once with maxTokens set to 1. The script then reads usage.inputTokens from the response [15]. That field is the one Bedrock bills on. On 2026-09-04, CloudWatch's InputTokenCount for the same calls added up to exactly the same values [16]. I like this design. It reads the billed number and checks it against a second meter.

The published guidance covers English. Anthropic's glossary says a token "approximately represents 3.5 English characters, though the exact number can vary depending on the language used" [12]. The CountTokens API reference says token counting is model-specific [13].

The English test sentence is 89 characters and the Thai one is 95 [17]. The Thai is about 7% longer in characters [4], yet the token count more than doubles. On Sonnet 5.5 the English works out to 2.3 characters per token [5]. That is below the glossary's 3.5, but the 39-token count includes whatever fixed framing makes a bare "Hi" count 11 [19]. Eleven tokens is a lot to pay for a greeting.

The 5.5 models also moved the baseline. Every count was exactly two above Claude Sonnet 5 [20]. That puts Sonnet 5 at 9, 37 and 89 tokens [2] and its Thai-to-English ratio at 2.4 [3]. The author guesses two extra fixed framing tokens on an unchanged tokenizer. The author also says that guess is not documented [20].

A token benchmark transfers only if your model, your text and the counting rules match the ones measured. The author tested the last condition by repetition. The main table was measured in Singapore on 2026-09-25 and came back identical on 2026-10-01. The Haiku 4.5 and Nova rows also held on 4, 14 and 21 September, from both Singapore and Bangkok [18]. The counts held across those dates while the model list and one documentation page changed in the same month [5]. The documentation was read on 2026-09-14 and again on 2026-09-30 [6]. Every measurement came from one AWS account [4].

The post defines an agent as "a loop of model calls that decides the next step and calls tools" [10]. "Whatever is true for a single call is true for every step of that loop, just more often," the author wrote [10]. The running example is an airline's customer chat with one Thai customer [11]. I think that is the right order of work for a team shipping to non-English users: count tokens on the exact call you will ship, in your own account, before you pick a framework. For an English-only product the language ratio matters less. The two-token drift between model versions still applies.

The author held the talk to a strict rule: "no number goes on a slide unless I measured it myself, and no rule goes on a slide unless I can point at the page it comes from" [3]. The post also asks readers to push back: "if your numbers differ from mine, don't assume one of us is wrong. Re-run the scripts and tell me what you got" [7]. The code, the checklist and the dated measurements are in the repo that accompanies the talk at AWS Community Day Thailand in Bangkok on 3 October 2026 [1].

What to watch

  • Whether AWS or Anthropic publish a characters-per-token figure for Thai or other non-Latin scripts.
  • Whether the two-token offset on Sonnet 5.5 and Opus 5.5 gets documented, or later Claude models on Bedrock change the counts again.
  • Reruns of the repo scripts from other accounts and regions that report counts different from the author's.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories