Skip to content

Build1 publisher3 min readPublished

Claude's new tokenizer and GPT-6's 272K price cliff break old LLM cost models

Claude 4.7 emits about 30 percent more tokens for the same text and GPT-6 bills roughly double above 272K input tokens, a dev.to digest reports. Budget checks built on old token counts now undercount, so prompt size needs a hard cap enforced in code.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Claude's new tokenizer and GPT-6's 272K price cliff break old LLM cost models
Generated illustration

What happened

  • Claude Managed Agents charge $0.08 per session hour on top of token fees, and web search is billed per search.
  • Claude 4.6 and later bill the full 1M-token context at standard rates, so a 900K request costs the same per token as a 9K one.
  • Google's Gemini 3.1 Pro tiers at 200K tokens instead of 272K, charging $4 and $18 above that line.
  • According to the digest, processing modes multiply the same rates: Batch is half price, Fast is double and Ultrafast is six times.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost An agent session left open around the clock runs up $57.60 a month in session fees before a single token is billed.
  • decision A router that spreads traffic across providers now has to count tokens before choosing, because the cheapest provider changes with prompt length.
  • exposure A Fast-mode switch flipped to answer a latency complaint can double spend with no change to any model string, so the mode setting needs the same cost review as a model swap.
  • cost For a GDPR-scoped deployment on a regional endpoint, moving to any model released on or after 5 March 2026 adds a 10 percent uplift the older model did not carry.

A token is the unit every vendor prices in. According to the digest, Anthropic changed the size of that unit. Sonnet 4.6 and earlier use the old tokenizer [2]. Claude 4.7 and later produce about 30 percent more tokens for the same text [1]. The author summed up an upgrade from Sonnet 4.6 to a 5.x model this way: "The price per million went down. The number of millions went up." [11]

The post lists three things that break in a typical .NET codebase: cost projections built on token counts measured on older models, middleware that counts tokens locally to enforce a budget before the call, and chunking logic that packs a prompt to a fixed token target [12]. I'd check the chunker first, because it fails without an error. A packer that fills a fixed target now fits about 77 percent of the text it used to, since the same text costs 1.3 times the tokens [1]. The budget middleware has the opposite fault. It approves calls whose real count runs about 30 percent above what it measured [1]. "Re-measure. Do not assume," the post says [13].

OpenAI's problem is a threshold [3]. The digest says a request that goes one token over 272K input tokens costs about twice as much per token [14]. A RAG pipeline whose chunk count varies with the query therefore has no stable cost per call, the post adds [14]. Taken literally, a 272,001-token request bills like about 544K tokens at the base rate [3]. "You need a hard cap in code, not a hope," the author wrote [15].

I agree, and I'd be specific about where it goes. The cap belongs in the prompt assembler, after retrieval and before the client call. It has to count with the target model's tokenizer. The Claude change shows that a count taken with one tokenizer is wrong for another, even inside a single vendor [1]. When the assembled prompt crosses the provider's line, drop the lowest-ranked chunks until it fits. Cutting the tail of the string is easier to write and removes whatever happened to be appended last.

Claude's flat long-context pricing is the easiest of the three to build against. For long-context work such as document analysis and whole-repository reads, the author wrote that "that single line is worth more than a headline benchmark score." [19] On those models the cap is a budget control, because there is no cliff to stay under [5].

The digest does not say whether processing-mode multipliers compound with OpenAI's above-272K rate. If they do, a Fast-mode request past the line pays four times the base input rate [4]. Every vendor figure in this piece comes from that one dev.to digest [18].

What to watch

  • The 3.8 Flash model's introductory rate runs through 31 December 2026, after which the digest says it doubles to $1.50 and $7.50.
  • Whether OpenAI's own price page shows Fast or Ultrafast multipliers compounding with the above-272K input rate.
  • A vendor-published token ratio by content type, such as code against prose, would replace the digest's single 30 percent figure for the Claude 4.7 tokenizer.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories