Skip to content

Build1 publisher3 min readPublished

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

The Engineer · Build desk

Illustration accompanying Anthropic prices its newer Sonnet a third below Sonnet 4.5

What happened

  • GPT-5 takes the maths and multimodal scores, at 94.6 percent on AIME 2025 without tools and 84.2 percent on MMMU, and accepts 272K input tokens inside a 400K window against Sonnet 4.5's 200K.
  • Anthropic's 77.2 percent was produced with parallel test-time compute, and its 61.4 percent OSWorld score and 30-hour autonomous coding run come from its own launch testing.
  • Anthropic shipped Claude Sonnet 5 on 30 June 2026 at $2 input and $10 output per million tokens, below both of Sonnet 4.5's list prices.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction The case for these two as the value tier rests on them being one generation old, but Anthropic's price list makes the newer Sonnet the cheaper Anthropic option, so the argument only holds on the OpenAI side.
  • cost The headline saving belongs to a 5:1 input-to-output mix; output-heavy traffic narrows it to 37.5 percent, so the finance case moves with the token ratio of your own workload.
  • decision Because the SWE-bench lead was set with parallel sampling, a team choosing on coding quality needs its own single-attempt run before routing agent work by vendor.
  • exposure Any team that evaluated on GPT-5 Pro and budgeted for base GPT-5 is understating output cost by 12x, and the error surfaces on the invoice.

The worked example checks out at list prices. GPT-5 at $1.25 per million input tokens and $10 per million output [6] costs $12.50 plus $20 on a month of 10 million input and 2 million output, so $32.50 [9]. Sonnet 4.5 at $3 and $15 [7] costs $30 plus $30, so $60 [9]. The dev.to comparison puts the saving at about 46 percent and scales it linearly, to roughly $325 against $600 at ten times that volume [8]. The division gives 45.8 percent [24].

How much you save depends on the token mix. A 5:1 ratio loads the bill onto the side where GPT-5 is furthest ahead, since input is $1.25 against $3 while output is only $10 against $15 [6][7]. Move the same 12 million tokens to an even split, 6 million each way, and Sonnet 4.5 costs $108 against GPT-5's $67.50, a saving of 37.5 percent [10]. A workload that emits long diffs from short prompts sits nearer that second figure.

Coding quality needs the same treatment. Anthropic's 77.2 percent on SWE-bench Verified was produced with parallel test-time compute, which means more than one attempt per task [1]. The comparison itself calls a 2.3 point gap over GPT-5's 74.9 percent noise on a single leaderboard [5][2]. It rests the coding case on Terminal-Bench 2, where Sonnet 4.5 scores 50.0 and GPT-5 43.8 on multi-step terminal work [3]. That is a 6.2 point gap [4], and it comes from a public leaderboard rather than a vendor page [3].

Anthropic also reports 61.4 percent on OSWorld, and describes Sonnet 4.5 sustaining autonomous multi-step coding for more than 30 hours in its own launch testing. dev.to labels the 30 hours a vendor claim and not an independent measurement [16].

GPT-5 Pro lists at $15 input and $120 output per million tokens [11]. An evaluation run on Pro with a budget written for base GPT-5 is off by 12x on output [12].

On the other axis, GPT-5 posts 94.6 percent on AIME 2025 without tools and 84.2 percent on MMMU [13]. It accepts 272K input tokens inside a 400K total window against Sonnet 4.5's 200K [14]. The dev.to write-up calls the context difference academic for most coding work, because retrieval beats stuffing a monorepo into a prompt. It stops being academic, the write-up says, when you reconcile several large documents in one pass [15]. Its recommendation is to pick Sonnet 4.5 for code inside an agent and GPT-5 for high token volumes, image-heavy inputs or maths-flavoured reasoning [21].

Anthropic's half of that split has a price problem. Anthropic released Claude Sonnet 5 on 30 June 2026 at $2 input and $10 output per million tokens [17], a third below Sonnet 4.5 on both lines [18]. Keeping Sonnet 4.5 for coding costs more per token than using the current model, the opposite of what a value tier defined by being one generation behind would predict [19]. dev.to reports Sonnet 5 at 63.2 percent on SWE-bench Pro against GPT-5.5's 58.6 percent [17]. But the Sonnet 4.5 number above is SWE-bench Verified, and the material does not put the two Anthropic models on the same benchmark [23].

These figures transfer only if your tasks look like SWE-bench Verified patches or Terminal-Bench 2 shell sessions and your traffic runs near five input tokens per output token. All of it comes from one comparison, last verified 1 September 2026 [20], with the headline scores reported by Anthropic and OpenAI [1][2].

What to watch

  • An independent SWE-bench Verified run for Sonnet 4.5 at single-attempt settings would show how much of the 2.3 point lead over GPT-5 came from parallel sampling.
  • Published Sonnet 5 and GPT-5.5 numbers on SWE-bench Verified or Terminal-Bench 2 would make the current generation comparable to this pair.
  • Caching and batch discounts sit outside the comparison, which uses list prices only, and either vendor's discount schedule would move the 46 percent.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories