Build1 distinct publisher3 min readPublished
The old flat rate became a peak rate that blends 3.6x higher, with a half-price window covering seventeen hours, so inference cost modelling now depends on which UTC hour the traffic lands in.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The index behind these numbers blends three input tokens to one output token, and that ratio carries more weight than it looks [2]. V4 Pro's old flat rate charged twice as much for output as for input; the peak rate charges three times [5][6]. Input-only traffic therefore goes up 3.03x. Output-only traffic goes up 4.55x [2]. At a 3:1 blend it lands on 3.64x, which is where the report's "roughly tripled" comes from [1][9].
Peak covers 01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours, leaving seventeen at exactly half [6][3]. Exactly half is the well-built part. There are no tiers or commitments underneath it, so a cron window and a multiplier of two describe the entire schedule [7]. According to the report the change appeared on the pricing page at 16:00 UTC on August 16 with no notification to callers [5][9], which makes that page a dependency with no change feed.
Both increases came from the same lab. V4 Flash took the same treatment the same day, 3.8x blended, with its output rate up 4.71x [15][6]. The floor line did not move, because it reads the cheapest listed price across every host serving the base model rather than the lab's own page [16]. DeepInfra's $0.09 / $0.18 blends to $0.1125, which is the $0.113 the report prints [4][14]. Two things have to hold before that number belongs in your model: GPQA Diamond at 70 has to proxy the work you actually send [13], and DeepInfra has to have the throughput and latency you need at that price. The report's own line is the right caution, that a lab's list price is a policy while a reseller's is a margin [17].
One ratio in the report does not reconcile. It puts GPT-4-class capability 37x under the flagship ceiling on September 1 [18], but the ceiling it names is Claude Opus 5 at $10.00 blended and the floor it prints is $0.113, which are 88.9x apart [19][14][5]. The two pairs of numbers do not match, so 37x describes some other comparison in the report, not this ceiling and this floor.
Sol's cut is the larger headline and the softer input. Minus 29% is the first time in the index's history that a constituent lowered the price of a model it had already shipped [10]. The sentence holding it up says the promotional rate is available at least through November 21, 2026, and OpenAI names no date on which it goes back up [11]. The sentence sets a minimum duration for the rate, not a date when it reverts.
The month netted minus 5.7%, and the index is down 9.4% since February 23 [3], so a spreadsheet tracking only the average saw a quiet August, helped by three flagship handovers where each successor inherited its predecessor's list price and moved nothing [4]. Underneath the average, two of ten flagships now print a number that is only the top of a range [12], and the flagship spread narrowed from 21x to 13x from both ends at once as Sol's cut pulled the top down and V4 Pro's peak rate lifted the bottom onto Mistral Large 3 at $0.75 [19]. An assumption that held across 40 daily readings [1] now needs a UTC clock and an expiry date to express.
Ranked by verification strength, evidence, and original report placement.
Until 16:00 UTC on August 16, DeepSeek V4 Pro billed a single flat rate of $0.435 per million input tokens and $0.87 per million output tokens.
DeepSeek's pricing page then split V4 Pro into a peak rate of $1.32 in / $3.96 out for 01:00-04:00 and 06:00-10:00 UTC, and an off-peak rate for every other hour at exactly half, $0.66 in / $1.98 out.
The off-peak blended rate of $0.99 is 82% above the old flat price; the author describes the change as a 3x increase with a discount window attached rather than a discount layered on the old price.
The floor held because it reads the cheapest listed price for a base model across every host that serves it, and DeepInfra kept serving V4 Flash at $0.09/$0.18, 5.9x below the lab's own peak rate.
The index is the equal-weight average of ten flagships' blended price per million tokens, blended 3 parts input to 1 part output, using list prices as printed on each vendor's own pricing page.
Three flagship handovers in August (Muse Spark 1.1 to 1.2, Grok 4.5 to 4.6, GLM-5.2 to 5.3) moved the index nothing, because each successor kept its predecessor's list price.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 5, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
DeepSeek open-sources the harness, then raises the price of the model4 distinct publishers
product
Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.63 distinct publishers
build
Epoch's first-place ranking for GPT-6 Astra rests on a single coding score2 distinct publishers
invest
Compute scarcity meters the model OpenAI says can fill out forms at superhuman speed1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One observer, self-consistent in parts
Every number traces to a single self-published report written by the person who also maintains the index, and no pricing page is archived or screenshotted alongside it. Where the figures can be checked against each other they hold up: DeepSeek's peak blend of $1.98 against the old $0.54375 gives the 3.64x claimed, and DeepInfra's $0.09/$0.18 blends to the $0.113 floor printed. Where they cannot — the $4.14 index level, the -9.4% since February, the assertion that no constituent had ever cut a list price — the reader has the author's word and no table. And one figure fails its own test: the 37x floor-to-ceiling gap sits in the same report as a $10.00 ceiling and a $0.113 floor, which are 88.9x apart.
Live on the price cards, silent on the meter
The repricings are concrete, dated and already in force: DeepSeek's split at 16:00 UTC on 16 August, Sol's cut on the 21st, two $10/$50 flagship launches in the first three days of September, and a third-party host holding its Flash price through all of it. Volume is the missing piece: there is no token count, no share of traffic falling inside the seven-hour peak windows, no sense of how many callers can batch and how many cannot. The tripled-bill figure is addressed to a reader rather than drawn from a workload, so the price moves are documented while their cost impact is only asserted.
Headline rides a methodology choice
'One lab tripled its rate' is true of the number the index chose to track. A caller whose traffic never touches the two windows pays 82% more, not 3x, and the report says so two paragraphs down while keeping the tripling in the title. Set against that, the restraint is genuine: no turn is called on a single reading, the sentence propping up Sol's cut is labelled contingent rather than banked, and the closing advice is to write the tier next to the price. The overstatement is in the framing of one figure, not in the body.
The index is the author's own product
The report and the thing it reports on belong to the same person, and the findings are phrased in the currency of that series: first list-price cut by a constituent, largest single move on record, first time above the February reading. A month that breaks the author's previous headline is better copy than a month that confirms it, and the piece leans into exactly that. No lab is quoted, no vendor relationship is disclosed in either direction, and nothing here reads as favourable to a particular seller — the pull is toward a more eventful index, not toward DeepSeek or OpenAI.
Trust the price cards, hold the index level
Rates printed on a page at a stated hour are the kind of thing one attentive observer usually gets right, and the internal arithmetic here mostly bears that out. This settles at a coin flip because the parts a reader most wants to lean on — the index level, the month's net, the historical firsts — are visible only to the author, and the single cross-check available on his September summary does not come out even.