Published Build3 min read
DeepSeek's "off-peak discount" raises every V4 price, with cache hits up 150%
The cheaper of DeepSeek's two new tiers still costs more than today's rates on every line item. Anyone who sized a V4 budget on current prices has three days to redo the arithmetic.
Written for builders.See today for builders

What happened
- DeepSeek will raise API prices for its V4 models on August 16 and introduce a two-tier rate card intended to move flexible computing jobs outside its busiest hours.
- The new prices take effect at 16:00 UTC on August 16, 2026.
- DeepSeek said off-peak rates will sit 50% below peak pricing; that framing describes the gap between the two new tiers.
- Every off-peak rate in the new table remains higher than the corresponding price DeepSeek charges today.
- DeepSeek defines peak hours as 01:00-04:00 UTC and 06:00-10:00 UTC, totaling seven hours each day; the other 17 hours are off-peak.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
DeepSeek will move its V4 API to a two-tier peak/off-peak rate card at 16:00 UTC on August 16, 2026 [1][2], and according to RuntimeWire every rate in the cheaper tier sits above what the company charges now [4]. The practical effect is that the discount is measured against a new, higher ceiling rather than against your current invoice, so a V4 budget built on today's numbers is already stale [3][4].
DeepSeek said off-peak rates will sit 50% below peak pricing [3]. That is a statement about the distance between the two new tiers, not about the change from today. On V4-Flash, current prices are $0.0028 per million cache-hit input tokens, $0.14 per million cache-miss input tokens and $0.28 per million output tokens [6]. The new off-peak figures are $0.007, $0.22 and $0.66, which RuntimeWire calculates as increases of 150%, 57% and 136% for a customer that keeps all of its traffic inside the cheap 17-hour window [7]. Peak Flash rates are $0.014, $0.44 and $1.32 [8]: five times, 3.1 times and 4.7 times the current prices [12][13][14].
V4-Pro moves further. DeepSeek's own pricing documentation, as of August 13, listed $0.003625 for cache-hit input, $0.435 for cache-miss input and $0.87 for output per million tokens [9]. Off-peak cache-hit input will land at roughly six times that level, with cache-miss up about 52% and output up about 128%; a peak-hour cache hit will cost just over 12 times the current rate [10]. Carrying the 50% tier gap through, peak Pro output arrives near 4.6 times today's price [21].
The largest relative jumps are on cache-hit input on both models [7][10]. That is the line item that heavy prompt-reuse designs depend on, and it is the one least visible in most cost dashboards.
Timing is not as simple as one window. Peak is defined as 01:00-04:00 UTC and 06:00-10:00 UTC, seven hours in total, with the remaining 17 hours off-peak [5]. The peak block is split, which means 04:00-06:00 UTC falls in the cheap tier and any time-based router needs two exclusion ranges rather than one [22]. For a workload spread evenly across all 24 hours, RuntimeWire puts the blended output price at about three times today's rate on either model [11].
Who absorbs this depends on how much control you have over arrival time. Batch evaluation, document processing and data enrichment can queue [15]. Consumer chat products and agents cannot, and an agent can turn a single user instruction into repeated calls across planning, tool use and retries, multiplying the per-token increase [15]. One structural break goes the other way: DeepSeek's peak windows fall outside most normal US working hours, so US teams can run daytime batch jobs at off-peak rates, while round-the-clock services will pay both tiers [16].
Three things to check before Friday. First, whether your queueing and routing can express two peak windows and a hard cutover at 16:00 UTC on August 16 [2][16][17]. Second, your realised blend, because the 3x figure assumes uniform traffic and real bills depend on arrival time, cache behaviour and output length [11]. Third, every internal cost-per-task number produced under the old schedule, none of which describes post-August-16 economics; a reproducible figure now has to name the model version, the cache behaviour and the time window [18].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
DeepSeek will raise API prices for its V4 models on August 16 and introduce a two-tier rate card intended to move flexible computing jobs outside its busiest hours.
- [2]
The new prices take effect at 16:00 UTC on August 16, 2026.
- [3]
DeepSeek said off-peak rates will sit 50% below peak pricing; that framing describes the gap between the two new tiers.
- [4]
Every off-peak rate in the new table remains higher than the corresponding price DeepSeek charges today.
- [5]
DeepSeek defines peak hours as 01:00-04:00 UTC and 06:00-10:00 UTC, totaling seven hours each day; the other 17 hours are off-peak.
- [6]
DeepSeek currently charges V4-Flash customers $0.0028 per 1 million cache-hit input tokens, $0.14 per 1 million cache-miss input tokens and $0.28 per 1 million output tokens.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRuntimeWire StaffAug 13DeepSeek sets peak V4 API rates at twice off-peak levels as prices rise
Additional citations
- RuntimeWire
- DeepSeek on X, via RuntimeWire
- DeepSeek pricing documentation, via RuntimeWire

