Published Product3 min read
DeepSeek's harness arrives with a 4.5x price rise, and the cheap-substitute case goes with it
A million output tokens on DeepSeek-V4-Pro goes from $0.87 to $3.96 at peak from 16 August, in the same week the company shipped a developer preview of its own Claude Code rival.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- From 16 August, a million output tokens through DeepSeek-V4-Pro costs $3.96 at peak, up from $0.87.
- DeepSeek V4-Flash output pricing goes from $0.28 to $1.32 per million tokens, Bloomberg reported.
- The new peak V4-Pro output price is about 4.55 times the old peak price.
- The new V4-Flash output price is about 4.71 times the old price.
- Off-peak rates are half the new peak, so $1.98 per million output tokens for V4-Pro and $0.66 for V4-Flash.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
DeepSeek released a developer preview of DeepSeek Harness v0.1 on Thursday, the scaffolding layer that reads files, edits code, browses the web and keeps working until a task is done, which is the layer Anthropic sells as Claude Code [6][7]. In the same week it repriced its flagship: from 16 August, a million output tokens through DeepSeek-V4-Pro costs $3.96 at peak, up from $0.87 [1], a rise of about 4.5 times [3] that removes most of the arithmetic teams were using to justify building on it instead.
V4-Flash moves from $0.28 to $1.32, Bloomberg reported [2], roughly 4.7 times [4]. Off-peak rates are half the new peak, so $1.98 and $0.66 [5]. That is the number worth staring at: the new discounted V4-Pro rate is about 2.3 times the old undiscounted peak [8], so a team that reschedules every workload into the quiet hours still pays more than twice what it paid at lunchtime [9]. On 100 million output tokens a month, peak billing goes from $87 to $396, an extra $309 [10].
The harness explains the pricing more than the model does. DeepSeek says the harness uses an open architecture in which users can plug in any component, including models from other companies, and framed that against American rivals it said hard-code their products [11][12]. That is a bid to own the workspace rather than the engine, and the workspace is where the money in agentic coding sits [7][13]. The signalling ran for days beforehand: a "DeepSeek Harness Team" account on WeChat, verified by Tencent and sitting under a Beijing entity that Chinese corporate records link to DeepSeek, plus job listings, one of which said the company wants to turn its models into cutting-edge agentic products, per Bloomberg [14][15].
What developers get for 4.5 times the price is a model its own maker positions at roughly frontier level on coding and behind elsewhere. The figures are vendor-reported and no independent evaluator has replicated them for this build [16]. DeepSeek's model card puts V4-Pro at maximum reasoning on 80.6% for SWE-bench Verified, level with Gemini 3.1 Pro and 0.2 points behind Claude Opus 4.6 at 80.8% [17][20]. On Terminal Bench 2.0 it shows 67.9% against GPT-5.4 at 75.1%, a 7.2-point gap [18][21]; on Humanity's Last Exam, 37.7% against Gemini 3.1 Pro at 44.4%, 6.7 points [19][22]. Early reaction to the 0813 build left developers underwhelmed on general capability and unhappy about the pricing, according to the South China Morning Post, though researchers were impressed in narrower areas such as cybersecurity [23].
The cost story is not that DeepSeek got more expensive to run. V4-Pro is a mixture-of-experts system with 1.6 trillion parameters, 49 billion active per token, about 3% of the total [24][25], and DeepSeek says its attention design cuts the compute for a single token to 27% of its previous generation [26]. TNW reads the fourfold rise as a decision about margin rather than a report about costs [27], while also arguing the headline multiple invites the wrong conclusion because $3.96 remains cheap in absolute terms [28].
One loose thread: V4-Pro-0813 shipped as the general-availability build this week, ending a preview of nearly four months, marked by a website statement saying the model offered "significantly enhanced agent capabilities" [29][30]. By Thursday afternoon the statement had been removed, the SCMP reported; DeepSeek has not explained why and did not respond to Bloomberg on the harness team [31][32].
Watch whether independent evaluators reproduce the 0813 numbers, whether the open harness actually attracts developers running rival models, and whether the off-peak window becomes the de facto price.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
From 16 August, a million output tokens through DeepSeek-V4-Pro costs $3.96 at peak, up from $0.87.
ReportedView cited source - [2]
DeepSeek V4-Flash output pricing goes from $0.28 to $1.32 per million tokens, Bloomberg reported.
- [5]
Off-peak rates are half the new peak, so $1.98 per million output tokens for V4-Pro and $0.66 for V4-Flash.
ReportedView cited source - [6]
On Thursday DeepSeek released a developer preview of DeepSeek Harness v0.1.
ReportedView cited source - [7]
A harness is the scaffolding wrapped around a model that lets an agent read files, edit code, browse the web and keep going until a task is finished; that is the layer Anthropic sells as Claude Code, and where the money in agentic coding sits.
ReportedView cited source - [9]
A team that moves every workload to the quiet hours still pays over twice what it paid at lunchtime.
ReportedView cited source
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- thenextweb.comAna Maria ConstantinAug 13DeepSeek built a Claude Code rival, then quadrupled its prices
Cited in this coverage: Bloomberg, via thenextweb.com
Cited in this coverage: thenextweb.com analysis
Cited in this coverage: South China Morning Post, via thenextweb.com
Additional citations
- DeepSeek
- DeepSeek model card



