Build4 distinct publishers3 min readPublished Updated
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
DeepSeek published Harness v0.1, an MIT-licensed runtime for building coding agents, on August 13th, the same day its changelog confirms DeepSeek-V4-Pro reached the app, web interface and API [1][2][3]. Three days later, at 16:00 UTC on August 16th, the company replaces flat international API rates with peak and off-peak pricing that is an increase at every level for V4-Pro [4][5][6].
The part that changes an operator's options is not the model. Harness, also called dsh, sits on the Cordis plugin framework under one design rule: practically every component can be replaced, including models, tools, skills, sessions, sandboxes, storage, agent loops, orchestration and user interfaces [7]. The model is itself a plugin, and the provider documentation covers Anthropic, OpenAI, Codex, Amazon Bedrock, Google Vertex AI, Microsoft Azure and custom endpoints alongside DeepSeek's own API [8]. Standard agent mode reads and edits files, runs shell commands, searches locally and on the web, keeps a plan, calls reusable skills, delegates to subagents and applies approval policies; there is a Python SDK and a local web interface via `npx @deepseek-ai/dsh web` [9][10].
That is not the same product as Claude Code or Codex, which Anthropic and OpenAI have built out across terminals, development environments, cloud tasks and repository workflows [11]. What DeepSeek is shipping is the layer underneath: something you can inspect, fork and rebuild [11]. DeepSeek says Harness is a developer preview and will take compatibility-breaking changes, so the MIT license removes the legal barrier while the unstable interfaces keep the engineering cost high [12]. The roughly 94,000 stars and 8,600 forks GitHub showed on August 14th measure launch attention, not installations [13].
There is a second reason to want the runtime published. In its July 31st V4-Flash note, DeepSeek said its public code-agent evaluations ran on the then-unreleased Harness in minimal mode, configured with only a persistent shell and a file-editing tool [14]. Tool definitions, prompts and execution loops move agent scores materially, so publishing the harness is the difference between reading a vendor number and reproducing it [15]. DeepSeek reports 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE and 60.0 on Humanity's Last Exam with tools for the GA build, all vendor-reported, with several comparisons drawing on internal test sets [16].
The pricing side is where the harness earns its keep. The current card lists $0.003625 per million cached input tokens, $0.435 uncached and $0.87 output, against a 1 million-token context and 384,000-token maximum output [17]. Cached input is therefore about 120 times cheaper than uncached today [18]. After August 16th, DeepSeek says even the discounted rate raises cached input by about 507% and uncached input by 52%, which puts off-peak uncached near $0.66 and cached near $0.022, narrowing the caching advantage to roughly 30x [5][19]. Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day, with everything else off-peak [4][20].
Watch three things. Whether the 0813 weights follow: Simon Willison notes weights exist for April's V4-Pro and July's V4-Flash-0731 but could not confirm plans for this build, and had to link OpenRouter because DeepSeek published no obvious announcement page [21][22]. Whether the rate card settles, since the pricing page still warns of a significant overall increase without a final number [23]. And whether anyone outside DeepSeek reproduces 87.9 on Terminal Bench using the published harness [16][15]. Meanwhile, `deepseek-v4-pro` is a moving alias with no published change log for 0813, so log the returned model version and rerun your own regressions [24].
Ranked by verification strength, evidence, and original report placement.
DeepSeek will replace its flat international API rates with peak and off-peak pricing at 16:00 UTC on August 16th. Peak periods run 01:00 to 04:00 UTC and 06:00 to 10:00 UTC; all other hours use the off-peak rate.
DeepSeek says the general-availability release scored 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE and 60.0 on Humanity's Last Exam with tools; these are vendor-reported results and several comparisons include DeepSeek's internal test sets.
DeepSeek released an open-source runtime for building coding agents on August 13th, pushing beyond model APIs into the developer tooling controlled by Anthropic's Claude Code and OpenAI's Codex.
VentureBeat reported that DeepSeek Harness v0.1 entered developer preview on August 13th under the MIT license.
For V4-Pro the pricing change is an increase at every level; even the discounted rate raises cached-input pricing by about 507% and uncached-input pricing by 52%.
DeepSeek Harness, also called dsh, is built on the Cordis plugin framework, and its core design rule is that practically every component can be replaced: models, tools, skills, sessions, sandboxes, storage, agent loops, orchestration and user interfaces.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary artefacts documented, key schedule single-sourced
The rate card, build name and serving limits come from DeepSeek's own pricing page, and the runtime's architecture, license and capabilities are described in concrete, checkable detail (repository, npx command, provider list). Weakening the picture: the peak/off-peak schedule and every percentage increase rest on one publisher, the capability claims are vendor-reported with internal test sets, and the sources disagree on whether an official change log for the 0813 build exists at all.
Distribution real, usage unproven
The model is genuinely shipping - listed in the API catalog, live on app, web and API, and reachable via OpenRouter - so commercial distribution exists. The runtime side has only attention: about 94,000 stars and 8,600 forks within a day, which the reporting itself distinguishes from installations or production use, and DeepSeek labels Harness a breaking-change developer preview. No named deployment, production user or usage disclosure appears in any source.
Mildly overstated framing, well-hedged detail
The 'challenge to Claude Code and Codex' framing runs ahead of the artefact: the same reporting concedes Harness is a preview in a different product tier from managed coding products, and the strongest performance figures are vendor-reported with no independent replication. The overstatement is modest rather than severe because each source hedges in the same breath - stars are labelled attention, benchmark provenance is questioned, and the missing change log is flagged - and the hardest facts, the price increases, are the least flattering to the vendor.
Open runtime as distribution for a repriced endpoint
DeepSeek's commercial interest is legible in the sequencing itself: an MIT-licensed, model-agnostic runtime that seeds usage in the layer rivals monetise, released three days before every V4-Pro price tier rises, with the cache discount that made long agent sessions cheap compressing sharply. The reporting also notes DeepSeek gains a natural distribution channel for V4-Pro and V4-Flash, and that its own published agent evaluations were produced with this harness - a vendor incentive to define the measurement environment as well as the tooling.
Moderate - documented artefacts, thin corroboration
Two of three legs are anchored in DeepSeek's own published surfaces (pricing page, OpenRouter listing) and the third gives unusually specific, falsifiable runtime detail. Confidence is held to the middle by single-source dependence on the pricing schedule and deltas, an unresolved contradiction about the change log and GA date, absence of any independent benchmark or deployment evidence, and the note that open weights for the 0813 build remain unconfirmed.
build
DeepSeek shipped the harness: dsh makes the runtime, not the model vendor, the choice2 distinct publishers
build
Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test2 distinct publishers
build
NVIDIA's safety teams put the agent security boundary in the runtime, not the model2 distinct publishers
build
The agent harness is the product: DeepSeek ships a runtime that routes to its rivals1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 16, 2026
1 article · August 12, 2026
1 article · August 14, 2026
1 article · August 12, 2026