Invest1 publisher3 min readPublished
DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks
Composio ran the leaderboard-topping model through 240 agent runs against live Gmail, GitHub and Slack tools, and 129 passed. Harness choice moved the outcome more than price did.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- DeepSeek's V4 Flash launched July 31, was called a "total monster" by developers, and shot to the top of multiple AI leaderboards.
- V4 Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens.
- V4 Flash's pricing undercuts comparable models by roughly tenfold; at $0.14 per million input tokens it is roughly a tenth the cost of comparable models from major US competitors.
- When testing firm Composio ran V4 Flash through a battery of real-world agent tasks, the model managed a 53.8% pass rate.
- Out of 240 total runs spanning 30 deliberately difficult, multi-step workflows, only 129 passed.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
The testing firm Composio put DeepSeek's V4 Flash through 30 deliberately difficult multi-step agent workflows and recorded a 53.8% pass rate: 129 of 240 runs [4][5]. That number belongs in any deck that proposes migrating production agents to V4 Flash on the strength of its pricing, which the model's promoters have priced at roughly a tenth of comparable models [3][11].
The sticker price is real. V4 Flash lists at $0.14 per million input tokens and $0.28 per million output tokens, which the report describes as undercutting comparable models by roughly tenfold [2][3]. The model launched on July 31, was called a "total monster" by developers, and went to the top of multiple leaderboards [1]. It is also still in public beta, with a broader adjustment to DeepSeek's API pricing scheduled for August 16, 2026 [10].
What Composio measured is closer to what operators actually buy. The 30 tasks used live tools -- Gmail, GitHub, Slack and Google Sheets -- and were run across eight agent harnesses including Claude Code, Codex and OpenCode [7][6]. Thirty workflows over eight harnesses is 240 runs, so each workflow was attempted once per harness [13]. Of those 240 attempts, 111 failed [12]. Only six of the 30 workflows were completed successfully by every harness tested, which is 20% of the task set that can be called reliable independently of integration choices [9][15].
The spread across harnesses is the more useful finding. Pi Agent completed 20 of 30 tasks, a 66.7% rate, while the aggregate across all eight sat at 53.8% [8][14][4]. Per Composio's framing, the same underlying model looks brilliant or mediocre depending on how it is wired into a workflow, and thoughtful integration can partially compensate for model limitations [8][16]. That cuts both ways: a procurement decision made on model price alone is being made on the smaller of the two variables.
Now the arithmetic that a migration memo should include. At a 53.8% pass rate, you need about 1.86 attempts per completed task, which puts the effective cost of a successful run at roughly $0.26 per million input tokens and $0.52 per million output tokens [17]. That is still well under a competitor at ten times the sticker price, so the cost case does not vanish on retries alone. It vanishes on the parts Composio did not price: failed runs consume compute, developer debugging time, and end-user trust, per the report's own accounting [11]. A workflow that touches a customer's Gmail or a production GitHub repo and fails 46% of the time is not a cheaper version of one that works [7][12].
Three things to watch. First, the August 16, 2026 API pricing change, which will tell you whether the tenfold gap is a launch posture or a durable position [10]. Second, whether anyone reruns the Composio suite once V4 Flash exits beta, since the beta label is DeepSeek's own signal that this is unfinished [10]. Third, your own harness: if Pi Agent's 66.7% is the ceiling in this test set, the honest question is not which model to buy but whether your integration is closer to the best harness or the worst [14][8].