Invest1 distinct publisher3 min readUpdated
Composio ran the leaderboard-topping model through 240 agent runs against live Gmail, GitHub and Slack tools, and 129 passed. Harness choice moved the outcome more than price did.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The testing firm Composio put DeepSeek's V4 Flash through 30 deliberately difficult multi-step agent workflows and recorded a 53.8% pass rate: 129 of 240 runs [4][5]. That number belongs in any deck that proposes migrating production agents to V4 Flash on the strength of its pricing, which the model's promoters have priced at roughly a tenth of comparable models [3][11].
The sticker price is real. V4 Flash lists at $0.14 per million input tokens and $0.28 per million output tokens, which the report describes as undercutting comparable models by roughly tenfold [2][3]. The model launched on July 31, was called a "total monster" by developers, and went to the top of multiple leaderboards [1]. It is also still in public beta, with a broader adjustment to DeepSeek's API pricing scheduled for August 16, 2026 [10].
What Composio measured is closer to what operators actually buy. The 30 tasks used live tools -- Gmail, GitHub, Slack and Google Sheets -- and were run across eight agent harnesses including Claude Code, Codex and OpenCode [7][6]. Thirty workflows over eight harnesses is 240 runs, so each workflow was attempted once per harness [13]. Of those 240 attempts, 111 failed [12]. Only six of the 30 workflows were completed successfully by every harness tested, which is 20% of the task set that can be called reliable independently of integration choices [9][15].
The spread across harnesses is the more useful finding. Pi Agent completed 20 of 30 tasks, a 66.7% rate, while the aggregate across all eight sat at 53.8% [8][14][4]. Per Composio's framing, the same underlying model looks brilliant or mediocre depending on how it is wired into a workflow, and thoughtful integration can partially compensate for model limitations [8][16]. That cuts both ways: a procurement decision made on model price alone is being made on the smaller of the two variables.
Now the arithmetic that a migration memo should include. At a 53.8% pass rate, you need about 1.86 attempts per completed task, which puts the effective cost of a successful run at roughly $0.26 per million input tokens and $0.52 per million output tokens [17]. That is still well under a competitor at ten times the sticker price, so the cost case does not vanish on retries alone. It vanishes on the parts Composio did not price: failed runs consume compute, developer debugging time, and end-user trust, per the report's own accounting [11]. A workflow that touches a customer's Gmail or a production GitHub repo and fails 46% of the time is not a cheaper version of one that works [7][12].
Three things to watch. First, the August 16, 2026 API pricing change, which will tell you whether the tenfold gap is a launch posture or a durable position [10]. Second, whether anyone reruns the Composio suite once V4 Flash exits beta, since the beta label is DeepSeek's own signal that this is unfinished [10]. Third, your own harness: if Pi Agent's 66.7% is the ceiling in this test set, the honest question is not which model to buy but whether your integration is closer to the best harness or the worst [14][8].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Pi Agent was the strongest performer, completing 20 out of 30 tasks; Composio's reading is that the same underlying model can look brilliant or mediocre depending entirely on how it is integrated into a workflow.
DeepSeek is offering V4 Flash in public beta, with a broader adjustment to API pricing scheduled for August 16, 2026; the report argues the beta label signals DeepSeek views the model as a work in progress rather than ready for mission-critical deployment.
The report concludes that the choice of agent harness matters as much as the choice of model, and that thoughtful integration can partially compensate for model limitations.
DeepSeek's V4 Flash launched July 31, was called a "total monster" by developers, and shot to the top of multiple AI leaderboards.
V4 Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens.
When testing firm Composio ran V4 Flash through a battery of real-world agent tasks, the model managed a 53.8% pass rate.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Concrete third-party numbers, single unverifiable relay
The cluster carries specific, internally consistent measurements (240 runs, 129 passes, 30 workflows, eight harnesses, Pi Agent 20/30, six universal passes) that reconcile arithmetically. But everything comes from one publisher republishing another outlet's write-up of one testing firm's evaluation, with no methodology link, no per-harness table, no comparator baseline and no named competitor price behind the tenfold claim.
Beta availability and one eval; no production usage
Observable adoption is limited to availability and testing signals: a July 31 launch in public beta, a scheduled August 16, 2026 API pricing adjustment, and one third-party evaluation. Leaderboard placement and a developer quote are reception, not deployment. No customer, workload, traffic or revenue disclosure appears anywhere in the cluster, and the vendor's own beta label argues against production placement.
Leaderboard framing overstates agent-task reliability
The gap runs in the direction of overstatement: a model called a 'total monster' and topping multiple leaderboards fails roughly 46% of complex live-tool agent runs and ships as beta. The story itself documents that gap, which limits how far the hype charge extends to this cluster's own framing. Offsetting slightly, the article's counter-claim that the price savings 'evaporate' also overshoots its own arithmetic, since a retry-adjusted effective price near $0.26/$0.52 per million tokens still sits well under a tenfold premium.
Vendor-run eval relayed by a syndicating outlet
Three incentive layers are visible in the supplied material. DeepSeek benefits from leaderboard placement and aggressive token pricing ahead of an API pricing change. Composio is the sole author of the evaluation and its published conclusion is that harness and integration choice is decisive, a framing aligned with agent-tooling interests, yet no methodology, per-harness data or interest disclosure accompanies it. The publisher is a crypto-sector outlet republishing content credited 'via cnet.com', so editorial distance from the original reporting is unclear.
Coherent figures, one publisher, unresolved date and scope gaps
Confidence is limited mainly by structure rather than internal contradiction: one publisher, one relayed source, one testing firm, no primary document. The quantitative core is arithmetically consistent, which supports the specific pass-rate and scope claims at moderate confidence. Weakening factors are the unnamed price comparators, the withheld per-harness results, and the piece publishing on the same date it describes an API pricing adjustment as still upcoming.
build
GLM-5.3 kept the base model and bought ten times the environments instead2 distinct publishers
build
DeepSeek Harness makes the agent loop a plugin, so pick your seam before you write code1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
build
A 12MB Go binary bets agent cost control is cache stickiness, not a dashboard1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
1 article · August 16, 2026