Build1 publisher3 min readPublished
Together AI rebuilt its $3.35-per-task coding cascade from hidden-test benchmark trials
Together AI's 83% success at $3.35 per coding task came from replaying benchmark trials, not a live router, engineer Zain Hasan told Let's Data Science. The figure counts tokens only, so each team's own escalation check and review time set the price of an accepted patch.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- When DeepSeek's patch failed even one held-out test, GPT got the same problem from scratch, with no access to the first patch, its test output or any feedback.
- The replay drew on 113 tasks with four trials per model, 904 rollouts in all, using deepseek-v4-pro-0813 and gpt-5.6-sol at max settings.
- Hasan's illustrative review bill, 83 patches at 15 minutes each and an assumed $60 an hour, comes to $1,245.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams trying a cheap-model-first cascade have to build and measure their own accept-or-escalate check, because that check's error rate sets both success and spend in a live router.
- cost At Hasan's assumed $60-an-hour review rate, checking a solved patch costs about 3.7 times the tokens that produced it, so a cheaper first model trims the smaller part of the bill.
- exposure Patches from either model that clear functional tests can still carry performance or security regressions, so escalating to the pricier model still leaves a review to pay for.
"We reconstructed the cascade from the official pass@4 trials rather than running it live," said Zain Hasan, a Staff AI/ML Engineer at Together AI [2][13]. In that replay, the benchmark's own held-out tests decided when to escalate [3]. Those tests gave an exact accept signal [16]. According to Let's Data Science, a real team's tests may be incomplete, and passing them does not show that a patch meets requirements they do not cover [16]. The 83% [1] is success under benchmark scoring. It is not an observed production acceptance rate [16].
Most repositories do not ship with an answer key. With an exact grader, every passing first patch is kept and every failing one goes to GPT [3]. A real check can get this wrong in two ways. It can keep a bad DeepSeek patch. Or it can throw out a good one and pay for a fresh GPT attempt that may fail. Neither error raises the success rate, so I'd treat 83% as a ceiling for this two-step design on these tasks [3]. Spend can go up or down, depending on which error the check makes more often [3]. Hasan said model-based judges can help pick candidates, but their accuracy and speed need to be measured too [12]. A verifier that approves the wrong patch undermines the routing decision, and a slow one can become the bottleneck [12].
Hasan's example of a bad accept involves a profile cache. Users do not see their profile changes right away, and the agent removes the cache instead of fixing invalidation [10]. The visible bug goes away and the functional tests pass, but every page load now queries the database [10]. He gave it as a teaching example, not a failure seen in the benchmark or a customer incident [10]. His fix is to write tests before the coding attempt, add regression tests, and run security and performance checks where they apply, on output from both models [11]. Newly found failure cases go into the evaluation as they turn up [14].
"These figures cover only inference, meaning the token costs of running the models," Hasan said [7]. The figure leaves out sandbox time, verification, human review and the wait for a fix [8]. When LDS followed up, he confirmed both the $4.04-per-solved-task division and the review calculation [6][9]. I think vendors should publish cost figures this way. His review example comes to $15 a patch [1]. Add the $4.04 of tokens and a solved task costs about $19.04 before any sandbox or verifier time [4]. The review minutes and the hourly rate are his assumptions, not measured staffing costs [9]. A task the benchmark counts as solved is also not necessarily a patch a real codebase would accept [15].
What to watch
- Whether Together AI runs the cascade live with a visible-test or model-judge accept rule and publishes how often that rule accepts failing patches.
- A version where GPT receives DeepSeek's patch and test output, measuring repair of the first attempt instead of a second independent try.
- Measured review, sandbox and verification costs from a team running the cascade against its own repository.