The mechanism is arithmetic, not intelligence. Cost is token price multiplied by tokens consumed [12], and the scaffolding sets the second term: how many turns it takes to finish, and whether it pays for a tool interface at all. The two harnesses in the arXiv study that ship no MCP support completed every run over the plain command line [2], which is the cheap path by construction. The comparison the authors actually set out to make came apart in their hands: thirteen strictly paired MCP-to-CLI ratios run from 0.43x to 29x, with outliers on both sides [5].
Take the reported band at its geometric middle and the scaffolding penalty is about 11.8x [3]. That is the number a rail card rounds to 10x, and it is the one an operator can act on, because the scaffold is a procurement choice made in an afternoon.
Harvey is working the same seam from the other end. Tenet is a Kimi K3 base post-trained with Fireworks for long-horizon legal work, and Harvey says it also made harness improvements to training and task execution [9]. K3 is open weight, 2.8T parameters, activating 16 of 896 experts with a 1M-token context [13], so the base is available to anyone; Harvey's stated second goal is letting law firms build their own specialised models and own their intelligence [15]. The gains held up on two benchmarks the model had not seen, across other harnesses [11].
The two sources are not measuring the same axis, and that is the honest read. Harvey reports roughly double the held-out task completions against base K3 [10] and says reward shaping favoured shorter trajectories at equal performance, keeping cost stable [12], but publishes no dollar figure. The arXiv work reports 139x cost variation for one 27-billion-parameter model that completed the task under every scaffolding [4]. A model upgrade bought a doubling in quality. A scaffolding choice can move the bill by two orders of magnitude without changing whether the work gets done.