Leadership1 publisher2 min readPublished
Kognitos's founder traces the enterprise token bill to undocumented work
Binny Gill argues in a Forbes council column that firms are paying reasoning-model rates for rule-following tasks. The two studies he cites measure consultants and a research router. Neither one measured an enterprise bill.
The Board Room · Leadership desk

What happened
- Kognitos chief executive Binny Gill writes that firms burning too many tokens are making one or both of two mistakes: using a model overqualified for the work, or handing it work nobody documented.
- A 2026 Organization Science study of 758 consultants found that on a task outside the AI's reliable capabilities, people using it were 19% less likely to reach the correct answer than those without.
- Stanford's FrugalGPT experiments, which pick a model per query, matched the performance of the best individual model while cutting costs by up to 98%.
- Gill reports that a CIO at a large financial services firm has stopped giving AI to everyone, on the grounds that most staff do not yet know how to use it carefully.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- cost Documenting a workload is analyst time and tokens spent up front, and the team that funds the writing is rarely the team whose token line falls.
- constraint Routing by task class means classifying a task before the call, and a cheap model handed judgement work reproduces the same mismatch penalty in the other direction.
- decision Budget owners have to choose between re-allocating workloads they already run and waiting on price cuts that would apply to the current model mix unchanged.
The second of the two mistakes is the expensive one. An undocumented task makes the model work out the procedure again on every call, so a decision that needed making once gets billed at reasoning rates every time it runs [4]. "Spend the tokens once to write the rules down, then use a cheaper form of intelligence to execute them," Gill wrote [5]. Someone has to sit down and specify those steps before a cheaper model can follow them, and Gill's own framing puts the planning cost first and the execution saving after [14].
Spend at 2% of the best single model's cost is a fifty-fold reduction [13]. Those were research experiments in which a router picked among models query by query [7], and Gill wrote that the exact savings will differ by workload and that businesses do not need the most capable model for every job [8].
The consulting study points the same way from the other end. It measured people using an AI tool on a task the tool could not handle reliably [6]. Mismatched capability degraded the output. The experiments used no router, so they did not test what routing saves.
Gill names the instinct: blame the technology, or wait for prices to fall [3]. Re-allocating the work changes which rate applies; a price cut applies to whichever model you are already calling. A firm running a reasoning model on rule-following work would keep paying the reasoning rate, discounted.
Gill's forecast is also a vendor's forecast. "I posit that cheaper, specialized models will gain popularity in the coming months, as they'll produce less bias and higher accuracy at constrained tasks," he wrote [9]. He is the founder and chief executive of Kognitos and holds nearly 100 patents in computer science [1], and the piece ran as a Forbes Tech Council column [15]. He also warns in the other direction: whether a smarter system is a safer one depends on the job [12], and an overqualified worker in a rules-based role may question the process and change it where the business needed the rules followed [11].
The choice in front of a budget owner this quarter is whether to spend analyst time documenting workloads that already run. If they do, next quarter's question is a routing configuration. If they wait, any price cut lands on the same allocation of work.
What to watch
- Whether a named enterprise publishes routing savings on its own workload rather than on research benchmark queries.
- Whether the cheaper specialized models Gill predicts actually ship into enterprise stacks, and at what measured accuracy on constrained tasks.
- Whether access rationing of the kind the unnamed CIO described becomes a written entitlement policy anywhere it can be audited.