Skip to content

Leadership1 publisher2 min readPublished

Decisions CEO blames enterprise AI cost creep on routing logic run through the model

Giles Whiting, CEO of orchestration vendor Decisions, says enterprises pay inference to re-answer fixed questions millions of times in production. The same architecture also makes a routing decision hard to reproduce for an auditor.

The Board Room · Leadership desk

Illustration accompanying Decisions CEO blames enterprise AI cost creep on routing logic run through the model

What happened

  • Giles Whiting, CEO of AI orchestration company Decisions, wrote in Forbes that enterprises route already-known decisions through probabilistic models and pay inference to answer the same fixed questions thousands or millions of times.
  • He points to Gartner's May 2026 report on taming the AI cost curve, which he says observed expected AI value eroding through cost creep and a silent "token tax".
  • In some architectures, he wrote, the model rereads prior context at each step to decide what comes next, so a longer process costs more for reconstructing its own state rather than for new reasoning.
  • His worked example is an insurer whose eligibility requirements, thresholds and escalation paths could run as deterministic rules, leaving the model the adjuster's unstructured notes.
  • The column ends with five tests for technology leaders, starting with estimating cost at production volume because pilot economics say little once transaction counts and model calls rise sharply.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint The window to change this closes early: by the time volume makes the invoice material, Whiting says the architecture is already difficult to change, so the commitment is made while the numbers still look trivial.
  • exposure A firm that has to explain a decision, reproduce it and show a policy was enforced is reconstructing a sampled path when routing lives inside the model.
  • decision Where the governing logic sits decides whether the same input always follows the same route.
  • contradiction The prescription comes from a company that sells deterministic orchestration, and the column gives no spend figures, so the cost half of the case is an argument about mechanism, and readers have to work out the spend themselves.

The line Whiting draws is between reasoning and orchestration. Reasoning, in his definition, covers summarizing unstructured information, interpreting ambiguous inputs and handling cases that cannot be fully anticipated; orchestration covers sequencing, routing, policy enforcement and other known control logic, and where the path is predictable he says it can be defined explicitly instead of regenerated through inference [11]. The cost of getting that line wrong is folded into the rest of the bill. Whiting wrote that the orchestration tax "is buried inside AI consumption, mixed together with the inference that is actually creating value" [6].

A skeptic can compress the whole piece into one line: a company that sells deterministic orchestration reports that enterprises need deterministic orchestration [1]. That is a fair caution on the cost advice, and it applies to the supporting citation too, which reaches the reader through Whiting's own summary of a report dated four months before the column ran [15]. It bears much less on the predictability argument, which holds at any token price. The column does not say what the deterministic path costs to build and keep current when a policy changes [17].

There is also a practical reason the cost side is the weaker half. Whiting's own account gives two distinct drivers: the fixed question re-asked on every run, and the model that rereads prior context to work out its next action [4][5]. Only the second one compounds with process length, and how much it compounds depends on the workflow.

For a team deciding this quarter, the unit of work is the individual workflow step. Whiting proposed auditing each one by asking: "Does this step require genuine judgment, or is it a known decision being treated like a reasoning problem?" [13] That question is answerable at design time, before any call is billed at volume. The cost claim in this column is plausible and unmeasured; the predictability claim can be tested inside a company's own logs by sending the same input through twice and comparing the paths [7].

What to watch

  • Whether Gartner's cost-curve work, or any enterprise, puts a figure on how much of an inference bill goes to control flow.
  • Whether a named enterprise publishes before-and-after inference costs from moving routing out of the model into explicit rules.
  • Whether auditors and regulators in claims, credit or benefits start asking firms to reproduce a routing decision step by step.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories