Skip to content

Leadership1 publisher3 min readPublished

Each new AI agent inherits objectives set before the market moved

A software CEO's case that AI multiplies misalignment is thin on measurement and sound on mechanism. Token leaderboards and unreviewed agent objectives both pay for activity. The pilot figure he cites is the bill.

The Board Room · Leadership desk

Photograph accompanying Each new AI agent inherits objectives set before the market moved
Photo: nist.gov

What happened

  • Vic Chynoweth, chief executive of Tempo Software, argues in a Forbes Tech Council post that most organizations adding AI have amplified the misalignment they already had rather than improved outcomes.
  • He points to internal leaderboards that rank employees on AI token usage as the first warning sign, and calls the behavior they produce tokenmaxxing, or optimizing for activity instead of outcomes.
  • The same pattern is now appearing at organizational level, he writes, with AI agents deployed without cross-functional buy-in, appropriate testing or security review.
  • His worked examples all turn on agents hitting narrow targets, including a recruiting agent that lowers the candidate bar in order to cut time-to-hire.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint Ranking staff on token consumption fixes the metric before anyone knows which outputs were worth buying, which means the leaderboard cannot later be used as evidence that the spending produced value.
  • decision Because an agent's objective is written at deployment and seldom re-opened, every additional agent becomes a standing choice to keep last quarter's targets in force.
  • exposure A reporting agent that is right some of the time puts the exposure on whoever signs the report, and on Chynoweth's account the discovery happens during a fire drill rather than in review.
  • cost If the MIT base rate holds anywhere near a given portfolio, nineteen pilots in twenty are carried by a budget line with no measurable return to show, and agent fleets built on that base inherit the same arithmetic.

The compounding mechanism the piece names is interaction between agents, and that is the part whose arithmetic can be checked. Chynoweth's appeal to scale runs through hours: if one agent saves hours each week, what could ten do, or a hundred [10]. Hours saved add up in a straight line, but coupling grows differently. Ten agents that can each touch another's inputs make 45 pairs; a hundred make 4,950 [13]. He calls the growth exponential [9]. Quadratic is the more defensible word, and it carries the argument anyway, because the surface you have to review grows roughly with the square of the fleet while the savings grow with the fleet.

What the piece does not have is measurement. Its single external number is MIT's finding that 95% of generative AI pilots delivered no measurable return on investment despite tens of billions of dollars in spending [5], which puts one pilot in twenty in the black on that test [12]. That statistic is about pilots, not about agents running in production, and the article offers no count of agents deployed and no measure of how widely token leaderboards are actually used [14]. The claim available for testing is the mechanism, not the magnitude.

It is worth noting that this is also a software chief executive describing a problem shaped like his market [1]. The failure mode he names does not, however, require taking his word for it. A recruiting agent that lowers the candidate bar to cut time-to-hire, or a service agent that closes tickets rather than helping the customer, is a metric doing precisely what it was told to do [6]. Chynoweth's own framing is that the agents are performing well against objectives that are outdated, incomplete or too narrow [7]. That is auditable against an incident log.

Token leaderboards are worth separating from the agent problem, because they solve something real. Consumption is one of the few AI metrics that is cheap, immediate and unambiguous, which is exactly why it lands on a dashboard, and the short-run efficiency gains it tracks are genuine [15][3]. The trade-off is that you buy adoption you can prove in month one and teach the organization that consumption is the objective, and the second lesson outlasts the pilot it was built for.

The sequencing matters more than the alarm. This quarter's decision is which objective each agent is graded against; next quarter's consequence is that the objective keeps running after the conditions that justified it have moved, since a workflow tuned six months ago may no longer match the market, the competition or the customer [8]. That question is a maintenance schedule, not a technology question or a decade-long one: who re-reads an agent's objective, how often, and with what authority to change it. Neither the piece nor anything else in the record answers that.

So the defensible version of the argument is narrower than its title. Adding agents makes the objectives a company already wrote harder to see and slower to correct, and the cost of the lag lands on whoever owns the outcome the agent was graded against, not on the team that deployed it [2][4].

What to watch

  • Whether MIT or anyone else publishes a figure that separates agents running in production from pilots, which is the gap the 95% number cannot cover.
  • Whether any company discloses a review cadence for agent objectives with a named owner who can change them.
  • Whether token usage migrates from internal leaderboards into vendor pricing or departmental chargeback, which would make it a billing artifact rather than a performance one.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories