Product1 distinct publisher3 min readPublished
A Cutlefish post argues the token meter is precise about the spend and silent about the return, which is how hours and capacity mispriced product work long before AI arrived. The open question is who signs the R.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A token count is the cleanest number in the building. It arrives daily, it has an invoice behind it, and nobody argues about the definition. The other side of the ratio needs a person to say what changed and for whom, and there is no feed for that. The Cutlefish post's image catches the mood: an organisation that used to eat at a buffet is now buying each dumpling separately and re-litigating the value of every one [5]. What the finer-grained bill does not fix is the older problem the same post raises about hours, that where effort lands tells you nothing about the quality or efficacy of that effort, and that when the whole system is waiting on one specialist or one review, spending time elsewhere has negative leverage [8].
The arithmetic is worth doing slowly. Take the post's switching example: alternate between two tasks every ten minutes across a six-hour day, and at least three hours go to switching and reorienting [9]. Three of six is half the day gone [10], while the timesheet still books six hours against two tasks, because time accounting adds where coordination multiplies [9]. The post's other comparison lands in the same place: ten people at ten hours each and one person at a hundred hours both total 100 [11], and the ledger cannot tell them apart [8].
Token reporting has that property too. Spend per seat climbing tells you the meter is running. It does not tell you whether the same people were still inside that workflow six weeks later, or how long a new user took to get one output they were willing to ship. Those numbers are harder to pull, and they are the R.
Here is the forcing function I would use. Draw two axes: how precisely you can measure the input, and how precisely you have defined the return. Tokens land where hours and points landed, with the input measured to four decimal places and the return described in a sentence nobody will put their name to. The fix is not a better meter, it is deliberately inverting which axis gets the rigour. In practice that means a token budget does not clear review without one named workflow, the people working inside it, and an observable change with a date attached. If nobody will sign that sentence, the meter is doing the justifying.
That matters more than usual right now because the post is candid that it cannot tell what sits underneath the current enthusiasm, a genuine attempt to understand how AI augments people, or a calculation about how many people can safely be let go, or both [15]. Under that ambiguity, an undefined R does not stay undefined. It gets filled in later by whoever is holding the invoice.
Ranked by verification strength, evidence, and original report placement.
The post lists three recurring problems with such proxies: spending more time on the "I" side than the "R" side of ROI; myopically choosing shorter-term, easier-to-attribute use cases for the "R" side; and gravitating toward whatever is easiest to measure, with tokens being very easy to measure.
The post says "we've been here before" with hours and with capacity, and that the underlying measurement problems are remarkably universal.
The post compares the change to a world shifting from a buffet model to buying each dumpling one at a time, with everyone now worried about the ROI of each dumpling.
The post calls hours, as commonly used to understand investment, a major construct validity problem, noting that time is not a fungible thing that can be infinitely allocated or re-allocated across skill sets, team context, or even across a normal day.
The post says where you allocate time tells you nothing about the quality, efficiency or efficacy of that expenditure; that ten people spending ten hours each is not equal to one person spending one hundred hours; and that when the system is constrained by a single specialist, decision, dependency or review, spending time elsewhere has negative leverage.
The post's example: alternating between two tasks every ten minutes for a six-hour day means spending at least three hours on context switching and calibrating around the new task, because coordination is multiplicative while time accounting is additive.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Meta ran pods-plus-agents for a year and shelved it. Its own scoreboard says why1 distinct publisher
build
Opus 5 absorbed your verify prompts. The reading is still on your desk.1 distinct publisher
build
Elastic's IT team says AI ROI has to be a query, and it instrumented every event to get one1 distinct publisher
product
Salesforce says agents per org went 5 to 13 and build time fell to 1.9 days. It sells the agents.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One practitioner's reasoning, zero measurements
Everything here rests on a single newsletter post, and the post's strength is internal: the switching-tax arithmetic, the ten-by-ten versus one-by-hundred equivalence and the constraint argument all hold up on their own terms without needing a source. What is entirely absent is anything external — no token spend, no team data, no named organisation, not even a cited study — and the text we have breaks off before the token argument it advertises. Strong reasoning, unverified world.
Nothing here to count
This is an argument about how to measure, not an event with users. No release, deployment, benchmark or spend disclosure appears in the reporting, and inventing an uptake figure for an essay would be exactly the error the essay is complaining about.
Careful in the middle, sweeping at the edges
The measurement argument is stated with unusual care — Cutlefish concedes there is nothing inherently wrong with counting hours, labels flow metrics 'helpful, when focused', and flags its own uncertainty about motive. The overreach sits at the margins, in the two most quotable lines: that layoffs have failed to cover token budgets, and that the whole world has moved to pricing dumplings. Those are assertions dressed as observations, and they are what a reader will carry away.
Vendors' motives named, the writer's own left blank
The piece is unusually direct about whose interests bend the metric: vendors like return on tokens while the news is good, and executives need the sums to work for a board. Both are stated, neither is documented. Missing from the account is the writer's own position — this is a personal newsletter whose subject matter is precisely the value-architecture work companies are said to have given up on, and nothing in the text addresses that overlap.
Firm on the logic, thin on the world
We can say with confidence what this piece argues and that its core reasoning about hours survives inspection. We cannot say whether the market conditions it opens with are real, whether anyone is acting on the critique, or how the token argument resolves, since the text stops before it. One voice, no corroboration, an unfinished case.