Product1 distinct publisher3 min readPublished
The firm's rule is to meter at the highest unit you can measure and defend, which sounds obvious until you try to instrument a completed unit of work and find your billing system only knows how to count tokens.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
product
Rillet's $100M reads as proof mid-market ERP is rip-and-replace, mostly at the cheap end1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
product
OpenAI's plan to hand everyone a coding agent leaves the hard part to the model1 distinct publisher
A support lead can tell you how many conversations her team handled last quarter and be close. Ask the same person to forecast context length, retrieval volume, retries, reasoning time and output tokens for next quarter and she is guessing, which is a16z's core complaint: token pricing turns a straightforward ROI question into a separate compute-forecasting exercise for every application a company deploys [17].
When a buyer asks about tokens, the post argues, they are usually asking for a benchmark, or for a way to trace spend to a department, project, client or invoice [15]. Those are different requests and only one touches your price. a16z calls the benchmark version false precision, on the grounds that a token flowing through one product does not produce the same work as a token flowing through another [16].
The figure that will get quoted is the survey: 27 of 50 technical AI buyers preferred credits tied to recognizable work, and 14 preferred tokens [8]. That is 54 percent against 28 percent, a ratio of about 1.9 to 1 [9] [11], leaving nine respondents whose preference the post does not report [10]. It is stated preference, sample of 50, collected by a firm that describes the recommendation as coming from its own work with companies [4]. What would actually settle it is duller data: whether credit-priced vendors renew better, whether customer forecasts land near the invoice, whether accounts expand after a repricing. None of that is in the post.
The mechanism holds up without the survey, though. If your rate card is denominated in tokens, every price cut at the model layer reaches you as a smaller invoice rather than a wider margin, and you have taught the customer to value your data, workflow and orchestration at zero [3] [6].
The test worth taking into the pricing meeting turns on two things: whether you can measure the unit the same way for every customer, every month, and whether you can attribute the result to your product in a way that survives an argument with procurement.
Yes to both means you can price the outcome [5]. Yes to measurement and no to attribution means you price the work delivered, which is the box credits are built for [5]. No to both means you are selling model access, and tokens are the honest unit [5]. Strong attribution paired with weak measurement gives you a case study, not a pricing model, and you cannot invoice it monthly.
The failure inside the middle box is the one to plan for. Credits that are cost-plus tokens under a friendlier name still meter infrastructure, and the customer discovers this the first time a heavy month lands [7]. A real work unit, an account brief completed or a code change implemented, means you absorb the variance when one instance costs ten times another [12]. That is the tradeoff, stated plainly: you take on cost volatility in exchange for a price the buyer can forecast, so the p50 and p95 cost of one completed unit is the number you need before the rate card exists, not after. a16z's own warning is that this decision is difficult to undo [19], and it is easier to publish a token price on Monday than to explain to the same account in eighteen months why the unit changed.
Ranked by verification strength, evidence, and original report placement.
Token pricing began at the model layer: when OpenAI launched its API in 2020, charging for the computation a model consumed was a sensible way to meter raw inference.
ChatGPT's debut two years after the OpenAI API launch helped spark a wave of applications built on that infrastructure that combine proprietary data, tools, orchestration, integrations and workflow logic to complete work on a customer's behalf.
a16z says that based on its work it is often a mistake to carry the model layer's pricing logic into the application layer, and that companies should price at the highest layer of value they can reliably measure, attribute and defend.
a16z's tier rule: if you sell model access, price tokens; if you turn models into useful work, price the recognizable value unit, often through credits; if you deliver a clear and attributable business result, price the outcome.
a16z says credits do not solve the problem on their own if they are merely cost-plus tokens, because they still meter infrastructure rather than value.
In a16z's survey of 50 technical AI buyers, 27 preferred credits tied to recognizable work while only 14 preferred tokens.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One firm, one unpublished survey
The historical framing is checkable — the 2020 API did meter inference in tokens — and almost nothing after that is. The 27-versus-14 split carrying the headline arrives with no sample frame, no fielding date, no question wording, and nine of the fifty respondents simply missing from the tally. The strongest assertions, on margin lock-in and on irreversibility, rest on the phrase 'based on our work', which is a credential rather than an evidence trail.
No pricing moves on the record
A stated preference in a survey is not adoption. Nobody in this reporting is named as having moved from tokens to credits or outcomes, no dated pricing page changes, no revenue or retention effect. Voice AI and copilots appear as expected trajectories, and the account brief, the implemented change and the completed pipeline run are hypotheticals used to define a unit — there is nothing here to count.
Doctrine ahead of data
The rule is stated with the assurance of settled practice, and our own headline inherits that assurance: 'nearly two to one' is honest arithmetic performed on a figure no one outside a16z has seen, and the nine silent respondents could compress the ratio considerably. The prescription is coherent and may well be right — pricing to a unit whose cost is collapsing is a genuine trap — but a survey sentence and a set of hypothetical agents are being asked to carry a whole pricing doctrine.
Advice from an investor in the advised
This is published on Andreessen Horowitz's own site, aimed at founders the firm funds or hopes to, and the conclusion points toward higher-margin pricing at the application layer — which is where venture returns are made. The survey used to support it is also the firm's. None of that makes the reasoning wrong; it does mean the only party attesting to the argument benefits if the argument wins.
Clear source, narrow base
What a16z said and where each number came from is unambiguous, which makes this story easy to describe accurately. Judging whether it is true is harder with a single interested publisher, no corroboration and no observed outcomes — so treat our attribution as firm and our verdict on the substance as provisional.