Product1 publisher2 min readPublished
An AI chip insider argues inference should get cheap enough that teams stop counting tokens
A SiliconAngle opinion piece written from inside the AI semiconductor business says commodity inference would grow the market. Teams that ration tokens today still have to decide whether their caps are configuration or architecture.
The Product Desk · Product desk

What happened
- A SiliconAngle opinion piece written by someone who says they work in the AI semiconductor business argues inference should be commoditized instead of preserved as a scarce premium capability with high unit economics.
- The same piece says engineering teams already ration token usage, throttle API calls and cap deployments to keep cloud bills under control, and adds that Microsoft reportedly limits AI usage internally.
- For precedent it points to electricity, broadband, cloud computing and storage, all of which became cheaper and easier to deploy while demand grew instead of shrinking.
- Its central image is a $27 handcrafted truffle set against a 99-cent chocolate bar, where the cheap product widens the market because millions of people can afford it.
- It also argues that generated lines of code, requests per second and benchmark scores measure engineering activity and not the business outcome a buyer is actually paying for.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision Every team with a monthly token ceiling now has a planning choice: treat the ceiling as a permanent design constraint, or as a budget guard with an owner and a date to revisit it.
- constraint The piece supplies direction without magnitude or timing, so no product manager can move a launch date on it; the price in the current contract is still the one the roadmap has to survive.
- capability If per-call cost stops binding, the first features to come back off the shelf are the ones killed for volume: background agents, always-on assistants, anything that has to run continuously to be worth using.
- precedent A semiconductor insider arguing in public against premium unit economics hands buyers language for the next negotiation, where the thing to price is consistent delivered throughput.
A token cap starts as a finance decision and ends up as a product one. Someone writes the limit, and someone else answers on Friday for the feature that hit it.
The author of the SiliconAngle piece is explicit about whose margins the argument cuts against. "That may sound odd coming from someone in the AI semiconductor business," the author wrote [2]. The description of the current market is blunt: "AI is currently priced, marketed and deployed as a luxury product" [5]. On the fear that cheap inference shrinks the business, the piece says: "This isn't a race to the bottom. It's the dawn of a larger market" [11].
The price gap in the chocolate comparison is about 27 to 1, which is $27 divided by $0.99 [9]. That ratio describes a consumer good with a century of volume data behind it, and the piece uses it as an argument about direction [8]. It does not include inference price data or a date by which costs fall [16], and its statement that Microsoft reportedly limits AI usage arrives unsourced [7].
So sort the features you have capped or shelved against two questions. Does the value of the feature grow with call volume? And is per-call cost the thing actually stopping it, or is latency, accuracy or user trust the real blocker?
Volume-scaling and cost-blocked is the only quadrant where this argument changes a plan. Build those so the limit is a configuration value with an owner, and the piece's own list of candidates is ambient intelligence, autonomous systems and always-on assistants [10]. Volume-scaling but blocked by quality means cheaper tokens let you ship a weak feature more often. For everything that does not scale with volume, falling cost turns up as margin on what already ships; the piece argues each efficiency gain reduces operating costs for applications already in production [15].
Generated lines of code, requests per second and benchmark scores are useful engineering metrics and not business outcomes, according to the piece [12]. The product-side version of that mistake is calls per user, which tells you what the cap is doing and nothing about whether anyone got value. Two numbers do answer it: whether the people who hit the ceiling came back the following week, and how long a new user waits for one useful result.
Until a vendor publishes pricing a buyer can plan two years against, the cap stays. The choice available now is whether it sits in a config file with a named owner and a review date, or inside the design where nobody can move it.
What to watch
- Whether any inference vendor publishes volume pricing tiers or per-token commitments a buyer can plan a two-year roadmap against.
- Whether Microsoft confirms or quantifies the internal AI usage limits the piece attributes to it without a source.
- Whether procurement conversations start turning on delivered work per dollar instead of peak benchmark results.