Build1 distinct publisher3 min readPublished
The sharpest complaints about AI coding tools this fortnight came from heavy users, and they landed on metering and product behaviour rather than model quality. One of them is a test you can run on your own traffic.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
AI's 4x code generation ships with a doubled review cycle and tripled post-merge fixes1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
build
Waku 0.1.0 bets the product is the control plane, not another coding agent1 distinct publisher
Start with the field. Copilot's own model calls carried `X-initiator: agent`, and by j0selit0's account those calls did not draw down premium quota [5]. He then set the same value on requests he sent himself. The model answered, but the quota did not move and the request never appeared in billing analytics [6]. dev.to describes this as one developer's own testing rather than an audit, and characterises the mechanism as a billing system keying off a header the client controls [7].
That last phrase is the engineering story. A server handling that request already knows the authenticated seat and the model it is about to invoke. It does not need to be told who initiated the call, because the caller is the least reliable party to ask. Asking the client to declare whether its own request is billable is a generous default.
Whether this reaches your invoice depends on your billing setup and on whether the behaviour holds. Your org has to be paying for premium requests rather than working inside an included allowance, and the endpoint has to still behave the way one traffic capture said it did, which is an observation about someone else's session and not a measurement of yours [7]. Both are checkable in an afternoon with a logging proxy in front of the editor, which is also the only place a count originates that did not come from the vendor.
The Cursor complaints in the same roundup are usually read as taste, and one of them is not. jmuguy's list includes the editor changing his model to the latest Grok without prompting [3]. That is a provenance problem: a diff produced last Tuesday was produced by whatever model the client had selected at that moment, and if the client can reselect silently, the commit no longer identifies its own generator. Your review record ends up unable to tell you which model to credit or blame.
palmotea, commenting on the money thread, put the incentive plainly: leadership cares about velocity and cost, and would replace software development with worse AI software development if the numbers favoured it [9]. Cost, in that sentence, means the number on the dashboard. Two of the four grievances in this batch are about whether that number is a measurement.
dev.to is careful about what these are: each quote was located at its comment permalink and reproduced verbatim with username, platform and date, and framed as one practitioner's experience rather than a verdict [10]. The publication's own framing is that this batch came from the people using these tools most [2]. Of the specific claims, exactly one is a test rather than an impression, and reproducing it costs a proxy and an afternoon.
Ranked by verification strength, evidence, and original report placement.
A Hacker News user posting as jmuguy wrote on 20 August that Cursor's flagship app 'is getting enshittified at a surprising clip', constantly pops up and interrupts work pushing new features, 'changes your model to whatever the latest Grok is without prompting', has a 'mystery meat UI that is constantly changing', pushes cloud agents in ways designed to trick you, and that he was actively looking at alternatives.
dev.to states that every quote was located at its comment permalink and reproduced verbatim, listed with username, platform and date, and framed as experiences rather than verdicts.
A dev.to roundup titled 'Enshittified at a Surprising Clip' reviewed a week to fortnight of Hacker News comments on AI coding assistants and found the grievances were not existential but about the bill, the UI, and the effort of reading what the model just wrote.
dev.to says this batch of complaints came from the people who use AI coding tools most, in contrast with fortnights when complaints come from people who barely use it.
On a thread about Cursor's new GitHub-style features, a commenter posting as nikolay said a feature gated behind payment where he expected an open one was 'reminding me to cancel my Cursor subscription. I am paying for it but not using it because I don't need it as it offers me nothing.'
Commenting on the Codex billing story, a user posting as palmotea wrote that leadership does not care about dumb bugs, cares about velocity and cost, and 'would totally replace all software development with worse AI software development in a heartbeat.'
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 28, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quotes verified, mechanisms not
dev.to's discipline is about quotation, not about facts: permalinks, usernames, dates, verbatim text, and an explicit refusal to treat a forum post as a filed complaint. That gets you a reliable record of what people said and nothing more. Nobody re-ran j0selit0's header swap, GitHub was never asked whether X-initiator gates premium quota, and the tenfold Bedrock bill arrives with no primary account behind it. The strongest claims are the ones about who said what.
No usage figures of any kind
A handful of comments on one forum cannot tell you how many Copilot seats bill this way, how often the header relabelling works, or whether the Cursor behaviours reach anyone beyond the people describing them. There are no seat counts, no vendor disclosures, no rates. Converting five quotes into a prevalence would be inventing the number.
Title harder than the reporting under it
Our own headline says a client-set header decided whether a request spent premium quota; the reporting says a developer says so and calls it 'not an audit'. That gap opens in the framing, not the analysis — dev.to hedges the pairing, separates the commenter's verdict from his specifics, and refuses to treat any of it as dispositive. Small overreach, and it is at the top of the page rather than in the substance.
A downside beat, and no vendor in the room
This runs on a channel whose whole premise is AI's downside, and its witnesses are people mid-cancellation or already shopping for alternatives. Cursor, GitHub and AWS appear only as subjects. None of that makes the header finding wrong — it does mean every incentive present pushes the same direction, and the one party able to settle the metering question has not been asked to.
One publisher, one forum, one afternoon from certainty
We are confident about the record of the conversation and unconfident about the world it describes. A single publisher, a single forum, pseudonymous witnesses, no vendor voice. What keeps this from the floor is that the key claim is falsifiable by anyone with a Copilot subscription and a proxy — the moment someone runs it, this assessment should move sharply in one direction or the other.