Build1 publisher3 min readPublished
A Claude user reports Opus 4.6 handing subtasks to Opus 5 subagents that burn the limits
A dev.to roundup of paid-tier complaints on Reddit turned up ads inside a paid ChatGPT plan, an annual unlimited allowance with an end date, and one latency figure from one developer. Only the delegation report is checkable from a trace.
The Engineer · Build desk
What happened
- A dev.to roundup of complaints from people who pay for AI tools quoted six Reddit posters across four subreddits, and said the quotes were verified against the live posts.
- Ads appeared in u/Loganh1976's ChatGPT Go conversations months after they subscribed, and they responded by trying Claude, Gemini and Grok for the first time.
- On r/cursor, u/General-History-5917 described the annual Cursor Pro plan's unlimited auto mode ending, which forces a choice about which more-metered tier to move to next.
- u/SemiMagnum warned on r/ClaudeAI that choosing Opus 4.6 does not stop the main model handing subtasks to Opus 5 subagents that burn the limits.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision When a flat allowance inside an annual plan has a scheduled end, renewal becomes a metering choice, and a budget line that was fixed last year turns into a function of usage.
- exposure If delegation works as u/SemiMagnum described, the model picker is not a spending control, and the team that chose the older careful model pays for the newer model's tokens.
- constraint None of the week's posts include a prompt, a model version or a request count, so a buyer who wants to know whether its own latency moved has to time pinned versions itself.
- precedent Advertising inside a tier people already pay for makes the same move easier to defend one rung further up the price ladder, after ads reached the free plan in Europe.
Delegation is the one complaint in the week a team can check against its own records. u/SemiMagnum, posting on r/ClaudeAI, wrote that picking "the good old Opus 4.6" for careful work is no protection, because the main model may still hand subtasks to "verbose and token-consuming Opus 5 subagents ruining your work and burning the limits" [10]. If that is how the orchestrator behaves, the selector sets the model for the top-level call and the system chooses the rest. The model id returned on each call settles it, next to the tokens billed against it.
One developer supplied the week's only number, and it measured latency. Short requests on r/GeminiAI "used to take between 4 and 15 seconds" and were taking "50s or more" over the past week, the developer wrote [11]. At the endpoints that is 12.5 times the old best case and 3.3 times the old worst case [13]. The cause was offered as a guess: "Google reduces capacity for older models to push people to move to the latest" [12]. The post gives two recalled figures. For that number to say anything about another team's workload, the comparison would have to hold the prompt, the model version and the region fixed, and it would need a measured before-and-after distribution.
Pricing complaints are plain in the posts themselves. "I have spent most of this year exclusively with Claude," wrote u/-AMARYANA- on r/cursor on 26 August. "I just feel ripped off at this point. I see the value is dropping each month as other models close the gap." [7] dev.to describes the pattern behind the expiry as a generous flat allowance introduced to win the habit, then converted to usage-based pricing once the habit is formed [15].
The odd reports came from the top of the price list. u/SweatyActuator2119, on the "20x" plan, said on 26 August: "I am seeing a lot of hallucination and scope drift, and extremely slow execution," describing new models that "constantly over-engineer to the point where scope is an afterthought" [9]. dev.to concludes that the priciest and most thoughtful modes are also the slowest [16].
Ads inside a plan sold as the clean paid alternative make a second charge. dev.to describes that as an old subscription move: you get monetised again once you are inside [17]. "I genuinely think putting intrusive advertising into a paid plan is one of the worst strategic decisions OpenAI could have made," wrote u/Loganh1976 [5].
Six posters across four subreddits measure nothing [14], and the write-up says as much: it quotes real users "to show the texture of the frustration, not to pretend a forum is a poll" [2]. Of the complaints collected, only the model id on the subagent calls leaves a record a buyer can pull for itself.
What to watch
- Whether Anthropic documents which model serves subagent calls and whose limits those tokens draw down.
- Whether Cursor publishes the metering of the tier that annual subscribers land on when unlimited auto mode ends.
- Whether anyone posts pinned-version latency timings for older Gemini models, since the throttling explanation is currently one developer's guess.