Build1 distinct publisher3 min readPublished
The Plus window meters agent work rather than elapsed time, so tests and browser verification bill far faster than the clock runs. That makes the tier you sit on, not the model you picked, the thing you have to plan around.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
OpenAI's $8 reset button turns a flat ChatGPT seat into an unbudgeted line item1 distinct publisher
build
Codex learns to click: the coding agent stops typing patches and starts operating the machine1 distinct publisher
build
Codex remembers by popularity: who actually decides what your agents forget1 distinct publisher
build
A decision rule sorts Codex and CodeRabbit by the unit of work each owns to completion1 distinct publisher
The meter counts work, not seconds. According to the dev.to writeup, agentic usage tracks context size, reasoning, tool calls, tests and commands [7], and OpenAI's own documentation points at model, task complexity and context [8]. All of those scale with tokens moved, not with time passed. A verification loop moves a lot of them, because each build, each lint run and each browser step resubmits state to the model [3]. Wall clock is the wrong denominator for that work, which is why the label and the experience come apart.
The arithmetic is worth doing once. Five hours is 300 minutes, and the allowance emptied in about 24 [4]. That works out to 12.5 minutes of budget spent per minute at the keyboard [1]. Read the five-hour figure as a refill interval and the confusion clears up, but the number also loses its use as a planning unit.
Two live buckets are where this bites. The author reports being locked out with substantial weekly usage still on the account [6], and says multiple Plus users on OpenAI's community forum describe the same pattern [10]. When a small bucket and a large one are both in force, the small one governs, and the honest question becomes how many feature-sized tasks fit inside one window. On this evidence, one [2][4]. If the windows are fixed and consecutive, a day holds 4.8 of them [2].
Before anyone ports that number: this is a single task in one Markdown notes app [2]. For 24 minutes to be your figure too, you would need comparable repo size, the same model, tests genuinely executing, browser verification in the loop, and similar context on every turn. Change any of those and the burn rate changes. What this measures is one task shape, and it should not be mistaken for a rate card.
The part I would treat as a design defect rather than a pricing complaint is the resume path. When the wall arrived, the model still held the change set and the check it was running, and the user cannot hand the job back until the limit resets [5]. So the cheapest state to occupy is "not started" and the most expensive is "almost done", and the metering offers no way to prefer the first.
The author also notes that Plus briefly ran on the weekly limit alone, and that he preferred it [9]. That is the more useful datum for anyone sizing a tier, because it shows the five-hour window is a policy setting that has already been switched off once, and it is worth remembering as such rather than treating it as a fixed fact about capacity. In my context the sane response is measurement: spend one window on a representative task, record where it stopped, and size the plan from that rather than from the number in its name.
Ranked by verification strength, evidence, and original report placement.
The author states that agentic coding usage depends on context size, reasoning, tool calls, tests, commands and other factors.
The writeup says OpenAI's current documentation states that usage depends on things like the model, task complexity and context.
OpenAI recently brought back the 5-hour usage limit for Codex on the Plus tier, according to a dev.to writeup by a Plus user.
The author gave Codex one task in a Markdown notes app: add Find and Replace, including Find in the Edit menu, keyboard shortcuts, match navigation, Replace and Replace All.
Codex implemented most of the feature, wrote tests, ran the build and linting, and had started browser verification of the UI.
The author's entire 5-hour allowance was consumed after about 24 minutes of actual elapsed time.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 29, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One stopwatch, no receipts
Everything measurable here is one person's recollection of one session. The 24-minute figure has no usage dashboard, log line or error screenshot behind it, the documentation is paraphrased rather than quoted or linked, and OpenAI is neither cited announcing the window nor asked about it. The internal detail is good — the agent's steps are named in order, which makes the burn plausible — but plausible is not verified.
One desk, plus a chorus we cannot hear
Real usage is present — a paying subscriber ran a real feature against a real app and hit a real wall — but it is one desk. The forum users said to be reporting the same thing are described, not counted or linked, and no figure appears for how many Plus subscribers use Codex or how often they hit the window. What does come through is that the metering is live on the tier today.
The measurement is narrower than the verdict
The headline verdict — practically unusable — is drawn from one task on one afternoon, and the author's own concession that agent usage varies with context, tools and complexity cuts against generalising from it. The specific finding is modest and interesting; the frame around it is a product obituary. The tighter statement is the one our own coverage leads with: the window bills work, not time, so the tier you sit on is what you plan around.
An unhappy customer writing on his own platform
No vendor is speaking here and nothing is being sold. The pull is the ordinary one behind a paying subscriber's public complaint: pressure on OpenAI to loosen the cap, and the attention a sharp title earns on a developer blogging platform. That cuts both ways — the author has no reason to invent a session he actually lived, and every reason to describe it in its worst light.
Believable, unconfirmed, and narrow
We hold the mechanism with reasonable confidence and the magnitude with little. Metering by work rather than clock time is consistent with the author's own reading of the documentation, and the stranded mid-verification session is the kind of detail people do not invent. But with one publisher, one session, no artefacts and no vendor response, the specific 24-minute figure and the sweeping conclusion drawn from it both sit on a single unchecked account.