Leadership1 distinct publisher3 min readPublished
A testing-tools CEO writing for Forbes Tech Council argues that the check on an autonomous coding loop matters more than the model inside it, which puts the definition of done back where a manager has to write it.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
A stop condition is a piece of specification work, and it is not the kind that delegates well. In Jiao's account the check that governs an autonomous loop does not read source code at all; it opens the running software, uses it the way a person would and reports what actually happened [12]. His example is an image tag carrying perfect accessible alt text over a dead source URL, which passes every static check and fails the only question that matters, whether a person can see the image [11]. Somebody has to decide, in advance, which of those two signals the loop is allowed to believe.
The failure mode he names is drift. An open loop with no meaningful check can report done even when the code does not work, because nothing in the system is permitted to disagree with it [10]. Jiao adds two properties that read more like procurement requirements than engineering taste: the check has to resist gaming, since an agent will satisfy a loose check rather than the underlying goal, and it has to preserve what already works, so this morning's fix does not reintroduce last week's bug [12]. Both are constraints on the people writing the check rather than on the model doing the work.
Andrew Ng's three nested loops give the timing problem its shape, with agentic coding measured in minutes, developer feedback in hours and external feedback over days or weeks [5]. Put an inner turn at ten minutes and an external cycle at one week, and the arithmetic is 7 x 24 x 60 divided by 10, or roughly a thousand inner iterations before the slowest signal returns [1]. The exact figure follows from the assumption, but the ratio is the managerial content: whatever the outer loop would have caught arrives about a thousand turns after the agent started acting on it.
A skeptic will point at the byline. Jiao is chief executive of a company selling AI-powered testing tools [1], and the claim that models and orchestration are becoming cheap and interchangeable while the verifier remains the scarce, load-bearing piece is exactly the sentence that business needs to be true [8]. The mechanism can still be inspected on its merits. What the provenance should push to the back of the queue is the pricing claim, that inside a strong loop a cheaper model can close the gap with a frontier model on whether the software works [13]. An incentive is not a refutation, but it is a reason to read that particular paragraph last.
One detail deserves to survive the argument. Agents are becoming primary users of developer tooling, and an agent cannot click through a dashboard or glance at a chart while it can reliably run a command and read the result [14]. That is a tooling constraint with a budget attached, because the systems a loop can verify against are the ones that speak in text.
The decision available this quarter is narrow. Who owns the definition of done for each agent loop, and whether that definition lives somewhere it can be audited rather than in the habits of whoever set the loop running. The consequence arrives later, once those definitions have hardened into the infrastructure every team's agents are graded against, and amending one means renegotiating what shipping means.
Ranked by verification strength, evidence, and original report placement.
Google Chrome engineering leader Addy Osmani gave the pattern a name and structure, and Claude Code creator Boris Cherny says he no longer writes prompts but writes the loop.
Andrew Ng framed the idea as three nested loops: an agentic coding loop measured in minutes, a developer feedback loop in hours and an external feedback loop over days or weeks.
An open loop with no meaningful check drifts and can report 'done' even when the code does not work, because nothing in the system is allowed to disagree.
A verifier that only reads source code is easy to fool: an image tag can have perfect, accessible alt text but a dead source URL, passing every static check while failing the test of whether a person can see the image.
At one inner-loop turn every ten minutes, roughly 1,000 agentic coding iterations can run inside a one-week external feedback cycle.
Jiao claims that inside a strong loop the model matters less than the feedback around it, and that a cheaper model can close the gap with a frontier model on the metric of whether the software works.
Distinct publishers with included, body-backed reporting in this cluster.
forbes.com
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
A twelve-word joke became a discipline, and one seven-step chain had no loop to remove1 distinct publisher
build
Harness choice moved token use 83-fold with the model held constant1 distinct publisher
security
An agent guard that runs on your laptop, and cannot tell you whether anyone keeps it on1 distinct publisher
product
ChatGPT Work's real ask is your Slack, and somebody has to say yes on everyone's behalf1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor essay, argued rather than measured
Every load in this story rests on a single Forbes Technology Council post. The one part that survives independent inspection is the dead-image-URL example, which is true by construction: a static reader of source cannot resolve a link. Everything above it — a month-long diffusion, a verifier that decides loop quality, agents outnumbering humans as tool users — is reasoning from the author's vantage point, and the four practitioners cited as corroboration reach the reader only in his paraphrase.
Nothing countable to count
For a story about a practice allegedly spreading overnight, there is not one number: no repositories, no CI runs, no tool installs, no customer named, not even a rough share of developers working this way. Named practitioners saying they now write loops is testimony about four people, and the author's company is never identified, so its usage cannot be examined either.
The pricing claim outruns its proof
The framing is calibrated in places and overreaches in one: telling readers that two years of assuming better software needs a pricier model is wrong, then conceding in the same story that no comparison was run, is a large claim resting on an analogy about drafts. Add a one-month timeline no one can verify and a headline declaration that this is the most important new software skill, and the rhetoric sits a good distance ahead of what is shown. The gap is not cynical — the mechanism arguments are sound — it is simply unmeasured.
The scarce piece is the thing the author sells
Read the argument backwards and it is a market map: models commoditise, orchestration routinises, and the one irreplaceable, quality-determining component turns out to be automated behavioural testing — which is the category the author's company builds in. Forbes discloses his role in the first line and the invitation-only council format at the end, so nothing is hidden; the shape of the interest is just unusually tight. The unnamed company is the notable omission, since it keeps the reader from connecting the thesis to a specific product.
Confident about what this is, not about whether it is true
We can be fairly sure of the document's nature and its limits, because it labels both itself: the author's role is stated up front and the missing comparison is conceded in the text. That makes our reading of provenance and incentive solid. What we cannot judge from here is whether the loop thesis holds at scale, and a single-source story with no numbers leaves no room to test it.