Build1 distinct publisher3 min readPublished
The adjusted figure is the honest one, but it still only counts whether calls came back. That is why 20 customers who topped up in more than one month, carrying 70.8% of lifetime revenue, is the harder number.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The arithmetic on the gap is the finding. At 95.5% of 299,668 August calls, about 13,485 requests were not successes [5] [17]. At 98.9%, about 3,296 were not [6] [18]. The difference is roughly 10,190 calls, near three quarters of the whole failure bucket, and those were customer prompts the upstream model refused on policy grounds rather than anything the gateway did wrong [19].
For that split to be a measurement instead of an opinion, the forwarding layer has to distinguish an upstream policy refusal from an upstream 5xx and from its own timeout. The post attributes the difference to refusals but does not describe how they are identified [24]. If refusals arrive as generic HTTP errors from some providers and as structured refusal payloads from others, 98.9% is a judgement call with a decimal point on it.
The July-to-August bridge has the same shape of problem. Volume grew 2.6x, from 116k to 300k calls, while the reported success rate went from 93.9% to 98.9% [7]. But 98.9% is the moderation-excluded number and 95.5% is the raw one, and the post does not say which basis the July figure uses [21]. If July is raw and August is adjusted, part of that improvement is a change in definition. That matters, because "volume nearly tripled and stability went up" is the load-bearing claim of the whole systems section.
The retention number does something a counter cannot. The author wired a very cheap upstream for a popular video model, and several users tried it once and never came back, with no errors and every request returning 200 [10]. Running the same prompt through the official channel and comparing output side by side showed a watered-down clip, most likely a lower resolution upscaled or a cheap tier passed off as the premium one [11]. A 200 came back regardless of whether the pixels behind it were right. The rule that came out of it is official model channels only, never a reseller priced far below market [12], and the instrument that caught the substitution was retention, not the error rate [13].
The 70.8% figure needs care before anyone borrows it. It is a share of all revenue ever collected, not of the $8,880 month, and the cumulative total is not published, so you cannot back out dollars per repeat customer [8] [22]. What survives is a count of people who chose to pay a second time, which no dashboard configuration can flatter.
The build decision follows from the same place. The product is one API key in front of 182 models behind an OpenAI-compatible interface [9], and off-the-shelf open-source gateways were rejected because price and stability both live in the forwarding logic between request and upstream, which is exactly the layer such a gateway occupies [14]. A bug in a dependency means filing an issue and waiting; a bug in your own code gets fixed that afternoon [15]. That trade only pays if you can genuinely ship that afternoon, and here every mechanism was retrofitted after a specific production failure rather than designed up front [16], with an AI coding agent writing it under a product manager's direction [1].
Whether 98.9% transfers to your gateway depends on whether you can prove which bucket a given call landed in.
Ranked by verification strength, evidence, and original report placement.
The author is a product manager who states he cannot write code, and says every line of the product was written by an AI coding agent working under his direction.
Revenue taken from the production database with admin accounts excluded was $30.32 six months ago and $8,880 this month.
Trailing 30-day revenue is $9,180, and the author says he has not yet crossed $10k per month, expecting it in September.
In August the gateway served 299,668 API calls at a 95.5% success rate.
Excluding content-moderation rejections, described as a customer's prompt refused upstream rather than the author's system failing, the August success rate is 98.9%.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
One status code, two opposite remedies: the 429 that cost ARGUS seven minutes1 distinct publisher
build
OpenAI-compatible image APIs normalize transport, not fallback routing1 distinct publisher
build
Your token ratio, not the leaderboard, decides which model is cheap1 distinct publisher
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One byline, one private database
Every number this story stands on — $8,880, 299,668 calls, 70.8% — was queried by the man who sells the product and published under his own name on dev.to. What an outsider can check is only the internal arithmetic, and it holds: 95.5% and 98.9% of 299,668 really do bracket a ten-thousand-call moderation bucket, and $30.32 to $8,880 really is about 293x. Consistency is not corroboration. The two places the post declines to show its work are the two that carry its argument: which basis the July 93.9% sits on, and how a moderation refusal gets distinguished from his own gateway dropping a call.
Real traffic, narrow base
Three hundred thousand calls and roughly nine thousand dollars in a month is a live business with paying users, not a demo — and the growth from 116k to 300k calls month over month is the kind of movement that is hard to stage. It is also thin where it counts: twenty accounts that have ever paid twice hold 70.8% of everything earned, so most of the demand sits with a group small enough to name. All of it is disclosed by the operator; no customer, upstream provider or third-party tracker appears anywhere in this reporting.
Headline outruns a mostly candid post
The body of this is more restrained than its own title. He publishes $9,180 and says plainly he has not crossed $10k, names the unadjusted 95.5% next to the flattering 98.9%, and spends a full section arguing that his best-looking metric lied to him. Then the framing takes it back: 293x is what happens when your denominator is $30.32, and the marquee 93.9%-to-98.9% jump quietly compares a number of unstated basis against the adjusted one, which is the one place a skeptic would push. Overstated at the edges, honest in the middle.
Founder selling, in a genre that rewards it
This is the operator of apimodels.app writing about apimodels.app on a platform where the from-zero-to-revenue post is a known traffic instrument, and he alone chose which rows to publish. Two secondary incentives deserve naming. The build-it-myself argument doubles as a differentiation pitch against anyone running the open-source gateways he declined. And the substitution story asks readers to accept, on description alone, that an unnamed discount upstream shipped degraded video — a serious accusation that also happens to justify his premium pricing.
Checkable arithmetic, uncheckable books
We can be fairly sure what was claimed and reasonably sure the claims are consistent with each other; we have no way to test whether the underlying rows exist as described, and one publisher gives us nothing to triangulate against. The operational lesson about silent substitution stands on its own logic regardless — a 200 response is not a quality signal — which is why our read of that part is firmer than our read of the numbers.