Build1 publisherNot yet confirmed elsewhere3 min readPublished
The agent shipped working code and five products in one codebase
A post-race audit of PricePulse, built unsupervised by Claude, found features that mostly worked and an organisation that did not. The failures were startup failures, not model failures.
The Engineer · Build desk
What happened
- The author expected the audit to find broken code, a pile of half-working features and sloppy logic, but almost everything Claude built actually worked when taken piece by piece.
- The project audited is called GetPricePulse, a SaaS pricing intelligence product, and it is Claude's entry from The $100 AI Startup Race.
- The $100 AI Startup Race is a season-long challenge run by the author in which seven AI agents each get $100 and full autonomy to build a real startup from scratch, with no human coding and no product manager in the loop.
- Each agent picked its own idea; Claude picked SaaS pricing intelligence, named it PricePulse, and kept building on it for the entire race.
- The piece is a full production audit of Claude's build, PricePulse, done after the race and before the author would let anyone treat it as a real business.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
An audit of a SaaS product built end to end by an autonomous coding agent found that almost everything worked when taken piece by piece, and that the product as a whole did not cohere [1] [13]. The person who ran the audit went in braced for broken code and came out with an organisational finding instead [1], which moves the interesting risk in agent-built software from correctness to product judgement.
The artefact is GetPricePulse, a SaaS pricing intelligence product, and it is Claude's entry in The $100 AI Startup Race [7]. According to the organiser, the race gives seven AI agents $100 each and full autonomy to build a real startup from scratch, with no human coding and no product manager in the loop [8]. That is $700 of total capital across the field [16]. Claude picked pricing intelligence, named it PricePulse, and kept building on it for the whole race [9]. The audit was run afterwards, before the organiser would let anyone treat the thing as a real business [10].
The output inventory is not small: more than 1,300 HTML files, hundreds of content pages, a pricing database, calculators, monitoring features, full authentication, Stripe payment integration, and email infrastructure [11]. The organiser writes that he would have expected months from a solo developer asked to build that scope on a normal timeline, and that Claude did it across the race's running sessions [17].
Feature by feature, it held up. Authentication let people sign up and log in [2]. The pricing database was real content rather than placeholder, and the calculators worked [4]. Stripe, in the auditor's own phrasing, processed at least one pricing tier correctly [3], which is a sentence worth reading twice: it confirms a path, not a billing system. There were real, concrete engineering bugs, but the organiser says they were not the headline finding [12].
The headline finding was the seams. The biggest problems sat in the relationships between features, the places where five individually reasonable decisions added up to something incoherent [13]. By the end of the race PricePulse had become five different products sharing one codebase: a SaaS pricing publication and database, a monitoring SaaS, a FinOps toolkit, a competitive intelligence product, and one more [14]. From the commit history, the organiser reads the agent as optimising for speed, feature creation, shipping and monetisation experiments, and not for correctness, coherence, or whether a decision still makes sense three weeks later [5].
That is a familiar failure and it is not a model failure. Nobody was deciding what PricePulse should be; nobody was saying "we have enough pricing tiers now" or "this feature doesn't belong here" [15]. The organiser's conclusion is that these were startup mistakes rather than AI mistakes, the kind any fast team makes when velocity is the only metric [6]. The agent did not lack the ability to build. It lacked the function that kills work.
Two things to watch. First, the engineering bugs, which the organiser has flagged as interesting on their own terms but had not yet detailed [12]; a seam-level story is only complete once you know what the code-level defects cost. Second, whether the other six entries fail the same way, given that all seven agents, asked independently what agents still cannot do, converged on the same answer [18]. If unsupervised agents reliably fan out into product portfolios nobody asked for, then "no product manager in the loop" is not an experimental condition. It is the defect.