Build1 publisher3 min readPublished
Armin Ronacher's unattended 35-hour GPT-6 Astra run wrote 75,000 lines of unusable code
Armin Ronacher let GPT-6 Astra code alone for 35 hours, spending about $1,200 on a net 75,000 lines of "absolutely nothing of value". His traces show habits from training that pays for finished tasks, and they reach the code when nobody reviews it.
The Engineer · Build desk

What happened
- The agent made 79 commits and consumed around 1 billion tokens in API terms before Ronacher stopped it.
- The single prompt asked for a Python with virtual threads and lexical scoping, and the agent ran it as a software factory whose subagents traded about 1,400 messages.
- To run a single Windows test, the agent chained Python, subprocess, prlctl exec, Node.js and PowerShell.
- Its task names drifted from 1, 2, 3 to '8b2c2b2b checkpoint1', and the code carried magic indexes and a switch listing case 30 through case 72.
- Earendil, Ronacher's company, measured agent code as about twice as verbose and twice as eroded as human repositories on SlopCodeBench.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost At about $15 a commit, the invoice for an overnight run looks like a human's, so the spend alone will not reveal that the output is unmaintainable.
- exposure String-replace edits fail without a report when the anchor text is missing or duplicated, so an unreviewed agent can pile up broken changes that surface only when someone reads the diff.
- decision Teams running agents overnight have to decide where a human reads the output, since in Ronacher's account nothing else pushes back on code quality.
The edit habit is the one I would worry about first. When a subagent needed a new function in a C file, it skipped the edit tool. It piped in a Python heredoc that read the file, spliced the code in with str.replace and wrote it back [5]. The dev.to write-up of Ronacher's post notes that this works until the anchor string appears twice or not at all, and then nothing tells you [6]. In a run with no human review, nobody was placed to notice [1].
The tests were trimmed for tokens too. The committed unit tests had no whitespace, a format Ronacher measured as "10 % more token efficient" than the same file after ruff format [8]. He tweeted it as "Astra is a code golfer." [9]
Ronacher's explanation is about what training pays for. He wrote that the model is "greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for 'shitty code.'" [11] The dev.to account extends it. Terse tool calls help the agent's own token budget, so the habit leaks from the tool calls into the codebase [12]. The fewer humans look, the less anything pushes back [12]. Ronacher titled that section "It's AGI If You Don't Look." [13] On Hacker News, the commenter nojs described labs moving reinforcement learning "from 'being rated as useful according to human feedback' to 'succeeds at long horizon tasks'" [14].
The bill looks ordinary. $1,200 over 79 commits is about $15.20 a commit [1]. The dev.to write-up rounds that to $15.50, calls it roughly a human rate, and adds that a human would have stopped and asked a question somewhere around hour three [17]. Averaged over the run, each commit carried about 950 net lines, and the agent added about 2,100 net lines an hour [2][3].
How far this transfers depends on the work. The evidence is one prompt and one run, on a goal Ronacher made ambitious on purpose [1][4]. The only wider measurement is Earendil's, and Earendil is Ronacher's company [15]. For the result to apply to a team, its agent tasks have to be long-horizon and unread until the end. Short tasks with a reviewer in the loop were not tested.
In my view, the change this supports is a scheduled human read of agent output, at commit boundaries or on a clock. The run shows what zero review produces on a hard task. It does not measure which review interval recovers usable code. Ronacher's own summary on X was "Astra is really, really cool but I cannot current trust it for my present day engineering" [16]. The tweet reached 182,000 views [16].
What to watch
- An independent SlopCodeBench run outside Earendil that confirms or contradicts the twofold verbosity and erosion figures.
- A repeat of a long Astra task with scheduled human reads, to test whether checkpoints recover usable code.
- Any change in how GPT-6 Astra is trained or scored on code quality, the gap Ronacher blames.