Build1 distinct publisher3 min readPublished
Headlong is Apache-2.0, so a team can read the loop before running it. What nobody can read is a benchmark, and the shared thought stream is also the privacy problem.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Who decides how many model calls happen matters more than the line count. The main loop, called Thinker, repeatedly invokes a component named shellm, which asks a language model for reasoning, Bash commands or both, executes the commands, and continues until the model sets a FINAL environment variable [15]. There is no checklist unless the agent writes one, and each next thought is selected from earlier thoughts and new observations [13]. Iteration count is therefore a property of the agent, not of the request log [20], which is the mechanism under runtimewire's warning that the continuous loop can increase token spending [6]. Tiered compaction keeps recent events in full detail and summarizes older history [16], which shrinks the context per call and does nothing to the number of calls.
The shared stream has the same double character. Audel, the agent Laude has been running for several weeks, works as one agent over one stream reached through Slack, Telegram and a mobile app [8]. The behavior Laude puts forward as the point depends on that arrangement: following what different people are doing, returning to an old subject, contacting someone with no new prompt [9]. runtimewire flags the same stream as a cross-user privacy risk [7]. One mechanism, two names, and the team that installs it is choosing on behalf of everyone in the channel.
Shell access is not the delta. OpenAI's Codex CLI can read, modify and run code locally [17], and GitHub Copilot CLI can execute shell commands subject to permissions or approval [18]. An approval prompt assumes a human is present at the moment the command runs. Laude's clearest example, dated August 5th, happened while nobody was interacting with the agent: it checked a background recall process it had built, found the code reading an environment variable that was never set, searched the codebase to confirm, rewrote the process to read from a pipe, verified the repair, and the fix was merged, with the log putting the sequence at 48 minutes [12]. Attendance is what changed, not privilege.
What stands in for evaluation is that log. More than 50 commits from Audel's fork were pulled into the main repository [11]. On its first day the agent audited eight stale Git branches and came back 10 minutes later to correct its own count [10]. There are no independent performance results, and Laude says it judges changes to Audel mainly by qualitative observation [14], so every quantified figure on offer is first-party [19]. Laude's grantmaking is anchored by Konwinski's $100 million pledge [5], which is part of why the artifact turned up as Apache-2.0 code on GitHub instead of a product with a score attached [3].
The compensating handle is instrumentation. Thoughts and actions are stored in a directed graph of JSONL files that supports forks and merges [15], and the core is 9,900 lines of Bash in bin/ and thinkers/ [4]. Absent an eval, that graph is the only way to reconstruct what an unattended agent spent and what it touched, after the fact.
Ranked by verification strength, evidence, and original report placement.
Andy Konwinski's Laude Institute and MIT released Headlong on August 25th, an open-source harness that allows an AI agent to keep generating thoughts, choosing projects and taking actions after its users have stopped talking to it.
The August 25th launch post calls the design "persistent agency": messages become observations inside a continuous thought stream instead of opening isolated sessions, and the agent decides whether to answer, wait, investigate something else or contact a user later.
Developers can inspect and modify the Apache-2.0 Headlong code on GitHub.
Laude says Headlong's core is currently less than 10,000 lines of Bash, specifically 9,900 lines in bin/ and thinkers/.
Laude says its grantmaking and research programs are anchored by Konwinski's $100 million pledge.
Headlong's continuous loop can increase token spending, according to runtimewire.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed mechanism, first-party outcomes only
The architectural description is concrete and checkable in public Apache-2.0 code (Thinker/shellm loop, FINAL termination, JSONL thought graph, tiered compaction, 9,900-line Bash core), which lifts evidence above pure announcement. But every claim about the system being useful -- the eight-branch audit, 50-plus merged commits, the 48-minute self-repair -- rests on Laude's own deployment and logs, there are no independent performance results, and Laude evaluates Audel mainly through qualitative observation. One publisher covers the story.
One internal deployment plus an open-source drop
Observed adoption is a single Apache-2.0 release and a single first-party deployment: Laude running the shared agent Audel internally across Slack, Telegram and a mobile app for several weeks, with more than 50 agent-authored commits merged upstream. No external users, downstream projects, download or star counts, or third-party deployments are disclosed in the supplied material.
Big framing, vendor-only proof
The headline concept -- 'persistent agency', an agent that keeps working after users stop asking -- is a strong claim about behavior over days, and the supporting record is entirely first-party anecdote with no benchmark and qualitative evaluation. That gap is real but modest, because the code is open for inspection and the covering article itself foregrounds the missing independent results, the token-spend consequence of a loop bounded only by its own termination, and the shared-stream privacy problem rather than amplifying the claim.
Institutional interest in a research-to-deployment showcase
All quantified usefulness figures originate with Laude, which has a direct reputational stake: the report frames Headlong as a literal expression of the research-to-deployment thesis behind grantmaking and research programs Laude says are anchored by Konwinski's $100 million pledge. Evaluation being self-administered and qualitative compounds that alignment. Offsetting factors are the permissive Apache-2.0 license, the small readable codebase and Laude's own disclosure of weaknesses such as Audel being bad at keeping secrets.
Single publisher, self-critical, unverified outcomes
Confidence is limited by a one-publisher cluster whose factual base is a vendor launch post and vendor logs. It is not lower because the mechanism claims are specific, license and line counts are checkable, external vendor documentation is cited for the shell-access comparison, and the coverage explicitly marks what is unproven.
build
One agent, 119 blog heroes, and the scaffolding that made them shippable1 distinct publisher
build
Claude Code's agent-team panes need tmux, and Anthropic says Windows Terminal is out1 distinct publisher
build
Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses1 distinct publisher
build
A cost monitor overcounted 4.9x, then went dark for a week when set -e did its job1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026