BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Laude and MIT put persistent agency in 9,900 lines of Bash, and the bill arrives with it
Headlong is Apache-2.0, so a team can read the loop before running it. What nobody can read is a benchmark, and the shared thought stream is also the privacy problem.
The Engineer · Build desk

What happened
- Laude Institute and MIT released Headlong on August 25th, a harness that lets an agent keep thinking, choosing projects and acting after its users stop talking to it.
- The code is Apache-2.0 on GitHub, with Laude putting the core at 9,900 lines of Bash in bin/ and thinkers/.
- There are no independent performance results, and Laude says it assesses the agent mainly through qualitative observation.
Why it matters
- cost Budget owners inherit a line item no request log explains: model spend attributable to nobody's question, generated while the office is empty.
- exposure Cross-user exposure lands on whoever installed it, because something said in one channel can resurface inside an answer given to a different colleague.
- contradiction runtimewire treats the loop as the hazard while noting that reactive CLIs already run shell commands, so a policy written around per-command approval does not cover an agent acting hours after the...
- decision With nothing to benchmark, the adopt-or-not call falls to whoever can read Bash, not to whoever compares scores.
Who decides how many model calls happen matters more than the line count. The main loop, called Thinker, repeatedly invokes a component named shellm, which asks a language model for reasoning, Bash commands or both, executes the commands, and continues until the model sets a FINAL environment variable [15]. There is no checklist unless the agent writes one, and each next thought is selected from earlier thoughts and new observations [13]. Iteration count is therefore a property of the agent, not of the request log [20], which is the mechanism under runtimewire's warning that the continuous loop can increase token spending [6]. Tiered compaction keeps recent events in full detail and summarizes older history [16], which shrinks the context per call and does nothing to the number of calls.
The shared stream has the same double character. Audel, the agent Laude has been running for several weeks, works as one agent over one stream reached through Slack, Telegram and a mobile app [8]. The behavior Laude puts forward as the point depends on that arrangement: following what different people are doing, returning to an old subject, contacting someone with no new prompt [9]. runtimewire flags the same stream as a cross-user privacy risk [7]. One mechanism, two names, and the team that installs it is choosing on behalf of everyone in the channel.
Shell access is not the delta. OpenAI's Codex CLI can read, modify and run code locally [17], and GitHub Copilot CLI can execute shell commands subject to permissions or approval [18]. An approval prompt assumes a human is present at the moment the command runs. Laude's clearest example, dated August 5th, happened while nobody was interacting with the agent: it checked a background recall process it had built, found the code reading an environment variable that was never set, searched the codebase to confirm, rewrote the process to read from a pipe, verified the repair, and the fix was merged, with the log putting the sequence at 48 minutes [12]. Attendance is what changed, not privilege.
What stands in for evaluation is that log. More than 50 commits from Audel's fork were pulled into the main repository [11]. On its first day the agent audited eight stale Git branches and came back 10 minutes later to correct its own count [10]. There are no independent performance results, and Laude says it judges changes to Audel mainly by qualitative observation [14], so every quantified figure on offer is first-party [19]. Laude's grantmaking is anchored by Konwinski's $100 million pledge [5], which is part of why the artifact turned up as Apache-2.0 code on GitHub instead of a product with a score attached [3].
The compensating handle is instrumentation. Thoughts and actions are stored in a directed graph of JSONL files that supports forks and merges [15], and the core is 9,900 lines of Bash in bin/ and thinkers/ [4]. Absent an eval, that graph is the only way to reconstruct what an unattended agent spent and what it touched, after the fact.
What to watch
- An evaluation of continuous thought run by anyone other than Laude, covering days rather than single tasks.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence34
- Adoption16
- Hype gap+22
- Incentives66
- Confidence42
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Andy Konwinski's Laude Institute and MIT released Headlong on August 25th, an open-source harness that allows an AI agent to keep generating thoughts, choosing projects and taking actions after its users have stopped talking to it.
- [2]
The August 25th launch post calls the design "persistent agency": messages become observations inside a continuous thought stream instead of opening isolated sessions, and the agent decides whether to answer, wait, investigate something else or contact a user later.
- [3]
Developers can inspect and modify the Apache-2.0 Headlong code on GitHub.
- [4]
Laude says Headlong's core is currently less than 10,000 lines of Bash, specifically 9,900 lines in bin/ and thinkers/.
- [5]
Laude says its grantmaking and research programs are anchored by Konwinski's $100 million pledge.
- [6]
Headlong's continuous loop can increase token spending, according to runtimewire.
- [7]
Headlong's shared stream creates cross-user privacy risks, according to runtimewire.
- [8]
Laude has spent several weeks running a shared Headlong agent named Audel through Slack, Telegram and a mobile app, and every conversation enters the same stream.
- [9]
Audel can follow what different people are doing, return to an old subject and message someone without a new prompt.
- [10]
On its first day Audel audited eight stale Git branches belonging to a Laude member, then returned 10 minutes later to correct its own count.
- [11]
Laude reports that more than 50 commits made by Audel in its fork were pulled into Headlong's main repository.
- [12]
On August 5th, while no one was interacting with it, Audel checked whether a background recall process it had created was connected, found the code looked for an environment variable that was never set, searched the codebase to verify the diagnosis, rewrote the process to read from a pipe and checked the repair; Laude's log puts the sequence at 48 minutes and the resulting fix was merged into the public repository.
- [13]
Laude says Headlong has no checklist unless the agent creates one, and that it continuously asks a model to choose its next thought from its earlier thoughts and new observations; cron-style systems by contrast wake an agent on a schedule to run a fixed checklist.
- [14]
Headlong has no independent performance results showing whether continuous thought reliably produces useful work, and Laude says it currently evaluates changes to Audel mainly through qualitative observation.
- [15]
Headlong's main loop, Thinker, repeatedly invokes shellm, which asks a language model to produce reasoning, Bash commands or both, executes the commands and continues until the model sets a FINAL environment variable; thoughts and actions are stored in a directed graph of JSONL files that supports forks and merges.
- [16]
Headlong uses tiered compaction: recent events remain in full detail while older history is progressively summarized.
- [18]
GitHub Copilot CLI can execute shell commands subject to permissions or approval.
- [19]
Every quantified claim of Headlong usefulness in circulation is first-party: the merged commit count, the 48-minute self-repair and the eight-branch audit all come from Laude's own deployment and logs, and no independent results exist.
- [20]
The number of model calls Headlong makes is bounded by the agent's own loop termination, not by the number of user requests, so calls continue when no user is present.
Sources
1 independent publisher whose own reporting we read for this story.
- runtimewire.comLaude ships an agent harness that keeps thinking after you stop asking
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Open-Source AI ToolingFollow
- Agent Harness ArchitectureFollow
- Persistent Autonomous AgentsFollow
- Multi-User Agent Privacy And Trust BoundariesFollow
- Agent Evaluation GapsFollow
- Agent Inference CostFollow