Product1 distinct publisher3 min readUpdated
Ben's Bites rebuilt its author's personal agent from scratch. The lesson from version one was that automatic memory made the agent steer him instead of helping him.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Retained memory failed here in a specific way, and it is worth being precise about it. The previous folder told the agent to write down anything that looked like important context about him or his work [13]. What came back, when he wanted to explore a new direction, was steering: the agent kept citing what the file said he liked and staying in that lane [14]. The instruction worked exactly as written. Written was the problem.
The same mechanism shows up inside a single session, at a shorter timescale. He floated auto-saving early on, then later asked the agent to review recent chats and describe what he actually uses it for. It came back still holding the auto-save idea and proposing another, more complicated auto-commit instruction [17]. Text in the context window outweighs the thing you asked for two messages ago, whether that text arrived from a memory file or from your own discarded suggestion.
The file layout is where the argument gets settled. Five files went into the sketch [5]. The agent contested two of them: merge the building preferences into the main instructions, and drop the session log because git history already records the work [9][11]. He took the second and refused the first, which leaves four [1][2]. His reason for refusing is the useful part. He is working in Codex, whose system instructions say it is a coding agent, so its advice about how to organise a workspace came out shaped like a coding workspace [10]. About half the readers who answered his poll wanted to know how this works in Claude specifically [2]; his answer is that the vendors work much the same way, being files, folders, tools and instructions [3]. On that reading, one line in AGENTS.md saying questions get answers, not changes [6] is doing more work than the choice of model behind it.
One loose end is left open, and it is the one an operator will hit first. The log file was dropped because git already stores versions of files, diffs and commits [12], and a duplicate is not worth the space [11]. But the request he says he makes most often is "what did we talk about last week re: [thing]" [18], and a commit history of file changes is not a record of conversation. He notes that agents save all chat sessions to files, and the published text breaks off mid-sentence there [19], with a dedicated memory post promised later [20].
What survives is a rule rather than a stack: pick the smallest files with the least context, know what is in them, and edit when something changes or when the agent starts saying things you do not like [16]. Memory, on this evidence, is not storage. It is a standing instruction about what gets re-read, and the first version of his agent read too much [4].
Ranked by verification strength, evidence, and original report placement.
Around 700 Ben's Bites readers told the author they either use or want to use a personal agent.
About 50% of respondents to the previous day's poll wanted to know how a personal agent works in Claude.
He notes that agents save all your chat sessions in files; the supplied text is truncated mid-sentence at that point.
He says his own existing personal agent is pretty messy and often mentions stuff that is irrelevant, so he set up a fresh one.
Before starting, he sketched five files for the new folder: AGENTS.md, code.md (building preferences such as Vercel and Supabase), todos.md (current work), memory.md, and log.md (a log of every session).
AGENTS.md holds the main instructions: who he is, how he wants the agent to work with him ("questions get answers, not changes"), and pointers to the other files.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One first-person build log, no corroboration
All claims trace to a single newsletter post by the practitioner himself, written as a live session narrative. The mechanical details (file layout, git behaviour, Codex system-prompt bias, session files on disk) are internally consistent and specific, which is why this is not near-zero. But the load-bearing behavioural claim - that automatic memory made the agent steer him - is an unmeasured impression with no before/after comparison, no transcripts, and no second practitioner or vendor corroboration. The cross-tool equivalence claim is contradicted within the same article, and the supplied body is truncated mid-sentence so the final configuration is incompletely documented.
One personal deployment plus soft reader interest
Adoption evidence is limited to the author's own single-user setup and a self-reported reader tally. The ~700 figure conflates people who already use a personal agent with people who merely want to, comes from an AI newsletter's own audience, and carries no response base or methodology. There is no organisational deployment, no third-party usage, no retention or scale data. The pattern is real and running for exactly one person.
Mostly hedged, over-generalised in the framing
The body is unusually deflationary for the genre: the author admits he does not know what he wants from memory, flags that agents are agreeable and their suggestions can be wrong, defers a proper memory treatment to a later post, and leaves the SQLite idea unresolved. That pulls the gap close to aligned. The small positive comes from framing that outruns the evidence - 'they work pretty much the same way' across Claude and ChatGPT, and the general lesson that automatic memory steers you rather than helps, both generalised from a single session on one machine and immediately qualified by tool-specific exceptions in the same piece.
Newsletter audience-growth and follow-up-post incentives
The publisher is the author's own subscription newsletter, and the piece is structured around audience engagement: it opens by quantifying reader replies and poll results, invites readers to write in about the SQLite approach, and promises a dedicated memory post later, all of which serve retention and open rates. He also references his fund work as part of the context he feeds the agent, so an adjacent professional interest in the personal-agent category exists. Offsetting this, no product, vendor or paid placement is being sold, the named tools (Codex, Claude, Vercel, Supabase, git, SQLite) are mentioned incidentally, and the conclusion argues against buying more machinery rather than for it.
Confident about the recipe, not about the lesson
Confidence is moderate and split. The descriptive layer - what files he sketched, which two suggestions the agent made, which one he accepted, that he dropped log.md for git history, that sessions persist in dot-file directories - is directly attested and low-risk to restate. The interpretive layer - that automatic memory causes steering and that this generalises across tools and users - rests on one unmeasured anecdote from a single publisher with engagement incentives, and the supplied text is truncated before the setup is fully described.
Follow any of these and your For You feed starts watching them — no settings page required.
build
A coding agent deleted the rule its own rename had broken1 distinct publisher
build
The Context Tax: Your Developers Are Doing Unpaid Platform Work Every Session1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
1,500 submissions in 14 days: what a 12th-place GPU kernel says about agent loops1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026