Product1 distinct publisher3 min readPublished
Replatforming notes from the cate.blog post on Twill show what happens when your scaling artifacts have to hold up for something that remembers nothing between sessions, and where the unwritten doc bills you.
The Product Desk · Product desk

build
A 548KB CLAUDE.md burned 150,000 tokens before work started. The fix was a commit gate1 distinct publisher
product
The personal agent is a folder, not a model: four files and less memory than you thought1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
Perf work stopped being a specialist queue item, and slow endpoints became a choice1 distinct publisher
Compiled by The Product DeskSomething wrong?How this is made
The strictness in that merge queue comes from what the reviewing agent lacks, not from any upgrade. An agent that helped build a branch carries the reasons the branch is fine. An agent that meets the same branch at a gate has the repository and the rules and nothing else, so the rules are the only thing it can check against, which is what the cate.blog post credits for the stricter review [8]. Its author says she has given up on metaphors and treats that contextlessness as a neutral property rather than a strength or a weakness [16]. Operationally it behaves like a dial you get to set. Inside the work, an agent moves fast and forgives its own decisions. Outside it, the agent enforces what you wrote down and nothing you did not.
That is where the documentation bill lands. Scaling has always meant converting tribal knowledge into documentation, automation and process; only automation used to be written for a machine [1], and now all three are [1]. The doc that was a courtesy to a teammate is an input to a work product, and the process that lived in someone's head has to be parseable [1]. With people, knowledge sits in the people and in the systems; with agents, only the systems half survives the session boundary [2].
The test I would put on a backlog item tomorrow has two axes. Axis one is whether the knowledge is written somewhere an agent actually reads. Axis two is whether it is checked at a gate the work cannot route around.
- Written and gated: the merge queue, which serialises merges, files follow-ups and returns PRs that are not ready [6]. - Written, not gated: the encoded instruction to dispatch a review agent, which helped for months and still leaked the odd issue and untracked follow-up [5]. - Gated, not written: a rule that blocks work for reasons nobody on the team can reconstruct. Cheap to add, expensive the first time it is wrong. - Neither: the tribal knowledge, re-derived from scratch each session, in ways you find out about on Friday.
The human pacing rule was one careful change a month, on the reasoning that no effective person wants a job that feels more like process than impact [12]. That constraint was about people's tolerance, and it does not govern an artifact a machine reads. The post's own example is CLAUDE.md as a living team document whose churn is treated as a good sign [10]. On an earlier platform, the team moved off the affectionate "intern Claude" framing, put guardrails in, and started treating the thing as a system to be designed rather than a novelty to be supervised [9].
The uncomfortable part is the one the author names herself: people will do for a machine what they were never willing to do for their team, and she declines to claim moral high ground over anyone for it [11]. Automation used to arrive after an incident, because the incident was what made the cost of not doing it legible to everyone who had to approve it [13]. Now a contextless agent supplies that same argument every morning, for the price of one wasted session.
Ranked by verification strength, evidence, and original report placement.
Scaling a team is the work of turning tribal knowledge into documentation, automation and process; it used to be that only the automation was for the machine, but now documentation is machine critical and process is machine readable.
With people, knowledge lives in two places, in the people and in the systems; with AI only the systems part persists, because every session starts contextless.
While deep in replatforming Twill, the author was back to being a developer, operating an orchestrator and dispatching agents on the backlog, and found that habits she reached for felt familiar from team scaling.
'Dispatch a review agent' was encoded in the team's process for months and was very helpful, but the author would still find the odd issue or untracked follow-up.
Tired of Claudes colliding on CI and wasting build minutes, the author built a merge queue that serialises merges across concurrent sessions, files the follow-ups, and sends PRs back when they are not ready.
The next round of reviews dispatched after the merge queue was in place came back clean, and so did the rounds after that.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-hand and specific, corroborated by nobody
The queue, the streak of clean review rounds, the stricter reviewer — every fact this story stands on comes from the person who built the thing, published on her own blog, with no counts, dates, defect baseline or build-minute figures attached. The texture is credible in the way practitioner writing is credible: she names the mechanism, the file, the rule she added after wasting build minutes. But the sentence the story is titled on is an inference about why a change worked, drawn by the one person running the change, in a setup where the queue also added serialisation and follow-up filing at the same time.
Two codebases, both her own
Real usage, narrow blast radius. The practices are live on DRI and Twill and have survived long enough to change — the review-agent step ran for months, CLAUDE.md churns constantly, the queue is in the loop now. Nothing here has been picked up outside those two teams, and no other engineer's experience of the setup appears, not even Jean's.
The framing outruns a post that hedges
Slightly overstated, and mostly in the packaging rather than the writing. cate.blog explicitly refuses to make contextlessness a virtue or a defect and lands on 'a neutral'; she flags that the reason a rule was missing could be clarity rather than process. The claim that a merge queue made the reviewer stricter is nevertheless carried as a finding when it is one team's before-and-after with two variables changed at once — and the broader assertion that automation economics have flipped arrives with no numbers behind it at all.
No pitch, but she is grading her own build
There is nothing being sold: no product launch, no vendor, no funding, no benchmark to win. What is present is ordinary self-assessment risk — the author is on both teams she describes, built the queue she credits, owns the way of working she is arguing for, and is simultaneously making a reputational case that engineering managers should get closer to the work, which her own account happens to vindicate.
Internally coherent, externally untested
We are confident about what happened on one project and unconfident about what it means anywhere else. The account hangs together and the mechanics are checkable in principle by anyone who builds the same queue; a second team reporting the same before-and-after, or a count of what the queue actually bounced, would move this sharply. Until then the management observations travel better than the causal one.