Skip to content

Build1 publisher3 min readPublished

A 548KB CLAUDE.md burned 150,000 tokens before work started. The fix was a commit gate

One team cut its agent context file to 34KB without deleting a single rule, then wired a size check and a structure test so it cannot quietly regrow.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A 548KB CLAUDE.md burned 150,000 tokens before work started. The fix was a commit gate
Generated illustration

What happened

  • The team's CLAUDE.md was 548KB, and every session, including every subagent, loaded all of it before doing any work.
  • One measured headless run wrote about 150,000 tokens to cache before the actual task started, and the CLAUDE.md file was the dominant contributor.
  • The team cut CLAUDE.md from 548KB to 34KB resident without deleting a single obligation.
  • The cut from 548KB to 34KB is a reduction of about 94 percent in resident size.
  • The official memory docs (code.claude.com/docs/en/memory.md, checked 2026-08-18) say files are loaded into the context window at session start and advise targeting under 200 lines per CLAUDE.md file, because "longer files consume more context and reduce adherence."

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A team writing under the Rulestack account on dev.to published measurements this week showing their CLAUDE.md had reached 548KB and was loading in full at the start of every session, before any work began [1]. In one measured headless run, according to the write-up, roughly 150,000 tokens went to cache before the actual task started, and the file itself was the dominant contributor [2].

That is the useful reframing: the agent context file is a fixed cost paid on every invocation, not documentation that sits inert until someone reads it. The team cut it to 34KB without deleting a single obligation [3], a reduction of about 94 percent in resident size [4].

The official memory docs, which the authors say they checked on 2026-08-18, tell you the shape of the budget directly: target under 200 lines per CLAUDE.md file, because "longer files consume more context and reduce adherence" [5]. Theirs peaked above 2,200 lines [6], roughly eleven times that target [7]. Nobody approved that; it accreted one incident postmortem and one owner instruction at a time [8].

The trap worth naming is the reorganisation that feels like progress. Splitting a large file into ten imported files does nothing for startup cost, because the docs state that imported files "still load and enter the context window at launch" [9]. Imports are for organisation and deduplication [10].

What actually moves cost off the every-session line is conditional loading. Files in .claude/rules/ with a paths frontmatter field "only apply when Claude is working with files matching the specified patterns"; a rule without paths loads at launch like CLAUDE.md, so the frontmatter is the entire difference [11]. Skills load in two stages: the description is always in context so the agent knows the skill exists, but the SKILL.md body loads only on invocation [12]. Nine operational runbooks became nine skills, and their combined body text left the every-session budget [13]. Block-level HTML comments are stripped before injection, so maintainer notes are free; the team says they had been paying for notes-to-self for months without knowing [14].

The sorting rule that emerged is a triage policy, not a style guide: decision content stays resident, procedures become skills, code conventions become path-scoped rules, reference material goes to docs/ to be found by grep, and the full pre-migration text goes to an archive file so history stays greppable without being resident [15].

The multiplier is the part most teams will feel first. CLAUDE.md loads into every subagent as well [16], so for fan-out workloads the resident size is charged once per parallel agent, and shrinking it cuts fixed overhead across the whole spawn.

Deletion is not the interesting move here; enforcement is. The same day, the team added a size check in their health monitor that warns at 45KB and alerts at 60KB, evaluated every session [17] - eleven kilobytes of headroom above current size before the first warning fires [18]. They also added a structure commit gate: a test that fails the commit if CLAUDE.md references a skill directory that does not exist, or if a skill exists and no trigger in CLAUDE.md points at it [19]. That converts a document nobody owns into a build artifact with a failing test.

Two things went wrong during the migration, by their own account: one was caught by the commit gate, and one reached production behaviour [20]. That second one is the honest warning in the piece. Moving obligations into conditionally loaded files changes when they apply, and the gate checks references, not semantics.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories