Build1 publisher3 min readPublished
Config rot has a latency bill: a 70-line weekly audit for agent environments
One developer found seven dead MCP connections and could not say when they broke. His answer was measurement: a weekly snapshot that turns "feels slow lately" into a dated diff.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author built a weekly health check for his Claude Code environment consisting of two files: a 70-line diagnostic script (~/.claude/scripts/env-audit.sh) and a launchd job config that fires it weekly.
- The launchd job com.shun.env-audit is started every Sunday at 09:00 and executes ~/.claude/scripts/env-audit.sh.
- An MCP server whose auth token has expired will sit in the config as Failed to connect forever.
- A plugin's command files can still be on disk while the plugin is gone from enabledPlugins in settings.json.
- An auto-skill written as a good-enough-for-now procedure can still be sitting in ~/.claude/skills/auto/ six months later, with the author paying the context tax of Claude loading it every single time.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing on dev.to has published his fix for a problem most operators feel but cannot date: a 70-line bash script plus a launchd job, two files in total, that snapshot a Claude Code environment every week [1]. The job, com.shun.env-audit, fires at 09:00 each Sunday and runs the script [2]. It matters because every failure mode he catalogues is a silent one.
An MCP server whose auth token has expired does not announce itself. It sits in the config reading Failed to connect, indefinitely [3]. A plugin's command files can stay on disk after the plugin itself is gone from enabledPlugins in settings.json [4]. An auto-skill written as a good-enough-for-now procedure is still in ~/.claude/skills/auto/ six months later, loaded on every run [5].
The case for treating this as performance rather than tidiness rests on the context window being finite. According to the author, the total volume of settings, skills and hooks loaded at startup directly affects the quality of the first response [6]. Enable 50 plugins and their metadata rides along in context every time [7]. Leave 10 MCP servers stuck at Failed and connection-attempt timeouts stretch startup [8].
His own incident is the useful part. Three months into serious use, one week's first response was noticeably sluggish, and claude mcp list returned seven Failed to connect lines [9]. Five of the seven he had no memory of installing; they had been added automatically by plugins [10]. That is roughly 71 percent of the broken surface arriving without a human decision [11]. If plugins can install MCP servers, your dependency count is not something you can recall from memory.
He frames the cost plainly: three weeks running with a broken connection is cumulative minutes of timeout waiting, on a problem that takes 10 seconds to fix once you notice it [12].
Detection is the whole design. Reports accumulate as ~/.claude/logs/env-audit-YYYYMMDD.md, so a diff dates the regression instead of describing it [13]. His argument for automating rather than watching: daily use raises your own threshold for "heavy", so something 20 percent slower than three months ago just becomes normal, and without weekly snapshots there is no baseline to compare against [14].
What the script counts is unremarkable, which is the point. jq reads the length of enabledPlugins out of settings.json, and find counts the command files, SKILL.md files and agent files actually present under ~/.claude/plugins [15]. Then claude mcp list runs under a 25-second timeout with Connected, Needs auth and Failed tallied, jq dumps the hooks config, ls lists the auto-skills, and ccusage blocks --active pulls recent cost [16]. The gap between the enabledPlugins count and the on-disk file count is his zombie-file indicator [17], and he is careful to note those orphaned files do not themselves consume context [18].
Two things to watch. First, cadence. Weekly bounds worst-case time-to-detect at seven days [19], which is tolerable for token waste and expensive if an expired token is blocking work you need done Monday. Second, whether agent runtimes begin reporting startup context load and failed-connection counts as first-class output. Until they do, the instrumentation is a shell script you maintain yourself, writing to /tmp/claude-env-audit.md unless the scheduler hands it a date-stamped path [20].