Build1 publisher3 min readPublished
Claude Code's plugin eval caught skills firing 55.6% of the time once 89 were loaded
Claude Code's new plugin eval showed one developer's skills firing in 5 of 9 relevant runs once all 89 were loaded, down from every run with one skill. The test covers three prompts in one project, but it gives teams that keep rules in skills a way to measure how often those rules get consulted.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Skills run only when the model decides from their descriptions to invoke one, and a skipped skill produces no error, warning or log line.
- By default, claude plugin eval also runs a second arm with the plugin not loaded and reports the difference between the two scores.
- The author reused four prompts written in July, before knowing what would be measured, to avoid writing test sentences the skill happens to catch.
- In three re-runs with traces kept, the skill never fired, and the model made only Glob and Grep file searches without invoking any other skill.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint In this project a rule kept in a skill was skipped in 4 of 9 runs where it applied, so a step like a migration GRANT cannot depend on the skill alone.
- decision A test plugin holding one skill reports a hit rate the production setup will not reach, so evals have to load the full competing skill set to be meaningful.
- cost Explaining a miss depends on sandbox traces the tool advises deleting, and deleting them before reading them means paying for another run, $0.82 in this case.
On a coffee e-commerce project with 89 skill directories, the developer's first eval came back 12 for 12, for $2.49 and 273 seconds [1][12]. That was four cases, each run three times [1]. The author wrote the result off immediately [13]. The test plugin held one skill, and when the only candidate's name matches the task, picking it costs the model nothing [13]. I think the replacement fixture is the right design. An assembly script builds a throwaway plugin from all 89 skills at run time. It keeps no copies in version control, because copies drift from the originals and this project has already paid for that once [14].
The second run kept the prompts and the model and changed one thing: 1 skill or 89 [15]. The grader asked a single yes-or-no question, whether the target skill was invoked [11]. The positives came back 5 of 9, or 55.6%, and the negative case stayed correct [15].
That figure describes this project. For it to carry over, another team would need a similar number of skills and prompts phrased the way its own users type. The sample is also small. Treating the nine runs as independent, a 95% Wilson interval on 5 of 9 runs from about 27% to 81% [3]. The author's July hand counts, Sonnet at 3 of 4 and Haiku at 1 of 4, land in the same range, and the author called the agreement of two independent methods "mildly reassuring" [5][16].
Before concluding that 89 skills dilute attention and should be pruned, the author wanted to see what the model had invoked in the missed runs [17]. A better-suited skill would mean correct delegation. A skill with an overlapping description would mean the descriptions needed editing [17]. The traces were already gone. The tool leaves a sandbox directory per run and prints "remove it when you're done", since the directories may hold agent-written content, and the author had cleared them all [18]. "I deleted something far more valuable than tidiness, in order to be tidy," the author wrote [19].
A re-run with --keep-temp cost $0.82 [20]. With no other skill in the traces, the author ruled out both delegation and overlap [22]. Each final reply opened, in the author's translation, "Based on this project's skill descriptions, here's what I expect" [23]. The model referred to the skill descriptions in its answer and still called no skill [21][23]. The author wrote that this reply changed the conclusion. The post's text breaks off inside that quote, before the new conclusion is stated [23].
The tool gets one thing right by default. Its documentation says: "If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass." [8] "Vendors don't usually ship a default that proves their users' work is useless," the author wrote [9]. Testing is cheap, too. The one-skill run worked out to about 21 cents per case-run, and the traced re-run to about 27 cents a run [4][5]. The alternative is the one the author described at the start: "You just watch the AI start working, and find out two days later that it skipped the rule." [4]
What to watch
- The author's revised conclusion from the final replies, and whether it points at description wording or at how the model chooses between skills and file search.
- Larger runs of the 89-skill set, with more prompts or more repeats per case, that would narrow the 27% to 81% interval.
- Trigger rates from other projects with large skill sets, to show whether 55.6% is typical or specific to this one.