Build1 publisher3 min readPublished
Before you spend quota on an agent skill, make it pass an eval harness
A Google AI series on dev.to shows how Inspect AI turns "is this MCP server worth my tokens" into a measured question, using a cheap grader model and three runs per test.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A dev.to post published under Google AI's account, titled "Designing AI Evals: Clarity Now and Visualization Next", is the opening of a series on designing objective evaluations of AI tooling such as MCP servers, agent skills and agent plugins.
- The post frames the problem as: modern LLMs can likely one-shot many specific tasks, but that specificity may require wasting tokens and time repeatedly prompting them with the same resources, descriptions and scripts, raising the question of whether a skill is worth your time, tokens and quota.
- The post names Inspect AI and Harbor as open source eval frameworks for evaluating agent skills.
- In the companion codelab, the author ran evals with Gemini CLI, Inspect and Inspect SWE in an isolated Docker sandbox to understand how well each skill aids the agent in answering the same question.
- The author used three models for the demo: google/gemini-3.5-flash-lite and google/gemini-3.6-flash as solvers (the models under evaluation) and google/gemini-3.1-flash-lite as a grader (the model rating the runs).
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A series published on dev.to under Google AI's account walks through evaluating an agent skill with open source eval frameworks instead of adopting it on impression, naming Inspect AI and Harbor as the tooling for the job [1][3]. That framing is the useful part for anyone running agents in production: the cost of a mediocre skill is not a bug report, it is quota burned re-prompting a model with the same schemas, descriptions and scripts over and over [2].
The mechanics are ordinary, which is the point. The companion codelab runs evals with Gemini CLI, Inspect and Inspect SWE inside an isolated Docker sandbox, asking the agent the same question with each skill in turn so the skill is the variable [4]. The author reports building the sweep around three architectural dimensions: external configs, isolated sandboxing so parallel attempts cannot corrupt each other's state, and multidimensional rubrics [9]. The `--epochs` flag is set to 3 so each test runs three times and non-deterministic output gets averaged out, with metrics aggregated by the mean of those runs [10].
The cost engineering is worth copying. According to the post, two models act as solvers under evaluation, google/gemini-3.5-flash-lite and google/gemini-3.6-flash, while a previous-generation model, google/gemini-3.1-flash-lite, does the grading [5]. Grading was deliberately pushed onto the older model: with rubric criteria reduced to strict binary decisions and the reduction applied programmatically, the author says it is robust enough without consuming solver quota [6]. For a production eval system the same post recommends newer, more capable graders, on the grounds that they will likely produce narrower confidence intervals [7]. Note what that arithmetic implies: two solvers at three epochs is six solver runs per test case before you count grader calls, which is exactly why the grader model choice is a budget decision and not a taste one [16].
Reproducibility gets real attention. The runs were benchmarked on Python 3.13 with inspect-ai 0.3.247, inspect-swe 0.2.66, inspect-viz 0.4.1 and pandas 3.0.3, with exact version pins in the README in case upstream PyPI releases break something [8]. Followers along at home are told to clone Google's public skills repository into a local `google-skills/` directory before running the sweeps [13], to install `inspect view` in advance [11], and to expect evals to take a few minutes depending on machine and quota [11]. `inspect view` then serves the results locally and opens up individual traces [11].
Two honesty notes. The published text reports the last run taking "8 minutes and 21 minutes", which is not a readable number, and the bullet describing multidimensional rubrics repeats the sandboxing description verbatim, so the rubric design is asserted rather than shown [12][15]. Harbor is named in the framing but the demonstrated stack in this installment is Inspect and its companion packages [17]. The post also discloses that its diagrams are AI-generated alongside screenshots and hand-drawn edits, with AI assisting minor copy editing [14].
What to watch: the promised follow-up extending the evaluation into visualizations via Google Sheets and Data Studio [18], whether the pinned dependency set survives upstream churn [8], and whether anyone publishes rubric definitions concrete enough that a second team can reproduce a skill's score rather than just its run command [9].