Build1 distinct publisher3 min readUpdated
SkillEvaluator scores each skill with and without installation and calls the gap Skill Lift. The more interesting figure is the control: 39 to 46 out of 100.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
NVIDIA has published a first benchmark run for its verified agent skills: more than 300 skills across more than 30 products, each evaluated on two independent agent harnesses [3]. The measurement is a controlled A/B, where the same task is run once with the skill installed and once without, and the difference in score is reported as Skill Lift in points [4].
That reframes what a "skill" is. Until now, packaging instructions, examples, and tool guidance for an agent has been an act of taste [18]. NVIDIA's definition is narrower: a verified skill is a packaged, signed capability descriptor saying what a product does, when to invoke it, and how to call it, and the word verified refers specifically to the measurement that decides it is ready to ship [2]. SkillEvaluator, the open source tool that does the measuring, combines static checks with real task runs on both sides of the comparison [1].
The gate has three tiers that can each run alone. Tier 1 is safety and structure: schema and frontmatter validation, quality scoring, security scanning for prompt injection and data exfiltration, secret and PII detection, license checks, and script linting [6]. Tier 2 uses embedding similarity to find duplicated guidance inside one skill and overlapping coverage across the catalog [7]. Tier 3 runs a live agent against generated tasks in an isolated sandbox, with and without the skill, and measures the delta [8]. Tier 3 sits on Harbor, an open source framework for repeatable isolated agent evaluations, with SkillEvaluator handling setup, task conversion, sandbox execution, collection, and the arithmetic [9]. Inside a harness, prompt, model, task inputs, and grading criteria are held constant, so installation state is the only variable [10].
The number worth staring at is not the lift, it is the control. Average without-skill scores, taken from runs on Codex and Claude Code, ranged from 39 to 46 out of 100 across Correctness, Discoverability, Effectiveness, and Efficiency, four of the five dimensions Table 1 defines [14][15]. Every one of those baselines is below half marks [16]. Lift measured off a floor that low is real, but it says as much about how badly an unaided agent handles specialised library work as it does about the packaging.
Two caveats operators should carry. The evaluation dataset is generated by the same tool, via create-eval-dataset, producing cases with IDs, prompts, expected outputs, and optional assertions, with the full mode adding explicit, implicit, contextual, and negative cases for the author to review [11]. The exam is written by the pipeline being graded. And scores are macro-averaged across published skill-harness pairs with equal weight per pair, so a product shipping many thin skills moves the headline as much as one shipping a few load-bearing ones [12].
The reproducibility story is better than the usual vendor chart. Figures come from a dated snapshot of benchmarks.json at a named commit, and the catalog is evaluated continuously with current numbers in the nvidia/skills repository [12][13]. That is a file a sceptic can diff against.
What to watch: whether lift compresses as base models improve and the 39 to 46 floor rises; whether Tier 2 starts rejecting skills as catalog overlap grows; and whether anyone outside NVIDIA runs SkillEvaluator against skills distributed through Skills.sh, ClawHub, or Hermes Hub, or through the Claude Code, Codex, and Cursor plugins [5].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The post shares the first benchmark results for more than 300 verified skills across over 30 NVIDIA products, with each skill evaluated on two independent harnesses.
Each run executes in its own isolated sandbox with the same prompt, model, task inputs, and grading criteria, so within each harness the only experimental variable is whether the skill is installed.
Table 1 defines five scoring dimensions; average baseline scores ranged from 39 to 46 out of 100 across Correctness, Discoverability, Effectiveness, and Efficiency, indicating substantial room for improvement.
NVIDIA SkillEvaluator is an open source tool for measuring how skills affect agent performance through static checks and real-world task runs with and without each skill.
NVIDIA verified Skills are packaged, signed capability descriptors that tell an agent exactly what an NVIDIA product does, when to invoke it, and how to call it; the verified part is the measurement that determines it is ready.
For each harness, Skill Lift was calculated by comparing scores from runs with and without the skill installed, and is reported in points as the with-skill score minus the without-skill score.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed vendor methodology, no outside check
The methodology is unusually specific for a vendor post: a three-tier gate, Harbor-backed sandboxes, an explicit single-variable control design, a named metric definition, and figures pinned to a dated benchmarks.json commit that readers can inspect. It is also entirely self-measured and self-graded on NVIDIA's own skills for NVIDIA's own products, with no confidence intervals and 85% of published skills scored on a single attempt per task, and no second publisher in the cluster corroborates any number.
Catalog scale shipped, real usage undisclosed
There is concrete shipping evidence: 300+ verified skills across 30+ products, evaluated continuously in a public repository, plus plugins for Claude Code, Codex, and Cursor and listings on Skills.sh, ClawHub, and Hermes Hub. What is missing is any demand-side signal: no installs, downloads, active users, customer deployments, or third-party adopters are disclosed, so measured adoption reflects distribution surface rather than uptake.
Mildly overstated by the 'verified' framing
Positive but small. The word 'verified' and the claim that skills improved performance across both harnesses carry more authority than single-attempt, vendor-graded runs without confidence intervals can support, and the Security baseline of 97 measures absence of regression rather than security strength. The gap is held down by the post's own limitations section, which discloses run-to-run variance, the 85/15 attempt split, and the fact that Discoverability and Efficiency partly score actions unavailable in the control arm.
Vendor grading its own skills for its own products
NVIDIA authored the tool, authored the skills, defined the metric, ran the harnesses, graded the runs, and published the results, all for skills that steer agents toward NVIDIA products and libraries. Wider skill adoption directly supports use of NVIDIA's own stack, and distributing through third-party agent tools and hubs extends that reach. Publishing the tooling as open source and pinning results to an inspectable commit partially offsets, but does not remove, the conflict.
Facts clear, significance unverified
Confidence in what was said is high: the cluster contains one detailed primary source and the descriptive claims about the tool, tiers, metric, and baselines are unambiguous. Confidence in what it means is moderate at best, because a single vendor publisher, self-grading, a mostly single-attempt sampling design, no confidence intervals, and no adoption metrics leave the significance of the reported lift untested.
build
A 12MB Go binary bets agent cost control is cache stickiness, not a dashboard1 distinct publisher
build
Waku 0.1.0 bets the product is the control plane, not another coding agent1 distinct publisher
build
Agent-written docs need a paper trail, not a confidence score1 distinct publisher
product
LangChain's dcode and NVIDIA's NemoClaw sell controls, not code quality1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026