Build1 publisher2 min readPublished
NVIDIA writes BlueField's API contracts into SKILL.md files for coding agents
NVIDIA put DOCA agent skills on GitHub and says agents using them met 100% of checklist items on 65 BlueField prompts, up from 19%. NVIDIA ran and graded that test, so the score carries over only as far as a team's own DOCA tasks resemble its prompts.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Each skill is scoped to one DOCA component and carries its real function signatures, pkg-config module names, build-container constraints and known failure modes.
- The skills span the DOCA library, including the Flow, GPUNetIO and PCC components.
- In NVIDIA's test, agents without skills misused APIs or flags on 59 of 65 prompts and failed to verify hardware capability on 46.
- Unaided agents also routed to the wrong tool on 39 prompts, skipped smoke tests on 34 and guessed versions on 30.
- With skills loaded, NVIDIA says agents run preflight checks, plan rollbacks and account for cold power cycles before touching hardware.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost With API or flag misuse on roughly nine unaided prompts in ten, a developer using a general agent on DOCA today has to review nearly every answer before it reaches a build.
- decision Because NVIDIA wrote and graded the test, a team weighing the skills needs its own DOCA tasks run against real hardware before the 19-to-100 gap tells it anything about its schedule.
- capability Domain knowledge now sits in an open file the agent loads, so a team can add or correct DOCA guidance without waiting for a model to be retrained.
DOCA is NVIDIA's software platform for BlueField DPUs, covering networking, storage, security, telemetry and lifecycle management [14]. NVIDIA explains why general agents struggle with it: the library surface is large, changes quickly and is hardware-specific in ways general training data does not capture [12]. The contents NVIDIA chose for each skill go straight at that gap. A pkg-config module name or a build-container constraint has exactly one right answer, and an agent that guesses wrong gets a failed build. Putting those values in a plain SKILL.md file the agent loads [3] is cheaper than hoping a model memorised them. Each file covers one component or workflow, so an agent writing Flow code reads Flow material [5]. NVIDIA's post calls each skill "a machine-readable specification the agent can reason against directly" [6].
NVIDIA ran the comparison itself. Its 65 prompts ranged from one-line questions to detailed multi-requirement tasks, and each was graded pass/fail against a required-answer checklist [8]. The results passage does not say which agents or models were tested, or how many checklist items a prompt carried.
The failure categories show what the checklist rewarded. Two of the five are about process: verifying hardware capability and running smoke tests [7][13]. The skills steer the agent toward that kind of step, checking what the device supports before writing code [11] and running preflight checks before touching hardware [10]. A checklist that awards points for those steps will score well any agent that has read a file prescribing them. A perfect score on a vendor's own checklist is the least surprising number in the post [9].
For the result to hold on a team's own work, two things would have to be true. Its tasks would need to resemble those 65 prompts. The grading would also need to track code that builds and runs on a BlueField, because an answer can tick every required item and still fail on the card.
The files also have to keep pace with the software they describe. NVIDIA says the DOCA library is rapidly evolving [12]. A SKILL.md with a stale signature hands the agent a wrong answer labelled as verified. The agent then has less reason to doubt that answer than it had to doubt its own guess.
What to watch
- Whether the GitHub skills are tagged to specific DOCA releases, and how quickly they change after a new DOCA version ships.
- A customer-run or independent evaluation that scores build and run success on BlueField hardware instead of checklist items.
- Whether NVIDIA publishes the 65 prompts and their checklists so others can rerun the comparison with different agents.