Build1 publisher3 min readPublished
NVIDIA is documenting Holoscan for the agent, not the engineer
A vendor walkthrough has a coding agent build an endoscopic segmentation app from HoloHub examples, with the CLI as the shared execution surface and named skills as the spec.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- NVIDIA's post walks through building a real-time endoscopic tool segmentation application using an AI coding agent, with HoloHub examples and documentation providing implementation patterns and development skills guiding the agent through the HoloHub development process.
- The agents were additionally provided with: the Holoscan CLI with Bash execution permission; the HoloHub repository with documentation in a progressive disclosure pattern via agents.md; and HoloHub development skills, including holohub-app-lifecycle and holohub-debug-build-run.
- NVIDIA Holoscan is a platform for building real-time AI applications at the edge, from medical imaging to robotics; HoloHub is its companion repository, a growing collection of reference applications and components.
- NVIDIA states it wanted to explore how a general-purpose coding agent could use the same examples, documentation, and development tools available to an engineer in an actual development task.
- The overall objective was an end-to-end endoscopic tool segmentation application: real-time inference with live visualization of the segmentation masks, along with statistical analysis rendering.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
NVIDIA has published a walkthrough in which a general-purpose coding agent builds a Holoscan application using the same repository, documentation and command-line tooling an engineer would use [1]. The notable part is not the demo application but the inventory of things the agent had to be handed first: the Holoscan CLI with Bash execution permission, a repository whose documentation is arranged in a progressive disclosure pattern via agents.md, and named development skills [2].
Holoscan is NVIDIA's platform for real-time AI applications at the edge, from medical imaging to robotics, and HoloHub is its companion repository of reference applications and components [3]. The stated aim was to see whether a general-purpose coding agent could use those same examples, docs and development tools on an actual task [4]. The task was an end-to-end endoscopic tool segmentation application: real-time inference, live visualization of segmentation masks, and statistical analysis rendering [5]. The team reused the existing MONAI endoscopic tool segmentation model and a Holoscan sample video, and first confirmed the existing monai_endoscopic_tool_seg application ran locally [6]. The new work was visualization, runtime telemetry and repeatable benchmarking layered on the existing segmentation pipeline [7].
The CLI is the load-bearing piece here. Invoked through a ./holohub wrapper, it is described as the shared execution interface: the agent discovers development operations through it, and the engineer can inspect and repeat the same commands [8]. In practice that meant the agent dry-ran and then invoked ./holohub create to generate and register the standard scaffold, then implemented the application graph, execution modes, tests and documentation using existing Holoscan operators [9]. If the agent's work is expressed as CLI calls, it is reviewable in the same terms as a human's.
Two development skills are named, holohub-app-lifecycle and holohub-debug-build-run [2][10]. The first prompt invoked $holohub-app-lifecycle by name while constraining reuse: reuse the MONAI model, sample data, preprocessing and inference, show mask, coverage, timeline and uncertainty in a HoloViz overlay, do not train or modify the model weights, and make the sample video work end to end [11]. A skill referenced as a dependency in a prompt is an interface, with the versioning and maintenance cost that implies.
The review loop stays human. The engineer defines the goal and constraints, the agent inspects examples, implements the application-specific code and runs CLI operations, and the engineer reviews code, outputs and tests before setting the next goal [12]. NVIDIA's guidance is to decompose rather than one-shot the objective, into iterations guided by uncertainty and evidence [13], framed as questions: is the environment configured, does the model and video run end to end, is the visualization meaningful, can latency be measured repeatedly, can rendering throughput improve without feature regressions [14].
Read it as a vendor account of its own tooling. The workflow is called agent-agnostic but was run with Codex and GPT-5.6 "sol max mode", with agent processing times described only as approximate [15]. Three existing HoloHub applications are singled out as particularly useful references [16][17], which suggests the leverage came from the corpus rather than the model.
Worth watching: whether agents.md and skills files start getting tested and versioned like shipped code, and whether the CLI grows agent-facing affordances such as dry-run everywhere and machine-readable output.