Build1 publisher3 min readPublished
Passing Schema Validation Says Nothing About Whether Your Agents Can Act Safely
A NexFlow write-up splits structural, cross-reference and runtime checks. The first two only prepare a configuration for execution; the third, which the author calls future, is what stops an action.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A manifest can pass JSON Schema validation and still describe an AI team that cannot work safely; every required field may be present, the kind supported, and permissions an array in the right place while one permission points to a nonexistent capability, a task names an unknown owner, and a risky action has no approval gate.
- NexFlow describes AI developer teams through YAML manifests, which are readable in reviews but require a stricter contract for tools.
- Each supported NexFlow manifest kind maps to a JSON Schema.
- The current repository validator parses YAML safely, converts each document into a JSON-compatible structure, selects a schema from its kind, and reports errors with file and instance paths.
- At the reviewed repository checkpoint, npm run validate validates 113 manifests against 17 schemas.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post on dev.to about NexFlow, a YAML manifest format for describing AI developer teams, makes a distinction most configuration tooling quietly collapses: a manifest can pass JSON Schema validation and still describe an AI team that cannot work safely [1][2]. That matters because the word "validation" is doing three different jobs, and only one of them is shipped.
The mechanics are ordinary and worth reading precisely. Each supported manifest kind maps to a JSON Schema [3]; the repository validator parses YAML safely, converts each document to a JSON-compatible structure, picks a schema by kind, and reports errors with file and instance paths [4]. At the reviewed repository checkpoint, `npm run validate` checks 113 manifests against 17 schemas [5] - roughly 6.6 manifests per schema kind [6], which tells you the manifest surface is broad: project, actors and agent definitions, tasks, workflows, handoffs, capabilities, permissions, context, memory, providers, model profiles, prompt sets, retrieval profiles, events, and extensions [7].
That layer earns its keep on errors that need no knowledge of another file. A permission effect is constrained to `allow`, `deny`, or `approval_required`, so `full_auto_magic` fails immediately [8], as do missing required fields, an unsupported specification version, an unknown kind, a wrong type, or a malformed identifier [9]. Reviewers should not be spending attention on those [10].
The gap opens the moment one manifest points at another. An agent definition listing `permissionRefs: [implementation_branch_work, delete_everything]` is, to JSON Schema, a valid array of well-formed identifiers; the schema for that file does not necessarily know whether `delete_everything` exists in `permissions.yaml`, what capabilities it grants, or whether an approval gate covers it [11]. Resolving that requires treating the manifest set as a connected model: task owners, workflow dependencies, handoff artifacts, permissions, capabilities, context sources, memory scopes, events, and extensions, with each reference required to find exactly one valid target of the right kind without contradicting adjacent policy [12][13]. NexFlow currently ships a bounded semantic reference smoke check; at the same checkpoint, `npm run semantic-smoke` passes for seven example projects, which the author presents as evidence of current coverage rather than complete semantic validation [14][15].
The third layer is not there yet. The post refers to a future runtime as the thing that must decide what happens when an agent attempts a real action [16], and enumerates what schemas cannot do: isolate credentials, inspect actual access in GitHub, Linear, or a filesystem, pause execution before a risky change, revoke a previously granted permission, or record a real agent action in an audit log [17]. Runtime enforcement is where real permissions, credentials, approval gates, isolation, and audit rules apply [18]. The first two layers prepare a configuration for execution; they do not execute it and do not replace runtime controls [19].
The practical failure mode is a tool that prints `Configuration valid` and leaves the user unable to tell whether it parsed YAML, applied schemas, resolved a subset of references, or intended the word to be read as a safety guarantee [20][21].
Watch whether the output names the layer that ran, whether semantic coverage moves past seven curated examples toward the full 113-manifest set, and whether the runtime arrives with credential isolation and audit before anyone points these manifests at a production repository.