Skip to content

Build1 publisher2 min readPublished

ServiceNow's AutoSynthData turns an agent's failed tasks into verified training data

ServiceNow's CoreAI team generates new agent training tasks, each with its own verifier, aimed at gaps a stronger teacher model can already solve. Adopting it means running a seedable copy of the environment and writing verifiers that accept any valid solution.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying ServiceNow's AutoSynthData turns an agent's failed tasks into verified training data
Generated illustration

What happened

  • Each generated task bundles three parts: a system specification, a user prompt and a verifier that decides whether the agent's trajectory succeeded.
  • A useful task must be feasible in the environment and realistic, and the target must not yet solve it consistently, because tasks it already solves reliably add little training signal.
  • The loop runs the target on diagnostic tasks, has a stronger teacher show which failures are solvable, generates tasks, checks each in the environment, then post-trains on the ones accepted.
  • ServiceNow demonstrates the pipeline on EnterpriseOps Gym (Malay et al., 2026), using that benchmark's released dataset.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The loop can only teach skills the chosen teacher already performs in the environment, so picking the teacher sets the upper limit on what the target can learn from it.
  • exposure Verifier defects become reward. ServiceNow says a lax check can reward incorrect behavior and a strict one can penalize valid solutions, so adopters answer for every bad check the generator produces.
  • cost Adopters pay again for teacher runs, environment validation and post-training in every round, so the compute bill grows with the number of rounds the target model needs.

A task's specification can seed a database state or load knowledge articles, so the generator writes starting states as well as prompts [4]. Feasibility is then tested against the live environment [6]. At least one trajectory has to satisfy the prompt while respecting the specification [6]. That excludes tasks that need unavailable tools, inaccessible knowledge, impossible state transitions or actions the policy prohibits [6].

Each candidate task is checked by running it in the environment before its sample goes to post-training [11]. In my view this is where most of the adoption cost sits. A team needs a copy of its own systems that it can seed, drive through the same tools and reset for every candidate [4][11].

One rule in the specification section is good engineering, and I would want it in any synthetic-data pipeline. Instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty [5]. A generator asked for tasks the target fails has an easy route to them: add a rule no real user would impose. The realism requirement closes the same route from the prompt side. "The space of executable behaviors is usually much larger than the space of realistic workflows," the CoreAI team writes [7].

Completeness is the expensive verifier property. A verifier has to accept valid solutions instead of encoding one particular reference trajectory, while still rejecting trajectories that violate constraints [9]. Replaying the teacher's path is the cheap check to write. It would mark a correct agent wrong whenever the agent took a different valid route [9].

Targets come from tasks the target model fails and the teacher completes [2]. After each post-training round the updated model is evaluated again. The curriculum then shifts toward what it still finds difficult [2][11]. The pipeline description does not include a compute budget, a before-and-after score or the teacher model's name [11].

The worked example uses the released EnterpriseOps Gym dataset [12]. A gain on that dataset would carry over to a given enterprise only if that enterprise's systems can be seeded, executed and verified the same way [4][11].

What to watch

  • Before-and-after EnterpriseOps Gym scores for a named target model, with the number of generation rounds each gain took.
  • The identity of the teacher model, and whether ServiceNow releases the generator and verifier code.
  • Any measured rate of generated verifiers later found lax or overly strict, which would size the reward-error risk.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories