Build1 publisher3 min readPublished
A guard that reads workflow files proves the schedule is in git, not that the job ran
A reader argued a static check only proves a declaration exists in source control. The canary that replaced it went red on its first manual run against production, and one failure was two days old.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Heinrich Neb, writing on dev.to, had shipped a guard requiring that any script whose own header says it runs daily must appear in a workflow file that has a schedule.
- A reader named Mads Hansen replied that the guard does not go far enough, because a guard that reads workflow files proves only that a declaration exists in source control.
- Hansen suggested a deployed canary instead: run the thing on the real path, and check that it uses the same identity, secrets and result sink as production.
- Neb built a second daily probe and, on Hansen's advice, ran it once by hand against production instead of waiting for its first scheduled run.
- The manual run went red immediately, then went red again for a different reason; the second failure was not in the new probe but in an older probe that had been running for weeks.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A static check that reads workflow files can only prove that a schedule is declared in source control, which is what a reader named Mads Hansen told Heinrich Neb after Neb shipped exactly that guard: any script whose own header claims a daily cadence must appear in a workflow file carrying a schedule [1][2]. Hansen's counter-proposal was a deployed canary on the real path, using the same identity, secrets and result sink as production; Neb built one, ran it by hand against production instead of waiting for its first scheduled run, and it went red immediately, then red again for a different reason [3][4][5].
The two failures were wrong permissions on the token and a missing notifier, and both took under a minute to find [6]. Neither is a declaration in a YAML file, so neither was visible to the guard that had just passed [1].
The missing notifier is the part worth borrowing, because the mechanism is ordinary. The probe keeps its history on a separate branch so the main branch does not collect one commit per data point: the step checks that branch out, appends a line, and pushes [7]. That branch contains exactly one file, so checking it out removes every other file from the working tree, including the notifier the next step calls [8]. On the day the step was written the branch did not exist, so the code took the other path, the one that creates an orphan branch and leaves the working tree alone, and the run was green [9]. The next morning the branch existed, the first path ran, the notifier was gone, the daily message stopped arriving, and the only trace was a red run nobody read [10]. That cost two days of silence, on a probe that had been running for weeks [11][5].
The general shape is a create-if-missing path: the first run takes one branch of the code and every later run takes the other, so the version that matters was never the version anyone reviewed [12]. A change that passes the day it ships and fails the day after is the hardest kind to catch, because the review, the test and the memory of it all come from day one [13]. Neb's fix is ordering, not cleverness: move the notification before the push, because a failed push must never be allowed to silence the message [14].
Two of his three checks are mechanical. Find scheduled workflows that also push with `grep -l 'schedule:' .github/workflows/*.yml | xargs grep -l 'git push'`, then read each hit by hand and ask whether any step after that push runs a script from the repo [15]. The third is the one no file can answer: `gh workflow run <name>.yml && gh run watch`, against production, and watch it [16]. Related failure class in the same post: a notifier called through plain curl, where curl exits 0 as soon as the request is made, including a 401 that delivered nothing [17]. Neb's new probe aborts when its API key is missing, before it touches the network, and exits non-zero rather than printing a tidy empty report, on the grounds that found nothing and never asked must not look the same [18].
What to watch on your own pipelines: whether the canary can actually go red. Neb's advice is to break it once on purpose, remove the token or point the notifier at a wrong channel, and confirm both that the run fails and that a human hears about it [19]. His older probe is sending its daily message again, and he knows that because he asked, not because a check told him [20].