Build1 publisher3 min readPublished
Scheduling a prompt is not monitoring: keep the LLM upstream of the cron
A dev.to post argues that recurring natural-language analysis confounds the metric with the instrument. The remedy is a frozen query, reviewed as a diff, with the model on either side of it.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Every major assistant now ships some version of scheduled prompts: a user writes something like "every Monday at 9am, analyze last week's signup funnel and tell me what changed," picks a cadence, and it runs.
- The dev.to post argues that scheduled prompts are good for a news briefing but a mistake for analysis, and that the problem is not the scheduler but where the LLM sits relative to it.
- People schedule analysis because they want to know what changed, not because they want a report: a one-off analysis answers "what is the state of things" while a recurring analysis answers "is the state of things moving, and in which direction."
- Detecting change requires that the instrument hold still.
- If the query is regenerated from a natural language prompt on every run, a movement in the number has two possible causes, the world changed or the measurement changed, with no way to separate them; the post calls this textbook confounding built into the foundation of monitoring.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Every major assistant now ships some version of the scheduled prompt: you write "every Monday at 9am, analyze last week's signup funnel and tell me what changed," pick a cadence, and it runs [1]. A post on dev.to argues that this is fine for a news briefing and wrong for analysis, and that the fault lies not with the scheduler but with where the LLM sits relative to it [2]. The case is about measurement, not taste. Nobody schedules an analysis because they want a report: a one-off analysis tells you what the state of things is, while a recurring one tells you whether it is moving and in which direction [3]. That second question only works if the instrument holds still [4]. Regenerate the query from English on every run and a move in the number has two candidate causes, the world or the measurement, with no way to separate them [5]. That is confounding installed in the foundation of your monitoring [5]. The obvious rebuttal is that models are good at SQL now. The post's answer is that accuracy is not the issue and consistency is, because a model that writes correct SQL every single time can still write differently correct SQL each run [6]. Its worked example: run one counts distinct users with a LEFT JOIN to subscriptions over created_at >= '2026-08-01' and < '2026-09-01'; two weeks later the same prompt yields an INNER JOIN, so users with no subscription silently vanish, plus > and <= boundaries that shift the window by a day [7]. Neither query throws an error, both are defensible readings of the same English sentence, and the number moves by a few percent [8]. The author's recurring suspects are LEFT JOIN quietly becoming INNER JOIN, boundary operators flipping, NOT IN versus NOT EXISTS and their NULL handling, COUNT(*) versus COUNT(DISTINCT) on a fanned-out join, and UTC versus local day boundaries [9]. This is worse than an outright failure because it is silent, plausible, and unreconstructable: six months on, the query that produced March's number no longer exists anywhere [10]. The proposed fix is positional. Move the LLM from runtime to build time and split the work into five phases: explore, freeze, review, run, interpret [11]. Explore stays interactive and LLM-heavy, and nondeterminism there is a feature [14]. Freeze means the model emits SQL or a script that goes into version control, so the metric definition becomes an artifact with a name and a hash [12]. Review is the step that looks bureaucratic and, per the author, is the single most valuable one: a human reads the diff, which converts "the metric definition changed" from an invisible accident into an explicit, attributable, reviewable event, something a scheduled prompt cannot offer because the change happens inside a sampling distribution [13]. The run itself is deterministic code on a deterministic schedule, with the model returning downstream to interpret [18]. Two of the five phases use the model, and neither of them is the run [1]. The BI vendors got here first. The post points to dbt's Semantic Layer, Omni, Dremio and Cortex Analyst as sharing one thesis, that metric definitions must be codified so the LLM chooses which metric rather than how to compute it [15], with dbt's own framing being that fixed logic stops a model producing correct-looking numbers that differ subtly between runs [16]. What the author finds odd is that this argument is conducted in data engineering vocabulary about text-to-SQL and governance, while the "schedule an AI task" conversation happens in productivity blogs, and the two have not met [17]. The test to run on any recurring AI task you already operate is whether you can produce the exact query that generated last month's number [10]. If the answer is no, the time series is not a time series.