Skip to content

Build1 publisher3 min readPublished

Three freshness rules, 8,900 characters, and a fortnight-old ETF shutdown shipped as fact

An automated news watch republished an August 3 announcement on August 17 with a verified-fact label. The useful fix was not a longer prompt but six lines that grade the output.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The top item on the feed read "Last trading day for the first US spot Bitcoin ETF to shut down", was tagged as verified fact, and carried three sources all dated August 3rd.
  • The date at the time of observation was August 17th.
  • Fourteen days separated the most recent source from the event the item announced.
  • The author states he had written a rule a week earlier saying in plain words never to copy a two-week-old announcement as if it were still true.
  • The site runs an automated watch: every 24 hours a scheduled task wakes a Node script that calls a language model with web search access, asks it to sweep the last few hours of crypto and financial news, and demands strict JSON in return.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

On August 17, the top item on an automated crypto and finance watch page read "Last trading day for the first US spot Bitcoin ETF to shut down", tagged as verified fact, with three sources all dated August 3 [1][2]. Fourteen days separated the newest source from the event it announced, and the site's operator says he had written a rule a week earlier forbidding precisely that [3][4].

The pipeline is ordinary. Every 24 hours a scheduled task wakes a Node script that calls a language model with web search access, asks it to sweep the last few hours of crypto and financial news, and demands strict JSON back [5]. That JSON renders a public page of items, each with a title, a summary, its sources and their publication dates [6]. The instruction text is 8,900 characters covering categories, preferred sources, output format and freshness rules [7], and it requires a certainty label on every item, from FAIT_VERIFIE when two independent primary sources agree down to SPECULATIF for a hypothesis [8].

The stored rule is not vague. If an item's sources are all older than the cycle window, the model may not publish it as is: search for a recent source, add it and keep FAIT_VERIFIE, or downgrade to PROBABLE and say in the summary that the deadline has not been reconfirmed [12]. The output honoured neither branch. No recent source, no downgrade, no caveat, published with the confidence of something confirmed that morning [13]. At a 24-hour cycle, sources 14 days old sit 336 hours out, fourteen times the window the rule was written to police [17].

There is a structural reason the failure is hard to see. The prompt is a template: double-brace markers such as {{DATE}}, {{FREQUENCY_HOURS}} and {{PRICES}} are filled in by the script immediately before the call, so the text the model receives is never exactly the text in the editor [9][10]. The author places the defect at that fill-in step, and notes the model only ever sees the rendered output, so anything that goes wrong there surfaces nowhere else [11]. His own headline calls them three rules that never reached the model [22]. The published account breaks off mid-sentence before the mechanism and the measured comparison arrive [21], so treat the specific corruption as unresolved.

The instinct in that situation is to blame disobedience and push harder: more capitals, the instruction repeated twice [14]. He says that instinct cost him the most time [14]. The prompt was genuinely badly built, to be fair: four blocks each declaring itself top priority, a freshness rule and a reconfirmation rule that contradict each other on edge cases, and exactly one concrete example across 8,900 characters [15]. The rewrite added section tags, an explicit priority ranking, a four-branch numbered freshness procedure and two worked examples, reaching 11,800 characters [16]. That is 2,900 more characters of specification, roughly a third longer, and still not one line of enforcement [18].

The enforcement is the part worth copying. Because the output is JSON, the rule is gradeable without anyone reading the prose: parse the source dates, take the newest, mark the item stale if it falls outside the window, test the summary for a hedge, and flag a violation when a stale item is either labelled FAIT_VERIFIE or carries no caveat [19][20]. Six lines, and a run stops being an aesthetic judgement about prompt quality and becomes a count of items that break the rule [20][23].

What to watch: whether the comparison of old prompt against new ever produces numbers, since the account stops before it [21]; and whether the check runs as a publish gate or only as an after-the-fact audit. A rendered prompt is an artefact you can assert against, the same way the JSON is.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories