Build1 publisher3 min readPublished
The fence was fine, the test was green, and the injection still worked
A maintainer pointed his own multi-model CLI at its own repository and found both a prompt-injection hole and the passing test that was supposed to prove the hole did not exist.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- llm-council is a small CLI that puts one question to several models, hides the authorship, and has the models rank each other's answers; the maintainer uses it as an adversarial reviewer.
- On 26 July the maintainer pointed llm-council at its own repository, and it found a prompt-injection hole in its own prompts.
- The second finding was that a test had already been written for exactly that hole, and the test was green.
- In llm-council, stage 1 collects answers, stage 2 asks a model to rank them, stage 3 asks for a synthesis, and every stage feeds the previous stage's text, written by an untrusted party, into a new prompt.
- The author describes the situation as OWASP LLM01 in its plainest form, with the standard mitigation being fencing: wrapping untrusted content in delimiters and telling the reader that anything inside is quoted data, never instructions.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
On 26 July the maintainer of llm-council, a CLI that puts one question to several models and has them rank each other's answers with authorship hidden, pointed the tool at its own repository and it reported a prompt-injection hole in its own prompts [1][2]. The more useful finding was the second one: a test for exactly that hole already existed, and it was passing [3].
The setup is the ordinary one for chained models. Stage 1 collects answers, stage 2 asks a model to rank them, stage 3 asks for a synthesis, and each stage feeds the previous stage's text, written by an untrusted party, into a new prompt [4]. The author identifies this as OWASP LLM01 and the mitigation he applied as fencing: wrap the untrusted content in delimiters and tell the reader that anything inside is quoted data, not instructions [5].
The delimiters were fixed strings of the form `<<<{kind}_{label}_BEGIN>>>` and `<<<{kind}_{label}_END>>>`, sitting in a public repository [6]. So a hostile voter, or a model that had read the repo during training, could write `<<<RESPONSE_A_END>>>` in the middle of its own answer; to the model reading downstream the block is now closed, and everything after it reads as orchestrator text [7].
Now the test. It was named `test_a_voter_cannot_forge_another_fence_boundary`, and what it actually asserted was that the output contained the string `<<<RESPONSE_B_END>>>` exactly twice and that the forged text appeared before `<<<RESPONSE_A_END>>>` [8]. Both of those hold whether or not the attack works; as the author puts it, the test verifies that string concatenation concatenated [9]. It was not empty and not skipped. It ran, it exercised real code, and it would have caught a genuine refactoring mistake, while never touching the property in its own name, which is the part everyone reads when deciding whether an area is covered [10]. The suite was at 100% coverage, a number about lines executed, not about where assertions are aimed [11].
The repair moves the defence off the shape of the markers and onto something the attacker has not seen: a per-run nonce baked into the marker template [12]. The nonce comes from `secrets.token_hex(8)`, with the code comment noting that `secrets` rather than `random` is the point, because a predictable PRNG returns exactly what the nonce was meant to remove [13]. That is 8 random bytes, 16 hex characters, 64 bits per run [16]. The comment in the fixed code states the principle plainly: the nonce is the defence, not the shape of the markers [14].
The rewritten test is `test_forged_markers_never_match_the_run_nonce`, feeding marker-shaped payloads through `stage3_prompt` [15]. Note what it still does not do: it asserts a property of the prompt text, not that the downstream model treated the fenced region as data [17]. That is the honest ceiling here. Unguessability is checkable locally; model obedience is not.
Worth doing this week: grep your own suites for test names that promise a security property, and read the assertions underneath them. Then check whether any delimiter in a public repo of yours is a fixed string, and whether removing the fence entirely makes a single test go red [6][11]. The write-up was submitted to DEV's Summer Bug Smash [18].