Build1 publisher3 min readPublished
One sanitizer strips invisible Unicode from every traveller's text before it reaches a prompt
One developer's travel site deletes all 128 Unicode tag characters, plus zero-width and bidi controls, from user text before a model reads it. That closes the hidden-text channel outright, though typed-out injection still has to be caught by code checks downstream.
The Engineer · Build desk

What happened
- Two zero-width joiners survive only in the function whose output returns to travellers, because some scripts and emoji need them.
- Untrusted text is wrapped in tags with an eight-character random boundary that changes on every request, so it cannot be closed from inside.
- Every quote the model returns is searched for in the report it cites, and any passage that is not found gets dropped.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability An instruction spelled in tag characters is gone before the prompt is assembled, so this layer holds whether or not the model follows its security rules.
- constraint Coverage ends at the listed ranges, so an invisible code point outside them reaches the model until someone adds it to the regex.
- decision Any team that echoes user text back has to decide per code path whether joiners stay, since stripping them breaks some scripts and emoji.
- cost Catching typed-out injection means writing a code check for every output field's shape, type, length and allowed values before anything is stored.
I checked the order of the calls in sanitizeForLLM first, and it is right. Text is normalized to NFC. Then the tag block, zero-width and bidi ranges are deleted, then control characters, then repeat runs, and finally the text is cut to length [2]. The same function sits in front of reports, questions, answers and edits [1]. Deletion runs before the repeat collapse, so a run of one letter padded out with zero-width characters becomes a plain run and gets cut to ten [13]. The author caps runs because a thousand repeated "a" characters is a token bomb that costs money and does nothing else [5].
The first range is the Unicode tag block. Its characters map one to one onto ASCII and render as nothing [3]. "You can write a full sentence in it. Your screen shows a blank. A model reads the sentence," the author wrote [3]. The range in the regex, U+E0000 to U+E007F, is 128 code points [4]. All of them are deleted before any prompt is built.
The joiner flag is the detail I would question at review, and it holds up. With preserveJoiners set, the zero-width pattern skips U+200C and U+200D [7]. Some scripts and emoji sequences need those two, so only the function whose output goes back to the traveller keeps them [8]. The analysis paths strip them. "There, a joiner is only useful to an attacker," the author wrote [9].
Two smaller lines show the same care. slice() cuts UTF-16 code units, so the function drops a trailing lone high surrogate that the cut can leave behind [10]. The post says control characters go except newline and tab [11]. The ranges in the code also leave U+000D, the carriage return, in place [12].
The post is clear about where stripping stops. It does nothing against "ignore previous instructions" typed in plain letters [14]. For that case, each input is wrapped in tags named with eight characters from crypto.randomUUID(), fresh on every request [15]. To close the fence from inside the text, an attacker has to guess those characters first [15]. The prompt tells the model that everything inside the tags is data authored by users and never instructions [16]. The author does not rely on that rule alone. "Prompts are suggestions," the author wrote [17].
Most of the protection is code on the output side. Every field the model returns is checked for shape, type, length and allowed values before it touches the database, and anything that fails is dropped [18]. On the question page, each quoted passage is lowercased, stripped of punctuation and searched for in the report it cites. A miss is dropped [19]. According to the code comment, the model can be talked into lying but cannot make a quote appear in the source text [20].
The evidence supports stripping at input as the one layer that removes the hidden channel. The author's own words for the goal are "Not resisted. Useless." [24] The claim that a prompt rule fails is the author's judgement. The post does not report attack attempts, test cases or failure rates [23]. The function is also a denylist: a code point outside the listed ranges passes through unchanged [21]. For a site whose users write trip reports, the deleted ranges hold nothing a traveller needs, according to the author [22]. In that context I would adopt the function as written and add ranges when attempts using other characters turn up.
What to watch
- Whether the author publishes attack attempts or test cases showing what reached the model with and without the sanitizer.
- Whether the denylist grows to cover invisible code points outside the current ranges as new attempts turn up.