Skip to content

Build1 publisher3 min readPublished

Filtering hidden prompt-injection text takes a separate rule for each kind of invisible Unicode

In a dev.to test, deleting every invisible character from a chat reply removed 25 hidden tag characters and the two joiners that hold a family emoji together. The post's fix is to detect Unicode's whole default-ignorable set, then decide class by class what to delete.

The Engineer · Build desk

Illustration accompanying Filtering hidden prompt-injection text takes a separate rule for each kind of invisible Unicode

What happened

  • A hand-written regex matching U+200B to U+200D and U+FEFF deleted the joiners but left all 25 tag characters, so the emoji broke and the hidden instruction still reached the model.
  • Unicode 18.0's DerivedCoreProperties.txt lists 4,174 default-ignorable code points, built from format characters, variation selectors and other listed code points minus white space.
  • No-break, narrow no-break and ideographic spaces are category Zs and take up width, so they sit outside the default-ignorable set and a check on that set does not flag them.
  • Trojan Source, CVE-2021-42574 from Nicholas Boucher and Ross Anderson, uses bidirectional controls to make source code display in a different order than the compiler reads it.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A four-character deny-list leaves 4,170 of the 4,174 default-ignorable code points unfiltered, so its coverage depends on attackers choosing the four characters it knows.
  • constraint A sanitizer that deletes the whole default-ignorable set cannot keep ZWJ emoji such as the family or rainbow flag, or Persian text that needs ZWJ and ZWNJ to join letters.
  • decision The same bidirectional mark has to be rejected in code, file names and identifiers and kept in Arabic and Hebrew prose, so the cleaner has to know what kind of field it is handling.

The viewer in the post got detection right. On 2026-09-29 it labeled both joiners and all 25 tag letters in the test reply. Then its one-click strip button returned the family emoji as three separate people [4]. The hidden sentence spelled out a harmless canary word, BANANA. In an attack it would be an instruction to the model [2].

The regex fails in a different way, and the post calls it the worse of the two approaches [5]. This is the line:

``` const naive = reply.replace(/[\u200B-\u200D\uFEFF]/g, ''); ```

Its character class holds four code points: U+200B through U+200D, plus U+FEFF [10]. The zero width joiner, U+200D, is one of the four [1]. The tag characters sit at 0xE0000 and above, and the post's check finds them there with `codePointAt(0) >= 0xE0000` [5].

Unicode already publishes a set built for detection. UAX #44 describes Default_Ignorable_Code_Point as "for programmatic determination of default ignorable code points" [6]. Most of its entries are reserved. The range U+E01F0..U+E0FFF alone holds 3,600 unassigned code points [9], about 86 percent of the set [1]. UAX #44 also says that "new characters that should be ignored in rendering (unless explicitly supported) will be assigned in these ranges" [7]. A property lookup covers those characters when they arrive. A hand-written list has to be edited each time one is assigned [7].

"DI is the correct set to detect. It is the wrong set to delete," the post says [11]. It sorts the invisible characters into four families, each with its own policy, and offers a cleaning function meant to remove hidden text and keep emoji [20]. I think that split is right for any pipeline that passes user text to a model. The test reply shows both ways to fail inside 73 code points, 27 of them invisible [1].

For bidirectional controls, the right policy also depends on the field. PropList.txt lists 12 Bidi_Control code points [16]. In stretched-string.js from the Trojan Source repository, an RLO, an LRI, a PDI and an LRI make part of a string literal look like a closing quote and a comment [22]. With the controls removed, the real comparison is visible. The variable holds "user" and the literal is longer, so the two never match and the admin branch runs [22]. Boucher and Anderson recommend that compilers and build pipelines "throw errors or warnings for unterminated bidirectional control characters in comments or string literals" [18]. MITRE ATT&CK T1036.002 describes attackers who use U+202E "to disguise a string and/or file name to make it appear benign" [19].

Blank characters that take up width need their own rule. In April 2025 Rumi reported U+202F in o3 and o4-mini output and quoted OpenAI's reply that it is "a quirk of large-scale reinforcement learning" [13]. U+2800 BRAILLE PATTERN BLANK is category So, and it takes up width too [12]. The post puts these characters under a separate whitespace policy [12].

What to watch

  • Whether online invisible-character viewers change their one-click strip so it keeps ZWJ emoji sequences while still removing tag characters.
  • New Unicode versions assigning characters in the default-ignorable ranges, which a property-based check picks up and a fixed regex does not.
  • Whether LLM providers change model output that carries U+202F, which OpenAI described as a quirk of large-scale reinforcement learning.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories