Skip to content

Build1 publisher3 min readPublished

Tokenize-then-compress works, but the win is window coverage, not the 45% smaller stream

A 452-configuration benchmark of the parmar pre-filter splits two effects that looked like one: denser tokens buy a flat 15% on gzip, while the gain that scales comes from window expansion.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Tokenize-then-compress works, but the win is window coverage, not the 45% smaller stream
Generated illustration

What happened

  • Two developers, @ronak-create and @u84u, built parmar, a subword-tokenization pre-filter for byte-level compressors, plus a benchmarking rig, and published the results on dev.to.
  • The premise tested: replacing UTF-8 prose with BPE token IDs (the same tokenization used to feed LLMs) yields a token stream roughly 45% smaller than the source text.
  • Byte-level compressors such as LZMA2 and zstd find repeated patterns inside a fixed-size sliding dictionary window measured in bytes.
  • A 64 MiB dictionary window that normally covers about 64 MB of prose can cover roughly twice as much prose once the prose has been pre-shrunk by tokenization.
  • The window effect is only testable once the corpus is bigger than the compressor's window; on a 5 MB file everything already fits inside the window, so there is no expansion to measure and pre-tokenization buys basically nothing.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Two developers writing on dev.to, @ronak-create and @u84u, built a subword-tokenization pre-filter called parmar and then spent most of their effort on the rig that tests it: 452 benchmark configurations sweeping tokenizer, packing, compressor backend, transport, chunking and threading across four corpus tiers [1][8]. The idea holds up, but the mechanism is narrower than the intuition, and the narrow version changes who should bother.

The intuition is that BPE token IDs are a denser way to write English than UTF-8, roughly 45% smaller as a stream [2]. The mechanism that actually scales is a different one. LZMA2 and zstd look for repeats inside a fixed-size sliding dictionary window measured in bytes [3], so a 64 MiB window that covers about 64 MB of prose covers roughly twice as much prose once that prose has been pre-shrunk [4]; at a 45% reduction the arithmetic is about 1.8x [17]. None of that is measurable until the corpus is bigger than the window. On a 5 MB file everything already fits and pre-tokenization buys basically nothing [5]. So the deliverable is not a ratio, it is the curve of parmar ratio minus raw-byte ratio as corpus size grows [6].

The control that made the result readable was gzip. With a 32 KiB window gzip is saturated at every tier, and its +15% advantage is completely flat from 64 MB to 4 GB [12]: pure representation density, no window effect at all. On LZMA and zstd variants with multi-megabyte windows the gap climbs from 64 MB up through roughly 1 GB, then flattens once the corpus is well past the dictionary size [13]. zstd --long, with a 2 GiB window, was still climbing at the 4 GB tier, because 4 GB is only twice its window [14]. The authors say plainly that without the tiny-window control they would have credited window expansion for gains that were only density [15]. Worth noting how small the density effect is on its own: 15 points of advantage out of a representation that is 45% smaller, about a third of the shrink [20].

The harness discipline is the part worth copying. A literal cartesian product is 8,000-plus valid cells per tier, weeks of runtime, so they cut it to a 51-cell ratio grid plus a one-factor-at-a-time sweep for axes that should only move speed, while still logging ratio on the OFAT cells so that a supposedly speed-only axis that quietly moves ratio surfaces as a contradiction instead of being averaged away [10]. That is 204 ratio-grid cells across four tiers, with the remaining 248 of the 452 coming from the OFAT sweep [18], and roughly 1.4% of the full grid [19]. Every decompression was actually executed and checked against a sha256 in the archive footer, with failing cells excluded from averages and reported separately: 452 cells, 452 verified round trips, zero failures [11]. The corpus was PG-19, about 10,600 documents, tiered from 64 MB to 4 GB [9]. The pipeline never materializes the input: tokenize with tiktoken, pack IDs as LEB128 or a fixed 2-byte width, pipe into xz, zstd, gzip or bzip2 stdin [7].

It is not universal. On bzip2 pre-tokenization is a flat loss of about -3.9% at every corpus size, which the authors attribute to the Burrows-Wheeler transform [16]. Combined with the small-file result [5], the honest scope is: large-window backends, corpora well past the window, and a check that your backend is not bzip2.

What to watch is whether the zstd --long curve flattens the way LZMA's did once someone runs it at 8 GB or 16 GB, which is the prediction the plateau story makes [13][14], and whether the +15% density floor survives changes of tokenizer and packing, since only the ratio grid was built to move ratio at all [10].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories