Build1 publisher3 min readPublished
A million-row CSV sort ran 140x slower in Electron's renderer than the same code in Node
Sorting a million-row CSV in an Electron renderer took over a minute, 140x slower than the same code in a Node script, its developer found. The slowdown showed up only in a renderer already holding a large file, so a profile taken there was the only way to find it.
The Engineer · Build desk

What happened
- The developer first blamed the DOM, the comparator and the data size; swapping per-comparison localeCompare for a shared Intl.Collator helped, but by nothing close to the size of the gap.
- A renderer profile put a single line on top of the flame graph: a Map.set that counts each cell value using the cell string as the key.
- Run twice inside the app, hashing a million cell strings into a brand-new Map went about 30 times faster on the second pass.
- The author attributes the cost to V8 hashing each string lazily on its first use as a key, when every parsed cell is a fresh, never-hashed slice of the file buffer.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Electron teams have to time hot paths inside the renderer with production-sized data loaded, because a renderer holding only a small file ran the same first pass at normal speed.
- exposure Any renderer code that counts or dedupes freshly parsed strings in a Map can hit the same first-touch cost, and the author cannot say which renderer state triggers it.
- cost The workaround leaves the app owning a hand-rolled hash table with collision buckets, built around an engine behaviour its author cannot fully explain.
- contradiction The post's headline says 30x against Node while its body measures 140x; the body's 30x compares two passes inside the renderer, so quoting 30x as the Node gap understates it.
The post's debugging is careful, and two controls make the diagnosis convincing. The test Map was created fresh on each pass [10]. So whatever made the second pass cheaper had to be something the strings carried over from the first. The second control ruled out slow strings in general: `length`, `charCodeAt`, `trim()` and `Number()` over the same values ran at normal speed on the first pass [11].
According to the author, V8 computes a string's hash the first time the string is used as a key and caches it in the string itself [13]. The CSV parser returns every cell as a new slice of the file buffer. A million-row sort therefore starts with a million first-time hashes [14]. On the second pass, each string already held its hash [13].
That accounts for the gap between passes. It does not account for the gap with Node, and the author does not claim to. "I do not have a confident explanation for why the same operation is so much cheaper in a Node process, or exactly which part of the renderer's state makes first-touch hashing expensive," the developer wrote [15]. The one clue is state. The slow first pass appeared only after a large file was loaded in that renderer. A fresh renderer with a small file ran it fine [12].
For the half-second Node result to predict the renderer, the Node process would have to reproduce the renderer state that makes first-touch hashing expensive [3]. The author has not identified that state [15]. A renderer benchmark on a small fixture would also have passed [12]. The profile that found the problem came from the renderer with the real 55 MB file loaded [1][9].
The post's figures do not agree with each other. Its headline says 30x against Node [5]. Its body puts the sort at 140x [4], and uses 30x for first pass against second pass inside the renderer [10]. On a half-second baseline, 140x is about 70 seconds [1]. That matches a freeze of over a minute [2].
The fix stops V8 from hashing the strings at all. A `StringTable` class runs FNV-1a over `charCodeAt` with `Math.imul`, masks the result with `0x3fffffff`, and keys the Map by that integer. Strings are compared only when two values collide [17]. The mask is 2^30 minus 1, so every key fits in 30 bits [2]. The developer called the approach "reimplementing a hash map on top of a hash map" [18]. I would flag that class in review, then approve it after seeing the flame graph. The author's argument for it is that an integer-keyed Map skips string hashing entirely, a `charCodeAt` loop is fast, and collisions are rare [19].
Having the distinct values also allows a cheaper sort. Each row becomes one number in a `Float64Array`. For a text column that number is the value's rank in collation order. The collator then runs only over distinct values, which the author says are usually a tiny fraction of the rows [20]. The available text of the post stops partway through that code, before any timing for the fixed sort. "I know what I measured and that avoiding it fixed the problem," the developer wrote [16].
What to watch
- An explanation from V8 or Electron engineers of why first-touch string hashing is expensive in a loaded renderer; a heap-size cause would make a large-heap Node run a usable proxy.
- A reproduction in plain Node with a comparably large heap: if the slowdown appears there, the cause is heap state, not Electron.
- Published timings for the StringTable-based sort on the same million-row, 55 MB file.