Build1 publisher3 min readPublished
Excel Guessed the Alphabet, and 4,685 Names Went Wrong Without an Error
A dev.to writeup traces mojibake in a CSV import to Excel silently assuming the wrong encoding. The mangled names then fail search, joins and grouping downstream.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- In the build behind the series, the encoding problem was found late, in a finished dashboard: two of the top fifteen games displayed as garbage, and a count against the file put the damage at 4,685 of 82,956 names. Nothing had errored at any point.
- 4,685 corrupted names out of 82,956 is about 5.6 percent of the name column.
- A text file on disk is numbered codes, one per character. The encoding is the codebook saying which number means which letter. The file does not carry its codebook visibly, so the program opening the file has to know, or guess.
- The build's CSV was written in UTF-8, which covers every language by spending two or more codes on any letter beyond plain English. Excel's legacy open path guessed the old Western European codebook instead, which reads one code per character, always.
- Because the wrong codebook reads one code per character, every two-code letter was read as two separate wrong characters; an e-acute stored as two codes came out as A-tilde followed by a copyright-style symbol.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A build described on dev.to shipped a finished dashboard before anyone noticed that two of its top fifteen game titles were rendering as strings of symbols; a count against the source file put the damage at 4,685 of 82,956 names, and nothing had errored at any point [s1c1]. That is roughly 5.6 percent of the name column silently wrong [1], which matters because the wrong values do not stay in the display layer.
The mechanism is dull and worth knowing. A text file on disk is numbered codes, one per character, and the encoding is the codebook saying which number means which letter; the file does not carry its codebook visibly, so the program opening it must know or guess [s1c2]. The CSV was written in UTF-8, which spends two or more codes on any letter beyond plain English, while Excel's legacy open path guessed the old Western European codebook, which reads one code per character, always [s1c3]. Every two-code letter therefore came back as two wrong characters: an e-acute stored as two codes surfaced as the familiar A-tilde pair [s1c4]. A Romanian title with four accented letters becomes eight wrong characters, which is how "Aventura Copilului Albastru" flagged the problem in the build's top fifteen [s1c5].
Calling this a display bug gets the direction backwards. A font problem means the stored value is correct and the pixels are wrong; here the pixels faithfully show a stored value that is now wrong [s1c6]. The consequences arrive in a predictable order, according to the writeup. Search fails, because the user types the real title and the cell holds the mangled one, with no match and no hint why [s1c7]. Joins fail: a lookup against a clean source list matches on name, so 4,685 rows match nothing, and a failed match does not error, it just drops or blanks the rows [s1c8]. Grouping splits, because the same publisher spelled cleanly in one file and mangled in another becomes two publishers, which is an entity-resolution problem manufactured out of nothing [s1c9].
The detection cost is near zero. Because the wrong codebook maps UTF-8's lead codes to a small set of characters, the corruption has a fingerprint, and Ctrl+F for the A-tilde character is the whole audit, taking about ten seconds [s1c10]. The source says every mapping in its table was re-run through the actual conversion before publishing [s1c11].
The prevention is one dropdown: Data > From Text/CSV rather than a double-click, then set File Origin in the preview to 65001: Unicode (UTF-8), and check a row you know has accents before loading [s1c12]. Paired with setting identifier columns to type Text, that is the whole defensive import, and it costs under a minute [s1c13].
Repair after the fact was cheaper than a full re-import in this case because the corruption is mechanical and reversible where the original file survives, and often reversible even from the mangled text, since the wrong reading was consistent [s1c14]. The build went back to the raw CSV, which existed because raw files never get overwritten, re-read only the damaged column under the correct encoding, exported it with a byte order mark so the paste source opens correctly on any machine, pasted it over the mangled column, and refreshed the pivots [s1c15].
Worth watching in your own pipelines: whether the raw file is still on disk, and whether anything between the import and the dashboard would have raised a single complaint. In this account, nothing did [s1c1].