← technical essays
[ESSAY]
No. 4.75 Aug 3, 2026 short essay

UTF-8 Is How Unicode Usually Arrives

Unicode is the inventory. UTF-8 is the shipping crate.

[ essay ]

Unicode is the catalog of characters. UTF-8 is how that catalog usually shows up as bytes. Mixing those jobs is how you get mojibake: the inventory was fine; the crate was opened with the wrong tool.

I publish text. I do not implement Unicode. mystic-bytes taught me the difference when a helper treated files as latin-1 because the bytes looked like English until they did not. An em dash became â€" and macrons in Tāmaki Makaurau turned into junk. Nothing was wrong with the character set. The pipeline had decoded UTF-8 as latin-1, or encoded latin-1 and then labeled the result UTF-8 — the classic lie, because western European text is a subset until the first character that is not.

Declare UTF-8 at every boundary: HTTP charset, XML prolog, editor, git attributes. Refuse unlabeled bytes. Never “fix” mojibake by guessing. If a file is not valid UTF-8, fail loud. Latin-1 compatibility is how corruption becomes a default instead of an incident.

Encoding is a parse, not a vibe. Unicode names the characters. UTF-8 is how they usually arrive. Keep those sentences apart.

— JV · Dark Heart Labs.

№ 4.75 — JV · Dark Heart Labs.