Unicode
computing industry standard for the consistent encoding, representation and handling of text expressed in most of the world's writing systems
Episodes
-
#5488: How Emoji Went From 176 Characters to 3,600From a 176-character set built for Japanese pagers to a 3,600-character Unicode standard — and why your emoji breaks in production.Main topic -
#5487: Why Your Spreadsheet Mangles José's NameA deep dive into character sets, from ASCII to UTF-8, and why José becomes "José" in your spreadsheet.Main topic -
#5486: Why Hebrew URLs Turn Into Percent-Sign GibberishHebrew renders fine in the domain but explodes into percent-hex in the path. Two different standards explain the split. -
#5461: TTS Can't Pronounce Hebrew Inside EnglishYour TTS reads Hebrew words with English phonetics. Here's why — and why the obvious fix doesn't work yet.