UTF-8
variable-width encoding (into one to four bytes) and transformation format of code points for the universal character set defined by ISO/IEC 10646 and The Unicode® Standard, compatible with ASCII
Episodes
-
#5487: Why Your Spreadsheet Mangles José's NameA deep dive into character sets, from ASCII to UTF-8, and why José becomes "José" in your spreadsheet.Main topic -
#5488: How Emoji Went From 176 Characters to 3,600From a 176-character set built for Japanese pagers to a 3,600-character Unicode standard — and why your emoji breaks in production. -
#5486: Why Hebrew URLs Turn Into Percent-Sign GibberishHebrew renders fine in the domain but explodes into percent-hex in the path. Two different standards explain the split.