HotShard
Absolute basics

Text Encoding & Unicode

Why the length of a string is a trick question, and how to sort names correctly.

Ask a program how long the text "café" is and you can get 4, 5, or 6. All three are correct. They answer different questions. The reason is one idea most people mix up: the difference between a character and the bytes that store it. This page pulls those two apart, shows how UTF-8 turns one into the other, and answers what an interviewer asks: which encoding, how long is a string, and how do you compare and sort text.

~5 min read

Start here: how long is "café"?#

TL;DRthe 30-second version
  • A character set gives every character a number, called a code point. That's Unicode's job. An encoding writes that number as bytes. That's UTF-8's job.
  • UTF-8 won the web. It uses 1 to 4 bytes per character, and the first 128 characters are one byte each, the same as old ASCII.
  • "String length" has four honest answers: bytes, code units, code points, and grapheme clusters (what a person sees as one character). They only agree for plain English.
  • The same visible text can be two different byte sequences. Normalize before you compare, hash, or store. And never sort names by byte value. Use a locale-aware collation.

The confusion has one root. We treat a character and the bytes that store it as the same thing. For plain English on an old system they nearly were, so a generation of code assumed one character equals one byte. That assumption is now wrong for most of the text on Earth: names, currencies, accents, emoji. "Length" is four different questions wearing one name.

Two layers, then four lengths#

There is the character itself: the letter "A", the euro sign "€", the grinning face. And there is the number we agree to call it by. Unicode is a giant table that gives every character a unique number, called a code point. We write code points as U+ followed by the value in hexadecimal. "A" is U+0041, which is 65. The grinning face 😀 is U+1F600. Unicode has room for about 1.1 million code points and has assigned over 150,000. A code point is just a number. It says nothing yet about bytes.

The encoding is the second layer. It's the rule that turns a code point into bytes on disk or on the wire, and back again. This step has choices. U+0041 could be stored as one byte, or two, or four, depending on the encoding, and the reader has to use the same encoding to get "A" back.

UTF-8 is the encoding that won. A character takes between 1 and 4 bytes, and UTF-8 picks the smallest that fits. Code points U+0000 to U+007F, the original 128 ASCII characters, are one byte with the exact value ASCII used. So every ASCII file ever written is already valid UTF-8, and English text costs one byte per letter.

CharacterCode pointUTF-8 bytesByte count
AU+0041411
éU+00E9C3 A92
€U+20ACE2 82 AC3
😀U+1F600F0 9F 98 804

UTF-8 is not the only encoding. UTF-16 uses 2 bytes for most characters and 4 for the rest. It's the in-memory string format of JavaScript, Java, C#, and Windows. Any character above U+FFFF, which includes every emoji, needs two 16-bit units. That pair is called a surrogate pair, and it's why a JavaScript string reports the length of 😀 as 2.

  • Bytes: how much storage the text takes in its encoding. "café" is 5 bytes in UTF-8.
  • Code units: the fixed-size pieces of the encoding, 8-bit for UTF-8 and 16-bit for UTF-16. This is what most languages' built-in length returns.
  • Code points: the count of Unicode characters, ignoring how they're stored.
  • Grapheme clusters: what a person sees as one character. "café" is 4 grapheme clusters however the é is stored.
TextBytes (UTF-8)Code pointsGrapheme clusters
A111
café (é as one code point)544
café (é as e + accent)654
😀411
👨‍👩‍👧‍👦 (family: 4 emoji joined by invisible joiners)2571
PredictA form limits a display name to 10 characters and enforces it with the string's built-in .length, which counts UTF-16 code units. A user pastes a name made of 6 emoji. What happens?

The form rejects it. Each emoji above U+FFFF is a surrogate pair, so 6 emoji count as 12 code units, over the limit of 10. The user typed 6 characters and got told they typed 12. A limit a person judges by eye should count grapheme clusters.

If this comes up in an interview#

The one-linerUnicode gives every character a number. UTF-8 writes that number as 1 to 4 bytes and is the default everywhere. String length depends on what you count. Normalize before comparing, and use a collation to sort.
Is "Unicode" an encoding?

No. Unicode is the character set: the table that gives every character a code point. UTF-8, UTF-16, and UTF-32 are the encodings: the rules for turning those code points into bytes. "Save as Unicode" in some old software actually meant "save as UTF-16". Say which encoding you mean.

Which encoding should I use?

UTF-8, for files, databases, JSON, HTML, and every network protocol. It's ASCII-compatible, it's the smallest for English and markup, and it's what everyone else expects. UTF-16 is something you inherit in memory in JavaScript, Java, and C#, not something you'd pick for the wire.

Why does JavaScript say an emoji has length 2?

JavaScript strings are UTF-16, and the built-in length counts 16-bit code units. An emoji above U+FFFF is a surrogate pair, so it counts as 2. When asked about string length, ask back: length in what, bytes, code points, or characters a person sees? For a user-facing limit, count grapheme clusters.

Why does "café" not equal "café"?

Because é can be written two ways: the single code point U+00E9, or a plain "e" (U+0065) followed by a combining acute accent (U+0301). Both look identical on screen, but they're different bytes. Normalization rewrites text into one canonical form. NFC (composed) is the usual choice; NFD (decomposed) is the other. Normalize before you compare, hash, or store. Skip it and a passphrase typed on a different keyboard hashes differently, and the login fails.

How do you sort a list of names?

Not by bytes. A byte sort puts "Zebra" before "apple", because "Z" is code point 90 and "a" is 97. Worse, the right order depends on the language: "ä" sorts next to "a" in German but after "z" in Swedish. The tool is collation, a locale-aware ordering. The Unicode Collation Algorithm turns each string into a sort key that weighs case, accents, and letter identity in the right priority. Databases expose it as a collation setting on a column.

UTF-8 vs UTF-16 vs UTF-32
UTF-8UTF-16UTF-32
Bytes per character1–42 or 4always 4
ASCII-compatibleYesNoNo
Size for English textSmallest (1 byte each)2× (2 bytes each)4× (4 bytes each)
Fixed-width indexingNoNo (surrogate pairs)Yes
Where you meet itWeb, files, APIs, everywhereJS/Java/C#/Windows in memorySome internal processing
  • Many systems keep strings in UTF-16 in memory but still read and write UTF-8 on disk and on the network.
  • UTF-32 makes character N sit at byte 4N, which is handy for some text-processing internals. But it's 4× the size of UTF-8 for English, so it almost never leaves the program that uses it.
  • One caveat for Chinese, Japanese, and Korean text: those characters are 3 bytes in UTF-8 but 2 in UTF-16. For text that's overwhelmingly CJK, UTF-16 can be smaller.
Pitfalls & gotchas
  • Mojibake (garbled text like é‚) is text written with one encoding and read with another. The bytes are fine; the reader used the wrong decoding rule. Declare the encoding on both ends with a charset header.
  • Reversing or truncating a string by bytes or code units splits multi-byte characters and surrogate pairs, and you get replacement boxes. Operate on grapheme clusters.
  • macOS has historically stored filenames in a decomposed (NFD-like) form while Linux stores whatever bytes you give it, so a file named "café" on one can fail to match the same name typed on the other until both are normalized.
  • Homoglyphs are different code points that look the same, like the Latin "a" (U+0061) and the Cyrillic "а" (U+0430). An attacker registers a lookalike domain or username. Identity-critical text is normalized to one form and often restricted to a safe set of characters before it's trusted.
References
References

Feedback on this topic →