Start here: how long is "café"?#
TL;DRthe 30-second version
- A character set gives every character a number, called a code point. That's Unicode's job. An encoding writes that number as bytes. That's UTF-8's job.
- UTF-8 won the web. It uses 1 to 4 bytes per character, and the first 128 characters are one byte each, the same as old ASCII.
- "String length" has four honest answers: bytes, code units, code points, and grapheme clusters (what a person sees as one character). They only agree for plain English.
- The same visible text can be two different byte sequences. Normalize before you compare, hash, or store. And never sort names by byte value. Use a locale-aware collation.
The confusion has one root. We treat a character and the bytes that store it as the same thing. For plain English on an old system they nearly were, so a generation of code assumed one character equals one byte. That assumption is now wrong for most of the text on Earth: names, currencies, accents, emoji. "Length" is four different questions wearing one name.
Two layers, then four lengths#
There is the character itself: the letter "A", the euro sign "€", the grinning face. And there is the number we agree to call it by. Unicode is a giant table that gives every character a unique number, called a code point. We write code points as U+ followed by the value in hexadecimal. "A" is U+0041, which is 65. The grinning face 😀 is U+1F600. Unicode has room for about 1.1 million code points and has assigned over 150,000. A code point is just a number. It says nothing yet about bytes.
The encoding is the second layer. It's the rule that turns a code point into bytes on disk or on the wire, and back again. This step has choices. U+0041 could be stored as one byte, or two, or four, depending on the encoding, and the reader has to use the same encoding to get "A" back.
UTF-8 is the encoding that won. A character takes between 1 and 4 bytes, and UTF-8 picks the smallest that fits. Code points U+0000 to U+007F, the original 128 ASCII characters, are one byte with the exact value ASCII used. So every ASCII file ever written is already valid UTF-8, and English text costs one byte per letter.
| Character | Code point | UTF-8 bytes | Byte count |
|---|---|---|---|
| A | U+0041 | 41 | 1 |
| é | U+00E9 | C3 A9 | 2 |
| € | U+20AC | E2 82 AC | 3 |
| 😀 | U+1F600 | F0 9F 98 80 | 4 |
UTF-8 is not the only encoding. UTF-16 uses 2 bytes for most characters and 4 for the rest. It's the in-memory string format of JavaScript, Java, C#, and Windows. Any character above U+FFFF, which includes every emoji, needs two 16-bit units. That pair is called a surrogate pair, and it's why a JavaScript string reports the length of 😀 as 2.
- Bytes: how much storage the text takes in its encoding. "café" is 5 bytes in UTF-8.
- Code units: the fixed-size pieces of the encoding, 8-bit for UTF-8 and 16-bit for UTF-16. This is what most languages' built-in length returns.
- Code points: the count of Unicode characters, ignoring how they're stored.
- Grapheme clusters: what a person sees as one character. "café" is 4 grapheme clusters however the é is stored.
| Text | Bytes (UTF-8) | Code points | Grapheme clusters |
|---|---|---|---|
| A | 1 | 1 | 1 |
| café (é as one code point) | 5 | 4 | 4 |
| café (é as e + accent) | 6 | 5 | 4 |
| 😀 | 4 | 1 | 1 |
| 👨👩👧👦 (family: 4 emoji joined by invisible joiners) | 25 | 7 | 1 |
PredictA form limits a display name to 10 characters and enforces it with the string's built-in .length, which counts UTF-16 code units. A user pastes a name made of 6 emoji. What happens?
The form rejects it. Each emoji above U+FFFF is a surrogate pair, so 6 emoji count as 12 code units, over the limit of 10. The user typed 6 characters and got told they typed 12. A limit a person judges by eye should count grapheme clusters.
If this comes up in an interview#
Is "Unicode" an encoding?
No. Unicode is the character set: the table that gives every character a code point. UTF-8, UTF-16, and UTF-32 are the encodings: the rules for turning those code points into bytes. "Save as Unicode" in some old software actually meant "save as UTF-16". Say which encoding you mean.
Which encoding should I use?
UTF-8, for files, databases, JSON, HTML, and every network protocol. It's ASCII-compatible, it's the smallest for English and markup, and it's what everyone else expects. UTF-16 is something you inherit in memory in JavaScript, Java, and C#, not something you'd pick for the wire.
Why does JavaScript say an emoji has length 2?
JavaScript strings are UTF-16, and the built-in length counts 16-bit code units. An emoji above U+FFFF is a surrogate pair, so it counts as 2. When asked about string length, ask back: length in what, bytes, code points, or characters a person sees? For a user-facing limit, count grapheme clusters.
Why does "café" not equal "café"?
Because é can be written two ways: the single code point U+00E9, or a plain "e" (U+0065) followed by a combining acute accent (U+0301). Both look identical on screen, but they're different bytes. Normalization rewrites text into one canonical form. NFC (composed) is the usual choice; NFD (decomposed) is the other. Normalize before you compare, hash, or store. Skip it and a passphrase typed on a different keyboard hashes differently, and the login fails.
How do you sort a list of names?
Not by bytes. A byte sort puts "Zebra" before "apple", because "Z" is code point 90 and "a" is 97. Worse, the right order depends on the language: "ä" sorts next to "a" in German but after "z" in Swedish. The tool is collation, a locale-aware ordering. The Unicode Collation Algorithm turns each string into a sort key that weighs case, accents, and letter identity in the right priority. Databases expose it as a collation setting on a column.
UTF-8 vs UTF-16 vs UTF-32
| UTF-8 | UTF-16 | UTF-32 | |
|---|---|---|---|
| Bytes per character | 1–4 | 2 or 4 | always 4 |
| ASCII-compatible | Yes | No | No |
| Size for English text | Smallest (1 byte each) | 2× (2 bytes each) | 4× (4 bytes each) |
| Fixed-width indexing | No | No (surrogate pairs) | Yes |
| Where you meet it | Web, files, APIs, everywhere | JS/Java/C#/Windows in memory | Some internal processing |
- Many systems keep strings in UTF-16 in memory but still read and write UTF-8 on disk and on the network.
- UTF-32 makes character N sit at byte 4N, which is handy for some text-processing internals. But it's 4× the size of UTF-8 for English, so it almost never leaves the program that uses it.
- One caveat for Chinese, Japanese, and Korean text: those characters are 3 bytes in UTF-8 but 2 in UTF-16. For text that's overwhelmingly CJK, UTF-16 can be smaller.
Pitfalls & gotchas
- Mojibake (garbled text like é‚) is text written with one encoding and read with another. The bytes are fine; the reader used the wrong decoding rule. Declare the encoding on both ends with a charset header.
- Reversing or truncating a string by bytes or code units splits multi-byte characters and surrogate pairs, and you get replacement boxes. Operate on grapheme clusters.
- macOS has historically stored filenames in a decomposed (NFD-like) form while Linux stores whatever bytes you give it, so a file named "café" on one can fail to match the same name typed on the other until both are normalized.
- Homoglyphs are different code points that look the same, like the Latin "a" (U+0061) and the Cyrillic "а" (U+0430). An attacker registers a lookalike domain or username. Identity-critical text is normalized to one form and often restricted to a safe set of characters before it's trusted.
References
- The Absolute Minimum Every Software Developer Must Know About Unicode (Joel Spolsky) — The classic essay on the character-set-vs-encoding split.
- RFC 3629 — UTF-8, a transformation format of ISO 10646 — The formal definition of UTF-8's byte layout.
- UAX #15 — Unicode Normalization Forms — NFC / NFD / NFKC / NFKD, and when each is canonical.
- UTS #10 — Unicode Collation Algorithm — How locale-aware sorting turns strings into comparable sort keys.