Characters, code points and bytes

Unicode gives every character of every script a number, the code point, written as U+ followed by hexadecimal digits: A is U+0041, è is U+00E8, the euro sign U+20AC, the smiling face U+1F600. There are over a million positions, about 150,000 of them in use.

The code point then has to become bytes. UTF-8, the encoding of the web, uses one to four bytes: one for ASCII, two for accented letters, three for almost everything else, four for emoji and rare symbols. UTF-16, used by JavaScript, Java and Windows, uses two bytes or a surrogate pair of four.

That is why the length of a string depends on what is counted: "😀".length is 2 in JavaScript, the character takes 4 bytes in UTF-8 and to the eye it is one. And an accented letter can be a single code point or a letter plus a combining accent: they look the same but are not.

Common mistakes

  • Cutting a UTF-8 string at a fixed number of bytes: a character may be split in half.
  • Counting characters with .length in JavaScript: emoji and characters outside the basic plane count as 2.
  • Comparing texts without normalising them: a precomposed é and e plus a combining accent differ to the computer.

Frequently asked questions

What is the difference between Unicode and UTF-8?

Unicode is the list of characters and their numbers; UTF-8 is one way of writing those numbers as bytes. The same Unicode text can be saved as UTF-8, UTF-16 or UTF-32.

Why do I see odd characters like é instead of é?

Because a UTF-8 text is being read as Latin-1: the two bytes of é, C3 A9, become two separate characters. Declaring the right encoding fixes it.

How do I write a Unicode character in HTML?

With a numeric entity: é in decimal or é in hexadecimal. If the page is in UTF-8, you can also type the character directly.

How this calculation works

UTF-8: up to U+007F one byte 0xxxxxxx; up to U+07FF two bytes 110xxxxx 10xxxxxx; up to U+FFFF three bytes 1110xxxx 10xxxxxx 10xxxxxx; beyond, four bytes 11110xxx and three continuation bytes. UTF-16: up to U+FFFF one 16-bit unit; beyond, 0x10000 is subtracted and the 20 bits are split into two surrogates, 0xD800 + the high 10 bits and 0xDC00 + the low 10.