Computer science
Unicode, UTF-8 and UTF-16 codes of a text
Paste a text and see each character with its code point, its UTF-8 and UTF-16 bytes and its HTML code. Or type the codes, like U+1F600 or é, and get the characters.
Characters, code points and bytes
Unicode gives every character of every script a number, the code point, written as U+ followed by hexadecimal digits: A is U+0041, è is U+00E8, the euro sign U+20AC, the smiling face U+1F600. There are over a million positions, about 150,000 of them in use.
The code point then has to become bytes. UTF-8, the encoding of the web, uses one to four bytes: one for ASCII, two for accented letters, three for almost everything else, four for emoji and rare symbols. UTF-16, used by JavaScript, Java and Windows, uses two bytes or a surrogate pair of four.
That is why the length of a string depends on what is counted: "😀".length is 2 in JavaScript, the character takes 4 bytes in UTF-8 and to the eye it is one. And an accented letter can be a single code point or a letter plus a combining accent: they look the same but are not.
Common mistakes
- Cutting a UTF-8 string at a fixed number of bytes: a character may be split in half.
- Counting characters with .length in JavaScript: emoji and characters outside the basic plane count as 2.
- Comparing texts without normalising them: a precomposed é and e plus a combining accent differ to the computer.
Frequently asked questions
What is the difference between Unicode and UTF-8?
Unicode is the list of characters and their numbers; UTF-8 is one way of writing those numbers as bytes. The same Unicode text can be saved as UTF-8, UTF-16 or UTF-32.
Why do I see odd characters like é instead of é?
Because a UTF-8 text is being read as Latin-1: the two bytes of é, C3 A9, become two separate characters. Declaring the right encoding fixes it.
How do I write a Unicode character in HTML?
With a numeric entity: é in decimal or é in hexadecimal. If the page is in UTF-8, you can also type the character directly.
How this calculation works
UTF-8: up to U+007F one byte 0xxxxxxx; up to U+07FF two bytes 110xxxxx 10xxxxxx; up to U+FFFF three bytes 1110xxxx 10xxxxxx 10xxxxxx; beyond, four bytes 11110xxx and three continuation bytes. UTF-16: up to U+FFFF one 16-bit unit; beyond, 0x10000 is subtracted and the 20 bits are split into two surrogates, 0xD800 + the high 10 bits and 0xDC00 + the low 10.
Related calculators
IEEE 754 floating point
How a decimal is stored as a 32 or 64-bit float: sign, exponent, mantissa and the error.
Base64 converter
From text, integer, hexadecimal or binary to Base64 and back, with the 6-bit groups explained.
Data size and download time
From bytes to gigabytes and gibibytes, and how long a file takes to download at a given speed.
Colour converter
From HEX to RGB, HSL, HSV and CMYK and back, with the WCAG contrast of white and black text.