What Base64 is and when it is used

Base64 is a way of writing arbitrary bytes with only 64 safe characters: upper- and lower-case letters, digits, plus and slash. It is needed whenever binary data has to travel through a channel built for text: email attachments, images embedded in a page as data URIs, keys and certificates in PEM files, JWT tokens.

The mechanism is simple: take three bytes, that is 24 bits, and split them into four groups of six. Six bits range from 0 to 63, and each value maps to a character. When one or two bytes are left over at the end, they are padded with zeros and the = sign marks how many characters carry no data. That is why a Base64 string is always about a third longer than the original data.

Base64 is not encryption: anyone can decode it without a key. It converts, it does not protect. The URL-safe variant replaces + and / with - and _, because the first two mean something in web addresses, and often drops the padding.

Common mistakes

  • Treating Base64 as a way to hide a password: it is an encoding anyone can reverse and offers no security at all.
  • Decoding to text data that is not text: an image or a key in Base64 gives bytes that are not valid UTF-8, and that is expected.
  • Mixing the two alphabets: a URL-safe string with - and _ given to a standard decoder is rejected, or misread if the decoder skips unknown characters.

Frequently asked questions

Why does a Base64 string end with one or two = signs?

Because the data was not a multiple of three bytes. One byte left over gives two characters and two =, two bytes give three characters and one =. The padding tells the decoder how many bytes to rebuild; many decoders, this one included, accept it missing.

How much longer does a file get in Base64?

Every three bytes become four characters, so about 33% more, rounded up to a multiple of four. A 30 kB image becomes a string of about 40 kB.

What does the big-endian integer mean?

The bytes are read as one number, with the first byte as the most significant digit in base 256. The bytes 01 00 are worth 256. It is the same order used in network protocols and in the large numbers of cryptography.

Why does the decoded text show odd characters or none at all?

Because the bytes are not UTF-8 text: they may be an image, a compressed file or text in another encoding such as Latin-1. In that case the hexadecimal is what counts, since it shows the bytes as they are.

How this calculation works

Each group of three bytes b₁ b₂ b₃ forms the number N = b₁·65536 + b₂·256 + b₃. The four characters are the positions ⌊N / 262144⌋, ⌊N / 4096⌋ mod 64, ⌊N / 64⌋ mod 64 and N mod 64 in the alphabet ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/. Decoding runs the other way. The integer is the sum of the bytes weighted by 256 raised to their position counted from the right.