Base64 is the duct tape of data transport: it lets arbitrary bytes travel through systems that were built for plain text. It is also the most misused encoding on the web, mostly by people who mistake it for a lock. Here is what it actually does, what it costs, and when to reach for something else.
How it works
Base64 reads input three bytes at a time. Three bytes are 24 bits, and 24 bits split evenly into four groups of six. Each 6-bit group indexes into a 64-character alphabet: A to Z, a to z, 0 to 9, then + and /. So every 3 bytes of input become exactly 4 characters of output:
Hi! = 01001000 01101001 00100001
= 010010 000110 100100 100001
= S G k h → "SGkh"When the input does not divide into threes, the output is padded with = so its length stays a multiple of four. That is the whole trick; there is no key, no secret, and nothing to crack.
What it is for
Text-only channels. Email attachments travel as Base64 because SMTP was designed for 7-bit text. Small images become data: URIs inside CSS and HTML. Binary values like keys, hashes, and signatures get embedded in JSON and XML, which cannot hold raw bytes. And JWTs use the URL-safe variant, which swaps + and / for - and _ so tokens survive URLs and cookies without extra escaping.
What it is not for
Protection. Base64 is an encoding, which means it is a reversible, keyless transformation; decoding it requires no skill and no time. A password stored as Base64 is a password stored in plain text with extra steps. If data must be unreadable to others, it needs encryption; if it must be tamper-evident, it needs a signature or an HMAC. Base64 often carries the output of those operations, which is exactly why the two get conflated.
The cost
Size. Three bytes in, four characters out means output is one third larger than the input, before padding. That is negligible for a 200-byte token and painful for images: a 300 KB photo inlined as a data URI becomes 400 KB that cannot be cached separately, cannot be lazy-loaded sensibly, and bloats the document that embeds it. The rule of thumb: inline tiny assets, an icon or a 1-pixel placeholder, and serve anything bigger as a normal file.
The gotchas that fill bug reports
- Padding mismatches: some producers omit the
=padding, some parsers demand it. A tolerant decoder re-pads before decoding. - Line wrapping: the email flavor inserts a line break every 76 characters. Whitespace inside Base64 is noise and should be stripped before decoding.
- Variant confusion: a standard decoder chokes on
-and_from the URL-safe alphabet. Good tooling accepts both. - Binary output: decoding valid Base64 can yield bytes that are not text at all, an image or compressed data. If your decoded text looks like static, the encoding was fine and the content was never text to begin with.