Base64: What It Is For, and the 33% You Pay

Not encryption, not compression. A way to move binary through channels that only accept text — at a measurable cost.

Base64 represents binary data using 64 printable ASCII characters: A–Z, a–z, 0–9, + and /, with = as padding. It exists because a lot of infrastructure — email bodies, JSON strings, URLs, XML — was designed for text and mangles or rejects arbitrary bytes.

The mechanism

Take three bytes: 24 bits. Split into four groups of six. Each 6-bit group indexes the alphabet, giving four characters. Three bytes in, four characters out — which is the whole story, including the cost.

Worked on the word Man: bytes 77, 97, 110 are 010011010110000101101110. Regrouped into six-bit pieces: 010011, 010110, 000101, 101110 = 19, 22, 5, 46 = TWFu.

When the input is not a multiple of three, the last group is padded and = marks how many bytes were real. One leftover byte produces two characters plus ==; two leftover bytes produce three characters plus =.

The 33% overhead is not negotiable

Four characters for every three bytes is a 1.333× expansion, before any encoding of the result. A 3 MB image becomes 4 MB of base64. Inlined into JSON and served without compression, that is 1 MB of extra transfer for no functional gain.

This is the argument against inlining images as data URIs by default. It is sometimes still right — a handful of tiny icons avoid separate requests — but for anything sizeable, a normal URL beats base64 comfortably. Gzip recovers some of the overhead, though base64 compresses worse than the original binary did, so you do not get it all back.

It provides no security whatsoever

Base64 is a public, reversible encoding with no key. Decoding is one function call. Credentials "hidden" in base64 are in plain text wearing a hat — which is precisely what HTTP Basic authentication does, and why it is only acceptable over TLS.

The URL-safe variant

Standard base64 uses + and /, both of which mean something in a URL: + can decode as a space in a query string and / is a path separator. The URL-safe alphabet (RFC 4648 ยง5) substitutes - and _, and usually drops the padding.

This is the single most common source of "it works in testing and fails in production" with base64 — a token encoded one way and decoded the other. JWTs use the URL-safe form without padding, which is why pasting a JWT segment into a standard decoder sometimes fails on padding.

Line breaks

MIME (RFC 2045) mandates a line break every 76 characters. Strict decoders that do not expect them reject the input; lenient ones ignore them. When base64 from an email context fails to decode elsewhere, embedded newlines are the first thing to check.

When to use it

Encode and decode → · Image to base64 → · Back to all articles