Base64 shows up everywhere: in data URIs, HTTP Authorization headers, email attachments,
JSON Web Tokens and Kubernetes secrets. It solves a narrow problem well: carrying arbitrary bytes
through channels that only handle text. It's also widely misunderstood, most often by being treated as a
way to hide data.
This article explains exactly how the encoding works at the bit level, why the output is a third
larger than the input, how Base64URL differs, the Unicode bug in JavaScript's btoa, and how
to use Base64 correctly from the shell, Python and Java. The definitive reference is RFC 4648.
The alphabet
Base64 represents data using 64 characters: A–Z (values 0–25),
a–z (26–51), 0–9 (52–61), + (62) and
/ (63), plus = for padding. These characters were chosen because they pass
unchanged through practically every text system, including old mail gateways, terminals and XML
documents. 64 is 26, so each character carries exactly 6 bits.
How 3 bytes become 4 characters
Bytes have 8 bits and Base64 characters carry 6. The lowest common multiple is 24 bits, so the
encoder takes 3 bytes at a time and splits them into four 6-bit groups. Take the ASCII string
Man:
| Character | M | a | n |
|---|---|---|---|
| Byte value | 77 | 97 | 110 |
| Bits | 01001101 | 01100001 | 01101110 |
Concatenate the 24 bits, 010011010110000101101110, and cut them into four groups of six:
| 6-bit group | 010011 | 010110 | 000101 | 101110 |
|---|---|---|---|---|
| Value | 19 | 22 | 5 | 46 |
| Character | T | W | F | u |
So Man encodes to TWFu. Decoding reverses the steps: look up each
character's 6-bit value, concatenate, and read the bits back as bytes. Nothing about the process depends
on the data being text, so the same steps work for images, ZIP files or encrypted ciphertext.
Padding and the 33% overhead
Input lengths that aren't a multiple of 3 leave one or two bytes at the end. The encoder fills the
last 6-bit group with zero bits and marks the missing bytes with =:
- Two bytes left (16 bits):
Ma→010011 010110 000100→TWE=. Three characters plus one=. - One byte left (8 bits):
M→010011 010000→TQ==. Two characters plus two=.
Encoded output is therefore always a multiple of 4 characters when padded: 4 × ceil(n / 3)
for n input bytes. That is 4 bytes out for every 3 in, an overhead of about 33%:
| Input bytes | Base64 characters |
|---|---|
| 1, 2 or 3 | 4 |
| 10 | 16 |
| 100 | 136 |
| 1,000 | 1,336 |
| 1,048,576 (1 MiB) | 1,398,104 |
MIME email adds a line break every 76 characters, which adds a little more. The overhead is why inlining large images as data URIs, or sending large files as Base64 strings inside JSON, is usually a bad trade: the payload grows by a third and has to be decoded before use.
Base64 vs Base64URL
Two characters of the standard alphabet cause trouble in URLs: + means a space in
form-encoded query strings, and / is a path separator. The = padding clashes
with key=value syntax. RFC 4648 section 5 defines the "URL and filename safe" variant,
usually called Base64URL, which swaps + for - and / for
_. Padding is often omitted, and JWTs require it to be omitted.
The bytes FB FF BF show the difference, since they hit both characters:
standard: +/+/
base64url: -_-_
The two variants aren't interchangeable. A standard decoder rejects or misreads - and
_, and a token with its padding stripped needs the padding restored before a strict
decoder will accept it. If you put standard Base64 in a URL, percent-encode it. Our URL encoding guide explains why.
Where you will meet Base64
- Data URIs:
data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciLz4=embeds a file directly in HTML or CSS. For SVG, URL-encoded text is often smaller than Base64. - HTTP Basic authentication (RFC 7617):
alice:s3cretis sent asAuthorization: Basic YWxpY2U6czNjcmV0. Anyone who sees the header can decode it, which is why Basic auth is only acceptable over HTTPS. - Email (MIME): attachments and non-ASCII bodies use
Content-Transfer-Encoding: base64, with lines of at most 76 characters (RFC 2045). - JSON Web Tokens: the header and payload are Base64URL-encoded JSON. Our JWT guide decodes one in detail.
- Binary data in JSON: JSON has no byte type, so APIs send files and hashes as
Base64 strings, for example
{"file":"report.pdf","data":"JVBERi0xLjc="}. - PEM files: certificates and keys are Base64 DER data wrapped at 64 characters
between
-----BEGIN ...-----lines.
JavaScript: the btoa Unicode trap
The browser's btoa() and atob() predate TextEncoder and work on
"binary strings", where each character must have a code point from 0 to 255. Pass anything outside that
range and they throw. In Chrome:
btoa('✓ done')
// Uncaught InvalidCharacterError: Failed to execute 'btoa' on 'Window':
// The string to be encoded contains characters outside of the Latin1 range.
Node.js throws InvalidCharacterError: Invalid character. The quieter bug is worse:
btoa('café') succeeds and returns Y2Fm6Q==, because é (U+00E9)
fits in Latin-1. But that encodes the single byte E9, not the UTF-8 bytes
C3 A9 that every other system expects. The correct UTF-8 encoding is
Y2Fmw6k=, and a server decoding Y2Fm6Q== as UTF-8 gets an invalid
byte.
The fix is to convert the string to UTF-8 bytes first:
function toBase64(str) {
const bytes = new TextEncoder().encode(str); // UTF-8 bytes
let binary = '';
for (const b of bytes) binary += String.fromCharCode(b);
return btoa(binary);
}
function fromBase64(b64) {
const binary = atob(b64);
const bytes = Uint8Array.from(binary, (c) => c.charCodeAt(0));
return new TextDecoder().decode(bytes); // back to a string
}
toBase64('✓ café 🚀'); // "4pyTIGNhZsOpIPCfmoA="
fromBase64('4pyTIGNhZsOpIPCfmoA='); // "✓ café 🚀"
In Node.js, use Buffer, which handles UTF-8 and Base64URL directly:
Buffer.from(str).toString('base64'), .toString('base64url'), and
Buffer.from(b64, 'base64').toString('utf8'). Newer JavaScript engines also add
Uint8Array.prototype.toBase64() and Uint8Array.fromBase64(), which remove the
binary-string step entirely. Check MDN's compatibility table before relying on them; they weren't
available in the Node 24 release we tested with.
Base64 on the command line
The base64 command exists on both Linux (GNU coreutils) and macOS, but the flags differ.
First, watch out for echo, which adds a trailing newline that becomes part of the
data:
$ echo 'hello' | base64
aGVsbG8K
$ printf '%s' 'hello' | base64
aGVsbG8=
The differences that matter:
- Decoding: GNU uses
-dor--decode. Recent macOS accepts-d,-Dand--decode; older macOS releases documented only-D.--decodeworks on both. - Line wrapping: GNU wraps output at 76 characters by default, which corrupts
values you paste into a header or YAML file. Use
base64 -w 0for one line. macOS doesn't wrap unless you pass-b 76. -imeans different things: on GNU it's--ignore-garbage; on macOS it names the input file.
Neither tool understands Base64URL or missing padding. On macOS, decoding the unpadded JWT segment
eyJzdWIiOiJ1c2VyXzg0MTIifQ printed {"sub":"user_8412" and exited with status 0,
silently dropping the last byte. GNU base64 reports invalid input in that case.
This shell function converts the alphabet and restores padding first:
b64url_decode() {
local s
s=$(printf '%s' "$1" | tr '_-' '/+')
case $(( ${#s} % 4 )) in
2) s="$s==" ;;
3) s="$s=" ;;
esac
printf '%s' "$s" | base64 -d
}
b64url_decode eyJzdWIiOiJ1c2VyXzg0MTIifQ # {"sub":"user_8412"}
Base64 in Python
Python's base64 module works on bytes, so encode strings explicitly:
import base64
data = "✓ café".encode("utf-8")
encoded = base64.b64encode(data) # b'4pyTIGNhZsOp' (bytes, not str)
text = base64.b64decode(encoded).decode("utf-8")
base64.urlsafe_b64encode(bytes([0xFB, 0xFF, 0xBF])) # b'-_-_'
def b64url_decode(s: str) -> bytes:
return base64.urlsafe_b64decode(s + "=" * (-len(s) % 4))
Two behaviors to know. Unpadded input raises binascii.Error: Incorrect padding, which is
why the helper above restores the = characters. And by default
b64decode silently discards characters outside the alphabet:
b64decode("TWFu!!") returns b'Man'. Pass validate=True to get
binascii.Error: Only base64 data is allowed instead. Use it for anything that crosses a
trust boundary.
Base64 in Java
Since Java 8, java.util.Base64 provides three encoder and decoder pairs:
import java.nio.charset.StandardCharsets;
import java.util.Base64;
byte[] data = "✓ café".getBytes(StandardCharsets.UTF_8);
String std = Base64.getEncoder().encodeToString(data); // 4pyTIGNhZsOp
String url = Base64.getUrlEncoder().withoutPadding()
.encodeToString(new byte[]{(byte) 0xFB, (byte) 0xFF, (byte) 0xBF}); // -_-_
String back = new String(Base64.getDecoder().decode(std), StandardCharsets.UTF_8);
Always pass an explicit charset to getBytes and new String. The no-argument
versions use the platform default, which has caused many "works on my machine" bugs. The basic decoder
rejects line breaks (IllegalArgumentException: Illegal base64 character d for a CR), so
decode wrapped MIME or PEM content with Base64.getMimeDecoder().
Base64 is not encryption
Base64 has no key. Anyone can reverse it with a one-line command, and it's recognizable on sight; a
string ending in = or == is a strong hint. Encoding a secret only stops
someone reading it over your shoulder. Common mistakes:
- Storing "encrypted" passwords in config files that are only Base64-encoded.
- Assuming Kubernetes
Secretvalues are protected because they look scrambled. They're Base64-encoded, and encryption at rest is a separate cluster setting. - Putting sensitive data in a JWT payload because it "looks encoded".
Use the tool that fits the job:
- Passwords: a slow password hash such as Argon2id, scrypt or bcrypt, never reversible encoding and never a single fast hash.
- Data that must be read back: authenticated encryption such as AES-GCM through a vetted library (the Web Crypto API, libsodium, or your platform's crypto module), with keys kept in a secrets manager.
- Integrity checks: a cryptographic hash such as SHA-256, or an HMAC when the check must prove who created the data. The Hash Generator is handy for comparing checksums.
- Data in transit: TLS.
Encrypted data is often Base64-encoded afterwards so it can travel as text. That's the right place for Base64: as packaging around real cryptography, not instead of it.