Every hash function maps input to a fixed-size output. A cryptographic one adds three guarantees: it is infeasible to find an input for a given hash (preimage resistance), to find a second input with the same hash as a known one, or to find any two inputs that collide. Two properties are visible right away:
- Deterministic and fixed length. The same bytes always give the same digest. SHA-256 is always 256 bits (64 hex characters), whether you hash one byte or a 4 GB disk image.
- Avalanche effect. Changing a single bit of the input flips about half of the output bits, so similar inputs give unrelated-looking hashes:
SHA-256("hello world") = b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9
SHA-256("hello world!") = 7509e5bda0c762d2bac7f90d758b5b2263fa01ccbc542ab5e3df163be08e6ca9
The flip side is that invisible differences matter. The most common "wrong hash" report is
a trailing newline: echo hello | sha256sum hashes hello\n, not
hello. Use printf '%s' hello or echo -n. Windows line
endings (\r\n), a byte-order mark, or a different text encoding change the
hash in the same way. The byte count next to the input tells you exactly how many bytes
were hashed.