When you type a ñ on your phone, the keyboard does not store the letter: it stores a number. And that number does not necessarily fit in a single byte (a unit of 8 bits, the smallest block a computer uses to represent data). For every language on Earth to coexist in the same memory — from Spanish to Chinese, from Arabic to Devanagari — a global pact is needed: it is called Unicode.
The problem of 256 characters
The earliest encodings, such as ASCII (American Standard Code for Information Interchange), assigned each character a 7-bit value: just 128 slots, only Latin uppercase and lowercase letters, digits and basic symbols. With 8 bits (256 values), Latin-1 added accented vowels and the ñ, U+00F1. But every country used its own table: byte 0xE9 meant “é” in one and “Ù” in another. The same file read differently depending on the operating system’s language. That chaos is called mojibake: text turned into gibberish when decoded with the wrong encoding.
Unicode: one number for every character in the world
Unicode solves this by assigning every character — also emojis, punctuation and technical symbols — a single code point: a number written in hexadecimal with the U+ prefix, such as U+00F1 for the ñ. The repertoire is organized into planes: the first one, the Basic Multilingual Plane (BMP, from 0 to U+FFFF), holds 65,536 positions and contains nearly all modern alphabets. The remaining planes, up to U+10FFFF, provide more than a million slots for historical scripts, unified CJK ideographs and emojis.
Unicode is the dictionary, not the representation. The dictionary says “code point U+00F1 is a ñ”, but it does not say how to turn that number into bytes. That is decided by the encodings: UTF-8, UTF-16 and UTF-32.
UTF-8: the byte that adapts
UTF-8 (Unicode Transformation Format, 8-bit) is the most widespread encoding on the web. Its genius is variable length: it uses 1 to 4 bytes depending on the code point.
For characters from 0 to 127 (all of ASCII) it uses a single byte, 0xxxxxxx, identical to ASCII. So any old English text is, unchanged, valid UTF-8; that is why compatibility was total. For the following ranges, prefixes mark how many bytes follow:
- 2 bytes (00–07FF, like the ñ U+00F1): pattern
110xxxxx 10xxxxxx. The ñ is encoded asC3 B1. - 3 bytes (0800–FFFF, nearly the whole BMP): pattern
1110xxxx 10xxxxxx 10xxxxxx. - 4 bytes (10000–10FFFF, emojis and supplementary planes):
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx.
The leading bit of each byte reveals everything: if it starts with 0, it is a single-byte character; if it starts with 110/1110/11110, it is the start of a 2/3/4-byte one; and inner bytes always begin with 10. This property makes UTF-8 self-synchronizing: if a byte is lost or corrupted, the damage stays local and does not destroy the rest of the text.
UTF-16 and the enigma of emojis
UTF-16 uses 16 bits (2 bytes) per character. Since a single 16-bit value only reaches U+FFFF, characters above the BMP (emojis, for instance) are represented with surrogate pairs: two 16-bit units that together encode a code point between U+10000 and U+10FFFF. So, technically, an emoji “occupies” two 16-bit characters: a detail that broke more than one length counter.
UTF-16 also suffers from endianness: is the high byte written first or the low one? To resolve this there is the byte order mark (BOM, an invisible character at the start), whereas in UTF-8 the BOM is optional and causes no ambiguity.
More than meets the eye: normalization
Unicode complicates programmers’ lives with equivalent representations. The é can exist as a precomposed character (U+00E9) or as the combination of an e plus a combining accent (U+0065 + U+0301). Both look identical but have different bytes. To compare, index or search correctly you must normalize the text: the NFC (compose) and NFD (decompose) forms transform strings so two visually identical texts are also byte-identical. Ignoring this causes hard-to-diagnose search, sorting and duplicate errors.
Why it matters
Multi-encoding is not a curiosity: today your browser, your database and your API negotiate the character set in every HTTP request through the header Content-Type: text/html; charset=utf-8. Understanding how characters are encoded explains why an Arabic keyboard, your ñ and an emoji can coexist in the same file — and why that coexistence was, for decades, computing’s biggest headache. Unicode was not a technical decision: it was a peace treaty for every alphabet on the planet.





