On the wire, a URL is pure ASCII: a small set of characters standardized in 1963, with no room for é, ß, 東, or emoji. Yet URLs carry all of those every day. The trick is that non-ASCII text is first converted to bytes by a character set, and those bytes are then percent-encoded. Which character set does the converting is the detail that makes or breaks international URLs.
One letter, two encodings
Take the letter é. In ISO-8859-1 (Latin-1), the dominant Western encoding of the 1990s, é is the single byte 0xE9, so it percent-encodes as %E9. In UTF-8, the encoding of the modern web, é is the two-byte sequence 0xC3 0xA9, which encodes as %C3%A9. Both are "é, URL-encoded", and they are completely incompatible:
ISO-8859-1: caf%E9
UTF-8: caf%C3%A9A decoder must know which convention produced the bytes. Feed %E9 to a strict UTF-8 decoder and it fails, because a lone 0xE9 byte is not valid UTF-8. Feed %C3%A9 to a Latin-1 decoder and you get two characters instead of one.
Mojibake, explained in one sentence
The garbled text called mojibake, where é appears as é, is UTF-8 bytes being read as if they were Latin-1: the two bytes 0xC3 0xA9 are rendered as the two Latin-1 characters à and © instead of being combined into one UTF-8 character. The reverse mistake usually shows up as question marks or replacement characters, because Latin-1 bytes rarely form valid UTF-8. Once you can recognize the two failure shapes, you can tell from the wreckage which side of the conversation was wrong.
Where Windows-1252 sneaks in
Windows-1252 is Latin-1's near-identical sibling, with one difference that matters: the 32 byte values that Latin-1 reserves for control codes are assigned to printable characters, including the euro sign, the trademark sign, and the curly quotes that word processors insert automatically. Text copied from a document into a form is a classic source of these. If a URL decodes almost correctly except for quotes, dashes, or a euro sign, decode it as Windows-1252 and it will usually snap into place.
The modern rule
The WHATWG URL standard, which browsers implement, settled the question: new systems percent-encode using UTF-8, full stop. JavaScript's encodeURIComponent, every modern framework, and every browser address bar produce UTF-8 percent-encoding. The legacy character sets survive only at the edges, in old backends, exported spreadsheets, decades-old feeds, and email systems, which is precisely where a converter that can speak them earns its keep.
Practical diagnosis
%C3,%C2, or%E2sequences in a URL are a strong sign of UTF-8.- Single escapes in the
%E0to%FFrange that fail strict UTF-8 decoding usually mean Latin-1 or Windows-1252. - Ã followed by another odd character means UTF-8 was read as Latin-1 somewhere upstream; fix the reader, not the data.
The encoder on this site has a character set selector for exactly these situations: decode the same string as UTF-8, ISO-8859-1, and Windows-1252, and let the readable one identify the culprit.