Unicode Character Lookup
Character ↔ code point with UTF-8/UTF-16/HTML/URL/JS forms, plus block browsing.
Debugging a mojibake problem starts with the character's full encoding dossier: is the code point U+4E2D or U+1F600, how many UTF-8 bytes, is UTF-16 a surrogate pair, what HTML entity, what does it become in a URL. This tool answers with one input — a character, or a code point in reverse — and shows all seven representations.
The other daily need is "find that special character": arrows, math operators, box drawing, emoji. Browsing by block beats digging through an input method. Twelve common blocks ship built in (Latin-1, arrows, math, box drawing, emoji, everyday CJK), and clicking any character loads its details.
UTF-8 and UTF-16: one character set, two encodings
UTF-8 is variable-length 1-4 bytes — ASCII single-byte, CJK three bytes, emoji four — the default for files and the internet. UTF-16 uses 2 or 4 bytes (2 inside the BMP) and lives inside JavaScript strings and Windows APIs. Code points beyond the BMP — most famously emoji — become surrogate pairs in UTF-16, which is why string.length returns 2 for one emoji.
Seven spellings of the same character
The HTML entity 😀 for source code where a raw character is unwelcome; %F0%9F%98%80 in URLs and logs; \u{1F600} in JavaScript. When a log shows %F0%9F, paste the fragment here and meet the character it really was.
Frequently asked questions
- Does it show the character's official name?
- Not yet — the Unicode name database is too heavy for a browser tool's loading model. With the code point in hand, any Unicode reference site gives you the formal name.
- Why do some characters render as boxes?
- Your system font lacks the glyph — an encoding-independent rendering gap. The code point exists; coverage varies by OS and installed fonts.
- How is this different from the JSON Unicode tool?
- This page is one character's complete dossier; the JSON tool family's Unicode page bulk-converts whole texts between characters and \uXXXX escapes. The two interlink.