Text encoding converter
Drop a file to sniff and decode; encode text back into GBK and other legacy encodings — local, honest about loss.
Opening an old file to a wall of "é" characters, or feeding a system that only accepts GBK — encoding is the classic CJK trap. The essence: one byte sequence interpreted under different rules. UTF-8's three 3-byte characters read two-by-two as GBK turn into different characters entirely. Repair has exactly two directions: decode bytes with the right encoding, or re-encode text into the target one.
This tool does both. Drop a file on the decode side and it sniffs candidates ranked by lossless validation plus CJK character share — switch with one click, and replacement characters are reported plainly as "this file is not that encoding". The encode direction covers what browsers never offered natively (TextEncoder speaks only UTF-8): the tool reverse-maps every valid byte sequence of the target encoding into a character table, supporting GBK, Big5, Shift_JIS, EUC-KR and more, listing unmappable characters honestly and exporting hex, percent-encoding or a byte file.
Sniffing is guessing — confirm the candidate yourself
No detector is certain: the same bytes are often valid in both GBK and Big5. Ranking here uses two hard signals — a strict decode with zero replacement characters (lossless), and a CJK character score. Every √-marked candidate decodes without loss, but lossless does not mean correct: a Japanese file can decode "cleanly" into nonsense hanzi. Your eyes make the final call.
How the encode direction works
Browsers only ship byte→text decoding; text→legacy-bytes has no API. The trick here: decode every legal byte sequence of the target encoding (256×256 pairs for double-byte encodings) once, register those that yield real characters into a reverse table — tens of milliseconds on first use, resident afterwards. No data files, and by construction anything it encodes decodes back identically.
Frequently asked questions
- Why can "é"-style mojibake be repaired?
- That is UTF-8 text displayed as GBK. Drop the original file (never copy-paste the mojibake characters off the screen) and decode as UTF-8 to recover the text. The rule: repair always operates on the original bytes — copying displayed mojibake has already destroyed them.
- How do gb18030 and GBK relate?
- GB2312 ⊂ GBK ⊂ GB18030. GBK covers twenty-odd thousand hanzi in two bytes; GB18030 is the mandatory national standard that adds a four-byte range covering all of Unicode. Browsers route the "gbk" label through the GB18030 decoder (a superset, backward compatible) — for everyday use they are equivalent.
- Which encodings are supported? Are files uploaded?
- Decoding covers UTF-8, UTF-16, GB18030/GBK, Big5, Shift_JIS, EUC-KR, Windows-1252/1251 and KOI8-R; the encode direction supports the same set including UTF-16 and double-byte CJK. Everything runs locally — files never leave the device.