HTML entity escape
Escape and unescape HTML entities in three modes; unknown entities pass through instead of being dropped.
Getting text safely into HTML comes down to three habits: the five characters & < > " ' must always be escaped or you risk broken markup at best and XSS at worst; characters like ©, — and CJK can use named entities (©) or numeric ones (©); and a text full of &lt; needs reliable unescaping to show the original.
This tool splits escaping into the three real-world modes: "required only" handles the five characters the spec mandates — the browser-parser baseline, nothing more; "named entities" also rewrites self-explanatory symbols into readable forms; "numeric entities" turns every non-ASCII character into &#xHH; that survives any legacy channel. Decoding understands all three notations and keeps unknown entities as-is — silently deleting something you have never seen is worse than leaving it visible.
Which mode when
Embedding dynamic text into attributes or body content: required-only — fewer escapes keep the source readable with zero safety loss, since those five characters are the entire injection story. Handing source to humans or email templates: named entities. Moving data across systems (CMS APIs, legacy encodings, ASCII-only channels): numeric.
Escaping and XSS
XSS happens when data is mistaken for code. Escaping < turns a would-be <script> into harmless text — it is the cornerstone of output-side injection defense. Know the boundary though: this tool gets text safely "into" HTML; complete defense also needs the right escaping per context (JavaScript strings and URL parameters have their own rules) and judgment about where output lands.
Frequently asked questions
- What is the difference between &nbsp; and a normal space?
- &nbsp; is a no-break space (U+00A0): HTML collapses runs of normal spaces into one, but nbsp neither collapses nor allows line breaks at that point — it is also the classic way to indent line starts. Decoding here restores it to the real U+00A0 character.
- Why were some entities not decoded?
- HTML defines thousands of named entities; this tool covers the two hundred or so in common use. Anything outside the table (including misspellings) is left untouched and remains visible in the result — better than silently emitting wrong content. Numeric entities (© / ©) are always supported.
- How is this different from JSON escaping?
- Different escaping systems entirely: HTML entities resolve the conflict between characters and markup (< means a literal <), while JSON escapes resolve the conflict between string content and syntax (a backslash-quote means a literal quote). Putting HTML inside a JSON string requires both, HTML rules first. There is a dedicated JSON escape tool on this site.