HTML entity escape

Escape and unescape HTML entities in three modes; unknown entities pass through instead of being dropped.

Getting text safely into HTML comes down to three habits: the five characters & < > " ' must always be escaped or you risk broken markup at best and XSS at worst; characters like ©, — and CJK can use named entities (&copy;) or numeric ones (&#xa9;); and a text full of &amp;lt; needs reliable unescaping to show the original.

This tool splits escaping into the three real-world modes: "required only" handles the five characters the spec mandates — the browser-parser baseline, nothing more; "named entities" also rewrites self-explanatory symbols into readable forms; "numeric entities" turns every non-ASCII character into &#xHH; that survives any legacy channel. Decoding understands all three notations and keeps unknown entities as-is — silently deleting something you have never seen is worse than leaving it visible.

Which mode when

Embedding dynamic text into attributes or body content: required-only — fewer escapes keep the source readable with zero safety loss, since those five characters are the entire injection story. Handing source to humans or email templates: named entities. Moving data across systems (CMS APIs, legacy encodings, ASCII-only channels): numeric.

Escaping and XSS

XSS happens when data is mistaken for code. Escaping < turns a would-be <script> into harmless text — it is the cornerstone of output-side injection defense. Know the boundary though: this tool gets text safely "into" HTML; complete defense also needs the right escaping per context (JavaScript strings and URL parameters have their own rules) and judgment about where output lands.

Frequently asked questions

What is the difference between &amp;nbsp; and a normal space?
&amp;nbsp; is a no-break space (U+00A0): HTML collapses runs of normal spaces into one, but nbsp neither collapses nor allows line breaks at that point — it is also the classic way to indent line starts. Decoding here restores it to the real U+00A0 character.
Why were some entities not decoded?
HTML defines thousands of named entities; this tool covers the two hundred or so in common use. Anything outside the table (including misspellings) is left untouched and remains visible in the result — better than silently emitting wrong content. Numeric entities (&#169; / &#xa9;) are always supported.
How is this different from JSON escaping?
Different escaping systems entirely: HTML entities resolve the conflict between characters and markup (&lt; means a literal <), while JSON escapes resolve the conflict between string content and syntax (a backslash-quote means a literal quote). Putting HTML inside a JSON string requires both, HTML rules first. There is a dedicated JSON escape tool on this site.

Related tools