Skip to content
AITroveRead. Build. Understand.
Make this comfortable

HTML character references: escape data before inserting it into markup

Last updated: 5 Oct 20266 min read
tutorial
IntermediateBy AITrove Editorial

The HTML parser treats less-than signs and ampersands as syntax unless character references represent them as text.

Use it for a real task

A receipt label is untrusted display data. Its characters belong in a text node, never as a string concatenated into an HTML template. The literal lesson shows how markup text must be represented; application code should use a renderer that escapes data for the exact output context.

html
<p>Customer note: 47 &lt; 73 &amp; review remains open.</p>
<p>Attribute value: &quot;pending&quot;</p>

What the markup guarantees

Escaping for a text node and escaping for a quoted attribute are related but distinct contexts. Do not turn already-escaped data into HTML again or it will show entity text; do not decode received text and then inject it as markup.

Cost and limits

Entity parsing is linear in the input length. The larger risk is a trust-boundary bug: treating user text as HTML can create elements or attributes the author never intended.

Common Mistakes

  • Do not build HTML by joining untrusted strings.
  • Do not assume URL encoding is HTML escaping.
  • Escape at the output boundary rather than mutating the stored value.

Connected lessons

Related: HTML iframe sandbox: isolate an embedded preview and name its purpose.

Related: HTML lang, dir, and bdi: keep text pronunciation and direction stable.

Related: HTML hidden versus aria-hidden: rendering and accessibility are different scopes.

Continue with HTML pre and code: preserve source layout without hiding the meaning.

Continue with HTML mark: identify relevant text without calling it a warning.

Continue with HTML kbd, samp, and var: separate input, output, and variables.

html
core
Storage details