The HTML parser treats less-than signs and ampersands as syntax unless character references represent them as text.
HTML character references: escape data before inserting it into markup
Use it for a real task
A receipt label is untrusted display data. Its characters belong in a text node, never as a string concatenated into an HTML template. The literal lesson shows how markup text must be represented; application code should use a renderer that escapes data for the exact output context.
<p>Customer note: 47 < 73 & review remains open.</p>
<p>Attribute value: "pending"</p>What the markup guarantees
Escaping for a text node and escaping for a quoted attribute are related but distinct contexts. Do not turn already-escaped data into HTML again or it will show entity text; do not decode received text and then inject it as markup.
Cost and limits
Entity parsing is linear in the input length. The larger risk is a trust-boundary bug: treating user text as HTML can create elements or attributes the author never intended.
Common Mistakes
- Do not build HTML by joining untrusted strings.
- Do not assume URL encoding is HTML escaping.
- Escape at the output boundary rather than mutating the stored value.
Connected lessons
- HTML anchors: use href for navigation and stable fragment targets
- HTML forms: choose GET for retrieval and POST for a state change
- JavaScript Tutorial: Core Concepts
Related: HTML iframe sandbox: isolate an embedded preview and name its purpose.
Related: HTML lang, dir, and bdi: keep text pronunciation and direction stable.
Related: HTML hidden versus aria-hidden: rendering and accessibility are different scopes.
Continue with HTML pre and code: preserve source layout without hiding the meaning.
Continue with HTML mark: identify relevant text without calling it a warning.
Continue with HTML kbd, samp, and var: separate input, output, and variables.
