Description
sanitize-html re-emits the text content of raw-text elements (textarea, xmp, style, script, option) without HTML-escaping it, on the assumption that a browser will always re-read that content as plain text. That assumption is false inside SVG and MathML foreign content. sanitize-html parses with htmlparser2, which marks an element as raw-text purely by its tag name and ignores the namespace. A real browser's HTML5 parser does the opposite: inside an <svg> or <math> subtree, <textarea> and <xmp> are ordinary foreign elements, not raw-text, so their contents are parsed as live markup. An attacker places active markup inside a raw-text element nested in an SVG or MathML root. sanitize-html treats that markup as inert text belonging to the raw-text element, never sees the inner tag or its event-handler attribute, and passes it through unescaped. When the sanitized output is rendered, the browser parses the inner tag as a real element whose event handler runs. I confirmed arbitrary JavaScript execution in Chrome, in both the server-side-rendering case and the client-side element.innerHTML case. No allowVulnerableTags warning is produced, because none of svg, math, textarea, or xmp is on sanitize-html's internal vulnerableTags list, and the project documentation presents allowing SVG as supported usage and never flags textarea or xmp as dangerous.
Root Cause
In index.js, the ontext handler emits the content of any tag in nonTextTagsArray verbatim, with no escaping:
// nonTextTagsArray = ['script', 'style', 'textarea', 'option', 'xmp']
} else if ((options.disallowedTagsMode === 'discard' || options.disallowedTagsMode === 'completelyDiscard') && (nonTextTagsArray.indexOf(tag) !== -1)) {
// htmlparser2 does not decode entities inside raw text elements like
// textarea and option. The text is already properly encoded, so pass
// it through without additional escaping to avoid double-encoding.
result += text;
}
The comment states the security assumption directly: the text is treated as "already properly encoded" because htmlparser2 saw it as raw text. That holds only in the HTML namespace. When the same textarea or xmp element sits inside an <svg> or <math> root, the browser does not treat it as raw text, so the "encoded text" is re-interpreted as executable markup. htmlparser2 itself never switches namespaces for the raw-text decision (its Tokenizer selects raw-text mode by tag name in stateBeforeSpecialS / stateBeforeSpecialT), which is the source of the differential.
Proof of Concept
Minimal reproduction showing the allowlist is bypassed. The configuration allows only svg and textarea, and allows no attributes at all, so neither the img element nor the onerror attribute is permitted:
const sh = require('sanitize-html');
const out = sh('<svg><textarea><img src=x onerror=alert(1)></textarea></svg>',
{ allowedTags: ['svg', 'textarea'], allowedAttributes: {} });
console.log(out);
// => <svg><textarea><img src=x onerror=alert(1)></textarea></svg>
The sanitized output contains an intact <img> element carrying an onerror handler that the configuration never allowed. Rendering that output in any HTML5 browser runs the handler. I verified this end to end by embedding the sanitized string in a normal page and loading it in headless Chrome: the img fails to load, the onerror fires, and attacker JavaScript runs. In the run I captured, the handler set document.title to "XSS-fired" and called alert(document.domain), both of which executed. During all of this sanitize-html printed no allowVulnerableTags warning.
Other payloads I confirmed executing, each producing no warning:
<math><textarea><img src=x onerror=alert(1)></textarea></math>
<svg><xmp><img src=x onerror=alert(1)></xmp></svg>
<math><xmp><img src=x onerror=alert(1)></xmp></math>
<svg><textarea><svg onload=alert(1)></textarea></svg>
Negative controls that sanitize-html handles correctly, which isolate the exact cause as the raw-text-in-foreign-content confusion: an img placed after </textarea> inside svg has its onerror stripped as normal; <option> (which htmlparser2 does not treat as raw-text) has its onerror stripped; and a textarea placed inside an HTML integration point such as <mtext> or <foreignObject> has its content correctly HTML-escaped, because the browser and htmlparser2 agree on the namespace there.
Impact
An attacker who can submit HTML that an application sanitizes with sanitize-html can bypass the allowlist and run JavaScript in the origin of any user who views that content. This yields session cookie theft, actions performed as the victim, CSRF-token exfiltration, and account takeover, in the standard stored-XSS manner. The precondition is a configuration whose allowedTags include an SVG or MathML root (svg or math) together with a raw-text container (textarea or xmp). This is not the default configuration, and applications on defaults are not affected. It is, however, a configuration a developer can reach while believing it is safe: the documentation shows allowing svg as a supported scenario and does not warn against textarea or xmp, and the library emits no warning for this combination. The bug is a bypass of the library's core guarantee that markup outside the allowlist, and all event-handler attributes, are removed.
Description
sanitize-html re-emits the text content of raw-text elements (textarea, xmp, style, script, option) without HTML-escaping it, on the assumption that a browser will always re-read that content as plain text. That assumption is false inside SVG and MathML foreign content. sanitize-html parses with htmlparser2, which marks an element as raw-text purely by its tag name and ignores the namespace. A real browser's HTML5 parser does the opposite: inside an
<svg>or<math>subtree,<textarea>and<xmp>are ordinary foreign elements, not raw-text, so their contents are parsed as live markup. An attacker places active markup inside a raw-text element nested in an SVG or MathML root. sanitize-html treats that markup as inert text belonging to the raw-text element, never sees the inner tag or its event-handler attribute, and passes it through unescaped. When the sanitized output is rendered, the browser parses the inner tag as a real element whose event handler runs. I confirmed arbitrary JavaScript execution in Chrome, in both the server-side-rendering case and the client-sideelement.innerHTMLcase. No allowVulnerableTags warning is produced, because none of svg, math, textarea, or xmp is on sanitize-html's internal vulnerableTags list, and the project documentation presents allowing SVG as supported usage and never flags textarea or xmp as dangerous.Root Cause
In
index.js, the ontext handler emits the content of any tag innonTextTagsArrayverbatim, with no escaping:The comment states the security assumption directly: the text is treated as "already properly encoded" because htmlparser2 saw it as raw text. That holds only in the HTML namespace. When the same textarea or xmp element sits inside an
<svg>or<math>root, the browser does not treat it as raw text, so the "encoded text" is re-interpreted as executable markup. htmlparser2 itself never switches namespaces for the raw-text decision (its Tokenizer selects raw-text mode by tag name instateBeforeSpecialS/stateBeforeSpecialT), which is the source of the differential.Proof of Concept
Minimal reproduction showing the allowlist is bypassed. The configuration allows only svg and textarea, and allows no attributes at all, so neither the img element nor the onerror attribute is permitted:
The sanitized output contains an intact
<img>element carrying an onerror handler that the configuration never allowed. Rendering that output in any HTML5 browser runs the handler. I verified this end to end by embedding the sanitized string in a normal page and loading it in headless Chrome: the img fails to load, the onerror fires, and attacker JavaScript runs. In the run I captured, the handler set document.title to "XSS-fired" and called alert(document.domain), both of which executed. During all of this sanitize-html printed no allowVulnerableTags warning.Other payloads I confirmed executing, each producing no warning:
Negative controls that sanitize-html handles correctly, which isolate the exact cause as the raw-text-in-foreign-content confusion: an img placed after
</textarea>inside svg has its onerror stripped as normal;<option>(which htmlparser2 does not treat as raw-text) has its onerror stripped; and a textarea placed inside an HTML integration point such as<mtext>or<foreignObject>has its content correctly HTML-escaped, because the browser and htmlparser2 agree on the namespace there.Impact
An attacker who can submit HTML that an application sanitizes with sanitize-html can bypass the allowlist and run JavaScript in the origin of any user who views that content. This yields session cookie theft, actions performed as the victim, CSRF-token exfiltration, and account takeover, in the standard stored-XSS manner. The precondition is a configuration whose allowedTags include an SVG or MathML root (svg or math) together with a raw-text container (textarea or xmp). This is not the default configuration, and applications on defaults are not affected. It is, however, a configuration a developer can reach while believing it is safe: the documentation shows allowing svg as a supported scenario and does not warn against textarea or xmp, and the library emits no warning for this combination. The bug is a bypass of the library's core guarantee that markup outside the allowlist, and all event-handler attributes, are removed.