Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 9 additions & 4 deletions .changeset/sanitize-html-svg-mathml-rawtext-xss.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,18 @@
Security: fixed an XSS/allowlist bypass in which the contents of a raw-text
element (`textarea` or `xmp`) nested inside an `svg` or `math` root were
re-emitted without HTML-escaping. `sanitize-html` treated that content as inert
raw text because `htmlparser2` classifies raw-text elements by tag name and
ignores the namespace, but a real HTML5 parser treats `textarea`/`xmp` as
raw text because `htmlparser2` 10.x classified raw-text elements by tag name and
ignored the namespace, but a real HTML5 parser treats `textarea`/`xmp` as
ordinary foreign elements inside SVG/MathML and re-parses their contents as live
markup. As a result, markup and event-handler attributes that the allowlist
never permitted (for example `<svg><textarea><img src=x onerror=alert(1)>`)
could survive sanitization and execute in the browser. Raw-text content is now
HTML-escaped whenever it appears anywhere inside an `svg` or `math` subtree. The
could survive sanitization and execute in the browser. This is now fixed on two
fronts: `htmlparser2` was upgraded to 12.x, which is namespace-aware and parses
`textarea`/`xmp` inside SVG/MathML as ordinary elements, so their
non-allowlisted children (such as the injected `img`) are dropped by the
allowlist instead of being preserved as raw text; and any raw-text content
`sanitize-html` still emits for these tags (at HTML integration points such as
`foreignObject`/`mtext`, or outside foreign content) is always HTML-escaped. The
default configuration is not affected; the precondition is an `allowedTags` that
includes `svg` or `math` together with `textarea` or `xmp`. Thanks to
[khoadb175](https://github.com/khoadb175) for responsibly disclosing the
Expand Down
5 changes: 5 additions & 0 deletions .changeset/sanitize-html-textarea-xmp-solidus-xss.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"sanitize-html": patch
---

Security: fixed a mutation-XSS / `allowedTags` bypass affecting configurations that allow the `textarea` or `xmp` raw-text tags. `htmlparser2` 10.x did not recognize an end tag with a trailing solidus (e.g. `</textarea/>`) as closing the element, so it kept the following markup as raw text, but a spec-compliant browser treats `</textarea/>` as a valid close and parses that markup as a live element. Because raw-text content was re-emitted without escaping, a payload such as `<textarea></textarea/><img src=x onerror=...>` could smuggle non-allowlisted, executable markup through the sanitizer. The default configuration was not affected. This is now defended at two layers: `htmlparser2` was upgraded to 12.x, whose tokenizer closes these end tags correctly, and the raw text sanitize-html emits for these tags is always escaped so no `<` can reopen a tag when the output is re-parsed (`textarea`, an RCDATA element whose entities `htmlparser2` decodes, is escaped like normal text, while `xmp`, a raw-text element, has only its angle brackets escaped to avoid double-encoding already-encoded entities). Because `htmlparser2` is ESM-only from version 11 onward, `sanitize-html` now requires Node.js `>=22.12.0` (the first 22.x release in which `require()` of an ES module is available unflagged). Thanks to [bibu123456](https://github.com/bibu123456) for reporting the vulnerability and [Kayiz-PT](https://github.com/Kayiz-PT) for coordinating the disclosure (GHSA-jxwj-j7wr-gfrw).
37 changes: 20 additions & 17 deletions claude-tools/run-sanitize-html-tests.sh
Original file line number Diff line number Diff line change
@@ -1,12 +1,14 @@
#!/bin/bash
# Run the sanitize-html package mocha suite and log full output to
# claude-tools/logs/sanitize-html.log. Prints a summary of failing tests.
# Run the sanitize-html mocha suite, logging full output to
# claude-tools/logs/sanitize-html.log and printing a summary of ONLY the
# specific tests that failed (so we never have to re-run from scratch to find
# out what broke). Optional first arg is a mocha --grep filter, e.g.:
#
# Usage:
# ./claude-tools/run-sanitize-html-tests.sh # whole suite
# ./claude-tools/run-sanitize-html-tests.sh "<grep string>" # filter by title
# ./claude-tools/run-sanitize-html-tests.sh 'GHSA-jxwj' # just matching tests
#
# Runs one suite at a time only. Does not run lint (use `npm test` for that).
# NEVER run test suites in parallel — they are designed to run one at a time
# and the host has limited resources.

set -u
grep_filter="${1:-}"
Expand All @@ -19,24 +21,25 @@ log="$logdir/sanitize-html.log"

cd "$root/packages/sanitize-html"

echo "=== sanitize-html tests ${grep_filter:+(grep: $grep_filter) }($(date -Is)) ===" | tee -a "$log"
echo "=== sanitize-html mocha ${grep_filter:+(grep: $grep_filter) }($(date +%Y-%m-%dT%H:%M:%S%z)) ===" | tee -a "$log"

if [[ -n "$grep_filter" ]]; then
./node_modules/.bin/mocha --reporter spec --grep "$grep_filter" >> "$log" 2>&1
./node_modules/.bin/mocha --grep "$grep_filter" >> "$log" 2>&1
else
./node_modules/.bin/mocha --reporter spec >> "$log" 2>&1
./node_modules/.bin/mocha >> "$log" 2>&1
fi
code=$?

echo "=== exit=$code ===" | tee -a "$log"

# Surface the passing/failing counts and any failing test titles.
echo "--- summary ---"
grep -E "passing|failing|pending" "$log" | tail -3
if [[ "$code" -ne 0 ]]; then
echo "--- failing tests ---"
# Mocha lists failures as a numbered block after the spec output.
awk '/^ [0-9]+\) /{flag=1} flag' "$log" | grep -E "^\s+[0-9]+\)" || true
echo
echo "----- passing/failing summary -----"
# mocha default (spec) reporter: passing/failing counts and the failure list.
grep -E '[0-9]+ (passing|pending|failing)' "$log" || true
if grep -qE '[1-9][0-9]* failing' "$log"; then
echo
echo "----- FAILED TESTS -----"
# Numbered failure headers look like " 1) sanitizeHtml ... : <title>"
grep -E '^[[:space:]]+[0-9]+\)' "$log" || true
fi
echo "Full log: $log"
echo "(full log: $log)"
exit "$code"
84 changes: 34 additions & 50 deletions packages/sanitize-html/index.js
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,6 @@ const mediaTags = [
];
// Tags that are inherently vulnerable to being used in XSS attacks.
const vulnerableTags = [ 'script', 'style' ];
// Tags that establish an SVG or MathML "foreign content" subtree. Inside such
// a subtree the HTML5 parser treats raw-text elements like <textarea> and
// <xmp> as ordinary foreign elements rather than raw-text elements, so a
// browser parses their contents as live markup instead of plain text.
// htmlparser2 decides raw-text purely by tag name and ignores the namespace,
// so we must not re-emit raw-text content unescaped when it appears anywhere
// inside one of these roots (see the `ontext` handler).
const foreignContentRootTags = [ 'svg', 'math' ];

function each(obj, cb) {
if (obj) {
Expand Down Expand Up @@ -585,24 +577,40 @@ function sanitizeHtml(html, options, _recursing) {
// your concern, don't allow them. The same is essentially true for style tags
// which have their own collection of XSS vectors.
result += text;
} else if (tag && tagAllowed(tag) && ((options.disallowedTagsMode === 'discard') || (options.disallowedTagsMode === 'completelyDiscard')) && ((tag === 'textarea') || (tag === 'xmp')) && !insideForeignContent()) {
// htmlparser2 treats <textarea> and <xmp> as raw text elements and
// does NOT decode entities inside them. The text is already properly
// encoded, so pass it through without additional escaping to avoid
// double-encoding. Other "nonTextTags" like <option> are not raw text
// elements in htmlparser2, so their contents are decoded and must be
// escaped below like any other text (important to prevent XSS via
// entity-encoded payloads such as
// <option>&lt;script&gt;...&lt;/script&gt;</option>).
//
// This raw-text pass-through is ONLY safe in the HTML namespace. Inside
// an <svg> or <math> foreign-content subtree the HTML5 parser treats
// <textarea>/<xmp> as ordinary foreign elements, so a browser re-parses
// their contents as live markup. `insideForeignContent()` detects that
// case and forces the text through the escaping branch below, closing
// an allowlist/XSS bypass such as
// `<svg><textarea><img src=x onerror=alert(1)></textarea></svg>`.
result += text;
} else if (tag && tagAllowed(tag) && (options.disallowedTagsMode === 'discard' || options.disallowedTagsMode === 'completelyDiscard') && (tag === 'textarea' || tag === 'xmp')) {
// <textarea> and <xmp> hold text that must not be re-emitted verbatim:
// if a raw `<` survives into the output it can reopen a tag when the
// result is re-parsed by a browser, smuggling non-allowlisted markup
// through the allowlist (mutation-XSS, GHSA-jxwj-j7wr-gfrw — e.g. the
// `</textarea/>` solidus mis-close). We therefore ALWAYS escape this
// content rather than passing it through. Escaping unconditionally also
// closes the related SVG/MathML foreign-content bypass (where the HTML5
// parser treats <textarea>/<xmp> as ordinary foreign elements and a
// browser re-parses their contents as live markup, e.g.
// `<svg><textarea><img src=x onerror=alert(1)></textarea></svg>`): since
// we never re-emit raw text for these tags, no namespace check is
// needed. The two tags need different escaping because htmlparser2
// tokenizes them differently:
if (tag === 'xmp') {
// <xmp> is a raw-text (CDATA) element: entities are NOT decoded, so
// its content reaches us as raw source that is already entity-encoded.
// Escape only the angle brackets so a literal `<` cannot reopen a tag,
// while leaving `&` untouched to avoid double-encoding entities that
// are already encoded in the source.
result += text.replace(/</g, '&lt;').replace(/>/g, '&gt;');
} else {
// <textarea> is an RCDATA element: htmlparser2 (>= 11) decodes
// entities inside it, so its content reaches us as plain decoded text.
// It must therefore be fully escaped like any other text — escaping
// `&` as well as `<`/`>` — so entities round-trip faithfully (no
// double-encoding, see the CVE-2026-40186 regression tests) and no
// `<` can reopen a tag.
result += escapeHtml(text, false);
}
// Other "nonTextTags" like <option> are not raw text elements in
// htmlparser2, so their contents are decoded and must be escaped below
// like any other text (important to prevent XSS via entity-encoded
// payloads such as <option>&lt;script&gt;...&lt;/script&gt;</option>).
} else if (!addedText) {
const escaped = escapeHtml(text, false);
if (options.textFilter) {
Expand Down Expand Up @@ -729,30 +737,6 @@ function sanitizeHtml(html, options, _recursing) {
skipTextDepth = 0;
}

// True when the currently open element is inside an <svg> or <math>
// foreign-content subtree. We check the effective *output* tag name of each
// ancestor (`frame.name`, set when a tag is renamed by transformTags, else
// `frame.tag`) so that a tag renamed to svg/math is caught and a svg/math
// renamed to something else is not. We deliberately do NOT treat HTML
// integration points (e.g. <foreignObject>, <mtext>) as escapes from foreign
// content: such a point only re-establishes the HTML namespace if it survives
// in the output, and it may be dropped (not allowlisted) or altered, which
// would leave the raw-text element as a direct child of svg/math again. Since
// escaping is always safe, treating the whole svg/math subtree as foreign is
// the robust, bypass-free choice.
function insideForeignContent() {
for (let i = stack.length - 1; i >= 0; i--) {
const frame = stack[i];
// The effective *output* tag name: frame.name is set when transformTags
// renamed the element, otherwise fall back to the original parsed tag.
const outputTag = frame.name || frame.tag;
if (foreignContentRootTags.includes(outputTag)) {
return true;
}
}
return false;
}

function escapeHtml(s, quote) {
if (typeof (s) !== 'string') {
s = s + '';
Expand Down
5 changes: 4 additions & 1 deletion packages/sanitize-html/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -25,10 +25,13 @@
],
"author": "Apostrophe Technologies, Inc.",
"license": "MIT",
"engines": {
"node": ">=22.12.0"
},
"dependencies": {
"deepmerge": "^4.2.2",
"escape-string-regexp": "^4.0.0",
"htmlparser2": "^10.1.0",
"htmlparser2": "^12.0.0",
"is-plain-object": "^5.0.0",
"launder": "workspace:^",
"parse-srcset": "^1.0.2",
Expand Down
Loading
Loading