Skip to content

Commit eae1fb2

Browse files
authored
fix for </textarea/> vulnerability (#5501)
* wip * fix for math/svg vulnerabilities
1 parent d7b6b85 commit eae1fb2

6 files changed

Lines changed: 279 additions & 118 deletions

File tree

‎.changeset/sanitize-html-svg-mathml-rawtext-xss.md‎

Lines changed: 9 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -5,13 +5,18 @@
55
Security: fixed an XSS/allowlist bypass in which the contents of a raw-text
66
element (`textarea` or `xmp`) nested inside an `svg` or `math` root were
77
re-emitted without HTML-escaping. `sanitize-html` treated that content as inert
8-
raw text because `htmlparser2` classifies raw-text elements by tag name and
9-
ignores the namespace, but a real HTML5 parser treats `textarea`/`xmp` as
8+
raw text because `htmlparser2` 10.x classified raw-text elements by tag name and
9+
ignored the namespace, but a real HTML5 parser treats `textarea`/`xmp` as
1010
ordinary foreign elements inside SVG/MathML and re-parses their contents as live
1111
markup. As a result, markup and event-handler attributes that the allowlist
1212
never permitted (for example `<svg><textarea><img src=x onerror=alert(1)>`)
13-
could survive sanitization and execute in the browser. Raw-text content is now
14-
HTML-escaped whenever it appears anywhere inside an `svg` or `math` subtree. The
13+
could survive sanitization and execute in the browser. This is now fixed on two
14+
fronts: `htmlparser2` was upgraded to 12.x, which is namespace-aware and parses
15+
`textarea`/`xmp` inside SVG/MathML as ordinary elements, so their
16+
non-allowlisted children (such as the injected `img`) are dropped by the
17+
allowlist instead of being preserved as raw text; and any raw-text content
18+
`sanitize-html` still emits for these tags (at HTML integration points such as
19+
`foreignObject`/`mtext`, or outside foreign content) is always HTML-escaped. The
1520
default configuration is not affected; the precondition is an `allowedTags` that
1621
includes `svg` or `math` together with `textarea` or `xmp`. Thanks to
1722
[khoadb175](https://github.com/khoadb175) for responsibly disclosing the
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
---
2+
"sanitize-html": patch
3+
---
4+
5+
Security: fixed a mutation-XSS / `allowedTags` bypass affecting configurations that allow the `textarea` or `xmp` raw-text tags. `htmlparser2` 10.x did not recognize an end tag with a trailing solidus (e.g. `</textarea/>`) as closing the element, so it kept the following markup as raw text, but a spec-compliant browser treats `</textarea/>` as a valid close and parses that markup as a live element. Because raw-text content was re-emitted without escaping, a payload such as `<textarea></textarea/><img src=x onerror=...>` could smuggle non-allowlisted, executable markup through the sanitizer. The default configuration was not affected. This is now defended at two layers: `htmlparser2` was upgraded to 12.x, whose tokenizer closes these end tags correctly, and the raw text sanitize-html emits for these tags is always escaped so no `<` can reopen a tag when the output is re-parsed (`textarea`, an RCDATA element whose entities `htmlparser2` decodes, is escaped like normal text, while `xmp`, a raw-text element, has only its angle brackets escaped to avoid double-encoding already-encoded entities). Because `htmlparser2` is ESM-only from version 11 onward, `sanitize-html` now requires Node.js `>=22.12.0` (the first 22.x release in which `require()` of an ES module is available unflagged). Thanks to [bibu123456](https://github.com/bibu123456) for reporting the vulnerability and [Kayiz-PT](https://github.com/Kayiz-PT) for coordinating the disclosure (GHSA-jxwj-j7wr-gfrw).
Lines changed: 20 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -1,12 +1,14 @@
11
#!/bin/bash
2-
# Run the sanitize-html package mocha suite and log full output to
3-
# claude-tools/logs/sanitize-html.log. Prints a summary of failing tests.
2+
# Run the sanitize-html mocha suite, logging full output to
3+
# claude-tools/logs/sanitize-html.log and printing a summary of ONLY the
4+
# specific tests that failed (so we never have to re-run from scratch to find
5+
# out what broke). Optional first arg is a mocha --grep filter, e.g.:
46
#
5-
# Usage:
67
# ./claude-tools/run-sanitize-html-tests.sh # whole suite
7-
# ./claude-tools/run-sanitize-html-tests.sh "<grep string>" # filter by title
8+
# ./claude-tools/run-sanitize-html-tests.sh 'GHSA-jxwj' # just matching tests
89
#
9-
# Runs one suite at a time only. Does not run lint (use `npm test` for that).
10+
# NEVER run test suites in parallel — they are designed to run one at a time
11+
# and the host has limited resources.
1012

1113
set -u
1214
grep_filter="${1:-}"
@@ -19,24 +21,25 @@ log="$logdir/sanitize-html.log"
1921

2022
cd "$root/packages/sanitize-html"
2123

22-
echo "=== sanitize-html tests ${grep_filter:+(grep: $grep_filter) }($(date -Is)) ===" | tee -a "$log"
24+
echo "=== sanitize-html mocha ${grep_filter:+(grep: $grep_filter) }($(date +%Y-%m-%dT%H:%M:%S%z)) ===" | tee -a "$log"
2325

2426
if [[ -n "$grep_filter" ]]; then
25-
./node_modules/.bin/mocha --reporter spec --grep "$grep_filter" >> "$log" 2>&1
27+
./node_modules/.bin/mocha --grep "$grep_filter" >> "$log" 2>&1
2628
else
27-
./node_modules/.bin/mocha --reporter spec >> "$log" 2>&1
29+
./node_modules/.bin/mocha >> "$log" 2>&1
2830
fi
2931
code=$?
3032

3133
echo "=== exit=$code ===" | tee -a "$log"
32-
33-
# Surface the passing/failing counts and any failing test titles.
34-
echo "--- summary ---"
35-
grep -E "passing|failing|pending" "$log" | tail -3
36-
if [[ "$code" -ne 0 ]]; then
37-
echo "--- failing tests ---"
38-
# Mocha lists failures as a numbered block after the spec output.
39-
awk '/^ [0-9]+\) /{flag=1} flag' "$log" | grep -E "^\s+[0-9]+\)" || true
34+
echo
35+
echo "----- passing/failing summary -----"
36+
# mocha default (spec) reporter: passing/failing counts and the failure list.
37+
grep -E '[0-9]+ (passing|pending|failing)' "$log" || true
38+
if grep -qE '[1-9][0-9]* failing' "$log"; then
39+
echo
40+
echo "----- FAILED TESTS -----"
41+
# Numbered failure headers look like " 1) sanitizeHtml ... : <title>"
42+
grep -E '^[[:space:]]+[0-9]+\)' "$log" || true
4043
fi
41-
echo "Full log: $log"
44+
echo "(full log: $log)"
4245
exit "$code"

‎packages/sanitize-html/index.js‎

Lines changed: 34 additions & 50 deletions
Original file line numberDiff line numberDiff line change
@@ -12,14 +12,6 @@ const mediaTags = [
1212
];
1313
// Tags that are inherently vulnerable to being used in XSS attacks.
1414
const vulnerableTags = [ 'script', 'style' ];
15-
// Tags that establish an SVG or MathML "foreign content" subtree. Inside such
16-
// a subtree the HTML5 parser treats raw-text elements like <textarea> and
17-
// <xmp> as ordinary foreign elements rather than raw-text elements, so a
18-
// browser parses their contents as live markup instead of plain text.
19-
// htmlparser2 decides raw-text purely by tag name and ignores the namespace,
20-
// so we must not re-emit raw-text content unescaped when it appears anywhere
21-
// inside one of these roots (see the `ontext` handler).
22-
const foreignContentRootTags = [ 'svg', 'math' ];
2315

2416
function each(obj, cb) {
2517
if (obj) {
@@ -585,24 +577,40 @@ function sanitizeHtml(html, options, _recursing) {
585577
// your concern, don't allow them. The same is essentially true for style tags
586578
// which have their own collection of XSS vectors.
587579
result += text;
588-
} else if (tag && tagAllowed(tag) && ((options.disallowedTagsMode === 'discard') || (options.disallowedTagsMode === 'completelyDiscard')) && ((tag === 'textarea') || (tag === 'xmp')) && !insideForeignContent()) {
589-
// htmlparser2 treats <textarea> and <xmp> as raw text elements and
590-
// does NOT decode entities inside them. The text is already properly
591-
// encoded, so pass it through without additional escaping to avoid
592-
// double-encoding. Other "nonTextTags" like <option> are not raw text
593-
// elements in htmlparser2, so their contents are decoded and must be
594-
// escaped below like any other text (important to prevent XSS via
595-
// entity-encoded payloads such as
596-
// <option>&lt;script&gt;...&lt;/script&gt;</option>).
597-
//
598-
// This raw-text pass-through is ONLY safe in the HTML namespace. Inside
599-
// an <svg> or <math> foreign-content subtree the HTML5 parser treats
600-
// <textarea>/<xmp> as ordinary foreign elements, so a browser re-parses
601-
// their contents as live markup. `insideForeignContent()` detects that
602-
// case and forces the text through the escaping branch below, closing
603-
// an allowlist/XSS bypass such as
604-
// `<svg><textarea><img src=x onerror=alert(1)></textarea></svg>`.
605-
result += text;
580+
} else if (tag && tagAllowed(tag) && (options.disallowedTagsMode === 'discard' || options.disallowedTagsMode === 'completelyDiscard') && (tag === 'textarea' || tag === 'xmp')) {
581+
// <textarea> and <xmp> hold text that must not be re-emitted verbatim:
582+
// if a raw `<` survives into the output it can reopen a tag when the
583+
// result is re-parsed by a browser, smuggling non-allowlisted markup
584+
// through the allowlist (mutation-XSS, GHSA-jxwj-j7wr-gfrw — e.g. the
585+
// `</textarea/>` solidus mis-close). We therefore ALWAYS escape this
586+
// content rather than passing it through. Escaping unconditionally also
587+
// closes the related SVG/MathML foreign-content bypass (where the HTML5
588+
// parser treats <textarea>/<xmp> as ordinary foreign elements and a
589+
// browser re-parses their contents as live markup, e.g.
590+
// `<svg><textarea><img src=x onerror=alert(1)></textarea></svg>`): since
591+
// we never re-emit raw text for these tags, no namespace check is
592+
// needed. The two tags need different escaping because htmlparser2
593+
// tokenizes them differently:
594+
if (tag === 'xmp') {
595+
// <xmp> is a raw-text (CDATA) element: entities are NOT decoded, so
596+
// its content reaches us as raw source that is already entity-encoded.
597+
// Escape only the angle brackets so a literal `<` cannot reopen a tag,
598+
// while leaving `&` untouched to avoid double-encoding entities that
599+
// are already encoded in the source.
600+
result += text.replace(/</g, '&lt;').replace(/>/g, '&gt;');
601+
} else {
602+
// <textarea> is an RCDATA element: htmlparser2 (>= 11) decodes
603+
// entities inside it, so its content reaches us as plain decoded text.
604+
// It must therefore be fully escaped like any other text — escaping
605+
// `&` as well as `<`/`>` — so entities round-trip faithfully (no
606+
// double-encoding, see the CVE-2026-40186 regression tests) and no
607+
// `<` can reopen a tag.
608+
result += escapeHtml(text, false);
609+
}
610+
// Other "nonTextTags" like <option> are not raw text elements in
611+
// htmlparser2, so their contents are decoded and must be escaped below
612+
// like any other text (important to prevent XSS via entity-encoded
613+
// payloads such as <option>&lt;script&gt;...&lt;/script&gt;</option>).
606614
} else if (!addedText) {
607615
const escaped = escapeHtml(text, false);
608616
if (options.textFilter) {
@@ -729,30 +737,6 @@ function sanitizeHtml(html, options, _recursing) {
729737
skipTextDepth = 0;
730738
}
731739

732-
// True when the currently open element is inside an <svg> or <math>
733-
// foreign-content subtree. We check the effective *output* tag name of each
734-
// ancestor (`frame.name`, set when a tag is renamed by transformTags, else
735-
// `frame.tag`) so that a tag renamed to svg/math is caught and a svg/math
736-
// renamed to something else is not. We deliberately do NOT treat HTML
737-
// integration points (e.g. <foreignObject>, <mtext>) as escapes from foreign
738-
// content: such a point only re-establishes the HTML namespace if it survives
739-
// in the output, and it may be dropped (not allowlisted) or altered, which
740-
// would leave the raw-text element as a direct child of svg/math again. Since
741-
// escaping is always safe, treating the whole svg/math subtree as foreign is
742-
// the robust, bypass-free choice.
743-
function insideForeignContent() {
744-
for (let i = stack.length - 1; i >= 0; i--) {
745-
const frame = stack[i];
746-
// The effective *output* tag name: frame.name is set when transformTags
747-
// renamed the element, otherwise fall back to the original parsed tag.
748-
const outputTag = frame.name || frame.tag;
749-
if (foreignContentRootTags.includes(outputTag)) {
750-
return true;
751-
}
752-
}
753-
return false;
754-
}
755-
756740
function escapeHtml(s, quote) {
757741
if (typeof (s) !== 'string') {
758742
s = s + '';

‎packages/sanitize-html/package.json‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -25,10 +25,13 @@
2525
],
2626
"author": "Apostrophe Technologies, Inc.",
2727
"license": "MIT",
28+
"engines": {
29+
"node": ">=22.12.0"
30+
},
2831
"dependencies": {
2932
"deepmerge": "^4.2.2",
3033
"escape-string-regexp": "^4.0.0",
31-
"htmlparser2": "^10.1.0",
34+
"htmlparser2": "^12.0.0",
3235
"is-plain-object": "^5.0.0",
3336
"launder": "workspace:^",
3437
"parse-srcset": "^1.0.2",

0 commit comments

Comments
 (0)