GHSA-g53g-w8rj-fmg7
Dashboard / Vulnerabilities / GHSA-g53g-w8rj-fmg7
Summary: xmldom PI grammar regex ReDoS: quadratic backtracking on unterminated processing instructions
Details: ## Summary `@xmldom/xmldom`'s processing-instruction (PI) grammar regex exhibits quadratic-time backtracking (ReDoS) when parsing an **unterminated** processing instruction. A single small XML document containing `<?` + a target + a long run of whitespace and no closing `?>` forces the regular expression engine into O(n²) work, stalling the Node.js event loop. The input is parsed with `DOMParser.parseFromString` under **default options**, so it is reachable from unauthenticated, network-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs). ## Details The PI production in `lib/grammar.js` compiles (flags `mu`) to: ``` ^<\?(NameChars)(?:[\x20\x09\x0D\x0A]+([Char]*?))?\?> ^^^ S+ greedy ^^^ Char*? lazy ``` - `lib/grammar.js` line 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261 In the optional tail `(?:S+(Char*?))?`, both the greedy separator `S+` and the lazy data `Char*?` match XML whitespace. When the required trailing `?>` is absent, the engine must ultimately fail — but first it tries every partition of the whitespace run between `S+` and `Char*?`, which is O(n²) in the length of the trailing whitespace. The regex is executed against the **entire remaining source string** in two places in `lib/sax.js`, so the whole whitespace tail is scanned: - `parsePI` — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691 - `parseProcessingInstruction` — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879 ## Affected Versions Only the `0.9.x` line is affected. `lib/grammar.js` (and this PI regex) was introduced in commit `726b471` ("fix!: preserve DOCTYPE internal subset (#498)"), first released in **0.9.0-beta.9**, and is unchanged through **0.9.10**. The `0.8.x` line (≤ 0.8.13) and the unscoped `xmldom` package (≤ 0.6.0) parse PIs via a different code path bounded by `indexOf('?>')` — they do **not** contain this regex and are **not affected** by this issue. (They were not separately tested for a *different* PI ReDoS; the scope here is the specific `grammar.js` regex.) | Line | PI code path | Affected? | |---|---|---| | `0.9.x` (0.9.0-beta.9 … 0.9.10) | `grammar.js` `PI` regex over full remaining source | **Yes** | | `0.8.x` (≤ 0.8.13) | `parseInstruction`, bounded by `indexOf('?>')` | No | | unscoped `xmldom` (≤ 0.6.0) | older `indexOf('?>')`-bounded parsing | No | ## Proof of Concept ```js const { DOMParser } = require('@xmldom/xmldom'); const n = 32 * 1024; const payload = '<a><?p' + ' '.repeat(n); // unterminated PI, no `?>` console.time('parse'); new DOMParser().parseFromString(payload, 'text/xml'); console.timeEnd('parse'); ``` Measured (Node 18), trailing whitespace after `<?p`, no `?>` — time quadruples per doubling of input length (canonical O(n²)): | Trailing whitespace | `g.PI.exec` | `parseFromString` | |---|---|---| | 2 KB | 4.4 ms | 5.1 ms | | 4 KB | 16.9 ms | 17.0 ms | | 8 KB | 111.4 ms | 66.3 ms | | 16 KB | 263.8 ms | 336.5 ms | | 32 KB | 1073.1 ms | — | ## Impact Availability only: a single parse of a small crafted document blocks the Node.js event loop for the duration of the quadratic scan (≈1 s at 32 KB; multi-second with larger inputs). No memory blow-up, no data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and parsed with default options, one request can stall a server. ## Fix Applied Fixed in `@xmldom/xmldom` **0.9.11** (`0.9.x`-only; the `0.8.x` LTS line and the unscoped `xmldom` package use a different, bounded PI code path and are not affected). PR [#1039](https://github.com/xmldom/xmldom/pull/1039) inserts a fixed-width negative lookahead `(?!\s)` immediately after the greedy `S+`, so the separator can no longer hand whitespace back to the lazy data group: ``` - var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(', Char, '*?)'), '?', /\?>/); + var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(?!', _SChar, ')(', Char, '*?)'), '?', /\?>/); ``` The change is correct, minimal, and behavior-preserving: it produces identical `[target, data]` captures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl cases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead is fixed-width and cannot itself backtrack — a strict improvement with no new parsing risk. ## Severity note The complexity is **quadratic**, not exponential, so a multi-second stall requires tens-to-hundreds of KB of input. `VA:H` reflects that xmldom applies **no input-size limit** and the path runs on default-options parsing, so a single unbounded parse can fully stall the event loop.
References: https://github.com/xmldom/xmldom/security/advisories/GHSA-g53g-w8rj-fmg7, https://nvd.nist.gov/vuln/detail/CVE-2026-83606, https://github.com/xmldom/xmldom/pull/1039, https://github.com/xmldom/xmldom/commit/73df6b8bdbd86f904b9e8c3ab9c49aa54ef2802e, https://github.com/xmldom/xmldom, https://github.com/xmldom/xmldom/releases/tag/0.9.11
Affected packages
Package
Name: @xmldom/xmldom
Purl: pkg:npm/%40xmldom/xmldom
Affected ranges
Type: SEMVER
Events:
