Five CVSS 8.7 denial-of-service CVEs hit vLLM's disaggregated serving connectors, and every fix pull request was still open on 3 October.
Six denial-of-service flaws in the KV-transfer and KV-offload paths of vLLM, the open-source LLM serving engine, were published to NVD on 21 September 2026 as CVE-2026-94622 through CVE-2026-94627. Five carry a CVSS v4.0 score of 8.7 from VulnCheck; CVE-2026-94625 scores 6.9. The headline case, CVE-2026-94622, lets a request with an incomplete kv_transfer_params entry kill the decode engine of a prefill/decode disaggregated deployment, and every routed request fails until someone restarts it. As of 3 October, the fix pull requests cited by these records were all still open and unmerged, and no record we read names a fixed vLLM release.
What was published
The NVD records for CVE-2026-94622, 94623, 94624, 94625, 94626 and 94627 all went public on 21 September and all say "through 0.29.0", so vLLM 0.29.0 (released 9 September) is itself affected. vLLM 0.30.0 followed on 22 September. NVD lists no NVD-own score for these records; the scores below are VulnCheck's, who assigned the CVE ids.
| CVE | Issue | Score (VulnCheck) | Fix PR on 3 Oct |
|---|---|---|---|
| CVE-2026-94622 | NIXL: incomplete kv_transfer_params cause an uncaught KeyError; decode engine exits | CVSS v4 8.7, v3.1 7.5 | #54807, open |
| CVE-2026-94623 | NIXL: multi-prompt completions trigger an assertion failure; decode worker exits | v4 8.7, v3.1 7.5 | #51505, open |
| CVE-2026-94624 | P2P KV offload: arbitrary peer host and port exhaust the ZeroMQ context; EngineCore crashes | v4 8.7, v3.1 7.5 | #51504, open |
| CVE-2026-94626 | tp_size in kv_transfer_params unvalidated; kernel OOM-kills the decode worker | v4 8.7, v3.1 7.5 | #51137, open |
| CVE-2026-94627 | Mooncake: concurrent child requests sharing one transfer ID leak GPU KV blocks | v4 8.7, v3.1 7.5 | #49796, open |
| CVE-2026-94625 | Mooncake: ownerless transfer placeholders delay valid requests | v4 6.9, v3.1 5.3 | #51236, open |
We checked each of the six pull requests against GitHub's API on 3 October: all show state: open, merged: false, targeting main. A merged fix on main would still need a release before most operators could run it. The v0.30.0 release notes list other NIXL and Mooncake connector fixes, but none of these pull request numbers appears in them, and the Security section does not mention these flaws. So 0.30.0 is not confirmed as a fix for any of the six. Treat the question as open until the maintainers say otherwise.
How CVE-2026-94622 works
The NIXL connector reads a kv_transfer_params dictionary supplied with the request. According to the record, entries with missing keys raise an uncaught KeyError during EngineCore scheduling. The decode engine terminates, and all routed requests fail until a manual restart. VulnCheck's vector is CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N, with high availability impact (VA:H): reachable over the network, low complexity, no privileges and no user interaction. The CVSS v3.1 vector is AV:N/AC:L/PR:N/UI:N with high availability impact, scoring 7.5. VulnCheck's advisory credits Mingkai Yu, Jiapeng Li and Jiajia Liu.
In a disaggregated deployment, a prefill instance computes the KV cache and hands it to a decode instance that generates tokens. The decode side is where one crash takes out serving capacity for every tenant behind it. Whether an attacker can reach the decode instance with a crafted request depends on how the deployment is exposed; the CVSS vectors assume it is reachable.
The sibling flaws
Four of the records (94622, 94623, 94626, 94627) name prefill/decode disaggregated deployments; the other two concern the Mooncake connector and P2P KV offloading. The failure modes differ.
- CVE-2026-94623 (NIXL). A completion request carrying several prompts of different lengths triggers an assertion failure in
NixlBaseConnectorWorker._apply_prefix_caching. The decode worker terminates and stays down until restarted. - CVE-2026-94626.
tp_sizeinkv_transfer_paramsis not validated, so an attacker can drive unbounded memory allocation until the kernel OOM-kills the decode worker. - CVE-2026-94627 (Mooncake). Concurrent child requests that share one transfer ID mismanage GPU KV block ownership. Multi-prompt completion requests exhaust GPU memory, and orphaned blocks accumulate until the process restarts.
- CVE-2026-94625 (Mooncake). Rejected prefill requests leave transfer placeholders that are never reclaimed. Valid requests can be delayed by up to 480 seconds while health checks keep returning success, which makes this one hard to spot from a load balancer's view.
- CVE-2026-94624 (P2P KV offload). This applies when
OffloadingConnectoris configured withTieringOffloadingSpecand a peer-to-peer secondary tier. Attacker-chosen remote host and port values create unreachable peer sessions that hold ZeroMQ sockets until the context quota runs out. The resulting uncaught ZMQError crashes EngineCore and stops all inference.
What defenders should do
The mitigations below are ThreatFrontier's suggestions, inferred from the flaw descriptions. They are not vendor guidance, and we have not tested them against these CVEs.
- Inventory exposure. Find deployments that run vLLM 0.29.0 or earlier with the NIXL connector, the Mooncake connector, or P2P KV offloading in a prefill/decode split. The records do not describe a plain single-instance setup without these connectors as affected.
- Keep prefill and decode endpoints off untrusted networks. Reach them only from your own router or proxy. The CVSS vectors assume a network attacker with no privileges, so network position is the main control you have today.
- Filter
kv_transfer_paramsat the edge. Clients of a public API have no reason to send this field. Strip or reject it at a gateway before the request reaches vLLM, and consider rejecting completion requests that carry multiple prompts if your clients do not use them. - Watch the decode tier for crash loops. Alert on EngineCore exits, decode-worker restarts and OOM kills, and add a probe that sends a real request through the decode path. For the Mooncake placeholder flaw, a health check that returns success is not evidence of health.
- Track the pull requests. Follow #54807, #51505, #51504, #51137, #49796 and #51236, and the vLLM release notes, for the merge and the first release that includes them. Do not assume an upgrade to 0.30.0 fixes this group.
Other vLLM records in the same wave
NVD's keyword search for "vllm" over 19 September to 3 October returns 16 records, separate bugs, none scoring above the 8.7 of the connector group. Two are worth a line for operators of high-concurrency serving.
CVE-2026-100647 needs no authentication: the cache_salt field has a minimum length but no maximum, and it is hashed with sha256(pickle.dumps(...)) on the single EngineCore scheduler thread. The GitHub advisory (GHSA-wpww-v874-ph2p) measured a 100 MB salt at about 599 ms of added latency against a baseline near 9 ms, and with six concurrent attackers unrelated requests reached a 90th-percentile latency of about 1.8 seconds. vLLM applies no HTTP body limit by default. It scores 6.9 (CVSS v4) and 5.3 (v3.1) from VulnCheck, and the advisory lists 0.29.0 as patched. A proxy-level body limit blunts it.
GHSA-85xf-c7hm-whqw, which has no CVE id and scores 6.5 (CVSS v3.1), covers three structured-output paths where an ordinary request raises an engine-fatal exception that denies service to all tenants. Its metadata lists 0.30.0 or later as patched. Both advisories are separate from the connector flaws above.
Sources
- NVD, CVE-2026-94622: https://services.nvd.nist.gov/rest/json/cves/2.0?cveId=CVE-2026-94622
- NVD, CVE-2026-94623 to 94627 (same API,
cveIdparameter): https://services.nvd.nist.gov/rest/json/cves/2.0?cveId=CVE-2026-94626 - VulnCheck advisory for CVE-2026-94622: https://www.vulncheck.com/advisories/vllm-through-0.29.0-denial-of-service-via-incomplete-nixl-kv-transfer-metadata
- vLLM pull request #54807: https://github.com/vllm-project/vllm/pull/54807
- vLLM pull request #51505: https://github.com/vllm-project/vllm/pull/51505
- vLLM pull request #51504: https://github.com/vllm-project/vllm/pull/51504
- vLLM pull request #51137: https://github.com/vllm-project/vllm/pull/51137
- vLLM pull request #49796: https://github.com/vllm-project/vllm/pull/49796
- vLLM pull request #51236: https://github.com/vllm-project/vllm/pull/51236
- vLLM v0.30.0 release notes: https://github.com/vllm-project/vllm/releases/tag/v0.30.0
- GHSA-wpww-v874-ph2p (CVE-2026-100647): https://github.com/vllm-project/vllm/security/advisories/GHSA-wpww-v874-ph2p
- GHSA-85xf-c7hm-whqw: https://github.com/vllm-project/vllm/security/advisories/GHSA-85xf-c7hm-whqw