PatchSiren cyber security CVE debrief
CVE-2026-94627 vllm-project CVE debrief
The vLLM Mooncake connector through 0.29.0 is vulnerable to a GPU memory exhaustion attack. This occurs because the connector fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. An attacker can trigger this vulnerability by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart. This accumulation can eventually prevent legitimate requests from executing, leading to a denial-of-service condition. Defenders should prioritize verifying exposure, especially in prefill/decode disaggregated deployments, and assess the
- Vendor
- vllm-project
- Product
- vllm
- CVSS
- HIGH 8.7
- CISA KEV
- Not listed in stored evidence
- Original CVE published
- 2026-09-21
- Original CVE updated
- 2026-09-29
- Advisory published
- 2026-09-21
- Advisory updated
- 2026-09-29
Who should care
Defenders responsible for vLLM Mooncake connector deployments, especially those using prefill/decode disaggregated deployments, should assess exposure and prioritize verification and potential updates or mitigations.
Why it matters
CVE-2026-94627 allows attackers to trigger GPU memory exhaustion in vLLM Mooncake connector deployments, potentially causing denial-of-service conditions. Defenders should prioritize verifying exposure, especially in prefill/decode disaggregated deployments, and assess the need for updates or mitigations.
- GPU memory exhaustion can prevent legitimate requests from executing
- Potential for denial-of-service (DoS) conditions in affected deployments
- Need for verification of exposure in prefill/decode disaggregated deployments
- Priority on updating or mitigating affected vLLM Mooncake connector versions
Technical summary
The vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. This can allow attackers to trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart and eventually preventing legitimate requests from executing.
Defensive priority
Defenders should prioritize verifying exposure in vLLM Mooncake connector deployments, especially those using prefill/decode disaggregated deployments, and assess the need for updates or mitigations.
Recommended defensive actions
- Verify vLLM Mooncake connector version and assess exposure in prefill/decode disaggregated deployments.
- Review and apply patches or updates from the vendor, if available.
- Monitor GPU memory usage and implement mitigations to prevent memory exhaustion.
- Perform an inventory of assets using the vLLM Mooncake connector to identify potential exposure.
- Consider rollback or change windows for updating affected versions.
- Track the status of mitigations and patches through source tracking.
- Implement compensating controls for exposed systems while remediation is scheduled and verified.
Evidence notes
The CVE record and NVD entry provide details on the vulnerability in vLLM Mooncake connector through 0.29.0, describing a failure in managing GPU KV cache block ownership. Official sources include CVE Program and NIST NVD records.
Sources and references
Verified primary and authoritative sources
-
CVE-2026-94627 CVE Program record
Publisher, destination, and source semantics verified
URL: https://www.cve.org/CVERecord?id=CVE-2026-94627
CVE Program - Official CVE Program record with source-provided CVE metadata.
-
CVE-2026-94627 NVD vulnerability detail
Publisher, destination, and source semantics verified
URL: https://nvd.nist.gov/vuln/detail/CVE-2026-94627
NIST National Vulnerability Database - Official NIST NVD detail page and source-specific vulnerability assessment.
Supplemental references
-
Source reference
Unverified legacy reference
URL: https://github.com/vllm-project/vllm
[email protected] - Product
-
Source reference
Unverified legacy reference
URL: https://github.com/vllm-project/vllm/blob/v0.29.0/vllm/distributed/kv_transfer/kv_connector/v1/mooncake/mooncake_connector.py
[email protected] - Product
-
Source reference
Unverified legacy reference
URL: https://github.com/vllm-project/vllm/pull/49796
[email protected] - Issue Tracking, Patch
-
Source reference
Unverified legacy reference
URL: https://www.vulncheck.com/advisories/vllm-through-0.29.0-gpu-kv-cache-leak-via-mooncake-transfer-id-collision
[email protected] - Third Party Advisory
Methodology and review provenance
AI-assisted synthesis based on stored public vulnerability evidence. System validation, approval state, and publication status do not by themselves establish human review of this revision. PatchSiren helps prioritize defensive review and does not prove exposure or remediation on any system.