PatchSiren

PatchSiren cyber security CVE debrief

CVE-2026-94627 vllm-project CVE debrief

The vLLM Mooncake connector through 0.29.0 is vulnerable to a GPU memory exhaustion attack. This occurs because the connector fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. An attacker can trigger this vulnerability by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart. This accumulation can eventually prevent legitimate requests from executing, leading to a denial-of-service condition. Defenders should prioritize verifying exposure, especially in prefill/decode disaggregated deployments, and assess the

Vendor
vllm-project
Product
vllm
CVSS
HIGH 8.7
CISA KEV
Not listed in stored evidence
Original CVE published
2026-09-21
Original CVE updated
2026-09-29
Advisory published
2026-09-21
Advisory updated
2026-09-29

Who should care

Defenders responsible for vLLM Mooncake connector deployments, especially those using prefill/decode disaggregated deployments, should assess exposure and prioritize verification and potential updates or mitigations.

Why it matters

CVE-2026-94627 allows attackers to trigger GPU memory exhaustion in vLLM Mooncake connector deployments, potentially causing denial-of-service conditions. Defenders should prioritize verifying exposure, especially in prefill/decode disaggregated deployments, and assess the need for updates or mitigations.

  • GPU memory exhaustion can prevent legitimate requests from executing
  • Potential for denial-of-service (DoS) conditions in affected deployments
  • Need for verification of exposure in prefill/decode disaggregated deployments
  • Priority on updating or mitigating affected vLLM Mooncake connector versions

Technical summary

The vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. This can allow attackers to trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart and eventually preventing legitimate requests from executing.

Defensive priority

Defenders should prioritize verifying exposure in vLLM Mooncake connector deployments, especially those using prefill/decode disaggregated deployments, and assess the need for updates or mitigations.

Recommended defensive actions

  • Verify vLLM Mooncake connector version and assess exposure in prefill/decode disaggregated deployments.
  • Review and apply patches or updates from the vendor, if available.
  • Monitor GPU memory usage and implement mitigations to prevent memory exhaustion.
  • Perform an inventory of assets using the vLLM Mooncake connector to identify potential exposure.
  • Consider rollback or change windows for updating affected versions.
  • Track the status of mitigations and patches through source tracking.
  • Implement compensating controls for exposed systems while remediation is scheduled and verified.

Evidence notes

The CVE record and NVD entry provide details on the vulnerability in vLLM Mooncake connector through 0.29.0, describing a failure in managing GPU KV cache block ownership. Official sources include CVE Program and NIST NVD records.

Sources and references

Verified primary and authoritative sources

  • CVE-2026-94627 CVE Program record

    Publisher, destination, and source semantics verified

    URL: https://www.cve.org/CVERecord?id=CVE-2026-94627

    CVE Program - Official CVE Program record with source-provided CVE metadata.

  • CVE-2026-94627 NVD vulnerability detail

    Publisher, destination, and source semantics verified

    URL: https://nvd.nist.gov/vuln/detail/CVE-2026-94627

    NIST National Vulnerability Database - Official NIST NVD detail page and source-specific vulnerability assessment.

Supplemental references

Methodology and review provenance

AI-assisted synthesis based on stored public vulnerability evidence. System validation, approval state, and publication status do not by themselves establish human review of this revision. PatchSiren helps prioritize defensive review and does not prove exposure or remediation on any system.