PatchSiren

PatchSiren cyber security CVE debrief

CVE-2026-100652 vllm-project CVE debrief

A denial-of-service vulnerability exists in vLLM versions 0.22.0 through 0.23.0. Attackers can submit requests with out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures, leaving EngineCore in a fatal state requiring service restart. This issue arises from the failure to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Exploitation involves submitting requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids.

Vendor
vllm-project
Product
vllm
CVSS
HIGH 8.2
CISA KEV
Not listed in stored evidence
Original CVE published
2026-09-26
Original CVE updated
2026-10-07
Advisory published
2026-09-26
Advisory updated
2026-10-07

Who should care

Defenders responsible for deploying and managing vLLM should assess exposure and prioritize patching or mitigation to prevent potential denial-of-service attacks. This includes reviewing the current version of vLLM in use, verifying if it falls within the affected range (0.22.0 through 0.23.0), and taking appropriate actions to secure the system. Additionally, defenders should ensure that input validation and sanitization for stop_token_ids are properly in

Why it matters

A denial-of-service vulnerability in vLLM versions 0.22.0 through 0.23.0 can be exploited by submitting malicious requests, potentially disrupting service availability. Defenders should prioritize verifying affected versions and applying patches or mitigations to prevent exploitation.

  • Denial-of-service attacks can be launched by submitting malicious requests, potentially disrupting service availability.
  • Defenders need to verify affected versions and apply patches or mitigations to prevent exploitation.
  • The vulnerability requires verification of input validation and sanitization for stop_token_ids to prevent out-of-vocabulary token IDs from reaching MinTokensLogitsProcessor.
  • Service restart may be required to recover from a fatal state in EngineCore.

Technical summary

The vulnerability exists in vLLM versions 0.22.0 through 0.23.0, where the failure to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends allows out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. This can be exploited by submitting requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids, triggering CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.

Defensive priority

Defenders should prioritize verifying affected versions and applying patches or mitigations to prevent potential denial-of-service attacks.

Recommended defensive actions

  • Verify if the vLLM version is within the affected range (0.22.0 through 0.23.0) and apply patches or mitigations as needed.
  • Implement input validation and sanitization for stop_token_ids to prevent out-of-vocabulary token IDs from reaching MinTokensLogitsProcessor.
  • Monitor for potential denial-of-service attacks and adjust security controls accordingly.
  • Review compensating controls for exposed systems while remediation is scheduled and verified.
  • Check relevant monitoring, detection, and logs for exposed assets that need extra review.
  • Track exceptions, retest remediated assets, and close the item only after evidence is documented.
  • Confirm whether affected product deployments exist in managed environments and assign an owner for follow-up.

Evidence notes

The vulnerability is caused by a failure to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends. This allows out-of-vocabulary token IDs to reach MinTokensLogitsProcessor, triggering CUDA tensor indexing failures.

Sources and references

Verified primary and authoritative sources

  • CVE-2026-100652 CVE Program record

    Publisher, destination, and source semantics verified

    URL: https://www.cve.org/CVERecord?id=CVE-2026-100652

    CVE Program - Official CVE Program record with source-provided CVE metadata.

  • CVE-2026-100652 NVD vulnerability detail

    Publisher, destination, and source semantics verified

    URL: https://nvd.nist.gov/vuln/detail/CVE-2026-100652

    NIST National Vulnerability Database - Official NIST NVD detail page and source-specific vulnerability assessment.

Supplemental references

  • PYSEC-2026-4185

    Unverified legacy reference

    URL: https://storage.googleapis.com/osv-vulnerabilities/PyPI/PYSEC-2026-4185.json

    osv_dev

  • Source reference

    Unverified legacy reference

    URL: https://www.vulncheck.com/advisories/vllm-0.22.0-through-0.23.0-denial-of-service-via-stop-token-ids

    Supplemental source

  • Source reference

    Unverified legacy reference

    URL: https://github.com/vllm-project/vllm/security/advisories/GHSA-qff2-492f-9fm4

    Supplemental source

Methodology and review provenance

AI-assisted synthesis based on stored public vulnerability evidence. System validation, approval state, and publication status do not by themselves establish human review of this revision. PatchSiren helps prioritize defensive review and does not prove exposure or remediation on any system.