These pages are published after PatchSiren validates generated defensive summaries against stored public CVE and source evidence.
A vulnerability in vLLM's multimodal cache can lead to a desync between the frontend and engine core processes. This desync allows a later request reusing the same media hash to trip a receiver assertion in the engine core. The vulnerability arises when a request is rejected after the frontend has rendered and hashed the multimodal input but before the engine core receives the item. As a result, the front [truncated]
A vulnerability in the vLLM package allows an ordinary API request to cause an uncaught, engine-fatal exception, leading to a denial of service for all concurrent and subsequent tenants of the engine. The issue arises from the absence of a per-request exception boundary around structured-output grammar/token handling. This vulnerability affects vLLM ≤ 0.25.1 and is fixed in version 0.30.0. Defenders shoul [truncated]
A vulnerability in the vLLM package allows for a cross-request integrity break and induced errors on `/score` and `/rerank` due to the caching of query embeddings under a caller-controlled request ID. This issue affects vLLM versions ≤ 0.25.1 and is remotely reachable. The vulnerability arises from the worker caching per-request query embeddings under a key derived from the caller-controlled `X-Request-Id [truncated]
The vLLM inference and serving engine for large language models, versions from 0.24.0 until 0.30.0, contains a vulnerability in the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes. These classes accept and use request-level values for video.max_frames and video.fps without enforcing server-side limits. An unauthenticated user can submit these values to the /tokenize endpoint, causing the video sampler [truncated]
The CVE-2026-105754 issue in vLLM, a large language model inference and serving engine, allows an attacker to supply tensors, cache identifiers, and multimodal field processors that can lead to a denial of service or unintended behavior by accepting caller-supplied features without proper validation. This issue is fixed in version 0.30.0. Affected deployments should be identified, and owners assigned for [truncated]
A vulnerability was found in vllm-project vLLM up to 0.26.0, which affects unknown code in the Gemma4UnifiedParser component. This issue can lead to a denial of service and may be exploited remotely. The exploit has been published and can be used. Upgrading to version 0.29.1rc0 resolves this issue with patch 3439bad37e68ba9755a46f4f6b44a4aeaf1f60a9. Users of the affected component should upgrade.
A denial-of-service vulnerability exists in vLLM versions 0.22.0 through 0.23.0. Attackers can submit requests with out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures, leaving EngineCore in a fatal state requiring service restart. This issue arises from the failure to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token [truncated]
A remote attacker can cause the API server or batch-runner process of vLLM through 0.29.0 to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The vulnerability exists across multiple ingress paths, including the shared media-acquisit [truncated]
The vLLM Mooncake connector through 0.29.0 is vulnerable to a GPU memory exhaustion attack. This occurs because the connector fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. An attacker can trigger this vulnerability by submitting completion requests with multiple prompts, causing orphaned KV cache [truncated]
AI-assisted PatchSiren debrief based on the supplied source corpus. The CVE record was published on 2026-09-21T22:17:01.587Z and has not been modified since then. The vulnerability in vLLM through 0.29.0 allows attackers to cause memory exhaustion by supplying arbitrary tp_size values, impacting systems with prefill/decode disaggregated deployments. Defenders should assess exposure and prioritize patching [truncated]
CVE-2026-94625 is a resource exhaustion vulnerability in vLLM through version 0.29.0. Rejected prefill requests can create ownerless transfer placeholders that are never reclaimed, causing valid requests to be delayed by up to 480 seconds. Health checks continue returning success despite this issue. Defenders should assess exposure and prioritize verification and remediation efforts. The vulnerability req [truncated]
AI-assisted PatchSiren debrief based on the supplied source corpus. The CVE record was published on 2026-09-21T22:17:01.280Z and has not been modified since then. The vulnerability is a denial-of-service issue in vLLM through 0.29.0, specifically in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. This allows attackers to supply arbitra [truncated]
CVE-2026-94623 is a high-severity denial-of-service vulnerability affecting vLLM through version 0.29.0. The vulnerability is located in the NIXL connector's prefix caching implementation and can be triggered by submitting completion requests with multiple prompts of varying lengths, causing the decode worker to terminate. This vulnerability can cause significant disruption to services, and defenders shou [truncated]
A denial of service vulnerability exists in vLLM versions through 0.29.0 in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. An attacker can send requests with incomplete kv_transfer_params dictionary entries to trigger an uncaught KeyError in EngineCore scheduling, causing the decode engine to terminate and making all routed requests fail until manual restart.
CVE-2026-93989 is a vulnerability in vLLM through version 0.29.0, where the failure to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer() allows attackers to supply out-of-bounds token indices. This can corrupt logits memory of concurrent requests, causing different in-flight HTTP requests to return incorrect tokens.
CVE-2026-93436 debrief: vLLM through 0.29.0 is vulnerable to memory exhaustion via rejected requests. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts. This vulnerability affects vLLM deployments, particularly those with prefill/decode disaggregated deployments. Defenders should assess exposure and prioritize verification and up [truncated]
CVE-2026-57173 is a vulnerability in the vLLM inference and serving engine for large language models. The input_audio handling path for /v1/chat/completions calls in vLLM prior to 0.24.0 can lead to an out-of-memory worker crash due to a small compressed audio input expanding into a large float32 PCM allocation. This issue affects deployments serving an audio-capable model. The vulnerability is fixed in v [truncated]
A vulnerability was found in vllm-project vllm up to 0.29.0, affecting some unknown functionality of the file vllm/v1/sample/thinking_budget_state.py. The manipulation results in inefficient algorithmic complexity, allowing remote attackers to launch an attack. The pull request to fix this issue awaits acceptance. Defenders should assess exposure and prioritize verification of the vulnerability's presence [truncated]
CVE-2026-37237 is a high-severity vulnerability in the vLLM project, allowing remote attackers to cause a Denial of Service (DoS) via memory exhaustion. The vulnerability is caused by the lack of a maximum response size limit when fetching user-supplied media URLs using aiohttp in the AsyncMediaIO.fetch_audio and AsyncMediaIO.fetch_image functions. This allows an attacker to exhaust server memory by provi [truncated]
CVE-2026-78684 is a medium-severity vulnerability in vLLM before 0.27.0 that allows unauthenticated attackers to cause partial denial of service by bypassing resource controls through the DeepStream GPU backend. The vulnerability can be triggered by activating DeepStream at request time to initialize the process-wide GPU decode pool and submit video that bypasses resource controls. Defenders should assess [truncated]
CVE-2026-73560 is a vulnerability in the vLLM inference and serving engine for large language models. The MiMoV2OmniMultiModalProcessor in vllm/transformers_utils/processors/mimo_v2_omni.py passes attacker-controlled image and audio strings through _fetch_image, requests.get, and Image.open instead of MediaConnector, bypassing allowed_media_domains and allowed_local_media_path protections. This allows ser [truncated]
CVE-2026-71486 is a vulnerability in the vLLM inference and serving engine for large language models. An authenticated API client can cause excessive CPU and memory consumption and produce oversized responses by supplying malformed GenerateResponse objects to the /v1/completions/derender and /v1/chat/completions/derender endpoints. This issue is fixed in version 0.26.0.
CVE-2026-73559 is a vulnerability in vLLM, an inference and serving engine for large language models. The issue allows an authenticated API client to exhaust system resources with a single request, potentially leading to service unavailability. The vulnerability is caused by the /v1/completions endpoint accepting an unbounded list of strings or integers, which can be used to create multiple engine generat [truncated]
AI-assisted PatchSiren debrief based on the supplied source corpus. The CVE record was published on 2026-08-13T15:20:18.220Z and has not been modified since then. The vulnerability is an integer overflow issue in vLLM prior to version 0.27.0, which could allow a request processed in the same inference batch to receive a partial or complete copy of another user's inference result. Defenders should assess e [truncated]
CVE-2026-73557 is a medium-severity vulnerability in the vLLM inference and serving engine for large language models. The issue arises from a race condition in the safe_load_prompt_embeds function, which can be exploited through concurrent submissions to the POST /v1/chat/completions endpoint. This vulnerability has been fixed in version 0.26.0. Affected deployments should be identified, and owners assign [truncated]
CVE-2026-73556 is a vulnerability in the vLLM inference and serving engine for large language models, specifically affecting versions prior to 0.26.0. The issue arises from the structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py being passed to lmformatenforcer.RegexParser without proper validation or compilation with a timeout. This allows an unauthenticated /v1/ [truncated]
CVE-2026-73555 debrief based on the supplied source corpus. The CVE record was published on 2026-08-13T15:20:17.773Z and has not been modified since then. This information disclosure vulnerability in vLLM allows unauthenticated attackers to access sensitive information such as OS usernames, home and virtual-environment paths, Python version, internal package structure, line numbers, and endpoint handler n [truncated]
AI-assisted PatchSiren debrief based on the supplied source corpus. The CVE record was published on 2026-07-06T21:16:57.347Z and has not been modified since then. The vLLM inference and serving engine for LLMs is vulnerable to a denial-of-service attack due to an issue with the structured_outputs.regex API parameter. Prior to version 0.24.0, the API parameter passes a user-supplied regular expression stri [truncated]
A vulnerability in the vLLM library for LLM inference and serving, from version 0.12.0 to before 0.24.0, allows remote users to induce a crash via a /v1/completions request with a model using M-RoPE. This issue arises from a failure in the EngineCore assertion when a pure prompt embeds payload in the request. The vulnerability is fixed in version 0.24.0. This issue has a high impact on the security of the [truncated]
A high-throughput and memory-efficient inference and serving engine for LLMs, vLLM, is vulnerable to a denial of service attack. Prior to version 0.24.0, a specific multi-request speculative decoding workload can cause the engine to crash, resulting in a service-wide denial of service for other clients until the worker is restarted. The vulnerability exists in the vLLM engine's handling of speculative dec [truncated]