© 2025-2026 PySpect
First version without a known vulnerability: 0.30.0
vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
Fixed in: 0.28.0
vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle
Fixed in: 0.30.0
vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach
Fixed in: 0.30.0
vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
Fixed in: 0.30.0
vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Fixed in: 0.30.0
vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Fixed in: 0.30.0
vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`
Fixed in: 0.30.0
vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP deployments — uncaught downstream `ValueError` denial of service
Fixed in: 0.30.0
A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attack may be launched remotely. The exploit has been published and may be used. Upgrading to version 0.29.1rc0 is able to resolve this issue. This patch is called 3439bad37e68ba9755a46f4f6b44a4aeaf1f60a9. Upgrading the affected component is advised.
Fixed in: 0.27.0
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.
Fixed in: 0.24.0
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths — the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route — the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or, on some paths, never applies it). A remote attacker can therefore cause the API server or batch-runner process to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. There is no code execution or data disclosure impact.
Fixed in: 0.30.0
vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, allowing attackers to allocate unbounded memory. Attackers can supply arbitrary tp_size values in prefill/decode disaggregated deployments to exhaust memory and trigger kernel OOM-kill of the decode worker process.
Fixed in: 0.30.0
vLLM through 0.29.0 contains a resource exhaustion vulnerability in MooncakeConnector where rejected prefill requests create ownerless transfer placeholders that are never reclaimed. Attackers can send rejected requests to exhaust sender task pools, causing valid requests to be delayed by up to 480 seconds while health checks continue returning success.
Fixed in: 0.30.0
vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can supply arbitrary remote host and port values in kv_transfer_params to create unreachable peer sessions that retain ZeroMQ sockets until the context quota is exhausted, causing an uncaught ZMQError that crashes EngineCore and stops all inference.
Fixed in: 0.30.0
vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix caching implementation that fails to properly validate block counts across multi-prompt completion requests in prefill/decode disaggregated deployments. Attackers can trigger an assertion failure in NixlBaseConnectorWorker._apply_prefix_caching by submitting completion requests with multiple prompts of varying lengths, causing the decode worker to terminate and become unavailable until restarted.
Fixed in: 0.30.0
vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entries to trigger an uncaught KeyError in EngineCore scheduling, causing the decode engine to terminate and making all routed requests fail until manual restart.
Fixed in: 0.30.0
vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of concurrent requests, causing different in-flight HTTP requests to return incorrect tokens.
Fixed in: 0.30.0
vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal audio requests with tokens equal to vocabulary size, causing out-of-bounds writes that corrupt concurrent requests' sampler state and alter repetition penalty behavior.
Fixed in: 0.30.0
vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing LogitBiasState to corrupt GPU logits state and allow concurrent requests to sample tokens outside their allowlists.
Fixed in: 0.29.0
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
Fixed in: 0.28.0
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
Fixed in: 0.30.0
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Fixed in: 0.28.0
vLLM: Unauthenticated audio decompression-bomb DoS in /v1/chat/completions
Fixed in: 0.24.0
vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor loader that ignores the trust_remote_code parameter when loading remote processor classes. Attackers can craft a malicious model with arbitrary code in processing_llava_onevision2.py that executes with vLLM process authority even when trust_remote_code is set to False.
Fixed in: 0.28.0
vLLM: SSRF + arbitrary local file read in MiMoV2OmniMultiModalProcessor `_fetch_image` and audio loader bypass MediaConnector protections
Fixed in: 0.26.0
vLLM: Cross-User Data Leak Vulnerability
Fixed in: 0.27.0
vLLM: Incomplete CVE-2025-62164 remediation can be bypassed by concurrent prompt parts
Fixed in: 0.26.0
vLLM: ReDoS via structured_outputs.regex in the lm-format-enforcer backend (no compile timeout) — missed sibling of GHSA-rwxx-mrjm-wc2m
Fixed in: 0.26.0
vLLM: Unauthenticated Internal Path and Username Disclosure via Validation Error Messages
Fixed in: 0.26.0
vLLM: Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds
Fixed in: 0.26.0
vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the process-wide GPU decode pool and submit video that bypasses resource controls, causing partial denial of service for concurrent requests.
Fixed in: 0.27.0
vLLM: Completion prompt lists fan out into unbounded engine requests
Fixed in: 0.26.0
vLLM denial of service via prompt embeds on M-RoPE models
Fixed in: 0.24.0
vLLM: Speech-to-text upload size limit is enforced after full UploadFile read
Fixed in: 0.24.0
vLLM: ReDoS via structured_outputs.regex compiled without timeout in xgrammar and outlines backends
Fixed in: 0.24.0
vLLM has Remote DoS via Invalid Recovered Token Reinjection
Fixed in: 0.24.0
vLLM: Processing differential in multi-channel audio downmixing enables hidden-input/moderation bypass for audio models
Fixed in: 0.18.0
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM Dockerfile is vulnerable to a dependency confusion attack through the flashinfer-jit-cache package. The package is installed from a custom index (flashinfer.ai/whl/) using --extra-index-url, but the package name was not registered on PyPI, and UV_INDEX_STRATEGY="unsafe-best-match" is set globally. An attacker who registers flashinfer-jit-cache on PyPI with version 0.6.11.post2 can execute arbitrary code as root during the Docker build and backdoor every resulting container image, enabling exfiltration of all user prompts, API credentials, and model data from production vLLM deployments This vulnerability is fixed in 0.22.1.
Fixed in: 0.22.1
vLLM: OOM Denial of Service via Audio Decompression Bomb
Fixed in: 0.24.0
vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router
Fixed in: 0.24.0
vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving
Fixed in: 0.24.0
vLLM: image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
Fixed in: 0.24.0
vLLM: temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels
Fixed in: 0.24.0
vLLM: OpenAI auth bypass
Fixed in: 0.22.0
vLLM: Security Check Bypass via assert Statement in Activation Function Loading Allows Arbitrary Code Execution
Fixed in: 0.22.0
vLLM is vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attack due to unbounded frame count processing in the `VideoMediaIO.load_base64()` method
Fixed in: 0.19.0
vLLM's Artifact Pin Decay allows pinned deployments to load unpinned code, weights, and processors
Fixed in: 0.22.0
vllm has Improper Resource Shutdown or Release
vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters
Fixed in: 0.20.0
vLLM Vulnerable to Remote DoS via Special-Token Placeholders
Fixed in: 0.20.0
vLLM makes Use of Uninitialized Resource
Fixed in: 0.19.1
vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
Fixed in: 0.19.0
vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
Fixed in: 0.19.0
vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
Fixed in: 0.19.0
vLLM has Hardcoded Trust Override in Model Files Enables RCE Despite Explicit User Opt-Out
Fixed in: 0.18.0
vLLM has SSRF Protection Bypass
Fixed in: 0.17.0
vLLM has RCE In Video Processing
Fixed in: 0.14.1
vLLM vulnerable to Server-Side Request Forgery (SSRF) through MediaConnector
Fixed in: 0.14.1
vLLM affected by RCE via auto_map dynamic module loading during model initialization
Fixed in: 0.14.0
vLLM is vulnerable to DoS in Idefics3 vision models via image payload with ambiguous dimensions
Fixed in: 0.12.0
vLLM introduced enhanced protection for CVE-2025-62164
Fixed in: 0.13.0
vLLM vulnerable to remote code execution via transformers_utils/get_config
Fixed in: 0.11.1
vLLM vulnerable to DoS via large Chat Completion or Tokenization requests with specially crafted `chat_template_kwargs`
Fixed in: 0.11.1
vLLM vulnerable to DoS with incorrect shape of multimodal embedding inputs
Fixed in: 0.11.1
vLLM deserialization vulnerability leading to DoS and potential RCE
Fixed in: 0.11.1
vLLM is vulnerable to Server-Side Request Forgery (SSRF) through `MediaConnector` class
Fixed in: 0.11.0
vLLM: Resource-Exhaustion (DoS) through Malicious Jinja Template in OpenAI-Compatible Server
Fixed in: 0.11.0
vLLM is vulnerable to timing attack at bearer auth
Fixed in: 0.11.0
vLLM has remote code execution vulnerability in the tool call parser for Qwen3-Coder
Fixed in: 0.10.1.1
vllm API endpoints vulnerable to Denial of Service Attacks
Fixed in: 0.10.1.1
vLLM Tool Schema allows DoS via Malformed pattern and type Fields
Fixed in: 0.9.0
vLLM allows clients to crash the openai server with invalid regex
Fixed in: 0.9.0, 08bf7840780980c7568c573c70a6a8db94fd45ff
vLLM DOS: Remotely kill vllm over http with invalid JSON schema
Fixed in: 0.9.0, 08bf7840780980c7568c573c70a6a8db94fd45ff
vLLM has a Weakness in MultiModalHasher Image Hashing Implementation
Fixed in: 0.9.0, 99404f53c72965b41558aceb1bc2380875f5d848
Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching
Fixed in: 0.9.0, 77073c77bc2006eb80ea6d5128f076f5e6c6f54f
vLLM vulnerable to Regular Expression Denial of Service
Fixed in: 0.9.0
vLLM has a Regular Expression Denial of Service (ReDoS, Exponential Complexity) Vulnerability in `pythonic_tool_parser.py`
Fixed in: 0.9.0, 4fc1bf813ad80172c1db31264beaef7d93fe0601
vLLM Allows Remote Code Execution via PyNcclPipe Communication Service
Fixed in: 0.8.5
Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration
Fixed in: 0.10.0
vLLM: Quadratic Time Complexity in Input Token Processing leads to denial of service
Fixed in: 0.8.5
vLLM Vulnerable to Remote Code Execution via Mooncake Integration
Fixed in: 0.8.5, a5450f11c95847cf51a17207af9a3ca5ab569b2c
Data exposure via ZeroMQ on multi-node vLLM deployment
Fixed in: 0.8.5
CVE-2025-24357 Malicious model remote code execution fix bypass with PyTorch < 2.6.0
Fixed in: 0.8.0
vLLM vulnerable to Denial of Service by abusing xgrammar cache
Fixed in: 0.8.4
vLLM allows Remote Code Execution by Pickle Deserialization via AsyncEngineRPCServer() RPC server entrypoints
vLLM deserialization vulnerability in vllm.distributed.GroupCoordinator.recv_object
vLLM Deserialization of Untrusted Data vulnerability
vLLM Allows Remote Code Execution via Mooncake Integration
Fixed in: 0.8.0, 288ca110f68d23909728627d3100e5a8db820aa2
vLLM denial of service via outlines unbounded cache on disk
Fixed in: 0.8.0
vLLM uses Python 3.12 built-in hash() which leads to predictable hash collisions in prefix cache
Fixed in: 0.7.2, 432117cd1f59c76d97da2eaff55a7d758301dbc7
vllm: Malicious model to RCE by torch.load in hf_model_weights_iterator
Fixed in: 0.7.0, d3d6bb13fb62da3234addf6574922a4ec0513d04
vLLM denial of service vulnerability
Fixed in: 0.5.5
vLLM Denial of Service via the best_of parameter