© 2025-2026 PySpect
First version without a known vulnerability: 0.26.0
vLLM: Completion prompt lists fan out into unbounded engine requests
Fixed in: 0.26.0
vLLM denial of service via prompt embeds on M-RoPE models
Fixed in: 0.24.0
vLLM: Speech-to-text upload size limit is enforced after full UploadFile read
Fixed in: 0.24.0
vLLM: ReDoS via structured_outputs.regex compiled without timeout in xgrammar and outlines backends
Fixed in: 0.24.0
vLLM has Remote DoS via Invalid Recovered Token Reinjection
Fixed in: 0.24.0
vLLM: Processing differential in multi-channel audio downmixing enables hidden-input/moderation bypass for audio models
Fixed in: 0.18.0
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM Dockerfile is vulnerable to a dependency confusion attack through the flashinfer-jit-cache package. The package is installed from a custom index (flashinfer.ai/whl/) using --extra-index-url, but the package name was not registered on PyPI, and UV_INDEX_STRATEGY="unsafe-best-match" is set globally. An attacker who registers flashinfer-jit-cache on PyPI with version 0.6.11.post2 can execute arbitrary code as root during the Docker build and backdoor every resulting container image, enabling exfiltration of all user prompts, API credentials, and model data from production vLLM deployments This vulnerability is fixed in 0.22.1.
Fixed in: 0.22.1
vLLM: OOM Denial of Service via Audio Decompression Bomb
Fixed in: 0.24.0
vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router
Fixed in: 0.24.0
vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving
Fixed in: 0.24.0
vLLM: image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
Fixed in: 0.24.0
vLLM: temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels
Fixed in: 0.24.0
vLLM: OpenAI auth bypass
Fixed in: 0.22.0
vLLM: Security Check Bypass via assert Statement in Activation Function Loading Allows Arbitrary Code Execution
Fixed in: 0.22.0
vLLM is vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attack due to unbounded frame count processing in the `VideoMediaIO.load_base64()` method
Fixed in: 0.19.0
vLLM's Artifact Pin Decay allows pinned deployments to load unpinned code, weights, and processors
Fixed in: 0.22.0
vllm has Improper Resource Shutdown or Release
vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters
Fixed in: 0.20.0
vLLM Vulnerable to Remote DoS via Special-Token Placeholders
Fixed in: 0.20.0
vLLM makes Use of Uninitialized Resource
Fixed in: 0.19.1
vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
Fixed in: 0.19.0
vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
Fixed in: 0.19.0
vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
Fixed in: 0.19.0
vLLM has Hardcoded Trust Override in Model Files Enables RCE Despite Explicit User Opt-Out
Fixed in: 0.18.0
vLLM has SSRF Protection Bypass
Fixed in: 0.17.0
vLLM has RCE In Video Processing
Fixed in: 0.14.1
vLLM vulnerable to Server-Side Request Forgery (SSRF) through MediaConnector
Fixed in: 0.14.1
vLLM affected by RCE via auto_map dynamic module loading during model initialization
Fixed in: 0.14.0
vLLM is vulnerable to DoS in Idefics3 vision models via image payload with ambiguous dimensions
Fixed in: 0.12.0
vLLM introduced enhanced protection for CVE-2025-62164
Fixed in: 0.13.0
vLLM vulnerable to remote code execution via transformers_utils/get_config
Fixed in: 0.11.1
vLLM vulnerable to DoS via large Chat Completion or Tokenization requests with specially crafted `chat_template_kwargs`
Fixed in: 0.11.1
vLLM vulnerable to DoS with incorrect shape of multimodal embedding inputs
Fixed in: 0.11.1
vLLM deserialization vulnerability leading to DoS and potential RCE
Fixed in: 0.11.1
vLLM is vulnerable to Server-Side Request Forgery (SSRF) through `MediaConnector` class
Fixed in: 0.11.0
vLLM: Resource-Exhaustion (DoS) through Malicious Jinja Template in OpenAI-Compatible Server
Fixed in: 0.11.0
vLLM is vulnerable to timing attack at bearer auth
Fixed in: 0.11.0
vLLM has remote code execution vulnerability in the tool call parser for Qwen3-Coder
Fixed in: 0.10.1.1
vllm API endpoints vulnerable to Denial of Service Attacks
Fixed in: 0.10.1.1
vLLM Tool Schema allows DoS via Malformed pattern and type Fields
Fixed in: 0.9.0
vLLM allows clients to crash the openai server with invalid regex
Fixed in: 0.9.0, 08bf7840780980c7568c573c70a6a8db94fd45ff
vLLM DOS: Remotely kill vllm over http with invalid JSON schema
Fixed in: 0.9.0, 08bf7840780980c7568c573c70a6a8db94fd45ff
vLLM has a Weakness in MultiModalHasher Image Hashing Implementation
Fixed in: 0.9.0, 99404f53c72965b41558aceb1bc2380875f5d848
Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching
Fixed in: 0.9.0, 77073c77bc2006eb80ea6d5128f076f5e6c6f54f
vLLM vulnerable to Regular Expression Denial of Service
Fixed in: 0.9.0
vLLM has a Regular Expression Denial of Service (ReDoS, Exponential Complexity) Vulnerability in `pythonic_tool_parser.py`
Fixed in: 0.9.0, 4fc1bf813ad80172c1db31264beaef7d93fe0601
vLLM Allows Remote Code Execution via PyNcclPipe Communication Service
Fixed in: 0.8.5
Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration
Fixed in: 0.10.0
vLLM: Quadratic Time Complexity in Input Token Processing leads to denial of service
Fixed in: 0.8.5
vLLM Vulnerable to Remote Code Execution via Mooncake Integration
Fixed in: 0.8.5, a5450f11c95847cf51a17207af9a3ca5ab569b2c
Data exposure via ZeroMQ on multi-node vLLM deployment
Fixed in: 0.8.5
CVE-2025-24357 Malicious model remote code execution fix bypass with PyTorch < 2.6.0
Fixed in: 0.8.0
vLLM vulnerable to Denial of Service by abusing xgrammar cache
Fixed in: 0.8.4
vLLM allows Remote Code Execution by Pickle Deserialization via AsyncEngineRPCServer() RPC server entrypoints
vLLM deserialization vulnerability in vllm.distributed.GroupCoordinator.recv_object
vLLM Deserialization of Untrusted Data vulnerability
vLLM Allows Remote Code Execution via Mooncake Integration
Fixed in: 0.8.0, 288ca110f68d23909728627d3100e5a8db820aa2
vLLM denial of service via outlines unbounded cache on disk
Fixed in: 0.8.0
vLLM uses Python 3.12 built-in hash() which leads to predictable hash collisions in prefix cache
Fixed in: 0.7.2, 432117cd1f59c76d97da2eaff55a7d758301dbc7
vllm: Malicious model to RCE by torch.load in hf_model_weights_iterator
Fixed in: 0.7.0, d3d6bb13fb62da3234addf6574922a4ec0513d04
vLLM denial of service vulnerability
Fixed in: 0.5.5
vLLM Denial of Service via the best_of parameter