NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2776 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release 12 days ago
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
Breaking changes : scale-out endpoints are opt-in on plain vllm serve via --enable-scale-out , replacing VLLM_ENABLE_SCALE_OUT_ENDPOINTS ( #54579 , #5…
This release features 762 commits from 315 contributors (104 new)!
--load-format ipc_cache instead of reloading from disk (#54921), now covering FP4 checkpoints (#55465) and multi-node TP (#55468).HiSparseConnector (#53781), with Prometheus counters (#56061), a host cache shared across TP ranks (#56629), and the attention config inferred from the connector (#57041).--return-sampling-mask compacted on GPU, fixing an about 2x RL step-time regression (#54901).--engram-config (#54371), and torch.compile removed from the NVIDIA implementation so FP8 fits on a single GB300 (#55272).quantization_config.targets (#51285) and on partially pre-quantized checkpoints from any quant method (#51392), W4A16 DSA with the nvfp4_fp8_ds_mla KV cache (#51724), FlashInfer CuTeDSL NVFP4 W4A16 default over Marlin on SM100/103 (#53014), NVFP4 in the torch linear backend (#53319), per-quantization linear backend overrides (#51204), AutoRound 2/3/5/6/7-bit on CUDA (#52890), and DeepSelect top-k for the DSA sparse indexer (#56464).vllm serve via --enable-scale-out, replacing VLLM_ENABLE_SCALE_OUT_ENDPOINTS (#54579, #55176); GPTQ activation ordering (g_idx) removed (#54809); items deprecated for 0.29 removed, including the VLLM_PREFIX_CACHE_RETENTION_INTERVAL and VLLM_MM_HASHER_ALGORITHM env vars (#55353); the all Mamba cache mode deprecated (#55041); python -m vllm.entrypoints.grpc_server deprecated in favor of vllm serve --grpc (#56746); YaRN aligned with Transformers so vendor YaRN aliases no longer re-scale max_model_len (#56446).| Platform | Install |
|---|---|
| PyPI (CUDA 13.0) | pip install vllm |
| PyPI (CUDA 13.0, uv) | uv pip install vllm --torch-backend=auto |
| ROCm | pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/0.30.0/rocm723 |
| XPU | uv pip install vllm --extra-index-url https://wheels.vllm.ai/0.30.0/xpu --extra-index-url https://download.pytorch.org/whl/xpu --index-strategy unsafe-best-match |
| Platform | Docker Image |
|---|---|
| CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.30.0 |
| CUDA 12.9 | docker pull vllm/vllm-openai:v0.30.0-cu129 |
| CUDA 13.0 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.30.0-ubuntu2404 |
| CUDA 12.9 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.30.0-cu129-ubuntu2404 |
| ROCm | docker pull vllm/vllm-openai-rocm:v0.30.0 |
| CPU | docker pull vllm/vllm-openai-cpu:v0.30.0 |
| XPU | docker pull vllm/vllm-openai-xpu:v0.30.0 |
Pre-built release artifacts are available in the Assets section at the bottom of this page, including:
--kv-cache-dtype auto resolving to fp8_ds_mla with FlashMLA (#45091), prefill sparse index workspace sentinel seeded (#55299), dequant gather grid sized by rows (#55061), optional Q-norm in the fused MLA epilogue and group_size=32 packed FP8 quant (#56215), sparse settings read from the text config for composite models (#56160), MegaMoE startup without EP fixed (#55914), and a Triton iHC pre/post fallback for HY V4 (#55059).glm5next NVIDIA subtree now packaged in wheels (#55214).VLLM_SPARSE_INDEXER_MAX_LOGITS_MB (#54513), fused PLE kernels and merged K/V projections (#54517), sparse GQA skipping padded indices (#54873), FP8 indexer cache (#54890), fused PLE residual and QSA output gate (#55309), UVA PLE offload and Engram tensor parallelism via --engram-config (#54371), torch.compile removed from the NVIDIA path (#55272), compact indexer logits workspace (#54915), HC combine-norm reused for MTP input (#54687), PLE MTP metadata transfers (#55054), Hopper LL-GEMM tuning table (#54560), B200 FP8 MoE tuning (#55890), and fixes for FP8 PLE weight scales (#54722, #54882) and fused PLE conv strides (#55375).n_predict from the text config (#55369), Qwen3 DSpark padded-vocab drafts (#55133), and GLM-OCR MTP weight loading (#49869) and position masking under CUDA graphs (#56447).audio_backend in --media-io-kwargs (#51826) with automatic decoding kept soundfile-first (#55642), torchaudio as the default resampler (#52598), media_io_kwargs in multimodal hashes (#54241), cache hash kwargs scoped by modality (#54918), empty video URLs with multimodal UUIDs (#54220), cached audio inputs with UUIDs (#56310), encoder CUDA graphs for MiniCPM-V 2.5/2.6/4.0 (#42785), Voxtral Realtime with FULL_DECODE_ONLY graphs (#51167), Triton/FlashInfer composite attention for multimodal prefixes (#56305), pruned sliding-window tiles for Gemma 4 multimodal prompts (up to 3-4x E2E, #53147), SDPA for BLIP-2 Q-Former (#55285), baddbmm Conformer scores (#55062), fused DeepEncoder relative bias (#55629), Gemma3n sparse GELU Triton kernel (#48498), and no duplicate text embedding in Qwen2.5-Omni (#55415).Note truncated.
Breaking changes : ten deprecated model architectures removed ( #53608 ); FlexOlmo, Olmo3 and Hunyuan V1/VL migrated to the Transformers modeling back…
This release features 594 commits from 277 contributors (91 new)!
extract_hidden_states speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694). MRV1 remains in use for a few ROCm models and features MRV2 does not yet support.eh_proj (12.9-25.2% kernel speedup, #53942), GEMM-RS extended to GEMM-AR (#53053), MLA gate merged into the QKV-A projection (#54015), K3 DCP with DSpark (#52188) and DCP partial prefix cache hits (#50493); DeepSeek V4 shared experts fused into MegaMoE (#53040), adaptive top-k width re-landed (#52823), a native SwiGLU clamp kernel for Humming MoE (#53685), and an opt-in FlashInfer moe_ep expert backend (#49636).--per-request-spec-decode-metrics (#48915), adaptive verification extended to logprobs (#52242), SM100 sparse MLA for GLM-5.2 (#52783) and DeepSeek V4 on SM90 (#52795), Qwen3-Omni DSpark drafts (#52560), PLaMo3 EAGLE-3/DFlash (#54239), and DFlash2 loading from speculators format (#53797).sharded_rdt P2P backend where each worker pulls only its TP/EP slice over NIXL or Ray Direct Transport (#43375), rank-local IPC weight updates (#52497), sparse checkpoint-coordinate updates through native weight loaders (#50723, #53751), and routed expert loading for gpt-oss (#52209).prefix_cache_retention_interval is now a CLI argument defaulting to 0 (#52216), with dense retention automatically restored for hybrid models using EAGLE/MTP (#55760, #55861).VLLM_ALLREDUCE_USE_FLASHINFER=0 (#52998); prefix-cache NONE_HASH is deterministic by default so distributed KV cache users no longer need to pin PYTHONHASHSEED (#51875); new --max-num-queued-reqs / --max-num-queued-tokens admission-control flags (#49445).python -m vllm.entrypoints.openai.api_server deprecated in favor of vllm serve (#52131); VLLM_TEST_FORCE_FP8_MARLIN (#52182) and VLLM_ROCM_USE_AITER_FP4_ASM_GEMM (#53141) removed.Now that Model Runner V2 is used by default, we are considering Model Runner V1 deprecated and are targeting v0.32 for its removal. We do not intend to accept any more MRV1-specific improvements or optimizations.
Some features are not yet supported in MRV2 but we are planning for these gaps to be closed within the next 2-3 weeks. These include sequence parallelism, dual-batch overlap, elastic expert parallellism, custom logits processors and certain speculative decoding methods. For now, vLLM will still fall back to use MRV1 if any of these features are configured.
| Platform | Install |
|---|---|
| PyPI (CUDA 13.0) | pip install vllm |
| PyPI (CUDA 13.0, uv) | uv pip install vllm --torch-backend=auto |
| ROCm | pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/0.29.0/rocm723 |
| XPU | uv pip install vllm --extra-index-url https://wheels.vllm.ai/0.29.0/xpu --extra-index-url https://download.pytorch.org/whl/xpu --index-strategy unsafe-best-match |
| Platform | Docker Image |
|---|---|
| CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.29.0 (v0.29.0-cu130 also works) |
| CUDA 12.9 | docker pull vllm/vllm-openai:v0.29.0-cu129 |
| CUDA 13.0 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.29.0-ubuntu2404 |
| CUDA 12.9 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.29.0-cu129-ubuntu2404 |
| ROCm | docker pull vllm/vllm-openai-rocm:v0.29.0 |
| CPU | docker pull vllm/vllm-openai-cpu:v0.29.0 |
| XPU | docker pull vllm/vllm-openai-xpu:v0.29.0 |
Pre-built release artifacts are available in the Assets section at the bottom of this page, including:
RMSNormFuser.fuse performance fixed (#52766).lm_head is loaded even when the config claims tied embeddings (#51665, #53170).mm_processor_kwargs honored (#53808), Qwen3-VL profiling honors cap_pixels_per_frame (#54380), video frame sampling respects MP4 edit-list trims (#48608), oversized items skip the MM processor cache instead of crashing (#53016), encoder cache entries survive until their last use (#54284), and repeated multimodal requests with the SHM cache no longer terminate the engine (#54994).extract_hidden_states (#49811), padded FULL cudagraph dispatch (#53407), DP-sync skipping for drafts (#53694), decoupled draft/target gumbel noise streams (#54282), encoder-only path split out (#53176), and memory released correctly on shutdown and sleep (#53508, #54246, #54162, #53955, #53682).DSparkDraftModel configs for Qwen3 (#52197), and Gemma4 MTP under CUDA graphs (#53884).prefix_cache_retention_interval argument (#52216, #55760, #55861), deterministic NONE_HASH (#51875), queue admission control (#49445), KV null block reserved when validating max_model_len (#47272), spec decode no longer padded up to max_model_len (#53962), and negative external block allocation prevented (#52707).sharded_rdt P2P weight sync (#43375), rank-local IPC updates (#52497), sparse checkpoint updates (#50723, #53751), gpt-oss routed expert loading (#52209), stable DeepSeek V4 mHC broadcast buffers across weight sync (#52626), and packed weight transfer stream reuse capping reserved-memory waste (#52951).trace_decode_token_ids for deterministic decode replay (#46701), per-arch tuned batch-invariant matmul configs (about 3x decode kernels on RTX 4090D/H20, #53247), Blackwell autotuning with 33.6% E2E latency reduction (#53649), deterministic MoE combine under DP+EP (#45683), and fuse_allreduce_rms disabled under batch invariance (#51292).--linear-backend flashinfer_cutedsl (#50572), replicated embedding and norm fusion for DSV3 flat models (#48484), standardized fused shared-expert selection (#51695), GPT-OSS MoE topk metadata reuse (#45457), tuned cooperative topk (#53382), and tuned FP8 fused_moe for Qwen3.5 on L40S (+7%, #53819).KVCacheLayout enum (#51718), JIT warmup provider registry (#50174), --cpu-offload-params now reaches vision/audio towers (#53120), attention backend probe failures no longer crash init (#51703), FlashInfer XQA falls back on unsupported head_dim (#53111), FlashInfer prefill LSE normalized before merging to fix prefix-cache logits divergence (#52796), seed preserved when a batch mixes seeded and unseeded requests (#51866), startup thread allocation accounts for local DP workers (#52385), int32 overflow fixes in fused SiLU block quant (#53409) and LoRA kernels (#53034), a shared-memory race in fused groupwise RMSNorm quantization (#54111), BLHNC addressing for FlashInfer sparse MLA (#54465), Mamba state copy race (#50729), and a start_profile no-op after auto-stop (#51839).Note truncated.
Your coding agent can read these notes before it upgrades. Set up the MCP server →