NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2222 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release today
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
Breaking changes : bitsandbytes support migrated to an out-of-tree plugin ( #43529 ); Transformers bumped to 5.15.0 ( #51668 ); the deprecated calcula…
This release features 584 commits from 270 contributors (76 new)!
thinking_token_budget support (#46727).module_path (#51007), partial secondary-tier load results (#50321), tiering metrics (#48798), and a canonical CPU layout for parallelism-agnostic offload (#48414).max_num_batched_tokens raised from 8192 to 16384 (#51726), prefix caching enabled by default for Mamba models (#50991), and the Blackwell CUDA graph capture default raised to 1024 (#49390).calculate_kv_scales runtime KV scale calculation was removed (#49389); override_attention_dtype was removed (#48684).| Platform | Install |
|---|---|
| PyPI (CUDA 13.0) | pip install vllm |
| PyPI (CUDA 13.0, uv) | uv pip install vllm --torch-backend=auto |
| ROCm | pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/0.28.0/rocm722 |
| Platform | Docker Image |
|---|---|
| CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.28.0 (v0.28.0-cu130 also works) |
| CUDA 12.9 | docker pull vllm/vllm-openai:v0.28.0-cu129 |
| CUDA 13.0 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.28.0-ubuntu2404 |
| CUDA 12.9 + Ubuntu 24.04 | docker pull vllm/vllm-openai:v0.28.0-cu129-ubuntu2404 |
| ROCm | docker pull vllm/vllm-openai-rocm:v0.28.0 |
| CPU | docker pull vllm/vllm-openai-cpu:v0.28.0 |
| XPU | docker pull vllm/vllm-openai-xpu:v0.28.0 |
Pre-built release artifacts are available in the Assets section at the bottom of this page, including:
CuMemAllocator.discard() for tag-selective GPU memory release (#52514), level-2 sleep/wake/reload fixed with LoRA enabled (#39935), and rewritten weight-transfer docs with standardized examples (#51729).get_open_port() livelock on DP-reserved ports was fixed (#50965), and NVML is no longer re-initialized on every device-capability check (#50393).module_path (#51007), partial secondary-tier load results (#50321), tiering offloading metrics (#48798), data-parallel topology exposed to offloading backends (#51879), a canonical CPU layout for parallelism-agnostic offload (#48414), and quadratic ARC batch eviction avoided (#50992).count_reasoning_tokens in the streaming parser engine (Note truncated.
This is a patch release on top of v0.27.0.
This is a patch release on top of v0.27.0.
This release features 561 commits from 242 contributors (64 new)!
This release features 561 commits from 242 contributors (64 new)!
vllm-bench integrated into the vllm CLI (#48930).sm_107 target for NVIDIA Rubin (#49387) with NVLink all-reduce paths on SM107 (#49647), and ROCm gfx1250 architecture enabled (#46516).fx tracer (#49957), fused residual-add + RMSNorm compilation pass (#48757), and fixes for MLA padding + grouped topk routing (#49982), MQA with TP (#49987), and Qwen3-VL M-RoPE (#49292)./wake_up crash on hybrid models (#41602).sample_from_anchor loaded from speculators config (#48639), earliest-completing stop string selected (#49391).CustomOp.forward_native compiled for ReLU^2 (#50244), HF config used for HF tokenizers (#49907), batch-invariant RMSNorm via pinned block size (#48391).reduce_scatter regression fix restoring 5% E2E throughput (#48763), non-grouped bias-less topk routing dispatched to the fused path (#49618), tuned LL BF16 router GEMM (#48774) with warmup skipped for non-MoE models (#49659), Triton tensor-descriptor path for fused MoE via VLLM_TRITON_USE_TD (#42436), cudagraph/DP padding skipped in topk (#48979), coalesced HBM access in the Marlin INT4-FP8 AWQ preprocess kernel (#47268).sm_107 for Rubin (#49387), NVLink all-reduce paths on SM107 (#49647), fixed CUDA arch detection producing kernel-less builds on SM121 (#49904).Note truncated.
This release features 411 commits from 212 contributors (61 new)!
This release features 411 commits from 212 contributors (61 new)!
fused_topk_bias (1.5–2x kernel, #47463), and redundant repeat/copy removal (1.8% E2E TPOT, #48137), plus ROCm two-stage compressor for HCA prefill (#47718), sparse decode/prefill optimizations (#48519, #48788, #46275), and DSpark speculative decoding on AMD (#47419) and XPU (#47677).lm_head for generation models via head_dtype (#48390), extended to the LoRA path (#48525) and given a ROCm torch.mm fast path (#48688), improving accuracy for generation heads.vllm-bench port (#48107).lm_head on the LoRA path (#48525), optimized TrtLlmLoRAExperts (#48759).torch.compile (#48901).lm_head for generation models via head_dtype (#48390); lower memory for capturing large CUDA graph sizes (#48483); opt-in persistence and reuse of the memory-profiling result across boots (#47388); improved InstantTensor loading (#46868).kv_cache_dtype for speculative_config (#48787).blocks_per_chunk config for heterogeneous KV groups (#48878), P2P default host/port env vars (#47636).new_block_ids (#44490), DSv3.2 + MTP + sequence-parallel accuracy (#48036).fused_topk_bias 1.5–2x (#47463), redundant repeat/copy removal (1.8% TPOT, #48137)._copy_mamba_state_block to uint64 (#48110), stop upcasting logits to fp32 in the sampler (#48641).head_dtype torch.mm fast path (#48688), DSv4 two-stage compressor kernel (#47718), sparse decode/prefill optimizations (#48519, #48788, #46275), DSv3.2 sparse MLA KV-split heuristic (#46832) and MTP CUDA-graph mode (#45149), MXFP8 GEMM for MiniMax-M3 (#46117), AITER sparse paged attention + spec decode for MiniMax-M3 (#47287, #47984), MiniMax-M2 fused QK-norm + all-reduce via AITER (#44849), HybridW4A16 linear kernel (#40977), Qwen3-30B-A3B QK-Norm+RoPE+KV runtime fusion (#42749).nvfp4_per_token online MoE quantization (#48538), CuTe-DSL FlashInfer MXFP4 quantization (#48417); bounded peak memory when repacking FP4 MoE weights for Marlin (#47851) and for NVFP4 MoE weight loading (#46276).kv_cache_dtype_skip_layers support (#47309).vllm-bench port (#48107), continue_final_message handling with renderer sentinel (#47844).bad_words in /v1/completions (#46793), expose logprob_token_ids on Python OpenAI endpoints (#43463), include_reasoning param for non-Harmony models (#44301), populate num_cache_creation_tokens on Messages responses (#48535)./abort_requests on the RLHF dev API router (#47173); Deepstream video decoding backend (#42424); overlap preprocessing and computation for pooling models in offline inference (#47699).Note truncated.
This release features 2 commits from 2 contributors (1 new)!
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
import torchcodec raised a RuntimeError at import time when system FFmpeg was missing, which blocked startup (e.g. vllm serve Qwen/Qwen3-VL-2B-Instruct) even when TorchCodec was not in use. The error is now deferred to runtime so it only surfaces if TorchCodec is actually needed.!!!!! tokens. A dtype-match guard now routes incompatible mixed-dtype graphs to the safe path, while same-dtype models retain the full allreduce + RMSNorm + quant fusion.@Isotr0py, @hugo-cen
FP8: weights padding for per-block online quantization (#44763); deprecated the old FP8 online MoE quantization class (#44514).
This release features 558 commits from 232 contributors (64 new)!
tok_sparse_select from MSA replacing Triton kernels (#47502).mm_token_type_ids fix (#46552), tied-embedding lm_head.bias fix (#46835).architectures crash fix (#46037), DeepSeek-V2 hidden-size and aux-hidden-state fixes (#46986, #46973).FLASH_ATTN_MLA_SPARSE Hopper sparse-MLA backend (#46189), DCP + FP8 KV cache in MLA decode (#44044), XQA decode kernels (#43232).LookupResult enum (#46363).VLLM_GPU_SYNC_CHECK env var (#44800), VRAM semaphore infrastructure (#44465), skip detokenization in online beam search (#46422), several int32-overflow fixes in sampler/attention kernels (#46560, #47383, #47671).fused_indexer_q_rope_quant Triton kernel (1.9–3.3% E2E throughput) (#46862), reduce-scatter MoE all-reduce (3.1–3.2% E2E) (#46635), op fusion for GLM5/DSV3.2 (#46876), token_to_req_indices cache for DSv4 (5–6x kernel speedup) (#47474), better DSv4 MXFP8 kernel (#47229), redundant-op removal (#47198, #46651).fused_qk_norm_rope (#44010) and silu_and_mul_per_block_quant (#43994), Triton MLA logits workspace (#46819), swap-AB optimization for fused MoE (#36559), vectorized fp32 moe_sum supporting any top-k (#46643), blocking CUDA events to avoid busy-polling the driver lock (#47081).ROCM_AITER_FA (#45033); fused shared-expert for GLM-4.5/6/7 (#44313) and MiniMax-M3 (#46474, #46545); AITER MoE optimization for DeepSeek-V4 (#46122); AITER custom all-reduce in CudaCommunicator (#46065); INT3 quantization for quickreduce (#45666).get_memory_info (#47134).get_memory_info (#44825).kv_transfer_params merging (#46777), usage field exposed for disaggregated serving (#42748).metrics field on Chat/Completions responses (#46768), token offsets on render endpoints (#44226), return_loss_mask for training-data generation (#46846), HTTP 422 for unprocessable image URLs (#47165).process_eos() flush (#46437), raw-output recovery on non-terminal parse (#47062, #47379).repetition_detection sampling param (#46684), unified/combined parser interface (#46583), reduced multimodal tensor copies (#47581), plus many parser and validation fixes.vllm chat (#46775), model_class_overrides for development/debugging (#47148).thinking_token_budget re-entry #43757); rejection of invalid config values (#44070, #44002, #46612) and degenerate structured_outputs that crash EngineCore (#45346).split_audio with NaN audio samples (#46463).truncation_side is set (#47007).api_server.py moved to the examples directory (#46783); gptq_marlin removed from supported ROCm quant schemes (#46655).Thank you to all the contributors who made this release possible!
@AndreasKaratzas, @njhill, @BugenZhao, @hmellor, @yewentao256, @WoosukKwon, @Sunt-ing, @micah-wil, @mgoin, @reidliu41, @peizhang56, @mawong-amd, @TheEpicDolphin, @jeejeelee, @taneem-ibrahim, @chaunceyjiang, @chaojun-zhang, @divakar-amd, @fxmarty-amd, @LopezCastroRoberto, @wzhao18, @mayuyuace, @jperezdealgaba, @noooop, @yzong-rh, @jikunshang, @zxd1997066, @bigPYJ1151, @yma11, @hickeyma, @benchislett, @xianbaoqian, @andakai, @NickLucche, @ivanium, @joerowell, @EazyReal, @mganczarenko, @majunze2001, @hongxiayang, @WindChimeRan, @Rohan138, @tjtanaa, @bbrowning, @thisjiang, @Fangzhou-Ai, @blasrodri, @Isotr0py, @zhenwei-intel, @zyongye, @frida-andersson, @muhammadfawaz1, @lcheng321, @spandantiwari, @Palaiologos1453, @soaringk, @Lynn-hh, @fadara01, @djramic, @Liangliang-Ma, @ronensc, @aarushjain29, @HDCharles, @qianlihuang, @AgenticSpark, @charlifu, @cleonard530, @shen-shanshan, @xaguilar-amd, @xiaohongchen1991, @varun-sundar-rabindranath, @gau-nernst, @tahsintunan, @GirasoleY, @hclsys, @Yejing-Lai, @LucasWilkinson, @matteso1, @akii96, @atalman, @lucianommartins, @I3eg1nner, @rahulssv-ibm, @ZichenYuan, @tanpinsiang, @hillelda, @Srinivasoo7, @Etelis, @Rukhaiya2004, @Oxygen56, @Priyjain-amd, @GuyStone, @nholmber, @CienetStingLin, @xinyu-intel, @JartX, @esmeetu, @hhhhhhhhhhhhhhhhho, @harsha20032020, @walterbm, @Acaciasama, @jessiewei7, @ashwin-phadke, @shivampr, @cyq1017, @kjiang249, @orestis-z, @xyang16, @tianmu-li, @mgehre-amd, @aaarkai, @guybd, @wcynb1023, @Josephasafg, @qyYue1389, @russellb, @haoyangli0109, @sfeng33, @mikekg, @EanWang211123, @ovidiusm, @ItsMatti4, @hyeongyun0916, @qli88, @juliendenize, @calvarado2004, @tdoublep, @brandonpelfrey, @davispuh, @weizhoublue, @jasonozuzu-cohere, @wentian-byte, @skajre, @gty111, @omirosh, @decarpentierg, @fjosw, @ilmarkov, @yuwenzho, @JisoLya, @JohnLangford, @aldenlobo, @bnellnm, @jasonlizhengjian, @zufangzhu, @izhuhaoran, @MatthewBonanni, @deng451e, @ashwing, @sriganesh123, @linitra24, @liranschour, @umarkovi-amd, @aman0603, @adobrzyn, @jwzheng96, @eicherseiji, @ArsalanShakil, @tc-mb, @imargulis, @fangyuchu, @puririshi98, @JeanPaulShapo, @VectorPeak, @tarjan1, @qiching, @Achyuthan-S, @ZJY0516, @lucifer1004, @cinnamonica02, @jmamou, @almayne, @hao-aaron, @Jyothirmaikottu, @andylolu2, @AIvashov, @stevenkuang-tencent, @lcskrishna, @Aneureka, @wan-danfeng, @chengzheng345, @pranavthakur0-0, @zRzRzRzRzRzRzR, @DanBlanaru, @adamkbaranowski, @wendyliu235, @eparshut, @yangyang-cs95, @kalyanamdewri, @maxdebayser, @fenghourun, @tpopp, @okorzh-amd, @labAxiaoming, @sychen52, @ekagra-ranjan, @gausah01, @yuyue0225sc, @cpersson-amd, @lslusarczyk, @alex101-ops, @Zhenzhong1, @velonica0, @zhongjing123, @zhou9402, @llsj14, @majian4work, @akinsella, @BadrBasowid, @afierka-intel, @ayush1399, @LiJzd, @jesco-absolut, @Laurent-Zhang, @Kevin-XiongC, @NathanielMcVicar, @askliar, @ACEEE-1222, @jinzhen-lin, @SherryC41, @simondanielsson, @nv-nedelman-1, @yisustc, @kylesayrs, @jialoop-git, @NicolasHug, @guan404ming, @HumphreySun98, @danielafrimi, @gcanlin, @robertgshaw2-redhat
Your coding agent can read these notes before it upgrades. Set up the MCP server →