NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2222 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release 4 days ago
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
Dependencies: Upgrade Starlette to ≥ 1.0.1 to fix CVE-2026-48710 (#45675).
This release features 571 commits from 256 contributors (77 new)!
next_n > 2 on SM100 (#45322). It is now enabled on SM120 alongside GLM-5.1 (#43477), with XPU (#44144, #44517, #45240) and ROCm (#44899, #45103, #45681) attention/MoE paths added./tokenize + /detokenize (#44222), /pause /resume /is_paused (#44499), /abort_requests (#44382), /get_world_size (#44801), thinking_token_budget (#46137), a Python bridge for Rust tool parsers (#44624), and many new parsers and validation paths.CUDA_VISIBLE_DEVICES internally; a new device_ids argument is provided instead (#45026). On ROCm, a deprecation window for CUDA_VISIBLE_DEVICES has begun (#46636).mm_prefix support (#42175); many parser/serving fixes — forced-JSON skip for required/named tool choice (#45795), parsing with thinking disabled (#45832), streaming reasoning-state init (#45852), reasoning rendering on assistant turns (#45867), offline-parser truncation/token-leak fix (#45553); legacy Gemma4 parsers replaced with an engine-based implementation (#45588).num_heads_q (#45564), LoRA warmup fix (#35536), more accurate FP32 Gumbel sampling (#45996), min_tokens off-by-one fix in the V2 GPU sampler (#46243), plus assorted model/config compatibility fixes (#45868).ParallelLoader for weight loading (#40183), release of cached device memory under pressure on UMA GPUs (#45179), structured outputs for beam search (#35022), device_ids arg / no internal CUDA_VISIBLE_DEVICES (#45026), graceful fallback when numactl --membind is blocked (#45438), config-class registration before tokenizer init (#40299), async scheduling with prompt embeds for multimodal models (#45673).P2pNcclConnector (#44854).fused_moe FP8 for Qwen3-Next-80B on H100 (+25%) (#44830), native DSA indexer decode on SM100 (#45322), cluster-cooperative topK for DeepSeek low-latency (#43008), PDL support for DeepGEMM (#46006), FlashInfer cutedsl NVFP4 GEMM (#42235) and cute-dsl MXFP8 linear kernel (#46393), new Helion kernels for FP8/RMSNorm quant (#36902, #33790, #36895, #34432)._C library migration [12/n] (#45415).CUDA_VISIBLE_DEVICES on ROCm (#46636).VLLM_TRITON_FORCE_FIRST_CONFIG to skip Triton autotuning (#42425), Triton recompile detection (#45631), fused multi-group block-table staged writes (#44944).modelopt_mixed support extended to Ampere/SM80-86 (#45306) and Turing/SM75 (#45375).flashinfer_cutlass allowed as a clamped NVFP4 MoE backend (#46492), NVFP4/OCP MX MoE emulation fix (#46254), FP8 MoE re-enabled on NVIDIA Thor (#46339).fp8_e5m2 KV cache allowed for non-fp8 checkpoints (#45040)./v1/embeddings support for messages + chat_template_kwargs (#45173), multimodal token counts in usage.prompt_tokens_details (#45458), omit empty tool_calls from chat responses (#44105), Responses API streaming function_call id fix (#44608), Harmony refactor of streaming/non-streaming paths (#45171, #45104)./v1/messages (#40912), mid-conversation system-message handling (#46025), inline system-message position preserved for prefix caching (#44602), tool_use argument-dropping fix (#45287)./tokenize + /detokenize (#44222), /pause /resume /is_paused (#44499), /abort_requests (#44382), /get_world_size (#44801), thinking_token_budget (#46137), parallel_tool_calls=false (#44760), continuous usage stats (#43965), model metadata in /v1/models (#45950), Python bridge for Rust tool parsers (#44624), dedicated runtime for HTTP/ZMQ (#46051), and many validation/correctness fixes.vllm:tool_call_parser_invocations_total (#44448), group-aware KV cache capacity in vllm:cache_config_info (#42206), MLA attention metrics for DeepSeek MFU estimation (#39457)./v2/embed input exclusivity (#45640), non-negative rerank top_n (#46119), matryoshka embedding dimension bounds (#46313).vllm bench serve (#42457), multi-turn benchmark api_key/custom headers (#44516), tokenizer-mismatch auto-correction (#44708).This release ships another coordinated security-hardening batch (much of it from security researcher @jperezdealgaba).
prompt_embeds on M-RoPE models (#45252), regex-compilation timeout guard in structured outputs (#45118), audio upload size limit before full materialization (#45510), audio decode duration limit in the chat-completions path (#45908).temperature/repetition_penalty (#45116), sanitize_message applied to Anthropic and STT error paths (#45119).mistral_common is now optional via deferred import (#45305); CUDA Dockerfiles upgraded from GCC 10 to GCC 12 for C++20 (#44923); spinloop extension skipped on Python < 3.11 (#44783).CUDA_VISIBLE_DEVICES on ROCm (#46636); general deprecations for v0.23/v0.24 (#44992).Thank you to everyone who made this release possible!
@yewentao256, @Sunt-ing, @jperezdealgaba, @AndreasKaratzas, @BugenZhao, @sfeng33, @njhill, @micah-wil, @bbrowning, @mgoin, @jeejeelee, @hmellor, @tlrmchlsmth, @xianbaoqian, @mmangkad, @jikunshang, @Dao007forever, @zhenwei-intel, @noooop, @Isotr0py, @ivanium, @reidliu41, @varun-sundar-rabindranath, @chaunceyjiang, @WoosukKwon, @mawong-amd, @zxd1997066, @chaojun-zhang, @NickLucche, @bigPYJ1151, @ZJY0516, @charlifu, @yzong-rh, @divakar-amd, @khluu, @cleonard530, @wseaton, @xiaohongchen1991, @ywang96, @taneem-ibrahim, @mikekg, @itayalroy, @Alex-ai-future, @sahilsGit, @bnellnm, @littlecircle0730, @majian4work, @ricky-chaoju, @ronensc, @Fangzhou-Ai, @lucianommartins, @Srinivasoo7, @zyongye, @Rohan138, @Etelis, @wentian-byte, @ekagra-ranjan, @LucasWilkinson, @tahsintunan, @waynehacking8, @gau-nernst, @tuukkjs, @stefankoncarevic, @Palaiologos1453, @lucifer1004, @jmamou, @liulanze, @Terrencezzj, @Change72, @LopezCastroRoberto, @he-yufeng, @benchislett, @juliendenize, @s3woz, @panpan0000, @ilmarkov, @zixi-qi, @wcynb1023, @fynnsu, @ZhanqiuHu, @yuwenzho, @tdoublep, @MatthewBonanni, @hickeyma, @majunze2001, @mrn3088, @Yejing-Lai, @vllmellm, @Saddss, @DarkLight1337, @hongxiayang, @m4r1k, @qli88, @jonathanc-n, @felix0080, @djramic, @aoshen02, @fxmarty-amd, @simon-mo, @llsj14, @akii96, @walterbm, @dmaniloff, @zlxi02, @grYe99, @jeffye-dev, @parthash0804, @qyYue1389, @sagearc, @maeehart, @TanNgocDo, @cinnamonica02, @zucchini-nlp, @tykow, @mganczarenko, @yangdian96, @jimmy-evo, @YellowFoxH4XOR, @yzhan1, @shenoyvvarun, @yufufi, @laviier, @xiaohuguo2023, @EanWang211123, @JartX, @shantipriya-amd, @askliar, @hallerite, @appleparan, @effi-ofer, @angelayi, @TheCodeWrangler, @DanBlanaru, @ankrovv, @velonica0, @pjdurden, @cyyever, @wjinxu, @kliukovkin, @x41lakazam, @Jasen2201, @r-barnes, @tc-mb, @nataliepjlin, @KaletoAI, @WineChord, @fangyuchu, @vraiti, @nascheme, @jjppp, @sasindharan, @xiaguan, @snadampal, @chfeng-cs, @thillai-c, @guan404ming, @sridhar-3009, @vincentzed, @j-i-l, @rjrock, @abinggo, @anony-mous-e, @Achyuthan-S, @Harry-Chen, @mfylcek, @amd-asalykov, @noa-neria, @maobaolong, @TheEpicDolphin, @FAUST-BENCHOU, @martin-kukla, @xin3he, @ZiguanWang, @youkaichao, @factnn, @llx-08, @xx-thomas, @gitbisector, @Bortlesboat, @thisisjimmyfb, @JOSH1024, @wendyliu235, @wangxiyuan, @shen-shanshan, @HanHan009527, @amd-lalithnc, @netanel-haber, @fuscof-ibm, @AjAnubolu, @carlyou, @abcd1927, @CienetStingLin, @kouroshHakha, @alexbi29, @jesse996, @sungsooha, @andakai, @cquil11, @nehmathe2, @liangel-02, @hello-args, @j9smith, @nikhilesh-csa, @ruocco, @oguzhankir, @yiliu30, @xaguilar-amd, @amirkl94, @danisereb, @wangjiaxin99, @shanjiaz, @Oseltamivir, @alexeldeib, @wzhao18, @coder3101, @lyd1992, @markmc, @ashishpatel26, @HumphreySun98, @ByteFlowing1337, @nv-nedelman-1, @JaredforReal, @sammshen, @okorzh-amd, @muhammadfawaz1, @vadiklyutiy, @JasonLi314, @SumanthRH, @Sirius29, @tjtanaa, @zhangshuoming990105, @amanchugh89, @umut-polat, @srajabos, @junkang1991, @pst2154, @WindChimeRan, @Zedong-Liu, @gq112, @sunnweiwei, @athrael-soju, @EazyReal, @Liangliang-Ma, @jinzhen-lin, @V-3604, @aarushjain29, @ZewenShen-Cohere, @Bot1822, @BowenBao, @MichaelCao0, @tanpinsiang, @QwertyJack, @nagisa-kunhah, @Meihan-chen, @robertgshaw2-redhat
…failure handling (#43659), scheduled-function deprecations (#43358).
Please note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
This release features 408 commits from 200 contributors (63 new)!
torch.compile (#43746, #43891), its attention and RoPE paths were refactored (#44569, #44262, #43926), and an XPU attention decode path was added (#42953).generate endpoint (#43779), dynamic LoRA endpoints (#43778), /version (#43854) and /server_info (#43942) endpoints, a server-router extension hook (#43774), request-ID headers (#43883), and many new tool parsers (InternLM2 #43481, hy_v3 #43872, Phi-4-mini #44213, Gemma4 #43850).on_new_request lifecycle hook (#43205).Parser.parse() interface (#44267), with the Responses parser migrated to it (#42977).fetch_audio for transformers≥5.10 (#44559).torch.compile (#43617), EVS for Qwen3-VL (#44205), GLM-5.1 PP loading (#42944), GLM-4.1V processor logits (#43575), GLM-4.6V video loader (#44417), OlmoHybrid init (#43846), HyperCLOVAX remote-code removal (#43860), Bailing-MoE rotary factor (#43770), Step3 PP residual KeyError (#37622), MiniCPM-V-4.6 video (#44509), MiniCPM-O audio unpadding (#38053), MiniCPM-V batched preprocessing (#44609), FunASR-Nano init (#44215), Cohere routing method (#44021), Kimi-K2.5 FlashInfer ViT metadata (#44493).extra_repr() for pooler classes (#44805), LoRA-adapter-name pooling fix (#44410), resettled generative scoring entrypoint (#44153), expanded pooler unit tests (#43818, #44471).max_seq_len for attention metadata (#43991), rejection-sampling acceptance-rate fix (#40651), KVConnector + PP cleanup (#43732), speculator-prefill warmup/capture (#44253).num_heads_q for drafts (#43543), EAGLE/MTP lookahead caching in the SWA prefix-cache mask (#44082).do_not_specialize (#43803), Qwen3.5 mixed prefill+decode split routing (#44700), MiniMax-M2 gate kernel (#38445).KVCacheSpec (#37505), scheduler_block_size threaded into KVCacheManager/Coordinator (#44165), max_concurrent_batches moved to VllmConfig (#44274), config validation rejecting 0/negative knobs (#43794, #44057, #44207), KV-cache scale boilerplate removed from weight loading (#43167).on_new_request) (#43205) and on_schedule_end() hook (#44206), token-offset selective offload (#39983), skip decode-phase blocks in CPU offload (#43797), page-size block alignment (#43689), Triton fast-path for small CPU→GPU swap_blocks_batch (#42212), stale sliding-window block fix (#42959).kv_both role deprecation cycle (#43874), Mooncake fixes (#43742, #44103, #42694), LMCache LMCacheMPConnector (#42865), EC connector shutdown API (#42423) and non-blocking lookup (#41627), KV-transfer tokens excluded from iteration_tokens_total (#43346).Fp8BlockScaledMM new_empty() optimization (#43677), TurboQuant shared dequant buffers (#40941), tuned selective_state_update for H200/RTX PRO (#44251), Inductor fast-path fallback for vLLM/AITER custom ops (#42129), Gemma RMS all-reduce fusion (#42646), NUMA auto-binding on DGX B300 (#43270).permute_cols for ROCm (#44674), blocks-first KV layout for AMD (#43660), N=5 wvSplitK for spec decode (#40687), MoRI connector improvements (#43303, #41751, #40344).block_fp8_moe (#42139), block-scaled W8A8 FP8 path (#39968), WNA16 oracle for GPTQ sym-int4 (#41426), rms_norm/act quant fusions (#43963), GDN-attention MTP (#43565), Triton selective-scan op (#43421), transparent sleep mode (#37149), CPU/tiering offloading on XPU (#36423), DeepSeek-V4 attention decode path (#42953).cpu_awq folded into awq_marlin (#43841), RISC-V RVV WNA16 helpers (#42730), fused GDN gated-delta-rule kernels (#43534), PowerPC SHM communicator (#43754), arm64 CI image (#41303)._has_module trial-import verification (#44035).supports_expert_map (#43108) and the inplace fused-experts mechanism (#43727).system_fingerprint field (#40537), streaming tool/function calling with required (#40700), chat_template_kwargs in Responses (#43761), developer-to-system conversion in the HF renderer (#43590), unstreamed tool-call-args streaming fix (#44348).Parser.parse() (#44267), Responses parser migrated to the unified interface (#42977), unstreamed tool-arg flush moved into the parser (#44017); new/fixed tool parsers — MiniCPM5 XML (#43175), Qwen3 XML JSON-args-first (#43243), DeepSeek DSML incremental streaming (#42879), first-args-chunk serializer fix (#42683), tool_choice="none" honored in streaming (#42752), null-tool-args crash fix (#43862).thinking_token_budget validation (#43402), GPT-OSS instruction rendering (#44330), Harmony stop_token_ids cleanup (#44009), consistent VLLMValidationError in chat/completion validators (#36254), consolidation of dev entrypoints (#44170) and online-serving utils (#44479).generate endpoint (#43779), dynamic LoRA endpoints (#43778), /version (#43854) and /server_info (#43942), server-router extension hook (#43774), --enable-request-id-headers (#43883), recursive tool-parameter conversion (#44299), include_reasoning=false (#44391), --language-model-only skips the multimodal processor (#44500), per-engine batch auto-abort (#44591), UTF-8 char-boundary detokenizer fix (#44620), HF chat-template fixes (#44311), cross-DP aggregation of is_sleeping/reset_prefix_cache (#43429); new tool parsers — InternLM2 (#43481), hy_v3 (#43872), Phi-4-mini JSON (#44213), Gemma4 (#43850).vllm bench serve (#39795), reasoning-model (thinking) benchmarking via --chat-template-kwargs (#44244).thinking_token_budget values (#43402), non-positive ParallelConfig integer knobs (#44057), zero-valued config fields (#43794), and out-of-range max_num_scheduled_tokens (#44207).libcublas-dev (#39855), and CUTLASS DSL cu13 install order (#45204).JAISLMHeadModel (#43784).kv_both role (#43874).Thank you to everyone who made this release possible!
@AndreasKaratzas, @WoosukKwon, @BugenZhao, @yewentao256, @hmellor, @khluu, @njhill, @sfeng33, @bnellnm, @vadiklyutiy, @NickLucche, @JartX, @lucianommartins, @cleonard530, @wzhao18, @yma11, @simondanielsson, @jeejeelee, @zyongye, @chaunceyjiang, @bigPYJ1151, @ronensc, @taneem-ibrahim, @LucasWilkinson, @MatthewBonanni, @mmangkad, @chunyang-wen, @yzong-rh, @JaredforReal, @zixi-qi, @Isotr0py, @noooop, @chaojun-zhang, @Xunzhuo, @ivanium, @zufangzhu, @DaoyuanLi2816, @CienetStingLin, @aoshen02, @akii96, @benchislett, @MengqingCao, @rshavitt, @kliuae, @omerpaz95, @willamhou, @Majid-Taheri, @micah-wil, @ricky-chaoju, @mikekg, @mgoin, @mayuyuace, @Etelis, @ilmarkov, @tlrmchlsmth, @UranusSeven, @bedeks, @izhuhaoran, @ZJY0516, @fadara01, @pschlan-amd, @wangxiyuan, @Oxygen56, @charlifu, @varun-sundar-rabindranath, @shen-shanshan, @TheEpicDolphin, @adobrzyn, @XuZhou26, @tjtanaa, @Terrencezzj, @zhejiangxiaomai, @ILikeIneine, @yubofredwang, @chfeng-cs, @ThibaultCastells, @linzm1007, @javierdejesusda, @meenchen, @zhewenl, @xyang16, @angelayi, @nholmber, @zhangtao2-1, @adityasingh2400, @sts07142, @jatseng-ai, @fallintoplace, @andakai, @he-yufeng, @ignaciosica, @JINO-ROHIT, @tonyliu312, @QwertyJack, @animeshtrivedi, @jzakrzew, @juliendenize, @zexplorerhj, @ruocco, @mgehre-amd, @jasonboukheir, @MaciejBalaNV, @JohnQinAMD, @huanghua1994, @rajkiranjoshi, @rasmith, @harshaljanjani, @ltd0924, @wdhongtw, @yintong-lu, @tianmu-li, @jikunshang, @JMonde, @MHYangAMD, @frida-andersson, @gau-nernst, @Wauplin, @czhu-cohere, @gagandhakrey, @nemanjaudovic, @Liangliang-Ma, @liulanze, @sphinx07, @aadwived, @nightcityblade, @umut-polat, @jeffreywang88, @wcynb1023, @zzt93, @shadeMe, @Dao007forever, @alec-flowers, @Krishnachaitanyakc, @orozery, @BWAAEEEK, @cinnamonica02, @albertoperdomo2, @Rukhaiya2004, @mfylcek, @shreyas269, @Gruner-atero, @TomerBN-Nvidia, @wjinxu, @IdoAtadTD, @xiaozcy, @brian-dellabetta, @zhenwei-intel, @adotdad, @Kartavyasonar, @lesj0610, @ECMGit, @cakeng, @william-rom, @qiching, @NolanHo, @andylolu2, @xwu-intel, @linitra24, @hoobnn, @Dymasik, @wanghenshui, @maobaolong, @oguzhankir, @Jie-Fang, @okorzh-amd, @Kevin-XiongC, @jiahanc, @garrygale, @dsikka, @QiliangCui2023, @wjabbour, @zvik, @tc-mb, @jwzheng96, @divakar-amd, @tushar00jain, @galletas1712, @hanlin12-AMD, @tuukkjs, @viiccwen, @Sunt-ing, @HueCodes, @tianyu-z, @adhithyamulticoreware, @rishitdholakia13, @effi-ofer, @Vikrantpalle, @walterbm, @devin-lai, @Yadan-Wei, @amd-fuweiy, @maeehart, @qyYue1389, @BramVanroy, @SunskyXH, @Holworth, @majian4work, @xaguilar-amd, @Rohan138
This release features 8 commits from 6 contributors (1 new)!
This release features 8 commits from 6 contributors (1 new)!
v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions.
fmin compatibility issue that broke initialization (0decac0d).OlmoHybridForCausalLM failing to initialise after the checkpoint changed rope_parameters from None to {"rope_type": None} (#43846).transformers >= 5.9.0): register the hyperclovax model_type so vLLM uses its vendored config instead of the stale auto_map (#43860).num_api_servers > 1 by excluding the Ray DP backend from the deferred (kernel-assigned) port allocation introduced in #42585 (#43864).flashinfer-jit-cache via --extra-index-url while it is quarantined on PyPI, fixing image builds (#44366).ImportError: libcudart.so.12 when importing nixl_ep on CUDA 13 images (#44266).@khluu, @vadiklyutiy, @aadwived, @shadeMe, @alec-flowers, @hmellor
Removed deprecated MLA prefill arguments (#42555).
This release features 459 commits from 230 contributors (63 new)!
vllm/models/deepseek_v4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385). A large set of fused kernels (MegaMoE, mhc, Q-norm, indexer, sparse MLA) and ROCm parity fixes landed alongside accuracy fixes (#42810, #43710).update_config (#42783), and shared KV-cache layers (#35045), plus many correctness fixes.hf_overrides docs (#42163); EXAONE-4.5 aligned with Transformers update (#42246).extract_hidden_states (#39949), non-MTP speculation for NemotronH (#43130), shared MTP weights in MRv2 (#42538).mm_projector dtype fix (#42081).anyOf/oneOf/$ref resolution re-land (#37831), shared coerce_to_schema_type across MiniMax-M2 / DeepSeek-V3.2 / Seed-OSS parsers (#43006, #43019, #43140).update_config (#42783), shared KV-cache layers (#35045), FP32 gumbel sampling (#41775), auto-fallback to MRv1 with connectors (#42955), logprob_token_ids correctness (#43125, #41761), prompt-logprobs size fix (#42778).reset_cache() (#41956), per-request tracking (#42507), store-deferral fix (#41945).ExpertMapManager (#41046), experts moved to experts/ (#42334), RoutedExperts alias for FusedMoE (#40735), EPLB refactoring for FusedMoE (#41055).head_dim=512 for FlashInfer TRTLLM attention (#38822), FlashInfer Blackwell GDN prefill (#40717), GDN prefill kernel for SM100 (#43273).compute_prefill_context / _v_up_proj optimizations (#42460, #42561), penalties Triton kernel (#40657), do_not_specialize in fused FP8 RoPE (#42849), FULL CUDA graph capture for TRITON_MLA decode (#42885).--cpu-distributed-timeout-seconds (#42968).X-data-parallel-rank header (#42330).torch.accelerator.synchronize() (#40733).quantization_config to use QuantKey with activation override (#41566), MoE W4A8 CT migrated to oracle (#42680), AWQ Marlin MoE onto modular WNA16 oracle (#42483), GPTQ consolidation (gptq_marlin → auto_gptq) (#38288).AuthenticationMiddleware path extraction (#43426).chat_template_kwargs support (#42272), message-merging fix (#42189), empty channel/recipient harmony fix (#35540).thinking_token_budget support (#42116) with inverted-condition fix (#41674); map reasoning_effort to enable_thinking (#43401).reasoning_content → reasoning (#42664), reworked fastokens integration (#43168), consolidated Speech-to-Text entrypoints (#42370, #42274), beam-search consolidation via BeamSearchMixin (#42946), score/rerank chat-template instructions (#42412)./v2 endpoints (#42594).PoolingOfflineMixin (#42267), split offline inference APIs/utils (#43553).manylinux_2_28 base (#41668).nvidia-cutlass-dsl to 4.5.2 (#42991, #43230, #43745); llguidance to 1.7 (#42150); triton_kernels downgraded to v3.5.1 for gpt-oss (#43135).setuptools-rust dependency (#43287, #43377), pinned protoc in rust-build stages (#43292).vllm-openai target (#40275), build mooncake-transfer-engine from source (#42114), AINIC & Thor NIC support (#40453); Python-only installation made optional (#42293).humming MoE backend dependency added, reverted, then restored with CuPy runtime fix (#42540, #43492, #43530).get_tokenizer and resolve_hf_chat_template (#35024).--moe-backend / --linear-backend (#43148).@yewentao256, @haosdent, @njhill, @mgoin, @jeejeelee, @AndreasKaratzas, @NickLucche, @sfeng33, @noooop, @WoosukKwon, @khluu, @taneem-ibrahim, @Dao007forever, @vadiklyutiy, @bnellnm, @ivanium, @tjtanaa, @mmangkad, @hmellor, @DarkLight1337, @hickeyma, @zhenwei-intel, @jikunshang, @ronensc, @benchislett, @hao-aaron, @arpera, @zyongye, @gau-nernst, @frida-andersson, @ZhanqiuHu, @cleonard530, @akii96, @bedeks, @Isotr0py, @JasonKeyiL, @bigPYJ1151, @zhewenl, @weizhoublue, @zxd1997066, @gnovack, @chaojun-zhang, @majian4work, @chaunceyjiang, @pschlan-amd, @amitz-nv, @yma11, @dsikka, @tc-mb, @shanjiaz, @jperezdealgaba, @yzong-rh, @viktorpusTT, @TheEpicDolphin, @MatthewBonanni, @shen-shanshan, @hallerite, @zufangzhu, @bbrowning, @divakar-amd, @ianliuy, @esmeetu, @rasmith, @louie-tsai, @pmaybank, @liulanze, @ZJY0516, @TheDuyIT, @wzhao18, @jinzhen-lin, @BugenZhao, @ashwing, @fuergaosi233, @hqhq1025, @shaharmor98, @pisceskkk, @lkm2835, @noa-neria, @Rohan138, @whx-sjtu, @vrdn-23, @alexagriffith, @Flink-ddd, @jeffreywang-anyscale, @skyloevil, @ymoslem, @Lucaskabela, @kg6-sleipnir, @woernfl, @tdoublep, @GOavi101, @jmamou, @PeaBrane, @KaivalyaMDabhadkar, @BWAAEEEK, @MrZ20, @afierka-intel, @JoursBleu, @hissu-hyvarinen, @mwawrzos, @CynicDora, @NoeliaBentancor, @johncalesp, @fynnsu, @fxmarty-amd, @walterbm, @liangel-02, @lgeiger, @he-yufeng, @abinggo, @KrxGu, @hks-9697-v2, @Sarah-Salah, @rebklee, @aoshen02, @haic0, @libinta, @Zhenzhong1, @xhx1022, @b-mu, @WindChimeRan, @tpopp, @charlifu, @chengyinie, @ricky-chaoju, @lyd1992, @daniel-devlab, @paulyu12, @bobofang11235, @laudney, @BadrBasowid, @maeehart, @PatchouliTIS, @chunxiaozheng, @blake-snc, @southfreebird, @rbrugaro-amd, @rasdani, @dusthunter, @qizzzh, @ProExpertProg, @qianlihuang, @alec-flowers, @JisoLya, @gaozihao-shy, @rishaps, @xyang16, @wendyliu235, @hlin99, @tianmu-li, @yuwenzho, @inisis, @kfirtoledo, @roikoren755, @liranschour, @vllm-agent, @blancsw, @netanel-haber, @BowenBao, @czhu-cohere, @amitport, @tuukkjs, @revit13, @ofirzaf, @qyYue1389, @junyanxu, @gracie-guo, @sagearc, @xinyu-intel, @yiwen101, @DomBrown, @tomeras91, @Dogacel, @maxdebayser, @fadara01, @Terrencezzj, @izikgo, @wangrui6, @kebe7jun, @rishitdholakia13, @j9smith, @meena-at-work, @dllehr-amd, @alexeldeib, @sonusflow, @lucianommartins, @AAISSJ, @DaoyuanLi2816, @zexplorerhj, @zhangxin81, @velonica0, @fuscof-ibm, @anishesg, @zhengluo-nv, @ylangtsou, @fangyuchu, @zx3xyy, @simondanielsson, @ruizhang99, @zixi-qi, @xwu-intel, @yufufi, @wdhongtw, @mrjunwan-lang, @wangxiyuan, @wasnertobias, @ilmarkov, @sychen52, @zhandaz, @russellb, @SandishKumarHN, @juhi10071998, @itayalroy, @djmmoss, @SumanthRH, @mayuyuace, @zhougit86, @meenchen, @lucifer1004, @popkart-EZ, @jzakrzew, @ffggs, @huanghua1994, @orozery, @danisereb, @rshavitt, @Yihuki, @QingZhou-YangHY, @Jie-Fang, @bbartels
Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389). Users should migrate to transformers v5.
This release features 367 commits from 202 contributors (49 new)!
transformers v4 support (#40389). Users should migrate to transformers v5.max reasoning effort (#40982), disaggregated serving fixes (#41957).hidden_act variant support (#40588), pipeline parallelism fix (#40786), MoE fixes (#41206, #41574, #41401), tool parser crash fix (#41991, #42188).logprob_token_ids support (#40559).rope_type checkpoint support (#41734).max_split_size_mb during model loading (#41268).AllPool.forward 51% faster (#41163), GPU<->CPU sync elimination in pooling (#41433) and attention (#41434), numpy zero-copy embedding serialization (#41681), multimodal processor skip for text-only (#41246), FlashInfer FP8 async TP fusion (#39505), NVFP4 all-gather GEMM fusion for AsyncTP (#41882), re-enable allreduce+RMS fusion for DP/PP (#41458), DeepSeek bf16→fp32 via torch.mm (#41300), persistent MLA for sparse backend (#41990), configurable safetensors checkpoint prefetch (#41499), fused mhc_post_pre kernel (#41536), 2D-grid W8W8 group quant kernel (#42153), relaxed memory ordering for KV cache swaps (#39306).required (#40700) and named tool/function choice (#41110), resubmitting output items with missing fields (#41355).system_fingerprint field in responses (#40537), prompt_embeds content part support (#40720), defer_loading and tool_reference support (#40190), rendered prompt text in chat completion response (#42052), tolerate empty content in forced tool choice (#40148)./start_weight_update and /finish_weight_update APIs (#39212).VLLM_SKIP_MODEL_NAME_VALIDATION env var (#34676), configurable model weights loading tracking (#41086), Triton JIT compilation monitor (#40137).@AndreasKaratzas, @haosdent, @khluu, @yewentao256, @stecasta, @mgoin, @Isotr0py, @hmellor, @chaunceyjiang, @jeejeelee, @noooop, @MatthewBonanni, @njhill, @zyongye, @yzong-rh, @ronensc, @NickLucche, @chaojun-zhang, @dzhengAP, @chfeng-cs, @TheEpicDolphin, @esmeetu, @wzhao18, @ZJY0516, @juliendenize, @kylesayrs, @fadara01, @Etelis, @tianmu-li, @arpera, @ekagra-ranjan, @orozery, @wxsIcey, @jikunshang, @izhuhaoran, @rasmith, @russellb, @Lucaskabela, @Harry-Chen, @alec-flowers, @pmaybank, @Terrencezzj, @hickeyma, @Baekpica, @itej89, @fxmarty-amd, @WoosukKwon, @juhi10071998, @sychen52, @baonudesifeizhai, @vllmellm, @johncalesp, @the-david-oy, @lucianommartins, @bittoby, @Dao007forever, @lyd1992, @yuwenzho, @lesj0610, @sfeng33, @micah-wil, @akii96, @yma11, @SoluMilken, @mmangkad, @SiluPanda, @ojhaanshika, @zhandaz, @bhoomit, @simon-mo, @msanft, @angelayi, @anthonsu, @artem-spector, @zhangxin81, @benoittgt, @joerowell, @yangrz7, @chelnnexy, @liangel-02, @walterbm, @rishitdholakia13, @SKRohit, @BugenZhao, @JaredforReal, @amd-lalithnc, @frgossen, @h-avsha, @DarkLight1337, @danisereb, @laithsakka, @Bortlesboat, @wangluochao902, @Rohan138, @hao-aaron, @puririshi98, @roikoren755, @heachary, @UranusSeven, @dsingal0, @ChenxiQ, @snadampal, @ilmarkov, @wendyliu235, @lequytra, @JisoLya, @LuisRobaina, @sniper35, @eicherseiji, @Yuyi-Ao, @raviguptaamd, @sungsooha, @ganyi1996ppo, @andylolu2, @FredericOdermatt, @ProExpertProg, @rbrugaro-amd, @mcsantiago, @hnt2601, @jinzhen-lin, @taneem-ibrahim, @tomeras91, @alex-jw-brooks, @Aktsvigun, @HanFa, @netanel-haber, @JasonKeyiL, @gshtras, @joa-stdn, @Seven-Streams, @JartX, @xuechendi, @BowenBao, @Akashcodes732, @jeffreywang-anyscale, @czhu-cohere, @zhewenl, @marvinzh, @Lidang-Jiang, @gcanlin, @whx-sjtu, @S1ro1, @liulanze, @Dhruvilbhatt, @laviier, @wi-adam, @aaab8b, @yuankaichen-amd, @ZhanqiuHu, @QwertyJack, @viktorpusTT, @divakar-amd, @starkwj, @benchislett, @jcyang43, @JLiu4Coding, @xy3xy3, @hongxiayang, @amd-mghanimi, @wenyili, @bigPYJ1151, @s-yanev, @AlonKejzman, @noobHappylife, @TomerBN-Nvidia, @MeganEFlynn, @liuzijing2014, @jbuchananr, @lokashrinav, @ssam18, @dllehr-amd, @gmagogsfm, @tpopp, @tjtanaa, @simondanielsson, @zhenwei-intel, @HiroakiMikami, @nholmber, @SumanthRH, @LucasWilkinson, @maeehart, @rishaps, @r-barnes, @gau-nernst, @Kermit-C, @tdoublep, @aoshen02, @Naveassaf, @wangxingran222, @cvan20191, @AbhiOnGithub, @abdulrahman-cohere, @jmamou, @Flink-ddd, @bnellnm, @hqhq1025, @gnovack, @wangxiyuan, @princepride, @jiahanc, @LCAIZJ, @ovidiusm
This release features 6 commits from 6 contributors (0 new)!
This release features 6 commits from 6 contributors (0 new)!
This is a small patch release with bug fixes for DeepSeek V4, gpt-oss, and Qwen3-VL
max_seq_len, fixing the MTP=1 hang on DeepSeek V4 (#41665, revert of #41605).hidden_dim_unpadded through the moe_forward fake op so MXFP4 works under torch.compile on v0.20.x (#42002, backport of #41646).@ywang96, @zyongye, @stecasta, @wzhao18, @Isotr0py, @khluu
This is a patch release on top of v0.20.0 primarily focused on DeepSeek V4 stabilization and performance improvements, along with several important bu
This is a patch release on top of v0.20.0 primarily focused on DeepSeek V4 stabilization and performance improvements, along with several important bug fixes.
VLLM_MULTI_STREAM_GEMM_TOKEN_THRESHOLD (#41526).cvt instruction for faster FP32->FP4 conversion (#41015).head_compute_mix_kernel) for optimized head computation (#41255).max_num_batched_token not being captured in CUDA graph (#40734).num_gpu_blocks_override not accounted for in max_model_len checks (#41069).expandable_segments around cumem memory pool (#40812).input_ids and expert_map args for Quark W4A8 GPT-OSS (#41165).@BugenZhao, @chaunceyjiang, @gau-nernst, @ghphotoframe, @Isotr0py, @jeejeelee, @khluu, @njhill, @Rohan138, @wzhao18, @youkaichao, @ywang96, @ZJY0516, @zixi-qi, @zyongye
…— XPU is no longer pinned to 2.10. This is a breaking change for environment dependency.
This release features 752 commits from 320 contributors (123 new)!
vllm/vllm-openai:v0.20.0 image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669). As a general rule of thumb, our CUDA version policy follows PyTorch's. We highly recommend to install vLLM with uv and use --torch-backend=cu129 if you are on CUDA 12.9.transformers>=5 (#30566), with vision-encoder torch.compile bypass (#30518) and continued v4/v5 compat fixes including PaddleOCR-VL image processor max_pixels (#38629), Mistral YaRN warning (#37292), and Jina ColBERT rotary inv_freq recompute (#39176).SharedFusedMoE removed (#35782), DefaultMoERunner split (#35326) and later combined back into MoERunnerBase (#40560), shared/fused expert output sum moved into MoERunnerBase (#35949), ZeroExpertFusedMoE in new framework (#35549), compressed_tensors_moe.py split (#38960), GPTQMarlinMoEMethod reworked with MK (#37990), XPU & CUTLASS MoE relocated to fused_moe/experts/ (#40568, #40574), make_expert_params_mapping renamed (#40671), MoE LoRA refactor (#40338), and MoE DP chunking removed (#39107).seq_lens_cpu GPU→CPU sync (#40654); cache InductorPass.hash_source (#39328); skip FX-graph deserialization on loading for faster warm compile (#40151); CUDAGraph memory profiling enabled by default for clearer startup memory accounting (#38284).mamba_ssm_cache_dtype=float32 with NemotronHNanoVLV2 auto-hook (#39032); new TP plan styles for the Transformers backend (#40467); GLM-5.1 fix on ROCm (#40763).donate_graph_module=True for standalone_compile (#39733), skip FX graph deserialization on loading (#40151), include Inductor & functorch configs in compile-cache key (#40627), respect TORCH_COMPILE_DISABLE at vLLM config level (#40715), disable Sequence Parallelism for piecewise compilation (#38373).concat_mla_q half-types only (#37892), batch-invariance-aware backend auto-selection (#40193), avoid seq_lens_cpu GPU→CPU sync (#40654).shutdown() on OffloadingConnector (#39182), request context passed through KV offload (#39185), sliding-window lookup (#36645), multi-group worker transfer (#38453), multi-KV-group lookup/load/store (#39401, #39402, #39403).VLLM_MEDIA_CACHE media URL caching (#37123), safe request abort when FSM fails to advance (#38663), KV connector prioritized over internal registry (#38301), CUDAGraph memory profiling on by default (#38284), shared-expert overlap restored (#39222), CONFIG_REGISTRY config-class lookup fix when on-disk model_type differs (#39554), workspace-resize GPU memory leak fix (#39226), SWA/chunked-local runtime admission capped to startup pool-sizing bound (#40946).selective_state_update (#36162).get_num_embed overhead reduced (#40143), request_id on FinishedRequestStats (#39710).--enable-vit-cuda-graph for VLM examples (#40580), default max_frames_per_batch auto-infer for ViT CG video (#40445), fused FP8 output quantization into merge_attn_states (#36518), batched KV-cache swap via cuMemcpyBatchAsync (#38460), sm_110 (Jetson Thor) added to CUDA 13.0 build targets (#39233).TritonInt8ScaledMMLinearKernel (#38501), fused_silu_mul_block_quant enabled (#38817), KV-cache shuffle for paged_attention_common (#32914), MLA decode output zero-fill removed in AITER (#37539), MLA dual RMS norm fusion pass for DeepSeek/Kimi-K2 (#39242, with older-AITer guard #40386), AITER MLA + Eagle3 spec decode (#39616), DFlash on ROCm (#39703), wvSplitK FP8 path for RDNA (#37712), GPU↔NUMA-node detection (#40015), non-causal attention in ROCM_ATTN (#40176), engine-shutdown GPU memory leak fix (#38503), score-correction-bias dtype cast for DeepSeek/Kimi-K2 (#39999).round_int8 for Intel Triton (#38825), MoE Triton in online FP8 quantization fix (#40109), current_platform.supports_fp8() updated for TritonExperts (#40132), NIXL import on XPU fix (#40430), fusion-pattern support disabled on XPU (#39789).cpu_attn (#38676), gelu in cpu_fused_moe (#38770), OMP replacement (#36487), BF16 GELU LUT on ARM (#37469), W4A16 Autoround on CPU (#38192), CPU affinity/memory mgmt refactor (#39781), IBM Z s390x torch 2.11 builds (#39910), faster exp routine for lower-precision dtypes (#38112), inter-node pipeline parallel fix (#40150), RISC-V multiple RVV VLEN targets (#39478), RISC-V platform detection fix (#40427), exp() input clamp to prevent NaN on CPU/RISC-V (#40428).InductorPass.hash_source cached (#39328), humming quantization kernel (#34556).TransferMetadata consolidation (#37341), Async EPLB synchronization refactor (#37601), asyncio infrastructure removed from Async EPLB (#40730), replica-selection bias fix in fused_moe router (#40810), Async EPLB integration test added (#40168).num_lmcache_extra_cached_token in KVTransferParams (#39843), offload all KV blocks during prefill in P/D (#40346), DP control bundle pinned to first GPU's node on Ray (#39167), FlashInfer NVLink MNNVL workspace sized to EP group (#40893).TpKVTopology + HeteroTPTransferConfig unified into TransferTopology (#39529), NIXL EP treated as batched experts in fused_moe (#40412).reshape_and_cache_flash (#37332), batch-invariant NVFP4 linear (#39322), FlashInfer CuteDSL batched-experts backend for NVFP4 MoE (#38251), special GptOssMxfp4MoeMethod (#39604), W4A8_FP8 MoE TP>1 correctness fix (#40310), NVFP4 CUTLASS MoE OOB-read fix for non-multiple-of-4/16 expert counts (#40351), RMS norm + quant fusion fix on DeepGEMM UE8M0 path for B200 (#40552), Gemma4 quantized MoE (#39045).CompressedTensorsW8A8Mxfp8) (#38815), CT W8A8 in Oracle structure (#39187), layerwise reloading of attention/KV quantized models (#38995), experts_int8 consolidated with FP8 online quant (#38463), MXFP8 online quant on the new frontend (#40152).current_platform.supports_fp8() updated for TritonExperts on XPU/ROCm (#40132).TritonInt8ScaledMMLinearKernel on ROCm (#38501), Quark W8A8 INT8 MoE inference (#36320).presence_penalty / frequency_penalty on Responses API (#38613), Responses API streaming migrated to unified parser (#38755), tool_choice / tools validation on Responses to match OpenAI (#40399), Mistral Grammar factory (#38150), multimodal support on /inference/v1/generate (#38405), max_tokens_per_doc in rerank (#38827), Generative Scoring (#34539), MaxSim re-enabled on GPU (#38620), chat_template_kwargs on Anthropic /v1/messages (#40125), auto-detection of reasoning_config when only reasoning_parser is set (#38214), reasoning parsers can access model config via adjust_request (#37848, #39027), effective chat-template kwargs passed to reasoning parsers (#40460), reasoning parsers expose reasoning_start_str/reasoning_end_str (#40566).logit_scale added to PoolerConfig (#39435), then renamed logit_bias/logit_scale → logit_mean/logit_sigma for affine score calibration (#39530) — breaking. LLM.reward deprecated; use LLM.encode instead (#40688).grpc.health.v1 health check for Kubernetes-native probes (#38016).<tool_call> as implicit reasoning end in Qwen3 (#35687), is_reasoning_end_streaming() override for GptOssReasoningParser (#35745), Mistral tool parser HF-tokenizer fix (#39294), Mistral pre-v11 tool parser trailing-output fix (#40531), Gemma4 streaming HTML duplication / JSON corruption / null-as-string fixes (#38909, #38992, #39114, #39679), HF tokenizer concurrent-borrow fix in tool parsers (#40059), HYV3ReasoningParser no longer mutates chat_template_kwargs (#40713).mm_kwargs with cache injection (#39502), PyAV video backend for concurrent decoding (#39986), custom video metadata for pre-extracted frame sequences (#40133), image+video mixed inputs (per prompt) for VLM examples (#40335), deepstack buffer optimized for Qwen3 multimodal (#40145), readonly multimodal processor warmup during renderer startup (#40797), mm_processor_kwargs forwarded in offline generate APIs (#40251), normalize malformed dict prompts that carry token IDs in prompt (#40339), hotwords for FunASR (#39674), bundle get_generation_prompt() params into SpeechToTextParams (#36268).--omni delegates to vLLM Omni (#40744); avoid eager import of mistral_common (#40043).LLM.chat (#39352), use_audio_in_video passable at vllm serve for nemotron-nano-vl (#38538), deferred imports save ~2s CLI startup (#40056), improved MM-input-too-long error message (#39409), warning when FP8 KV cache misses prefill query quant (#39752), clearer DCP error message (#28443), --model deprecation warning updated (#39518), Mimo reasoning/tooling parsers mapped (#40089), human-readable k/K/m/M… suffix in JSON CLI args (#40473).'align' cache mode for Mamba-based models when speculative decoding is enabled (#40454).SpecDecodeBaseProposer moved out of eagle.py (#40732); DSA + MTP IMA fix (#40772).download_bytes_from_url (#38482).resampy dependency dropped (#39524), librosa direct dependency dropped (#39079), pyav and soundfile moved to common requirements (#39997).vllm:prompt_tokens_recomputed removed (#38709); num_cached_tokens / num_external_computed_tokens replaced with PrefillStats (#37460).logit_bias/logit_scale → logit_mean/logit_sigma (#39530).LLM.reward deprecated, use LLM.encode (#40688); cprofile / cprofile_context deprecated (#39100); V0 accept output buffer deprecated (#39125).accept output buffer in attention (#39125), cprofile / cprofile_context (#39100), LLM.reward offline API (#40688).This is a patch release on top of v0.19.0 with Transformers v5.5.3 upgrade and bug fixes for Gemma4:
This is a patch release on top of v0.19.0 with Transformers v5.5.3 upgrade and bug fixes for Gemma4:
Deprecations: --calculate-kv-scales (#37201), score task (#37537), pooling multi-task support (#37956), reasoning_content message field removed (#3748…
This release features 448 commits from 197 contributors (54 new)!
transformers>=5.5.0. We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.--lora-target-modules to restrict LoRA to specific modules (#34984), language_model_only respected (#37375), Mistral3 fix (#36928), Qwen3.5 fix (#36976), out-of-tree ops replacement (#37181).--speculative-config (#37880), Eagle3 drafter quant_config propagation (#37280), Eagle3 norm_before_fc propagation (#38111)./v1/chat/completions/batch for batched chat completions (#38011).--lora-target-modules (#34984), -sc shorthand for --speculative-config (#38380).--calculate-kv-scales (#37201), score task (#37537), pooling multi-task support (#37956), reasoning_content message field removed (#37480).VLLM_MAX_N_SEQUENCES environment variable to enforce sequence limits (#37952).--disable-frontend-multiprocessing (#37612).This is a patch release on top of v0.18.0 to address a few issues:
This is a patch release on top of v0.18.0 to address a few issues:
Upgrade xgrammar for security fix (#36168).
CUBLAS_STATUS_INVALID_VALUE and had to use a workaround in v0.17.0, you can reinstall torch 2.10.0. PyTorch published an updated wheel that addresses this bug.This release features 445 commits from 213 contributors (61 new)!
--grpc flag (#36169), enabling high-performance RPC-based serving alongside the existing HTTP/REST interface.vllm launch render command (#36166, #34551) enables GPU-less preprocessing and rendering, allowing separation of multimodal preprocessing from GPU inference.--enable-ep-weight-filter CLI option (#37351) for faster EP model loading.--enable-ep-weight-filter for faster EP loading (#37351).--grpc flag for gRPC serving (#36169).vllm launch render for preprocessing-only serving (#36166), vllm launch for GPU-less preprocessing (#34551).--distributed-timeout-seconds (#36047), --attention-backend auto (#35738), reasoning_effort=none (#36238), PyTorch profiler schedule (#35240).trust_remote_code setting in NemotronVL and KimiK25 (#36192).kaldi_native_fbank made optional (#35996).This is a patch release on top of v0.17.0 to address a few issues:
This is a patch release on top of v0.17.0 to address a few issues:
PyTorch 2.10 Upgrade: This release upgrades to PyTorch 2.10.0, which is a breaking change for environment dependencies.
Known Issue: If you are on CUDA 12.9+ and encounter a CUBLAS_STATUS_INVALID_VALUE error, this is caused by a CUDA library mismatch. To resolve, try one of the following:
/usr/local/cuda) from LD_LIBRARY_PATH, or simply unset LD_LIBRARY_PATH.uv pip install vllm --torch-backend=auto.pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129 (change the CUDA version to match your system).This release features 699 commits from 272 contributors (48 new)!
--performance-mode Flag: A new --performance-mode {balanced, interactivity, throughput} flag (#34936) simplifies performance tuning for common deployment scenarios.count_tokens API (#35588), tool_choice=none (#35835), and streaming/image handling fixes.aiter package renamed to amd-aiter (#35198).min_tokens support with speculative decoding (#32642).--performance-mode {balanced, interactivity, throughput} (#34936), --moe-backend for explicit kernel selection (#33807), --language-model-only for hybrid models (#34120), --enforce-eager clarification (#34523).generation_config max_tokens treated as default not ceiling (#34063).Patch protobuf for CVE-2026-0994 (#34253).
Please note that this release was branch cut on Feb 8, so any features added to vLLM after that date is not included.
This release features 440 commits from 203 contributors (7 new)!
--disable-access-log-for-endpoints option (#30011).repo_id:quant_type syntax (#33371), DeepSeek ReasoningParser with thinking enabled by default (#33221), remove noisy CT warning (#33273), early tokenization validation (#31366), reasoning_content backward compatibility (#33635), only include Authorization header when OPENAI_API_KEY is set (#33488).reasoning_content message field (#33402).VLLM_ALL2ALL_BACKEND environment variable (#33535).v0.15.1 is a patch release with security fixes, RTX Blackwell GPU fixes support, and bug fixes.
v0.15.1 is a patch release with security fixes, RTX Blackwell GPU fixes support, and bug fixes.
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.15.0...v0.15.1
Metrics: Removed deprecated vllm:time_per_output_token_seconds metric - use vllm:inter_token_latency_seconds instead (#32661).
This release features 335 commits from 158 contributors (39 new)!
--async-scheduling now works with pipeline parallelism (#32359).--enable-prefix-caching --mamba-cache-mode align. Achieves ~2x speedup by caching Mamba states directly (#30877).StreamingInput objects while maintaining KV cache alignment (#28973).include_stop_str_in_output tuning (#32383), prompt_cache_key support (#32824).skip_special_tokens configuration (#32345).data_1/data_2 and queries/documents (#32577).avg_logprob and compression_ratio in verbose_json segments (#31059).--ssl-ciphers CLI argument (#30937).api_server_count based on dp_size (#32525), wheel variant auto-detection during install (#32948), custom profiler URI schemes (#32393).vllm:time_per_output_token_seconds metric - use vllm:inter_token_latency_seconds instead (#32661).Full Changelog: https://github.com/vllm-project/vllm/compare/v0.14.1...v0.15.0
This is a patch release on top of v0.14.0 to address a few security and memory leak fixes.
This is a patch release on top of v0.14.0 to address a few security and memory leak fixes.
Deprecated quantization schemes have been removed (#31688, #31285).
This release features approximately 660 commits from 251 contributors (86 new contributors).
Breaking Changes:
--no-async-scheduling.
Key Improvements:
--max-model-len auto (#29431): Automatically fits context length to available GPU memory, eliminating OOM startup failures.VLLM_LOG_MODEL_INSPECTION=1 or by simply printing the LLM object.logit_bias/allowed_token_ids/min_tokens support (#32163).
New Model Architectures:
LoRA Support Expansion:
Model Enhancements:
enable_thinking: false (#31788)CUTLASS MoE Optimizations:
Other Performance:
Hardware Configs:
Platform:
New Features:
--max-model-len auto (#29431)attention_config in LLM() (#30710)reasoning_effort parameter (#31956)Tool Calling:
CLI:
-ep for --enable-expert-parallel (#30890)--input-len (#30816)--enable-log-deltas (renamed) (#32020)--default-chat-template-kwargs (#31343)API:
/server_info env info (#31899)/embeddings continue_final_message (#31497)weights_only=True in torch.load (#32045)--enforce-eager (#31643)seed_everything deprecated (#31659)Full Changelog: https://github.com/vllm-project/vllm/compare/v0.13.0...v0.14.0
Additional protection for CVE-2025-62164 (#30649).
This release features 442 commits from 207 contributors (61 new contributors)!
Breaking Changes: This release includes deprecation removals, PassConfig flag renames, and attention configuration changes from environment variables to CLI arguments. Please review the breaking changes section carefully before upgrading.
drop_thinking logic (#30490), DeepSeek V3.2 top-k fix (#27568).selective_state_update spec decode (#29488).compile_ranges for selective kernel compilation (#24252).TRITON_MLA without prefix-caching (#29125).FULL_DECODE_ONLY CUDA graph (#30072), CPU backend support (#30062).group_topk kernel: 1.9% throughput, 2.1% TPOT improvement (#30159)/reset_prefix_cache (#27170), KV events (#28309), failure recovery config (#26813).AttentionConfig replaces VLLM_ATTENTION_BACKEND env var (#26315).encoding_format=bytes_only (#30249), multiple image/audio per request (#29988), tokenization_kwargs override (#29794).VLLM_ATTENTION_BACKEND replaced with --attention-backend (#26315)-O.xx flag (#29991)embed_input_ids/embed_multimodal fallbacks (#30458)merge_by_field_config (#30035, #30170), --convert reward → --convert embed (#30463)Full Changelog: https://github.com/vllm-project/vllm/compare/v0.12.0...v0.13.0
Breaking Changes: This release includes PyTorch 2.9.0 upgrade (CUDA 12.9), V0 deprecations including xformers backend, and scheduled removals - please…
This release features 474 commits from 213 contributors (57 new)!
Breaking Changes: This release includes PyTorch 2.9.0 upgrade (CUDA 12.9), V0 deprecations including xformers backend, and scheduled removals - please review the changelog carefully.
Major Features:
GPU Model Runner V2 (Experimental) (#25266): Complete refactoring of model execution pipeline:
max_model_len and num_kv_groupsPrefill Context Parallel (PCP) (Preparatory) (#28718): Partitions the sequence dimension during prefill for improved long-sequence inference. Complements existing Decode Context Parallel (DCP). See RFC #25749 for details.
RLHF Support: Pause and Resume Generation for Asynchronous RL Training (#28037).
KV Cache Enhancements: Cross-layer KV blocks support (#27743), KV cache residency metrics (#27793).
Audio support: Audio embeddings support in chat completions (#29059).
Speculative Decoding:
use_aux_hidden_states (#27688)Configuration: Flexible inputs_embeds_size separate from hidden_size (#29741), --fully-sharded-loras for fused_moe (#28761).
NVIDIA Performance:
AMD ROCm:
CPU:
Attention: FlashAttention ViT support, now default backend (#28763).
Long Context: Optimized gather_and_maybe_dequant_cache kernel for extremely long sequences (#28029).
Multi-NUMA: Enhanced NUMA functionality for systems with multiple NUMA nodes per socket (#25559).
Docker: Image size reduced by ~200MB (#29060).
model.load_weights (#26327).parallel_tool_calls param compliance (#26233)verbose_json and timestamp features for transcription/translation (#24209).SamplingParams (#28914).repo_id:quant_type syntax (#29137).-O0, -O1, -O2, -O3 allow trading startup time for performance, more compilation flags will be added in future releases (#26847)rope_parameters (#28542).Removed Parameters:
num_lookahead_slots (#29000)best_of (#29090)Deprecated:
xformers backend (#29262)seed=None (#29185)Scheduled Removals (will be removed in future release):
ParallelConfig's direct child EPLB fields (#29324)guided_* config fields (#29326)override_pooler_config and disable_log_requests (#29402)CompilationConfig.use_inductor (#29323)Other Breaking Changes:
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.11.1...v0.12.0
This release includes 4 bug fixes on top of v0.11.1:
This release includes 4 bug fixes on top of v0.11.1:
[V0 Deprecation][Models] Remove all V0 condition for mm embeddings merge (@Isotr0py #25331)
This release includes 1456 commits from 449 contributors (184 new contributors)!
Key changes include:
torch==2.9.0+cu129, enabling Inductor partitioning and landing multiple fixes in graph-partition rules and compile-cache integration.torch.compile: Generalized batch-invariant support across attention and MoE backends, with explicit support for DeepGEMM and FlashInfer on Hopper and Blackwell GPUs.--async-scheduling to be enabled by default in the next release./v1/messages endpoint, allowing users to interact with vllm serve using Anthropic-compatible clients.Detailed release notes will be updated in the next few days.
image_size for phi4_multimodal (@Renovamen #25796)get_input_embeddings_v0 (@DarkLight1337 #25857)vllm.worker and update according imports (@aarnphm #25901)VllmConfig from config/__init__.py to config/vllm.py (@hmellor #25271)logical_not (@zhoukezi #25925)vision_feature_select_strategy into resolve_visual_encoder_outputs (@DarkLight1337 #25938)inputs_embeds (@DarkLight1337 #25922)cos_sin_cache in Llama4VisionRotaryEmbedding (@cjackal #25889)__syncwarp on ROCM (@zhewenl #25996)v4.56.2 (@hmellor #24638)_apply_feature_select_strategy (@DarkLight1337 #26003)n=1 and n>1 (@kmaehashi #26005)merge_by_field_config for MM models (A-C) (@DarkLight1337 #26073)merge_by_field_config for MM models (D-F) (@DarkLight1337 #26076)vllm.entrypoints.openai.api_server entrypoint with vllm serve command (@DarkLight1337 #25967)gemm_afp4wfp4 failure on AMD (@zhewenl #26068)merge_by_field_config for MM models (G) (@DarkLight1337 #26117)FusedMoE support for the Transformers backend (@hmellor #22650)prompt field (@DarkLight1337 #26097)reshape_and_cache CUDA Kernel (@ZJY0516 #25955)merge_by_field_config for MM models (InternVL family) (@DarkLight1337 #26153)merge_by_field_config argument (@DarkLight1337 #26214)executor.apply_model (@DarkLight1337 #26215)_reqs_to_process leak on abort (@NickLucche #26012)merge_by_field_config for MM models (H-L) (@DarkLight1337 #26230)--skip-tokenizer-init with echo and return_token_ids (@DarkLight1337 #26238)ruff instead of yapf + isort (@hmellor #26247)yapf as it's no longer used (@hmellor #26251)fmt: on/off (@hmellor #26253)ruff pre-commit hooks version (@hmellor #26255)merge_by_field_config for MM models (Llava family) (@DarkLight1337 #26280)DotsOCR tensor type (@what-in-the-nim #26281)LLM.set_tokenizer (@DarkLight1337 #26333)VLLM_USE_V1 from docs and scripts (@DarkLight1337 #26336)merge_by_field_config for MM models (Ovis family) (@Isotr0py #26308)LRUCache into its own file (@DarkLight1337 #26342)VLLM_USE_V1 from tests (@DarkLight1337 #26341)test_mxfp4_moe.py to test_ocp_mx_moe.py (@fxmarty-amd #26364)vllm bench throughput (@DarkLight1337 #26395)vllm/config/__init__.py to only add classes and functions (@hmellor #26405)w8a8 (@yewentao256 #25293)vllm bench ... on CPU-only head nodes (@Aydin-ab #25283)current_platform.import_kernels (@NickLucche #26286)RMSNorm substitution for Transformers backend (@hmellor #26353)_validate_and_reshape_mm_tensor (@lgeiger #26426)QKVCrossParallelLinear implementation (@Isotr0py #26475)pad with cat for better performance (@lgeiger #26486)prev_sampled_token_ids_invalid_indices input batch field (@njhill #26514)GPUModelRunner._update_states() (@njhill #26508)cu_seqlens on CPU to remove (@lgeiger #26496)pre-commit hook versions (@hmellor #26591)mypy==1.18.2 (@hmellor #26596)is_reasoning_end (@chaunceyjiang #25735)Optional[x] -> x | None and Union[x, y] to x | y (@hmellor #26633)fast_pos_embed_interpolate (@lgeiger #26647)git blame (@hmellor #26690)multimodal.py config (@andycandy #26629)vllm/distributed (@yewentao256 #26593)max_transformers_version for Qwen-VL test (@DarkLight1337 #26792)typos to fix by default (@hmellor #26785)SupportsV0Only interface and update supported models docs (@DarkLight1337 #26783)VLLM_DISABLE_PAD_FOR_CUDAGRAPH (@yewentao256 #26743)vllm/executor (@yewentao256 #26845)isort and yapf ignores (@DarkLight1337 #26888)vllm.utils.func (@DarkLight1337 #26904)vllm.utils.async_utils (@DarkLight1337 #26913)VLLM_ALLREDUCE_USE_SYMM_MEM by Default (@yewentao256 #26925)utils submodules (@DarkLight1337 #26920)vllm.utils.collections (@DarkLight1337 #26990)HF_TOKEN (deprecates HUGGING_FACE_HUB_TOKEN) (@yankay #27020)set in the CLI generation (@hmellor #27031)has to is (@yewentao256 #27032)get_m_alignment_for_contiguous_layout to get_mk_alignment_for_contiguous_layout (@yewentao256 #26935)random-input-len / random-output-len (@yewentao256 #26834)vllm.utils.import_utils (@DarkLight1337 #27022)execute_model_with_error_logging() to be a ctx manager (@njhill #27060)get_input_positions in MRotaryEmbedding (@MengqingCao #27088)rst style double-backtick with md single-backtick (@hmellor #27091)PolyNorm layer (@Isotr0py #27110)test_failure more stable for batch invariance (@yewentao256 #27054)vllm.utils.mem_utils (@iAmir97 #27143)get_kv_cache_spec into AttentionLayerBase (@NickLucche #26587).contiguous() calls (@lgeiger #27106)vllm.utils (@Isotr0py #26908)test_lm_eval_accuracy_v1_engine[google/gemma-3-1b-it] (@LucasWilkinson #27111)StatLoggerBase implementations (@ptovam #22456)vllm.utils.network_utils (@iAmir97 #27164)kv_role instead of kv_disagg_role (@hyongtao-code #27166)Distributed Tests (4 GPUs) async_sched+ray test (@NickLucche #27195)modules_to_not_convert field (@Isotr0py #26909)apache-tvm-ffi for flashinfer (@hmellor #27262)level references not removed #26355 (@hmellor #27260)LLM(data_parallel_size=k) single-process DP Usage (@yewentao256 #27282)get_mrope_input_positions instance methods (@DarkLight1337 #27342)kv_output_aggregator collect size (@NickLucche #26734)merge_by_field_config=True with tensor schema support (@Isotr0py #27361)has_ functions in vllm.utils (@jonathanc-n #27372)vllm.utils.platform_utils.py (@jonathanc-n #27374)PatchEmbed's conv3d to linear layer (@Isotr0py #27418)encoding_format="bytes" (@DarkLight1337 #27467)cudagraph_capture_sizes related improvements (@fhl2000 #26016)enable_reasoning parameter (@yyzxw #27550)merge_by_field_config=False (@DarkLight1337 #27551)max_model_len when it is not specified by users (@shen-shanshan #27556)utils.counter and move utils.Device to engine (@DarkLight1337 #27588)LayerBlockType a Literal instead of Enum (@DarkLight1337 #27658)supports_torch_compile on generic nn.Module and demonstrate speedup on Qwen Vision model (@Lucaskabela #23207)rocm_unquantized_gemm_impl (@zhewenl #27605)tools/pre_commit (@DarkLight1337 #27657)MultiModalDataParser (@Isotr0py #27664)vllm bench sweep to CLI (@DarkLight1337 #27639)test_two_responses_with_same_prev_id test (@NickLucche #27745)http_address (@yewentao256 #27488)use_aot_compile should respect VLLM_DISABLE_COMPILE_CACHE (@BoyuanFeng #27698)assert self.batched_router_logits.size(-1) == full_router_logits.size(-1) Bug (@yewentao256 #27682)SharedStorageConnector (@njhill #27719)vllm/v1/core and vllm/v1/engine (@yewentao256 #27108)record_sleep_state logic in PrometheusStatsLogger if not in dev mode (@SumanthRH #27789)VLLM_DEEPEP_LOW_LATENCY_ALLOW_NVLINK (@yewentao256 #27750)SpeculativeConfig.enable_chunked_prefill (@njhill #27826)KeyeForConditionalGeneration (@tjtanaa #27895)Self (@DarkLight1337 #27918)Note truncated.
Metrics: V1 TPOT histogram (#24015), hidden deprecated gpu_ metrics (#24245), KV cache GiB units (#25204, #25479).
This release features 538 commits, 207 contributors (65 new contributors)!
Note: In v0.11.0 (and v0.10.2), --async-scheduling will produce gibberish output in some cases such as preemption and others. This functionality is correct in v0.10.1. We are actively fixing it for the next version.
xm.mark_step in favor of torch_xla.sync (#25254).cpu_attn.py:_run_sdpa_forward for better memory access by @ignaciosica in https://github.com/vllm-project/vllm/pull/24701--enable-log-outputs does not match the documentation by @kebe7jun in https://github.com/vllm-project/vllm/pull/24626_validate_and_reshape_mm_tensor by @lgeiger in https://github.com/vllm-project/vllm/pull/24742supports_kw by @lgeiger in https://github.com/vllm-project/vllm/pull/24773s3_utils type hints with BaseClient by @Zerohertz in https://github.com/vllm-project/vllm/pull/24825stop in reasoning content by @gaocegege in https://github.com/vllm-project/vllm/pull/14550kv_output_aggregator support heterogeneous by @LCAIZJ in https://github.com/vllm-project/vllm/pull/23917MultiModalConfig from config/__init__.py to config/multimodal.py by @hmellor in https://github.com/vllm-project/vllm/pull/24659HuggingFace -> Hugging Face in Integration with Hugging Face docs by @sergiopaniego in https://github.com/vllm-project/vllm/pull/24889is_flashmla_supported Check Error by @yewentao256 in https://github.com/vllm-project/vllm/pull/24774n_groups % tp_size == 0 by @tomeras91 in https://github.com/vllm-project/vllm/pull/24593SpeculativeConfig from config/__init__.py to config/speculative.py by @hmellor in https://github.com/vllm-project/vllm/pull/24904EngineCoreRequest arguments in tests and fix extra kwargs by @qthequartermasterman in https://github.com/vllm-project/vllm/pull/24987CpuGpuBuffer for block table tensors by @njhill in https://github.com/vllm-project/vllm/pull/24795AutoModelForVision2Seq by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25065cutlass_mla hang by @alexm-redhat in https://github.com/vllm-project/vllm/pull/24966MultiModalCache by @lgeiger in https://github.com/vllm-project/vllm/pull/25006sliding_window from text config in Gemma3 MM by @hmellor in https://github.com/vllm-project/vllm/pull/25085StructuredOutputsConfig from config/__init__.py to config/structured_outputs.py by @hmellor in https://github.com/vllm-project/vllm/pull/25153validate-config pre-commit check by @hmellor in https://github.com/vllm-project/vllm/pull/25157vllm bench serve by @ywang96 in https://github.com/vllm-project/vllm/pull/25138apply_grammar_bitmask() method from ModelRunner to structured output utils by @shen-shanshan in https://github.com/vllm-project/vllm/pull/21999conv1d metadata to GDN attn by @vadiklyutiy in https://github.com/vllm-project/vllm/pull/25105returned_lse not Defined issue by @yewentao256 in https://github.com/vllm-project/vllm/pull/25106PoolerConfig from config/__init__.py to config/pooler.py by @hmellor in https://github.com/vllm-project/vllm/pull/25181KVTransferMetrics and aggregation strategy by @NickLucche in https://github.com/vllm-project/vllm/pull/22188get_input_embeddings interface by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25242ModelConfig from config/__init__.py to config/model.py by @hmellor in https://github.com/vllm-project/vllm/pull/25252pip-compile pre-commit hook so it runs on MacOS by @hmellor in https://github.com/vllm-project/vllm/pull/25273MIN_BLOCK_PER_SM by @yewentao256 in https://github.com/vllm-project/vllm/pull/25193LLM.apply_model by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18465fast_pos_embed_interpolate by @ywang96 in https://github.com/vllm-project/vllm/pull/25337fast_pos_embed_interpolate by @Isotr0py in https://github.com/vllm-project/vllm/pull/25347MultiModalPlaceholderMap by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25366xm.mark_step in favor of ``torch_xla.sync` by @NickLucche in https://github.com/vllm-project/vllm/pull/25254mypy behave like a proper pre-commit hook by @hmellor in https://github.com/vllm-project/vllm/pull/25313CPUOffloadingSpec by @orozery in https://github.com/vllm-project/vllm/pull/24251clear_connector_metadata by @NickLucche in https://github.com/vllm-project/vllm/pull/25397per_block_cast_to_fp8 by @yewentao256 in https://github.com/vllm-project/vllm/pull/24611_set_default_args_v0 function by @Isotr0py in https://github.com/vllm-project/vllm/pull/25409compile_size is None case. by @jikunshang in https://github.com/vllm-project/vllm/pull/25433tie_word_embeddings by @Isotr0py in https://github.com/vllm-project/vllm/pull/25454GuidedDecodingParams by @hmellor in https://github.com/vllm-project/vllm/pull/25422reshape_and_cache_flash by @bringlein in https://github.com/vllm-project/vllm/pull/24503examples/spec_decode.py and prevent breaking Acceptance Length by @ekagra-ranjan in https://github.com/vllm-project/vllm/pull/24531VLLM_NVTX_SCOPES_FOR_PROFILING=1 to enable nvtx.annotate scopes by @coreylowman in https://github.com/vllm-project/vllm/pull/25501_chunk_cumsum_fwd_kernel by @tdoublep in https://github.com/vllm-project/vllm/pull/25197UnquantizedLinearMethod by @kylesayrs in https://github.com/vllm-project/vllm/pull/23036DeviceConfig, ObservabilityConfig, SpeechToTextConfig to their own files by @hmellor in https://github.com/vllm-project/vllm/pull/25564fail_on_warning for the docs build in CI by @hmellor in https://github.com/vllm-project/vllm/pull/25580--help for enhanced user experience by @hmellor in https://github.com/vllm-project/vllm/pull/24903BasevLLMParameter.torch_function calling disabled super() by @yewentao256 in https://github.com/vllm-project/vllm/pull/25613v1/outputs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25629is by @nicole-lihui in https://github.com/vllm-project/vllm/pull/25641video_grid_thw typing by @ywang96 in https://github.com/vllm-project/vllm/pull/25646guided_... API by @hmellor in https://github.com/vllm-project/vllm/pull/25615BasevLLMParameter.torch_function calling disabled super()" by @mgoin in https://github.com/vllm-project/vllm/pull/25681merge_by_field_config MM interface by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25676test_argsort_mm_positions by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25690InputPreprocessor by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25702get_model_architecture by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/25682MsgpackEncoder._encode_tensor to avoid hidden performance regressions by @qthequartermasterman in https://github.com/vllm-project/vllm/pull/25738dbo_decode_token_threshold always (and ignoring dbo_prefill_token_threshold) by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/25622Full Changelog: https://github.com/vllm-project/vllm/compare/v0.10.2...v0.11.0
Breaking Changes: This release includes PyTorch 2.8.0 upgrade, V0 deprecations, and API changes - please review the changelog carefully.
This release contains 740 commits from 266 contributors (97 new)!
Breaking Changes: This release includes PyTorch 2.8.0 upgrade, V0 deprecations, and API changes - please review the changelog carefully.
aarch64 support: This release features native support for aarch64 allowing usage of vLLM on GB200 platform. The docker image vllm/vllm-openai should already be multiplatform. To install the wheels, you can download the wheels from this release artifact or install via
uv pip install vllm==0.10.2 --extra-index-url https://wheels.vllm.ai/0.10.2/ --torch-backend=auto
--model-impl terratorch support.--safetensors-load-strategy for NFS based file loading acceleration (#24469), critical CUDA graph capture throughput fix (#24128), scheduler optimization for single completions (#21917), multi-threaded model weight loading (#23928), and tensor core usage enforcement for FlashInfer decode (#23214).openai<1.100 to unblock CI by @mgoin in https://github.com/vllm-project/vllm/pull/23118propose_draft_token_ids non-blocking for lower TTFT by @WoosukKwon in https://github.com/vllm-project/vllm/pull/23041/collective_rpc API endpoint by @22quinn in https://github.com/vllm-project/vllm/pull/23075--mm-encoder-tp-mode by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23190_convert_tokens_to_string_with_added_encoders by 13.7x by @misrasaurabh1 in https://github.com/vllm-project/vllm/pull/20413prompt_token_ids arg fallback in LLM.generate and LLM.embed by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18800MinPLogitsProcessor.update_states() by @njhill in https://github.com/vllm-project/vllm/pull/23401target and content for prompt updates by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23411tokenizer explicitly instead of binding to prompt update by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23542tolist to avoid unintentional copy ops blocking across different CUDA streams, improving disagg TTIT/TTFT by @liuzijing2014 in https://github.com/vllm-project/vllm/pull/22760docs/api/summary.md by @hmellor in https://github.com/vllm-project/vllm/pull/23637rocm tag by @vllmellm in https://github.com/vllm-project/vllm/pull/20988mkdocs build by @Zerohertz in https://github.com/vllm-project/vllm/pull/23649tests/kernels/quantization by @ZJY0516 in https://github.com/vllm-project/vllm/pull/23675BaseMultiModalProcessor by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23018_send_reconfig_message() in core_client.py by @njhill in https://github.com/vllm-project/vllm/pull/23127default_pooling_type interface by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23736get_attr_docs if generating help text by @hmellor in https://github.com/vllm-project/vllm/pull/23723SupportsMultiModalWithRawInput with SupportsMultiModal by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/23749mkdocs build (continued) by @Zerohertz in https://github.com/vllm-project/vllm/pull/23743torch.compile for dynamic rope models in Transformers backend by @hmellor in https://github.com/vllm-project/vllm/pull/23738VLLM_DISABLE_PAD_FOR_CUDAGRAPH to Avoid Hang Issue by @yewentao256 in https://github.com/vllm-project/vllm/pull/23595ReplicatedLinear for SequenceClassification head by @Isotr0py in https://github.com/vllm-project/vllm/pull/23836json_count_leaves utility function by @aditchawdhary in https://github.com/vllm-project/vllm/pull/23899Apertus and XIELU by @EduardDurech in https://github.com/vllm-project/vllm/pull/23068aiter to matching list of issue auto labeller for rocm tag by @vllmellm in https://github.com/vllm-project/vllm/pull/23942download_weights_from_hf more reliable by @hmellor in https://github.com/vllm-project/vllm/pull/23863multi_modal_uuids as multimodal identifiers. by @ywang96 in https://github.com/vllm-project/vllm/pull/23394<tool_call> format in streaming mode for XLAM Tool Parser by @DevonPeroutky in https://github.com/vllm-project/vllm/pull/22769transcriptions/translations endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/23735security-opt for docker/k8s by @panpan0000 in https://github.com/vllm-project/vllm/pull/24017routed_scaling_factor Double Mul Issue by @yewentao256 in https://github.com/vllm-project/vllm/pull/24119w4a8_mm_entry.cu by @yewentao256 in https://github.com/vllm-project/vllm/pull/23660PyNcclConnector by @panpan0000 in https://github.com/vllm-project/vllm/pull/24151tokenizers == 0.22.0 by @faaany in https://github.com/vllm-project/vllm/pull/24159custom_stat_loggers extend default logger list by @eicherseiji in https://github.com/vllm-project/vllm/pull/20952KVEventsConfig from config/__init__.py to config/kv_events.py by @hmellor in https://github.com/vllm-project/vllm/pull/24433vllm serve cli args documentatNote truncated.
This is a critical bugfix and security release:
This is a critical bugfix and security release:
Fix CUTLASS MLA Full CUDAGraph (#23200)
Limit HTTP header count and size (#23267): https://github.com/vllm-project/vllm/security/advisories/GHSA-rxc4-3w6r-4v47
Do not use eval() to convert unknown types (#23266): https://github.com/vllm-project/vllm/security/advisories/GHSA-79j6-g2m3-jgfw
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.10.1...v0.10.1.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →