NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2222 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release 5 days ago
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
Breaking change: Removed AQLM quantization support (#22943) - users should migrate to alternative quantization methods.
v0.10.1 release includes 727 commits, 245 committers (105 new contributors).
NOTE: This release deprecates V0 FA3 support and as a result FP8 kv-cache in V0 may have issues
reshape_and_cache_flash CUDA kernel (#22036), CPU transfer support in NixlConnector (#18293).pip install vllm[flashinfer] for flexible installation (#21959).Important: As part of the ongoing V0 engine cleanup, several breaking changes have been introduced:
--task with --runner and --convert options (#21470), deprecated --disable-log-requests in favor of --enable-log-requests for clearer semantics (#21739), renamed --expand-tools-even-if-tool-choice-none to --exclude-tools-when-tool-choice-none for consistency (#20544).SpecializedManager by @zhouwfang in https://github.com/vllm-project/vllm/pull/21407--expand-tools-even-if-tool-choice-none with --exclude-tools-when-tool-choice-none for v0.10.0 by @okdshin in https://github.com/vllm-project/vllm/pull/20544flashinfer to v0.2.8 by @cjackal in https://github.com/vllm-project/vllm/pull/21385cutlass_fp4_group_mm illegal memory access by @yewentao256 in https://github.com/vllm-project/vllm/pull/21465run-batch supports V1 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/21541site_url for RunLLM by @hmellor in https://github.com/vllm-project/vllm/pull/21564requirements/common.txt to run unit tests by @zhouwfang in https://github.com/vllm-project/vllm/pull/21572has_flashinfer_moe Import Error when it is not installed by @yewentao256 in https://github.com/vllm-project/vllm/pull/21634moe_align_block_size_triton by @yewentao256 in https://github.com/vllm-project/vllm/pull/21335torch.compile for bailing moe by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/21664--task with --runner and --convert by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/21470Ernie 4.5] Name Change for Base 0.3B Model by @vasqu in https://github.com/vllm-project/vllm/pull/21735metavar to list the choices for a CLI arg when custom values are also accepted by @hmellor in https://github.com/vllm-project/vllm/pull/21760dynamic_scaled_fp8_quant and static_scaled_fp8_quant by @yewentao256 in https://github.com/vllm-project/vllm/pull/21773_lazy_init() by @smarterclayton in https://github.com/vllm-project/vllm/pull/21472CompressedTensorsW8A8Fp8MoEMethod and CompressedTensorsW8A8Fp8MoECutlassMethod by @yewentao256 in https://github.com/vllm-project/vllm/pull/21775uv in GPU installation docs by @davidxia in https://github.com/vllm-project/vllm/pull/20277flashinfer_python to CUDA wheel requirements by @mgoin in https://github.com/vllm-project/vllm/pull/21389__nv_fp8_e4m3 instead of c10::e4m3 for per_token_group_quant by @yewentao256 in https://github.com/vllm-project/vllm/pull/21867ZE_AFFINITY_MASK for device select on xpu by @jikunshang in https://github.com/vllm-project/vllm/pull/21815per_token_group_quant by @yewentao256 in https://github.com/vllm-project/vllm/pull/21860concurrency by @hmellor in https://github.com/vllm-project/vllm/pull/21919get_required_kvcache_layout class method to kv connector api by @wxsms in https://github.com/vllm-project/vllm/pull/20433async_llm_streaming.py example for AsyncLLM streaming in python by @mgoin in https://github.com/vllm-project/vllm/pull/21763test_shared_storage_connector_hashes by @mgoin in https://github.com/vllm-project/vllm/pull/21973collective_rpc returns None by @njhill in https://github.com/vllm-project/vllm/pull/22006vllm[flashinfer] by @mgoin in https://github.com/vllm-project/vllm/pull/21959per_block_cast_to_fp8, Remove Dependencies of DeepGEMM by @yewentao256 in https://github.com/vllm-project/vllm/pull/21787speculators config support by @dsikka in https://github.com/vllm-project/vllm/pull/21345get_kwargs for case where type hint is list[Union[str, type]] by @hmellor in https://github.com/vllm-project/vllm/pull/22016uv in CPU installation docs by @davidxia in https://github.com/vllm-project/vllm/pull/22089--disable-log-requests and replace with --enable-log-requests by @hmellor in https://github.com/vllm-project/vllm/pull/21739ModelConfig.try_get_generation_config to prevent future confusion by @hmellor in https://github.com/vllm-project/vllm/pull/21526reshape_and_cache_flash CUDA Kernel by @yewentao256 in https://github.com/vllm-project/vllm/pull/22036VLLM_TARGET_DEVICE.lower() by @NickLucche in https://github.com/vllm-project/vllm/pull/22101vllm bench serve by @yeqcharlotte in https://github.com/vllm-project/vllm/pull/21696aiohttp connection pool for benchmarking by @eicherseiji in https://github.com/vllm-project/vllm/pull/21981store=True and process the request by default by @WoosukKwon in https://github.com/vllm-project/vllm/pull/22185KVConnectorBase code (1/2) by @lk-chen in https://github.com/vllm-project/vllm/pull/21785VLLM_NO_DEPRECATION_WARNING by @yewentao256 in https://github.com/vllm-project/vllm/pull/22199v4.55 by @hmellor in https://github.com/vllm-project/vllm/pull/21931kernel_unified_attention_2/3d caused by attention sinks by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/22368flashinfer-python==0.2.10 by @mgoin in https://github.com/vllm-project/vllm/pull/22389flash_attn_varlen_func interface on xpu by @jikunshang in https://github.com/vllm-project/vllm/pull/22350hf_xet pin to resolve hangs by @hmellor in https://github.com/vllm-project/vllm/pull/22356packed_modules_mapping to DeepseekV2ForCausalLM by @fxmarty-amd in https://github.com/vllm-project/vllm/pull/22352from_dict from SpeculativeConfig by @hmellor in https://github.com/vllm-project/vllm/pull/22451get_tensor_model_*_group by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/22494CompilationConfig from config.py by @hmellor in https://github.com/vllm-project/vllm/pull/22524--no-build-isolation by @tdoublep in https://github.com/vllm-project/vllm/pull/22541--help from the JSON tip by @hmellor in https://github.com/vllm-project/vllm/pull/22567ParallelConfig from config/__init__.py to config/parallel.py by @hmellor in https://github.com/vllm-project/vllm/pull/22565CacheConfig from config/__init__.py to config/cache.py by @hmellor in https://github.com/vllm-project/vllm/pull/22586vllm subcommands by @hmellor in https://github.com/vllm-project/vllm/pull/22601VLLM_USE_DEEP_GEMM_E8M0 Env to Control E8M0 Scale by @yewentao256 in https://github.com/vllm-project/vllm/pull/21968test_max_len.py to unblock CI by @tjtanaa in https://github.com/vllm-project/vllm/pull/22664hf_xet has been updated by @hmellor in https://github.com/vllm-project/vllm/pull/22666SpeculativeConfig from the CLI by @hmellor in https://github.com/vllm-project/vllm/pull/22652SchedulerConfig from config/__init__.py to config/scheduler.py by @hmellor in https://github.com/vllm-project/vllm/pull/22626test_remote_decode_lifecycle.py::test_short_prompt_lifecycle by @NickLucche in https://github.com/vllm-project/vllm/pull/22727SupportsEagle3 iNote truncated.
[V1][Metrics] Deprecate metrics with gpu_ prefix for non GPU specific metrics. by @sahelib25 in https://github.com/vllm-project/vllm/pull/18354
v0.10.0 release includes 308 commits, 168 contributors (62 new!).
NOTE: This release begins the cleanup of V0 engine codebase. We have removed V0 CPU/XPU/TPU/HPU backends (#20412), long context LoRA (#21169), Prompt Adapters (#20588), Phi3-Small & BlockSparse Attention (#21217), and Spec Decode workers (#21152) so far and plan to continued to delete code that is no longer used.
--async-scheduling flag to overlap engine core scheduling with GPU runner (#19970).get_tokenizer_info for tokenizer/chat-template information (#20575), cache_salt support for completions/responses (#20981).--help=page option for enhanced help documentation (#20961), default model changed to Qwen3-0.6B (#20335).MultiModalHasher.hash_prompt_mm_data by @lgeiger in https://github.com/vllm-project/vllm/pull/19422num_blocks check by @NickLucche in https://github.com/vllm-project/vllm/pull/19532enable_caching in connector-delayed kvcache load case by @njhill in https://github.com/vllm-project/vllm/pull/19435moe_permute_unpermute_kernel.inl by @yewentao256 in https://github.com/vllm-project/vllm/pull/19573num_heads/num_kv_heads divisibility check by @22quinn in https://github.com/vllm-project/vllm/pull/19339bad_words by @f14-bertolotti in https://github.com/vllm-project/vllm/pull/19564moe_align_block_size CUDA kernel by @yewentao256 in https://github.com/vllm-project/vllm/pull/19572InputBatch by @afeldman-nm in https://github.com/vllm-project/vllm/pull/19778vllm/triton_utils/importing.py in an extensible way by @afeldman-nm in https://github.com/vllm-project/vllm/pull/19783648764942e552a8bb5fe16026703716a81f05374 by @tjtanaa in https://github.com/vllm-project/vllm/pull/18990LLM.beam_search by @NekoMimiUnagi in https://github.com/vllm-project/vllm/pull/19301Value of type "SampleRequest" is not indexable by @b8zhong in https://github.com/vllm-project/vllm/pull/18032tp_size > num_kv_heads deployments by @NickLucche in https://github.com/vllm-project/vllm/pull/19691ReqId and EngineId for better readability by @lk-chen in https://github.com/vllm-project/vllm/pull/19880attn_temperature_tuning by @b8zhong in https://github.com/vllm-project/vllm/pull/19997ceil_div by @yewentao256 in https://github.com/vllm-project/vllm/pull/20023/v1/audio/translations OpenAI API endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/19615node_count function by @njhill in https://github.com/vllm-project/vllm/pull/20045tp_size exchange with rank0 by @NickLucche in https://github.com/vllm-project/vllm/pull/19413CUDA_VISIBLE_DEVICES='' in Platform.device_id_to_physical_device_id by @eicherseiji in https://github.com/vllm-project/vllm/pull/18979compressed-tensors by @dsikka in https://github.com/vllm-project/vllm/pull/20033v1/audio/transcriptions endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/20179has_deepgemm, has_deepep, has_pplx by @yewentao256 in https://github.com/vllm-project/vllm/pull/20187CachedRequestData Instance Across All Requests by @WoosukKwon in https://github.com/vllm-project/vllm/pull/20232-O/--compilation-config by @ProExpertProg in https://github.com/vllm-project/vllm/pull/20156is_global_first_rank by @mgoin in https://github.com/vllm-project/vllm/pull/19516numel() downcast in vllm/csrc/moe/moe_align_sum_kernels.cu +2 by @r-barnes in https://github.com/vllm-project/vllm/pull/17082hf_to_vllm_mapper by @kylesayrs in https://github.com/vllm-project/vllm/pull/20046stream=True by @NickLucche in https://github.com/vllm-project/vllm/pull/20271find_free_port by @yewentao256 in https://github.com/vllm-project/vllm/pull/20333VLLM_ENABLE_MOE_ALIGN_BLOCK_SIZE_TRITON by @yewentao256 in https://github.com/vllm-project/vllm/pull/20334tok_kwargs (#20058) by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20353test_streaming_response test by @NickLucche in https://github.com/vllm-project/vllm/pull/20363Unable to detect current VLLM config. Defaulting to NHD kv cache layout warning by @NickLucche in https://github.com/vllm-project/vllm/pull/20400benchmark_moe.py + Add Triton Fused MoE kernel config for FP8 E=16 on B200 by @b8zhong in https://github.com/vllm-project/vllm/pull/20516use_cross_encoder flag to use correct activation in ClassifierPooler by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20527gh-pr and gh-issue everywhere we can in the docs by @hmellor in https://github.com/vllm-project/vllm/pull/20564code and console admonitions so readers are less likely to miss them by @hmellor in https://github.com/vllm-project/vllm/pull/20585get_language_model for Keye-VL by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20631get_max_tokens_per_item for backward compatibility by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20630multiply_add with homogeneous_multiply_add to Address Clang Template Parameter Issue by @wenxin0319 in https://github.com/vllm-project/vllm/pull/20142Platform.set_device method by @jikunshang in https://github.com/vllm-project/vllm/pull/20262reasoning_content is None when Thinkng is enabled and tool_choice is set to 'required'. by @chaunceyjiang in https://github.com/vllm-project/vllm/pull/20662inc.md file by @hmellor in https://github.com/vllm-project/vllm/pull/20697--compress-mode to reduce binary size by 30% by @mgoin in https://github.com/vllm-project/vllm/pull/20694tool_choice='required' and $defs. by @chaunceyjiang in https://github.com/vllm-project/vllm/pull/20629VllmConfig() construction on all platforms by @njhill in https://github.com/vllm-project/vllm/pull/20695clear_metadata() at end of step by @njhill in https://github.com/vllm-project/vllm/pull/20756/invocations to be task-agnostic by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20764no_use_tqdm_on_load flag while capturing CUDA graph by @lk-chen in https://github.com/vllm-project/vllm/pull/20834Note truncated.
Metrics: Deprecate metrics with gpu_ prefix for non GPU specific metrics (#18354), Export NaNs in logits to scheduler_stats if output is corrupted
This release contains 452 commits from 167 contributors (31 new!)
NOTE: This is the last version where V0 engine code and features stay intact. We highly recommend migrating to V1 engine.
CachedRequestData objects and cached sampler‑ID stores deliver perf enhancements (#20232, #20291)./v1/audio/translations & revamped /v1/audio/transcriptions (#19615, #20179, #19597).LLM.beam_search and cached template‑resolution speed‑ups (#19301, #20065).llm.chat, tool‑choice expansion, and custom‑arg passthroughs enrich multi‑modal agents (#19635, #17177, #16862).-O/--compilation-config, batch‑size‑sweep benchmarking, richer --help, faster startup (#20156, #20516, #20430, #19941).MultiModalHasher.hash_prompt_mm_data by @lgeiger in https://github.com/vllm-project/vllm/pull/19422num_blocks check by @NickLucche in https://github.com/vllm-project/vllm/pull/19532enable_caching in connector-delayed kvcache load case by @njhill in https://github.com/vllm-project/vllm/pull/19435moe_permute_unpermute_kernel.inl by @yewentao256 in https://github.com/vllm-project/vllm/pull/19573num_heads/num_kv_heads divisibility check by @22quinn in https://github.com/vllm-project/vllm/pull/19339bad_words by @f14-bertolotti in https://github.com/vllm-project/vllm/pull/19564moe_align_block_size CUDA kernel by @yewentao256 in https://github.com/vllm-project/vllm/pull/19572InputBatch by @afeldman-nm in https://github.com/vllm-project/vllm/pull/19778vllm/triton_utils/importing.py in an extensible way by @afeldman-nm in https://github.com/vllm-project/vllm/pull/19783648764942e552a8bb5fe16026703716a81f05374 by @tjtanaa in https://github.com/vllm-project/vllm/pull/18990LLM.beam_search by @NekoMimiUnagi in https://github.com/vllm-project/vllm/pull/19301Value of type "SampleRequest" is not indexable by @b8zhong in https://github.com/vllm-project/vllm/pull/18032tp_size > num_kv_heads deployments by @NickLucche in https://github.com/vllm-project/vllm/pull/19691ReqId and EngineId for better readability by @lk-chen in https://github.com/vllm-project/vllm/pull/19880attn_temperature_tuning by @b8zhong in https://github.com/vllm-project/vllm/pull/19997ceil_div by @yewentao256 in https://github.com/vllm-project/vllm/pull/20023/v1/audio/translations OpenAI API endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/19615node_count function by @njhill in https://github.com/vllm-project/vllm/pull/20045tp_size exchange with rank0 by @NickLucche in https://github.com/vllm-project/vllm/pull/19413CUDA_VISIBLE_DEVICES='' in Platform.device_id_to_physical_device_id by @eicherseiji in https://github.com/vllm-project/vllm/pull/18979compressed-tensors by @dsikka in https://github.com/vllm-project/vllm/pull/20033v1/audio/transcriptions endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/20179has_deepgemm, has_deepep, has_pplx by @yewentao256 in https://github.com/vllm-project/vllm/pull/20187CachedRequestData Instance Across All Requests by @WoosukKwon in https://github.com/vllm-project/vllm/pull/20232-O/--compilation-config by @ProExpertProg in https://github.com/vllm-project/vllm/pull/20156is_global_first_rank by @mgoin in https://github.com/vllm-project/vllm/pull/19516numel() downcast in vllm/csrc/moe/moe_align_sum_kernels.cu +2 by @r-barnes in https://github.com/vllm-project/vllm/pull/17082hf_to_vllm_mapper by @kylesayrs in https://github.com/vllm-project/vllm/pull/20046stream=True by @NickLucche in https://github.com/vllm-project/vllm/pull/20271find_free_port by @yewentao256 in https://github.com/vllm-project/vllm/pull/20333VLLM_ENABLE_MOE_ALIGN_BLOCK_SIZE_TRITON by @yewentao256 in https://github.com/vllm-project/vllm/pull/20334tok_kwargs (#20058) by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20353test_streaming_response test by @NickLucche in https://github.com/vllm-project/vllm/pull/20363Unable to detect current VLLM config. Defaulting to NHD kv cache layout warning by @NickLucche in https://github.com/vllm-project/vllm/pull/20400benchmark_moe.py + Add Triton Fused MoE kernel config for FP8 E=16 on B200 by @b8zhong in https://github.com/vllm-project/vllm/pull/20516use_cross_encoder flag to use correct activation in ClassifierPooler by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/20527Full Changelog: https://github.com/vllm-project/vllm/compare/v0.9.1...v0.9.2
Remove metrics that were deprecated in 0.8
This release features 274 commits, from 123 contributors (27 new contributors!)
scaled_fp8_quant by increasing vectorization (#18844)LLM API: make use_tqdm accept a callable for custom progress bars (#19357)model when initializing LLM (#18802)inputs arg fallback in Engine classes (#18799)Qwen2EmbeddingModel (#18913)get_dummy_text and get_dummy_mm_data (#18796)vllm bench serve and sync with benchmark_[serving,datasets].py by @mgoin in https://github.com/vllm-project/vllm/pull/18566get_dummy_text and get_dummy_mm_data by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18796async_timeout by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18792None for fields which should never be None by @hmellor in https://github.com/vllm-project/vllm/pull/17985model when initializing LLM by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18802base by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18914Qwen2EmbeddingModel by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18913packed_modules_mapping for VLM with arbitrary components by @Isotr0py in https://github.com/vllm-project/vllm/pull/18987WeightsMapper for qwen2-vl/qwen2.5-vl by @Isotr0py in https://github.com/vllm-project/vllm/pull/19054_Backend enums by @NickLucche in https://github.com/vllm-project/vllm/pull/19081scaled_fp8_quant by increasing vectorization by @mgoin in https://github.com/vllm-project/vllm/pull/18844Optional and Annotated in CLI typing by @hmellor in https://github.com/vllm-project/vllm/pull/19093compressed-tensors by @dsikka in https://github.com/vllm-project/vllm/pull/19217generate() to handle Generator exits by @Adolfo-Karim in https://github.com/vllm-project/vllm/pull/19225inputs arg fallback in Engine classes by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18799max_model_len in V0 by @22quinn in https://github.com/vllm-project/vllm/pull/19348kv_sharing_target_layer_name argument to cutlass_mla backend by @pavanimajety in https://github.com/vllm-project/vllm/pull/19374use_irope by @YUNQIUGUO in https://github.com/vllm-project/vllm/pull/19134Full Changelog: https://github.com/vllm-project/vllm/compare/v0.9.0...v0.9.1
This patch release contains important bugfix for DeepSeek family of models on NVIDIA Ampere and below
This patch release contains important bugfix for DeepSeek family of models on NVIDIA Ampere and below (#18807)
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.9.0...v0.9.0.1
vLLM has upgraded to PyTorch 2.7! (#16859) This is a breaking change for environment dependency.
This release features 649 commits, from 215 contributors (82 new contributors!)
pip install https://download.pytorch.org/whl/cu128/flashinfer/flashinfer_python-0.2.5%2Bcu128torch2.7-cp38-abi3-linux_x86_64.whl then set VLLM_ATTENTION_BACKEND=FLASHINFER for better performance.top_k to be disabled with 0 (still accept -1 for now) (#17773)0 by default for V1 Engine, meaning that different vLLM runs now yield the same outputs even if temperature > 0. This does not modify the random state in user code since workers are run in separate processes unless VLLM_USE_V1_MULTIPROCESSING=0. (#17929, #18741)transformers (from source) to use Falcon-H1.torchrun (#17827)VLLM_ALLOW_INSECURE_SERIALIZATION env var (#17490)deprecated=True (#17426)chat_template_kwargs in LLM.chat (#17356), /classify endpoint (#17032), truncation control for embedding models (#14776), cached_tokens in response usage (#18149)nvidia/DeepSeek-R1-FP4 (#16362), Quark MXFP4 format (#16943), AutoRound (#17850), torchao models with AOPerModuleConfig (#17826), CUDA Graph support for V1 GGUF support (#18646)--enable-reasoning (#17452)tool_choice: required for Xgrammar (#17845), Structural Tag with Guidance backend (#17333)--torch-backend=auto (#18505)vllm.multimodal (#18031)ruff format (#17656, #18068, #18400)numel() downcast in fused_layernorm_dynamic_per_token_quant.cu by @r-barnes in https://github.com/vllm-project/vllm/pull/17316'<string>' filepath by @zou3519 in https://github.com/vllm-project/vllm/pull/17330pre-commit autoupdate by @hmellor in https://github.com/vllm-project/vllm/pull/17380chat_template_kwargs in LLM.chat by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17356cutlass_mla_decode for ROCm build by @tywuAMD in https://github.com/vllm-project/vllm/pull/17289python3 setup.py develop with standard pip install --e on TPU by @NickLucche in https://github.com/vllm-project/vllm/pull/17374ModelConfig by @hmellor in https://github.com/vllm-project/vllm/pull/17130logger.info_once by @hmellor in https://github.com/vllm-project/vllm/pull/17416ObservabilityConfig by @hmellor in https://github.com/vllm-project/vllm/pull/17453awscli dependency by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17532arg_utils.py to be in their final groups by @hmellor in https://github.com/vllm-project/vllm/pull/17531pt_load_map_location to allow loading to cuda by @jerryzh168 in https://github.com/vllm-project/vllm/pull/16869prompt_embeds behind a feature flag by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17607--enable-prompt-embeds by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17615dockerfilegraph pre-commit hook by @hmellor in https://github.com/vllm-project/vllm/pull/17698benchmark_serving.py by @dtransposed in https://github.com/vllm-project/vllm/pull/16839apply_rotary_emb from vllm_flash_attn for Qwen2-VL vision RoPE by @Isotr0py in https://github.com/vllm-project/vllm/pull/17726deprecated=True CLI kwarg by @hmellor in https://github.com/vllm-project/vllm/pull/17781MultiprocExecutor.workers error by @njhill in https://github.com/vllm-project/vllm/pull/17811vllm by @hmellor in https://github.com/vllm-project/vllm/pull/17810--disable-log-stats in V1 server mode by @njhill in https://github.com/vllm-project/vllm/pull/17600use_fast failing to be propagated to Qwen2-VL image processor by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17838>= 0.7.11) to avoid AttributeError (no StructTag) by @shen-shanshan in https://github.com/vllm-project/vllm/pull/17839 max_num_batched_tokens config by @inkcherry in https://github.com/vllm-project/vllm/pull/17853top_k to be disabled with 0 (still accept -1 for now) by @hmellor in https://github.com/vllm-project/vllm/pull/17773str passed to /v1/audio/transcriptions by @hmellor in https://github.com/vllm-project/vllm/pull/17909ModelConfig when default constructing VllmConfig by @hmellor in https://github.com/vllm-project/vllm/pull/17943transformers.Auto*.from_pretrained processors by @xinli-centml in https://github.com/vllm-project/vllm/pull/17948rocm_aiter_rms_norm by @tjtanaa in https://github.com/vllm-project/vllm/pull/17857"EngineClient" has no attribute "model_config" by @b8zhong in https://github.com/vllm-project/vllm/pull/17976KVTransferConfig properly from Python instead of using JSON blobs without CLI by @hmellor in https://github.com/vllm-project/vllm/pull/17994SchedulerConfig by @hmellor in https://github.com/vllm-project/vllm/pull/17995tool_choice: required when using Xgrammar as the StructuredOutputBackend. by @chaunceyjiang in https://github.com/vllm-project/vllm/pull/17845.buildkite to ruff format by @hmellor in https://github.com/vllm-project/vllm/pull/17656model_executor/layers by @hmellor in https://github.com/vllm-project/vllm/pull/18056vllm/profiler by @hmellor in https://github.com/vllm-project/vllm/pull/18057vllm/transformers_utils by @hmellor in https://github.com/vllm-project/vllm/pull/18058benchmarks to ruff format by @hmellor in https://github.com/vllm-project/vllm/pull/18068vllm/compilation by @hmellor in https://github.com/vllm-project/vllm/pull/18072vllm/adapter_commons by @hmellor in https://github.com/vllm-project/vllm/pull/18073NixlConnector by @njhill in https://github.com/vllm-project/vllm/pull/18102topk_weight loading by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/18080vllm/lora by @hmellor in https://github.com/vllm-project/vllm/pull/18128vllm/device_allocator and vllm/distributed by @hmellor in https://github.com/vllm-project/vllm/pull/18126platform, plugins, triton_utils, vllm_flash_attn by @hmellor in https://github.com/vllm-project/vllm/pull/18129AOPerModuleConfig by @jerryzh168 in https://github.com/vllm-project/vllm/pull/17826test_openai_schema.py pass by @davidxia in https://github.com/vllm-project/vllm/pull/17664models by @hmellor in https://github.com/vllm-project/vllm/pull/18132test_kv_cache_events() by @davidxia in https://github.com/vllm-project/vllm/pull/18183model_loader by @hmellor in https://github.com/vllm-project/vllm/pull/18130chunked_prefill_paged_decode as fallback for V1 attention on ROCm by @kliuae in https://github.com/vllm-project/vllm/pull/18093resolve_hf_chat_template by @fxmarty-amd in https://github.com/vllm-project/vllm/pull/18259an illegal memory access was encountered of marlin kernel + act_order by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/18245usedforsecurity=False in MD5 hashing to enable FIPS by @shaoyuyoung in https://github.com/vllm-project/vllm/pull/18319AutoWeightsLoader to skip loading weights with specific substr in name by @Isotr0py in https://github.com/vllm-project/vllm/pull/18358--device arg by @kebe7jun in https://github.com/vllm-project/vllm/pull/18399--enable-reasoning by @Zerohertz in https://github.com/vllm-project/vllm/pull/18404dp_rank==0 by @njhill in https://github.com/vllm-project/vllm/pull/18502ndarray.tobytes() directly instead of ndarray.data.tobytes() by @lgeiger in https://github.com/vllm-project/vllm/pull/18347test_openai_schema.py pass by @davidxia in https://github.com/vllm-project/vllm/pull/18224KVTransferConfig.engine_id in post_init by @lk-chen in https://github.com/vllm-project/vllm/pull/18576cuda hard code with current_platform by @shen-shanshan in https://github.com/vllm-project/vllm/pull/16983--torch-backend=auto by @mgoin in https://github.com/vllm-project/vllm/pull/18505requirements/cpu.txt by @yankay in https://github.com/vllm-project/vllm/pull/18542{func} with mkdocs style links by @hmellor in https://github.com/vllm-project/vllm/pull/18610chmod +x to cleanup_pr_body.sh by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18650cuda hard code with current_platform in Ray by @noemotiovon in https://github.com/vllm-project/vllm/pull/14668kernels/quantization/ tests by @mgoin in https://github.com/vllm-project/vllm/pull/18669gte-Qwen2-1.5B-instruct usage by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/18683Note truncated.
This post release contains two bug fix for memory leak and model accuracy
This post release contains two bug fix for memory leak and model accuracy
_cached_reqs_data (#17567)Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.5...v0.8.5.post1
This release contains 310 commits from 143 contributors (55 new contributors!).
This release contains 310 commits from 143 contributors (55 new contributors!).
This release features important multi-modal bug fixes, day 0 support for Qwen3, and xgrammar's structure tag feature for tool calling.
structural_tag support using xgrammar (#17085)v1/audio/transcriptions endpoint (#16591)vllm bench [latency, throughput] CLI commands (#16508)--enable-chunked-prefill, --multi-step-stream-outputs, --disable-chunked-mm-input can no longer explicitly be set to False. Instead, add no- to the start of the argument (i.e. --enable-chunked-prefill and --no-enable-chunked-prefill) (https://github.com/vllm-project/vllm/pull/16533)SchedulerConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16533max-num-batched-tokens is not a power of 2 by @NickLucche in https://github.com/vllm-project/vllm/pull/16596pyzmq version by @taneem-ibrahim in https://github.com/vllm-project/vllm/pull/16549vllm bench [latency, throughput] CLI commands by @mgoin in https://github.com/vllm-project/vllm/pull/16508compressed-tensors WNA16 to support zero-points by @dsikka in https://github.com/vllm-project/vllm/pull/14211backend_xgrammar.py by @shen-shanshan in https://github.com/vllm-project/vllm/pull/16578additional_dependencies: [toml] for pre-commit yapf hook by @yankay in https://github.com/vllm-project/vllm/pull/16405TokenizerPoolConfig + DeviceConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16603max-num-batched-tokens is not even by @NickLucche in https://github.com/vllm-project/vllm/pull/16726--compilation-config by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16729_validate_structured_output() by @shen-shanshan in https://github.com/vllm-project/vllm/pull/16748MultiModalConfig + PoolerConfig + DecodingConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16789nullable_kvs fallback by @hmellor in https://github.com/vllm-project/vllm/pull/16837v1/audio/transcriptions endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/16591CacheConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16835_update_states for GPU model runner by @SnowCharmQ in https://github.com/vllm-project/vllm/pull/16910SpeculativeConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16971collective_rpc timeout by @njhill in https://github.com/vllm-project/vllm/pull/17000tests/kernels/ based on kernel type by @mgoin in https://github.com/vllm-project/vllm/pull/16799pid passed to kill_process_tree is int for mypy by @hmellor in https://github.com/vllm-project/vllm/pull/17051CacheConfig.block_size should always be int when used by @hmellor in https://github.com/vllm-project/vllm/pull/17052@property and private field for data_parallel_rank_local by @hmellor in https://github.com/vllm-project/vllm/pull/17053TokenizerGroup by @hmellor in https://github.com/vllm-project/vllm/pull/16790LoRAModelRunnerMixin by @hmellor in https://github.com/vllm-project/vllm/pull/17104tool-calling github label by @russellb in https://github.com/vllm-project/vllm/pull/17118:markdownhelp: to EngineArgs docs so markdown docstrings render properly by @hmellor in https://github.com/vllm-project/vllm/pull/17124LoRAConfig + PromptAdapterConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16980SchedulerConfig args into scheduler config group in EngineArgs by @hmellor in https://github.com/vllm-project/vllm/pull/17131get_text_config() instead of checking for text_config by @hmellor in https://github.com/vllm-project/vllm/pull/17105LLM.chat() tokenization by @njhill in https://github.com/vllm-project/vllm/pull/16081-n in multi-image example by @Isotr0py in https://github.com/vllm-project/vllm/pull/17223structural_tag support using xgrammar by @russellb in https://github.com/vllm-project/vllm/pull/17085vllm_flash_attn during development mode by @aarnphm in https://github.com/vllm-project/vllm/pull/17228skip_tokenizer_init with num_scheduler_steps by @junstar92 in https://github.com/vllm-project/vllm/pull/9276stop_token_ids contents by @njhill in https://github.com/vllm-project/vllm/pull/17268PromptAdapterConfig by @hmellor in https://github.com/vllm-project/vllm/pull/17302get_language_model to new MLLMs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/17300platforms/interface.py by @NickLucche in https://github.com/vllm-project/vllm/pull/17307compressed-tensors quant method consistent across vLLM by @hmellor in https://github.com/vllm-project/vllm/pull/17255process_weights_after_loading. by @charlifu in https://github.com/vllm-project/vllm/pull/16854Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.4...v0.8.5
This release contains 180 commits from 84 contributors (25 new contributors!).
This release contains 180 commits from 84 contributors (25 new contributors!).
This release includes important accuracy fixes for Llama4 models, if you are using it, we highly recommend you to update.
--max-model-len (#16168)request_latency, time_to_first_token and time_per_output_token (#15202)language_model interface for getting text backbone in MM (#16410)auto by default (#15724)supports_structured_output() method to Platform (#16148)LoadConfig (#16422), ParallelConfig (#16332)tokenizer as kwarg to validate_guidance_grammar by @ywang96 in https://github.com/vllm-project/vllm/pull/16117request_latency, time_to_first_token and time_per_output_token by @yankay in https://github.com/vllm-project/vllm/pull/15202supports_structured_output() method to Platform by @shen-shanshan in https://github.com/vllm-project/vllm/pull/16148ChatGLMForConditionalGeneration by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16187max_num_seqs to V0 values for most hardware by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16158max-model-len cli arg by @NickLucche in https://github.com/vllm-project/vllm/pull/16181disable_chunked_mm_input arg to disable partial mm input prefill by @mgoin in https://github.com/vllm-project/vllm/pull/15837nvidia-smi topo -m output in collect_env.py by @imkero in https://github.com/vllm-project/vllm/pull/16272process_weights_after_loading for QKVCrossParallelLinear by @Isotr0py in https://github.com/vllm-project/vllm/pull/15328SupportsMultiModal.get_language_model interface by @NickLucche in https://github.com/vllm-project/vllm/pull/16007benchmark_throughput.py --backend=hf by @mgoin in https://github.com/vllm-project/vllm/pull/16352skip_special_use=False for MistralTokenizer by @gcalmettes in https://github.com/vllm-project/vllm/pull/14094BaseProcessingInfo.get_mm_max_tokens_per_item by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16408language_model interface for getting text backbone in MM by @NickLucche in https://github.com/vllm-project/vllm/pull/16410ParallelConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16332auto by default by @russellb in https://github.com/vllm-project/vllm/pull/15724ppc64le platform by @hmellor in https://github.com/vllm-project/vllm/pull/16470--disable_chunked_mm_input mandatory for serving MM models by @NickLucche in https://github.com/vllm-project/vllm/pull/16483LoadConfig by @hmellor in https://github.com/vllm-project/vllm/pull/16422Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.3...v0.8.4
[V1][Spec Decode] Remove deprecated spec decode config params by @ShangmingCai in https://github.com/vllm-project/vllm/pull/15466
This release features 260 commits, 109 contributors, 38 new contributors.
wake_up (#15500)huggingface_hub to enable Xet downloads (#15873)num_embeds by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15443SchedulerInterface type for engine scheduler field by @njhill in https://github.com/vllm-project/vllm/pull/15499TransformersModel by @hmellor in https://github.com/vllm-project/vllm/pull/15467backend when generate grammar by @aarnphm in https://github.com/vllm-project/vllm/pull/15317scatter_patch_features by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15559is_encoder_decoder_inputs with split_enc_dec_inputs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15620mm_registry in compute_encoder_budget by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15621tpu label. by @russellb in https://github.com/vllm-project/vllm/pull/15634mm_hashes forgetting to be passed by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15668embed_is_patch for Idefics3 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15696mm_counts for dummy data creation by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15703embed_is_patch mask for fuyu model by @kylehh in https://github.com/vllm-project/vllm/pull/15731_try_schedule_encoder_inputs for every request by @WoosukKwon in https://github.com/vllm-project/vllm/pull/15778transformers to v4.50.3 by @hmellor in https://github.com/vllm-project/vllm/pull/13905MultiModalDataParser by @Isotr0py in https://github.com/vllm-project/vllm/pull/15828format.sh as it's been unsupported >70 days by @hmellor in https://github.com/vllm-project/vllm/pull/15884format.sh and make pre-commit installation simpler by @hmellor in https://github.com/vllm-project/vllm/pull/15890k_index is int64 for apply_top_k_only by @b8zhong in https://github.com/vllm-project/vllm/pull/15907huggingface_hub to enable Xet downloads by @hmellor in https://github.com/vllm-project/vllm/pull/15873async_request_deepspeed_mii uses the OpenAI choices key by @b8zhong in https://github.com/vllm-project/vllm/pull/15926tool_choice='required' by @meffmadd in https://github.com/vllm-project/vllm/pull/13483huggingface-cli[hf-xet] -> huggingface-cli[hf_xet] by @hmellor in https://github.com/vllm-project/vllm/pull/15969Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.2...v0.8.3
This release contains important bug fix for the V1 engine's memory usage. We highly recommend you upgrading!
This release contains important bug fix for the V1 engine's memory usage. We highly recommend you upgrading!
disable-any-whitespace option support for xgrammar (#15316)auto fallback mode (#14779)fastsafetensors loader for loading model weights (#10647)TransformersModel (#12832)merge_async_iterators fast-path for single-prompt requests by @njhill in https://github.com/vllm-project/vllm/pull/15150misc issues with link to forum by @hmellor in https://github.com/vllm-project/vllm/pull/15226extra_body as a way top pass vLLM only parameters using the OpenAI client by @hmellor in https://github.com/vllm-project/vllm/pull/15240max_num_seqs is between cudagraph capture sizes by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/15308disable-any-whitespace option support for xgrammar by @russellb in https://github.com/vllm-project/vllm/pull/15316generation_config by default by @ywang96 in https://github.com/vllm-project/vllm/pull/15281fastsafetensors loader for loading model weights by @manish-sethi in https://github.com/vllm-project/vllm/pull/10647TransformersModel by @hmellor in https://github.com/vllm-project/vllm/pull/12832auto fallback mode by @russellb in https://github.com/vllm-project/vllm/pull/14779Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.1...v0.8.2
This release contains important bug fixes for v0.8.0. We highly recommend upgrading!
This release contains important bug fixes for v0.8.0. We highly recommend upgrading!
V1 Fixes
TPU
Model
AutoModelForImageTextToText to load image models in tests by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/14945logprobs in ChatCompletionRequest by @schoennenbeck in https://github.com/vllm-project/vllm/pull/14352do_rescale warning when passing dummy data by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15107tokenizer_mode by @aarnphm in https://github.com/vllm-project/vllm/pull/15040Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.0...v0.8.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →