NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2222 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release 7 days ago
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
…vllm:model_execute_time_milliseconds) has been deprecated and subject to removal
v0.8.0 featured 523 commits from 166 total contributors (68 new contributors)!
We have now enabled V1 engine by default (#13726) for supported use cases. Please refer to V1 user guide for more detail. We expect better performance for supported scenarios. If you'd like to disable V1 mode, please specify the environment variable VLLM_USE_V1=0, and send us a GitHub issue sharing the reason!
SupportsV0Only protocol for model definitions (#13959)We observe state of the art performance with vLLM running DeepSeek model on latest version of vLLM:
lm_head on last pp rank (#13833)pip install git+https://github.com/huggingface/transformers.git) to use this model. Also, there may be numerical instabilities for float16/half dtype. Please use bfloat16 (preferred by HF) or float32 dtype.seed is now None to align with PyTorch and Hugging Face. Please explicitly set seed for reproduciblity. (#14274)kv_cache and attn_metadata arguments for model's forward method has been removed; as the attention backend has access to these value via forward_context. (#13887)generation_config from model for chat template, sampling parameters such as temperature, etc. (#12622)vllm:time_in_queue_requests, vllm:model_forward_time_milliseconds, vllm:model_execute_time_milliseconds) has been deprecated and subject to removal (#14135)return_tokens_as_token_id as a request param (#14066)/is_sleeping (#14312)vllm bench CLI (#13993)AMD
TPU
Neuron
CPU
s390x
Plugins
Benchmarks
CI and Build
torch.compile integration (#14437)pre-commit's isort version to remove warnings by @hmellor in https://github.com/vllm-project/vllm/pull/13614requirements-test.txt by @hmellor in https://github.com/vllm-project/vllm/pull/13617mm_processor_kwargs to chat-related protocols by @ywang96 in https://github.com/vllm-project/vllm/pull/13644vllm:cache_config_info by @markmc in https://github.com/vllm-project/vllm/pull/13299--show-hidden-metrics-for-version CLI arg by @markmc in https://github.com/vllm-project/vllm/pull/13295--dataset from benchmark_serving.py by @ywang96 in https://github.com/vllm-project/vllm/pull/13708VLLM_ATTENTION_BACKEND is set by @NickLucche in https://github.com/vllm-project/vllm/pull/12513EngineArgs.create_engine_config by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/13734AsyncOutputProcessing Logs by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/13780/v1/audio/transcriptions Bad Request Error by @HermitSun in https://github.com/vllm-project/vllm/pull/13811MyGemma2Embedding test by @hmellor in https://github.com/vllm-project/vllm/pull/13820lm_head on last pp rank by @hmellor in https://github.com/vllm-project/vllm/pull/13833kv_cache and attn_metadata by @hmellor in https://github.com/vllm-project/vllm/pull/13887mergify by @b8zhong in https://github.com/vllm-project/vllm/pull/13944.pre-commit-config.yaml's exclude by @hmellor in https://github.com/vllm-project/vllm/pull/13967SupportsV0Only protocol for model definitions by @ywang96 in https://github.com/vllm-project/vllm/pull/13959pipeline_parallel_size to optimization docs by @b8zhong in https://github.com/vllm-project/vllm/pull/14059whisper and florence2 examples by @Isotr0py in https://github.com/vllm-project/vllm/pull/14050__repr__ to KVCacheBlock to avoid recursive print by @heheda12345 in https://github.com/vllm-project/vllm/pull/14081TransformersModel by @hmellor in https://github.com/vllm-project/vllm/pull/14147head_dim not existing in all model configs (Transformers backend) by @hmellor in https://github.com/vllm-project/vllm/pull/14141vllm:tokens_total by @markmc in https://github.com/vllm-project/vllm/pull/14134--generation-config is not None by @hmellor in https://github.com/vllm-project/vllm/pull/14223prompt_logprobs clamping for chat as well as completions by @hmellor in https://github.com/vllm-project/vllm/pull/14225envs.VLLM_USE_V1 in mm processing by @ywang96 in https://github.com/vllm-project/vllm/pull/14256best_of Sampling Parameter in anticipation for vLLM V1 by @vincent-4 in https://github.com/vllm-project/vllm/pull/13997ray list nodes command to troubleshoot ray issues by @ruisearch42 in https://github.com/vllm-project/vllm/pull/14318QKVParallelLinear computation by @NickLucche in https://github.com/vllm-project/vllm/pull/12325best_of for V0 by @hmellor in https://github.com/vllm-project/vllm/pull/14356cudaProfilerStop in benchmarks script by @b8zhong in https://github.com/vllm-project/vllm/pull/14183kv_caches and attn_metadata in OpenVINOCausalLM by @hmellor in https://github.com/vllm-project/vllm/pull/14271extra_args to SamplingParams by @akeshet in https://github.com/vllm-project/vllm/pull/13300generation_config from model by @hmellor in https://github.com/vllm-project/vllm/pull/12622use_tqdm_on_load to reduce logs by @aarnphm in https://github.com/vllm-project/vllm/pull/14407model_impl arg when explaining Transformers fallback by @hmellor in https://github.com/vllm-project/vllm/pull/14552Github -> GitHub by @hmellor in https://github.com/vllm-project/vllm/pull/14561second_per_grid_ts for Qwen2-VL & Qwen2.5-VL by @ywang96 in https://github.com/vllm-project/vllm/pull/14548VLLM -> vLLM by @hmellor in https://github.com/vllm-project/vllm/pull/14562--hf-overrides for Alibaba-NLP/gte-Qwen2 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/14609vllm bench CLI by @randyjhc in https://github.com/vllm-project/vllm/pull/13993QKVCrossParallelLinear implementation to support BNB 4-bit quantization by @Isotr0py in https://github.com/vllm-project/vllm/pull/14545not include_stop_str_in_output by @afeldman-nm in https://github.com/vllm-project/vllm/pull/14624VLLM_CPU_MOE_PREPACK to allow disabling MoE prepack when CPU does not support it by @gau-nernst in https://github.com/vllm-project/vllm/pull/14681SamplingParams.__post_init__() by @njhill in https://github.com/vllm-project/vllm/pull/14772SupportsMultiModal by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/14794/is_sleeping by @waltforme in https://github.com/vllm-project/vllm/pull/14312ccache with pip install -e . in doc by @vadiklyutiy in https://github.com/vllm-project/vllm/pull/14901main branch by @russellb in https://github.com/vllm-project/vllm/pull/14692--seed option to offline multi-modal examples by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/14934Full Changelog: https://github.com/vllm-project/vllm/compare/v0.7.3...v0.8.0
[ROCm] MI300A compile targets deprecation by @gshtras in https://github.com/vllm-project/vllm/pull/13560
🎉 253 commits from 93 contributors, including 29 new contributors!
transformers backend
transformers backend (#12960)torch_dtype in TransformersModel (#13088)/v1/audio/transcriptions OpenAI API endpoint (#12909)uses_mrope in GPUModelRunner by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12969local_rank as device index by @MengqingCao in https://github.com/vllm-project/vllm/pull/13027torch_dtype in TransformersModel by @hmellor in https://github.com/vllm-project/vllm/pull/13088tokenizer-mode=mistral by @rafvasq in https://github.com/vllm-project/vllm/pull/12332dynamic quantization per module override/control by @Qubitium in https://github.com/vllm-project/vllm/pull/7086/v1/audio/transcriptions OpenAI API endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/12909SupportsQuant to phi3 and clip by @kylesayrs in https://github.com/vllm-project/vllm/pull/13104DictEmbeddingItems by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13380SamplingType.BEAM by @hmellor in https://github.com/vllm-project/vllm/pull/13402transformers backend by @Isotr0py in https://github.com/vllm-project/vllm/pull/12960--use-v2-block-manager by @hmellor in https://github.com/vllm-project/vllm/pull/13492hf_list_repo_files for local model path by @Isotr0py in https://github.com/vllm-project/vllm/pull/13348LLM terminates by @njhill in https://github.com/vllm-project/vllm/pull/13565offline_inference into single basic example by @hmellor in https://github.com/vllm-project/vllm/pull/12737Full Changelog: https://github.com/vllm-project/vllm/compare/v0.7.2...v0.7.3
[Core] Silence unnecessary deprecation warnings by @russellb in https://github.com/vllm-project/vllm/pull/12620
transformers library at the moment (#12604)transformers backend support via --model-impl=transformers. This allows vLLM to be ran with arbitrary Hugging Face text models (#11330, #12785, #12727).torch.compile to fused_moe/grouped_topk, yielding 5% throughput enhancement (#12637)VLLM_LOGITS_PROCESSOR_THREADS to speed up structured decoding in high batch size scenarios (#12368)transformers backend support by @ArthurZucker in https://github.com/vllm-project/vllm/pull/11330uncache_blocks and support recaching full blocks by @comaniac in https://github.com/vllm-project/vllm/pull/12415VLLM_LOGITS_PROCESSOR_THREADS by @akeshet in https://github.com/vllm-project/vllm/pull/12368Linear handling in TransformersModel by @hmellor in https://github.com/vllm-project/vllm/pull/12727FinishReason enum and use constant strings by @njhill in https://github.com/vllm-project/vllm/pull/12760TransformersModel UX by @hmellor in https://github.com/vllm-project/vllm/pull/12785Full Changelog: https://github.com/vllm-project/vllm/compare/v0.7.1...v0.7.2
This release features MLA optimization for Deepseek family of models. Compared to v0.7.0 released this Monday, we offer ~3x the generation throughput,
This release features MLA optimization for Deepseek family of models. Compared to v0.7.0 released this Monday, we offer ~3x the generation throughput, ~10x the memory capacity for tokens, and horizontal context scalability with pipeline parallelism
For the V1 architecture, we
prompt_logprobs with ChunkedPrefill by @NickLucche in https://github.com/vllm-project/vllm/pull/10132pre-commit hooks by @hmellor in https://github.com/vllm-project/vllm/pull/12475suggestion pre-commit hook multiple times by @hmellor in https://github.com/vllm-project/vllm/pull/12521?device={device} when changing tab in installation guides by @hmellor in https://github.com/vllm-project/vllm/pull/12560cutlass_scaled_mm to support 2d group (blockwise) scaling by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/11868sparsity_config.ignore in Cutlass Integration by @rahul-tuli in https://github.com/vllm-project/vllm/pull/12517Full Changelog: https://github.com/vllm-project/vllm/compare/v0.7.0...v0.7.1
[Doc] Create a vulnerability management team by @russellb in https://github.com/vllm-project/vllm/pull/9925
VLLM_USE_V1=1. See our blog for more details. (44 commits).LLM.sleep, LLM.wake_up, LLM.collective_rpc, LLM.reset_prefix_cache) in vLLM for the post training frameworks! (#12361, #12084, #12284).torch.compile is now fully integrated in vLLM, and enabled by default in V1. You can turn it on via -O3 engine parameter. (#11614, #12243, #12043, #12191, #11677, #12182, #12246).This release features
Models
get_*_embeddings methods according to this guide is automatically supported by V1 engine.Hardwares
W8A8 (#11785)Features
collective_rpc abstraction (#12151, #11256)moe_align_block_size for cuda graph and large num_experts (#12222)weights_only=True when using torch.load() (#12366)Detokenizer and EngineCore input by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/11545RayExecutor support for AsyncLLM (api server) by @jikunshang in https://github.com/vllm-project/vllm/pull/11712StreamingResponse Exception Handling by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/11752Whisper by @ywang96 in https://github.com/vllm-project/vllm/pull/11784LLaVa-NeXT-Video by @ywang96 in https://github.com/vllm-project/vllm/pull/11798gte-Qwen2 models by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/11808ModelConfig when RunAI Model Streamer is used by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/11825W8A8 by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/11785print_*_once from utils to logger by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/11298UnspecifiedPlatform package name by @jikunshang in https://github.com/vllm-project/vllm/pull/11916get_ip function by @KuntaiDu in https://github.com/vllm-project/vllm/pull/11932benchmark_long_document_qa_throughput.py by @KuntaiDu in https://github.com/vllm-project/vllm/pull/11933transformers 4.46.0 by @Isotr0py in https://github.com/vllm-project/vllm/pull/9685OutputProcessor Abstraction by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/11973.transpose() and .view() consecutively. by @liaoyanqing666 in https://github.com/vllm-project/vllm/pull/11979get_tokenizer function by @e1ijah1 in https://github.com/vllm-project/vllm/pull/11982Engine Arguments document by @maang-h in https://github.com/vllm-project/vllm/pull/12045HTTPConnection by @zhouyuan in https://github.com/vllm-project/vllm/pull/12042OpenAI-Compatible Server documents by @maang-h in https://github.com/vllm-project/vllm/pull/12082head_size=256 for Deepseek v2 and v3 by @Isotr0py in https://github.com/vllm-project/vllm/pull/12067is not None check in VllmConfig.post_init by @heheda12345 in https://github.com/vllm-project/vllm/pull/12138deepseek_vl2 dependency by @Isotr0py in https://github.com/vllm-project/vllm/pull/12169pre-commit by @hmellor in https://github.com/vllm-project/vllm/pull/11975VllmRunner by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10353_get_cache_block_size by @heheda12345 in https://github.com/vllm-project/vllm/pull/12214attention to impl backend by @wangxiyuan in https://github.com/vllm-project/vllm/pull/12218HfExampleModels.find_hf_info by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12223MultiModalInputsV2 -> MultiModalInputs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12244benchmark_serving.py by @njhill in https://github.com/vllm-project/vllm/pull/12288reset_prefix_cache by @comaniac in https://github.com/vllm-project/vllm/pull/12284uncache_blocks by @comaniac in https://github.com/vllm-project/vllm/pull/12333process_after_weight_loading for W4A16 MoE Group Act Order by @dsikka in https://github.com/vllm-project/vllm/pull/11528RequestOutputs by @njhill in https://github.com/vllm-project/vllm/pull/12298Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.6...v0.7.0
This release restore functionalities for other quantized MoEs, which was introduced as part of initial DeepSeek V3 support 🙇 .
This release restore functionalities for other quantized MoEs, which was introduced as part of initial DeepSeek V3 support 🙇 .
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.6...v0.6.6.post1
Breaking change: X-Request-ID echoing is now opt-in instead of on by default for performance reason. Set --enable-request-id-headers to enable it.
Support Deepseek V3 (#11523, #11502) model.
vllm serve deepseek-ai/DeepSeek-V3 --tensor-parallel-size 8 --trust-remote-code --max-model-len 8192. The context length can be increased to about 32K beyond running into memory issue.Last mile stretch for V1 engine refactoring: API Server (#11529, #11530), penalties for sampler (#10681), prefix caching for vision language models (#11187, #11305), TP Ray executor (#11107,#11472)
Breaking change: X-Request-ID echoing is now opt-in instead of on by default for performance reason. Set --enable-request-id-headers to enable it.
QVQ and QwQ to the list of supported models (#11509)num_evictable_computed_blocks by @heheda12345 in https://github.com/vllm-project/vllm/pull/11310Molmo by @ywang96 in https://github.com/vllm-project/vllm/pull/11325Pixtral-Large-Instruct-2411 by @ywang96 in https://github.com/vllm-project/vllm/pull/11393input_positions creation for text-only inputs with mrope by @Isotr0py in https://github.com/vllm-project/vllm/pull/11434QVQ and QwQ to the list of supported models by @ywang96 in https://github.com/vllm-project/vllm/pull/11509Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.5...v0.6.6
[Misc] Remove deprecated names by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10817
torch.compile integration: Support for all attention backends, encoder-based models, dynamic FP8 fusion, shape specialization fixes, and performance optimizations (#10558, #10613, #10121, #10383, #10399, #10406, #10437, #10460, #10552, #10622, #10722, #10620, #10906, #11108, #11059, #11005, #10838, #11081, #11110).mark_step for hpu by @jikunshang in https://github.com/vllm-project/vllm/pull/10239AutoWeightsLoader by @Isotr0py in https://github.com/vllm-project/vllm/pull/10327get_default_attn_backend to Platform by @MengqingCao in https://github.com/vllm-project/vllm/pull/10358device_type in Platform by @MengqingCao in https://github.com/vllm-project/vllm/pull/10508dynamic_image_size as mm_processor_kwargs for InternVL2 models by @Isotr0py in https://github.com/vllm-project/vllm/pull/10518multi_modal_kwargs broadcast for CPU tensor parallel by @Isotr0py in https://github.com/vllm-project/vllm/pull/10541is_causal HF config field for Qwen2 model by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10621-O[0,3] with LLM entrypoint by @mgoin in https://github.com/vllm-project/vllm/pull/10677lm_head when loading embedding models by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10719driver_worker init by @NickLucche in https://github.com/vllm-project/vllm/pull/10779support_torch_compile by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10763MMMU-Pro vision dataset to serving benchmark by @ywang96 in https://github.com/vllm-project/vllm/pull/10804out in flash_attn_varlen_func by @WoosukKwon in https://github.com/vllm-project/vllm/pull/10811mistral_common version for tests and docs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10825AsyncLLM by @ywang96 in https://github.com/vllm-project/vllm/pull/10997async output check to platform by @wangxiyuan in https://github.com/vllm-project/vllm/pull/10768None in self.generators by @WoosukKwon in https://github.com/vllm-project/vllm/pull/11038deprecated decorator by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/11025-inf logprob values in prompt_logprobs by @rafvasq in https://github.com/vllm-project/vllm/pull/11073logits_processors as an extra completion argument by @bradhilton in https://github.com/vllm-project/vllm/pull/11150RMSNorm by @ywang96 in https://github.com/vllm-project/vllm/pull/11241Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.4...v0.6.5
This patch release covers bug fixes (#10347, #10349, #10348, #10352, #10363), keep compatibility for vLLMConfig usage in out of tree models
This patch release covers bug fixes (#10347, #10349, #10348, #10352, #10363), keep compatibility for vLLMConfig usage in out of tree models (#10356)
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.4...v0.6.4.post1
[Core] Deprecating block manager v1 and make block manager v2 default by @KuntaiDu in https://github.com/vllm-project/vllm/pull/8704
torch.compile support. Many models now support torch compile with TorchInductor. You can checkout our meetup slides for more details. (#9775, #9614, #9639, #9641, #9876, #9946, #9589, #9896, #9637, #9300, #9947, #9138, #9715, #9866, #9632, #9858, #9889)--task parameter for models that support both generation and embedding (#9424)fused_moe Performance Improvement (#9384)config.json via CLI (#5836)logger.exception by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9461mistral_common by @sasha0552 in https://github.com/vllm-project/vllm/pull/9446mistral_common tokenizer only once by @sasha0552 in https://github.com/vllm-project/vllm/pull/9468BERTModel (first encoder-only embedding model) by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9056mistral_common by @sasha0552 in https://github.com/vllm-project/vllm/pull/9457fork_new_process by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9484FetchContent multiple build issue by @ProExpertProg in https://github.com/vllm-project/vllm/pull/9596_init_vision_model in NVLM_D model by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9611image_url.detail by @mgoin in https://github.com/vllm-project/vllm/pull/9663illegal memory access error with chunked prefill, prefix caching, block manager v2 and xformers enabled together by @sasha0552 in https://github.com/vllm-project/vllm/pull/9532EngineArgs refactor on V1 by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9954MQLLMEngine hanging by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9973mm_processor_kwargs for Qwen2-VL by @li-plus in https://github.com/vllm-project/vllm/pull/10112MultiModalInputs to MultiModalKwargs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/10040tensor_parallel_size passing by @Isotr0py in https://github.com/vllm-project/vllm/pull/10161add_request_id middleware by @cjackal in https://github.com/vllm-project/vllm/pull/9594config.json via CLI by @KrishnaM251 in https://github.com/vllm-project/vllm/pull/5836tokenizer_mode and trust_remote_code for Detokenizer by @ywang96 in https://github.com/vllm-project/vllm/pull/10211swap_out function by @yyccli in https://github.com/vllm-project/vllm/pull/10212AsyncLLM Implementation by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9826--distributed-executor-backend by @russellb in https://github.com/vllm-project/vllm/pull/10231ValueError when tool_choice is set to the supported none option and tools are not defined. by @gcalmettes in https://github.com/vllm-project/vllm/pull/10000Detokenizer by @ywang96 in https://github.com/vllm-project/vllm/pull/10288tool_calls iterator before processing by mistral-common when MistralTokenizer is used by @gcalmettes in https://github.com/vllm-project/vllm/pull/9951Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.3...v0.6.4
[Core] Deprecating block manager v1 and make block manager v2 default by @KuntaiDu in https://github.com/vllm-project/vllm/pull/8704
_version.py not found issue (#9375)logger.exception by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9461mistral_common by @sasha0552 in https://github.com/vllm-project/vllm/pull/9446Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.3...v0.6.3.post1
[Bugfix] Deprecate registration of custom configs to huggingface by @heheda12345 in https://github.com/vllm-project/vllm/pull/9083
VLLM_TORCH_COMPILE_LEVEL to control torch.compile various levels of compilation control and integration (#9058). Along with various improvements (#8982, #9258, #906, #8875), using VLLM_TORCH_COMPILE_LEVEL=3 can turn on Inductor's full graph compilation without vLLM's custom ops.auto if at least one tool is provided by @chiragjn in https://github.com/vllm-project/vllm/pull/8568PromptInputs and inputs with backward compatibility by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8876max_num_seqs>1 by @Isotr0py in https://github.com/vllm-project/vllm/pull/8892BaichuanTokenizer by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8921continue_final_message parameter by @danieljannai21 in https://github.com/vllm-project/vllm/pull/8942cv2 via mistral_common[opencv] by @ywang96 in https://github.com/vllm-project/vllm/pull/8951lm-eval directly to requirements-test.txt by @mgoin in https://github.com/vllm-project/vllm/pull/9161get_vocab instead of vocab in tool parsers by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9188NotImplementedError by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/9218Dockerfile.cpu file's PIP_EXTRA_INDEX_URL Configurable as a Build Argument by @jyono in https://github.com/vllm-project/vllm/pull/9252is_first_step_output for TPUModelRunner by @allenwang28 in https://github.com/vllm-project/vllm/pull/9202Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.2...v0.6.3
⚠️ You will see the following error now, this is breaking change!
Support Llama 3.2 models (#8811, #8822)
vllm serve meta-llama/Llama-3.2-11B-Vision-Instruct --enforce-eager --max-num-seqs 16
Beam search have been soft deprecated. We are moving towards a version of beam search that's more performant and also simplifying vLLM's core. (#8684, #8763, #8713)
⚠️ You will see the following error now, this is breaking change!
Using beam search as a sampling parameter is deprecated, and will be removed in the future release. Please use the
vllm.LLM.use_beam_searchmethod for dedicated beam search instead, or set the environment variableVLLM_ALLOW_DEPRECATED_BEAM_SEARCH=1to suppress this error. For more details, see https://github.com/vllm-project/vllm/issues/8306
Support for Solar Model (#8386), minicpm3 (#8297), LLaVA-Onevision model support (#8486)
Enhancements: pp for qwen2-vl (#8696), multiple images for qwen-vl (#8247), mistral function calling (#8515), bitsandbytes support for Gemma2 (#8338), tensor parallelism with bitsandbytes quantization (#8434)
MQLLMEngine for API Server, boost throughput 30% in single step and 7% in multistep (#8157, #8761, #8584)IQ1_M quantization implementation to GGUF kernel by @Isotr0py in https://github.com/vllm-project/vllm/pull/8357MQLLMEngine to avoid asyncio OH by @alexm-neuralmagic in https://github.com/vllm-project/vllm/pull/8157dead_error property to engine client by @joerunde in https://github.com/vllm-project/vllm/pull/8574collect_env.py by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8649PromptInputs to PromptType, and inputs to prompt by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8673SequenceData and Sequence by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8675SequenceData.from_token_counts to create dummy data by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8687PromptInputs to PromptType, and inputs to prompt" by @simon-mo in https://github.com/vllm-project/vllm/pull/8750replace_parameters by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/8748PromptInputs and inputs, with backwards compatibility by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8760Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.1...v0.6.2
This release contains an important bugfix related to token streaming combined with stop string
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.1.post1...v0.6.1.post2
This release features important bug fixes and enhancements for
This release features important bug fixes and enhancements for
--max_num_batched_tokens 16384 with --max-model-len 16384Also
engine_use_ray (#8126)Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.1...v0.6.1.post1
Added support for Pixtral (mistralai/Pixtral-12B-2409). (#8377, #8168)
mistralai/Pixtral-12B-2409). (#8377, #8168)benchmark_serving.py by @afeldman-nm in https://github.com/vllm-project/vllm/pull/8191SqueezeLLM by @dsikka in https://github.com/vllm-project/vllm/pull/8220Paligemma by @Isotr0py in https://github.com/vllm-project/vllm/pull/8269post_layernorm in CLIP by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8155Full Changelog: https://github.com/vllm-project/vllm/compare/v0.6.0...v0.6.1
We are excited to announce a faster vLLM delivering 2x more throughput compared to v0.5.3. The default parameters should achieve great speed up, but w
--num-scheduler-steps 8 in the engine arguments. Please note that it still have some limitations and being actively hardened, see #7528 for known issues.
torch.compile: avoid Dynamo guard evaluation overhead (#7898), skip compile for profiling (#7796)RemoteOpenAIServer by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/7836qqq to use vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7805gptq_marlin_24 to use vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7762prefix from create_weights by @dsikka in https://github.com/vllm-project/vllm/pull/7825vllm/core by @jberkhahn in https://github.com/vllm-project/vllm/pull/7229max_model_len for multimodal models by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/7998max_num_batched_tokens for multimodal models by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8028GPTQ to use vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7976--async-engine option to benchmark_throughput.py by @njhill in https://github.com/vllm-project/vllm/pull/7964vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7972Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.5...v0.6.0
[Misc] Deprecation Warning when setting --engine-use-ray by @wallashss in https://github.com/vllm-project/vllm/pull/7424
--num-scheduler-steps 8 as a parameter to the API server (via vllm serve) or AsyncLLMEngine. We are working on expanding the coverage to LLM class and aiming to turning it on by defaultUltravoxModel (#7615, #7446)chat method in the LLM class (#5049)prompt_logprobs in Chat Completion (#7453)torch.compile: register custom ops for kernels (#7591, #7594, #7536)VLLM_PORT is set by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/7205BasevLLMParameter and weight_loader_v2 by @dsikka in https://github.com/vllm-project/vllm/pull/5874zmq frontend to IPC instead of TCP by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/7222merge_async_iterators is_cancelled arg optional by @njhill in https://github.com/vllm-project/vllm/pull/7282AsyncLLMEngine by @njhill in https://github.com/vllm-project/vllm/pull/7336stop_token_ids to InternVL example by @Isotr0py in https://github.com/vllm-project/vllm/pull/7354PerTensorScaleParameter weight loading for fused models by @dsikka in https://github.com/vllm-project/vllm/pull/7376compute_slot_mapping by @Yard1 in https://github.com/vllm-project/vllm/pull/7377model as both argument and option by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/7347gptq_marlin kernels by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/7323GB constant and enable float GB arguments by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/7416gptq_marlin to use new vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7281awq and awq_marlin to use vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7422compressed-tensors code reuse by @kylesayrs in https://github.com/vllm-project/vllm/pull/7277compressed-tensors code reuse by @kylesayrs in https://github.com/vllm-project/vllm/pull/7521torch.testing.assert_close by @jon-chuang in https://github.com/vllm-project/vllm/pull/7324MultiModalConfig initialization and profiling by @ywang96 in https://github.com/vllm-project/vllm/pull/7530zeromq Frontend by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/7279tie_word_embeddings for all models by @zijian-hu in https://github.com/vllm-project/vllm/pull/5724AttentionState abstraction by @Yard1 in https://github.com/vllm-project/vllm/pull/7663worker_class_fn argument in Executor by @Yard1 in https://github.com/vllm-project/vllm/pull/7707mm_limits initialization for CPU backend by @Isotr0py in https://github.com/vllm-project/vllm/pull/7735zeromq Frontend by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/7394blockReduce[...] functions with cub::BlockReduce by @ProExpertProg in https://github.com/vllm-project/vllm/pull/7233vLLMParameter by @dsikka in https://github.com/vllm-project/vllm/pull/7437vllm serve --load-format by @mgoin in https://github.com/vllm-project/vllm/pull/7784marlin to use vLLMParameters by @dsikka in https://github.com/vllm-project/vllm/pull/7803Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.4...v0.5.5
[Kernel] Fix deprecation function warnings squeezellm quant_cuda_kernel by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/6901
We are progressing along our quest to quickly improve performance. Each of the following PRs contributed some improvements, and we anticipate more enhancements in the next release.
zeromq. This brought 20% speedup over time to first token and 2x speedup over inter token latency. (#6883)get_seqs function, bring 2% throughput enhancements. (#7051)fp8 quantization by @mgoin in https://github.com/vllm-project/vllm/pull/6657transformers version for Llama 3.1 hotfix and patch Chameleon by @ywang96 in https://github.com/vllm-project/vllm/pull/6690reshape_and_cache_flash by @Yard1 in https://github.com/vllm-project/vllm/pull/6667fp8-marlin channelwise via compressed-tensors by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6524kv_cache_dtype=fp8 without scales for FP8 checkpoints by @mgoin in https://github.com/vllm-project/vllm/pull/6761conf.py by @hmellor in https://github.com/vllm-project/vllm/pull/6834allowed_token_ids decoding request parameter by @njhill in https://github.com/vllm-project/vllm/pull/6753multi_modal_kwargs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/6836get_seqs by @WoosukKwon in https://github.com/vllm-project/vllm/pull/7051zeromq by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6883max_model_len is greater than derived max_model_len by @fialhocoelho in https://github.com/vllm-project/vllm/pull/7080Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.3...v0.5.4
We fixed an configuration incompatibility between vLLM (which tested against pre-released version) and the published Meta Llama 3.1 weights
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.3...v0.5.3.post1
[Misc] Remove deprecation warning for beam search by @WoosukKwon in https://github.com/vllm-project/vllm/pull/6659
--cpu-offload-gb to control how much memory to "extend" the RAM with. (#6496)vllm CLI is now ready for testing. It comes with three commands: serve, complete, and chat. Feedback and improvements are greatly welcomed! (#6431)Attention.kv_scale into k_scale and v_scale by @mgoin in https://github.com/vllm-project/vllm/pull/6081vllm serve by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/6431torch.Tensor for type annotation by @WoosukKwon in https://github.com/vllm-project/vllm/pull/6505compressed-tensors by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6522compressed-tensors for Llama by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6515/metrics endpoint by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/6463fp8 by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6547fbgemm checkpoints by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6559fp8-marlin for fbgemm-fp8 models by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6606vocab_size field access in LLaVA models by @jaywonchung in https://github.com/vllm-project/vllm/pull/6624modules_to_not_convert in FBGEMM Fp8 quantization by @cli99 in https://github.com/vllm-project/vllm/pull/6665Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.2...v0.5.3
❗Planned breaking change ❗: we plan to remove beam search (see more in #6226) in the next few releases. This release come with a warning when beam sea…
llm-compressor by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6110object field in CompletionStreamResponse by @kczimm in https://github.com/vllm-project/vllm/pull/6196compressed-tensors integration by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6356vllm.__commit__ by @mgoin in https://github.com/vllm-project/vllm/pull/6386ClipVisionModel by @ywang96 in https://github.com/vllm-project/vllm/pull/6436Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.1...v0.5.2
Notably, it has a breaking change that all VLM specific arguments are now removed from engine APIs so you no longer need to set it globally via CLI. H…
--pipeline-parallel-size. This feature is in early stage, please let us know your feedback.<image> into the prompt instead of complicated prompt formatting. See more herecompressed-tensors supporting Marlin, W4A16 (#5435, #5385)is_quant_method_supported to control quantization test configurations by @mgoin in https://github.com/vllm-project/vllm/pull/5253is_quant_method_supported to control quantization test configurations" by @simon-mo in https://github.com/vllm-project/vllm/pull/5463w4a16 support for compressed-tensors by @dsikka in https://github.com/vllm-project/vllm/pull/5385cuda_device_count_stateless by @Yard1 in https://github.com/vllm-project/vllm/pull/5473perf-benchmarks label by @KuntaiDu in https://github.com/vllm-project/vllm/pull/5073is_in_the_same_node by @youkaichao in https://github.com/vllm-project/vllm/pull/5512compressed-tensors marlin 24 support by @dsikka in https://github.com/vllm-project/vllm/pull/5435gelu to CPU by @ywang96 in https://github.com/vllm-project/vllm/pull/5717RayTokenizerGroupPool by @Yard1 in https://github.com/vllm-project/vllm/pull/5748w4a16 compressed-tensors support to include w8a16 by @dsikka in https://github.com/vllm-project/vllm/pull/5794cutlass_scaled_mm by @ProExpertProg in https://github.com/vllm-project/vllm/pull/5560multi_modal_kwargs is broadcasted properly by @xwjiang2010 in https://github.com/vllm-project/vllm/pull/5880MLPSpeculator handling of num_speculative_tokens by @njhill in https://github.com/vllm-project/vllm/pull/5876min_tokens behaviour for multiple eos tokens by @njhill in https://github.com/vllm-project/vllm/pull/5849_get_logits_warper in Sampler Test by @ywang96 in https://github.com/vllm-project/vllm/pull/5922multi_modal_kwargs can broadcast properly with ring buffer. by @xwjiang2010 in https://github.com/vllm-project/vllm/pull/5905num_speculative_tokens is set too high by @tdoublep in https://github.com/vllm-project/vllm/pull/5894fp8_shard_indexer from Col/Row Parallel Linear (Simplify Weight Loading) by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/5928Attention.kv_scale if kv cache quantization is enabled by @mgoin in https://github.com/vllm-project/vllm/pull/5936eos_token_id from config.json by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5954SequenceStatus.is_finished by switching to IntEnum by @Yard1 in https://github.com/vllm-project/vllm/pull/5974get_min_capability by @dsikka in https://github.com/vllm-project/vllm/pull/5971process_weights_after_load (Simplify Weight Loading) by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/5940end_forward calls to flashinfer by @Yard1 in https://github.com/vllm-project/vllm/pull/6044image_input_type from VLM config by @xwjiang2010 in https://github.com/vllm-project/vllm/pull/5852compute_logits in Jamba by @ywang96 in https://github.com/vllm-project/vllm/pull/6093CompressedTensorsW8A8 by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/6113Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.0...v0.5.1
Fix crashes when using FlashAttention backend
is_quant_method_supported to control quantization test configurations by @mgoin in https://github.com/vllm-project/vllm/pull/5253is_quant_method_supported to control quantization test configurations" by @simon-mo in https://github.com/vllm-project/vllm/pull/5463w4a16 support for compressed-tensors by @dsikka in https://github.com/vllm-project/vllm/pull/5385cuda_device_count_stateless by @Yard1 in https://github.com/vllm-project/vllm/pull/5473Full Changelog: https://github.com/vllm-project/vllm/compare/v0.5.0...v0.5.0.post1
[Bugfix] Remove deprecated @abstractproperty by @zhuohan123 in https://github.com/vllm-project/vllm/pull/5174
tools support named functions (#5032)stream_options for OpenAI protocol (#5319, #5135)FSM to Guide (#4109)LLM.encode for non-generation Models by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/5184conftest.py by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5118tools support named functions by @br3no in https://github.com/vllm-project/vllm/pull/5032prompt_logprobs==0 by @toslunar in https://github.com/vllm-project/vllm/pull/5217HfRunner by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5251add_special_tokens to ChatCompletionRequest (default False) by @tomeras91 in https://github.com/vllm-project/vllm/pull/5278ProposerWorkerBase abstract class by @njhill in https://github.com/vllm-project/vllm/pull/5252FSM to Guide by @br3no in https://github.com/vllm-project/vllm/pull/4109stream_options in ChatCompletionRequest by @Etelis in https://github.com/vllm-project/vllm/pull/5135compressed-tensors config by @dsikka in https://github.com/vllm-project/vllm/pull/5350stream_options implementation also in CompletionRequest by @Etelis in https://github.com/vllm-project/vllm/pull/5319MultiprocessingGPUExecutor.check_health when world_size == 1 by @jsato8094 in https://github.com/vllm-project/vllm/pull/5254Full Changelog: https://github.com/vllm-project/vllm/compare/v0.4.3...v0.5.0
[Doc]Replace deprecated flag in readme by @ronensc in https://github.com/vllm-project/vllm/pull/4526
MultiprocessingGPUExecutor (#4539)asyncio.Task not being subscriptable by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/4623get_name method to attention backends by @WoosukKwon in https://github.com/vllm-project/vllm/pull/4685swap_blocks by @jikunshang in https://github.com/vllm-project/vllm/pull/4726test_utils.py to tests/utils.py by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/4425tensorizer to version 2.9.0 by @sangstar in https://github.com/vllm-project/vllm/pull/4208build/ directory by @mgoin in https://github.com/vllm-project/vllm/pull/4945max_seq_len_to_capture by @kerthcet in https://github.com/vllm-project/vllm/pull/4935clang-format by @mgoin in https://github.com/vllm-project/vllm/pull/4722Sequence in stop checker test by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5092merge_async_iterators by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5096git diff by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5097None values to Sequence.inputs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/5099seq_lens_tensor in model_runner.py by @ita9naiwa in https://github.com/vllm-project/vllm/pull/5129Full Changelog: https://github.com/vllm-project/vllm/compare/v0.4.2...v0.4.3
Your coding agent can read these notes before it upgrades. Set up the MCP server →