NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2222 most downloaded on PyPI
A high-throughput and memory-efficient inference and serving engine for LLMs
Last release 9 days ago
22 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
98 releases · first in 2023
One column per quarter.
Chunked prefill is ready for testing! It improves inter-token latency in high load scenario by chunking the prompt processing and priortizes decode
torch==2.3.0 (#4454)tensorizer==2.9.0 (#4467)engine to executor package by @njhill in https://github.com/vllm-project/vllm/pull/4347shutdown() method to ExecutorBase by @njhill in https://github.com/vllm-project/vllm/pull/4349get_tokenizer by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/4107DistributedGPUExecutor abstract class by @njhill in https://github.com/vllm-project/vllm/pull/4348min_tokens when eos_token_id is None by @njhill in https://github.com/vllm-project/vllm/pull/4389torch==2.3.0 by @mgoin in https://github.com/vllm-project/vllm/pull/4454num_readers, update version by @alpayariyak in https://github.com/vllm-project/vllm/pull/4467/metrics by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/4523multiproc_worker_utils for multiprocessing-based workers by @njhill in https://github.com/vllm-project/vllm/pull/4357tests directory from being packaged by @itechbear in https://github.com/vllm-project/vllm/pull/4552_force_log from being garbage collected by @Atry in https://github.com/vllm-project/vllm/pull/4567Full Changelog: https://github.com/vllm-project/vllm/compare/v0.4.1...v0.4.2
Support and enhance CommandR+ (#3829), minicpm (#3893), Meta Llama 3 (#4175, #4182), Mixtral 8x22b (#4073, #4002)
Features
tensorizer (#3476)Enhancements
mypy (#3816, #4006, #4161, #4043)Hardwares
__init__.py files for vllm/core/block/ and vllm/spec_decode/ by @mgoin in https://github.com/vllm-project/vllm/pull/3798is_cpu() by @njhill in https://github.com/vllm-project/vllm/pull/3804attention_bias Usage in Llama Model Configuration by @Ki6an in https://github.com/vllm-project/vllm/pull/3767guided_json parameter in OpenAi compatible Server by @dmarasco in https://github.com/vllm-project/vllm/pull/3945linear_weights directly on the layer by @Yard1 in https://github.com/vllm-project/vllm/pull/3977merge_async_iterators to utils by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/4026tensorizer by @sangstar in https://github.com/vllm-project/vllm/pull/3476tokenizer_revision when getting tokenizer in openai serving by @chiragjn in https://github.com/vllm-project/vllm/pull/4214EngineArgs by @hmellor in https://github.com/vllm-project/vllm/pull/4219EngineArgs by @hmellor in https://github.com/vllm-project/vllm/pull/4223autodoc directives by @hmellor in https://github.com/vllm-project/vllm/pull/4272Full Changelog: https://github.com/vllm-project/vllm/compare/v0.4.0...v0.4.1
v0.4.0 lacks support for sm70/75 support. We did a hotfix for it.
v0.4.0 lacks support for sm70/75 support. We did a hotfix for it.
__init__.py files for vllm/core/block/ and vllm/spec_decode/ by @mgoin in https://github.com/vllm-project/vllm/pull/3798Full Changelog: https://github.com/vllm-project/vllm/compare/v0.4.0...v0.4.0.post1
New models: Command+R(#3433), Qwen2 MoE(#3346), DBRX(#3660), XVerse (#3610), Jais (#3183).
--enable-prefix-caching to turn it on.json_object in OpenAI server for arbitrary JSON, --use-delay flag to improve time to first token across many requests, and min_tokens to EOS suppression.eos_token_id in Sequence for easy access by @njhill in https://github.com/vllm-project/vllm/pull/3166dynamic_ncols=True by @chujiezheng in https://github.com/vllm-project/vllm/pull/3242flash_attn optional by @WoosukKwon in https://github.com/vllm-project/vllm/pull/3269/tmp/ to ~/.cache/vllm/locks/ dir by @mgoin in https://github.com/vllm-project/vllm/pull/3241flash_attn in Docker image by @tdoublep in https://github.com/vllm-project/vllm/pull/3396dist.broadcast stall without group argument by @GindaChen in https://github.com/vllm-project/vllm/pull/3408lstrip() with removeprefix() to fix Ruff linter warning by @ronensc in https://github.com/vllm-project/vllm/pull/2958LRUCache by @njhill in https://github.com/vllm-project/vllm/pull/3511logits computation and gather to model_runner by @esmeetu in https://github.com/vllm-project/vllm/pull/3233_prune_hidden_states by @rkooo567 in https://github.com/vllm-project/vllm/pull/3539rotary_embedding.py file, get_device() -> device by @jikunshang in https://github.com/vllm-project/vllm/pull/3604_get_ranks in Sampler by @Yard1 in https://github.com/vllm-project/vllm/pull/3623Full Changelog: https://github.com/vllm-project/vllm/compare/v0.3.3...v0.4.0
Performance optimization and LoRA support for Gemma
benchmark_serving.py by @ronensc in https://github.com/vllm-project/vllm/pull/2934counter_generation_tokens by @ronensc in https://github.com/vllm-project/vllm/pull/2802aioprometheus to prometheus_client by @hmellor in https://github.com/vllm-project/vllm/pull/2730get_ip error in pure ipv6 environment by @Jingru in https://github.com/vllm-project/vllm/pull/2931AttributeError in OpenAI-compatible server by @jaywonchung in https://github.com/vllm-project/vllm/pull/3018Full Changelog: https://github.com/vllm-project/vllm/compare/v0.3.2...v0.3.3
This version adds support for the OLMo and Gemma Model, as well as seed parameter.
This version adds support for the OLMo and Gemma Model, as well as seed parameter.
sampling_params by @njhill in https://github.com/vllm-project/vllm/pull/2881vllm:prompt_tokens_total metric calculation by @ronensc in https://github.com/vllm-project/vllm/pull/2869Full Changelog: https://github.com/vllm-project/vllm/compare/v0.3.1...v0.3.2
This version fixes the following major bugs:
This version fixes the following major bugs:
Also with many smaller bug fixes listed below.
prefix_len. by @sighingnow in https://github.com/vllm-project/vllm/pull/2688device="cuda" to support more device by @jikunshang in https://github.com/vllm-project/vllm/pull/2503LlamaForCausalLM instead by @pcmoritz in https://github.com/vllm-project/vllm/pull/2854LLM class by @WoosukKwon in https://github.com/vllm-project/vllm/pull/2882Full Changelog: https://github.com/vllm-project/vllm/compare/v0.3.0...v0.3.1
Experimental multi-lora support
top_p and top_k Sampling by @chenxu2048 in https://github.com/vllm-project/vllm/pull/1885benchmark_serving.py by @hmellor in https://github.com/vllm-project/vllm/pull/2172group as an argument in broadcast ops by @GindaChen in https://github.com/vllm-project/vllm/pull/2522scheduler.running as deque by @njhill in https://github.com/vllm-project/vllm/pull/2523benchmark_serving.py by @hmellor in https://github.com/vllm-project/vllm/pull/2552include_stop_str_in_output and length_penalty parameters to OpenAI API by @galatolofederico in https://github.com/vllm-project/vllm/pull/2562--engine-use-ray by @HermitSun in https://github.com/vllm-project/vllm/pull/2664Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.7...v0.3.0
Fix Gradio example: remove deprecated parameter concurrency_count by @ronensc in https://github.com/vllm-project/vllm/pull/2315
concurrency_count by @ronensc in https://github.com/vllm-project/vllm/pull/2315Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.6...v0.2.7
Fast model execution with CUDA/HIP graph
quantization argument by @WoosukKwon in https://github.com/vllm-project/vllm/pull/2145Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.5...v0.2.6
Optimize Mixtral performance with expert parallelism (thanks to @Yard1)
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.4...v0.2.5
Mixtral model support (officially from @mistralai)
rope_scaling in config.json by @theFool32 in https://github.com/vllm-project/vllm/pull/1956Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.3...v0.2.4
Refactoring on Worker, InputMetadata, and Attention
pylint to ruff by @simon-mo in https://github.com/vllm-project/vllm/pull/1665input_is_parallel=False for ScaledActivation by @zhuohan123 in https://github.com/vllm-project/vllm/pull/1737max_num_seqs in latency benchmark by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1855echo for chat API by @Tostino in https://github.com/vllm-project/vllm/pull/1756Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.2...v0.2.3
Bump up to PyTorch v2.1 + CUDA 12.1 (vLLM+CUDA 11.8 is also provided)
test_models by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1366MptForCausalLM key in model_loader by @wenfeiy-db in https://github.com/vllm-project/vllm/pull/1526MPTConfig by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1529_get_and_verify_max_len by @irasin in https://github.com/vllm-project/vllm/pull/1617get_rope by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1633MptConfig by @megha95 in https://github.com/vllm-project/vllm/pull/1668Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.1...v0.2.2
This is an emergency release to fix a bug on tensor parallelism support.
This is an emergency release to fix a bug on tensor parallelism support.
fix vulnerable memory modification to gpu shared memory by @soundOfDestiny in https://github.com/vllm-project/vllm/pull/1241
tiiuae/falcon-rw-7b model name by @0ssamaak0 in https://github.com/vllm-project/vllm/pull/1226dtype arg to benchmarks by @kg6-sleipnir in https://github.com/vllm-project/vllm/pull/1228TORCH_CUDA_ARCH_LIST by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1239Full Changelog: https://github.com/vllm-project/vllm/compare/v0.2.0...v0.2.1
Up to 60% performance improvement by optimizing de-tokenization and sampler
max_model_len configurable by @Yard1 in https://github.com/vllm-project/vllm/pull/972--ipc=host in docker run for distributed inference by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1125max_tokens behavior with openai by @HermitSun in https://github.com/vllm-project/vllm/pull/852TORCH_CUDA_ARCH_LIST for selecting target GPUs by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1074max_num_batched_tokens by @WoosukKwon in https://github.com/vllm-project/vllm/pull/1198uvicorn by @danilopeixoto in https://github.com/vllm-project/vllm/pull/1166Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.7...v0.2.0
A minor release to fix the bugs in ALiBi, Falcon-40B, and Code Llama.
A minor release to fix the bugs in ALiBi, Falcon-40B, and Code Llama.
trust_remote_code=True by @Jingru in https://github.com/vllm-project/vllm/pull/871Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.6...v0.1.7
Note: This is an emergency release to revert a breaking API change that can make many existing codes using AsyncLLMServer not work.
Note: This is an emergency release to revert a breaking API change that can make many existing codes using AsyncLLMServer not work.
AsyncLLMEngine.generate by @Yard1 in https://github.com/vllm-project/vllm/pull/988Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.5...v0.1.6
Align beam search with hf_model.generate.
hf_model.generate.AsyncLLMEngine more robust & fix batched abort by @Yard1 in https://github.com/vllm-project/vllm/pull/969Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.4...v0.1.5
Fix for breaking changes in xformers 0.0.21 by @WoosukKwon in https://github.com/vllm-project/vllm/pull/834
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.3...v0.1.4
More model support: LLaMA 2, Falcon, GPT-J, Baichuan, etc.
KeyError when loading bloom-based models by @HermitSun in https://github.com/vllm-project/vllm/pull/441trust_remote_code in benchmark by @wangruohui in https://github.com/vllm-project/vllm/pull/518Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.2...v0.1.3
ChatCompletion endpoint in OpenAI demo server
Thanks to the following amazing people who contributed to this release:
@michaelfeil @WoosukKwon @metacryptom @merrymercy @BasicCoder @zhuohan123 @twaka @comaniac @neubig @JRC1995 @LiuXiaoxuanPKU @bm777 @Michaelvll @gesanqiu @ironpinguin @coolcloudcol @akxxsb
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.1...v0.1.2
Fix Ray node resources error by @zhuohan123 in https://github.com/vllm-project/vllm/pull/193
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.1.0...v0.1.1
Thanks @WoosukKwon @zhuohan123 @suquark for their contributions.
See our README for details.
Thanks @WoosukKwon @zhuohan123 @suquark for their contributions.
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →