NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #127 most downloaded on PyPI
SGLang is a fast serving framework for large language models and vision language models.
Last release 2 days ago
18 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Some releases are documented
notes for 27 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
157 releases · first in 2024
We're excited to announce SGLang v0.4.1, which now supports DeepSeek V3 - currently the strongest open-source LLM, even surpassing GPT-4o.
We're excited to announce SGLang v0.4.1, which now supports DeepSeek V3 - currently the strongest open-source LLM, even surpassing GPT-4o.
The SGLang and DeepSeek teams worked together to get DeepSeek V3 FP8 running on NVIDIA and AMD GPU from day one. We've also supported MLA optimization and DP attention before, making SGLang one of the best open-source LLM engines for running DeepSeek models.
Special thanks to Meituan's Search & Recommend Platform Team @ispobock @HandH1998 and Baseten's Model Performance Team @zhyncs for implementing the model, and DataCrunch for providing GPU resources.
Various improvements to the cache-aware sglang router, torchao integration, server termination
Added a standalone package sgl-kernel for supporting more custom kernels in the code base.
/add_worker api by @ByronHsu in https://github.com/sgl-project/sglang/pull/2369Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.0...v0.4.1
One column per month.
Nothing published for this version
Nothing published for this version
blog: https://lmsys.org/blog/2024-12-04-sglang-v0-4/
blog: https://lmsys.org/blog/2024-12-04-sglang-v0-4/
We’re excited to release SGLang v0.4, featuring significant performance improvements and new features:
Scheduler.update_running_batch by @merrymercy in https://github.com/sgl-project/sglang/pull/2154is not not != to test None by @WrRan in https://github.com/sgl-project/sglang/pull/2196Full Changelog: https://github.com/sgl-project/sglang/compare/v0.3.6...v0.4.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Deprecate --disable-flashinfer and --disable-flashinfer-sampling by @merrymercy in https://github.com/sgl-project/sglang/pull/2065
get_cuda_graph_seq_len_fill_value by @merrymercy in https://github.com/sgl-project/sglang/pull/1783ZMQ buffer size heuristic by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1801SGLANG_CPU_COUNT by @ByronHsu in https://github.com/sgl-project/sglang/pull/1803engine.generate by @ByronHsu in https://github.com/sgl-project/sglang/pull/1820use_thread in the run_program for easier debugging. by @liuyanyi in https://github.com/sgl-project/sglang/pull/1823zmq Version Requirement by @HuanzhiMao in https://github.com/sgl-project/sglang/pull/1982--disable-nan-detection to --enable-nan-detection by @merrymercy in https://github.com/sgl-project/sglang/pull/2066Full Changelog: https://github.com/sgl-project/sglang/compare/v0.3.4.post1...v0.3.6
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Hosted the first LMSYS online meetup: Efficient LLM Deployment and Serving.
logprobs_nums by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1548atexit hook to implicitly shutdown Runtime by @ByronHsu in https://github.com/sgl-project/sglang/pull/1595PortArgs.init_new by @glen-amd in https://github.com/sgl-project/sglang/pull/1611--num-continuous-decode-steps as an argument by @merrymercy in https://github.com/sgl-project/sglang/pull/1652CacheConfig import in all model files by @ByronHsu in https://github.com/sgl-project/sglang/pull/1658is_all_ready for overlap copy by @merrymercy in https://github.com/sgl-project/sglang/pull/1710max_req_len and max_req_input_len by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1748Full Changelog: https://github.com/sgl-project/sglang/compare/v0.3.2...v0.3.4.post1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Deprecate --disable-flashinfer and introduce --attention-backend by @merrymercy in https://github.com/sgl-project/sglang/pull/1380
model_override_args to launch_server via the CLI. by @kevin85421 in https://github.com/sgl-project/sglang/pull/1298undefined is_single in meth create_abort_task by @wcsjtu in https://github.com/sgl-project/sglang/pull/1370Full Changelog: https://github.com/sgl-project/sglang/compare/v0.3.0...v0.3.2
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Checkout the release blog post https://lmsys.org/blog/2024-09-04-sglang-v0-3/ to find detailed instructions and descriptions for the items below.
Checkout the release blog post https://lmsys.org/blog/2024-09-04-sglang-v0-3/ to find detailed instructions and descriptions for the items below.
Full Changelog: https://github.com/sgl-project/sglang/compare/v0.2.13...v0.3.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
New Feature: Support window attention for Gemma-2 (#1056 #1090 #1112), enable chunked-prefill by default (#1040 #984), support all sampling penalties
get_new_prefill_batch by @hnyls2002 in https://github.com/sgl-project/sglang/pull/948req_pool_indices on CPU by @hnyls2002 in https://github.com/sgl-project/sglang/pull/960InputeMetadata and ScheduleBatch by @hnyls2002 in https://github.com/sgl-project/sglang/pull/981input_ids && rename to fill_ids by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1021dtype to control generate by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1082stop_token_ids in sglang API by @hnyls2002 in https://github.com/sgl-project/sglang/pull/1092Full Changelog: https://github.com/sgl-project/sglang/compare/v0.2.9...v0.2.13
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
New feature: Chunked prefill (#800, #811)
--max-total-tokens by @hnyls2002 in https://github.com/sgl-project/sglang/pull/840/test/srt as unit tests by @Ying1123 in https://github.com/sgl-project/sglang/pull/875Full Changelog: https://github.com/sgl-project/sglang/compare/v0.2.5...v0.2.9
Nothing published for this version
Nothing published for this version
Nothing published for this version
We recently released a blog. Compared to TensorRT-LLM and vLLM, SGLang Runtime consistently delivers superior or competitive performance in both onlin
We recently released a blog. Compared to TensorRT-LLM and vLLM, SGLang Runtime consistently delivers superior or competitive performance in both online and offline scenarios, handling models from Llama-8B to Llama-405B, and on A100 and H100 GPUs, using FP8 and FP16. SGLang consistently outperforms vLLM, achieving up to 3.1x higher throughput on Llama-70B. It also often matches or sometimes outperforms TensorRT-LLM.
We have now automated the release processes for PyPI, Docker, and Release using GitHub workflows. Previously, because Release was not automated, GitHub Tags were not updated in time, leading to a jump from v0.2.0 directly to v0.2.5.
Welcome everyone to try using https://github.com/sgl-project/sglang, and also welcome everyone to actively participate in the community, including but not limited to issues, PRs, and discussions. Cheers!
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
We performed extensive engineering to improve the base performance. Compared to TensorRT-LLM and vLLM, SGLang now consistently delivers superior or co
global_server_args_dict by @hnyls2002 in https://github.com/sgl-project/sglang/pull/642TokenizerManager.context_len should inherit from `server_args.conte… by @shrirajh in https://github.com/sgl-project/sglang/pull/654Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.20...v0.2.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Code clean up: Remove deprecated prefill move InputMetadata to infer_batch.py by @merrymercy in https://github.com/sgl-project/sglang/pull/609
--enable-p2p-check option by @hnyls2002 in https://github.com/sgl-project/sglang/pull/599LogitsMetadata by @hnyls2002 in https://github.com/sgl-project/sglang/pull/604sgl.gen API by @huyiwen in https://github.com/sgl-project/sglang/pull/503Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.18...v0.1.20
Nothing published for this version
2x large batch prefill improvement with the new flashinfer kernels #579
Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.17...v0.1.18
Add speculative execution for OpenAI API #250
--disable-radix-cache by @hnyls2002 in https://github.com/sgl-project/sglang/pull/451Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.16...v0.1.17
Support more models: DBRX, Command-R, Gemma
model_rpc style improvement by @hnyls2002 in https://github.com/sgl-project/sglang/pull/293model_runner simplify by @hnyls2002 in https://github.com/sgl-project/sglang/pull/329DBRX support by @hnyls2002 in https://github.com/sgl-project/sglang/pull/337command-r by @ZhouXingg in https://github.com/sgl-project/sglang/pull/369fork(1) by @hnyls2002 in https://github.com/sgl-project/sglang/pull/375.isort.cfg by @hnyls2002 in https://github.com/sgl-project/sglang/pull/378sync() when fork(1) by @hnyls2002 in https://github.com/sgl-project/sglang/pull/412input_ids in the input of the /generate endpoint by @lolipopshock in https://github.com/sgl-project/sglang/pull/363Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.13...v0.1.16
Nothing published for this version
Nothing published for this version
Gemma Support by @hnyls2002 in https://github.com/sgl-project/sglang/pull/256
get_var(var_name) in text iter when stream is not enabled by @exceedzhang in https://github.com/sgl-project/sglang/pull/198agent_calls.jsonl download link by @hnyls2002 in https://github.com/sgl-project/sglang/pull/226set_var to interpreter.py by @1024th in https://github.com/sgl-project/sglang/pull/263server_args by @hnyls2002 in https://github.com/sgl-project/sglang/pull/277Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.12...v0.1.13
Output logprobs for decoding tokens
--disable-disk-cache by @hnyls2002 in https://github.com/sgl-project/sglang/pull/160Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.11...v0.1.12
Serve the official release demo of LLaVA v1.6 blog
is_multimodal_model judge by @hnyls2002 in https://github.com/sgl-project/sglang/pull/132Full Changelog: https://github.com/sgl-project/sglang/compare/v0.1.6...v0.1.11
Nothing published for this version
Nothing published for this version
Nothing published for this version
Add OpenAI-compatible API server (Completion and ChatCompletion)
sgl.selectFull Changelog: https://github.com/sgl-project/sglang/compare/v0.1.5...v0.1.6
Your coding agent can read these notes before it upgrades. Set up the MCP server →