NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #127 most downloaded on PyPI
SGLang is a fast serving framework for large language models and vision language models.
Last release today
18 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Some releases are documented
notes for 27 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
157 releases · first in 2024
[HiCache] Cleaning the deprecated host memory state by @xiezhq-hermann in https://github.com/sgl-project/sglang/pull/10778
init_incremental_detokenization by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10412--enable-custom-logit-processor: agree with cli arg by @thalahors in https://github.com/sgl-project/sglang/pull/10076get_load and shortest queue strategy by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10201--dataset-path in bench_one_batch_server by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10475kv_cache_quant_algo in ModelOpt checkpoints by @brayden-hai in https://github.com/sgl-project/sglang/pull/10336from sglang.python by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10493is_blackwell platform check by @b8zhong in https://github.com/sgl-project/sglang/pull/10498bench_one_batch_server by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10548ChatCompletionRequest and fix validation logic by @CatherineSue in https://github.com/sgl-project/sglang/pull/10675get_worker_urls_for_model in http/router.rs by @CatherineSue in https://github.com/sgl-project/sglang/pull/10754sglang.environ by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10771FutureMap by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10715environ into sglang.srt to avoid break SRT auto sync. by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10791build_constraint in grpc router by @CatherineSue in https://github.com/sgl-project/sglang/pull/10881grpc/client.rs to grpc_client/sglang_scheduler.rs by @CatherineSue in https://github.com/sgl-project/sglang/pull/10924get_ip by @merrymercy in https://github.com/sgl-project/sglang/pull/10883transformers: the error: AttributeError: 'TransformersForCausalLM' object has no attribute 'tp_size' by @vincentzed in https://github.com/sgl-project/sglang/pull/9614JsonParser and LlamaParser by @CatherineSue in https://github.com/sgl-project/sglang/pull/11073get_pooled in process_single_choice by @CatherineSue in https://github.com/sgl-project/sglang/pull/11079--repeat in run_eval by @hnyls2002 in https://github.com/sgl-project/sglang/pull/11101logprob_start_len = -1 by @CatherineSue in https://github.com/sgl-project/sglang/pull/11113.item() in paged allocator. by @hnyls2002 in https://github.com/sgl-project/sglang/pull/11156io_struct and base sglang io classes. by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10133thinking_mode in eval by @hnyls2002 in https://github.com/sgl-project/sglang/pull/11192skip_sample adjust by @hnyls2002 in https://github.com/sgl-project/sglang/pull/11225v1/responses to be more OpenAI-compatible. by @vincentzed in https://github.com/sgl-project/sglang/pull/9624Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.2...v0.5.3
One column per month.
Nothing published for this version
Nothing published for this version
[Vulnerability]feat(conn): set bootstrap server host by @jinmingyi1998 in https://github.com/sgl-project/sglang/pull/9931
response_format support for completion API by @cicirori in https://github.com/sgl-project/sglang/pull/9665topk>1 and page>1 for paged attention backends. by @hnyls2002 in https://github.com/sgl-project/sglang/pull/9784silu_and_mul_scaled_fp4_grouped_quant perf by @kaixih in https://github.com/sgl-project/sglang/pull/9556set_interal_state API by @hnyls2002 in https://github.com/sgl-project/sglang/pull/9850current_platform bug by @BBuf in https://github.com/sgl-project/sglang/pull/9919Router arguments passing and build it in docker image by @hnyls2002 in https://github.com/sgl-project/sglang/pull/9964get_fused_moe_impl_class function by @kaixih in https://github.com/sgl-project/sglang/pull/9764tokenizer_communicator_mixin by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10028is_alive timeout by @hnyls2002 in https://github.com/sgl-project/sglang/pull/10159Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.1...v0.5.2
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Improve LoRA Perf by Deprecating FlashInfer and Eliminating Redundant Tensor Ops by @lifuhuang in https://github.com/sgl-project/sglang/pull/8940
cuda < 12.4 by @hnyls2002 in https://github.com/sgl-project/sglang/pull/8701routed_scaling_factor is None by @hnyls2002 in https://github.com/sgl-project/sglang/pull/8709--attention-backend triton by @ch-wan in https://github.com/sgl-project/sglang/pull/8723renormalize=False in Triton kernels by @ch-wan in https://github.com/sgl-project/sglang/pull/8735once_cell with LazyLock in worker.rs and remove once_cell dependency from Cargo.toml by @htiennv in https://github.com/sgl-project/sglang/pull/8698get_col_major_tma_aligned_tensor for Blackwell deepgemm in EpMoE by @kaixih in https://github.com/sgl-project/sglang/pull/8955builtin_tools by @CatherineSue in https://github.com/sgl-project/sglang/pull/9114num_token_non_padded by @ch-wan in https://github.com/sgl-project/sglang/pull/9107sglang.srt.layers.quantization.scalar_types with sgl_kernel.scalar_type by @Hongbosherlock in https://github.com/sgl-project/sglang/pull/8951from python.sglang.srt -> from sglang.srt by @netanel-haber in https://github.com/sgl-project/sglang/pull/9268update_tensor_inplace function by @b8zhong in https://github.com/sgl-project/sglang/pull/6307CMakeLists.txt binary_dir by @EduardDurech in https://github.com/sgl-project/sglang/pull/7019--allow-auto-truncate argument in tokenizer manager. by @hnyls2002 in https://github.com/sgl-project/sglang/pull/9391required by @CatherineSue in https://github.com/sgl-project/sglang/pull/9525Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.10...v0.5.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This is a regular release with many new optimizations, features, and fixes. Please checkout the following exciting roadmaps and blogs
This is a regular release with many new optimizations, features, and fixes. Please checkout the following exciting roadmaps and blogs
--model as an alias for --model-path in server_args by @CatherineSue in https://github.com/sgl-project/sglang/pull/7505docs/references/production_metrics.md by @rudeigerc in https://github.com/sgl-project/sglang/pull/7741server_args.py & Fix session control & Other cleanups by @merrymercy in https://github.com/sgl-project/sglang/pull/7748pybase64 instead of base64 by @b8zhong in https://github.com/sgl-project/sglang/pull/7724get_model_info endpoint by @Arist12 in https://github.com/sgl-project/sglang/pull/7660srt/layer/quantization by @ch-wan in https://github.com/sgl-project/sglang/pull/7989select_experts by @ch-wan in https://github.com/sgl-project/sglang/pull/7966torch.compile in forward pass by @BBuf in https://github.com/sgl-project/sglang/pull/8353num_local_experts by @kaixih in https://github.com/sgl-project/sglang/pull/8453Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.8...v0.4.10
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Re-structured the OpenAI-compatible server to support production and enterprise environments. Key improvements include:
Re-structured the OpenAI-compatible server to support production and enterprise environments. Key improvements include:
Consistent metrics and logging for better observability and debugging.
Unified error handling, request validation, and processing logic for improved reliability and maintainability.
Improved request tracking across sessions and components.
Fixed bugs in embedding requests and reasoning parsers.
This work was a collaborative effort involving engineers from academic and industry institutions. Special thanks to the Oracle Cloud team and the SGLang team and community — including @slin1237, @CatherineSue, @key4ng, @JustinTong0323, @jhinpan, @yhyang201, @woodx9 and @whybeyoung — for their invaluable contributions.
Added support for DeepSeek R1 with FP4 and MTP on NVIDIA Blackwell GPU.
Integrated FlashInfer NVFP4 MoE, supporting TP, EP, and DP.
Supported 2-stream shared expert execution.
Achieved up to 90 TPS per user at isl/osl/bs = 1k/1k/16 on B200.
Further optimization in progress. Special thanks to the FlashInfer, NVIDIA Enterprise Products, Novita AI, DataCrunch, Google Cloud, and SGLang teams — especially @Alcanderian and @pyc96 — for their critical contributions.
The sglang/srt/openai_api directory has been removed and replaced with sglang/srt/entrypoints/openai.
Update your imports to the new module path. For example:
- from sglang.srt.openai_api.protocol import Tool
+ from sglang.srt.entrypoints.openai.protocol import Tool
INFO to DEBUG for dp and add force quit for tokenizer manager by @ishandhanani in https://github.com/sgl-project/sglang/pull/7251_normalize_rid before other normalization in io_struct by @CatherineSue in https://github.com/sgl-project/sglang/pull/7363openai_api with entrypoints/openai by @CatherineSue in https://github.com/sgl-project/sglang/pull/7351TokenToKVPoolAllocator by @hnyls2002 in https://github.com/sgl-project/sglang/pull/7414BaseFormatDetector.parse_streaming_increment by @CatherineSue in https://github.com/sgl-project/sglang/pull/7479Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.7...v0.4.8
Nothing published for this version
The PD disaggregation and large-scale EP functionalities from the blog post have now been fully merged into the latest release.
The PD disaggregation and large-scale EP functionalities from the blog post have now been fully merged into the latest release.
The blog has been successfully reproduced by over six industry teams, including the TensorRT LLM team.
SGLang’s large-scale EP is now actively used by leading organizations such as Cursor, Qwen, Alimama, Alibaba Cloud, iFlytek, and more. It has been deployed and validated at large scale, running on GPU clusters with thousands of devices.
PD disaggregation and large-scale EP, in addition to supporting DeepSeek V3/R1, now also support Qwen 3 in the latest release.
Full Blackwell support for DeepSeek V3/R1, Llama 4, and Qwen 3. Further optimizations are underway.
SGLang's DeepSeek V3/R1 now achieves 190 TPS on single H200, outperforming other frameworks by over 50%.
We extend our sincere thanks to the following contributors, listed in alphabetical order: Alibaba Cloud, AMD Team, Ant Group, Baseten Team, Cursor Team, Dynamo Team, EAGLE Team, FlashInfer Team, Google Vertex AI Team, iFlytek MaaS Team, Intel Team, LinkedIn Team, Meituan Team, Microsoft Copilot Team, Mooncake Team, NVIDIA Team, Oracle Team, Qwen Team, Voltage Park Team and open source community users. Your support and collaboration are deeply appreciated!
max_completion_tokens for OpenAIChatCompletions by @CatherineSue in https://github.com/sgl-project/sglang/pull/5857calculate_num_image_tokens from qwen2_vl.py by @JustinTong0323 in https://github.com/sgl-project/sglang/pull/5783chat_template_kwargs documentation by @vincentzed in https://github.com/sgl-project/sglang/pull/5679SGLANG_ and SGL_ environment variables by @b8zhong in https://github.com/sgl-project/sglang/pull/6206--moe-dense-tp-size=1 by @ch-wan in https://github.com/sgl-project/sglang/pull/5657launch_dummy_health_check_server to start inside of running asyncio loop by @ishandhanani in https://github.com/sgl-project/sglang/pull/6330return_hidden_states for the OpenAI API by @kyle-pena-kuzco in https://github.com/sgl-project/sglang/pull/6137return_hidden_states for the OpenAI API (#6137)" by @zhyncs in https://github.com/sgl-project/sglang/pull/6440Cargo.lock, add it into .gitignore by @hnyls2002 in https://github.com/sgl-project/sglang/pull/6438required and specific function mode by @CatherineSue in https://github.com/sgl-project/sglang/pull/6550parse_streaming_increment by @CatherineSue in https://github.com/sgl-project/sglang/pull/6715n_share_experts_fusion as num_fused_shared_experts by @ch-wan in https://github.com/sgl-project/sglang/pull/6735num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled by @ch-wan in https://github.com/sgl-project/sglang/pull/6736Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.6...v0.4.7
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Deprecate disable-mla by @Fridge003 in https://github.com/sgl-project/sglang/pull/5481
Thanks very much to LinkedIn team, Alibaba Cloud, Mooncake team, NVIDIA Team, AMD Team, Pytorch Team, Ant Group, Baseten Team, Oracle Team, Meituan Team, iFlytek MaaS team and the open source community users for their contributions!
We’re thrilled about these advancements and eager to hear your feedback! Join us on our Slack channel at slack.sglang.ai to connect and share your thoughts. Cheers!
bench_one_batch support enable_dp_attention by @fzyzcjy in https://github.com/sgl-project/sglang/pull/4058--enable-llama4-multimodal by @ch-wan in https://github.com/sgl-project/sglang/pull/5254compressed-tensors by @c8ef in https://github.com/sgl-project/sglang/pull/5640torch.full in DeepSeek by @fzyzcjy in https://github.com/sgl-project/sglang/pull/5601--n-share-experts-fusionwith R1 by @guoyuhong in https://github.com/sgl-project/sglang/pull/5707ArcticForCausalLM architecture (Snowflake/snowflake-arctic-instruct) by @b8zhong in https://github.com/sgl-project/sglang/pull/5078ArcticForCausalLM architecture (Snowflake/snowflake-arctic-instruct)" by @merrymercy in https://github.com/sgl-project/sglang/pull/5754404 Not Found error when running compile_deep_gemm.py in multi-node setups by @guoyuhong in https://github.com/sgl-project/sglang/pull/5720decrypted_config_file by @vincentzed in https://github.com/sgl-project/sglang/pull/5685sglang:v0.4.5.post3-rocm630. by @saienduri in https://github.com/sgl-project/sglang/pull/5697is not None instead of != None for None checks. by @vincentzed in https://github.com/sgl-project/sglang/pull/5687Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.5...v0.4.6
Nothing published for this version
Nothing published for this version
Nothing published for this version
The SGLang team is excited to the release of v0.4.5! This version introduces several significant features, including Llama 4 support, FlashAttention 3
The SGLang team is excited to the release of v0.4.5! This version introduces several significant features, including Llama 4 support, FlashAttention 3 backend, EAGLE3 speculative decoding, DeepEP integration, and disaggregated prefill and decoding.
Llama 4 Support: We supported Llama 4 model with accuracy matching official benchmark numbers, achieving a zero-shot score of 75.2 on the MMLU Pro dataset for Llama-4-Scout-17B-16E-Instruct model and 80.7 for Llama-4-Maverick-17B-128E-Instruct model. https://github.com/sgl-project/sglang/pull/5092
FlashAttention 3 Backend: Our implementation of the FlashAttention 3 backend delivers significant acceleration for long-context tasks. https://github.com/sgl-project/sglang/issues/4709
EAGLE3 Speculative Decoding: We’re proud to be the first to support EAGLE3 speculative decoding, offering substantial gains in decoding throughput. Learn more in our documentation and the EAGLE3 paper. https://github.com/sgl-project/sglang/pull/4247
DeepEP Integration: By incorporating DeepEP, we enhanced performance for MoE inference.
Disaggregated Prefill and Decoding: We introduced a prototype for disaggregated prefill and decoding, with plans for further optimizations.
Thanks very much to the NVIDIA team, LinkedIn team, EAGLE team, Oracle team, Meituan team, and our incredible open-source community for their invaluable contributions!
Disaggregated Prefill and Decoding: https://github.com/sgl-project/sglang/issues/4655
Llama 4 Optimization: https://github.com/sgl-project/sglang/issues/5118
EP Enhancement: https://github.com/sgl-project/sglang/issues/4734
FA3 Enhancement: https://github.com/sgl-project/sglang/issues/4709
We’re thrilled about these advancements and eager to hear your feedback! Join us on our Slack channel at slack.sglang.ai to connect and share your thoughts. Cheers!
torch.cat instead of torch.concat to prevent entering the Autograd backends. by @Alcanderian in https://github.com/sgl-project/sglang/pull/4466torch.inference_mode() instead of torch.no_grad() by @Alcanderian in https://github.com/sgl-project/sglang/pull/4372n in OpenAI API completions by @ChuyueSun in https://github.com/sgl-project/sglang/pull/3446SGLANG_LOGGING_CONFIG_PATH by @guoyuhong in https://github.com/sgl-project/sglang/pull/4592--attention-backend fa3 by @hebiao064 in https://github.com/sgl-project/sglang/pull/4680reply to replay in base_attn_backend.py by @Thysrael in https://github.com/sgl-project/sglang/pull/4784get_attention_sliding_window_size for attn init by @vhain in https://github.com/sgl-project/sglang/pull/4823gguf and torchvision by @vhain in https://github.com/sgl-project/sglang/pull/4826neuralmagic/gemma-2-2b-it-FP8 by @merrymercy in https://github.com/sgl-project/sglang/pull/4830self.worker assignment in TpModelWorker and refactor references by @JustinTong0323 in https://github.com/sgl-project/sglang/pull/4788mem_fraction_static for gemma3 vision test by @vhain in https://github.com/sgl-project/sglang/pull/4840import vllm in quantization/init.py by @merrymercy in https://github.com/sgl-project/sglang/pull/4834cuda_device_count_stateless by @Alcanderian in https://github.com/sgl-project/sglang/pull/5060Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.4...v0.4.5
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
The SGLang team is excited to announce the release of v0.4.4. We will keep improving DeepSeek V3/R1 performance. With the combination of FlashInfer, M
The SGLang team is excited to announce the release of v0.4.4. We will keep improving DeepSeek V3/R1 performance. With the combination of FlashInfer, MTP, DeepGEMM, and Torch Compile optimizations on H200, it can achieve nearly 100 tokens/s, which is currently the fastest open-source implementation. Look out for new optimizations coming soon!
Thanks very much to xAI Team, NVIDIA Team, AMD Team, LinkedIn team, Baseten Team, Oracle Team, Meituan Team and the open source community users for their contributions!
Regarding the use of SGLang for DeepSeek R1 inference acceleration, in addition to the users mentioned in the announcement, there are also teams such as Tencent and Ant Group. We are very happy to have received recognition and usage from these teams!
Though surely there will be bugs and fixes that we'll be discovering and quickly patching in the coming days, including today :) Let's build and ship. Please feel free to join our Slack channel https://slack.sglang.ai/ Cheers!
AMD Performance Leadership: SGLang is now the fastest LLM engine for DeepSeek V3/R1 inference on AMD hardware, as confirmed by AMD's technical blog
Enhanced FlashInfer MLA Support: Now fully compatible with radix cache, chunked prefill, and MTP optimizations - enable with
--enable-flashinfer-mla
Advanced MTP Capabilities: Both Triton and FlashInfer backends now offer comprehensive Multi-Token Prediction support, easily tunable via the bench_speculative script, compatible with radix cache and chunked prefill.
DeepGEMM Integration: Full integration of DeepGEMM for NVIDIA Hopper architectures - enable with
export SGL_ENABLE_JIT_DEEPGEMM=1
Pioneering INT8 Quantization: First industry implementation of INT8 support for DeepSeek R1 models:
Other Optimizations:
Blackwell architecture Block Scale FP8 GEMM support
Support page size greater than 1 https://github.com/sgl-project/sglang/pull/4356
Optimized W8A8 FP8 implementation with performance gains across all architectures (sm80, sm89, sm90), featuring 15%+ improvement specifically on sm89
Enhanced distributed parallelism capabilities (e.g., two-node configurations with DP 2, TP 16) https://github.com/sgl-project/sglang/pull/4390
Integrate Flash Attention https://github.com/sgl-project/sglang/issues/4385
Integrate FlashMLA https://github.com/sgl-project/sglang/issues/4384
EAGLE 2 optimization https://github.com/sgl-project/sglang/pull/4383
EAGLE 3 day one support https://github.com/sgl-project/sglang/pull/4247
Integrate DeepEP https://github.com/sgl-project/sglang/pull/4232
Prefill and Decoding Disaggregation
tl.range() in block GEMM kernels with num_stages set by host. by @whchung in https://github.com/sgl-project/sglang/pull/3535tl.range() in block GEMM kernels with `num_stage… by @zhyncs in https://github.com/sgl-project/sglang/pull/3632debug_tensor_dump_output_folder optional key missing by @Qubitium in https://github.com/sgl-project/sglang/pull/4046lm_head Quantization by @Qubitium in https://github.com/sgl-project/sglang/pull/3790__init__ function of model_runner.py shorter by @merrymercy in https://github.com/sgl-project/sglang/pull/4132aiohttp into public dependencies by @stevapple in https://github.com/sgl-project/sglang/pull/3980torch.compile by @junliu-mde in https://github.com/sgl-project/sglang/pull/3844Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.3...v0.4.4
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
The SGLang team is excited to announce the release of v0.4.3. We will keep improving DeepSeek V3/R1 performance. In the last six weeks, SGLang has bee
The SGLang team is excited to announce the release of v0.4.3. We will keep improving DeepSeek V3/R1 performance. In the last six weeks, SGLang has been the fastest engine running DeepSeek V3/R1 among all open-source LLM inference engines. We stay ahead by integrating FlashInfer MLA and optimizing further. Look out for new optimizations coming soon! Please feel free to join our Slack channel https://slack.sglang.ai Cheers!
padding circular import by @BBuf in https://github.com/sgl-project/sglang/pull/2624update_weights_from_tensor by @fzyzcjy in https://github.com/sgl-project/sglang/pull/2631update_weights_from_tensor by @fzyzcjy in https://github.com/sgl-project/sglang/pull/2695--enable-metrics by @CatherineSue in https://github.com/sgl-project/sglang/pull/2822-inf by @ByronHsu in https://github.com/sgl-project/sglang/pull/3212-inf by @ByronHsu in https://github.com/sgl-project/sglang/pull/3224Full Changelog: https://github.com/sgl-project/sglang/compare/v0.4.1...v0.4.3
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →