NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2365 most downloaded on PyPI
SGLang is a fast serving framework for large language models and vision language models.
Last release 3 days ago
01 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Some releases are documented
notes for 28 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
158 releases · first in 2024
Breaking Changes & Upgrade Notes
779 PRs from 227 contributors.
| Model | Type | Cookbook |
|---|---|---|
| DeepSeek-V4.1 Flash | LLM / VLM | link |
| GigaChat 3.5 | LLM / VLM | link |
| IQuest-Q1 | LLM / VLM | link |
| MiMo-V2.6 / MiMo-V2.6-Pro | LLM / VLM | link |
| Ling-3.0-flash-VL | LLM / VLM | link |
| DiffusionGemma | Diffusion | link |
| Qwen-Image 2.1 | Diffusion | link |
| Anima Base v1.0 | Diffusion | link |
| Ming-Image 0.1 Design / Design-Layer | Diffusion | link |
| FLUX 3 Action | Diffusion | link |
/v1/decisions) turns an LLM / VLM into a low-latency classifier and scorer (docs, #41208)./v1/score) scores all candidates in one request (#38965, #41188).uv pip install --prerelease=allow sglang==0.5.21| Platform | Docker image |
|---|---|
| NVIDIA (CUDA 13) | lmsysorg/sglang:v0.5.21 |
| AMD MI35x | lmsysorg/sglang:v0.5.21-rocm10-mi35x |
| AMD MI30x | lmsysorg/sglang:v0.5.21-rocm10-mi30x |
| Intel GPU | lmsysorg/sglang:v0.5.21-xpu |
| Intel CPU | lmsysorg/sglang:v0.5.21-xeon |
TreeSpeculativeSamplingTargetOnly: #35798none decode retraction backup and subclass seams in the PD queues: #41103/server_info: #39993cpp_extension loader: #40989RadixCache no-insert cleanup with kv_len_to_handle: #35204owned_kv_len on radix cache insert: #40075send_with_url, fix broken tests: #40502/v1/decisions route it builds on): #41208The experimental Rust router and frontend (experimental/sgl-router, the Rust renderer) keep moving toward parity with the Python server.
Note truncated.
One column per month.
Full release notes by category below; breaking changes are at the end.
713 PRs from 237 contributors.
New models in this release (see the cookbook for all supported models):
| Model | Type | PRs | Cookbook |
|---|---|---|---|
| GLM-5.3-Flash | Autoregressive | #36507, #38621 | link |
| Hy4-Preview | Autoregressive | #36805 | link |
| Qwen3.8-Flash-Next | Autoregressive | #37500 | link |
| K2 Horizon | Autoregressive | #37654, #38033 | link |
| Nanbeige4.2 | Autoregressive | #32151 | |
| SenseNova-U1.5-8B-MoT | Diffusion | #36606 | link |
| FastH3 (4-step MiniMax-H3 distill) | Diffusion | #37480 | link |
| VDN-H3 (hybrid-attention MiniMax-H3 distill) | Diffusion | #37903 | link |
Sampling masks for RL rollouts. With return_sampling_mask, each decode step returns the exact token support the sampler drew from and the log-probability of the sampled token under it, so a trainer can replay the rollout without reconstructing top-k or top-p (#36630). Masks now run under overlap scheduling: on Qwen3-8B, decode throughput is 17% higher at batch 1 and 52% higher at batch 64 than the previous implementation. Capacity is set by --sampling-mask-max-tokens (default 4096) (#36631). DisallowedTokensLogitsProcessor is supported alongside masks (#38279).
Unified radix tree. Branching-point caching for the SWA component keeps the sliding-window state at the point where requests fork from a shared prefix, so branches reuse it instead of recomputing. On DeepSeek-V4-Flash with a shared system prompt, token hit rate rises from 43.8% to 60.8% and mean TTFT falls from 1.57 s to 1.07 s (#34565). An opt-in external linker lets the tree address a shared global memory pool through Mooncake or UMBP (#37381).
DSpark under PD with decode context parallelism. A DCP1 prefill can now transfer its DSpark draft KV to a DCP-N decode, so hybrid models such as Kimi-Linear run DSpark in disaggregated, context-parallel serving. Verified on 8x B300 over NIXL and Mooncake up to 256K input (#37709).
Responses API storage is opt-in. /v1/responses no longer retains results in memory unless the server starts with --enable-response-store. Without it, retrieval, previous_response_id chaining, and background requests return 400; PD deployments cannot enable it (#39122).
SGLang Simulator. A CPU-only simulator runs the real scheduler, radix cache, and hierarchical cache with a latency predictor in place of the model forward. Against measured serving traces it predicts TTFT within about 6% on most traces (up to 10% on the longest 32K to 128K ones) and prefix reuse within 0.05 percentage points, for cache and scheduling studies without GPUs (#33824).
Prefill context parallelism v1 removed. The strategy-based implementation is now the only prefill CP path; the v1 runtime and its CLI options are gone. Prefill CP on HIP, NPU, and MUSA is rejected until those platforms are ported (#36228).
Faster model loading on ROCm. Large pageable host-to-device copies are staged rather than pinned in place, which stops the driver from suspending GPU queues on every eviction. GLM-5.2 at TP4 on 4x MI355X loads in 40.4 s instead of 505.7 s (#37720).
Intel XPU joins the release images. Every tagged release now builds and publishes lmsysorg/sglang:vX.Y.Z-xpu from the XPU Dockerfile, so Intel GPU users get a versioned image instead of relying on nightly builds (#37340).
DeepSeek-V4 on Blackwell. TRT-LLM attention kernels now cover DeepSeek-V4's CSA and HCA layers on SM100 and SM103: about 1.2x faster prefill and 1.45x faster decode than FlashMLA at the kernel level on B200 (#30805). FlashInfer MegaMoE is available as --moe-runner-backend flashinfer_megamoe; on DeepSeek-V4-Flash NVFP4 at TP4/DP4 it adds up to 11.9% prefill throughput at 8192 tokens per rank, with decode within 2% of the trtllm runner at saturation (#31470).
DeepSeek-V4 on RTX PRO 6000. On SM120 the sparse-MLA indexer now runs on DeepGEMM's paged-MQA kernel and the DeepGEMM FP4 MoE backend is enabled, replacing the torch fallback that was the only working path. On 4x RTX PRO 6000, DeepSeek-V4-Flash decode TPOT drops from 36.1 to 10.5 ms at batch 1 and TTFT falls 20% from 8K to 128K input. Opt-in through the environment flags in the PR (#29927).
Dependencies and images. The CUDA 12 lane is retired; v0.5.19 was the last release with -cu12x wheels and images (#38404). sglang-kernel moves to 0.4.7 (#39346) and sgl-deep-gemm to 0.2.0 (#39371). New images: ROCm 10 for MI30x and MI35x with a matching kernel wheel (#38763), gfx1151 for Strix Halo / Ryzen AI MAX+ (#33939), and Moore Threads MUSA (#36709). ROCm 7.0 CI, images, and kernel wheel are retired (#38632, #38767).
Full release notes by category below; breaking changes are at the end.
SGLANG_DISAGGREGATION_ENGINE_INIT_TIMEOUT: #37874DynamicChunkSizer scheduler component: #37674--schedule-policy hrrn; mean TTFT -69% and p99 -8% vs FCFS on a GLM-5.2 production trace): #32911all_reduce in check_hicache_events for PP (TTFT -7% on DeepSeek-V4-Flash with HiCache L3, PP4 TP2): #37562free_swa sync-free on page_size == 1: #36723torch.unique sync from the SWA page expansion: #37463free_segment and drop the boundary trim: #37729free_segment: #37876allocator/ and split the composites out: #38072page_size > 1: #38159<image> sentinel on artifact fast path: #39278Note truncated.
Unified radix tree by default. The unified tree is now the cache for every model, not just hybrid ones ( #35081 , see Breaking Changes). It also picke…
786 PRs from 214 contributors.
New models in this release (see the cookbook for all supported models):
| Model | Type | PRs | Cookbook |
|---|---|---|---|
| Qwen3.8 (2.4T-A95B) | Autoregressive | #35758 ⭐ | link |
| Qwen3.8-27B | Autoregressive | #34859 | link |
| dots3.note | Autoregressive | #33829 ⭐ | link |
| Ling-3.0-flash | Autoregressive | #33561 | link |
| Ling-3.0-tiny | Autoregressive | #33561 | link |
| Spark2.5 | Autoregressive | #35963 ⭐ | |
| MiniCPM-SALA | Autoregressive | #30360 | |
| Granite 4.2 | Autoregressive | #36286 | link |
| LongCat-Image-Edit & Edit-Turbo | Diffusion | #35829 |
Cookbook updates:
Beam search. SGLang can now do beam search. Pass beam_width in your request and you get back the n best sequences instead of a single sample. It works out of the box next to regular requests, though it does not yet mix with speculative decoding, disaggregation, DP attention, or HiCache (#31626).
DeepEP v2. DeepEP's new ElasticBuffer engine is available as --moe-a2a-backend deepep_v2 for DeepSeek-V3/V4 and Qwen3-MoE in FP8. Its buffers have a fixed size, so decode can run under CUDA graphs even across nodes. Performance is on par with the classic backend (#35634, #34923).
LayerNorm sequence parallelism. With --enable-layernorm-sp, each tensor-parallel rank normalizes only its own share of the prefill tokens instead of all of them. That takes 3.5% off Qwen3-8B prefill on H100 and 5.6% on B200, and the saving grows with the TP degree. Dense Qwen3 models only for now (#30915).
W4A8 MoE on Hopper. If you serve MXFP4 experts on Hopper, you can now quantize the activations to FP8 as well with --flashinfer-mxfp4-moe-precision fp8. DeepSeek-V4-Flash gains about 12% output throughput with no change in GSM8K accuracy. Needs FlashInfer 0.6.18 (#34967).
DCP on the default Blackwell MLA backend. Decode context parallelism now runs on trtllm_mla, not just CuTe DSL and Tokenspeed. It pays off at long context: at 128K input, plain TP stops scaling around 680 tokens per second on eight B200s, while DCP keeps going as concurrency grows (#33926).
Faster speculative kernels. DSA prefill top-k moves to the v2 kernel, 1.3 to 1.8 times faster on B200 (#35175). KDA models get an opt-in fused accept path, SGLANG_OPT_KDA_FUSED_ACCEPT_STATE=1, that cuts MTP verify-and-commit time by 45% to 63% on Kimi-Linear shapes with bit-identical output (#33722).
Unified radix tree by default. The unified tree is now the cache for every model, not just hybrid ones (#35081, see Breaking Changes). It also picked up three things this cycle: PD decode workers can reuse cached prefixes for SWA hybrid models like gpt-oss (#27770), you can attach or detach L3 storage on a running server (#35269), and pipeline parallelism with HiCache L3 stays consistent across ranks (#27010).
Lean attention on AMD. Long or uneven decode batches used to leave many compute units idle on MI300X and MI355X. The new persistent Lean kernel spreads the work across all of them, for up to 1.52x more throughput and up to 3.62x lower inter-token latency on MI355X. It turns on by itself where it helps, and SGLANG_DISABLE_LEAN_ATTENTION=1 turns it off (#33576).
DSA models on ROCm. Disaggregated GLM-5.2 serving now uses the fused top-k seed remap, which brings decode TPOT from 23 ms down to 8 ms on eight MI355Xs (#36714). DeepSeek-V4 gets the v2 top-k kernel, up to three times faster (#36684), and a shared-experts gate fix gives GLM-5.2 up to 16% better TPOT (#36124).
Dependencies. FlashInfer moves to 0.6.18 (#36954), sgl-deep-ep to 0.1.2 (#35450), sgl-deep-gemm to 0.1.7 (#37279), and mooncake to 0.3.13 (#36493). There is a new CUDA 13.4 preview image for Rubin (#36233) and new ROCm 10 images for gfx942, gfx950, and gfx1250 (#36434, #36871).
Full release notes by category below; breaking changes and known issues are at the end.
KimiK3ForConditionalGeneration: #36211<optional> include in AITER topk kernel: #36216Note truncated.
The first launch after upgrading recompiles once; see Breaking Changes ( #32434 ).
710 PRs from 212 contributors.
New models in this release (see the cookbook for all supported models):
| Model | Type | PRs | Cookbook |
|---|---|---|---|
| Muse Glimmer | Autoregressive (Multimodal) | #34262 | link |
| Intern-S2-Mobius | Autoregressive | #33691 | link |
| SANA-Video | Diffusion | #32921 | link |
| LingBot-Video-MoE | Diffusion | #32341 | link |
| LTX-2.5 | Diffusion | #34471 | link |
| Cosmos3 Edge & Distilled | Diffusion | #31590 | link |
| LongCat-Image | Diffusion | #23274 |
Plus cookbook recipes for the Qwen3.8 family, Ling-3.0, Nemotron 3.5 Lightning, Dots3-Note, and DeepSeek-V4-Pro-0813 (#34809).
Overlapped checkpoint staging at startup: Checkpoint pages now stage from storage while CUDA graphs capture. Qwen3-32B on H100 starts 8.6-11.7% faster than serial with prefetch, and 2.38x faster (35.6s vs 84.8s) than the plain default. Opt in with --startup-weight-load-mode overlap (#32017).
TP LMHead with All-to-All: The TP LMHead's allgather + scatter becomes a single all-to-all for pure-DP dp-attention. On DeepSeek-V4-Pro B200 decode, LMHead time drops 320us to 169us and TPOT improves 36.97ms to 35.67ms (#32313).
FlashInfer MNNVL for pure allreduce: Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL. DeepSeek-V4-Flash TP4 decode on Blackwell gains up to +6.9% at small batches. Auto-enabled for DeepSeek-V3/V3.2/V4; elsewhere --enable-flashinfer-pure-allreduce (#30700).
NVFP4 checkpoints run on AMD: --quantization quark_mxfp4 dequantizes ModelOpt and Quark NVFP4 weights and requantizes to MXFP4 at load, never holding a full-precision copy. 97.5-100.2% GSM8K recovery vs the NVFP4 reference across MiniMax-M2.7, GLM-5.1, Kimi-K2.6, Qwen3.5-397B, and DeepSeek-R1 (#29328).
Kimi K3 tuned for MI355X: a grouped-head MLA verify kernel replaces the MHA-shaped split-KV path that re-read the shared latent once per head, for 1.37-1.77x throughput and 1.45-2.42x ITL at concurrency 2-32; the AITER MLA prefill kernel now accepts K3's 12-head shape (TTFT up to -14.9%); a gfx950-tuned decode stage-1 geometry adds 46-73% ITL at 68k input, opt-in via SGLANG_MLA_DECODE_TUNE=1. GSM8K holds at 0.951-0.957 (#33981, #34261, #34837, #34580).
One compiled-kernel cache directory: Triton, FlashInfer, Inductor, DeepGEMM, and CUDA driver caches all move under SGLANG_CACHE_DIR. The first launch after upgrading recompiles once; see Breaking Changes (#32434).
Dependencies: torch 2.13.0 with triton 3.7.1 (#28836), flashinfer 0.6.17 (#33997), CuTeDSL 4.6.2, fixing an FA4 startup regression on Blackwell (#34372), DeepEP now installed from released sgl-deep-ep wheels (#33932), and sgl-kernel 0.4.6.post1 (#33842).
Full release notes by category below; breaking changes and known issues are at the end.
req.prefix_indices when the prefix cache is disabled: #34644is_image_understandable_model: #34217cuda-tile to 1.6.0rc5 to unblock Python 3.10 x86_64 installs: #34321Note truncated.
Your coding agent can read these notes before it upgrades. Set up the MCP server →