NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #127 most downloaded on PyPI
SGLang is a fast serving framework for large language models and vision language models.
Last release 13 days ago
04 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Some releases are documented
notes for 26 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
156 releases · first in 2024
…models ( #26997 , #26972 , #27463 ). Spec V1 is deprecated, with EAGLE/MTP now running on the unified V2 worker ( #25464 ), and topk = 1 drafting is f…
New Model Support:
Spec V2 is now the default speculative-decoding path: Tree drafting with topk > 1 is production-ready across the triton / FA3 / MLA / aiter backends, including page_size > 1 and Mamba/hybrid-linear models (#26997, #26972, #27463). Spec V1 is deprecated, with EAGLE/MTP now running on the unified V2 worker (#25464), and topk = 1 drafting is faster (#26397, #26424).
Lower per-step scheduler overhead: Unified async value passing through FutureMap plus moving prefill input transfer onto the forward stream reduced per-step launch overhead and improved stability under high concurrency (#25945, #25879, #26380).
Piecewise & Breakable CUDA Graph coverage: Piecewise (PCG) and Breakable (BCG) CUDA Graph capture more of the model to cut per-step kernel-launch overhead, now extended to DSA models, Kimi-K2.5, and DeepSeek V4: #23351, #26382, #25195.
Faster Qwen 3.5 on Blackwell: New FlashInfer Gated DeltaNet (GDN) kernels and a CuTeDSL GDN prefill kernel speed up Qwen 3.5 on Blackwell GPUs: #22921, #23273, #26200.
HiCache for hybrid models by default: HybridModel (SWA/Mamba) launches HiCache through UnifiedTree by default, bringing hierarchical KV-cache offload to sliding-window and Mamba hybrids out of the box: #27759.
Heterogeneous CPU + GPU EPD disaggregation (with Intel): Offload VLM vision encoding onto Intel Xeon CPUs alongside GPUs, with up to ~1.3x P99 TTFT and request-throughput gains under load. (blog)
MoRI on AMD Instinct MI355X (with AMD): Cost-competitive DeepSeek-R1 disaggregated inference via AMD's MoRI communication library, $0.169 per million tokens at 129 tok/s/user. (blog)
DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
flash_mla_sparse_fwd: #25418sgl_kernel.flashmla: #26421, #26132See the DeepSeek-V4 cookbook for tuned deployment commands.
SGLang-Diffusion — realtime & progressive resolution: OpenAI-style realtime video generation with msgpack frame streaming and a standalone browser WebUI (#26954, #26959), continuous camera controls + super-resolution controls (#27026, #27297), and progressive-resolution growing across FLUX / FLUX.2 / Qwen-Image / Wan / Z-Image (#27524).
Full release notes by category below.
flash_mla_sparse_fwd kernel: #25418page_size > 1 and Mamba/hybrid-linear, validated across triton / FA3 / MLA / aiter: #26997, #26972, #27463num_steps + observability metrics: #24055, #25940mha attention backend: #25002end_layer (Qwen3-Next disagg): #25476experimental_sgl_trtllm MoE backend for FP8 and NVFP4 models: #27329.any().item() guard in the LoRA MoE prefill path: #25531/v1/loads: #25440Realtime diffusion
Progressive resolution
Memory & residency
Quantization & backends
Performance
cuda_graph_kv_indices OOB under page_size > 1: #24587split_qkvgate_gemma_rmsnorm_rope for Qwen3.5 and Qwen3-Next: #23925torch.compile: #25256sgl_kernel.flashmla + DeepSeek V4 kernels): #26421, #26132No security-tagged PRs in this release.
All PRs included in this release: v0.5.12...v0.5.13
Note truncated.
One column per month.
v0.5.12.post1 is a stability patch on top of v0.5.12. It cherry-picks 12 fixes — primarily for DeepSeek V4 — onto the release branch.
v0.5.12.post1 is a stability patch on top of v0.5.12. It cherry-picks 12 fixes — primarily for DeepSeek V4 — onto the release branch.
deep_gemm UE8M0 scale-packing path by ceiling activation scales before packing): #25733--enable-nsa-prefill-context-parallel --nsa-prefill-cp-mode round-robin-split) in --disaggregation-mode prefill: scheduler crash at startup: #25396SGLANG_OPT_USE_COMPRESSOR_V2=1: GSM8K accuracy restored from 0.825 → 0.960: #25646pp_size=1 assertion): #25771--load-format dummy + FlashInfer mxfp4 hits CUDA illegal memory access during CUDA-graph capture (the integer HashTopK.tid2eid lookup table was left uninitialized by dummy load): #25892SGLANG_OPT_CACHE_SWA_TRANSLATION=1 returns stale translation indices after a cache rebuild, causing OOB writes / wrong outputs: #25889is_last; only expect state when truthy: #25699group arg in get_dp_buffer: #25585SGLANG_OPT_DEEPGEMM_HC_PRENORM=1 + SGLANG_OPT_USE_TILELANG_MHC_PRE=1 + hybrid SWA) to eliminate 20–40s cold-bucket forward stalls: #25810_dispatch_bf16_fp32_backend to cut runtime JIT compile cost: #25860[cu13] extra for nvidia-cutlass-dsl (default to CUDA 13; required for sm_103 / B300): #25576All PRs included in this release: v0.5.12...v0.5.12.post1
Full Changelog: v0.5.12...v0.5.12.post1
DeepGEMM deprecated in sgl-kernel; custom sgl-deep-gemm wheel + release workflow: #24268 , #24348 , #24385
DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
Day-0 Features: #23882
Post-Day-0 additions:
lmsysorg/sglang:v0.5.12 for all Nvidia GPUsSee the LMSYS blog and the DeepSeek-V4 cookbook for more details.
TokenSpeed MLA attention backend (Blackwell, FP8 KV cache): New MLA prefill/decode kernels integrated as an attention backend on SM100, with FP8 KV cache support for low-latency MLA serving: #24925
DSv3.2 / GLM-5 FP4 low-latency perf: PDL enabled across DSv3.2 / GLM-5 kernels, torch.mm for the DeepSeek V3.2 indexer GEMM, and a reland of the Cute-DSL FP4 dense GEMM — materially trimming low-latency overheads on FP4 paths: #23965, #23856, #23590, #25311
New Model Support: DeepSeek V4 #23882, Intern-S2-Preview #24875, MiniCPM-V 4.6 #24855, Laguna-XS.2 #24204, Ring-2.6-1T #25360, and Gemma 4 MTP #24436 — with cookbook recipes for tuned deployment commands. See docs.sglang.io/cookbook
HiCache + UnifiedRadixTree: HiCache framework support for UnifiedRadixTree (with SWA), HiCache for DeepSeek V4, SSD offload through Mooncake store, and stability fixes across cascade eviction, tombstone replay, and partial-match paths: #23316, #23391, #24691, #24277, #24943, #24972, #25068, #25277
Speculative Decoding V2 maturation: Adaptive Spec V2, EAGLE-3 SWA + newer drafters, Kimi K2.5 EAGLE-3 MLA, Gemma 3/4 + EAGLE-3, and an extensive naming / shape-handling refactor across draft-extend paths: #23336, #24663, #24664, #24826, #23976, #24859
CUDA 13 DeepEP migration: Gateway DeepEP source swapped from a community fork to deepseek-ai/DeepEP@hybrid-ep so DeepEP builds and runs cleanly on the CUDA 13 default; FlashInfer pinned at 0.6.11.post1 alongside a gpt-oss triton-kernel fix: #25113
Entries with a published cookbook recipe come first; entries whose cookbook page is still pending are grouped at the bottom.
EagleDraftExtendInput: #24859trtllm decode kernel for draft extend: #24566ngram metric off-by-1 in num_accepted_drafts_per_req_cpu: #24965bonus_tokens is None: #25204state_type branch: #24878PrefillDelayer: NCCL all-gather for cross-DP info sync: #24768update_status from cleared entries; fix abort update_status across KV backends: #24601, #24539, #24522_cascade_evict leaf determination fix: #25068cache_empty_result with RadixTree: #24779qkv_proj buffer sizing when tp_size > num_key_value_heads: #24420lora_id for multi-node --lora-paths: #24555sgemm_lora_a_graph_fwd due to invalid torch.mm(): #24760set_mla_kv_buffer (up to 12× over baseline): #25311return_hidden_states=false: #25155SGLANG_OPT_FP8_WO_A_GEMM on by default: #25181--prefill-only-disable-kv-cache to skip KV pool allocation: #23675SGLANG_USE_JIT_ALL_REDUCE → SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2: #24297MatchResult in RadixCache: #24470bs > 1: #24662scheduler_metadata precompute under DP attention: #24632aten::rms_norm / aten::mm.dtype registration in batch-invariant mode: #24459sglang:get_loads_duration_seconds Prometheus metric: #25163SGLANG_TRACE_LEVEL env for startup trace level: #24716fwd_occupancy metric in SchedulerStats + Prometheus collector: #24458/v1/tokenize chat-completion-style support: #23981--enable-strict-thinking: #23953reasoning.enabled mapping to thinking + enable_thinking: #23951az:// and *.blob.core.windows.net): #23995SGLANG_MAX_KV_CHUNK_CAPACITY env: #25120SGLANG_RADIX_FORCE_MISS env: #24726, #24950repetition_penalty=0 in SamplingParams.verify(): #24874--random-input-len for send_one.py: #24464dit_precision config respected (no hardcoded bf16): #24988torch.compile in native denoising: #25328out argument handling: #24688Note truncated.
CUDA 13 + Torch 2.11: Default CUDA version moves to 13.0 across SGLang, sgl-kernel, and Docker images, and PyTorch is upgraded from 2.9 to 2.11 — mode
CUDA 13 + Torch 2.11: Default CUDA version moves to 13.0 across SGLang, sgl-kernel, and Docker images, and PyTorch is upgraded from 2.9 to 2.11 — modernizing the build matrix and unlocking newer kernels: #21247, #24162, #24183, #23593 (tracking issue #21498)
Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
Decode Radix Cache for PD Disaggregation: Decode-side prefix caching now works under prefill/decode disaggregation, recovering radix-cache hit rates and TTFT savings for long shared prefixes in disaggregated deployments: #19746
Day-0 / New Model Support: Gemma 4, GLM-5.1, Qwen3.6, MiMo-V2.5 / V2.5-Pro, Ling-2.6-Flash, Mistral Medium 3.5, and Kimi-K2.6 — with cookbook recipes for tuned deployment commands. See docs.sglang.io/cookbook: #21952, #23808, #23811, #23851, #23947, #23486, #23394
DFLASH Speculative Decoding: New high-throughput spec-decode kernel from the kernel community, expanded across model backends and AMD ROCm: #22077, #22358, #22342, #23553
FA3 Kernels from the Kernel Community: Drop-in FA3 kernels contributed by the community, integrated alongside FA4 to give users a high-performance option that's easy to maintain: #20796
LoRA support for DeepSeek-V3 and Kimi-K2: LoRA now works on the largest MLA-based MoE models, including DeepSeek-V3 MLA LoRA and Kimi K2 — enabling adapter-based fine-tuning of frontier-scale models: #22323, #22381
Context Parallel (CP) Enhancements: All-reduce + RMSNorm fusion under CP for end-to-end speedups, plus support for moe_dp_size = 1 paired with arbitrary attention_cp_size so MoE and attention parallelism can be tuned independently: #21249, #22003
FlashInfer CuteDSL MoE Runner Backend: New dedicated FlashInferCuteDslMoE layer for the standard FP4 MoE path, giving an additional high-performance fused-MoE option: #21339
Entries with a published cookbook recipe come first; entries whose cookbook page is still pending are grouped at the bottom.
speculative_num_steps for EAGLE topk=1: #21599accept_length into num_accepted_drafts / num_accepted_tokens: #23962PrefillDelayer in disaggregated-prefill mode: #23588IntraNode NVLink, MTP-layer KV transfer, and disagg-prefill DP rank resolution: #23252, #23539, #22901, #22990moe_dp_size = 1 paired with arbitrary attention_cp_size: #22003reduce_scatterv for DP attention: #22642sgemm speedup with better grid selection: #22386scheduler_metadata to eliminate per-layer prepare cost: #21104gemma_weight to avoid redundant add on every forward: #22673RadixKey view for EAGLE bigram key: #23106get_load: #22480--page-size > 1 memory access fault with speculative decoding: #23596gemma4_rmsnorm_cpu kernel: #22842extend_attention_cpu / flash_attn_varlen_func NaN for large seq: #22434All PRs included in this release: https://github.com/sgl-project/sglang/compare/v0.5.10.post1...v0.5.11
Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.10.post1...v0.5.11
Bumps flashinfer from v0.6.7.post2 to v0.6.7.post3 to resolve an issue in its jit cubin downloader.
Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.10...v0.5.10.post1
Bumps flashinfer from v0.6.7.post2 to v0.6.7.post3 to resolve an issue in its jit cubin downloader.
Fix CVE-2026-3989: Replace unsafe pickle.loads with SafeUnpickler in replay_request_dump.py: #20904
Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
Elastic EP for Partial Failure Tolerance: Integrate Elastic NIXL-EP into SGLang, enabling partial failure tolerance for DeepSeek MoE deployments — when a GPU fails, the system redistributes expert weights and continues serving without full restart: #19248, #17374, #12068 blog
GPU Staging Buffer for PD Disaggregation: Gathers scattered head slices into contiguous memory for bulk RDMA transfer, reducing RDMA request count on GQA models by ~1000x. TPS/GPU on large concurrency increased by ~5x with Prefill TP4+Decode DEP4 on Qwen3.5: #19890
HiSparse for Sparse Attention: Integrate HiSparse sparse attention backend for efficient long-context inference with reduced compute through sparsity-aware attention: #20343
SGLang-Diffusion Update:
FlashInfer MXFP8 Kernel Support: Integrate FlashInfer mxfp8 kernels for GEMM and MoE operations, enabling mixed-precision FP8 inference with higher accuracy through microscaling for RL and general workloads: #19537
Transformers 5.3.0 Upgrade: Major upgrade from transformers 4.57.1 to 5.3.0, unlocking support for the latest model architectures and features from HuggingFace. GLM-5 model is now supported in this image instead of the custom built image: #17784
DeepSeek V3.2 / GLM-5 Optimization: GLM-5 runnable on main branch (with upgraded transformers). Fused Triton kernel for prefill KV cache fetching, NSA fuse store indexer for K cache, TRT-LLM prefill/decode DSA kernels as default on SM100/SM103, and IndexCache for improved throughput by more than 10% on high workloads: #19319, #19148, #20062, #21914, #21405
Qwen3.5 GDN/KDA Optimization: Transpose linear attention state layout from [N, HV, K, V] to [N, HV, V, K] and fuse split/reshape/cat ops in GDN projection with Triton kernel, plus CuTeDSL KDA decode kernel support for improved Qwen3.5 performance: #20283, #21019, #21203
LoRA Support for MoE Layers: Add LoRA fine-tuning support for Mixture-of-Experts layers with JIT alignment kernels, fused Triton kernels, TP support, CUDA graph support, and auto-detection of LoRA target modules — enabling efficient adapter-based tuning on MoE models like DeepSeek: #19710, #19711, #14105, #21439, #21647
Prefill Context Parallel for MHA (Qwen3): Enable context parallelism during prefill for multi-head attention models like Qwen3 MoE, distributing long sequences across GPUs to reduce per-GPU memory and accelerate prefill: #18233
Flash Attention 4 Official Library Support: Upgrade to the official FlashAttention 4 package, bringing the latest attention optimizations and Blackwell GPU support: #20303
Skip-Softmax Attention for FlashInfer TRT-LLM Kernels: Reduce computation overhead in attention layers by skipping redundant softmax normalization: #19089
Speculative Decoding with FA4 Backend: Enable speculative decoding for the FA4 attention backend, combining speculative inference with next-generation flash attention for faster generation: #21080
MM Attention FA4 Default on SM100: Multi-modal attention now uses FA4 by default on Blackwell hardware for improved VLM performance: #21595
Stronger Transformers Modeling Backend: Enhanced transformers backend with full TP, PP, MoE, VLM support, and torch.compile compatibility: #19163
sglang-kernel 0.4.1: Major kernel package release with renamed package (sgl-kernel → sglang-kernel), consolidated kernels, and cleanup of deprecated ops: #20440, #22009
Native MLX Backend for Apple Silicon: Add native MLX execution backend enabling SGLang to run inference directly on Apple Silicon Macs without CUDA: #20342
SGLANG_NSA_DENSE_ATTN_KV_LEN_THRESHOLD environ for controlling KV length threshold of applying sparse MLA attention kernel at prefill: #20062try_ensure_parallel_info in pending queue: #20785--stream-response-default-include-usage server flag: #16711NetworkAddress abstraction for IPv6-safe address handling: #20306--strict-ports option for predictable port assignment: #21320Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.9...v0.5.10
Fix CVE-2026-3989: Replace unsafe pickle.loads with SafeUnpickler in replay_request_dump.py: #20904
Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
Elastic EP for Partial Failure Tolerance: Integrate Elastic NIXL-EP into SGLang, enabling partial failure tolerance for DeepSeek MoE deployments — when a GPU fails, the system redistributes expert weights and continues serving without full restart: #19248, #17374, #12068 blog
HiSparse for Sparse Attention: Integrate HiSparse sparse attention backend for efficient long-context inference with reduced compute through sparsity-aware attention: #20343
SGLang-Diffusion Update:
FlashInfer MXFP8 Kernel Support: Integrate FlashInfer mxfp8 kernels for GEMM and MoE operations, enabling mixed-precision FP8 inference with higher accuracy through microscaling for RL and general workloads: #19537
Transformers 5.3.0 Upgrade: Major upgrade from transformers 4.57.1 to 5.3.0, unlocking support for the latest model architectures and features from HuggingFace. GLM-5 model is now supported in this image instead of the custom built image: #17784
DeepSeek V3.2 / GLM-5 Optimization: GLM-5 runnable on main branch (with upgraded transformers). Fused Triton kernel for prefill KV cache fetching, NSA fuse store indexer for K cache, and configurable KV length threshold for sparse MLA attention at prefill — boosting throughput for long-context DeepSeek V3.2 and GLM-5 serving: #19319, #19148, #20062
Qwen3.5 GDN/KDA Optimization: Transpose linear attention state layout from [N, HV, K, V] to [N, HV, V, K] and fuse split/reshape/cat ops in GDN projection with Triton kernel, plus CuTeDSL KDA decode kernel support for improved Qwen3.5 performance: #20283, #21019, #21203
LoRA Support for MoE Layers: Add LoRA fine-tuning support for Mixture-of-Experts layers with JIT alignment kernels, fused Triton kernels, TP support, and auto-detection of LoRA target modules — enabling efficient adapter-based tuning on MoE models like DeepSeek: #19710, #19711, #14105, #21439
Prefill Context Parallel for MHA (Qwen3): Enable context parallelism during prefill for multi-head attention models like Qwen3 MoE, distributing long sequences across GPUs to reduce per-GPU memory and accelerate prefill: #18233
Flash Attention 4 Official Library Support: Upgrade to the official FlashAttention 4 package, bringing the latest attention optimizations and Blackwell GPU support: #20303
sglang-kernel 0.4.0: Major kernel package release with renamed package (sgl-kernel → sglang-kernel), consolidated kernels, and cleanup of deprecated ops: #20440
Native MLX Backend for Apple Silicon: Add native MLX execution backend enabling SGLang to run inference directly on Apple Silicon Macs without CUDA: #20342
SGLANG_NSA_DENSE_ATTN_KV_LEN_THRESHOLD environ for controlling KV length threshold of applying sparse MLA attention kernel at prefill: #20062try_ensure_parallel_info in pending queue: #20785is_fully_idle: #20756is_fully_idle() before attach/detach: #20746pos_emb layer TP issue when DP encoder enabled for Qwen3 VL: #20788NetworkAddress for dist_init_method and loopback fallbacks: #20657NetworkAddress abstraction for IPv6-safe address handling: #20306IPV6_V6ONLY, IPv4-first, is_port_available all-family check: #20643--strict-ports option for predictable port assignment: #21320All PRs included in this release: https://github.com/sgl-project/sglang/compare/v0.5.9...v0.5.10rc0
Full Changelog: https://github.com/sgl-project/sglang/compare/v0.5.9...v0.5.10rc0
ROCm 7 standardization and ROCm 6.3 deprecation: #17785
LoRA Weight Loading Overlap with Computation: Overlap LoRA weight loading with computation during inference, reducing TTFT by ~78% and TPOT by ~34.88% on large adaptors: #15512
TRT-LLM NSA Kernel Integration for DeepSeek V3.2: Integrate TRT-LLM DSA kernels for Native Sparse Attention, boosting DeepSeek V3.2 performance by 3x-5x on Blackwell platforms with trtllm for both --nsa-prefill-backend and --nsa-decode-backend (with minor accuracy drop): #16758, #17662, #18389
Flashinfer All-to-All MoE Dispatcher: Add the Flashinfer all-to-all MoE dispatcher for efficient expert parallelism communication, enabling optimized routing in MoE models: #14668
FA4 (FP4 Attention) Support for Multimodal Encoder: Introduce FP4 attention backend and variable-length attention function for multimodal encoders, enabling lower-precision inference for vision-language models: #13539
Anthropic Compatible API Endpoint: Add native Anthropic API compatibility to SGLang, allowing direct integration with tools and clients built for the Anthropic API format: #18630
SGLang-Diffusion Advanced Optimizations: Production-ready improvements including token-level sequence sharding, parallel VAE decoding, fused kernels, Nunchaku and FP8 support, and multiple new models in the ComfyUI plugin: blog
Spec V2 Critical bug fix: Fix out-of-index bug caused by torch garbage collection in speculative decoding v2, improving reliability of speculative verification: #18958
Deploying DeepSeek on GB300 NVL72: Optimization work for long-context inference using prefill-decode disaggregation and other SGLang features on NVIDIA's latest GB300 platform: blog
Bump AITER version to 0.1.10.post3: Support FP8 Prefill/Decode/KV Cache
Commit-to-Version Lookup in docs.sglang.io: Easily find the earliest official version that includes a given PR or commit, streamlining release tracking for users and developers: #18450, link
torch.__version__ for PEP440 by @EduardDurech in https://github.com/sgl-project/sglang/pull/15682mm_fp4 backend by @b8zhong in https://github.com/sgl-project/sglang/pull/17369modelopt_quant.py -> flashinfer_trllm.py by @b8zhong in https://github.com/sgl-project/sglang/pull/16685load_file in multithread loader by @mmangkad in https://github.com/sgl-project/sglang/pull/18124America/Los_Angeles timezone, default to UTC by @mmangkad in https://github.com/sgl-project/sglang/pull/18121sglext and Prometheus metrics by @vladnosiv in https://github.com/sgl-project/sglang/pull/17648--fp4-gemm-backend documentation by @mmangkad in https://github.com/sgl-project/sglang/pull/18350quant_config in KimiK25 by @mmangkad in https://github.com/sgl-project/sglang/pull/18440quantize_config to _initialize_model by @klshuster in https://github.com/sgl-project/sglang/pull/18273--round-robin-split by @Fridge003 in https://github.com/sgl-project/sglang/pull/18613pr-test by @hnyls2002 in https://github.com/sgl-project/sglang/pull/18650/rerun-stage by @hnyls2002 in https://github.com/sgl-project/sglang/pull/18658Qwen3_5ForCausalLMMTP class implementation by @zju-stu-lizheng in https://github.com/sgl-project/sglang/pull/18538spec_accept_histogram request statistic by @scottjlee in https://github.com/sgl-project/sglang/pull/18332mrope_section with rope_type: "yarn" by @raayandhar in https://github.com/sgl-project/sglang/pull/13313channels_last_3d by @BBuf in https://github.com/sgl-project/sglang/pull/18540flashinfer_deepgemm to --fp8-gemm-backend by @mmangkad in https://github.com/sgl-project/sglang/pull/18982Note truncated.
Nothing published for this version
Fixed urllib and gpgv vulnerabilities: #17439
get_device by @rauletorresc in https://github.com/sgl-project/sglang/pull/14225transfer_engine_bench into maunal test by @hnyls2002 in https://github.com/sgl-project/sglang/pull/14429python.sglang by @hnyls2002 in https://github.com/sgl-project/sglang/pull/14577--fp8-gemm-backend by @b8zhong in https://github.com/sgl-project/sglang/pull/14379test_eagle_infer_beta_dp_attention.py by @hnyls2002 in https://github.com/sgl-project/sglang/pull/14831 PowerOfTwo policy by @ppraneth in https://github.com/sgl-project/sglang/pull/14823num_token_non_padded computation in prefill by @yuchengz816-bot in https://github.com/sgl-project/sglang/pull/14313server_fixtures in sglang.test by @hnyls2002 in https://github.com/sgl-project/sglang/pull/14899parameter to require_reasoning by @JustinTong0323 in https://github.com/sgl-project/sglang/pull/14922code field and unify error responses for router by @fzyzcjy in https://github.com/sgl-project/sglang/pull/15028bench_one_batch_server.py by @hnyls2002 in https://github.com/sgl-project/sglang/pull/15158per_token could not be properly recognized when the token count was 1. by @haoyangli0109 in https://github.com/sgl-project/sglang/pull/14415v_head_dim by @hnyls2002 in https://github.com/sgl-project/sglang/pull/15384truncate argument by @hnyls2002 in https://github.com/sgl-project/sglang/pull/14270sglang generate --perf-dump-path to include per-denoising-step timings by @BBuf in https://github.com/sgl-project/sglang/pull/15397MiMo-V2-Flash day0 support by @acelyc111 in https://github.com/sgl-project/sglang/pull/15207Note truncated.
Your coding agent can read these notes before it upgrades. Set up the MCP server →