NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2131 most downloaded on PyPI
FlashInfer: Kernel Library for LLM Serving
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Most releases are documented
notes for 41 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
103 releases · first in 2025
One column per month.
_These highlights are also published at flashinfer.ai/releases._
These highlights are also published at flashinfer.ai/releases.
This release brings MoE expert parallelism into serving engines, refreshes the Blackwell SM12x fused-MoE kernels with an FP4 accuracy fix, extends the unified MoE API to MXFP4 W4A8/W4A16 and shared experts, and adds decode coverage for Kimi K3 MLA and MiniMax-M3 sparse attention on vLLM.
MoE expert parallelism production-ready in vLLM
The MegaMoE path in flashinfer.moe_ep is ready for serving engines: full CUDA-graph capture and replay, a fused single-launch quantize-and-stage hot path, prequantized weight packs, symmetric-buffer workspaces pooled across layers, and a persistent knob cache that resolves tuned knobs by lookup, so production sessions run with no in-engine autotuning. An opt-in fault-tolerance rank mask, available over both NCCL-EP and NIXL-EP, masks and skips a peer that times out during dispatch or combine, keeping the job alive when a rank dies. Deployment gets simpler too: the CuTe-DSL runtime floor returns to 4.5.2, the version vLLM 0.25.1 pins, and a new BootstrapConfig.device lets the host framework pin each worker's CUDA device.
Blackwell SM12x fused MoE refreshed, with an FP4 accuracy fix
W4A4 serving on DGX Spark and RTX PRO parts (GB10, SM120/SM121) now delivers the output quality its benchmark scores imply, with two NVFP4 quantization bugs fixed and a new input_global_scale that lets integrators pass a checkpoint's weight scale directly. The SM12x fused-MoE families are synced to current b12x: the NVFP4 W4A4 backend reaches kernel parity across decode and prefill, and the W4A16 family adds cooperative persistent launches, a tensor-core decode path for small batches, and shape-stable route packing that keeps decode batch-size changes recompile-free.
Unified MoE API adds MXFP4 W4A8 and W4A16, shared experts, and SiTU
TRTLLM-gen MXFP4 weights now run through the unified MoELayer API against both MXFP8 activations (W4A8) and BF16 activations (W4A16), completing the FP8 series begun in 0.6.16, alongside per-tensor routed FP8. Routing adds an unpacked pre-routed FP4 mode that accepts topk_ids and topk_weights as separate contiguous tensors. TRTLLM-gen FP4 MoE also gains shared-expert fusion, extending to FP4 what 0.6.15 added for FP8, plus SiTU activation for MXFP4 x MXFP8 and NVFP4 x NVFP4.
Kimi K3 MLA decode, and MiniMax-M3 sparse attention under vLLM
Blackwell decode now covers Kimi K3's MLA geometry — 96 global query heads against one KV head, TP-local head counts down to 6, speculative query lengths up to 8, and context parallelism for long contexts — by packing query-token and query-head rows into shared CuTe-DSL tiles, adding compact variable-length Q, and extending TRTLLM-gen dense and sparse MLA to non-power-of-two head counts. Sparse MLA decode also serves the no-rotary-tail shape (kv_lora_rank=512, qk_rope_head_dim=0) natively. MiniMax Sparse Attention accepts vLLM's packed paged KV layout for MiniMax-M3, unblocking the vLLM integration on SM120/SM121, with lower per-call overhead on paged decode.
Ulysses sequence parallelism for long-context and video diffusion
Ulysses sequence parallelism gets its head-scatter / sequence-gather all-to-all as a public API, for video diffusion transformers and other long-sequence attention workloads. A fused NVLink-P2P kernel folds the layout permutation directly into the cross-GPU writes over CUDA IPC for a single coalesced push, with automatic NCCL fallback when P2P is unavailable.
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16rc5...v0.6.17
This tag was signed with the committer’s verified signature .
aleozlx Alex Yang
GPG key ID: 5E5BC1FBFCE3737D
Verified Learn about vigilant mode .
a0a6b01
This commit was signed with the committer’s verified signature .
aleozlx Alex Yang
GPG key ID: 5E5BC1FBFCE3737D
Verified Learn about vigilant mode .
These highlights are also published at flashinfer.ai/releases .
This release brings MoE expert parallelism into serving engines, refreshes the Blackwell SM12x fused-MoE kernels with an FP4 accuracy fix, extends the unified MoE API to MXFP4 W4A8/W4A16 and shared experts, and adds decode coverage for Kimi K3 MLA and MiniMax-M3 sparse attention on vLLM.
MoE expert parallelism production-ready in vLLM
The MegaMoE path in flashinfer.moe_ep is ready for serving engines: full CUDA-graph capture and replay, a fused single-launch quantize-and-stage hot path, prequantized weight packs, symmetric-buffer workspaces pooled across layers, and a persistent knob cache that resolves tuned knobs by lookup, so production sessions run with no in-engine autotuning. An opt-in fault-tolerance rank mask, available over both NCCL-EP and NIXL-EP, masks and skips a peer that times out during dispatch or combine, keeping the job alive when a rank dies. Deployment gets simpler too: the CuTe-DSL runtime floor returns to 4.5.2, the version vLLM 0.25.1 pins, and a new BootstrapConfig.device lets the host framework pin each worker's CUDA device.
#4079
#4183
#4101
#4348
Blackwell SM12x fused MoE refreshed, with an FP4 accuracy fix
W4A4 serving on DGX Spark and RTX PRO parts (GB10, SM120/SM121) now delivers the output quality its benchmark scores imply, with two NVFP4 quantization bugs fixed and a new input_global_scale that lets integrators pass a checkpoint's weight scale directly. The SM12x fused-MoE families are synced to current b12x: the NVFP4 W4A4 backend reaches kernel parity across decode and prefill, and the W4A16 family adds cooperative persistent launches, a tensor-core decode path for small batches, and shape-stable route packing that keeps decode batch-size changes recompile-free.
#3932
#4285
#4255
#4253
#4130
Unified MoE API adds MXFP4 W4A8 and W4A16, shared experts, and SiTU
TRTLLM-gen MXFP4 weights now run through the unified MoELayer API against both MXFP8 activations (W4A8) and BF16 activations (W4A16), completing the FP8 series begun in 0.6.16, alongside per-tensor routed FP8. Routing adds an unpacked pre-routed FP4 mode that accepts topk_ids and topk_weights as separate contiguous tensors. TRTLLM-gen FP4 MoE also gains shared-expert fusion, extending to FP4 what 0.6.15 added for FP8, plus SiTU activation for MXFP4 x MXFP8 and NVFP4 x NVFP4.
#4159
#4088
#4104
#4239
#4180
Kimi K3 MLA decode, and MiniMax-M3 sparse attention under vLLM
Blackwell decode now covers Kimi K3's MLA geometry — 96 global query heads against one KV head, TP-local head counts down to 6, speculative query lengths up to 8, and context parallelism for long contexts — by packing query-token and query-head rows into shared CuTe-DSL tiles, adding compact variable-length Q, and extending TRTLLM-gen dense and sparse MLA to non-power-of-two head counts. Sparse MLA decode also serves the no-rotary-tail shape ( kv_lora_rank=512 , qk_rope_head_dim=0 ) natively. MiniMax Sparse Attention accepts vLLM's packed paged KV layout for MiniMax-M3, unblocking the vLLM integration on SM120/SM121, with lower per-call overhead on paged decode.
#4178
#4108
#4039
#4324
Ulysses sequence parallelism for long-context and video diffusion
Ulysses sequence parallelism gets its head-scatter / sequence-gather all-to-all as a public API, for video diffusion transformers and other long-sequence attention workloads. A fused NVLink-P2P kernel folds the layout permutation directly into the cross-GPU writes over CUDA IPC for a single coalesced push, with automatic NCCL fallback when P2P is unavailable.
#3820
#4240
feat: close feature gap by wiring up per-tensor routed FP8 fused-moe by @jdebache in #4088
Revert PR 4122 by @jimmyzho in #4171
[GDN] improve sm100 GDN performance by @Observer007 in #4133
fix(gdn): support WY decode on SM121 by @kahyunnam in #4117
fix(norm): convert float2 to e4m3 directly in packed cast by @elwhyjay in #4167
perf(test): bulk-precompile XQA decode kernels to cut test wall time ~4x by @bkryu in #4119
[perf] Optimize TRT-LLM routing for high-expert, high-top-k workloads by @jiahanc in #4152
feat(xqa): ragged Q and per-row sliding-window masking for speculative decode by @yichengj0 in #4137
test(jit): assert BMM export symlink under GEN_SRC_DIR by @kahyunnam in #4187
Feat/ulysses p2p a2a by @forrestl111 in #3820
feat(moe_ep): MegaMoE framework integration ready: CUDA graph support, fused quant+stage launch, persistent knob cache, and prequantized weight packs by @mhoqueanik in #4079
docs: document CuTe prefill scheduling override by @kangbintNV in #4162
docs: add missing trtllm_fp8_per_tensor_scale_routed_moe API entry by @kangbintNV in #4175
docs(mamba): document checkpointing varlen arguments by @hebo1221 in #4129
Yanqinz/fix-gemm-and-grouped-mm-test-issue by @yanqinz2 in #4185
feat(comm): extend trtllm_allreduce to SM12x and fix lamport buffer pointer packing by @yichengj0 in #3903
feat(moe): add unified unpacked pre-routed FP4 mode by @feih-nv in #4104
feat(mla): support packed low-head and variable-Q decode by @PerkzZheng in #4178
bump version to 0.6.16 by @jimmyzho in #4142
[fix]fix xqa flaky test on spark by @qsang-nv in #4161
fix: make mxfp8 gemm test pass by having it quantize along the correct dimension by @jdebache in #3882
fix(moe): serialize CuTe DSL autotune replay by @zianglih in #4192
feat(moe_ep): fault-tolerance rank mask (NCCL-EP + NIXL-EP) by @Anerudhan in #4183
fix(xqa): fix PDL load ordering and SM90 fp8 draft-mask dispatch by @yichengj0 in #4199
feat(msa): accept K/V views split from a packed paged KV cache by @yichengj0 in #4039
feat(moe): add TRTLLM MXFP4 W4A8 and W4A16 unified API support by @feih-nv in #4159
[cli] add CLI helper for flashinfer-jit-cache and flashinfer-cubin wheel installs by @dierksen in #3142
[feat] Add SITU trtllmgen MOE by @jiahanc in #4180
feat(sm120): fused MoE (SwiGLU) via moe_gemm is_gated for cute SM120 groupwise GEMM by @CarstyYou in #4130
fix(moe): pad trtllm-gen route map by one element to avoid OOB read by @syuoni in #4237
perf(moe_ep): CuTe-DSL 4.5.2 mainloop WAR — drop the 4.6.1 runtime floor by @mhoqueanik in #4101
Fix the routing inconsistency for num_groups > 1 by @b8zhong in #3946
fix: support host global scale in CuTe-DSL NVFP4 quantization by @akurathiswaraj in #4138
fix/test(moe_ep): self-bootstrap 1-rank process group in dg mega oracle test by @mhoqueanik in #4221
feat: support native qk_rope_head_dim=0 sparse MLA decode in trtllm-gen by @JustinTong0323 in #4108
feat(topk): support separate page table row starts by @zianglih in #4169
feat(comm): make mixed-comm VMM workspaces checkpointable by @galletas1712 in #3910
SM 107 Reland + Merge Back from v0.6.16 Release Branch by @Vinnie6167 in #4280
test(sm103): fix FP4 autotuner cache inspection by @tiffany940107 in #4145
bump version to 0.6.17 by @aleozlx in #4283
feat(gemm): sync mm_fp4 SM120 NVFP4 dense GEMM kernel to b12x HEAD by @yichengj0 in #4253
fix(b12x): correct fp4 quantization numerics and add input_global_scale to decouple weight and activation scales by @yichengj0 in #3932
feat(moe): support shared expert fusion for trtllm-gen fp4 moe by @Aneureka in #4239
fix(test): repair MoEFinalizeConfig call site in b12x unified MoE tests ( #4395 ) by @aleozlx in #4408
revert(moe): revert #3738 SM90 CUTLASS MoE backend (+ dependents #4025 , #4080 ) on release-v0.6.17 by @aleozlx in #4411
@hebo1221 made their first contribution in #4129
@akurathiswaraj made their first contribution in #4138
@JustinTong0323 made their first contribution in #4108
Full Changelog : v0.6.16rc5...v0.6.17
Anerudhan, dierksen, and 27 other contributors
fix(test): repair MoEFinalizeConfig call site in b12x unified MoE tests (#4395) by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/4408
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.17rc4...v0.6.17rc5
feat: close feature gap by wiring up per-tensor routed FP8 fused-moe by @jdebache in https://github.com/flashinfer-ai/flashinfer/pull/4088
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16rc5...v0.6.17rc1
Restores import flashinfer.comm on Python 3.10 and 3.11. A type annotation in flashinfer/comm/fd_exchange.py evaluated only on Python 3.12+, and flash
Restores import flashinfer.comm on Python 3.10 and 3.11. A type annotation in
flashinfer/comm/fd_exchange.py evaluated only on Python 3.12+, and flashinfer.comm
imports that module at import time, so the package failed to import on interpreters
inside the supported range. Downstream packages that touch flashinfer.comm during
their own initialization were affected as well.
Upgrade to this release if you run FlashInfer on Python 3.10 or 3.11.
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16.post3...v0.6.16.post4
revert(moe): revert #3738 SM90 CUTLASS MoE backend (+ dependents #4025, #4080) on release-v0.6.16 by @bkryu in https://github.com/flashinfer-ai/flashi
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16.post2...v0.6.16.post3
FlashInfer 0.6.16.post2 includes the tvm-ffi v0.1.13-post2 hotfix regarding its ABI compatibility. We recommend upgrading to this latest version.
FlashInfer 0.6.16.post2 includes the tvm-ffi v0.1.13-post2 hotfix regarding its ABI compatibility. We recommend upgrading to this latest version.
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16.post1...v0.6.16.post2
⚠️ FlashInfer 0.6.16.post2 has picked up the latest tvm-ffi compatibility fix. We recommend upgrading to the latest version.
⚠️ FlashInfer 0.6.16.post2 has picked up the latest tvm-ffi compatibility fix. We recommend upgrading to the latest version.
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16...v0.6.16.post1
_These highlights are also published at flashinfer.ai/releases._
These highlights are also published at flashinfer.ai/releases.
This release delivers the MegaMoE kernels promised for expert parallelism in 0.6.15 and extends Blackwell model coverage — MiniMax-M3 sparse attention on Blackwell RTX and DGX Spark, and a unified MoE API that now executes per-tensor FP8, block-scale FP8, and B12x NVFP4/W4A16 — plus distributed serving under Confidential Computing with multicast-free all-reduce fusion. Under the hood, an on-disk JIT cache for CuTe-DSL kernels cuts cold-start compilation and shrinks the install.
⚠️ FlashInfer 0.6.16.post2 has picked up the latest tvm-ffi compatibility fix. ⚠️ FlashInfer 0.6.16.post4 adds a Python 3.10 compatibility hotfix. We recommend upgrading to the latest 0.6.16 post-release.
MegaMoE kernels land in expert parallelism (moe_ep)
The MoEEpLayer now unifies the split and mega execution paths under one entry point and adds three mega backends — deep_gemm_mega (FP8/FP4), nvfp4_cutedsl, and mxfp8_cutedsl — each fusing expert-parallel communication with the local MoE in a single symmetric-memory kernel (Blackwell SM100+, NVSHMEM). In a microbenchmark, on a GB200 node (EP=4, DeepSeek-V3-like geometry) the tuned CuTeDSL NVFP4 backend with in-register FP4 combine reaches up to 1.89× the throughput of deep_gemm_mega at 8192 tokens/rank. NIXL-EP transport gaps and combine deadlocks blocking the vLLM Fleet/Handle adapters are also closed.
MiniMax Sparse Attention (MSA) on Blackwell RTX and DGX Spark
MiniMax-M3's MSA — a proxy/top-k indexer plus sparse prefill, decode, and combine — now runs on Blackwell RTX and DGX Spark (SM120/121) GPUs. The tensor-core kernels are rebuilt on SM12x warp-level mma.sync, top-k and combine are rewritten in CuTe-DSL, and an optional NVFP4 indexer is available, with prefill/decode accepting FP8 or NVFP4 KV (paged or flat). New APIs live under flashinfer.msa_ops.
XQA decode adds sliding-window, attention sinks, and ragged Q for speculative decode
The XQA decode kernel — used on Blackwell RTX and DGX Spark (SM120/121) for models with attention sinks — expands coverage for sliding-window attention and ragged Q, across causal and non-causal draft-block mask modes and combinations of them. Ragged Q lets each request in a batch verify a different number of draft tokens, and sliding-window masking is now computed per draft-token row, making the kernel useful for speculative-decoding workloads across models on SM120/121.
Unified MoE API reaches parity with the legacy FP8 and NVFP4 paths
The unified MoELayer API continues to expand quantization formats to reach parity with legacy flat APIs: TRTLLM per-tensor FP8 (SM100/SM103, with Llama4 routing-scale-on-input), DeepSeek FP8 and MXFP8 block-scale (SM100/SM103), and SM120/SM121 B12x NVFP4 and W4A16 backends. In-kernel routing (FromLogits) is now wired through the unified API and fuzzer, so precomputed and in-kernel routing share one path with CUDA-graph and autotuning coverage.
Confidential Computing: multicast-free all-reduce fusion and FP8 AllReduce
The TRT-LLM AllReduce-fusion workspace now allocates a multicast-free IPC workspace when NVIDIA Confidential Computing is detected (is_confidential_compute(), overridable via FLASHINFER_CONFIDENTIAL_COMPUTE), so one-shot Lamport and two-shot sync fusion run under CC where cuMulticast setup otherwise fails. Separately, a new flashinfer.comm.quantized_all_reduce() halves AllReduce transfer volume by quantizing activations to FP8 before P2P transfer over symmetric memory (SM90+, NVSwitch).
Faster JIT cold-start and smaller install
CuTe-DSL kernels now persist to an on-disk cache (JitSpec gains a JitSpecCuteDsl backend) and reload via JITLink in about 3–30 ms instead of recompiling in every new process; mm_fp4 autotuning additionally compiles tactics in parallel and reuses the shared disk cache, cutting autotune wall time. Pruning architecture gencode that dispatch can never load removes roughly 1.6 GB of installed size from the CUDA-13 aarch64 JIT-cache wheel.
set_autotune_process_group to synchronize tactic choice across ranks by @thanhhao98 in https://github.com/flashinfer-ai/flashinfer/pull/3187topk_weights (copy-free) by @jdebache in https://github.com/flashinfer-ai/flashinfer/pull/3763Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.15rc4...v0.6.16
cherry-pick: #4129 and #4162 by @jimmyzho in https://github.com/flashinfer-ai/flashinfer/pull/4245
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16rc4...v0.6.16rc5
fix(moe): reject incompatible output-scale cubins by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/4213
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16rc3...v0.6.16rc4
release-v0.6.16: fix cubin checksum collision and unguarded Sm107a build break by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/4200
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.16rc2...v0.6.16rc3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.15...v0.6.15.post1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.15...v0.6.15.post1
_These highlights are also published at flashinfer.ai/releases._
These highlights are also published at flashinfer.ai/releases.
This release ships Expert Parallelism (moe_ep) in the default install and extends the TRTLLM-GEN MoE stack for large-model serving. It broadens Blackwell model coverage — Gemma 4 / MiniMax-M3 MoE on consumer / DGX-Spark SM12x, DeepSeek-class MLA decode on B300, and context-parallel GDN on SM120 — and adds Video Sparse Attention plus CUDA-graph-safe FP8 all-reduce fusion for distributed inference.
Unified MoE API with Expert Parallelism (Experimental API to try!)
FlashInfer's unified MoE compute API is now wired into expert parallelism. A new flashinfer.moe_ep.MoEEpLayer runs one MoE layer split across ranks as dispatch → per-expert grouped GEMM → combine, over pluggable transport (NCCL-EP via nccl.ep/nccl4py, and NIXL-EP). The expert GEMM reuses the unified flashinfer.fused_moe.MoELayer as a pure per-expert grouped GEMM — routing lives in dispatch/combine. moe_ep is part of the default install (CUDA 13+), with vLLM-facing APIs and checkpoint-safe MoE all-to-all graph VAs. This EP API is also the integration surface for MegaMoE kernels in upcoming releases — try it out and let us know how it works for you.
Gemma 4 and MiniMax-M3 NVFP4 MoE now run on Blackwell SM12x
Gemma 4 and MiniMax-M3 NVFP4 MoE now run on Blackwell SM12x (consumer / DGX Spark), enabled by two new NVFP4 MoE activation functions — gelu_tanh and swiglu_oai — added to the SM12x MoE path.
TRTLLM-GEN MoE adds shared experts, DeepSeek-V4 routing, and low-latency FP8
TRTLLM-GEN MoE gains shared-expert fusion for FP8 paths, hash-based DeepSeek-V4 routing (hash_topk), and a fused FP8 blockwise megakernel that cuts latency for small batches (BS ≤ 8).
Video Sparse Attention and DeepSeek MLA decode on Blackwell B300
Video Sparse Attention (VSA) is now integrated into the block-sparse attention API, bringing efficient long-context attention for video diffusion models to FlashInfer. DeepSeek-class MLA decode extends onto Blackwell B300 (SM103) with cluster-aware CUTLASS split_kv, and CuTe-DSL GQA decode adds sliding-window and attention-sink masking for newer attention variants.
Linear attention: context-parallel GDN on SM120 and faster KDA decode
Gated delta-rule (GDN) linear attention adds SM120 context-parallel delta rule support, extending the CuTe-DSL GDN rewrite from 0.6.14 onto Blackwell SM120 prefill, alongside SM90 context-parallel prefill optimizations. KDA recurrent-decode kernels are also optimized for lower decode latency.
Distributed comm: FP8 fusion and CUDA-graph checkpoint restore
Dynamic per-token FP8 quantization fuses into allreduce + residual + RMSNorm (TRT-LLM and MNNVL backends, CUDA-graph safe). All-reduce workspaces support checkpoint_prepare / checkpoint_restore so physical backing can be released and remapped at stable VAs across serving checkpoints.
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.14rc1...v0.6.15
_These highlights are also published at flashinfer.ai/releases._
These highlights are also published at flashinfer.ai/releases.
Please modify your install instructions as you bump the version to 0.6.14.
pip install flashinfer-python
pip install flashinfer-cubin --index-url https://flashinfer.ai/whl # this is the difference
pip install flashinfer-jit-cache --index-url https://flashinfer.ai/whl/cu129
# OR
pip install flashinfer-jit-cache --index-url https://flashinfer.ai/whl/cu130
See also https://github.com/flashinfer-ai/flashinfer/issues/3808
This release pushes FlashInfer's coverage onto Blackwell RTX PRO and DGX Spark silicon, lands the CuTe-DSL rewrite of the gated-delta-rule (GDN) kernels behind Qwen3.5/3.6, and makes those kernels production-ready inside vLLM.
Gemma 4 and W4A16 on Blackwell RTX Pro and DGX Spark
Blackwell SM12x parts (RTX PRO 6000, DGX Spark / GB10) previously trailed Blackwell GB200 NVL72 (and B300 SM103) on attention shape and quantization coverage. This release closes the gap for Gemma 4 and weight-only-quantized inference: head_dim=512 attention for Gemma 4's global layers now runs on SM120/121, the FMHAv2 prefill path gains head_dim=256/512 and sliding-window masking on SM120, the new mm_bf16_fp4 W4A16 GEMM is tuned for DGX Spark (completing W4A16 across both dense GEMM and MoE), and gated tanh-GELU brings Gemma 4 MoE onto the CUTLASS backend.
GDN / gated delta rule: CuTe-DSL overhaul (Qwen3.5/3.6)
The gated-delta-rule kernels behind the Qwen3.5/3.6 family were rewritten from CUTLASS C++ to CuTe-DSL (https://github.com/flashinfer-ai/flashinfer/issues/3491). The rewrite eliminates the C++ JIT compilation pain reported by customers and establishes the base for context-parallel delta-rule kernels — covering SM90 prefill and its context-parallel variant, SM120 prefill, and a ~20–25% GDN prefill speedup from mainloop efficiency work.
GDN production-ready in vLLM
Two gaps blocked GDN serving in vLLM (https://github.com/flashinfer-ai/flashinfer/issues/3602); both are now resolved. GDN kernels are compilation batch-size agnostic, so a single compiled cubin is reused across batch shapes instead of recompiling on every new batch size at inference time, and a new BF16 state recovery/decode kernel writes SSM state into preallocated space to supply the MTP-compatible spec-decode path vLLM needs.
DeepSeek-class sparse MLA on Blackwell, FP8 KV on Hopper
New sparse-MLA paged-attention kernels extend the DeepSeek-V4 (d_qk=512) and DeepSeek-V3.2 / GLM-5.1 (d_qk=576) families onto SM120/121 through the existing flashinfer.mla APIs, with DSv4 coverage broadened to 8/16/32 head counts. On Hopper SM90, native FP8 KV cache support eliminates SGLang's per-layer cast workaround, saving an HBM round-trip per layer on DeepSeek-V3/V4 while staying bit-identical to the BF16 path.
This release pushes FlashInfer's coverage onto Blackwell RTX PRO and DGX Spark silicon, lands the CuTe-DSL rewrite of the gated-delta-rule (GDN) kernels behind Qwen3.5/3.6, and makes those kernels production-ready inside vLLM.
FilteredTopK overflow refinement by @awgu in https://github.com/flashinfer-ai/flashinfer/pull/3529mm_fp4 cute-dsl backend when M is not a multiple of 8. by @b8zhong in https://github.com/flashinfer-ai/flashinfer/pull/3667Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.13rc2...v0.6.14
replace deprecated APIs: cute.make_fragment and cute.core.ThrMma by @brandon-yujie-sun in https://github.com/flashinfer-ai/flashinfer/pull/3430
bmm_fp8 and cuDNN bmm_fp8/mm_fp4 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3437nvfp4_quantize(backend='cuda') silently corrupts scale factors when global_scale is not float32 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3497Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc3...v0.6.13
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.13rc1...v0.6.13rc2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.13rc1...v0.6.13rc2
replace deprecated APIs: cute.make_fragment and cute.core.ThrMma by @brandon-yujie-sun in https://github.com/flashinfer-ai/flashinfer/pull/3430
bmm_fp8 and cuDNN bmm_fp8/mm_fp4 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3437nvfp4_quantize(backend='cuda') silently corrupts scale factors when global_scale is not float32 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3497Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc3...v0.6.13rc1
fix deprecation warnings from cute-dsl by @b8zhong in https://github.com/flashinfer-ai/flashinfer/pull/3333
simple mamba SSU kernel by @ishovkun in https://github.com/flashinfer-ai/flashinfer/pull/2962moe_output_memset_inplace dense memset wrapper by @leejnau in https://github.com/flashinfer-ai/flashinfer/pull/3328Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11rc1...v0.6.12
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc2...v0.6.12rc3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc2...v0.6.12rc3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc1...v0.6.12rc2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.12rc1...v0.6.12rc2
fix deprecation warnings from cute-dsl by @b8zhong in https://github.com/flashinfer-ai/flashinfer/pull/3333
simple mamba SSU kernel by @ishovkun in https://github.com/flashinfer-ai/flashinfer/pull/2962moe_output_memset_inplace dense memset wrapper by @leejnau in https://github.com/flashinfer-ai/flashinfer/pull/3328Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11rc1...v0.6.12rc1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11.post2...v0.6.11.post3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11.post2...v0.6.11.post3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11.post1...v0.6.11.post2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11.post1...v0.6.11.post2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11...v0.6.11.post1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11...v0.6.11.post1
trying this one character fix for main branch by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/3213
mm_bf16 and enable multi-tactic autotuning for FP8/MXFP8 runners by @vadiklyutiy in https://github.com/flashinfer-ai/flashinfer/pull/2914Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.10rc1...v0.6.11
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11...v0.6.11rc1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.11...v0.6.11rc1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.10...v0.6.10.post1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.10...v0.6.10.post1
Vendor CCCL v3.3.2 from GitHub instead of relying on CTK-bundled copy by @kahyunnam in https://github.com/flashinfer-ai/flashinfer/pull/3091
row_starts and dsa_graph_safe to topk by @zianglih in https://github.com/flashinfer-ai/flashinfer/pull/3133Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.9rc1...v0.6.10
Vendor CCCL v3.3.2 from GitHub instead of relying on CTK-bundled copy by @kahyunnam in https://github.com/flashinfer-ai/flashinfer/pull/3091
row_starts and dsa_graph_safe to topk by @zianglih in https://github.com/flashinfer-ai/flashinfer/pull/3133Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.9rc1...v0.6.10rc1
feat: Add backend="b12x" for mm_fp4 on SM120 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3051
trtllm_fp8_per_tensor_scale_moe_op by @pavanimajety in https://github.com/flashinfer-ai/flashinfer/pull/3094tie_break for filtered topk by @zianglih in https://github.com/flashinfer-ai/flashinfer/pull/3095autotune() by @vadiklyutiy in https://github.com/flashinfer-ai/flashinfer/pull/2958Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.8rc1...v0.6.9
feat: Add backend="b12x" for mm_fp4 on SM120 by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/3051
trtllm_fp8_per_tensor_scale_moe_op by @pavanimajety in https://github.com/flashinfer-ai/flashinfer/pull/3094tie_break for filtered topk by @zianglih in https://github.com/flashinfer-ai/flashinfer/pull/3095autotune() by @vadiklyutiy in https://github.com/flashinfer-ai/flashinfer/pull/2958Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.8rc1...v0.6.9rc1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.8...v0.6.8.post1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.8...v0.6.8.post1
Add to CODEOWNER by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/2875
trtllm_fp4_block_scale_moe causing "Unsupported hidden state scale shape" for EP32+ configs by @qiching in https://github.com/flashinfer-ai/flashinfer/pull/2853trtllm_fp8_block_scale_moe by @wzhao18 in https://github.com/flashinfer-ai/flashinfer/pull/2739Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post3...v0.6.8
Add to CODEOWNER by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/2875
trtllm_fp4_block_scale_moe causing "Unsupported hidden state scale shape" for EP32+ configs by @qiching in https://github.com/flashinfer-ai/flashinfer/pull/2853trtllm_fp8_block_scale_moe by @wzhao18 in https://github.com/flashinfer-ai/flashinfer/pull/2739Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post3...v0.6.8rc1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post2...v0.6.7.post3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post2...v0.6.7.post3
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post1...v0.6.7.post2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7.post1...v0.6.7.post2
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7...v0.6.7.post1
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.7...v0.6.7.post1
bump version to 0.6.7 & fix api breaking changes by @aleozlx in https://github.com/flashinfer-ai/flashinfer/pull/2832
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.6...v0.6.7
fix: move ArtifactPath/CheckSumHash imports inside gen_moe_utils_modu… by @dierksen in https://github.com/flashinfer-ai/flashinfer/pull/2681
cutlass_fused_moe mxfp8 by @zianglih in https://github.com/flashinfer-ai/flashinfer/pull/2581Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.5...v0.6.6
feat: BF16 GEMM benchmarking support by @raayandhar in https://github.com/flashinfer-ai/flashinfer/pull/2525
TensorView by @hypdeb in https://github.com/flashinfer-ai/flashinfer/pull/2602Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.4...v0.6.5
perf: add fp4 GEMM tile configs and streamK scheduler for SM120 by @Yuening-wa in https://github.com/flashinfer-ai/flashinfer/pull/2460
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.3...v0.6.4
ci: add permission control for public ci tests by @yongwww in https://github.com/flashinfer-ai/flashinfer/pull/2397
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.2...v0.6.3
chore: MoE benchmark effective BW fix for trtllm_block_scale_moe by @rosenrodt in https://github.com/flashinfer-ai/flashinfer/pull/2341
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.1...v0.6.2
Add deprecation and removal policy by @sricketts in https://github.com/flashinfer-ai/flashinfer/pull/2349
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.0...v0.6.1
feat: Add backend='auto' to mm_fp4 and enable autotune for backend='cudnn' by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/1979
mm_fp4 docstring by @b8zhong in https://github.com/flashinfer-ai/flashinfer/pull/2177rope_quantize_fp8_append_paged_kv_cache by @elvischenv in https://github.com/flashinfer-ai/flashinfer/pull/2255Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.3...v0.6.0
[feat] Integrate SGLang concat_mla_k kernel into flashinfer by @jiahanc in https://github.com/flashinfer-ai/flashinfer/pull/2237
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.6.0rc1...v0.6.0rc2
feat: Add backend='auto' to mm_fp4 and enable autotune for backend='cudnn' by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/1979
mm_fp4 docstring by @b8zhong in https://github.com/flashinfer-ai/flashinfer/pull/2177Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.3...v0.6.0rc1
[API change] deprecate tile_token_dim in trtllm_moe by @jiahanc in https://github.com/flashinfer-ai/flashinfer/pull/2086
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.2...v0.5.3
ci: Update cudnn version requirements in CI container by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/2039
test_green_ctx and test_jit_example on spark (sm_121) by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/1951Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.1...v0.5.2
test: Enable xfailed trtllm decode long seqlen tests and update microbenchmark by @bkryu in https://github.com/flashinfer-ai/flashinfer/pull/2018
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.0...v0.5.1
fix vllm graph register and add test by @NVShreyas in https://github.com/flashinfer-ai/flashinfer/pull/1894
release-ci-docker workflow by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/1944/usr/local/cuda as default CUDA_HOME if possible, like torch.utils.cpp_extension.CUDA_HOME by @netanel-haber in https://github.com/flashinfer-ai/flashinfer/pull/1948Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.4.1...v0.5.0
fix: ensure SM120/121 SFA/SFB contiguity by @yongwww in https://github.com/flashinfer-ai/flashinfer/pull/1963
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.0rc2...v0.5.0rc3
bugfix: fix regex in update wheel index script by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/2009
Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.5.0rc1...v0.5.0rc2
fix vllm graph register and add test by @NVShreyas in https://github.com/flashinfer-ai/flashinfer/pull/1894
release-ci-docker workflow by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/1944/usr/local/cuda as default CUDA_HOME if possible, like torch.utils.cpp_extension.CUDA_HOME by @netanel-haber in https://github.com/flashinfer-ai/flashinfer/pull/1948Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.4.1...v0.5.0rc1
fix: fix the failed sampling unittest on 5090 by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/1886
tests/utils/test_load_cubin_compile_race_condition.py from pytest by @yzh119 in https://github.com/flashinfer-ai/flashinfer/pull/1907ffi::TensorView instead of ffi::Tensor by @cyx-6 in https://github.com/flashinfer-ai/flashinfer/pull/1844Full Changelog: https://github.com/flashinfer-ai/flashinfer/compare/v0.4.0...v0.4.1
Compare
Compare
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →