NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2184 most downloaded on PyPI
NVIDIA cuDNN Frontend — Python and C++ Graph API with SOTA attention (SDPA / Flash Attention), MoE grouped GEMM fusions, and FP8/MXFP8 kernels for Hopper, Blackwell, and Rubin GPUs.
Last release 11 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Nearly every release is documented
notes for 35 of 35 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
35 releases · first in 2024
cuDNN Frontend v1.30.0 Release Notes
cuDNN Frontend v1.30.0 is the recommended version for cuDNN 9.26 and later releases.
v1.29.0 shipped paged KV caches on the FROST SDPA engine and noted that FlashInfer already builds exactly the graph the engine wants. v1.30.0 is the release that makes that claim hold end to end: the graphs FlashInfer actually builds are re-declared in-tree as conformance suites, every place the engine read the buffer instead of the declaration is fixed, and the remaining declines on FlashInfer's serving shapes are gone.
The graph declaration is the operand contract (#1028). The cuDNN backend reads a variant pack as pointers, so callers bind buffers whose own shape need not match the declaration — a 2-D matrix to a [1, m, k] tensor, a flat quantizer blob to a reordered scale tensor, a 0-d scalar to (1, 1, 1). Python engines that read the pack saw the buffer's shape instead, so the same execute() succeeded on a backend plan and refused on a FROST plan. _normalize now describes a slot from the declaration whenever the caller's geometry disagrees with it but covers its bytes, with a precise rule: a buffer carrying the declared extents keeps its own strides (the linear-attention engines serve strided inputs that way), and a buffer of other extents is re-described only when its slots form one dense run that covers the declared bytes — which a transposed view of a contiguous block, such as FlashInfer's column-major B, does. The declaration supplies the dtype too (#1035): FlashInfer binds packed fp4 data as uint8 and the E4M3 scale blob as uint8, so a buffer whose storage slots are as wide as the declaration's is read as the declared dtype, while a narrower or wider buffer keeps its own and is never re-described.
The SDPA THD gate is over (H, S, D), not BSHD (#1028). Under ragged offsets the batch stride of Q/K/V/O is never read — every sequence base comes from the offset table and the lowering binds the batch axis at extent 1 — yet the engine gate and the DSL adapter required BSHD-physical order over all four axes. FlashInfer declares the batch stride equal to the token stride (h * d), so its ragged prefill was declined at every b > 1. packed_layout_ok now checks head dim innermost, then heads, then tokens, and the same predicate is used by the mismatch gate and by both the SM100 and SM120 adapters.
FlashInfer's padded (b, s_max, h) Stats is served under THD (#1036). A THD graph whose Stats tensor has no ragged offsets is the per-batch padded form — FlashInfer's return_lse buffer. The packed path wrote token rows contiguously, so at b > 1 the form was declined and at b == 1 it coincided by luck with the rows past a sequence's length left unwritten, which is what failed 52 of FlashInfer's token_indptr tests once an earlier test had dirtied the allocator. The 20 SM100/SM107 THD templates now take lse_padded_rows and select the per-batch form on the static fake, and the adapter seeds the buffer with -inf on the launch stream, matching the backend's contract. Nothing is added to the kernel ABI.
THD templates compile with dynamic batch and head extents (#1039). A compile() pinned b, qh, kh and every stride in its fakes, so FlashInfer's serving shapes minted one 2.6 s kernel per (b, qh, kh, strides) — 126 forward kernels in FlashInfer's cuDNN attention tests. The kernels never needed it: _host reads B/QH/KH from problem_size at run time. compile(dynamic_bhk=True) rebinds them to cute.sym_int() right after the cache key is taken, giving one kernel per layout class.
Dense padded-Q trim on every kernel (#1037). FlashInfer's dense batched prefill hands per-batch Q lengths with every padded graph. The SM100 fp8/mxfp8 d128 and (192, 128) flavors previously either ran a host read (seq_len_q.min().item(), which breaks CUDA-graph capture) or were declined; all four remaining kernels now carry the same trim as prefill_d256_fp8, and the capability row and the execute-time host read retire.
Split-KV is also kept on paged decode when the declared KV max is not a 128-multiple (#1092) — FlashInfer passes its true max verbatim, and a paged graph is padded by construction, so the clause that declined it never applied.
cute.compile runs the whole DSL backend once per process per distinct kernel — 0.7 s for a FROST GEMM, 2.6 s for an SDPA prefill on SM100 — and the DSL's own file cache stores only MLIR bytecode. cudnn.frost.compiled_cache (#1031) keeps the exported tvm-ffi object of every kernel compiled with --enable-tvm-ffi and reloads it in milliseconds, across processes.
compile_cached(fn, *args, cache_key=, symbol=, **kwargs) is a drop-in for cute.compile; hit and miss hand back the same reloaded artifact.os.replace), and anything doubtful is a miss. An environment with any unknown manifest field neither loads nor persists.cute.compile sites across the 36 SDPA forward/backward template files route through it. GEMM suites: 121 tests, 44 s cold → 12 s warm, and a second process reloads and computes a bit-identical result.prune() (#1038) retires whole dead environment directories — dead schema roots first, then oldest first, never the process's own — under CUDNN_FRONTEND_COMPILED_CACHE_MAX_BYTES (4 GiB default, 0 disables), run once per process after its first export. Prune never follows symlinks, never leaves the root, and removes only directories this cache made.A Python (FROST) plan is now described, named and replayed exactly like a backend plan: (engine_id, {cudnn.knob_type: int}) (#1026).
knobs.h freezes the 33 backend-mirrored KnobType_t values and adds a frontend-only band from FRONTEND_KNOB_TYPE_BASE = 1000 (SCHED_POLICY, PACK_GQA, SPLIT_KV) that convert_to_backend_knob_type refuses. Both bands are append-only, pinned by static_assert.SdpaFwdKnobs and becomes the sdpa() / sdpa_fp8() op attribute softmax_precision, which is never forwarded to C++ and makes the node backend-unlowerable when set.get_engine_and_knobs_at_index always returns a dict, create_execution_plan accepts the public dict, and get_plan_name_at_index renders engine[TILE_M=128, ...].frost_gemm family gains the facts + heuristics hooks so its plan enters the ranked list with the TileConfig it will build; GemmKnobs ↔ TileConfig is lossless through the canonical config name, so a recorded (engine_id, knobs) replays the exact kernel. A record naming no canonical config is declined, never snapped.recommend() hook, so its plans list the tiles the lowering will pick rather than {}.pygraph.serialize() now applies key()'s rule and raises cudnnGraphNotSupportedError on a backend-unlowerable graph before any lowering — previously a SET softmax_precision serialized as the f32 pipeline and would deserialize and run with different numerics.Plan ranking improved alongside it: FROST SDPA plans are ranked by workload against measured shards (#1152), SM100 heuristics account for candidate launch geometry (#1158), and the FROST GEMM heuristic takes the 64-byte-K MMA and the 256-tall CTA pair on the parts that issue them (#1177, #1150).
The SM100 f16/bf16 row served paged decode with the prefill pipeline: 512 Q rows per cga2 cluster with 1..G of them live at S_q = 1, so every KV tile paid full dead-row BMM1/BMM2 and two softmax warpgroups of exp on zeros. FlashInfer's Qwen3-235B decode shape (b=32, H=64/4, d128, S_kv=4096, page 16, bf16) measured 118 µs against 44 µs for the cuDNN backend's decode engine.
sm100/decode_d128_f16.py, TILES_Q=1 at cga1, one softmax warpgroup, Q/O SMEM aliased so the K/V ring is three stages deep (224 KiB), and two S/P TMEM slots alternating per KV tile so BMM1(i+1) overlaps softmax(i). Same contract and ABI as the prefill kernel, paged pools included.sm100/decode_d256_f16.py, swap-AB, selected for decode-shaped graphs (S_q × PackGQA group ≤ 32 rows: Qwen3.5 32/2 at S_q ≤ 2, Qwen3-Next 16/2 at S_q ≤ 4, MHA at S_q ≤ 32). Same engine row and decline discipline as the prefill d256 tile.G = H_q / H_kv that does not divide TILE_M could not pack at all, so a 96/8 (G = 12) d128 paged decode ran unpacked with one live row per 512-row cluster and the KV head re-read 12×. The d128/d256 f16 kernels now pack p = gcd(G, TILE_M): 96/8 packs 4 heads per token row-group, 48/8 packs 2. A group sharing no factor with the tile still cannot pack, and a pinned PACK_GQA=1 stays declined rather than silently running unpacked.S_q == 1 (#1095), including paged KV and sliding window — the python-native validator's blanket "decode only mode not supported with sink_token" rule is dropped, since whether a sink at S_q = 1 is served is an engine's support-surface answer at planning time.The f16/bf16 SM100/SM107 forward kernels move onto an explicit pointer/int host entry, so every extent and stride is a runtime argument and the compile key is layout-only (#1119). lower_dsl_prefill resolves at plan time what the graph fixed — operand ids, IR layouts, which feature operands the facts demand, the quantized-operand id table, a raw-stream → CUstream cache — and _execute_resolved then does dictionary lookups and one api.execute(**kwargs).
VariantPackNative.override_many, #1132), with the graph keeping a uid → slot map next to its cached DeclaredLayout. On a FlashInfer-shaped THD graph, a bounded batch override goes 73–75 µs → ~35 µs; the backend plan pays the same loop and benefits alike. Dense launches join THD on the prepared positional entry in the same change.(batch, seq, head) stride elements as dynamic Int32 from the compile() placeholders, so a 16-bit operand with S·H·D == 2^27 elements per batch encoded a negative batch stride (launch aborts) and 2^28..2^30 encoded 0 (every batch silently aliases batch 0). The stride placeholders are now cutlass.Int64 and _bshd coerces its leaves, so every TMA descriptor is 64-bit whatever the caller bound..contiguous() before a descriptor was ever built — a full gather of Q on every execute, and exactly what a caller slicing a fused QKV projection hits. Nothing in the kernels needed it: they address Q/K/V only through TMA coordinates. Twelve kernels across SM100 and SM107 now decide per operand at compile time.The SM107 d256 epilogue gains a fused gate, O *= sigmoid(G), in the production kernels and as an op graph, with a composable gated attention block on top (#1102). It grew over the release into a full quantized block:
sdpa_mxfp8 also learns to infer O / Stats / Amax_O dims the way sdpa_fp8 does, so a caller that never asks for Amax_O no longer fails Tensor.validate().MxQuantSpec.w_qkvg_dtype puts MXFP4 weights on the mixed block-scale GEMM row, and MxQuantSpec.o_fp4 = NVFP4 | MXFP4 block-quantizes the gated O to feed an fp4 × fp4 block-scale out projection. New tile_dsl fp4 primitives (fp32_to_fp4_pack via cvt.rn.satfinite.e2m1x2, e4m3_scale_from_amax, e8m0_from_amax) and a quantize_fp4 kernel writing per-token e2m1 codes directly in the out-projection GEMM's F8_128x4 order. proj_gemm takes a per-operand dtype, block size 16|32 and E4M3|E8M0 scale dtype through one pairing table derived from the FROST block-scale catalog.The per-tensor FP8 forward can now emit O as FP4_E2M1 (two per byte, one E4M3 scale per 16 d elements) or as FP8_E4M3 with one UE8M0 scale per 32 d elements, writing the scale factors to a new optional sdpa_fp8 output sf_o in F8_128x4 atom order (#1088). The epilogue reuses the row-owning correction warps on the d128 SM100/SM107 kernels and a quad butterfly on the SM120 kernel; scale_o doubles as the FP4 global scale.
The MXFP8-input forward gets the same contract on the SM100 d128 kernels (#1180), including a python-only scale_o input — required for an FP4 O, since the E4M3 block scale alone cannot span its range — and rejected without sf_o.
A family of prepared FROST primitives for DSv4.1:
[FROST] Add sm100 conv forward (#961) brings the first convolution kernel to the FROST engine family on Blackwell.
There was no KDA path on Hopper at all: the FROST kernels are Blackwell-only by construction (42 tcgen05 and 84 tmem references in kda_prefill_f16.py alone), and the only other backend, cuTile, needs the cuda.tile runtime — so on an H100 all three linear-attention ops raised cudnnGraphNotSupportedError. Relaxing the arch gate cannot work, because Hopper has no Tensor Memory and no tcgen05 MMA: SM90 needs a different schedule, not a port.
[128, 128] state, and a mid-chunk anchor that keeps every exp argument inside ±40.dkg → d_kg followed in #1103.DeepSeek Sparse Attention (DSA).
(2,1,1)-cluster CuTe DSL kernel where the two CTAs split the 128 heads on cta_group::2 tensor cores, P and dS exchanged with shared::cluster bulk copies, and slot validity applied before every score and gradient operation so ignored top-k slots cannot contaminate gradients.Block Sparse Attention (BSA).
cudnn.jax and cudnn.torch (#1081, #1111).exp burst so BMM2 waits on softmax every iteration. The row-sum now rides the tensor core: one N=16 MMA of the e4m3 P operand against an all-ones SMEM tile. Zero-copy SF bind, scheduler claims and causal ranking land with it.cute.math.max lowers to arith.maxnumf, which the DSL → NVVM path emits as a COMPARE + SELECT pair, costing three ALU ops per O element in the epilogue's critical tail — the d512 fp8 kernel carried 1548 FSETP + 1541 FSEL per tile against the C++ reference's 6. tile_dsl.pointwise.fmax_f32 emits PTX max.f32, which ptxas fuses into FMNMX/FMNMX3 with |x| folded into the operand modifier: one instruction per element, NaN semantics unchanged.mbarrier.arrive. A hint-less ring-wait spin becomes a per-kernel opt-in on 7 Rubin prefill flavors, with d512 fp8/mxfp8 kernel hoists.MoeEp Python API with validated forward and backward contracts and lazy optional-dependency loading, vendored MegaMoE CuTeDSL communication/workspace/scheduling primitives, Rubin forward GLU and backward dGLU training kernels, and a runtime resource layer managing NVSHMEM lifecycle, symmetric workspaces and capability checks behind a lazy backend seam.PV_BF16 for the MXFP8 (QK) SDPA kernel (#983), avoiding causality leakage, with a D192 hybrid benchmark and FROST kernel-time profiling.graph_analyzer.thd_stats_packing becomes the one classifier for packed Stats — token_major, head_major or None — replacing the backward probe's, the lowering's and the two forward adapter sites' own copies, and the SM80 backward kernel reads packed Stats in both packings with an arch-agnostic lengths → cu_seqlens launch.ex2.approx per row, making MUFU the softmax warps' longest pipe while FP32 has slack. exp2_emul_pair evaluates two exp2 in 6 packed FP32/INT instructions with no MUFU — the same split the cuDNN backend kernel uses — and exp2_mixed routes a compile-time subset of a vector's pairs through it. Applied to d128 MXFP8, d128 FP8 and d192×128 bf16 prefill at cc 10.0.stats_use_log2 — LSE in base 2 ✨✨Flash-attention-style consumers (FlashInfer, FA2/FA3, TRT-LLM) define the LSE as max + log2(sum_exp); cuDNN returns max + ln(sum_exp), so callers convert with a separate elementwise pass — and a doubled conversion already shipped a silent 1.44×-wrong LSE with a bit-exact O (flashinfer-ai/flashinfer#4663). The convention is now owned at the graph level: SDPA_attributes::set_stats_use_log2 / graph.sdpa(stats_use_log2=...) (#931). The FROST forward engines serve it natively — every prefill kernel already keeps the softmax in the log2 domain, so each gets one const_expr-guarded lse *= log2(e).
Follow-ups fixed the FP8 Stats log-base binding, protected the SM120 split partials and covered native ragged decode GQA Stats writes (#1082). The attribute's final home is the unified softmax node (CUDNN_ATTR_OPERATION_SOFTMAX_STATS_LOG2), gated on cuDNN 9.27.0 (#1127); graph.sdpa(stats_use_log2=...) is unchanged.
state_indices for paged recurrent states (#1002). Serving stacks (vLLM, SGLang) keep the linear-attention recurrent state in a paged pool and hand the kernel one row id per sequence rather than a compact buffer in sequence order — FlashInfer's chunk_gated_delta_rule models this as state_indices. cuDNN had no equivalent, so a caller holding a pool had to gather the active rows in and scatter them back out around every call, costing more than the kernel saves. An optional mStateIdx is now threaded through the SM100 GDN prefill kernel; absent it the state row is the sequence, exactly as before.P = #SMs/(B·H) = 2 tuning (#1164).torch.cuda.ExternalStream when it is the default stream (#1165).(sum_m, k), B (l, n, k) C-contiguous, dense C-contiguous SFA/SFB buffers flat or in physical atom shape. Canonical SF buffers compile as flat 1-D pointers, since the kernels already rebuild the MMA-tiled SF layouts from the GEMM shapes. Pre-permuted kernel-facing inputs keep working unchanged.return_max_logit with preallocated outputs, stream ordering, compile-cache specialization and non-differentiable autograd results. A Q-stage-2 empty-tile bug that retained NaNs from reused TMEM is fixed by writing literal zeros.cudnn._torch_stream helper replacing eighteen local copies across sdpa, linear_attention, conv1d, flex, hstu, grouped GEMM, CSA, DSA and QAT; no plan-owned device memory in the GEMM/FROST MoE plans; dead ABI slots set to 0; and no host sync in the SWA execute path.nvidia-cutlass-dsl reads as a version skip rather than an import failure (#1121).row_tile_coord for the dX output loop fails with TYPE_UNSTABLE_JOIN when Quack 0.6.4 rewrites mixed constexpr/runtime guards under CUTLASS DSL 4.8. The dX loop gets its own dx_tile_coord; coordinate, predicate and arithmetic are unchanged.O*dO operand precision (#1134).Thanks to everyone who contributed to this release:
@Anerudhan, @Aneureka, @brandonfzhang, @C-TC, @Denny991, @egilliam-nv, @elfiegg, @hwanseoc, @icavan, @jhjpark, @jiayus-nvidia, @mdy666, @NVIDIA-JerryChen, @pmdavies-nv, @rmhaskarnvidia, @RomanAnders90, @thynics, @tiffany940107, @vedaanta, @XinboZhao, @YangXu1990uiuc, @yanqinz2, @yanzhuo607, @yihuawei, @yuweih205, @zach-ye0, @ZeYang1025, and @zhibinz-nv.
One column per month.
cuDNN Frontend v1.29.0 Release Notes
cuDNN Frontend v1.29.0 is the recommended version for cuDNN 9.26 and later releases.
import cudnn
from cudnn.hstu_attention import hstu_attn_varlen_funcHSTU (Hierarchical Sequential Transduction Unit) attention arrives as a complete CuTe DSL kernel family for Blackwell (#487) — packed variable-length forward and backward, FP16/BF16, head dimensions 64, 128, and 256, with full, causal, local, and arbitrary masks, paged-KV forward, and strided or preallocated gradient outputs. HSTU replaces softmax with a SiLU score transformation and derives block-sparse metadata automatically. Explicit-stream execution is allocation- and lifetime-safe, cross-device PyTorch streams are rejected, and output overlap is validated against all read-only metadata.
Built out over the release:
X/U inputs, optional SiLU, dropout, concatenated U/X outputs, and optional dWeight. Launch and vector configuration are chosen from the hidden dimension and the runtime device SM count, and changing a row stride does not trigger recompilation. Independent backward output strides followed in #962.cudnn.DSA.SparseAttentionForward and sparse_attention_forward_wrapper (#569) add the SM100 sparse forward path, so the DeepSeek Sparse Attention forward/backward workflow now completes inside the frontend-only CuTe DSL API instead of requiring an external FlashMLA forward. H64 D512/D576 and the H128 D512 small-top-k prefill specialization are covered, including arbitrary logical top-k lengths, invalid/out-of-bounds and duplicate indices, per-query lengths, attention sinks, and an optional indexer LSE. Supported on SM100-family capabilities 10.0, 10.3, and 10.7; decode, split-KV, regular H128, SM90, and FP8 cache paths are not included.
On the backward side:
import cudnn
mask_plan = cudnn.create_mask_plan(...)
out = cudnn.flex_attn_func(q, k, v, mask_plan)The experimental cudnn.flex_attention namespace (#775) brings CuTe DSL forward and backward kernels for SM90, SM100, and SM103, covering fixed-length and variable-length MHA/GQA workloads, with compact arbitrary-mask planning, scheduling, runtime compilation and cache support, autograd integration, and lazy top-level exports. The port also syncs the CUTLASS DSL 4.6.0/4.6.1 bulk-copy election fix that otherwise deadlocks the SM90 backward. An SM100 2-CTA mask-slot synchronization bug was fixed in #993.
torch.sdpa runs on the cuDNN Python API 🚀 🚀torch.sdpa is now served end to end — forward and backward, dense and varlen — by the cuDNN Python API (#554).
The "CUDNN" provider (python/cudnn/torch/) registers with torch.nn.attention's flash-impl registry (PyTorch 2.13+, the same mechanism FA3/FA4 use) and overrides the CUDA kernels of aten::_scaled_dot_product_cudnn_attention{,_backward}, so F.scaled_dot_product_attention under sdpa_kernel([CUDNN_ATTENTION]) and torch.nn.attention.varlen.varlen_attn run on pygraph plus the engine Router. Registration is passive; activation stays explicit.
Alongside it, cudnn::sdpa_bwd gained dense BHSD backward — it previously served only packed THD and raised NotImplementedError on dense. With that in place the provider's dense backward stops falling back, closing the last C++ hop in a dense training step. Three defects were fixed in the same change: 2·B blocking device-to-host syncs per varlen backward (the host loop is replaced by the device-side thd_lse_to_padded() conversion, which makes the path capturable), dense operands never being normalized to the innermost-dense / 16B-aligned-base contract, and the dense autograd path dropping is_deterministic so use_deterministic_algorithms(True) still built a non-deterministic backward.
The experimental torch op was replaced by cudnn::sdpa_fwd / cudnn::sdpa_bwd (#780), which then gained window_right and asymmetric band service (#901).
Paged KV caches are served by the existing SDPA graph API and the FROST SM100 engine — no new entry point (#964). A graph built with cuDNN's own paged-cache contract (paged_attention_k_table / paged_attention_v_table plus paged_attention_max_seq_len_kv) now lowers onto the d128 f16/bf16 kernel's new PAGED_KV specialization. FlashInfer's cudnn_batch_decode_with_kv_cache already builds exactly this graph, so with CUDNN_FRONTEND_ENABLE_FROST_ENGINES=1 vLLM and SGLang decode reach the FROST kernel with zero integration work.
[num_pages, H_kv, page_size, D]) and NHD ([num_pages, page_size, H_kv, D]) page layouts — the layout is only the pool's strides, read off the bound strides, with no new TemplateParams field.num_pages / max_pages never enter the compile key. graph.execute runs clean under torch.cuda.set_sync_debug_mode("error"), and the adapter path captures under torch.cuda.graph with seq_len_kv contents changing between replays.FP8/MXFP8 pools, sinks, unaligned page sizes, d > 256, a missing padding mask, and packed block tables are declined at plan time.
SM100 d=512 backward (#887). sdpa_bwd_sm100 serves the band d ∈ (256, 512] on Blackwell; before this, every sdpa_backward() there fell through to the cuDNN backend, which does not serve it. It is a three-stage chain, not one fused kernel — a fused d=512 backward needs 512 TMEM columns for dV and 512 more for dK against 512 per CTA, so S and dS go to a GMEM workspace and the gradients become three batched GEMMs (do_dot → bprop_d512_f16_sm100 → bprop_matmul_sm100). Stage 2 forks the forward d512 kernel's cga4x1 role split. Two user-visible consequences are documented in the support matrix: the workspace is 2·B·H_chunk·S_q·S_kv·2 B (the host loops over head chunks to hold it under 4 GiB), and the band is envelope-served, so d=264 pays d=512's MMA cost. It covers GQA/MQA, causal (top-left and bottom-right), SWA, right-band widening, arbitrary S_q/S_kv, and BSHD plus arbitrary dense stride order. THD/ragged support followed in #898, GQA/MQA under THD in #922 (retiring both THD conjunction flags), and a sXfer bank-conflict fix in #941.
Quantized and large head dims. Per-tensor FP8 d512 forward on SM100 (#845), SM100 D512/D512 MXFP8 prefill forward (#989), a specialized SM120 D512 F16 prefill kernel (#991), SM100 D256 FP8 and MXFP8 kernels (#860), the SM100 d=256 MXFP8 backward engine sdpa_bwd_sm100_mxfp8 (#904), a specialized SM120 f16 d256 kernel (#930), and completed SM100 D192/D128 feature support with performance tuning (#841). Split-KV partials are always stored in fp32 on SM100, and KV split is allowed with a quantized output (#891).
Rubin (SM107) prefill forward (#954) — nine new kernels joining the already-shipped d128 FP8 sibling behind three opt-in engine rows (sdpa_fwd_prefill_sm107, sdpa_fwd_prefill_sm107_fp8 widened to d256/d512, and sdpa_fwd_prefill_sm107_mxfp8), spanning f16/bf16, per-tensor FP8, and MXFP8 across d128, d192×d128, d256, and d512. Rubin gets its own config_sm107.py because it is a different kernel lineage, not the Blackwell kernels recompiled: routing the SM107 d256 FP8 kernel through the Blackwell make_cfg_d256 produced a scheduler mbarrier init count of 15 against 11 actual arrivers — an unreachable barrier that hung at every shape, down to B=1 H=1 S=128, with no fault and nothing for a sanitizer to find. Follow-ups added DSv3 quantized flavors, THD on every f16/FP8 flavor and arch-grouped forward kernels (#974), SCHED_LPT on the fp8 row (#850), SCHED_LPT honored with STAGES_KV restored as a knob (#1001), and SCHED_LPT claimed at d256 / d192×128 with heuristics ranking from the flavor's domain (#1020).
Ampere (SM80) backward. The q-loop is now bounded by the sliding window — 6–12× on gpt_oss-style SWA — with a window-aware deterministic relay (#866). THD plus attention sinks and THD plus deterministic dQ landed in #867, native strided-LSE reads and a THD max_s_kv grid hint in #766, the port onto SdpaBwdDsl + TemplateParams with sym_int THD extents in #765, workspace carving that removes per-execute allocation on the engine paths in #716, and conflict-free dQ staging with an L2-grouped causal grid order in #948.
THD, scheduling, and plan selection. THD LPT remap (#717), THD persistent grid (#848), split-KV ports (#768), native strides (#795), the CLC scheduler aligned with the canonical CUTLASS pattern (#894), and python-native validate() per engine family (#869), which defers the C++ lowering while a Python engine is still a candidate and fixes the issues that exposed on the SM80 backward and SM100 FP8. FROST GEMM and SDPA kernel symbols are now prefixed with cudnn (#854).
kda_cake (#912) — a third engine for kimi_delta_attention, hosting the CAKE-generated recurrent KDA training kernels. The kernel bodies are vendored byte-for-byte from FlashInfer (with UPSTREAM.md and SHA256SUMS), compiled with NVRTC at first use and launched through the driver API, with an on-disk cubin cache keyed by source digest, options, and NVRTC version. The engine is opt-in — check_support declines unless CUDNN_FRONTEND_ENABLE_FROST_ENGINES=1 is set. Training kernels ship from cuDNN Frontend and inference kernels from FlashInfer; hosting CAKE here next to kda_frost and kda_cutile keeps one NVIDIA training path for the operator and gives it autograd, the cudnn.fla drop-in, and the JAX path for free.B·H performance (#969), and gate + beta in F16 with PDL enabled (#819).headdim=32 for GDN (#994), and stale TMA descriptor flags for short varlen tails (#1015, #1013).tanh_clamp_scale soft clamp for the squared-ReLU grouped GEMM epilogues, forward and backward (#858), with its instruction count subsequently cut (#906); row/column-wise quantize optimization (#833); epilogue tile-size selection (#849); and 1-CTA FROST plans preferred for quantizing epilogues (#808).requires-python raised to >= 3.10 (#921). The OSS kernel engines (FROST SDPA, GEMM fusions, sparse and linear attention) are Python kernels that JIT through CuTeDSL — a headline part of the package rather than an add-on — so the runtime they need installs by default. The DSL version is additionally gated at runtime so an older DSL declines rather than fails (#917), with follow-through on the SM80 SDPA gates, a prerelease-aware floor, and a lazy-import message (#919).revision() / size() accessors (#909).Bias dim and stride are validated like every other SDPA I/O tensor (#823); the forward now accepts any Stats output layout, dropping the pre-9.26 packed-BHSD guard (#1023).BackendDescriptor's constructor avoids copies (#761).docs/*.md files moved into docs/utilities/, the RST-authored guides were converted and brought over, and fern check and publish workflows were added. Markdown across the tree was made MDX-safe (#968), and the docs site base path moved to /cudnn (#1016).O-stride semantics in the fp16_fwd sample were clarified (#756); and AGENTS.md codified the THD Stats packing rule, the editable-install gotcha, and a PR review checklist (#843).test/python/sdpa/suites, covering context, generation, and bprop across dtypes and model presets (#846). test_mhas_v2 gained flash-style references using roughly 16× less memory, sequences up to 8192, and more L0 configurations (#910), naturally produced deeply negative attention scores in every test (#743), and a mismatch budget for FP8 gradient compares at rounding boundaries (#947). The legacy test_mhas.py was removed as superseded (#986).test/api_index (#997) and then under test/python (#998), refreshed from the development CI wheel (#1008).cudnn_oss benchmarked for kimi_k3/deepseek_v4 and FA2 on Ampere (#852), the unused Amax_S output dropped from the FP8 graphs (#847), artifacts refreshed against backend 9.26.0.39 (#771), a benchmark config-selection fix (#870), and README image links fixed (#772).sm103 FROST SDPA failure fixed (#753), a conv ReLU execution-plan creation test fixed (#949), torch.matmul used instead of torch.einsum in test_mhas_v2 references to avoid SM107 worker crashes in cuBLAS (#952), causal-conv contract tests kept compatible with CuTe DSL 4.6.2 (#936), FP8 dprob reassociation drift allowed on Rubin (#939), zero-length sequences waived below cuDNN 9.25 (#944), a non-functional CUDA IMA guard removed (#791), guardword violations fixed (#929), and sensitive flags removed (#907, restored selectively in #923).P store raced the other warpgroup's S load. mhas_v2 forward now draws d=256 for FP8/MXFP8.mhas_v2 decode sweeps draw THD.test_mhas_v2 (#646, issue #624).O TMEM loads are guarded (#757).setmaxnreg budgets were rebalanced to fix d=256 register spills (#840); the cutlass prefix is kept for the SM120 backward (#903).create_tensor_map_tiled (#839), and another to avoid an overflow (#872).cudnn::sdpa_fwd / cudnn::sdpa_bwd (#780).test_mhas.py is removed, superseded by test_mhas_v2 and sdpa/suites (#986).Thanks to everyone who contributed to this release:
@Adnios, @adshen, @Anerudhan, @Aneureka, @brandonfzhang, @Butterfingrz, @cb521, @egilliam-nv, @fallintoplace, @hwanseoc, @icavan, @JacoCheung, @jhjpark, @jiayus-nvidia, @Jie-Fang, @ksivaman, @liujane-dev, @miyanyan, @pmdavies-nv, @RomanAnders90, @SolitaryThinker, @SuperGoodGame, @swalters22, @thlurte, @thynics, @tp5uiuc, @vedaanta, @wanyingw, @YangXu1990uiuc, @yanqinz2, @yanzhuo607, @yihuawei, @yujincheng08,
Note truncated.
cuDNN Frontend v1.28.0 Release Notes
cuDNN Frontend v1.28.0 is the recommended version for cuDNN 9.25.1 and later releases.
cudnn.fla — a drop-in accelerator for flash-linear-attention 🚀 🚀cudnn.fla (#596) monkeypatches the flash-linear-attention ops that cuDNN can serve onto cuDNN's Blackwell (SM100) kernels, with a transparent fallback to FLA everywhere else, so results never change:
import cudnn.fla
cudnn.fla.accelerate_fla() # before importing FLA layers/models
import fla # GatedDeltaNet / KDA now run on cuDNN where supportedchunk_gated_delta_rule) — the GDN convention (log-space decay g, post-sigmoid beta, GVA where HV > H) mapped onto cuDNN's native op, reproducing the fused-layer knobs use_gate_in_kernel, use_beta_sigmoid_in_kernel, and use_qk_l2norm_in_kernel.chunk_kda) — channel-wise gate plus scalar beta, l2norm forward and backward through cuDNN. BF16 only; FP16 declines and falls back.GatedMLP (#686) — an opt-in adapter (accelerate_fla(targets="gated_mlp")) backed by cudnn.gemm.ops.swiglu_mlp. The patch registry is target-selective, incremental, idempotent, and independently restorable via restore_fla(targets=...) / is_accelerated(target).Configurations cuDNN cannot serve raise cudnnGraphNotSupportedError / NotImplementedError and fall back to FLA — never a wrong answer. Correctness is pinned by test/python/linear_attention/test_fla_compat.py, which requires cuDNN to match FLA within FLA's own BF16 noise on the output and every gradient.
Underneath, the linear-attention stack gained KDA and GDN-2 backward support (#556), a safe beta guard for GDN-2 (#722), packed-QKV views for native GDN (#685), state-layout and convention alignment with FLA/FlashInfer plus a context/IMA fix (#644), and successive CPU-overhead and instruction-cache/numerics passes on the FROST linear-attention kernels (#616, #708, #759). See docs/fe-oss-apis/fla.md.
The GEMM CuTeDSL APIs are now type-erased (#529): every API under python/cudnn/gemm/cutedsl/ accepts JAX arrays alongside torch tensors, and the modules import and resolve their public symbols without torch installed — torch is imported only when torch tensors or dtypes are passed, and JAX only when JAX arrays are.
On top of that, cudnn.jax.call (#553) wraps CuTeDSL's native JAX integration (cutlass.jax.cutlass_call) and gives every JAX-reachable GEMM API a jax.jit entry point — gemm_amax, gemm_swiglu (including blockscaled MXFP8), gemm_srelu, gemm_dsrelu, gemm_proj_rope_mxfp8 (both BF16 and MXFP8 input paths), and the grouped and discrete-grouped families in their pointer-array modes. APIs without a JAX data path raise a clear error rather than failing obscurely. JAX outputs the kernel already writes are no longer zero-initialized first (#631).
cudnn.Handle 🚀 🚀cudnn.create_handle() now returns a Handle object that owns {backend handle, device, stream} instead of a bare int (#612). The per-handle state that had accreted as module-global side tables and per-engine device queries — the stream cache, and the three parallel device stacks used by the backend handle, pygraph, and FROST — unify behind Handle.stream and Handle.device. This matters because the Python engines (FROST, CuTeDSL, linear attention) need a device and a stream, not a cudnnHandle_t.
Backward compatibility is transparent for normal use: every handle-taking API (execute, set_stream, get_stream, destroy_handle, all graph methods) is Handle-aware, extracting the backend handle explicitly at each named handoff. Design notes and a full call-site inventory are in docs/handle_first_class_design.md.
from cudnn.gnn import CscGraph, agg_simplecudnn.gnn.agg_simple (#647) exposes the cuDNN GNN AggSimple backend as a PyTorch custom operator with autograd, fake-tensor, and torch.compile support, handling graph validation and backend invocation so callers never touch the low-level GNN structures. Requires cuDNN 9.26 or newer and compute capability 8.0+; not supported on Windows. See docs/operations/gnn/agg_simple.md.
The FROST engine family introduced in v1.27.0 now spans every architecture the frontend targets.
cudnn.sdpa adapters (#493), later ported to plan-time compilation with TemplateParams kernels and sym_int THD extents (#689). Previously the manifest had no engine below SM100.d_qk=192 / d_v=128 support (#507), and a backward engine (#486) extended to all head sizes ≤256 including d192/d256, non-compact layouts and deterministic dQ (#533), sliding-window attention (#505), GQA/MQA, padding masks, right-band widening and sink-token gradients (#557), deterministic 2-kernel mode and dBias (#707), and native service of declared strided layouts (#666).has_lse specialization and a static SMEM guard (#579), a fused LDTM row-max and row-sum-in-MMA epilogue (#580), and the softmax_precision knob axis lit up with an F16x2 exponent on the d128 sibling (#651).sdpa_bwd gains MLA support (#643).Head-dim envelopes and engine identity. Per-tensor FP8 now serves the dense head-dim ENVELOPE through the same TMA zero-padding path the F16/BF16 flavors use, and the engine table collapses to one engine per architecture × dtype family — head dims became a lowering concern (kernel-flavor selection) rather than an engine identity (#587).
Ragged / THD, with zero host reads. THD execute on SM100 (#606) and SM120 (#608) now performs no device-to-host reads at all — no .tolist() syncs, no host cumsum, no pageable H2D — building its metadata on device against a plan-time envelope grid, which makes the path CUDA-graph capturable (issue #552, with #543 binding host prep to the launch stream and plan-time-only THD compile keys). The FP8/MXFP8 SM100/SM107 engines were moved onto the same envelope design (#648) and the legacy pre-envelope THD leg removed (#622). Supporting work: native THD declared-stride support in the SM100/SM120 F16 forward kernels (#526), the cu_seq_len prefix-sum length form (#522), ragged stats on SM100 (#512) and SM120 (#508), and ragged S_kv tails served on the F16 rows via synthesized padding (#581).
Masking, splitting, and heuristics. Forward heuristics can now recommend the same engine several times under different complete knob assignments, which makes split-KV graph-reachable for the first time (#692); recommend() is a pure, backend-blind entry point that autotuners can call with hand-built graph facts. Split-KV also landed for the SM100/SM120 prefill kernels (#658), the KV split the heuristic chose now runs on the true cluster shape (#720), and pack_gqa is supported (#709). On masking, SM100 gained causal right-band widening with per-sequence THD bottom-right diagonals (#485), bottom-right diagonal plus sliding window (#584) — after which the bottom_right_with_swa notch was retired because every row serves BR + SWA (#623) — and full causal mask support (#498).
Other FROST SDPA work. SM100 MXFP8 for d192/d128 (#661); dense LSE written directly to non-contiguous, dense-compatible layouts (#712); an execute path made async where it can be, no longer re-deriving build-time facts (#570); FP8 scales folded in-kernel with a baked 2⁴ P-cast bias, removing Scale_S from below the graph (#619); the Amax_S output dropped from the FP8 kernels (#602); a has_lse specialization for the FP8/MXFP8 SM100 flavors (#574); and a strict LSE/sink/seq-lens execute contract with no torch.empty in execute (#484).
Removal: the legacy standalone SM100 d=256 forward and backward stacks, along with the
cudnn::sdpa_{fwd,bwd}_d256experimental torch ops; SM80 forward moved onto the sameSdpaFwdDsladapter path SM100/SM120 use, so one lowering function drives every forward cell (#682). d=256 remains available through the graph API on backend engines.
SDPA
max_total_seq_len_q / max_total_seq_len_kv on the forward node (#740). sdpa_backward has accepted these since cuDNN 9.6; the forward node never did, so a ragged graph could not express its packed token total and the FROST forward path had to infer a loose upper bound from the bound buffers' element span. A loose bound is memory-safe but not benign — masked rows are still multiplied, so an over-allocated, unwritten tail poisons whole tiles through 0 * NaN. Every framework already holds this number (q.shape[0] in vLLM, SGLang, TransformerEngine, Megatron-Core, PyTorch, FlashInfer); it can now be declared.Stats with a narrower dtype — explicitly, or implicitly by leaving it unset with a non-FP32 io_data_type — built and executed fine, and the kernel then wrote FP32 rows past the end of the caller's buffer, surfacing as silent corruption of adjacent allocations, illegal memory accesses, or driver launch failures. Stats is now set to FLOAT at creation, an unset dtype defaults to FLOAT, and a narrower declared dtype is rejected.post_validate_node (#642).Capabilities.bottom_right_padded_seq_q was retiled (#683).Serialization and plan management
Graph::deserialize(blob) overload (and pygraph.deserialize(blob)) rehydrates a serialized execution plan from a DeviceProperties descriptor instead of a cudnnHandle_t, enabling ahead-of-time compilation: build and serialize a plan on a GPU node, deserialize it later where no CUDA context or cuDNN handle exists. Requires cuDNN ≥ 9.8 at compile and runtime; the API compiles on older headers and returns a runtime error.Tensor_attributes::alignment is now serialized (#564).CUDNN_KNOB_TYPE_TILE_CGA is mapped (#729). An engine reporting a knob the mapping did not carry returned it as NOT_SET, and feeding that back through create_execution_plan() failed for every knob combination on that engine — making the engine impossible to drive through the explicit-plan API at all. KnobType_t::TILE_CGA is added and mapped in both directions, and exposed to the Python bindings.Python dispatch
propose_plans, the Router, and heuristics_sort are deleted. An engine cannot rank plans — it sees neither its siblings nor the backend's entries — and all four in-tree propose_plans were the base class's default copied verbatim. create_execution_plans() now goes straight to heuristics.rank(...), which delegates to each family's recommend().graph.execute(). This closes three cases where the same public call answered differently depending on which plan the heuristics happened to pick — bare device addresses, override_shapes on a FROST plan, and related identity-dependent behavior.Operations
cudnn.ops.fft_causal_conv1d(x, weight) following cuhyena's medium/long selection, padding, trimming and autograd behavior, preserved long-forward reserve space for the matching backward, C++ samples, notebooks, and docs. SM107 causal conv1d tests are skipped before cuDNN 9.26 (#632).GEMM and MoE fusions
cudnn.gemm.ops.swiglu_mlp (#609) — a dense BF16 autograd op for out = (silu(x @ Wg.T) * (x @ Wu.T)) @ Wd.T on SM100. The forward gate/up GEMMs, SiLU, and multiply run as one FORT-native runtime-fusion kernel that also emits gate and up, avoiding two recompute GEMMs in training; the backward fuses dh = dout @ Wd with the two-output dSwiGLU epilogue in one FROST kernel, keeping dh on chip. Unsupported layouts, architectures, and missing optional dependencies decline to nvjet plus pointwise.dprob for grouped GEMM dsrelu (#521) and FP32 row-scaled FP4 grouped GEMM (#461).grouped_gemm_wrapper_sm100 118.1 → 39.6 µs, grouped_gemm_glu_wrapper_sm100 161.1 → 40.6 µs, grouped_gemm_dglu_wrapper_sm100 185.0 → 51.1 µs (kernel time 18.4 µs).elect_one compilation hint optimized (#504).DSA (DeepSeek Sparse Attention)
backend="sm100_v2", a drop-in for the SM100 sparse indexer backward that is 1.16–1.92× faster on the GEMM stage (1.92× at topk=128, 1.31× at topk=1024, 1.17× at topk=2048; 1.31× end-to-end through the public wrapper at S=8192/topk=1024). A two-term BF16 hi/lo expansion of A = g·w additionally keeps d_index_k FP32-accurate at ~no cost for consumers that keep index_k in FP32. Scope is SM100 exactly, H=64, D=128, topk ∈ [128, 2048] in multiples of 128, sm_scale > 0, request-or-fail with no silent fallback; the default backend is untouched.d_qk=576, d_v=512 BF16 (#664).indexer_backward now validates the output and plan signature before kernel 1 on the default SM100/SM90 backends (#572), and range_constexpr was restored in eight kernel_gemm epilogue loops (#549).Block-sparse attention (BSA)
block_sparse_attention_fp8_forward API that quantizes contiguous BF16 BHSD inputs to FP8 E4M3 internally using the Sage recipe and returns contiguous BF16, gated on CUTLASS DSL 4.6.1 at runtime. Covers SM100/SM103 (blk64, with automatic split-KV selection) and a dedicated SM120 kernel with sequence tails, fixed or variable sparse counts, and batched block_sizes layouts. Persistent CLC scheduling now works together with split-KV: the scheduler's work-tile mapping explicitly encodes and decodes the split dimension.CSA (Compressor)
ratio=2 (#710) — ratio ∈ {2, 4, 128} with coff ∈ {1, 2}, the configuration used in production training for the model family this operation serves. No kernel changes; the previous gate encoded validation scope, not a kernel limitation. Review-response fixups for the ratio=128 kernels landed in #452.Toolchain
nvidia-cutlass-dsl ≥ 4.7.0 and check the version at support time, declining rather than failing when the installed DSL is older; the package itself is deliberately not pinned to that floor so it stays compatible with consumers holding the DSL back. The packed-FP4 wgrad layout workaround is now gated on cutlass-dsl < 4.8 (#764).pre-commit now runs as a GitHub Actions job.merge-requirements job fails while a PR has no Milestone or is not on any Project board, with a matching PR-template checkbox. Bot-authored PRs and PRs labeled cat-routine-update are exempt, and a failed Projects lookup reports a clear error rather than a false pass.setup.py defaults to a parallel extension build (#565), and a -Werror unused-parameter build break in init_gnn_submodule was fixed (#728).benchmark/attention_inference/ measures attention as served, in two phases: context (TFLOPS; full prefill and chunked prefill of 512/1024-token chunks against 64k/128k caches, bottom-right causal) and generation (GB/s and % of memory SOL; q_tokens = 1 + MTP for MTP 0–3 against a 128k cache). Two backends are swept and charted — cudnn on native backend engines, and cudnn_oss planned with heur_mode.OPENSOURCE plus CUDNN_FRONTEND_ENABLE_FROST_ENGINES=1 so only the frontend's open-source engines may serve it, recording the winning plan per case.QwenImageTransformer2DModel. These exercise the merged production paths: packed-QKV GDN (#685), the cudnn.fla GatedMLP shim (#686), and the backend-only public d256 SDPA after #682.cudnn_oss (FROST) backend with unified sustained-clock SOL and peak lines (#597), FA4 auto split-KV enabled by default (#607), a Qwen3-VL vision-encoder (ViT) config with GB300 results (#598, #629), SM-clock sampling on the GPU the benchmark actually runs on (#699), and an Ampere (sm80) row in the peak-MMA table (#715). GDN benchmarking was added in #501.test_mhas_v2 (#516) and sink tokens are drawn in the ragged backward suites (#630); DSA comparisons use only the effective top-k slice (#672); render-only tests were removed (#688); samples skip an unsupported case and drop an invalid cudaGraphDestroy (#734); the 02_low_level_api notebook uses an FP32 Stats tensor (#701).SDPA
sdpa_backward graph that set max_total_seq_len_q silently returned all-zero dQ/dK/dV when Stats was head-major — the forward O and stats were correct and no error was raised, so this corrupted training without ever surfacing. Head-major [h, total_q] is FlashAttention's and PyTorch varlen's softmax_lse layout, i.e. exactly the integrations that would reach for max_total_seq_len in the first place. Present unchanged in 9.22 through 9.26 on both sm90 and sm100.seq_len_kv sweeps (#575).O-descriptor row stride on the SM100 FP8 path (#577).Python and device handling
ensure_current_context also returned as soon as any context was current, so a thread bound to another GPU's context kept it. It now resolves the target context instead of accepting the incumbent. The missing ensure_current_context imports introduced with Note truncated.
Added a PEP 735 dev dependency group, so pip install --group dev works as a prerequisite to deprecating requirements.txt (#359).
cuDNN Frontend v1.27.0 is the recommended version for cuDNN 9.24.0 and later releases.
cudnn.pygraph 🚀 🚀cudnn.pygraph is now a Python-native graph class (#336). Graph structure — nodes, tensors, and parameters — lives in Python and is fully introspectable, while execution dispatches through pluggable backends: Python DSL engines and the cuDNN C++ backend.
engines.BaseEngine defines a propose_plans → build_plan → execute lifecycle, with each engine owning a stable engine_id in a reserved region.create_execution_plans() produces a ranked plan list mixing Python plans with a single delegating entry for the cuDNN backend. Plan indices are two-level and stable, so select_plan() and the classic at-index APIs keep working.validate(), the same conditional-output behavior, torch dtype acceptance, ragged (THD) offsets, and deserialize/build_plans passthrough.See docs/python_graph_and_execution_backends.md for the full design.
Note: the internal pybind class
pygraphwas renamed tobackend_graph(reachable only ascudnn._pybind_module.backend_graph). Nothing public imports the pybind name; the publiccudnn.pygraphis now the Python class.
Open-source cuDNN engines written with CUTLASS primitives (#476):
These engines are registered as Python engines and are selected through the new cudnn.pygraph router, so they are reachable from the graph API rather than only as standalone kernels — the linear-attention operations below are the first consumers of that path.
A new cudnn.linear_attention package (#476) provides gated linear-attention operations through the Graph API, as well as PyTorch custom operators, exported from cudnn.linear_attention.ops.
GDN/GDN_BWD).torch.library.custom_op, so they compose with autograd, torch.compile, and DDP.[total_tokens, heads, dim] tensors plus cu_seqlens boundaries — with grouped-value attention (GVA/GQA) and per-sequence recurrent state ports (initial state in, final state out).benchmark_single_linear_attention.py), with tests under test/python/linear_attention/.SDPA
use_deterministic_algorithm on sm90 can now route to an ordered-dQ engine whose workspace is linear — rather than quadratic — in sequence length. The faster dP-workspace path is still used when it fits the existing 256 MB limit (CUDNN_FRONTEND_ATTN_DP_WORKSPACE_LIMIT is honored). Long-sequence and THD/ragged deterministic training that previously failed to allocate now runs; results remain bitwise reproducible. No behavior change for cuDNN < 9.25 or other architectures.seq_len_*) or cumulative (cu_seq_len_*) representation. Paged-attention integrations that hold a cumulative Q prefix sum alongside per-batch KV lengths no longer need to materialize a KV-side prefix sum. Requires cuDNN 9.25.0 or later on the unified surface.cu_seq_len_q / cu_seq_len_kv are now exposed on the sdpa_fp8 Python binding (#366).Data types and operations
DOUBLE (F64) compute data type support for convolution and pointwise scaling attributes (#423).ValueErrors for NCW (2–256), NWH (2–128), B2B projection (2–32), and B2B mixer (2–256) (#465, #472). This prevents an unsupported width-specialized NWH launch that could fault the CUDA context on SM90.OperationBuilder_v8 path (#466, #467).Serialization and plan management
2.0 (#280). Integer UIDs are the tensor-table identity used by node references, ragged-offset references, and tensor dumps, so anonymous tensors are no longer dropped and duplicate names no longer collapse. Missing UIDs are assigned before validation while preserving user-supplied UIDs; malformed versions, duplicate identities, and dangling references are rejected with typed errors.Graph::serialize accepts a serialize_structure flag (default true), making it symmetric with the handle-based deserialize path and enabling plan-only round trips after deserialize(handle, ...) (#371).Build and integration
dynamic_cast has been removed from the public headers (#477).dev dependency group, so pip install --group dev works as a prerequisite to deprecating requirements.txt (#359).SDPA
d=256 flavor, the D_QK = 192 / D_V = 128 logical shape now has a dedicated Blackwell prefill kernel (BF16 and FP16, dense CGA2 classic pipeline). For top-left causal S=8192 it reaches 87% compute SOL / 820 useful TFLOPS — roughly 1.5× a padded d=256 proxy. The existing d128, d256, and d512 kernels are unchanged.sm_100, with samples skipped elsewhere (#473).GEMM fusions
_bf16in flavor, so recipes that project in either BF16 or MXFP8 are covered (#438).NotImplementedError rather than silently producing invalid results.DSA (DeepSeek Sparse Attention)
indexer_forward_top_k_wrapper produces Top-K indices, selected logits, optional fused softmax, and optional LSE without materializing the dense score tensor; deterministic=True resolves K-th-boundary ties toward the smallest local KV indices. Existing BF16 paths are preserved.indexer_forward_wrapper accepts an optional pre-allocated out tensor, avoiding repeated internal allocation in iterative calls (#470).CSA (Compressor)
+ APE, overlap-window transform, fp32 windowed softmax, gated weighted sum — collapses from roughly 39 forward and 51 backward kernel launches per call into one forward and one backward kernel. Two follow-on optimizations (32-bit vectorized forward access; kernel-side zero-writes in the backward) give 1.20–1.36× on the forward kernel and 1.21–1.50× on the backward region while leaving forward, dKV, and dScore bitwise unchanged.ex2.approx.ftz through fastmath= so it builds at the supported cutlass-dsl floor (#463).Block-sparse attention (BSA)
elect_one gate for 4.6.2 and 4.7 (#382, #453).Toolchain
nvidia-cutlass-dsl 4.6.0 (#368) and cleared CUTLASS DSL deprecation warnings across the CuTe DSL kernels, including the .ptr migration for cute.struct scalar fields (#365, #376).python -m cudnn.collect_env (#400) — a new environment-forensics tool for bug reports, wired into the issue template and README. It reports the frontend version with mismatch flags (imported cudnn.__version__ vs. pip metadata vs. torch.backends.cudnn.version()), traces the frontend's libcudnn dlopen search order, and distinguishes loaded from installed GPU libraries by parsing /proc/self/maps — flagging the version-confusion cases that dominate unreproducible issues. It is stdlib-only with individually guarded probes, so it still produces a report when import cudnn or torch is broken, and can be run standalone.AGENTS.md plus scoped guides under include/cudnn_frontend/, python/cudnn/, test/, and samples/; an llms.txt docs index; skills discovery for coding agents; and expanded CONTRIBUTING.md sections on development environment, testing, and formatting.AGENTS.md (#489).(B, S, H) tensor shapes directly instead of reshaping to (B*S, H, 1, 1) (#432).bench_moe benchmark (#348).test_mhas_v2: extended backward random head-dim coverage to d=256 (#425); removed ALiBi, score_max/sum_exp, dropout randomization (#435) and with_rope (#474) from the random forward/backward tests.C++ frontend
Engine_v8::Knob::getMaxValue(), which returned the minimum value (#443).ValueError in flatten_pass_by_value on malformed hex input (#343).Python / OSS kernels
_get_default_stream(None) now resolves to torch's current stream (#483).DSA
grad_loss is now consistently a single-element FP32 CUDA tensor (#354).head_dim = 576 (#396).indexer_top_k now falls back to scalar stores for odd top_k (#407), and out-of-bounds lanes in the variable-length indexer top-k are fixed (#410).license = "Apache-2.0 AND MIT".SPDX-License-Identifier tag. The complete per-file mapping — including the commit that introduced each surviving external line — is in LICENSING.md, alongside LICENSE.txt (Apache-2.0), LICENSE-MIT.txt, NOTICE, and THIRD_PARTY_LICENSES.txt.Thanks to everyone who contributed to this release:
@adshen, @Anerudhan, @bmanthos, @brandonfzhang, @chaseblock, @derdrdirk, @egilliam-nv, @fallintoplace, @hwanseoc, @hxbai, @JackRao123, @jhjpark, [@jiefan] @jiayus-nvidia, @kangbintNV, @kunlunl, @liujane-dev, @pmdavies-nv, @rmhaskarnvidia, @saltyminty, @sraman-rgb, @terminator123, @vedaanta, @vincejhan, @WanZzzzzz, @yanqinz2, @yanzhuo607, @YangXu1990uiuc, @yeliu-oss, and @zkyue.
Special thanks for the kernel contributions that came from outside this repository:
cuDNN Frontend v1.26.0 is the recommended version for cuDNN 9.24.0 and later releases.
cuDNN Frontend v1.26.0 is the recommended version for cuDNN 9.24.0 and later releases.
SDPA
Data types
BYTE_BOOLEAN frontend data type (#302). Boolean tensors automatically map to BYTE_BOOLEAN when running against cuDNN 9.25 and later (#339).Serialization and plan management
Graph::deserialize now accepts an enforce_precompiled option to require precompiled engine plans during deserialization (#323).run_warmup opt-out and a reuse-parsed-json overload to Graph::deserialize, reducing repeated parsing overhead (#329).Block-sparse attention - Video Sparse Attention
DSA Deepseek Sparse Attention
Grouped GEMM
grouped_gemm_quant_wrapper_sm100 now accepts an optional caller-provided output tensor (#338).cute.core.ThrMma and cute.make_fragment usage (#321) and switched the dGLU dbias reduction to a constexpr loop to fix a DSL 4.5 regression (#322).indexer_topk_wrapper (#312).reduce_dKV validity guard incorrectly comparing the top-k column position (#298).block_scale_quantize.h (#319).Thanks to everyone who contributed to this release:
@dimitar-asenov, @HollowMan6, @Hyaloid, @jiayus-nvidia, @Jie-Fang, @jiemingz, @NVIDIA-JerryChen, @phu0ngng, @shraiysh, @sraman-rgb, @szluyu99, @take-cheeze, @vincejhan, @Vinnie6167, and @zianglih.
cuDNN Frontend v1.25.0 Release Notes
cuDNN has moved completely to github for development. Please direct your PRs to develop and file issues in github.
cuDNN Frontend v1.25.0 is the recommended version for cuDNN 9.23.0 and later releases.
cu_seqlens in unified SDPA — the unified SDPA path now accepts cumulative sequence-length tensors, enabling variable-length (packed) batches without padding.CUDNN_ATTR_TENSOR_RAGGED_OFFSET_MULTIPLIER), letting ragged offsets be stored in coarser units and scaled back to element offsets by the engine. Exposed through Tensor_attributes (getters/setters, validation, serialization) and the Python tensor() bindings. Requires cuDNN 9.24.0.get_engine_and_knobs_at_index, which returns the structured (engine_id, {KnobType_t: value}) for a plan instead of a stringified tag, so a tuned plan can be persisted and replayed exactly via create_execution_plan(engine_id, knobs) even as plan enumeration drifts across versions. Available in C++ (Graph, Execution_plan_list) and Python.KnobType_t with SWAP_AB, INPUT_TMA_ENABLE, and OUTPUT_TMA_ENABLE.group_offset support to the reduction node (Reduction_attributes::set_group_offset), so cuDNN FE can express per-expert reductions for MoE grouped GEMM workloads. Wires CUDNN_ATTR_OPERATION_REDUCTION_GROUP_OFFSET_DESC with runtime version checks (cuDNN ≥ 9.24.0), and exposes the optional argument through the Python reduction binding.CUDNN_FRONTEND_CUDART_LIB_NAME, and the shim now warns instead of throwing when multiple libcudart libraries are found, improving robustness in containerized environments.getenv access and fixed C4996/C4005 compiler warnings on MSVC.sfd_col_d_srelu_tensor.cuDNN Frontend v1.24.1 Release Notes
cuDNN Frontend v1.24.1 is the recommended version for cuDNN 9.23.0 and later releases.
d=256. Requires cuDNN 9.23.0 or later.seq_lens.INT32_MAX.square_alpha scaling in dgeglu and dswiglu.The Native Sparse Attention forward-prop kernels, supporting head dim = 128 and optimized for the Blackwell architecture, were implemented in CuteDSL.
These kernels were a collaborative effort, jointly developed by: Jie Feng, Akash Mehra, Vincent Zhang, Dominik Ernst, Xinbo Zhao, Aditya Vavre, Vedaanta Agarwalla, Mingyang Wang, Anerudhan Gopal, Paul Springer, Yang Xu, and Nima Tajbakhsh.
cuDNN Frontend v1.24.0 Release Notes
cuDNN Frontend v1.24.0 is the recommended version for cuDNN 9.22.0 and later releases.
d=256. Requires cuDNN 9.23.0 or later.seq_lens.INT32_MAX.square_alpha scaling in dgeglu and dswiglu.The Native Sparse Attention forward-prop kernels, supporting head dim = 128 and optimized for the Blackwell architecture, were implemented in CuteDSL.
These kernels were a collaborative effort, jointly developed by: Jie Feng, Akash Mehra, Vincent Zhang, Dominik Ernst, Xinbo Zhao, Aditya Vavre, Vedaanta Agarwalla, Mingyang Wang, Anerudhan Gopal, Paul Springer, Yang Xu, and Nima Tajbakhsh.
cuDNN Frontend v1.23.0 is the recommended version for cuDNN 9.21.0 and later releases.
cuDNN Frontend v1.23.0 is the recommended version for cuDNN 9.21.0 and later releases.
cudnn-frontend now has pip wheels for python 3.14t.
y = activation(conv1d_causal(x, w) + b) Supports forward and backward passes with torch.autograd and torch.compile. (Not supported on Windows yet)Graph::transpose with Transpose_attributes(permutation, optional compute dtype, name)Slice_attributes with set_strides for per-axis slice steps; strided slices update inferred output shape and strides accordingly.pygraph.slice now honors each dimension's slice.stepConcatenate_attributes with set_in_place_index (optional). When unset, concatenate runs out-of-place per backend rules.ReshapeMode_t(VIEW_ONLY,LOGICAL) and Reshape_attributes::set_reshape_mode so reshapes can select view-style vs lexicographic logical reshape.cudnn.scalar_type(RUNTIME_PARAM,COMPILE_TIME_CONST) and Graph::tensor(scalar, ScalarType) overloads, so scalars can be execution-time variant-pack inputs or constants embedded in the plan.Tensor_attributes can be marked as a compile-time constant or a normal runtimepass-by-value scalar;16, and a per-CTA amax reduction.Fix block-scale quantize
The scale tensor uses a 128x4 reordered layout (TensorReordering_t::F8_128x4). When the reordering type is set on the scale tensor, the frontend will automatically pad the inferred scale dimensions to align with the 128x4 block structure (non-batch, non-axis dimensions are padded to multiples of 128, and the quantize axis dimension is padded to multiples of 4).
16, and a per-CTA amax reduction.Grouped GEMM APIs now default to dynamic MNKL compilation across GLU, dGLU, SwiGLU, dSwiGLU, SReLU, dSReLU, and quant wrappers. Set CUDNN_FE_GROUPED_GEMM_DYNAMIC_MNKL=0 to restore the previous M-only dynamic behavior.
Grouped GEMM wgrad wrapper APIs now support caller-provided output buffers (wgrad_tensor for dense, wgrad_ptrs for discrete)
Unused internal c_tensor removed from Grouped GEMM quant path
Grouped GEMM GLU bias compilation issue for 64B-aligned inputs with dynamic MNKL
Fix an issue with dropout in Blackwell when cudnn frontend 1.21 version is used with cudnn backend 9.21 and 9.22.
Kimi-K2.6, LTX-2, Qwen 2.5 , Wan2.2 to the benchmark results page.cuDNN Frontend v1.22.1 is the recommended version for cuDNN 9.20.0 and later releases.
cuDNN Frontend v1.22.1 is the recommended version for cuDNN 9.20.0 and later releases.
Introducing PyTorch custom operator wrapping cuDNN's MoE Grouped Gemm operation.
```python
def moe_grouped_matmul(
token: torch.Tensor,
weight: torch.Tensor,
first_token_offset: torch.Tensor,
token_index: Optional[torch.Tensor] = None,
token_ks: Optional[torch.Tensor] = None,
mode: str = "none",
top_k: int = 1,
) -> torch.Tensor
```
See test/python/test_moe_grouped_matmul_op.py for usage.
🕒 We will be rolling out new native custom torch ops in upcoming releases – stay tuned! 😃
nvidia-cutlass-dsl[cu13]==4.4.1GroupedGemmWgradSm100 and grouped_gemm_wgrad_wrapper_sm100 expose the grouped GEMM weight-gradient kernel. See grouped_gemm_wgrad.html for API reference moe_blockscaled_grouped_gemm_wgrad.py for samples.Blackwell sdpa fprop kernel supporting head dim = 256, written in cuteDSL kernel was jointly developed by Shengbin Di, Yuxi Chi, and Linfeng Zheng in close collaboration with Alibaba. We would like to extend special thanks to the core contributors from Alibaba: Siyu Wang, Haoyan Huang, Lanbo Li, Yun Zhong, Man Yuan, Minmin Sun, Yong Li, and Wei Lin for their significant contributions to this work.
cuDNN Frontend v1.22.0 is the recommended version for cuDNN 9.20.0 and later releases.
cuDNN Frontend v1.22.0 is the recommended version for cuDNN 9.20.0 and later releases.
Introducing PyTorch custom operator wrapping cuDNN's Scaled Dot-Product Attention (SDPA). scaled_dot_product_attention as the public entry point, closely
matching the signature of torch.nn.functional.scaled_dot_product_attention.
```python
def scaled_dot_product_attention(
query: torch.Tensor,
key: torch.Tensor,
value: torch.Tensor,
attn_mask: Optional[torch.Tensor] = None,
dropout_p: float = 0.0,
is_causal: bool = False,
scale: Optional[float] = None,
enable_gqa: bool = False,
*,
diagonal_alignment: int = 0,
left_bound: int = -1,
right_bound: int = -1,
seq_len_q: Optional[torch.Tensor] = None,
seq_len_kv: Optional[torch.Tensor] = None,
cumulative_seq_len_q: Optional[torch.Tensor] = None,
cumulative_seq_len_kv: Optional[torch.Tensor] = None,
) -> torch.Tensor:
```
Introduce a preindexed execute method, that reduces the CPU execution overhead.
Improve the reproducer tool to report and reproduce SDPA failures for fp8 data types as well.
🕒 We will be rolling out new native custom torch ops in upcoming releases – stay tuned! 😃
Blackwell sdpa bprop kernel supporting head dim = 256, written in cuteDSL. Support added through the torch-op above or callable as a standalone API. See samples for the API usage. Requires nvidia-cutlass-dsl[cu13]==4.4.1
Grouped Gemm + quantize kernels now support dynamic shape and layout. This is controllable via an environment toggle.
Grouped Gemm + Glu/Swiglu now supoprt optional bias fusion in both dense and discrete modes, including partial‑N support and optional bias‑gradient generation for discrete backward paths.
fp8 datatype with packed variable sequences (THD) is no longer supported for SM90 (Hopper) architecture.
Fix an issue where sdpa fp8 was failing when used with cuda toolkit 12.9
Blackwell sdpa bprop kernel supporting head dim = 256, written in cuteDSL kernel was jointly developed by Shengbin Di, Yuxi Chi, and Linfeng Zheng in close collaboration with Alibaba. We would like to extend special thanks to the core contributors from Alibaba: Siyu Wang, Haoyan Huang, Lanbo Li, Yun Zhong, Man Yuan, Minmin Sun, Yong Li, and Wei Lin for their significant contributions to this work.
cuDNN Frontend v1.21.0 is the recommended version for cuDNN 9.20.0 and later releases.
cuDNN Frontend v1.21.0 is the recommended version for cuDNN 9.20.0 and later releases.
Added new kernels for the GEMM fusions.
Grouped GEMM + GLU: Unified grouped GEMM GLU API supporting dense and discrete MoE weight layouts with optional bias. Grouped GEMM + dGLU: Unified grouped GEMM dGLU backward API supporting dense and discrete MoE weight layouts with optional bias. Discrete Grouped GEMM + SwiGLU: Per-expert-pointer SwiGLU grouped GEMM for MoE workloads without weight packing. Discrete Grouped GEMM + dSwiGLU: Per-expert-pointer dSwiGLU backward grouped GEMM for MoE workloads without weight packing. Uses dSwiGLU/dGeGLU backward epilogue. Grouped GEMM + dSwiglu: dSwiglu activation fused with Grouped GEMM Grouped GEMM + Quant: Grouped GEMM with output quantization for MoE FC2/dFC1 workloads
cuDNN Frontend v1.20.0 is the recommended version for cuDNN 9.20.0 and later releases.
cuDNN Frontend v1.20.0 is the recommended version for cuDNN 9.20.0 and later releases.
Allow GEMM + Amax, GEMM + SwiGLU, Grouped GEMM + SwiGLU, Grouped GEMM + dSwiglu, and NSA kernels to run on GB300.
Improve the reproducer tool to report and reproduce SDPA failures.
Pinning the pybind version to prevent failures with older versions.
Pinning the pybind version to prevent failures with older versions.
Restore support for cuda-12 toolkit that was accidentally dropped in 1.19.0 release.
cuDNN Frontend v1.19.0 is the recommended version for cuDNN 9.19.1 and later releases.
sm_version on the cuDNN graph.cuDNN Frontend v1.19.0 is the recommended version for cuDNN 9.19.1 and later releases.
cuDNN Frontend v1.19.0 is the recommended version for cuDNN 9.19.1 and later releases.
sm_version on the cuDNN graph.cuDNN Frontend v1.18.0 is the recommended version for cuDNN 9.18.1 and later releases.
cuDNN Frontend v1.18.0 is the recommended version for cuDNN 9.18.1 and later releases.
New open source kernel for Grouped Gemm and Swiglu fussion
New Features: Allows support for dynamic shapes for fprop. This will help reduce the graph building across different batch and sequence lengths.
Support Surface:
More samples:
moe_grouped_matmul. See cpp sample and documentation for API reference.cuDNN Frontend v1.17.0 is the recommended version for cuDNN 9.17.0 and later releases.
cuDNN Frontend v1.17.0 is the recommended version for cuDNN 9.17.0 and later releases.
Native Sparse Attention : The Native Sparse Attention (NSA) module implements Native Sparse attention as described in the Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. Samples of usage for Blackwell architecture in test/python/fe_api/nsa
Gemm/Swiglu : Gemm_Swiglu now supports block-scaled FP8/FP4 datatypes. API changes:
BufferError in non-pytorch tensors.cuDNN Frontend v1.16.0 is the recommended version for cuDNN 9.15.0 and later releases.
cuDNN Frontend v1.16.0 is the recommended version for cuDNN 9.15.0 and later releases.
This release introduces open-source implementations of commonly requested fused kernels for select architectures (Blackwell). These experimental kernels may require additional dependencies such as CuteDSL. The initial release includes:
Additional dependencies can be installed optionally using pip install nvidia-cudnn-frontend[cutedsl]. Usage examples and detailed documentation are available in the test/python/fe_api directory.
Please submit issue reports for additional kernel requests or bug reports.
Block Mask Support: Starting with cuDNN 9.14.0, SDPA attributes now support block masks to exclude tiles that do not require computation. Refer to the sample implementation for usage details.
Bug Fix: Resolved an invalid memory access (IMA) issue in SDPA backward propagation (fixed in cuDNN backend version 9.15.1 and later) that occurred when s_kv is not a multiple of 128, padding mask is disabled, and operations are performed in CUDA graph replay mode.
BehaviorNote_t::CUDNN_BEHAVIOR_NOTE_CUBLASLT_DEPENDENCY as a behavior note. This enables filtering of engine configurations (execution plans) that use cuBLAS as a backend, available starting with cuDNN version 9.15.0.Block Scale Quantization: Added Python bindings for block scale quantize operations (#173). Refer to the sample implementation for usage details.
Dependency Optimization: PyTorch is no longer a required dependency for cuDNN Frontend (#177).
Tensor Alignment: Enhanced tensor descriptor API to accept alignment as an attribute (#153).
Plan Generation Control: Updated cudnnGetPlan API to accept an optional maximum plan count parameter, enabling users to limit the number of plans built and autotuned.
cudnn frontend v1.15 is the preferred cudnn frontend version for cuDNN version 9.13.1 and above.
cudnn frontend v1.15 is the preferred cudnn frontend version for cuDNN version 9.13.1 and above.
cudnn.Graph API that enables interoperability between torch.tensors and the cudnn frontend API. Sample code for performing a matmul with bias addition:B, M, N, K = 16, 128, 128, 512
a_gpu = torch.randn(B, M, K, device="cuda", dtype=torch.bfloat16)
b_gpu = torch.randn(B, K, N, device="cuda", dtype=torch.bfloat16)
d_gpu = torch.randn(1, M, N, device="cuda", dtype=torch.bfloat16)
with cudnn.Graph(
intermediate_data_type=cudnn.data_type.FLOAT,
compute_data_type=cudnn.data_type.FLOAT,
inputs=["mm::A", "mm::B", "bias::bias"],
outputs=["bias::OUT_0"],
) as graph:
AB = graph.matmul(
name="mm",
A=a_gpu,
B=b_gpu,
)
C = graph.bias(name="bias", input=AB, bias=d_gpu)
C.set_output(True)
c_gpu = graph(a_gpu, b_gpu, d_gpu, handle=handle)
All notebooks under samples/python have been updated to showcase the flexibility of this API.
Graph now includes a warmup method that triggers kernel loading by performing a fake graph capture. This improves the startup time for running the initial kernel in the actual run and prevents deadlocks when used with other modules (e.g., NCCL).set_score_max and set_score_sum_exp to allow the kernel to output max attention score and sum of exponents.s_q==1 and s_kv==1.)COMPLEX_FP32 and COMPLEX_FP64 datatypes. (Requires cuDNN v9.14.0 or later.)fe::HeurMode_t::A over fe::HeurMode_t::FALLBACK.swish function now accepts a swish_beta parameter.📢 cuDNN Frontend v1.14.1 — Release Notes
📢 cuDNN Frontend v1.14.1 — Release Notes
9.11 and fixed in 9.13) affecting certain large head-dimension combinations of d_qk and d_v.fp16_fwd_with_sink_token.cpp
fp16_bwd_with_sink_token.cpp✅ Recommended Action: Upgrade to cuDNN Frontend v1.14.1 for full compatibility with cuDNN 9.13.0+, improved SDPA support, additional normalization support, and deviceless graph compilation features.
Preferred version for: cuDNN 9.12.0 and above Minimum Python version: 3.9 (previously 3.8, now obsolete) Updated pip wheels: Available for Python 3.13
Preferred version for: cuDNN 9.12.0 and above
Minimum Python version: 3.9 (previously 3.8, now obsolete)
Updated pip wheels: Available for Python 3.13
9.11) affecting certain large head-dimension combinations of d_qk and d_v.LayerNorm with ReLU bitmask dump.test_deviceless_aot_compilation.py✅ Recommended Action: Upgrade to cuDNN Frontend v1.14,0 for full compatibility with cuDNN 9.12.0+, improved SDPA support, additional normalization support, and deviceless graph compilation features.
When generate_stats is true, the output will contain the stats tensor. When migrating from is_inference (which is now deprecated), note that generate_…
cudnn frontend v1.13 is the preferred cudnn frontend version for cudnn version 9.11.0 and above.
Introduces device descriptor, which allows for device-less compilation of cudnn graph on a target GPU. See newly added sample and documentation.
Introduced generate_stats as a replacement for is_inference, to improve clarity. When generate_stats is true, the output will contain the stats tensor. When migrating from is_inference (which is now deprecated), note that generate_stats has the opposite meaning, so pass it the negation of the bool that was passed to is_inference.
Improved support checks for left and right diagonal bands in conjunction with the diagonal alignment.
Improved error handling for large head dimension (d > 128) in sdpa bprop.
Published improved SDPA training benchmarks for fp8 and fp16/bf16 graph patterns.
Enable int4 Weight only Quantization for matmul. See example
Allow block scale dequantize (required for low precision matmul) to take 2-D scale factor.
Allow reductions to accept deterministic as a attribute.
Added pybinds for block scale dequantize.
cudnn frontend v1.12 is the preferred cudnn frontend version for cudnn version 9.9.0 and above.
cudnn frontend v1.12 is the preferred cudnn frontend version for cudnn version 9.9.0 and above.
cudnn_frontend v1.12 is the minimum cudnn frontend version required to work with cuda 13.0 and above
Update the dlpack version and cmake minimum required version to be 3.18
Allows compilation and loading of cudnn frontend with cudnn-jit packages.
Introduce Adaptive Layernorm (fprop and bprop) operation in cudnn.
std::array<std::shared_ptr<Tensor_attributes>, 3>
adalayernorm(std::shared_ptr<Tensor_attributes>& input,
std::shared_ptr<Tensor_attributes>& scale,
std::shared_ptr<Tensor_attributes>& bias,
AdaLayernorm_attributes attributes);
std::array<std::shared_ptr<Tensor_attributes>, 3> adalayernorm_backward(
std::shared_ptr<Tensor_attributes> dy,
std::shared_ptr<Tensor_attributes> x,
std::shared_ptr<Tensor_attributes> scale,
AdaLayernorm_backward_attributes options);
Please refer to samples for usage.
cudnn.jit and cudnn.graph for simpler graph creation in python. Refer the matmul sample for usage.Allows large embedded dimension (d > 128) for fprop across Ampere, Hopper, and Blackwell architectures for bf16/fp16.
Added better validation checks for sliding window attention for cudnn version 9.9.0 and below.
Sliding windown attention now supports cases when s_q > s_kv
sdpa_fp8 operation now pads correctly with negative infinity on masking operation rather than high negative value. This improves the numerical stability of the sdpa operation with fp8 data type.
Paged attention now supports page tables in a packed format
Fixed the dlopen of cudart.so to look for the binary with version name.
Correctly fail when SDPA bprop is called on Blackwell with embedded dimension (d) > 128.
cudnn frontend v1.11 is the preferred cudnn frontend version for cudnn version 9.8.0 and above. With cuDNN frontend v1.11, the minimum supported cudnn
cudnn frontend v1.11 is the preferred cudnn frontend version for cudnn version 9.8.0 and above. With cuDNN frontend v1.11, the minimum supported cudnn version is 9.0.0.
Note: The FE will continue to build and run with cudnn_v8, until explicitly marked as compilation failure.
score_mod=partial(
custom_mask,
mod_tensor=mod_tensor,
neg_inf=neg_inf_tensor,
seq_len_q=seq_len_q,
seq_len_kv=seq_len_kv,
)
std::shared_ptr<Tensor_attributes>
concatenate(std::vector<std::shared_ptr<Tensor_attributes>>, Concatenate_attributes);
pip wheels compatible with windows x86_64 architecture are now available on pypi.
sdpa paged attention API now supports Q tensor to be ragged when used with cudnn version 9.7.0 and above.
Users can now pass the CMake flag -DCMAKE_CXX_FLAGS="-DNV_CUDNN_FRONTEND_DISABLE_LOGGING" to disable logging in the cuDNN frontend.
Adds a new sample to showcase native cudagraph creation from cudnn for sdpa bprop operation. Fixed a bug when using the update_cuda_graph API to update cuda graph for sdpa bprop operation.
Updates the create_container_and_page_table example function to use the layout that's desired for the more performant kernel."
Fixes memory leak in the test harness for some legacy tests that use ragged tensors.
Fixes a bug introduced in the benchmarking script that prevented the sdpa cudnn operation from being executed. This was because the use_padding_mask attribute was made mandatory for the sdpa operation. This has been fixed as well.
Updates the paged attention sample to not cause illegal memory access when changing the dimensions of the tensors in the sample.
Updates the DgradDReluBNBwdWeight sample to perform the right operation for the dgrad + drelu fusion.
cudnn frontend v1.10 is the preferred cudnn frontend to be used for cudnn backend 9.7.0 and later as it adds to the Blackwell specific features.
cudnn frontend v1.10 is the preferred cudnn frontend to be used for cudnn backend 9.7.0 and later as it adds to the Blackwell specific features.
cudnn Frontend v1.10 introduces two new operators, block_scale_quantize and block_scale_dequantize to specify the scaling and de-scaling of low precision datatypes supported from Blackwell GPU onwards.
create_execution_plan(int64_t const engine_id, std::unordered_map<KnobType_t, int64_t> const &knobs) allows creation
of a custom execution plan with hardcoded engine and knobs. Added a
sample in samples/cpp/misc/custom_plan.cpp to showcase how to work
with different Engine and Knobs.
Users can now query behavior notes of a particular execution plan
using get_behavior_notes(std::vector<BehaviorNote_t> ¬es) const and
get_behavior_notes_for_plan_at_index(int64_t const index, std::vector<BehaviorNote_t> ¬es) const functions.
SDPA operations now accept both left window and right window size with respect to diagonal. See Attention.md for more details.
SDPA operations now accept a diagonal alignment for the Attention
score matrix to be used describe the above window. When s_q != s_kv,
and causal mask is on this can be used to specify if the diagonal is top
left or bottom right.
Bottom right causal masking can now be enabled on the sdpa_fp8 operation.
SDPA_attributes and SDPA_bprop_attributes now accepts a score_mod function through set_score_mod and set_score_mod_bprop API. The function accepts a c
SDPA_attributes and SDPA_bprop_attributes now accepts a score_mod function through set_score_mod and set_score_mod_bprop API. The function accepts a custom chain of pointwise operations which operate on the Attention Score Matrix. Some common functors like causal mask, sliding window mask, soft capping etc. have been added to the headers as reference. More examples of usage have been added in samples for fprop and bprop.
Added support for THD format and sliding window mask.
Added support for THD format and Bottom right causal mask.
Added support for bottom right causal masking with sliding window mask
Added a new parameter called set_max_total_seq_len_q/set_max_total_seq_len_kv on the sdpa bprop node. This will help reduce the workspace size required when running with THD format.
Allow creation of serialized json for dgrad, wgrad and resample operations.
Added more diagnostic message when the compiled version of cudnn does not match the run-time version of cudnn.
Fixed an issue where log messages unparseable data at the end of messages.
Fixed an issue where while building the python pip wheel would hang.
Fixed natively creating cuda graphs for SDPA with alibi masks.
SDPA forward operation now supports paged attention on cudnn 9.5.0 and later by setting the appropriate page table descriptors. SDPA_attributes now ac
SDPA forward operation now supports paged attention on cudnn 9.5.0 and later by setting the appropriate page table descriptors. SDPA_attributes now accepts set_paged_attention_k_table and set_paged_attention_v_table to input these descriptors. Please refer to samples for usage : cpp samples, python samples. See docs for more API details. Paged attention allows for more efficient memory usage by storing K/V caches in non-contiguous memory, and using page tables to reconstruct them. For more information, refer to the cudnn_graph Library, and the Paged Attention paper
cudnn graph now allows user to directly build native cuda_graph for given sub_graph (requires cudnn 9.5.0). There are two APIs:
populate_cuda_graph : add the cudnn nodes to the empty cuda_graph provided as input.update_cuda_graph : update the populated cuda graph with necessary data pointers.
See docs and backend documentation for more details.Kernel cache for dynamic shapes are now supported in python. Added a sample to showcase usage.
graph.deselect_engines(str: ) has now a python equivalent through pybind11.
graph.tensor(...) can now accept int64_t scalars directly. (Previously limited to int32_t, float and fp16 data types).
fp8 sdpa attention now allows dropout and padding mask. Requires cudnn 9.5.0 and above.
More enhancements to pointwise output stride inferencing (for broadcast operation). For non-unary operands, the broadcasted tensor can now be either at IN_0 or IN_1.
SDPA backward operation now allows d upto 256 for Hopper. Requires cudnn 9.5.0 and above.
Fixed an issue while querying cudnnGetLastErrorString() from the backend. The error_t object will now have more meaningful message.
Fixed build issues seen with clang-19 compiler.
Fixed an issue where it was assumed a graph with bias in sdpa_bprop will always have a dbias.
Kernel Cache support for dynamic graphs Added New APIs to enable kernel cache support for graphs with dynamic shapes. Please refer to documentation fo
Added examples Convolution fprop dynamic shape, CSBR Graph dynamic shape, Matmul dynamic shape and Bias + Matmul dynamic shape to showcase use of dynamic shapes and kernel cache.
error_t
get_plan_name(std::string &name) const;
error_t
get_plan_name_at_index(int64_t plan_index, std::string &name) const;
Note: This name can be used later if you want to deselect_plan_by_name, if run into any potential errors.
query_tensor_with_uid(int64_t const uid, Tensor_attributes &tensor) const;sdpa fp16 bprop node can now compute dbias when padding mask is enabled (requires cudnn 9.4.0 and above).
sdpa fp8 (forward and bprop) nodes now support optional bias, dropout and padding mask(requires cudnn 9.4.0 and above).
Matmul fp8 node can now accept M,N,K overrides.
Added new python notebooks for implementing BatchNorm and BatchNorm bprop using cuDNN.
Updated benchmark numbers with cudnn 9.4.0 for fp16 and fp8 datatypes.
Fixed compilation issues when NV_CUDNN_DISABLE_EXCEPTION is enabled.
Fixed a crash when the output dimension of dgrad node is not specified. This now returns an error message instead.
Fixed incorrect SDPA stats stride inferencing.
Fixed a bug in sdpa test when sliding window attention is enabled and query sequence length (s_q) is greater than key length (s_kv). This case is now not supported.
Fixed an issue where custom dropout mask was not correctly applied.
-fvisibility=hidden for the pip wheels generated to avoid symbol conflicts with other modules that use cudnn frontend.c * d * h * w > 2 **31) tensors.Graph Slice Operation: Introduced the graph.slice operation for slicing input tensors. Refer to docs/operations/Slice.md for detailed documentation an
graph.slice operation for slicing input tensors. Refer to docs/operations/Slice.md for detailed documentation and samples/cpp/misc/slice.cpp for a C++ sample. Pybinds for this operation have also been added.set_sm_count(int32_t type) graph property to support the SM Carveout feature introduced in Ampere and Hopper GPUs. Engines that do not support SM_COUNT will return NOT_SUPPORTED.set_convolution_mode attribute to convolution attributes in forward propagation (fprop), data gradient (dgrad), and weight gradient (wgrad). Previously, this was hardcoded to CUDNN_CROSS_CORRELATION in the 1.x API.sdpa_fp8_backward node.graph.execute() by optimizing sub-node tree traversal, collected UIDs, workspace modifications, and workspace size.graph.validate() by deferring graph expansion to a later stage (build_operation_graph).create_execution_plans if called without the preceding build_operation_graph.graph.build() calls.CMAKE_SOURCE_DIR with PROJECT_SOURCE_DIR in CMake files for better integration. See the relevant pull request for more details.samples/python folder for more information.[Enhancement] Allows stride value of 0 indicating repetition of tensor in those dimensions.
[Enhancement] Allows stride value of 0 indicating repetition of tensor in those dimensions.
[Bug fix] Fixed an issue, where cudnn-frontend (1.5.0) when built with cudnn version 9.1.1 and below, runs into issues when run with 9.2.0 and above.
v1.5.1
[Bug fix] Fixed an issue, where cudnn-frontend (1.5.0) when built with cudnn version 9.1.1 and below, runs into issues when run with 9.2.0 and above.
[New feature] With cudnn backend 9.2.0 and above, Graph::check_support can determine support check for runtime engines without invoking the nvrtc comp
[New feature] With cudnn backend 9.2.0 and above, Graph::check_support can determine support check for runtime engines without invoking the nvrtc compiler. This allows users to check the support surface of cudnn without invoking the nvrtc compilation.
[New feature] Python pip wheel now contains the necessary c++ development headers.
[New feature] Sliding window attention is now supported as an attribute to the sdpa forward and bprop node. Usage: sdpa_attributes.set_sliding_window_length(window_length)
[New feature] Bottom right aligned causal masking is now supported as an attribute to the sdpa forward and bprop node. Usage: sdpa_attributes.use_causal_mask_bottom_right(true)
[New feature] SDPA bprop attributes can choose deterministic algorithm using the use_deterministic_algorithm API.
[New feature] Allow users to filter candidate execution plans of graph by its shared memory usage in cudnn 9.2.0 and later.
[Bug fix] A runtime error if chosen execution plan candidate is incorrectly set in the backend has been fixed. This would happen when check_support does not correctly filter by the workspace size.
[Bug fix] selecting/deselecting by behavior and numerical notes has now been fixed and works as intended.
[Debugging] A new tool for easy reproduction of a failure using the json representation of the graph can be found here.
[Samples] Restructured the cpp samples into categories for easier navigation.
[Samples] Added a sample to showcase how different plans can be built in parallel in separate threads.
[Compilation enhancement] Added a new macro CUDNN_FRONTEND_SKIP_NLOHMANN_JSON as compilation flag to not have nlohman::json as compilation dependency. Users lose access to certain API functions like print, key, serialize, deserialzie that depend on the library.
[Enhancement] Serialization of resample operation is now supported.
[Enhancement] Bug template has been added for new github issues
[New] Added a benchmark folder which contains a sample docker file to compare cudnn implementation of sdpa with that of the pytorch implementation.
[New] Added a benchmark folder which contains a sample docker file to compare cudnn implementation of sdpa with that of the pytorch implementation.
[Enhancement] Once an engine is de-selected by name, it will not be built as part of check support.
[Enhancement] The cudnn backend search order for wheels is as follows: (a) It will dlopen libcudnn.so.MAJOR_VERSION in the site packages. (b) It will try to dlopen unversioned libcudnn.so. This way pypi cudnn package nvidia-cudnn-cu* gets priority over default search path.
[Enhancement] Allow embedding dimension up to 256 (currently limited to 128) in sdpa fprop operation.
[Bug fix] Update the scale and bias shapes in batch norm sample.
[New API] Added new operations sdpa_fp8_forward and sdpa_fp8_backward to perform scaled dot prodcut attention of fp8 tensors. See more details in the
[New API] Added new operations sdpa_fp8_forward and sdpa_fp8_backward to perform scaled dot prodcut attention of fp8 tensors. See more details in the docs/operations/Attention.md and cpp sample in samples/cpp/mha.cpp. Pybinds for the fp8 nodes are also added.
[New API] Added new operation for resample forward operation. Add a new sample samples/cpp/resample.cpp to show its usage.
[New API] Add a new API deselect_engines(std::vector<std::string> const &engine_names) which blocks certain engine configs from running.
[New API] Add new APIs select_numeric_notes and select_behavior_notes to allow user select engine configs which have the selected numeric and behavior notes respectively.
[Python API] Added a custom exception cudnnGraphNotSupportedException to the python API to distinguish between graphs that are actually not supported as compared to programming errors.
[Python API] Added a new backend_version_string which returns the backend version in canonical form (eg. 9.1.0) instead of a version number.
[Bug Fix] Fixed issues with compilation on clang19 and c++20 standard.
[Bug Fix] Updated the workspace computation for sdpa fprop node. Previously, workspace was calculated for alibi slopes irrespective of whether alibi mask was turned on or not.
[Bug Fix] Fixed deserialization of fused scalars.
Your coding agent can read these notes before it upgrades. Set up the MCP server →