NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2556 most downloaded on PyPI
NVIDIA CUTLASS Python DSL
Last release 14 days ago
21 Sep 2026
Ships on a steady schedule
a new release about every 3 weeks
Nearly every release is documented
notes for 7 of 7 stable releases
Nothing withdrawn
no release was ever pulled
3 months old
8 releases · first in 2026
One column per month.
Existing GroupedGemmArguments is deprecated and will be removed in a future release.
Initial Rubin support to accelerate dense GEMMs. The following features are available:
NOTE: Executing Rubin kernels (SM107) requires the R615 driver which will be released
with CUDA Toolkit 13.4 GA. R610 from CUDA Toolkit 13.4 Developer Preview is not
sufficient.
CuTe DSL extensions has several new features:
cute_ext TMA load, store, multicast, and reduce-store operations. Explicit CTA-V maps remain supported as overrides.cute_ext GEMM mainloop and TMA epilogue helpers.This release includes an opt-in preview of the CuTe DSL extensions (cute_ext) compiler pipeline for ordinary Cute DSL kernels. This pipeline lets user mix cute_ext APIs directly into @cute.jit and @cute.kernel code and is required for kernels that mix the two API surfaces. You may test this feature with the following:
CUTE_DSL_USE_EXTENSION_COMPILER=1 python your_program.py
The pipeline is expected to preserve program behavior, but generated PTX/SASS may differ. The pipeline is planned to become the default in a future release.
Added examples for better control over Primitives' compiler warnings/errors introduced in 4.7.0. See the CuTeDSL/experimental/compiler_diagnostic/ directory.
IKET Profiler Tool
iket_enable_profiling=True). Task execution will generate an IKET range and individual pipeline stages in a schedule may generate separate ranges (iket_profiling_stages).A number of new examples were added in this release:
nvidia-cuda-nvdisasm is now an optional dependency of nvidia-cutlass-dsl via the optional [sass] extra. SASS dumping (CUTE_DSL_KEEP=sass / KeepSASS) now resolves nvdisasm from the bundled wheel (recommended since its version matches the DSL toolchain) or from a local CUDA Toolkit (CUDA_HOME/CUDA_PATH). A locally-provided nvdisasm must come from a CUDA Toolkit at least as new as the toolchain that produced the CUBIN. Installations that never dump SASS are unaffected.cutlass.jax.cutlass_callcute.autovec_copy emitted per-element instead ofThis release has been tested against the following packages:
Dense and blockscaled GEMMs in Operator API have preliminary Rubin support. These are provided as a preview and may need additional performance tuning.
Updated GEMMs include:
These kernels utilize the below new features in Rubin:
- Higher SMEM (328KB) and TMEM capacity (288KB)
- B-buffer reuse
- Enhanced mixed precision throughput
Operators can now be ranked by their estimated performance when nvMatmulHeuristics is available. See tutorial here
NOTE: This currently only supports Blackwell kernels as nvMatmulHeuristics does not yet support Rubin.
Standalone kernel implementations are now exposed through cutlass.kernels, in addition to those exposed through the Operator interface in cutlass.operators. This allows kernels to be called directly without looking them up first.
Custom Epilogue fusions now support partial (per-row or per-column) reductions.
IndexPtrGroupedGemmArguments is now used to represent Grouped GEMM with contiguous-offset/index-pointers. Existing GroupedGemmArguments is deprecated and will be removed in a future release.
sm_107a and sm_107f targets:
NOTE: Executing Rubin kernels (SM107) requires the R615 driver which will be released
with CUDA Toolkit 13.4 GA. R610 from CUDA Toolkit 13.4 Developer Preview is not
sufficient.
New features
Initial Rubin support to accelerate dense GEMMs. The following features are available:
CuTe DSL extensions has several new features:
cute_ext TMA load, store, multicast, and reduce-store operations. Explicit CTA-V maps remain supported as overrides.cute_ext GEMM mainloop and TMA epilogue helpers.This release includes an opt-in preview of the CuTe DSL extensions (cute_ext) compiler pipeline for ordinary Cute DSL kernels. This pipeline lets user mix cute_ext APIs directly into @cute.jit and @cute.kernel code and is required for kernels that mix the two API surfaces. You may test this feature with the following:
CUTE_DSL_USE_EXTENSION_COMPILER=1 python your_program.py
The pipeline is expected to preserve program behavior and performance, but generated PTX/SASS may differ. Note that this pipeline will become the default in the future, no earlier than 4.10.
Added examples for better control over Primitives' compiler warnings/errors introduced in 4.7.0. See the CuTeDSL/experimental/compiler_diagnostic/ directory.
IKET Profiler Tool
iket_enable_profiling=True). Task execution will generate an IKET range and individual pipeline stages in a schedule may generate separate ranges (iket_profiling_stages).A number of new examples were added in this release:
CuTe DSL now supports x86_64 Windows
CuTe DSL AoT now supports new host target: QNX8.0
Notebooks are restructured under examples/python/CuTeDSL/cute/notebooks and new notebooks for primitives will be added under examples/python/CuTeDSL/notebooks
Numpy is now not a default dependency
Bug fixes and improvements:
nvidia-cuda-nvdisasm is now an optional dependency of nvidia-cutlass-dsl via the optional [sass] extra. SASS dumping (CUTE_DSL_KEEP=sass / KeepSASS) now resolves nvdisasm from the bundled wheel (recommended since its version matches the DSL toolchain) or from a local CUDA Toolkit (CUDA_HOME/CUDA_PATH). A locally-provided nvdisasm must come from a CUDA Toolkit at least as new as the toolchain that produced the CUBIN. Installations that never dump SASS are unaffected.cutlass.jax.cutlass_callcute.autovec_copy emitted per-element instead oflink-libraries compile-option order so it is stable across processesIndexError on staged bool() with no argumentscute.compile on @cute.kernel with a user error instead of an ICEThis release has been tested against the following packages:
Dense and blockscaled GEMMs in Operator API have preliminary Rubin support. These are provided as a preview and may need additional performance tuning.
Updated GEMMs include:
These kernels utilize the below new features in Rubin:
Operators can now be ranked by their estimated performance when nvMatmulHeuristics is available. See tutorial here. NOTE: This currently only supports Blackwell kernels as nvMatmulHeuristics does not yet support Rubin.
Standalone kernel implementations are now exposed through cutlass.kernels. These kernels can be used directly, in addition to being discoverable and usable via the Operator interface in cutlass.operators.
Custom Epilogue fusions now support per-row or per-column reductions.
IndexPtrGroupedGemmArguments is now used to represent Grouped GEMM with contiguous-offset/index-pointers. Existing GroupedGemmArguments is deprecated and will be removed in a future release.
sm_107a and sm_107f targets:NOTE: Executing Rubin kernels (SM107) requires the R615 driver which will be released
with CUDA Toolkit 13.4 GA. R610 from CUDA Toolkit 13.4 Developer Preview is not
sufficient.
Nothing published for this version
Fixed a failed kernel compilation issue when setmaxnreg is used together with specific warp-specialized patterns ( !3420 )
cutlass.Numeric from Python value (!3443)This release has been tested against the following packages:
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a sta
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development velocity. Primitives are released as experimental and will evolve based on user feedback.
NOTE: Primitives is a transitional API until a CUDA Python-like solution is available.
Introduced the Task Scheduling framework. This provides static analysis of execution schedules for warp-specialized kernels. Compilation stops when known concurrency issues are detected. Also provides tools for visualizing resource/task dependencies and analyzing a kernel's schedule.
Improved compiler diagnostics. Register spills and use of local memory can now be reported at compile time with source line numbers. Preliminary support for detecting classes of NVVM synchronization and execution hazards at compile-time when using the Primitives API. Better reporting of compiler errors that previously did not include source line numbers.
This release has been tested against the following packages:
_e2m1_to_half_x2 and _e2m1_to_half_x4 by merging mask before prmt.Fixed a failed kernel compilation issue when setmaxnreg is used together with specific warp-specialized patterns ( !3382 )
cutlass.Numeric from Python value (!3443)This release has been tested against the following packages:
Reverted the TMA bulk copy elect_one change from 4.6.0 to reset the behavior to align with 4.5.x releases
This release has been tested against the following packages:
Fixed a compilation failure on Thor with 12.9 wheel
Long-deprecated API clean-up, including:
New features
preferred_smem_carveout to set manually.smem now defaults to None for auto-calculating kernel shared memory usage, which is recommended unless manual control is required.smem_merge_branch_allocs is provided to merge shared memory allocations across mutually exclusive code branches, which is recommended for inlined mega-kernels to reduce total footprint.Bug fixing and improvements
elect_one could be found in this PRpip install nvidia-cutlass-operators to get startedCollectiveMma and CollectiveBuilder specializations for MainloopSm120ArrayTmaWarpSpecialized, enabling ptr-array grouped GEMM (MoE expert dispatch) with tensor- and token-level FP8 scaling.DescriptorIterator::operator+ in mma_traits_sm100.hpp to use 32-bit arithmetic on CUDA toolkit version <= 13.3, preserving the high half of the smem descriptor.Your coding agent can read these notes before it upgrades. Set up the MCP server →