NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1958 most downloaded on PyPI
NVIDIA CUTLASS Python DSL
Last release 13 days ago
21 Sep 2026
Ships on a steady schedule
a new release about every 3 weeks
Nearly every release is documented
notes for 24 of 24 stable releases
1 version withdrawn
withdrawn after publishing
2 years old
33 releases · first in 2025
One column per month.
Existing GroupedGemmArguments is deprecated and will be removed in a future release.
New features
Initial Rubin support to accelerate dense GEMMs. The following features are available:
CuTe DSL extensions has several new features:
cute_ext TMA load, store, multicast, and reduce-store operations. Explicit CTA-V maps remain supported as overrides.cute_ext GEMM mainloop and TMA epilogue helpers.This release includes an opt-in preview of the CuTe DSL extensions (cute_ext) compiler pipeline for ordinary Cute DSL kernels. This pipeline lets user mix cute_ext APIs directly into @cute.jit and @cute.kernel code and is required for kernels that mix the two API surfaces. You may test this feature with the following:
CUTE_DSL_USE_EXTENSION_COMPILER=1 python your_program.py
The pipeline is expected to preserve program behavior and performance, but generated PTX/SASS may differ. Note that this pipeline will become the default in the future, no earlier than 4.10.
Added examples for better control over Primitives' compiler warnings/errors introduced in 4.7.0. See the CuTeDSL/experimental/compiler_diagnostic/ directory.
IKET Profiler Tool
iket_enable_profiling=True). Task execution will generate an IKET range and individual pipeline stages in a schedule may generate separate ranges (iket_profiling_stages).A number of new examples were added in this release:
CuTe DSL now supports x86_64 Windows
CuTe DSL AoT now supports new host target: QNX8.0
Notebooks are restructured under examples/python/CuTeDSL/cute/notebooks and new notebooks for primitives will be added under examples/python/CuTeDSL/notebooks
Numpy is now not a default dependency
Bug fixes and improvements:
nvidia-cuda-nvdisasm is now an optional dependency of nvidia-cutlass-dsl via the optional [sass] extra. SASS dumping (CUTE_DSL_KEEP=sass / KeepSASS) now resolves nvdisasm from the bundled wheel (recommended since its version matches the DSL toolchain) or from a local CUDA Toolkit (CUDA_HOME/CUDA_PATH). A locally-provided nvdisasm must come from a CUDA Toolkit at least as new as the toolchain that produced the CUBIN. Installations that never dump SASS are unaffected.cutlass.jax.cutlass_callcute.autovec_copy emitted per-element instead oflink-libraries compile-option order so it is stable across processesIndexError on staged bool() with no argumentscute.compile on @cute.kernel with a user error instead of an ICEThis release has been tested against the following packages:
Dense and blockscaled GEMMs in Operator API have preliminary Rubin support. These are provided as a preview and may need additional performance tuning.
Updated GEMMs include:
These kernels utilize the below new features in Rubin:
Operators can now be ranked by their estimated performance when nvMatmulHeuristics is available. See tutorial here. NOTE: This currently only supports Blackwell kernels as nvMatmulHeuristics does not yet support Rubin.
Standalone kernel implementations are now exposed through cutlass.kernels. These kernels can be used directly, in addition to being discoverable and usable via the Operator interface in cutlass.operators.
Custom Epilogue fusions now support per-row or per-column reductions.
IndexPtrGroupedGemmArguments is now used to represent Grouped GEMM with contiguous-offset/index-pointers. Existing GroupedGemmArguments is deprecated and will be removed in a future release.
sm_107a and sm_107f targets:NOTE: Executing Rubin kernels (SM107) requires the R615 driver which will be released
with CUDA Toolkit 13.4 GA. R610 from CUDA Toolkit 13.4 Developer Preview is not
sufficient.
Nothing published for this version
Fixed a failed kernel compilation issue when setmaxnreg is used together with specific warp-specialized patterns ( !3420 )
cutlass.Numeric from Python value (!3443)This release has been tested against the following packages:
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a sta
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development velocity. Primitives are released as experimental and will evolve based on user feedback.
NOTE: Primitives is a transitional API until a CUDA Python-like solution is available.
Introduced the Task Scheduling framework. This provides static analysis of execution schedules for warp-specialized kernels. Compilation stops when known concurrency issues are detected. Also provides tools for visualizing resource/task dependencies and analyzing a kernel's schedule.
Improved compiler diagnostics. Register spills and use of local memory can now be reported at compile time with source line numbers. Preliminary support for detecting classes of NVVM synchronization and execution hazards at compile-time when using the Primitives API. Better reporting of compiler errors that previously did not include source line numbers.
This release has been tested against the following packages:
_e2m1_to_half_x2 and _e2m1_to_half_x4 by merging mask before prmt.Fixed a failed kernel compilation issue when setmaxnreg is used together with specific warp-specialized patterns ( !3382 )
cutlass.Numeric from Python value (!3443)This release has been tested against the following packages:
Reverted the TMA bulk copy elect_one change from 4.6.0 to reset the behavior to align with 4.5.x releases
This release has been tested against the following packages:
Fixed a compilation failure on Thor with 12.9 wheel
Long-deprecated API clean-up, including:
New features
preferred_smem_carveout to set manually.smem now defaults to None for auto-calculating kernel shared memory usage, which is recommended unless manual control is required.smem_merge_branch_allocs is provided to merge shared memory allocations across mutually exclusive code branches, which is recommended for inlined mega-kernels to reduce total footprint.Bug fixing and improvements
elect_one could be found in this PRpip install nvidia-cutlass-operators to get startedCollectiveMma and CollectiveBuilder specializations for MainloopSm120ArrayTmaWarpSpecialized, enabling ptr-array grouped GEMM (MoE expert dispatch) with tensor- and token-level FP8 scaling.DescriptorIterator::operator+ in mma_traits_sm100.hpp to use 32-bit arithmetic on CUDA toolkit version <= 13.3, preserving the high half of the smem descriptor.Nothing published for this version
Fixed a compilation time regression issue in 4.5.0. Compilation times now match those in the 4.4 and 4.6 branches.
Python 3.14t is now supported with GIL enabled
New features
Bug fixing and improvements
Fixed following issues: https://github.com/NVIDIA/cutlass/issues/3219 https://github.com/NVIDIA/cutlass/issues/3218 https://github.com/NVIDIA/cutlass/
cvt.rn.bf16x2.e4m3x2 conversion instruction support to numeric_conversion.h.ab_dtype is deprecated in make_trivial_tiled_mma and make_blockscaled_trivial_tiled_mma from blackwell_helpers.py. Please specify a_dtype and b_dtype…
New features
block_copy() to simplify TMA and S2T copy. Users can ignore detail about multicast and 2CTA partition for TMA by block_copy() and need not to invoke tma_partition(). And users can remove bulk of S2T initialization to simplify S2T copy.C.remap_modes[:, 0, 1] subscript syntax (where : marks a broadcast dimension and integers select source mode indices). Covers scalar broadcast, row/column broadcast, and arbitrary mode permutations (e.g. transpose). The PyTorch reference evaluator mirrors the same transformations.Bug fixing and improvements
More examples of authorizing peak-performance kernels
API changes
ConstSubbyteReference__nv_atomic_load_n with volatile for CUDA 11.4 compatibility in subbyte referencePipelineStorage shadowing in SM100 complex epilogueNothing published for this version
CuTe DSL now supports Python 3.14 for both x86_64 and aarch64
Fixed a segfault issue with tvm-ffi on aarch64
Deprecate get_num_tmem_alloc_cols from blackwell_helpers.py. Use the one from tmem_allocator.py instead.
New features
More examples of authorizing peak-performance kernels
mixed_input_gemm. Common utility functions are also extracted into mixed_input_host_utils.py under the same folder.Bug fixing and improvements
cute.printf with f-stringAPI changes
Use 'Advanced control file' for mixed input gemm examples for better performance.
compute_memory_reordering_atom<tfloat32_t>()TmaGbasis in AuxTmaParams.
tma_gbasis.TmaGbasis parameter of AuxTmaParams and users are allowed to manually construct a dynamic gbasis.media/docs.Nothing published for this version
Nothing published for this version
Fixed the unexpected CPU overhead issue introduced by 4.3.4
Added PDL support along with example Kernel launch with Programmatic Dependent Launch
New features
Bug fixing and improvements
make_smem_layout_a in utils/hopper_helpers.pySupported namedtuple and kwargs for JIT function arguments in tvm-ffi
New features
Bug fixing and improvements
New env var CUTE_DSL_CACHE_DIR to specify the path for dumping caches
New features
CUTE_DSL_CACHE_DIR to specify the path for dumping cachesBug fixing and improvements
Multiple dependent DSOs in the wheel have been merged into one single DSO
Supported Apache TVM-FFI for further reduced host runtime overhead for JIT functions, better PyTorch and ML frameworks interopability
nsight profiling to correlate perf metrics with Python source code)PipelineProducer and PipelineConsumer to simplify code without explicit pipeline state management (Exiting APIs are still maintained)Baseline + XTensorSSA.reduce to support static value as initial valuemake_layout_tvis_staticPipelineAsyncSmemAllocatorpipeline, utils and cute.mathbatch, no_verif, cluster_shape and cluster_shape_fallback in example 89.moe_stride_utils is introduced to help setup strides in the kernel.problem_shapes_device and problem_shapes_hosts, a new problem shape struct called MoEProblemShape is introduced which takes in max_m, max_n, max_k and counts vector as input and deduce problem shapes internally whenever required.cutlass::int8_t and replace it with int8_t.wait_on_dependent_grids for PDL use case.bytes_with_problem_shape of block scaled profiler.Nothing published for this version
Fixed an issue when running DSL codes with cuda-python 13.0
More Python versions are now supported for both x86-64 and aarch64, including
cute.print_tensor for coordinate tensorcute.print for tuple of layoutssm103_ under GEMM device unit tests.nvidia-matmul-heuristics to find the best kernels for a given scenario.
get_unmasked_trip_count may return a negative value.CUTLASS_LIBRARY_INSTANTIATION_LEVEL to instantiate all possible combinations.CUTLASS_LIBRARY_KERNELS must be non-empty. Profiler will combine CUTLASS_LIBRARY_KERNELS and CUTLASS_LIBRARY_INSTANTIATION_LEVEL to instantiate specific kernels.cutlass to cutlass_cppgen and add Blackwell EVT support to legacy Python interface.
EpilogueDescriptors.nullspace implementation.cosize hacks.E<0,1> == 1@0@1.Add aarch64 support, you can now pip install nvidia-cutlass-dsl on GB200 systems!
CuTe DSL
nvidia-cutlass-dsl on GB200 systems!CUTLASS C++
subbyte_iterator with cute::recast_ptr when constructing logical iterators/arrays.get_layoutA|B|C_MN and friends from Atoms/TiledX.print_latex and friends and rewrite.print_svg and friends and rewrite.Nothing published for this version
CuTe DSL is a Python DSL centered around CuTe's abstractions
CuTe DSL
CuTe DSL is a Python DSL centered around CuTe's abstractions
CUTLASS C++
(old) cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_bf16_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma(new) cutlass3x_sm90_tensorop_gemm_bf16_bf16_f32_bf16_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma-DCUTLASS_LIBRARY_KERNELS, filter kernels in the CUTLASS profiler with --kernels), please update your uses accordingly, this is a breaking change.mma_promotion_interval has been removed from non-grouped GEMM to align with the grouped and Blackwell SM100 versions.fmha_gen sample only supports head dim 128.cute::copy_if so that the predicate tensor is also a true CuTe Tensor rather than a lambda and introduces transform-tensors to avoid any extra register or load/store overhead in using bool-tensors.Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →