NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #462 most downloaded on PyPI
A language and compiler for custom Deep Learning operations
Last release 1 months ago
28 Aug 2026
Ships on a steady schedule
a new release about every 2 months
Some releases are documented
notes for 7 of 22 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
22 releases · first in 2021
One column per quarter.
Aggregate types: @triton.aggregate and @gluon.aggregate are now public APIs. Aggregates support inherited fields, default values, generated constructo
@triton.aggregate and @gluon.aggregate are now public APIs. Aggregates support inherited fields, default values, generated constructors, immutable instances, and aggregate_replace() (#10095, #9572)tl.topk: Added a descending argument. Set descending=False to return the smallest values (#9355)tl.dot_scaled (#10311)tl.fdiv(..., ieee_rounding=True) now emits IEEE-rounded division, and floating-point atomic_min now returns a value-typed result (#10074, #10485)argmin, argmax, minimum, maximum, and clamp now match compiled behavior more closely when inputs contain NaNs (#10298, #10333, #10699)tma.store_wait now accepts a read_only argument. The default remains True; use read_only=False when the store must reach global memory before a release operation (#10415, #10419)RemoveLayoutConversions (#10646, #11029)tl.expect_zero (#9337, #9455, #9714, #10112, #10330, #10461)cdna5 as an alias for the gfx1250 target (#11383)gfx906 (#9628)gfx120x) targets (#10390, #10185)num_threads, num_warps, and smid to tl.extra.hip, and added clz and popc to the HIP libdevice (#9604, #10651)add4, sub4, mul4, and fma4 (#11084)GenericLinearEncodingAttr for supported swizzled and non-injective layouts, including local load/store lowering and permutation-matrix inference (#9765, #10122, #10515)GluonJITFunction (#9863)async_load and async_store names; the previous async_copy_* names remain available as aliases (#10083)triton_kernels adds an FP64 matmul path (#9393, #10634)Layout.storage_shape() for querying a layout's physical storage shape without materializing a conversion (#10554)do_bench_proton and do_bench_cudagraph_proton to reduce benchmark bias from CPU launch overhead and an unflushed L2 cache (#10149)roctracer to rocprofiler-sdk (#9704)tcgen05 code generation. CI now runs the full test suite (#10344, #10345, #10027)setup.py to CMake, added progress and resume support, and avoided repeated CUDA tool downloads (#9458, #9722, #10418, #10540)kernel_unload_hook and treated incomplete cache entries as misses (#9444, #9542, #10411)map_elementwise, tl.expect_zero, out-of-range float-to-integer casts, control-flow scoping, load semantics, and compiler-hint semantics (#10695, #10692, #10678, #10119, #10356)!tt.tensordesc<tensor<...>> must use the new !tt.tensordesc<..., #layout> form (#9851, #9984)tt.make_tensor_ptr and tt.advance IR operations were removed. Out-of-tree MLIR using those operations must migrate to regular pointer operations or tensor descriptors (#9668)buffer.load() should omit that argument; use buffer.get_reg_layout() when the layout is needed separately (#9594)float2 module was replaced by native packed operations such as add2, sub2, mul2, and fma2. Call these operations directly from the Blackwell module instead of importing blackwell.float2 (#11002, #11084)triton.tools.triton_to_gluon_translater to triton.tools.triton_to_gluon_translator. Update imports to use the corrected spelling (#9570)tl.dot output type: When a non-FP32 accumulator is supplied and out_dtype is omitted, the output now defaults to the accumulator dtype instead of tl.float32. Set out_dtype=tl.float32 explicitly to preserve the previous behavior (#10353)TRITON_PREFER_TMEM_16x256_LAYOUT environment variable (#10664)This release includes contributions from engineers at:
Special thanks to all contributors who submitted bug reports, feature requests, and code improvements!
Triton 3.7.1 is a patch release on top of 3.7.0. It fixes the following 2 regressions and contains no new features or API changes.
Triton 3.7.1 is a patch release on top of 3.7.0. It fixes the following 2 regressions and contains no new features or API changes.
Deprecation Warning for `make_block_ptr`: Emitted a deprecation warning when make_block_ptr is used
tl.squeeze / tl.unsqueeze: Added tl.squeeze and tl.unsqueeze operations to the standard library (#8924)constexpr values from JIT-compiled code (#8785)get_int_attr for Out-of-Tree Walk: Added get_int_attr to Operation to support out-of-tree IR walks (#8892)preload: Added optional device argument to preload and guardrails for cross-target preload (#8951, #8952, #9234)tl.cat(can_reorder=False): Added a non-reordering variant of tl.cat with broadcast support (#9312, #9163)desc.shape for FP4 Padded: Fixed desc.shape values for fp4-padded tensor descriptors (#9012)constexpr_functions (#8876)must_use_result for Methods: Fixed must_use_result check for methods (#8902)_semantic Default to None: Defaulted _semantic parameter to None (#8909)make_tensor_descriptor Error Typo: Fixed typo in make_tensor_descriptor error message (#8912)tl.cat Determinism: Made tl.cat deterministic via permute+reshape+join, then reverted (#9312, #8854, #8878)make_block_ptr: Emitted a deprecation warning when make_block_ptr is used (#9667)inspect.signature for builtins, lazily computed tuple type names, avoided find_paths_if and inspect.getclosurevars, removed outdated catch_warnings blocks — all to reduce JIT overhead (#8843, #8844, #8846, #8845, #8881)tcgen05.mma + Multicast: Added multicast support for tcgen05.mma (#9071)tcgen05.mma Verifier & Errors: Throw a clear error instead of miscompiling very large tcgen05.mma along N (#8915)LowerAref, per-partition asyncOp storage, explicit captures to WarpSpecializePartitionsOp, skip InsertTmemAref when WS isn't used (#8978, #9007, #9023, #9133, #9212)RegionBranchInterface: Made WarpSpecializePartitionsOp implement RegionBranchInterface (#8799)aref.get Filtering: aref.get creation now filters results not in the scheduled loop (#9114)tt.scan Layout Fixes: Fixed tt.scan with broadcasted layouts and additional scan layout issues (#9185, #9189)assert or print are no longer pipelined (#9055)wgmma wait(0) to first use of the accumulator (#9021, #9179)ext; improved robustness of ext slice rematerialization (#9194, #9019)AxisInfo for add/sub; reland of unvisited-operand handling (#9297, #8758)async_cp: Pick better layouts for small async_cp (#9183)getConvertBackwardSlice (#8291)memdesc_slice in Membar; extended membar with third-party ops via traits; AMD-aware membarFilter (#8755, #8798, #9265)ReduceOp lowering, later reverted on release branch (#9192, #9214)kReg smem Padding: Separated additive kReg shared-memory padding contribution (#9286)tcgen05.mma + multicast support and Generalized Encodings: continued generalization of TMEM and shared-memory layouts (#9071)SwizzledShared Layout, uniform hint on ttg.warp_id, and CGAEncoding rename (#9286, #9073, #8850, #9040, #9125)dotCanBeProperlyAsync when wgmma is not yielded by the loop and an associated infinite loop (#9274, #9282)FuncOpToLLVM Refactors: Moved handleArgPtrDatatype to Utility.h; support for LLVM struct/array types in DITypeAttr (#9120, #9124)JITFunction in preload: Support JITFunction in preload (#8794)3.7 is heavy on gfx1250 (RDNA4) maturation, warp specialization on AMD, Tensor Data Movement (TDM), and a new warp-pipeline path.
ttg.warp_id and AMD Conversions (#8659)v_permlane16_swap: Enabled for convert_layout and reduceOp on GFX1250 (#8724)AMDWmmaEncodingAttr, scalar-pointer cluster-load avoidance (#9342, #9340, #9129)AMDWMMALayout Rank Consistency (#9127)i8xi8xi32 v3, missing f64.16x16x4.f64, and clamp operand on WMMA int intrinsic (#9267, #9271, #9291, #9359)AsyncWait in UpdateAsyncWaitCount (#9352)dim > 2 (#8994)AsyncCopy by default for gfx950 and gfx1250 — later reverted on release/3.7.x (#9445, #9087)v_perm for convert_layout (#9014)ReorderInstructions: Removed sinkSecondLoad, sinkDotConversion, and moveUpTranspose optimizations (#9119, #9139, #9204, #9229)ReorderInstructions with MoveUpPrologueLoads (#9328)UpdateAsyncWaitCount: Support single-block execute regions (#9126)OptimizeLDSUsage Removal (#8282)finite/isfinited, rint, clampf via v_med3: libdevice and codegen additions (#9097, #9166, #9256)BlockPingpong Improvements: Debug messages and dot-dominates-predecessors fix (#8804, #9027)kWidth mandatory for WMMA v3 (#8783)copysign Replacement: Replaced LLVM copysign intrinsic (#8789)CanonicalizePointers: Support MakeTensorDescOp in CanonicalizePointers (#9228)PartitionedSharedEncodingAttr: Introduced and reverted (#9314, #9367)scf.if Combining: Added PrepareIfCombining pass (#9253)SinkLayoutConversions Pass (#9168)addOccurrence for proper LLVM-option disabling; ScopedNoAliasAAWrapperPass in MIR swap pipeline (#8711, #9311, #9309)atomic_cas Fixes: Wrong struct index for atomic-CAS pattern, ignored sem/scope, and atomic-CAS for non-int types (#8867, #9042, #9116)uniformSum Crash: Fixed null uniformSum in CanonicalizePointers (#8991)RangeAnalysis tripCount: Fixed trip-count calculation (#9383, #9944)BlockPingpong for non-MFMA dot (#9618, #9948)CanonicalizePointers Different Bases (#9541, #9950)tcgen05 MMA on sm110 (Jetson Thor) (#9160)tcgen05.ld.red on sm103: Implemented in Gluon (#9151)NVIDIA::canSkipBarSync: Resurrected (#9246)AsyncTMACopyGlobalToLocalOp, tensor-descriptor support, fix for tma load, and driver support (#9202, #9225, #9303, #9305)tt.split/join in WS Data Partition: Hopper WS support for tt.split/tt.join (#456, #9147)w_scale Mask: Fixed Hopper mask (#8974)get_view(): Added get_view() for Gluon layouts (#9270)PaddedSharedLayouts (#9336)to_linear_layout: Allow TM layouts in to_linear_layout for printing (#8682)SharedLinearEncoding: Continued lowering generalization (carry-over from 3.6 with backend updates).requires_persistent (#9198)Tensor.clone: Briefly added clone for triton_kernels.tensor.Tensor, then reverted (#9178, #9208)mxfp4→bf16 Conversion via mul.bf16x2 (#8967)ex2.approx.ftz for swiglu (#8801, #8905, #9164)distributed.py / bench_utils.py: Extracted common code from bench_mlp.py and distributed.py (#8866)num_stages Adjustment: For bf16/fp16 × mxfp (#8773)reduce_forward Metadata: Improved performance (#9068)p_matmul Asserts & Fixes (#9376)symm_mem_pool by Argument (#9092, #9155)deactivate / get_data Overhead Reduction: Especially for CUDA-graph profiling; exposed get_data_msgpack (#9030)get_data API: Export profile data directly in Python (#8928)clear_data API: Remove pre-deactivation data (#8971)finalize Cleanup: Clean up context source after teardown (#9069)GlobalScratchAllocOp Deprecation: Deprecated Proton's own op in favor of TritonGPU's, with a custom backend (#8976)TRITON_ENABLE_HW_TRACE in CuptiProfiler (#9324)fresh_knobs Default Behavior (#9184)tl.dot BF16xN Nondeterminism (#8818)DOCKER_API_VERSION (release/3.7.x) (#10244)assert_close: Propagate err_msg to numpy (#9170)test_line_info_ir_source Flake Fix (#9161)CMAKE_LIBRARY_OUTPUT_DIRECTORY: Fixed build with empty directory (#8810)llvm_update_compile_flags Removal (#9167)LLVM_BUILD_SHARED_LIBS Canonicalization (#8933)TRITON_EXT_ENABLED for Wheels (#9935, #9959)nvidia-toolchain-version.json Update.TRITON_DEFAULT_BACKEND: Control driver.active via this env var (#9144)TRITON_PTXAS_BLACKWELL_PATH: Allow override of ptxas-blackwell binary (#8945)topk in Plugin Example: Increment index in plugin example (#9315)link.py (#9084)AxisInfo (#9266)topk Operation: Added to language documentation (#9345)warp_specialize Docs: Updated gl.warp_specialize docs (#8553)LinearLayout Output Matrix Comment: Doc fix (#9243)triton_kernels matmul refactor (BC-breaking): The matrix-multiplication refactor introduces a backwards-incompatible API surface. Downstream users of triton_kernels.matmul_* should review call sites (#8765)tcgen05.cp Lowering Generalization & tcgen05.mma Encoding Acceptance: Continued from 3.6, with new verifier behavior and stricter encoding checks.GlobalScratchAllocOp Deprecated: Replaced with TritonGPU's GlobalScratchAllocOp + custom backend. Out-of-tree consumers must migrate (#8976)make_block_ptr Deprecated: A deprecation warning is now emitted; users should migrate to tensor descriptors (#9667)This release includes contributions from engineers at:
Special thanks to all contributors who submitted bug reports, feature requests, and code improvements!
Un-deprecated min/max (#8734): Un-deprecated min/max on scalar tensors
tl.trans and tl.dot operationsdot_scaled operationsdot_scaled Handling (#8658): Fixed missing handling for None acc in dot_scaledtl.cdiv (#8669): Optimized tl.cdiv for common case of 32-bit divisorsast.Numtcgen05.cp Lowering (#8225): Implemented generic lowering for tcgen05.cpatomic_cas operationtt.LoadOptcgen05.ld/st genericallytcgen05.mma to accept SharedLinearEncodingAttrgl.warp_specialize API for better usabilitynum_ctas in Gluonttgl.get_num_warps metafunctiongather and its layout teststl functions into gluon and expose catasync_copy to Gluon for gfx1250x in triton_kernelssplit_k on m * nBitmatrixMetadata and RaggedTensorMetadata; deprecated triton_kernels.routingy_indx and uniform distributionreduction_n=2 to bench_mlp.pysplit_k > 1 with fused scatterHopperValue layoutgl.warp_specialize APIThis release includes contributions from engineers at:
Special thanks to all contributors who submitted bug reports, feature requests, and code improvements!
This release is meant to fix the following issue:
This release is meant to fix the following issue:
Fix sm103 (GB300) support broken by Triton 3.5.0 release (https://github.com/triton-lang/triton/pull/8045)
Warp Specialization Enhancements (#8005): Made warp specialization require at least 4 warps with proper error messaging to prevent compiler crashes
mask parameter to tl.device_assert for easier debugging with masked operationsconstexpr_function to support cache invalidation and capability checkstl.float16 and other FP types@builtin or @core.extern functions modify their argumentscore.extern_elementwiseConstant{Int|Float}Op type and value orderTargetLibraryInfoImplconvert_layout that:
ldmatrix/stmatrix and transpose versionsselect and shuffle instructionsPaddedSharedEncoding with non-default orderlowerLdSt pathassert_trivial flag for performance validationnumel and nbytes properties (#7507)map_elementwise (#7564)tensor.sum (#7617)splat returning auto encoding (#7490)make dev-install-llvm in READMETRITON_DEBUG at import time (#7767)This release includes contributions from engineers at:
Special thanks to all contributors who submitted bug reports, feature requests, and code improvements!
FP8 Format Warnings - Enhanced warnings for deprecated FP8 formats
The Gluon framework has received major enhancements across all areas including new APIs, tensor memory management, layout operations, and synchronization primitives. Key additions include static_assert functionality, TensorDescriptor kernel arguments, async TMA operations, tensor memory implementation, thread synchronization barriers, and comprehensive tensor operations like split/join/reshape and reductions. (#7172, #7168, #7165, #7160, #7152, #7151, #7149, #7145, #7142, #7122, #7121, #7120, #7115, #7114, #7106, #7102, #7099, #7097, #7091, #7089, #7080, #7061, #7057, #7022, #7020, #7009, #7006, #7004, #7001, #6998, #6997, #6994, #6992, #6989, #6985, #6971, #6950)
@tl.aggregate decorator for autogenerating Triton types from Python classes (#6970).item() as syntactic sugar for .reshape([]) (#6873)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →