NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2341 most downloaded on PyPI
A tile level programming language to generate high performance code.
Last release 5 days ago
30 Sep 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 22 of 22 stable releases
1 version withdrawn
withdrawn after publishing
2 years old
24 releases · first in 2025
One column per month.
T.symbolic remains available as a deprecated alias; use T.dynamic for new code ( #3216 ).
T.gemm_blockscaled semantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support.enumerate, zip, comprehensions, and generator expressions.tilelang.ascend.language and target="ascend" for Huawei Ascend 950 (dav-3510).T.SimdVF and T.SimtVF regions.tvm_ffi and Cython execution, PyTorch NPU tensors and streams, and NPU profiling.See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.
TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples (#3059, #3185).T.fma and T.fmul, and complete additional FP16/BF16 math bridges (#3134, #3132, #3163).for ... in range(...) (#3230).else clauses (#3141, #3232, #3262, #3142).T.copy / T.async_copy coalesced_width as a hint and clamp it to the achievable vector width (#3246).uint64 argument mappings (#3207, #3229).sync and group parameters from T.Pipelined, and k_pack from T.gemm_sp (#3202).T.gemm(k_pack=...) should import tilelang.rocm.language (#3203).threads= explicitly in T.Kernel (#3186).T.symbolic remains available as a deprecated alias; use T.dynamic for new code (#3216).TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically (#3143, #3177).Full Changelog: v0.1.14...v0.1.15
AutoSchedule to AutoWarpSpecialization by @Yongqi-Zhuo in #3185Note truncated.
Reducer v2 ( #2940 , #3093 , #3043 , #3044 , #3079 , #3100 ): T.alloc_reducer reworked into first-class deferred reduction epochs whose physical lower
T.alloc_reducer reworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates.T.unroll(explicit=True) for early explicit unrolling (#2859)cluster_mask on T.tma_copy (#2932)T.gemm tile dimensions with a clear message (#3113), validate T.gemm k_pack arguments (#3094), reject non-positive arrive_count in alloc_barrier/alloc_cluster_barrier (#3112), reject T.Parallel indexing of local buffers (#3041), reject break in fully expanded loops (#3078)tcgen05.alloc arenas (#2831); support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st (#2880); fix ld/st segment pointer advancement in b32 columns (#2952)int4x2/uint4x2 codegen (#3036); 16-bit CUTLASS type overloads for fast-math, __ldg, and htan intrinsics (#3097, #3077, #3028, #2894)T.{reads,writes} for T.tma_{gather4,scatter4} (#3053); separate TMA atomic-add dtype support from layout encoding (#2846)tl.sync_warp on HIP (#2872), reject sub-wavefront block sizes instead of crashing (#2918), resolve versioned device properties in the HIP stub (#2919); emit #line directives for the HIP target (#3058)T.Parallel loops (#3121); scalarize Select in automatic vectorization (#3060)VerifyBufferInit, a general buffer-initialization check (#2956)#line directives from TIR spans (#3048)get_parent_locals frame self-reference leak (#2934)project() (#3102)tcgen05.alloc arenas by @Rachmanino in #2831T.unroll(explicit=True) for early explicit unrolling by @Yongqi-Zhuo in #2859Note truncated.
Breaking changes : this release removes several legacy APIs and packages. See Backend, API & Refactors before upgrading.
This release contains 138 commits (79 bug fixes plus features, refactors, and examples) accumulated since v0.1.12 (2026-07-08 → 2026-08-02).
The headline work is a multi-backend language-dialect refactor that replaces the runtime-activated language facade with static per-backend re-exports, alongside two major new hardware paths: SM120 NVF4 block-scale MMA for Blackwell and Metal 4 (M5) cooperative-tensor GEMM. On top of that, a large batch of correctness fixes landed across reductions/scans, atomics, TMA/copy lowering, loop-step preservation, and FP encoding edge cases.
Breaking changes: this release removes several legacy APIs and packages. See Backend, API & Refactors before upgrading.
T.mma_gemm_blockscaled now routes the packed-scale SM120 path internally through package-pingpong lowering, with an optimized non-persistent example (8192³ measured at ~1527 TFLOPS on SM120). The public micro_pipeline strategy knob was removed from the API.T.gemm (#2252) — TileLang-owned cooperative-tensor intrinsics, Metal 4 MPP matmul2d shader emission, and a shape-aware instruction selector that keeps the simdgroup fallback for fragment accumulators and unsupported tiles.from tilelang.cuda.language import * re-export; CUDA/Metal/ROCm dialects now build on tilelang.language.common with per-backend TIR overlays (details below).T.gemm now works on older architectures instead of erroring out.__hfma (#2769).sm_100a (#2691).fp32x2 ops usable as reducers (#2637).pass_profile pass-config option (with a configurable threshold) (#2622).lower-trace support for debugging (new doc: docs/tools/lower_trace.md) (#2725).T.assume conditions are now enforced at runtime (#2655).CanProve (#2772).The runtime-activated language facade has been replaced by a static re-export architecture:
.pyi stubs + generator, py.typed, the globals()-based __all__ scraping, and _activate_cuda_facade().cuda / metal / rocm) now build on tilelang.language.common with per-backend TIR overlays.mma/wgmma/mfma macro generators are pinned to the dtypes leaf and no longer touch the half-initialized facade during bootstrap.tilelang/cuda/intrinsics/sparse_layout.py (a dtypes-only leaf).Follow-up fixes: ROCm intrinsic resolution (#2779), rng_init (#2776), and shared-intrinsic resolution across backends.
tilelang.common package removed (#2810).tilelang package (#2761).apache-tvm-ffi 0.1.12, while keeping 0.1.11 compatibility (#2795); lower bound raised to >=0.1.11 (#2736).ptxas register-usage level is cast to int before building the nvcc command (#2641).Int64Promoter extracted into a common header (#2558).actions/setup-python 6 → 7 (#2773); transformers bumped in examples/bitnet-1.58b (#2658).For nodes (#2752).LoopUnswitching (#2585).warp_reduce no longer truncates int64/uint64 to 32 bits on sm_80+ (#2782).reduce_sum over-counts on straddle layouts fixed (#2424).nan_propagate honored in reduce max/min/absmax clear=False write-back (#2788).T.atomic_max/T.atomic_min no longer silently corrupt fp32 values (#2780).return_prev supported for scalar atomic_min/atomic_max (#2672), atomic_addx2 with BufferRegion destinations (#2753), and HIP vector atomic add (#2712).T.atomic_addx4 return type guarded for sliced destinations (#2590).T.infinity supported for float8_e5m2 (#2671).T.pow/T.power fixed for constant integer exponent y <= 0 (#2677).T.__exp computes e**x, not 2**x (docstring + CuTeDSL codegen) (#2696).x2 operand dtypes rejected (#2802); floating-point predicates rejected in vote intrinsics (#2800); alloc_var initializer dtype preserved (#2801); invalid dtypes rejected in T.dp4a (#2652).T.copy path casts to the destination dtype (#2771).== +1 (#2649); vectorized Select constraint handling fixed (#052e6741).st.bulk destination emitted as a shared write to fix a missing barrier and compilation-introduced races (#2700).increase_descriptor_offset guard (#2675).T.transpose swaps only the final two axes (#2757); contracting shared-buffer layouts rejected in T.annotate_layout (#2719); unused fragment buffers allowed without layouts (#2717); shared-TMEM buffer pointer types checked before dereference (#2794).vec_type in common.h for CPU codegen (#2768).Bind modeling fixed in parallel race checks (#2665).T.serial (#2674).T.gemm rejected instead of silently producing wrong results (#2724).DataType args no longer break compilation on ROCm (#2726).T.Kernel (#2653).fragment spelling corrected (#2695); BufferStore cast warning context improved (#2733).pass_configs supported in autotuning (#2496).None (#2657).topk_selector memory-access optimization with thread coarsening — ~1.9× faster with identical results (#2659).T.mma_gemm_blockscaled (#2364).e0f0ac90 [CUDA] Refactor TMA atomic add layout validation
6b81bb87 [BugFix] Correctly preserve loop step when unrolling loops (#2835)
09526a27 Revert "[BugFix] Preserve loop step when unrolling loops" (#2834)
56a0f729 [BugFix] Reject unsupported TMA atomic add dtypes (#2830)
bdb769ae [Enhancement] Aggregate VerifyParallelLoop race diagnostics with span (#2806)
e01c498b [JIT] Remove legacy DLPack execution backend (#2816)
2bb0def9 [BugFix] Fix two-instance modeling in ThreadSync cross-thread race checks (#2805)
8f34abf4 [CUDA] Add SM120 NVF4 block-scale MMA support (#2364)
3e4a0544 [Carver] Remove unused shape inference module (#2813)
e18d9699 [CUDA][Reduce] Simplify scalar AllReduce thread range analysis (#2814)
50481cce [Refactor] Remove intrinsic compatibility facade (#2812)
21e8c064 [BugFix] Reject unsupported fast-math input dtypes (#2804)
0bc1913d [BugFix] Resolve partial scalar reduce barrier participation (#2777)
5b1f3218 [BugFix] Reject uncovered warp partitions in T.gemm instead of producing silently wrong results (#2724)
7fd95363 [CUDA] Extend the GEMM FMA fallback to SM75 (#2811)
dd92b781 [Refactor] Remove unused tilelang.common package (#2810)
6c3dd971 [BugFix] Skip descriptor TMA for device-bound copy bases (#2803)
e41fadbe [BugFix] Fix warp_reduce truncating int64/uint64 to 32 bits on sm_80+ (#2782)
2a06036f [BugFix] Reject mixed packed x2 operand dtypes (#2802)
32e02e6c [Fix] Refine architecture guards (#2790)
bc9515fe [BugFix] Prevent autotuner cache reuse across different outputs and validation settings (#2793)
2c84f4f9 [BugFix] Check shared-TMEM buffer pointer types before dereference (#2794)
28f70338 [Metal] Add line-level threadgroup qualifier scanning (pass 5) (#2796)
51f88a88 [BugFix] Reject floating-point predicates in vote intrinsics (#2800)
b1b605da [BugFix] Preserve alloc_var initializer dtype (#2801)
1cb4d4f3 [CUDA] Support arbitrary TMEM layouts (#2785)
4086ba8e [FFI] Support apache-tvm-ffi 0.1.12 (#2795)
7ec5adbe [BugFix] Pass buffer row stride to 2D scan kernel to fix silent miscomputation (#2620)
6171343c [TIR][Language] Add typed vector lane extraction API (#2789)
1545f006 [Metal] M5 Cooperative Tensor T.gemm (#2252)
aaf68d2e [Metal] Add 16-byte alignment padding to shared/threadgroup memory (#2786)
940b1061 [BugFix] Honor nan_propagate in reduce max/min/absmax clear=False write-back (#2788)
a42bbc3c [TIR] Inject source spans into tirx IR and surface source locations in compiler errors (#2751)
500c3686 [CUDA][Transform] Fix PCWS index dtype handling (#2783)
eb31994a [BugFix] Preserve loop step when unrolling loops (#2784)
51fbfc7e [BugFix] Fallback non-16B cluster bulk copies (#2683)
b5e3eb93 [Enhancement] More compile-time guards for architecture-specific CUDA intrinsics (#2781)
2a17fffd [TileOP] Add SM70 GEMM FMA fallback (#2339)
aa7df867 [BugFix] Fix fp16/bf16 T.atomic_max/atomic_min silently corrupting fp32 values (#2780)
9fb75728 [BugFix] Fix ROCm intrinsic resolution after language dialect refactor (#2779)
92072ab2 [BugFix] Cast to the destination dtype in the scalar T.copy path (#2771)
a69708c1 [BugFix] Reject non-power-of-two AllReduce widths (#2611)
30aac1ac [CUDA][Reduce] Fix packed AllReduce workspace pointer (#2778)
fc517bdc [BugFix] Fix rng_init after language dialect refactor (#2776)
eceb0e66 [BugFix] Marshal NVRTC scalar parameters and dynamic strides (#2756)
5ef1500e [BugFix] Handle strided global buffers correctly in 1D TMA copies (#2746)
1dc86d71 [BugFix] Add arithmetic operators to vec_type in common.h for CPU codegen (#2768)
22a2452a [BugFix] Add threadgroup address space qualifier in Metal codegen for shared memory pointer arithmetic (#2770)
9609d3a5 [BugFix] Preserve explicit loop steps when transforms rebuild For nodes (#2752)
b049f87d [CI]: Bump actions/setup-python from 6 to 7 (#2773)
7cb4b1d9 [Enhancement] Fix nondeterministic CanProve (#2772)
c6294f07 [BugFix] Add pre-SM80 fallback for bf16 __hfma (#2769)
8ad82fa0 [Quality] Fixes typings in ast frontend (#2520)
8bb3300d [BugFix] Preserve if condition evaluation during fan-out (#2764)
e9240d68 [Refactor] Move example-only helpers out of tilelang package (#2761)
1591d368 [BugFix] Preserve re-evaluation of mutable if conditions (#2744)
f862dc38 [Language][Backend] Language dialect for multi-backends (#2734)
ab1d2df4 [BugFix] Support BufferRegion destinations in atomic_addx2 return_prev (#2753)
2b4dd803 [BugFix] Make T.transpose swap only the final two axes (#2757)
235077cb [BugFix] Preserve loop steps during unswitching (#2741)
390d208d [BugFix] Use callee global symbols for cross-target calls (#2740)
192ddea6 [BugFix] Gate TMEM and TMA builtins by CUDA architecture (#2743)
aae97c0e [TIR][Python] Add typing wrappers for DSL ops (#2739)
0c88682f [BugFix] Emit Metal barriers for dynamic shared memory (#2738)
bff1b9a3 [Feature] Support stochastic FP32 to FP16/BF16 casts (#2735)
22baf2e2 [FEATURE] Add block-causal attention for dLLM example (#2499)
a443dde9 [BugFix] Reject contracting shared-buffer layouts in T.annotate_layout (#2719)
dff136d4 [CUDA][Pipeline] Fix 1D TMA selection for versioned layouts (#2737)
f84825db [Build] Raise apache-tvm-ffi lower bound to 0.1.11 (#2736)
512d51f5 [BugFix] Honor explicit row strides in Metal GEMM (#2730)
322a9cbb [BugFix] Isolate cross-compiler options per invocation (#2728)
e4e110e5 [Example][DeepSeek-V3.2] Adaptive threads for sparse MLA backward (#2592)
24a023c6 [BugFix][Transform] Never classify side-effecting binds as replayable (atomics re-executed at every use site since v0.1.11) (#2651)
1ea7530f [Cherry][TIRx] Improve BufferStore cast warning context (#2733)
052e6741 [TIR][Analyzer] Fix vectorized Select constraint handling (#2731)
9d819c3f [BugFix] Fix MFMA DataType args causing compilation failure on ROCm (#2726)
cc106fa2 [Feature] Add lower-trace support for debugging & rebased (#2725)
96900c7d [Autotune] Support early_stop in decorator mode and add decorator example (#2729)
bd5ca2f0 [BugFix] Preserve buffer element offsets in access pointers (#2727)
923c8a7d [Autotune] Add early stop to skip slow configs during benchmark (#2723)
f8d8cd4b [BugFix] Use blockDim as workspace stride in batch AllReduce (#2621)
ac576c63 [BugFix] Fix FP4 dequant symbolic exponent clamp (#2656)
25c0a155 [Transform] Replace CPU fallback thread placeholder with a constant-zero logical thread index (#2718)
c4c5ec59 [Example][Opt] deepseek_v32 topk_selector kernel memory access optimization (~1.9× faster) (#2659)
8cfc90e0 [BugFix] Fix HIP predicated dword copy zero fill (#2721)
88e007d4 [BugFix] Guard T.atomic_addx4 return type for sliced destinations (#2590)
656c287a [BugFix] Fix wrong offset when scanning a non-zero-offset buffer sub-region (#2680)
ce9ff0c1 [BugFix] Gate stochastic FP4/FP8 casts on sm_100a (#2691)
39c5b4ef [BugFix] Correct IEEE math intrinsic names for fp64/fp16/bf16 (#2619)
cb26539a [BugFix] Fix T.pow/T.power for constant integer exponent y <= 0 (#2677)
ffeda9f3 [BugFix] Pack 32-lane 8-bit CUDA vectors correctly (#2701)
30221e20 [Bugfix] Emit st.bulk destination as a shared write to fix missing barrier and potential compilation-introduced races (#2700)
172f6fbf [TIR][Transform] Allow unused fragment buffers without layouts (#2717)
8cdd4d62 [BugFix] Prevent partial 1-D TMA stores from bypassing bounds checks (#2716)
46f3b31a [BugFix][WS] Fix pipeline replacement under persistent T.serial (#2674)
4433981c [Enhancement] Add local buffer reduction lowering (#2693)
335afcf8 [BugFix] Decode FP8 E4M3 special encodings correctly (#2710)
e3c3048f [BugFix] Correct the spelling of fragment (#2695)
134f9c2e [BugFix] Implement atomic load and stNote truncated.
TileLang v0.1.11 → v0.1.12 Changes
Summary of the main changes between v0.1.11 and v0.1.12 (91 commits).
pass_visualizer structure-tree pass browser (#2449) and pass-diff display for debugging (#2375)st.bulk shared-memory zero fill on SM100+ (#2403), stmatrix m16n8 on Blackwell (#2417), SM75 MMA dispatchers for FP16 accumulation and UINT8 (#2392)-ccbin support for choosing the C++ compiler (#2348), TILELANG_VERBOSE env var to control compile output (#2453)T.Persistent dropping tiles when the last dim isn't a multiple of group_size (#2455), sign-extension bugs in packed uint32 decode (#2500) and make_int negative int8 lanes (#2438), PTX v4 atomics for fp16/bf16 atomic_addx4 (#2492), vectorized atomic_add dtype mismatch (#2414), bf16 exp self-recursion (#2402) and rsqrt overload (#2386)do_bench gained a cache_size option (#2531)T.view / T.reshape enhancements (#2450), better T.assume/loop-bound handling to eliminate redundant boundary checks (#2502), improved diagnostics for T.serial fragment access (#2462)This release centers on the new LLVM backend and backend-registry refactor, major TMA/GMMA layout flexibility on CUDA, fp8/fp16/bf16 cast performance, and a large batch of correctness fixes across layout inference, atomics, and pipelining.
-ccbin arguments to specify C++ compiler by @Triang-jyed-driung in #2348st.bulk for shared zero fill on SM100+ by @Rachmanino in #2403__shfl_sync from tl_shuffle_elect by @Yongqi-Zhuo in #2445T.view and T.reshape by @bucket-xv in #2450TILELANG_VERBOSE environment var to control the compile output info by @bucket-xv in #2453Note truncated.
Fix atomic_load access_ptr lowering for dynamic indices by @VitalyAnkh in #2157
T.tma_copy for flexible TMA copy by @Rachmanino in #2205Full Changelog: v0.1.10...v0.1.11
Your coding agent can read these notes before it upgrades. Set up the MCP server →