NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2341 most downloaded on PyPI
A tile level programming language to generate high performance code.
Last release 5 days ago
30 Sep 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 22 of 22 stable releases
1 version withdrawn
withdrawn after publishing
2 years old
24 releases · first in 2025
One column per month.
This release focuses on broader backend support, new GPU instructions, compiler pipeline improvements, and release/build stability.
This release focuses on broader backend support, new GPU instructions, compiler
pipeline improvements, and release/build stability.
.shared::cta in TMA copy paths only on CUDA 12.8+ by @ColmaLiu in #2087T.sync_threads() by @bucket-xv in #2197do_not_specialize for autotune by @Rachmanino in #2084T.tma_copy SIMT fallback by @Rachmanino in #2242Note truncated.
Re-enable deprecated TL_DISABLE_TMA_LOWER pass config for TMA store by @LJC00118 in #2024
InjectFenceProxy by @Rachmanino in #1850ProveFragmentContains to avoid false layout conflicts by @LJC00118 in #1950wgmma_gemm and tcgen05_gemm functions by @LeiWang1999 in #1949gemm_streamk example on SM90 by @Rachmanino in #1969minBlocksPerMultiprocessor in __launch_bounds__ by @Rachmanino in #1979annotations parameter to alloc_buffer in tilelang/language/ast/ir.py by @Copilot in #1996TL_DISABLE_TMA_LOWER pass config for TMA store by @LJC00118 in #2024DecoupleTypeCast Pass by @LJC00118 in #2026Note truncated.
[Refactor] Enhance T.alloc_barrier with new features and deprecate legacy mbarrier related intrinsics by @Rachmanino in #1733
T.make_tensor not on the top of prim_func by @LeiWang1999 in #1412T.__ldg by @LeiWang1999 in #1414compile_flags to ffi compilation path with pass_configs by @LeiWang1999 in #1434pytest.mark.parameterize to speedup parallel testing by @kurisu6912 in #1447T.annotate_restrict_buffers by @LeiWang1999 in #1428test_tilelang_language_rand.py by @silentCoder-dev in #1464kDisableDynamicTailSplit and kDynamicAlignment as they are legacy by @LeiWang1999 in #1486alloc_local statement in examples and introduce processing for floating fragment buffers by @LeiWang1999 in #1495ctypes by @LeiWang1999 in #1510TargetIsCuda for all cuda target by @oraluben in #1522S_q != S_kv by @hukongyi in #1530local.var buffer as local by @LeiWang1999 in #1541T.Fill for local.var by @LeiWang1999 in #1543H in deepseek sparse mla backward via split-H by @Rachmanino in #1548ParallelOPNode and CopyNode by @LeiWang1999 in #1539tl_pipeline_sync. by @c8ef in #1566Note truncated.
[Pipeline] Refactor buffer allocation in Inject Pipeline Pass by @LeiWang1999 in #1525
S_q != S_kv by @hukongyi in #1530local.var buffer as local by @LeiWang1999 in #1541T.Fill for local.var by @LeiWang1999 in #1543H in deepseek sparse mla backward via split-H by @Rachmanino in #1548ParallelOPNode and CopyNode by @LeiWang1999 in #1539tl_pipeline_sync. by @c8ef in #1566test_tilelang_language_cooperative.py by @silentCoder-dev in #1593import tilelang on CPU-only machines without CUDA libraries by @XuehaiPan in #1481T.sync_warp & T.shfl_sync; change extern pdl into intrin by @silentCoder-dev in #1614ForwardRef usage in v2 frontend (#1619) by @kurisu6912 in #1621ConstrVisitor to src/transform/common/constr_visitor.h for reuse by @silentCoder-dev in #1622T.reduce_absmax to use less abs call by @kurisu6912 in #1626nvidia-cuda-nvcc as nvcc by @clouds56 in #1528k_dim==4 and open rocm-ci for gemmsr by @benenzhu in #1627examples/deepseek_v32/sparse_mla_fwd.py by @GoldenStain in #1634cp.reduce.async.bulk.tensor by @Rachmanino in #1667Threadsync with ConstrVisitor by @silentCoder-dev in #1631ParallelLoopTransformer by @LeiWang1999 in #1672Note truncated.
[Pipeline] Refactor buffer allocation in Inject Pipeline Pass by @LeiWang1999 in #1525
S_q != S_kv by @hukongyi in #1530local.var buffer as local by @LeiWang1999 in #1541T.Fill for local.var by @LeiWang1999 in #1543H in deepseek sparse mla backward via split-H by @Rachmanino in #1548ParallelOPNode and CopyNode by @LeiWang1999 in #1539tl_pipeline_sync. by @c8ef in #1566Full Changelog: v0.1.7.post1...0.1.7.post2
[Bugfix][Build] Update CMake configuration to remove project root injection for sys.path by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1
T.make_tensor not on the top of prim_func by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1412T.__ldg by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1414compile_flags to ffi compilation path with pass_configs by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1434pytest.mark.parameterize to speedup parallel testing by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/1447T.annotate_restrict_buffers by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1428test_tilelang_language_rand.py by @silentCoder-dev in https://github.com/tile-ai/tilelang/pull/1464kDisableDynamicTailSplit and kDynamicAlignment as they are legacy by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1486alloc_local statement in examples and introduce processing for floating fragment buffers by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1495ctypes by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1510TargetIsCuda for all cuda target by @oraluben in https://github.com/tile-ai/tilelang/pull/1522Full Changelog: https://github.com/tile-ai/tilelang/compare/v0.1.7...v0.1.7.post1
[Enhancement] Deprecate split&sum in attn bwd examples on Hopper by @Rachmanino in https://github.com/tile-ai/tilelang/pull/1065
seq_q<seq_kv in flash attention examples by @Rachmanino in https://github.com/tile-ai/tilelang/pull/864B[i,j] = c[i] + A[i,j] by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/798ExprDeepEqual instead of StructuralEqual when merge consecutive If stmt by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/876T.ieee_rsqrt and related high precision op by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/882atomic_add performance for bwd examples by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/940bfloat16 and user-defined sm_scale in attention sink examples by @Rachmanino in https://github.com/tile-ai/tilelang/pull/924pre-commit integration by @XuehaiPan in https://github.com/tile-ai/tilelang/pull/955CumSum1D by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/978T.alloc_var for AugAssign and AnnAsign by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/979InjectFenceProxy and expose some warp group primitives in frontend by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/977access_ptr("r") instead of access_ptr("w") for correct pipeline analysis by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/983torch.accelerator.synchronize() to torch.cuda.synchronize() by @yyttt6 in https://github.com/tile-ai/tilelang/pull/987LowerIntrin from tvm into tilelang by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/999LowerIntrin by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1035TL_LIBS by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1038T.get_warp_idx_sync and T.shuffle_elect for efficient thread election by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/989has_simt_copy to decide whether to insert set_max_nreg by @chengyupku in https://github.com/tile-ai/tilelang/pull/982LegalizeSafeMemoryAccess to support recursive load/store rewrite by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/1050T.Parallel with dynamic extents by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/990tileang.clear_cache() by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1077T.dynamic instead of T.symbolic by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1076T.reduce_ with shared memory input/output by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1080tilelang_cython and relocate its path by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1086tilelang.disable_cache() calls from examples and tests by @Rachmanino in https://github.com/tile-ai/tilelang/pull/1088TL_STORAGE_REWRITE_DETECT_INPLACE by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1089alloc_var(dtype, init=x) by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1092cuTensorMapEncodeIm2col call by @chengyupku in https://github.com/tile-ai/tilelang/pull/1094format.sh and add clang-tidy to GHA workflow by @XuehaiPan in https://github.com/tile-ai/tilelang/pull/1044ldmatrix and update mamba scan kernel by @chengyupku in https://github.com/tile-ai/tilelang/pull/1104format.sh by @XuehaiPan in https://github.com/tile-ai/tilelang/pull/1102T.ptr and T.Tensor by @xwhzz in https://github.com/tile-ai/tilelang/pull/1114fence_barrier_init primitive after mbarrier init by @chengyupku in https://github.com/tile-ai/tilelang/pull/1121format.sh and introduce loop carry thread sync unit test by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1153T.warpgroup_fence_operand for nvcc code motion by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/986cibuildwheel and reduce size of sdist by @oraluben in https://github.com/tile-ai/tilelang/pull/1171tl.infinity operator for infinity handling of bfloat16 by @Rachmanino in https://github.com/tile-ai/tilelang/pull/1175ccache for CIBW on Linux by @oraluben in https://github.com/tile-ai/tilelang/pull/1184T.serial with step and negative step by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/1188reduce.h by @LJC00118 in https://github.com/tile-ai/tilelang/pull/1204libtvm as a dep of libtilelang by @oraluben in https://github.com/tile-ai/tilelang/pull/1215CompleteBufferFragment by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1226builder.py by @LJC00118 in https://github.com/tile-ai/tilelang/pull/1235-inf instead of clearing accumulators. by @Rachmanino in https://github.com/tile-ai/tilelang/pull/1222from __future__ import annotations for python 3.8 by @oraluben in https://github.com/tile-ai/tilelang/pull/1273int64_t static and dynamic shape. by @Elevator14B in https://github.com/tile-ai/tilelang/pull/1218T.view/reshape by @SiriusNEO in https://github.com/tile-ai/tilelang/pull/1277T.print for bool type by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1279T.Tensor(n * 2 + 1) in function annotation by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/1285T.Var annotation by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/1291T.assume handling by @LJC00118 in https://github.com/tile-ai/tilelang/pull/1292T.print by @xwhzz in https://github.com/tile-ai/tilelang/pull/1329NormalizeToBufferRegion and MakeAccessPtrFromRegion to utils by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/1333not operator in frontend (#1347) by @kurisu6912 in https://github.com/tile-ai/tilelang/pull/1348T.gemm_sp_v2 on sm80 and sm89 by @botbw in https://github.com/tile-ai/tilelang/pull/1056Full Changelog: https://github.com/tile-ai/tilelang/compare/0.1.6...v0.1.7
Your coding agent can read these notes before it upgrades. Set up the MCP server →