NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2543 most downloaded on PyPI
CUDA Tile Compiler
Last release 25 days ago
10 Sep 2026
Release timing varies
gaps range from 2 weeks to 5 months
Nearly every release is documented
notes for 8 of 8 stable releases
Nothing withdrawn
no release was ever pulled
1 years old
19 releases · first in 2025
One column per month.
This release adds support for CTK 13.4 features, including programmatic dependent launch, unchecked memory accesses, ct.insert(), and float8_e5m3fnu s
This release adds support for CTK 13.4 features, including programmatic
dependent launch, unchecked memory accesses, ct.insert(), and
float8_e5m3fnu scale dtype support for FP4 block-scaled MMA on Rubin. It
also adds finer control over compilation and execution with static array
strides and portable TileIR bytecode export. It expands the supported Python
subset with user-defined context managers, dataclass methods, break in
non-static loops, and integer divmod. Autotuning, JAX interoperability,
diagnostics, and floating-point behavior are also improved.
ct.grid_dependency_control_launch_dependents() <cuda.tile.grid_dependency_control_launch_dependents>,
{py:func}ct.grid_dependency_control_wait() <cuda.tile.grid_dependency_control_wait>,
and the programmatic_dependent_launch option for {py:func}ct.launch() <cuda.tile.launch>.check_bounds option to {py:func}ct.load() <cuda.tile.load>,
{py:func}ct.store() <cuda.tile.store>, and the
{py:meth}TiledView.load() <cuda.tile.TiledView.load> and
{py:meth}TiledView.store() <cuda.tile.TiledView.store> methods. It defaults
to True. Setting it to False declares that every tile element is within
the array bounds and skips bounds checking for faster access.ct.insert() <cuda.tile.insert> and
{py:meth}Tile.insert() <cuda.tile.Tile.insert> to create a new tile by
replacing a subtile of a larger tile. This operation is the inverse
of {py:func}ct.extract() <cuda.tile.extract>. The original tile remains unchanged.float8_e5m3fnu scale dtype support for FP4 block-scaled MMA on Rubin (sm_107).ct.ArrayAnnotation <cuda.tile.ArrayAnnotation> as Annotated
metadata and list the dimensions to specialize, for example,
Annotated[ct.Array, ct.ArrayAnnotation(static_stride_dims=(0, 1))].
Once any static stride is specified, the dispatcher stops automatically
inferring stride == 1 for all dimensions, so list the contiguous dimension
explicitly if it should remain a compile-time constant.ct.ensure_constant() <cuda.tile.ensure_constant>, which
statically asserts that a value is a compile-time constant.astype() conversions.propagate_nan option to {py:func}ct.min() <cuda.tile.min>,
{py:func}ct.max() <cuda.tile.max>,
{py:func}ct.argmin() <cuda.tile.argmin>,
{py:func}ct.argmax() <cuda.tile.argmax>,
{py:func}ct.minimum() <cuda.tile.minimum>, and
{py:func}ct.maximum() <cuda.tile.maximum>. By default, NaN values are
ignored. When enabled, a NaN propagates, min/max/minimum/maximum return NaN and
argmin/argmax return the index of the first NaN.break in non-static for loops.@contextlib.contextmanager.__post_init__(),
__call__(), __getitem__(), __setitem__(), __repr__(), and __str__().ct.divmod() <cuda.tile.divmod>, as well as support for
built-in divmod(), for integer inputs only.ct.static_eval() <cuda.tile.static_eval>.ArrayAnnotation(static_stride_dims=...). This allows more permissive type
unification in control flow.ct.tune.exhaustive_search() <cuda.tile.tune.exhaustive_search>
to accept a function for its kernel argument that maps each configuration
to the kernel to tune. Passing a fixed kernel continues to work as before.export_kernel() <cuda.tile.compilation.export_kernel> to
export TileIR bytecode without specifying gpu_code when
output_format="tileir_bytecode". This requires bytecode version 13.3 or
later.Array.slice() <cuda.tile.Array.slice> to a compile-time constant
dimension when both start and stop are compile-time constants.x ** y and ct.pow(x, y) to integer-exponent
FPowI when x is an unrestricted float and y is an integer whose value
range fits in a signed 32-bit integer. Other cases continue to promote the
exponent to floating point and use FPowF.TuningResult.failures <cuda.tile.tune.TuningResult.failures> from
an exception class name to the exception class itself.ct.argmin() <cuda.tile.argmin> and
{py:func}ct.argmax() <cuda.tile.argmax> ignore NaN values by default,
consistent with ct.min() and ct.max().ct.launch() <cuda.tile.launch> compiling kernels for
CUDA device 0 rather than for the device being launched on. Kernels are now
compiled for the launch device's architecture. Devices with the same architecture
can still share a compiled kernel.AttributeError when compiler crash dumps are enabled with
CUDA_TILE_ENABLE_CRASH_DUMP=1.cutile-cache log, a command-line interface for inspecting compilation
history. Cache metadata includes mangled kernel names, compiler versions,
compilation dates and durations, and, with TileIR 13.4, tileiras compiler
remarks.{#release-1-5-0}
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This release adds finer control over kernel specialization by making selected array shape dimensions compile-time constants and declaring scalar divis
This release adds finer control over kernel specialization by making selected array shape dimensions compile-time constants and declaring scalar divisibility assumptions. Autotuning is faster and can isolate hanging kernels. The supported Python subset now includes tuple comprehensions, tuple-valued kernel arguments, enums, dictionaries, and variadic keyword parameters.
ct.assume_divisible_by(x, divisor) <cuda.tile.assume_divisible_by>,
a compiler hint that declares an integer scalar to be divisible by a constant.
The compiler propagates this fact through arithmetic, which can prove the
alignment of derived indices and pointer offsets and enable wider memory
operations.ct.ArrayAnnotation <cuda.tile.ArrayAnnotation> as Annotated
metadata and list the dimensions to specialize, for example,
Annotated[ct.Array, ct.ArrayAnnotation(static_shape_dims=(0, -1))].single_run_timeout_sec argument to
{py:func}ct.tune.exhaustive_search() <cuda.tile.tune.exhaustive_search> to prevent a hanging kernel from stalling
the entire search.ct.Constant[tuple[int, float]] makes the
whole tuple compile-time constant, while tuple[ct.Constant[int], float]
makes only the first element constant.in and not in operators on tuples.def foo(**kwargs)), and dictionary
unpacking (for example, foo(x, **y)).enum.Enum inside kernels.
Enum members can be compared with == and != and constructed from a
constant value, such as Color(0).ct.uint32(-3), ct.int8(300), and ct.float16(0.2)) were not
converted according to their dtype. Out-of-range integers are now wrapped to
the dtype's range, and floating-point values are rounded or clamped according
to the dtype's precision and range.functools.wraps().order arguments of load(), store(), and
num_tiles().ct.arange() <cuda.tile.arange> with optional start and
step arguments: ct.arange(size, start=0, step=1, dtype=...). size must
be a constant integer, while start and step may be dynamic values. For
example, ct.arange(8, start=7, step=-1, dtype=ct.int32)
creates a reversed range.calling convention v2 <cuda.tile.compilation.CallingConvention> for kernels with static shape
annotations or tuple arguments.{#release-1-4-0}
This release highlights compatibility with CTK 13.3 which introduces Hopper(sm_90) GPU support, block-scaled MMA, new float dtypes, pack/unpack operat
This release highlights compatibility with CTK 13.3 which introduces Hopper(sm_90) GPU support, block-scaled MMA, new float dtypes, pack/unpack operations, atomic store operations, load and store with advanced indexing, tiled view with gapped or overlapped tile access, and support for large arrays and scalars. It also adds support for Python 3.14 including free-threading, star "*" expression, and frozen dataclasses, as well as integration with JAX.
float8_e8m0fnu and float4_e2m1fn restricted float dtype.ct.mma_scaled() <cuda.tile.mma_scaled> operation for
block-scaled matrix multiply-accumulate.ct.pack_to_bytes() <cuda.tile.pack_to_bytes> operation that
flattens a tile and reinterprets its raw bytes as a 1D uint8 tile;
{py:func}ct.unpack_from_bytes() <cuda.tile.unpack_from_bytes>
is the inverse of {py:func}ct.pack_to_bytes() <cuda.tile.pack_to_bytes>.ct.load_advanced_indexing() <cuda.tile.load_advanced_indexing>
and {py:func}ct.store_advanced_indexing() <cuda.tile.store_advanced_indexing>
for gathering/scattering along one dimension while slicing on other
dimensions.atomic_store_add <cuda.tile.TiledView.atomic_store_add> and
more atomic methods on
{py:class}TiledView <cuda.tile.TiledView> for performing element-wise
atomic read-modify-write operations on a tiled view at a given tile index.Array.tiled_view() <cuda.tile.Array.tiled_view>
with traversal_steps for creating a tiled space with overlapped or spaced tiles.ByTarget <cuda.tile.ByTarget> now accepts a default value that
applies to all architectures (e.g. ByTarget(sm_100=8, default=2)), allowing
generated TileIR bytecode to be independent of the GPU architecture.ct.astile() <cuda.tile.astile> for creating a tile from a
scalar (yielding a 0-d tile) or a (possibly nested) tuple of scalars whose
nesting determines the tile's shape.ct.IndexedWithInt64 <cuda.tile.IndexedWithInt64> annotation
for array kernel parameters whose shape or stride values exceed the range
of a 32-bit integer. Arrays without the annotation continue to use
int32 for shape and stride.ct.ScalarInt64 <cuda.tile.ScalarInt64> annotation that
forces a scalar integer kernel parameter to be inferred as int64
instead of the default int32.use_fast_acc option to {py:func}ct.mma() <cuda.tile.mma>
to enable fast accumulation mode for FP8 inputs (float8_e4m3fn,
float8_e5m2) on Hopper GPUs.num_worker_warps to {py:class}ct.kernel <cuda.tile.kernel>.ct.atomic_add() <cuda.tile.atomic_add> now supports bfloat16
operands on Hopper (sm_90) and newer architectures.rounding_mode parameter for {py:func}ct.exp() <cuda.tile.exp>
(supports RoundingMode.FULL and RoundingMode.APPROX for f32).foo(a, *b, c).def foo(*args).ct.extract() <cuda.tile.extract> now raises a compile-time
TileTypeError when a constant index is out of bounds for the tile grid.
Dynamic indices are unaffected.ct.floordiv() <cuda.tile.floordiv> and the // operator now
support floating-point operands.ct.jax.cutile_call() <cuda.tile.jax.cutile_call>.{#release-1-3-0}
Add API for ahead-of-time compilation and export via {py:func}compilation.export_kernel() . See the {doc}Compilation and Export section for more detai
compilation.export_kernel() <cuda.tile.compilation.export_kernel>.<br>
See the {doc}Compilation and Export </compilation> section for more details.tune.exhaustive_search() <cuda.tile.tune.exhaustive_search> and the following helpers:
kernel.replace_hints() <cuda.tile.kernel.replace_hints> to get a new kernel with updated hints.compiler_timeout() <cuda.tile.compiler_timeout> for temporarily setting the
timeout on the tileiras compiler.<br>
See the {ref}Autotuning <autotuning> section for more details.Array.tiled_view() <cuda.tile.Array.tiled_view> to create a tiled view of an array
with a fixed tile shape and padding mode.memory_order and memory_scope on cuda.tile.load and
cuda.tile.store operations.print() to handle tuple and nested fstring.TileTypeError.cuda.tile.Constant.{#release-1-2-0}
Support Ampere and Ada (sm80 family) GPUs.
pip install cuda-tile[tileiras] to use tileiras from Python environment
without system-wide CTK installation.ct.atan2(y, x) operation for computing the arctangent of y/x.rounding_mode parameter for ct.tanh(), supporting RoundingMode.FULL and
RoundingMode.APPROX.TileUnsupportedFeatureError.opt_level=0 on ct.kernel is no longer required for ct.printf() and ct.print().ct.static_iter keyword that enables compile-time for loops.ct.static_assert keyword that can be used to assert that a condition is true at compile time.ct.static_eval keyword that enables compile-time evaluation using the host Python interpreter.ct.scan() for custom scan.ct.isnan().print() and ct.print() that supports python-style print and f-strings.mask parameter to ct.gather() and ct.scatter() for custom boolean masking.+ can now be used to concatenate tuples.a, (b, c) = t) and using square brackets
for unpacking (e.g., [a, b] = 1, 2).CUDA_TILE_CACHE_DIR and CUDA_TILE_CACHE_SIZE.nan != nan returns False.$retval" error when a helper function
returns after a while loop that contains no early return.a, b = 1, 2, 3
now raises an error instead of silently discarding the extra value.~x for const boolean x will raise a TypeError to prevent inconsistent
results compared to ~x on a boolean Tile.TileUnsupportedFeatureError to the public API.{#release-1-1-0}
Add support for nested functions and lambdas.
ct.reduce().Array.slice(axis, start, stop) to create a view of an array sliced along a single axis.
The result shares memory with the original array (no data copy).float('inf').TileRecursionError, thrown at compile time when the recursion limit
is reached during function call inlining.for loop, and then used after the loop,
it is now an error because the loop may take zero iterations, resulting
in a use of an undefined variable.TileError base class in the public API.ct.abs() for completeness.{#release-1-0-1}
Fix a bug in hash function that resulted in potential performance regression for kernels with many specializations.
__eq__ comparison logic.ct.cat().is not None comparison.{#release-1-0-0}
Initial release.
Initial release.
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →