NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3905 most downloaded on PyPI
A Python framework for high-performance simulation and graphics programming
Last release 1 months ago
31 Aug 2026
Ships on a steady schedule
a new release about every 4 weeks
Nearly every release is documented
notes for 52 of 52 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
52 releases · first in 2022
One column per quarter.
…types has been removed; see Removals and deprecations for migration guidance.
Warp v1.17 expands geometry queries with sphere and capsule searches over BVHs, exact sphere queries against mesh triangles, and direct access to a mesh's BVH. Tiles now support matrix-row indexing, CG and CR solvers can restart periodically, and new controls let you tune and inspect CUDA kernel resource use. The release also includes experimental native build hooks for external C++ and CUDA integrations, along with native CPU support when building Warp from source on Windows ARM64.
If you are upgrading, note that implicit conversion of Python and Warp numeric scalars to composite types has been removed; see Removals and deprecations for migration guidance.
A BVH can now be queried with a sphere or capsule instead of first converting the search region to an AABB. wp.bvh_query_sphere() finds item bounds that overlap a sphere using an exact sphere-AABB test. wp.bvh_query_capsule() searches for item bounds that overlap the volume swept by moving a sphere along a line segment. Capsule queries are conservative: they do not miss bounds within the requested radius, but they can return extra candidates near AABB corners (#1741).
A capsule query takes a start point and a direction rather than two endpoints. To query the segment from p0 to p1, pass p0 as the start and p1 - p0 as the direction to wp.bvh_query_capsule(). By default, traversal continues indefinitely along that direction. Pass max_dist=1.0 to wp.bvh_query_next() to limit the query to the full segment, including both endpoints. If p0 == p1, use wp.bvh_query_sphere() instead.
wp.mesh_get_bvh() exposes a mesh's internal BVH to the general wp.bvh_query_*() APIs. This makes BVH-only operations such as wp.bvh_query_capsule() available for meshes, with returned bound indices corresponding to triangle faces. Use wp.mesh_query_sphere() when you need exact triangle-sphere intersections rather than broad-phase candidates. wp.MeshQuery is now the common base type for AABB and sphere mesh queries, with wp.mesh_query_next() as the canonical iterator for both. The existing wp.MeshQueryAABB type and wp.mesh_query_aabb_next() alias remain available.
CUDA kernels now have controls for register allocation and shared-memory spilling, and their resource use can be inspected before launch. cuda_max_registers requires Warp to have been built with CUDA Toolkit 12.4 or newer and cannot be combined with launch_bounds. The Linux and Windows warp-lang wheels on PyPI are built with CUDA Toolkit 12.9, so they meet this requirement. enable_cuda_smem_spilling takes effect only when Warp itself was built with CUDA Toolkit 13.0 or newer. Warp 1.17's PyPI wheels use CUDA Toolkit 12.9 and ignore this option. CUDA 13.0 wheels for Linux x86-64, Linux ARM64, and Windows x86-64 are available from GitHub Releases, or you can build Warp from source with CUDA Toolkit 13.0 or newer. CUDA 13 builds require an R580-series or newer NVIDIA driver and a Turing-class GPU (compute capability 7.5) or newer. See CUDA 13 PyPI wheel timing for the planned PyPI transition. Shared-memory spilling is also ignored when a kernel uses dynamic shared memory (#1671).
wp.get_cuda_kernel_properties() compiles the requested kernel variant if necessary, without launching it, and reports its per-thread register count and local-memory use (#1805).
import warp as wp
@wp.kernel(cuda_max_registers=64, enable_backward=False)
def update(values: wp.array[float]):
i = wp.tid()
values[i] = wp.sin(values[i]) + wp.cos(values[i])
properties = wp.get_cuda_kernel_properties(
update,
device="cuda:0",
block_dim=128,
)
print(sorted(properties)) # ['local_memory_size', 'register_count']Resource counts depend on the GPU, toolchain, compiler options, and block size. Use them to investigate occupancy and spilling, then profile the kernel before changing its configuration.
Matrix-valued tiles now support chained row indexing such as tile[i, j, k][row]. Reads, writes, negative row indices, and adjoints work for tiles with one through four logical dimensions (#1028).
import numpy as np
import warp as wp
TILE_SIZE = 8
@wp.kernel
def extract_last_row(matrices: wp.array[wp.mat33], rows: wp.array[wp.vec3]):
i = wp.tid()
# Load eight 3x3 matrices into a one-dimensional tile.
matrix_tile = wp.tile_load(matrices, shape=(TILE_SIZE,))
# Select matrix i from the tile, then use -1 to select its last row.
rows[i] = matrix_tile[i][-1]
# Repeat [[1, 2, 3], [4, 5, 6], [7, 8, 9]] eight times.
data = np.tile(
np.arange(1.0, 10.0, dtype=np.float32).reshape(3, 3),
(TILE_SIZE, 1, 1),
)
matrices = wp.array(data, dtype=wp.mat33, device="cuda:0")
rows = wp.zeros(TILE_SIZE, dtype=wp.vec3, device="cuda:0")
wp.launch(
extract_last_row,
dim=TILE_SIZE,
inputs=[matrices],
outputs=[rows],
block_dim=TILE_SIZE,
device="cuda:0",
)
print(rows.numpy()[0]) # Last row of the first matrix: [7. 8. 9.]Over a long CG or CR solve, the recursively updated residual can drift away from the true residual b - A x. Set restart=N on warp.optim.linear.cg() or warp.optim.linear.cr() to recompute it and reset the search direction every N iterations. This matters most for float32 CUDA workloads, including batched and matrix-free solves. The restart path also works with CUDA graph capture (#1708).
A restart requires one extra matrix-vector product. Warp checks convergence and invokes callbacks only at cycle boundaries, so the solve can run up to restart - 1 iterations past maxiter. Omit restart to keep the existing recursive behavior.
Batched CUDA solves also get more accurate dot products as subproblem size grows. When the largest subproblem is known, pass max_batch_length to LinearOperator or aslinearoperator() to avoid unnecessary reduction work (#1700). Reusable CG, CR, BiCGSTAB, and GMRES states returned with run=False now allocate their temporary device memory once during construction instead of on each solve.
Kernel factories can now give closure-generated kernels distinct, predictable names for registration and ahead-of-time compilation. Set name on @wp.kernel to assign the registration key and the base of the generated native entry-point name (#1561).
import warp as wp
def make_scaler(factor: float, kernel_name: str):
@wp.kernel(name=kernel_name)
def scale(values: wp.array[float]):
i = wp.tid()
values[i] *= factor
return scale
double = make_scaler(2.0, "scale_by_two")
triple = make_scaler(3.0, "scale_by_three")
print(double.key, triple.key) # scale_by_two scale_by_threeNames must be valid C++ identifiers. With strip_hash=True, Warp uses the custom key without a hash suffix as the base of generated entry-point names.
Important
This is an experimental feature. The API may change without a formal deprecation cycle.
Add-on packages can now attach native C++ or CUDA headers to a Warp module and make header-defined types and functions available to Warp kernels. wp.ModuleBuildOptions supplies include directories, preambles, and dependency files. wp.build_experimental.add_native_type() and add_builtin() register the matching ABI types and functions, while wp.compile_aot_module() returns artifact paths for an external build or runtime system (#1575). We have not yet validated this as a complete production workflow.
In this example, addon_math.h comes from the add-on package, not Warp. The example assumes this package layout:
my_addon/
├── build_addon.py
└── include/
└── addon_math.h
addon_math.h defines the native function that the add-on exposes to Warp kernels. Warp includes its native headers first, so the add-on header can use CUDA_CALLABLE:
include/addon_math.h:
#pragma once
namespace addon {
CUDA_CALLABLE inline float square(float value)
{
return value * value;
}
} // namespace addonbuild_addon.py locates the header relative to its own file, registers addon::square as wp.addon_square, and tells Warp where to find the header when compiling the module:
from pathlib import Path
import warp as wp
# Resolve the header shipped with this add-on package.
header = (Path(__file__).parent / "include/addon_math.h").resolve()
# Map the C++ function to the name and signature used in Warp kernels.
wp.build_experimental.add_builtin(
"addon_square",
{"value": wp.float32},
wp.float32,
native_name="addon::square",
)
# add_builtin() registers no adjoint, so generate only the forward kernel.
@wp.kernel(enable_backward=False)
def square_kernel(values: wp.array[float], output: wp.array[float]):
i = wp.tid()
output[i] = wp.addon_square(values[i])
build_options = wp.ModuleBuildOptions(
extra_cuda_include_dirs=[header.parent], # Let #include find addon_math.h.
extra_cuda_preamble='#include "addon_math.h"', # Include it in generated CUDA source.
extra_build_dependencies=[header], # Recompile when the header changes.
)
# Apply the add-on's build inputs before compiling this kernel's module.
wp.set_module_options({"extra_build_options": build_options}, module=square_kernel.module)
# Generate PTX for the compute capability used by the external application.
artifacts = wp.compile_aot_module(
square_kernel.module,
arch=80, # Match the external application's target compute capability.
module_dir=Path(__file__).parent / "generated",
use_ptx=True,
strip_hash=True, # Keep the exported kernel name free of hash suffixes.
)
print(artifacts[0].suffix) # .ptxThis example stops after generating PTX. It does not include the external application that loads and launches the PTX.
External native value types do not gain arithmetic or differentiation automatically. Register those functions separately.
wp.transform() with no arguments and same-type copies. Unsupported Python arguments now raise TypeError instead of silently returning an all-zero transform, and scalar construction inside kernels fills all seven components (#1742, #1814).wp.empty() and wp.zeros() allocations now account for gaps correctly and reject negative or dimensionally invalid strides (#1703).wp.map() and wp.utils.create_warp_function() now have stable identities across processes. This allows persistent kernel-cache reuse when separate processes discover the callables in different orders (#1696).example_fdtd_3d.py is a three-dimensional finite-difference time-domain simulation on a Yee grid. It models a Luneburg lens that collimates waves from a point source and supports both interactive Matplotlib visualization and headless execution (#1772).
pip does not check the installed driver or GPU architecture and may install a wheel whose CUDA backend cannot run on the system.This updates the tentative plan announced in Warp v1.16.0. See NVIDIA's CUDA minor-version compatibility table for driver requirements.
wp.mat22(123) (#1721).wp.mat22(np.float32(123)). This conversion will be removed in a future feature release under Warp's standard deprecation timeline (#1721).wp.Texture.copy_from_array() and wp.Texture.copy_to_array() have been removed. Use wp.Texture.copy_from() and wp.Texture.copy_to() (#1722).wp.spatial_jacobian() and wp.spatial_mass() have been removed. Although previously listed as built-ins, neither function was callable from kernels or Python (#1768).Migration examples:
- wp.launch(consume_matrix, dim=1, inputs=[123])
+ wp.launch(consume_matrix, dim=1, inputs=[wp.mat22(123)])
- wp.launch(consume_matrix, dim=1, inputs=[np.float32(123)])
+ wp.launch(consume_matrix, dim=1, inputs=[wp.mat22(np.float32(123))])We thank the following contributors from outside the core Warp development team:
wp.empty() and wp.zeros() arrays (#1703).@wp.kernel(name=...) (#1561, #1570).For a complete list of changes, see the full changelog.
wp.bvh_query_sphere() and wp.bvh_query_capsule() for Minkowski-offset broad-phase queries over BVH bounds.
Sphere queries use exact sphere-AABB overlap tests, while capsule queries are conservative and may return extra
candidates near bounding-box corners. Add wp.mesh_query_sphere() for finding mesh triangles that intersect a sphere
and wp.mesh_get_bvh() for querying a mesh's BVH directly
(GH-1741).wp.ModuleBuildOptions,
the extra_build_options module option, wp.build_experimental.add_builtin(),
wp.build_experimental.add_native_type(), external kernel entry-point ABIs, and AOT artifact paths. These APIs may
change without deprecation in a future release (GH-1575).cuda_max_registers and enable_cuda_smem_spilling parameters to @wp.kernel to control CUDA register
allocation and shared-memory spilling. cuda_max_registers takes effect in builds using CUDA Toolkit 12.4 or later.
enable_cuda_smem_spilling takes effect in builds using CUDA Toolkit 13.0 or later. LLVM CUDA compilation ignores
both parameters (GH-1671).tile[i, j, k][r], including reads, writes,
negative row indices, and adjoints (GH-1028).@wp.kernel(name=...) to assign a custom name that Warp uses for kernel registration and as the base of generated
native entry-point names (GH-1561).wp.get_cuda_kernel_properties() to inspect a CUDA kernel variant's per-thread register count
and local-memory use. Warp compiles the requested variant if necessary but does not launch it, and returns the
properties as a dictionary (GH-1805).restart parameter to warp.optim.linear.cg() and warp.optim.linear.cr() to periodically
recompute the true residual and reset the search direction, limiting drift
from finite-precision arithmetic (GH-1708).wp.mat22(123) (GH-1721).wp.Texture.copy_from_array() and wp.Texture.copy_to_array() aliases; use
wp.Texture.copy_from() and wp.Texture.copy_to() instead (GH-1722).wp.spatial_jacobian() and wp.spatial_mass(). Although they were listed as built-ins, they were never
callable from kernels or Python (GH-1768).wp.mat22(np.float32(123)) (GH-1721).wp.MeshQuery as the common base type for AABB and sphere mesh queries while keeping wp.MeshQueryAABB as the
concrete AABB query type. Allow kernels to select query kinds at runtime, pass the resulting wp.BvhQuery or
wp.MeshQuery values through functions, and iterate them. Make wp.mesh_query_next() the canonical iterator for AABB
and sphere mesh queries; wp.mesh_query_aabb_next() remains as an alias
(GH-1741).wp.transform() with no arguments and copies of
an existing transform of the same type (GH-1742).wp.copy() or wp.array.assign() call in array-overwrite warnings
when wp.config.verify_autograd_array_access is enabled (GH-1727).wp.copy(), wp.clone(), or
wp.array.assign() is read again later in the forward pass. For non-overlapping same-device arrays and contiguous
CPU-CUDA copies with matching numeric dtypes, gradients now accumulate into the source without propagating to
overwritten destination values. Unsupported copy patterns now emit a warning
(GH-1728).wp.config.verify_autograd_array_access when reusing arrays
after wp.Tape.backward() or wp.Tape.reset() (GH-1727).wp.utils.radix_sort_pairs() and wp.utils.segmented_sort_pairs():
scratch-buffer growth during capture could corrupt the CUDA context or prevent graph instantiation; growth after
capture could invalidate previously captured graphs; concurrently launched graphs could race over shared scratch
storage; and sorts in conditional body graphs or on devices without memory pool support could fail or invalidate
capture (GH-1373).wp.empty() and wp.zeros() for custom-strided layouts with gaps. Newly allocated
arrays now reject negative strides and stride counts that do not
match the number of dimensions (GH-1703).AssertionError
or produce incorrect results (GH-1684).arr[i] / term[i] if flag else arr[i] (GH-1736).wp.Texture2D and wp.Texture3D sampling from optimized CUDA 12 CUBIN modules when texture handles
vary across lanes on Hopper and newer GPUs (GH-1811).wp.jax_kernel() and wp.jax_callable() from reusing cached JAX FFI wrappers when calls use different
in_out_argnames values or, for wp.jax_callable(), different stage_in_argnames or stage_out_argnames values
(GH-1215).wp.transform(...) at Python scope silently returning an all-zero transform for unsupported arguments. Scalar
inputs now fill all seven components, invalid inputs raise TypeError, and keyword arguments passed alongside all
seven positional components are no longer ignored (GH-1742).wp.transform(value) inside kernels. It now fills all seven components with value
instead of failing during compilation (GH-1814).wp.transform(other). Cross-dtype copies remain
unsupported and now report a clear error (GH-1814).max_batch_length argument to warp.optim.linear.LinearOperator and warp.optim.linear.aslinearoperator() to skip
redundant reduction levels when the maximum number of scalar degrees of freedom in any subproblem is known
(GH-1700).warp.optim.linear.gmres() and restart-enabled
warp.optim.linear.cg() and warp.optim.linear.cr() when check_every=0 and a CUDA graph-captured cycle performs
multiple iterations. The array now reports the actual number of
completed solver iterations (GH-1707).warp.optim.linear.cg(), warp.optim.linear.cr(), warp.optim.linear.bicgstab(), and
warp.optim.linear.gmres() with run=False now allocate the required memory once when created, allowing repeated
solves during CUDA graph capture.warp.sparse.bsr_compress(..., topology="compact") on CUDA treating unused trailing storage as active blocks,
which could incorrectly append those blocks
to the matrix's last row (GH-1769).max_workers option of wp.force_load()
or wp.load_module() (GH-1705).wp.force_load() and wp.load_module() calls
instead of returning successfully (GH-1749).libwarp-clang.dylib to Warp's native API
and required JIT-debugger symbols (GH-1758).wp.load_aot_module() failing to find ahead-of-time binaries when the module already had strip_hash=True
and the load call omitted strip_hash. Omitting the argument now preserves the existing module setting
instead of resetting it to False (GH-1730).wp.map() or wp.utils.create_warp_function() being
recompiled in every new process. Generated callable identities are now stable across processes, allowing persistent
kernel cache reuse and preventing redundant cache entries (GH-1696).wp.Tape.backward() failing with CUDA error 719 for tile kernels in debug mode when the same module also
contained wp.float16 or wp.bfloat16 vector or matrix code, even if that code
was never launched (GH-1809).wp.autograd.gradcheck() and wp.autograd.jacobian_fd() producing incorrect results for Python functions
that mutate array inputs, and prevent wp.autograd.jacobian() from leaving those inputs modified. All three
functions now accept restore_inputs, which defaults to True and restores top-level Warp array inputs before each
evaluation and before returning (GH-1726).enable_backward=False: warp.autograd.jacobian() now rejects kernels disabled
at either module or kernel scope, warp.autograd.gradcheck_tape() skips them, and warp.autograd.jacobian_fd()
works with backward-disabled kernels because it uses only
forward launches (GH-1718).wp.autograd.gradcheck(), wp.autograd.jacobian_fd(), and
wp.autograd.gradcheck_tape() when input and output arrays use different floating-point precisions, such as
wp.float32 inputs with a wp.float64 output (GH-1726).wp.fabricarray and wp.indexedfabricarray values
to be used as fields in @wp.struct types (GH-1818).TypeError when a Warp array assigned to an @wp.struct field does not match the dimensionality
declared by the field's annotation. This applies to regular, indexed, Fabric, and indexed Fabric arrays.2**31 when a kernel uses scalar wp.tid(),
preventing the signed 32-bit thread index from wrapping. The validation applies to wp.launch(),
wp.Launch.set_dim(), recorded launches, and JAX integration (GH-1799).wp.launch_tiled(). Errors now report the
kernel's usable dynamic shared-memory budget and the relevant block_dim, and appear on every failed launch instead
of degrading to a generic CUDA error after the first attempt (GH-1701).wp.capture_begin(..., apic=False) causing some host applications to retain caller frames
and large caller-owned objects after the first CUDA graph capture. Warp now loads its API Capture (APIC)
implementation for CUDA captures only when apic=True (GH-1829).wp.Tape integration in custom torch.autograd.Function implementations,
including zero-copy tensor aliasing, external gradient-buffer ownership, repeated backward calls, gradcheck,
and CUDA stream handling (GH-1737).@wp.func. Return loop accumulators
to the calling kernel, and perform operations whose derivatives depend on their final values there,
to avoid incorrect gradients (GH-1754).wp.config.launch_array_access_mode controls cross-device Warp array validation for kernel launches.
Explain why the default RELAXED mode can allow inaccessible arrays to crash a kernel, what CHECKED can verify, and
which valid cross-device accesses STRICT rejects (GH-1693).block_dim, and correct the
documented overflow error (GH-1699).warp/examples/core/example_fdtd_3d.py, a 3-D finite-difference time-domain (FDTD) simulation of a Luneburg lens
that collimates waves from a point source (GH-1772).wp.capture_if() and wp.capture_while() require
Python callback bodies for CPU capture and CUDA capture with apic=True; wp.Graph.get_param_ptr() returns
a host address for CPU graphs; and .wrp files store an APIC operation stream rather than
a prebuilt CUDA graph (GH-1639).Python badges to built-in overloads that can be called from Python scope while leaving
kernel-only overloads unmarked (GH-1825).Important This is an experimental feature. The API may change without a formal deprecation cycle.
Warp v1.16 adds in-place rebuilding for fixed-capacity NanoVDB volumes, including during CUDA graph capture. This lets fluid simulations update sparse-grid topologies without allocating new volumes. The release also adds grouped HashGrid queries for multi-environment workloads, NumPy-style tile slicing, CPU support for JAX FFI wrappers, experimental support for replaying more operations in CPU graphs and saved .wrp graphs, and CUDA profiler range controls.
Previously, a volume's sparse topology was fixed at allocation time. Simulations whose active grid changed from one step to the next, such as the affine particle-in-cell (APIC) fluid example, had to allocate a new volume. That allocation required host synchronization and could not be replayed inside a CUDA graph. Warp 1.16 adds fixed-capacity, in-place rebuilding so simulations can reuse the same volume and dependent FEM topology buffers across graph replays (#1606).
The excerpt below follows that example's allocation, rebuild, and replay flow. It omits particle initialization, capacity estimation, FEM space construction, and the APIC transfer and solve.
import warp as wp
import warp.fem as fem
# API-shape excerpt. Simulation setup and solver work are omitted.
# particle_q, voxel_size, grid_capacity, and frame_count come from that setup.
grid_status = wp.zeros(1, dtype=wp.uint32, device=particle_q.device)
volume = wp.Volume.allocate_by_voxels(
voxel_points=particle_q,
voxel_size=voxel_size,
device=particle_q.device,
rebuildable=True,
**grid_capacity,
status=grid_status,
)
grid = fem.Nanogrid(volume, rebuildable=True)
# ... build linear_basis_space and strain_space once from grid ...
def simulate():
# Rebuild from the particle positions updated by the previous step.
grid.rebuild(particle_q, status=grid_status)
linear_basis_space.topology.rebuild()
strain_space.topology.rebuild()
# ... transfer particles to the grid, solve, and advect particles ...
with wp.ScopedCapture(particle_q.device) as capture:
simulate()
for _ in range(frame_count):
wp.capture_launch(capture.graph)
# Host-side status and topology queries stay outside capture.
status = int(grid_status.numpy()[0])
stats = volume.get_active_stats()
print(f"rebuild status=0x{status:x}, active voxels={stats.voxel_count}")Capacity does not grow automatically. Pass a one-element wp.uint32 status array to the allocation method and to rebuild(), then check for wp.Volume.REBUILD_SUCCESS, a REBUILD_*_CAPACITY_EXCEEDED flag, or REBUILD_INVALID_INPUT. A tile-allocated volume must be rebuilt from tile positions, while a voxel-allocated volume must be rebuilt from voxel positions. The point_mask parameter can exclude input points, and Warp deduplicates repeated positions.
On CUDA, rebuildable allocation and rebuilding can run inside graph capture when memory-pool allocation is enabled. Exact, non-rebuildable allocation and host-side queries such as get_active_stats() must remain outside capture.
See example_apic_fluid.py for the complete capacity estimation, topology construction, and captured simulation loop.
API changes:
| Status | API | Change in Warp 1.16 |
|---|---|---|
| Existing class method | wp.Volume.allocate_by_tiles() |
Adds CPU support and the rebuildable, max_tiles, max_lower_nodes, max_upper_nodes, status, and point_mask parameters. |
| Existing class method | wp.Volume.allocate_by_voxels() |
Adds CPU support and the rebuildable, max_active_voxels, max_leaf_nodes, max_lower_nodes, max_upper_nodes, status, and point_mask parameters. |
| Existing constructor | warp.fem.Nanogrid() |
Adds the rebuildable parameter for retaining capacity-sized topology buffers. |
| New instance methods | wp.Volume.rebuild() and warp.fem.Nanogrid.rebuild() |
Rebuild the volume directly or rebuild it through a Nanogrid that also refreshes its geometry buffers. |
| New topology methods | topology.rebuild() on compatible Nanogrid function spaces |
Refresh topology buffers after rebuilding the geometry. |
| New inspection APIs | wp.Volume.is_rebuildable, wp.Volume.get_rebuild_info(), and wp.Volume.get_active_stats() |
Report whether a volume can be rebuilt, its reserved capacity, and its current active topology. |
| New status constants | wp.Volume.REBUILD_* |
Report success, exceeded capacities, or invalid input through the optional status array. |
| New kernel builtin | wp.volume_voxel_count() |
Returns the current active voxel or index count inside kernels and captured graphs. |
Independent simulation environments can now share one spatial grid without visiting candidates from neighboring worlds. Assign each point an int32 group in HashGrid.build(), then pass the desired group to wp.hash_grid_query() (#1579).
import warp as wp
@wp.kernel
def count_neighbors(
grid: wp.uint64,
points: wp.array[wp.vec3],
groups: wp.array[wp.int32],
counts: wp.array[wp.int32],
):
i = wp.tid()
count = int(0)
# Restrict the query to candidates in the current point's environment.
for j in wp.hash_grid_query(grid, points[i], 0.5, groups[i]):
if wp.length(points[j] - points[i]) <= 0.5:
count += 1
counts[i] = count
points = wp.array([(0.0, 0.0, 0.0), (0.4, 0.0, 0.0)] * 2, dtype=wp.vec3)
# The repeated two-point scene represents environments 0 and 1.
groups = wp.array([0, 0, 1, 1], dtype=wp.int32)
counts = wp.zeros(4, dtype=wp.int32)
grid = wp.HashGrid(8, 8, 8)
grid.build(points, radius=0.5, groups=groups)
wp.launch(count_neighbors, dim=4, inputs=[wp.uint64(grid.id), points, groups], outputs=[counts])
print(f"same-group neighbor counts: {counts.numpy()}") # same-group neighbor counts: [2 2 2 2]Group IDs may be any int32 values and can change between rebuilds, including during graph replay. If the query omits the group argument, it visits all candidates. Before capturing a grouped CUDA rebuild without a warm-up build, call grid.reserve(num_points, with_groups=True). Hash-grid queries return cell candidates, so kernels should still test the actual distance.
Tile kernels now use NumPy-style indexing to select, reverse, stride, gather, and assign subregions. An integer index collapses a dimension and may be negative. Use wp.tile_slice_indexed() for one-axis gathers such as tile[indices, :] (#1176).
import numpy as np
import warp as wp
@wp.kernel
def slice_tile(src: wp.array2d[float], dst: wp.array2d[float]):
tile = wp.tile_load(src, shape=(4, 4))
# Reverse the rows and keep columns 0 and 2.
wp.tile_store(dst, tile[::-1, ::2])
src = wp.array(np.arange(16, dtype=np.float32).reshape(4, 4), device="cuda:0")
dst = wp.zeros((4, 2), dtype=float, device="cuda:0")
# Launch one cooperative 32-thread block for the single tile.
wp.launch_tiled(slice_tile, dim=[1], inputs=[src], outputs=[dst], block_dim=32, device="cuda:0")
print(f"sliced tile (reversed rows, even columns):\n{dst.numpy()}")Output:
sliced tile (reversed rows, even columns):
[[12. 14.]
[ 8. 10.]
[ 4. 6.]
[ 0. 2.]]
The source tile in a slice assignment must have a matching shape. Tile views do not support compound assignment, and slices do not support scalar broadcasting. Indexed gathers select one axis at a time and require full : slices on the other axes.
The out-of-place wp.tile_lower_solve() now supports gradients for vector and matrix right-hand sides. This makes lower-triangular solves available in wp.Tape backward passes (#1378).
JAX can now run wp.jax_kernel() and wp.jax_callable() wrappers on CPU as well as CUDA. JAX selects the device, so a CPU jax.jit program invokes the Warp callback without routing buffers through a GPU (#1661).
import jax
import jax.numpy as jnp
import numpy as np
import warp as wp
@wp.kernel
def triple(x: wp.array[float], out: wp.array[float]):
i = wp.tid()
out[i] = 3.0 * x[i]
run = wp.jax_kernel(triple)
with jax.default_device(jax.devices("cpu")[0]):
# JAX places both the arrays and the Warp FFI callback on CPU.
(result,) = jax.jit(run)(jnp.arange(4, dtype=jnp.float32))
print(f"Warp FFI result on CPU: {np.asarray(result)}") # Warp FFI result on CPU: [0. 3. 6. 9.]JAX 0.5.0 or newer is required. JaxCallableGraphMode.NONE and JaxCallableGraphMode.JAX work on CPU. The Warp-managed graph modes remain CUDA-only. example_jax_kernel.py shows more call shapes.
JAX FFI wrappers created with wp.jax_kernel() now accept block_dim, letting a CUDA tile kernel choose how many threads cooperate on each tile. The value is fixed when the wrapper is constructed and is also used for generated adjoint launches (#1436).
import jax
import jax.numpy as jnp
import numpy as np
import warp as wp
ROW_COUNT = 4
TILE_SIZE = 256
TILE_THREADS = 64
@wp.kernel
def row_sum(values: wp.array2d[float], output: wp.array[float]):
row = wp.tid()
# Threads in one CUDA block cooperate to reduce each row.
tile = wp.tile_load(values[row], shape=TILE_SIZE)
wp.tile_store(output, wp.tile_sum(tile), offset=row)
# JAX FFI uses wp.launch(), so append the block width to the logical row count.
jax_row_sum = wp.jax_kernel(
row_sum,
launch_dims=(ROW_COUNT, TILE_THREADS),
output_dims=(ROW_COUNT,),
block_dim=TILE_THREADS,
)
with jax.default_device(jax.devices("gpu")[0]):
values = jnp.arange(ROW_COUNT * TILE_SIZE, dtype=jnp.float32).reshape(ROW_COUNT, TILE_SIZE)
(row_sums,) = jax.jit(jax_row_sum)(values)
print(f"CUDA tile row sums: {np.asarray(row_sums)}") # CUDA tile row sums: [ 32640. 98176. 163712. 229248.]TILE_SIZE is the logical data shape, while TILE_THREADS is a CUDA execution choice. Keep output_dims equal to the logical result shape rather than including the block width. See CUDA block dimensions and tile kernels for more detail.
.wrp graphsImportant
This is an experimental feature. The API may change without a formal deprecation cycle.
API Capture can now record wp.utils.array_sum() and wp.utils.array_inner() on CPU and CUDA, allowing saved .wrp graphs to recompute those reductions from current inputs during replay. Live CPU graphs can also rebuild wp.HashGrid data and refit or rebuild wp.Bvh trees from their current arrays (#1663, #1664, #1665).
import warp as wp
values = wp.array([1.0, 2.0, 3.0], dtype=float, device="cpu")
total = wp.zeros(1, dtype=float, device="cpu")
dot = wp.zeros(1, dtype=float, device="cpu")
with wp.ScopedCapture(device="cpu") as capture:
wp.utils.array_sum(values, out=total)
wp.utils.array_inner(values, values, out=dot)
wp.capture_launch(capture.graph)
print(
f"replayed reductions: sum={total.numpy()[0]}, "
f"inner product={dot.numpy()[0]}"
) # replayed reductions: sum=6.0, inner product=14.0Non-empty reductions recorded through API Capture require an explicit output array. Negative strides are unsupported. Counts and layout strides must fit a signed 32-bit integer. Because resources use process-local handles, graphs that record HashGrid.build(), Bvh.refit(), or Bvh.rebuild() are replay-only. wp.capture_save() rejects these graphs instead of writing a non-portable file. See CPU Graphs and Saving and Loading Graphs for the full lists of operations and limitations.
External profilers can now skip initialization and warm-up work and capture only the steady-state region. wp.cuda_profiler_start(), wp.cuda_profiler_stop(), and wp.ScopedCudaProfiler expose CUDA profiler range controls from Python (#1596).
import warp as wp
@wp.kernel
def increment(values: wp.array[float]):
values[wp.tid()] += 1.0
values = wp.zeros(1024, dtype=float)
# Compile and warm up before starting profiler collection.
wp.launch(increment, dim=values.size, inputs=[values])
wp.synchronize_device()
with wp.ScopedCudaProfiler():
for _ in range(10):
wp.launch(increment, dim=values.size, inputs=[values])
# Finish asynchronous launches before profiler collection stops.
wp.synchronize_device()The warm-up launch stays outside the capture range. The synchronization inside the scope ensures that the asynchronous launches complete before wp.cuda_profiler_stop() marks the end of collection. CUDA does not guarantee that the stop call synchronizes the device.
Configure the external profiler to honor these calls. For example:
nsys profile --capture-range=cudaProfilerApi python my_app.pyFor Nsight Compute, use --profile-from-start off. Warp invokes the profiler-control calls with the selected device's CUDA context current; the profiler and its configuration determine which activity is collected during that interval.
wp.quat_twist_angle_signed() recovers the signed rotational coordinate represented by a quaternion. For small float32 rotations, the existing unsigned wp.quat_twist_angle() is now more accurate (#1631).wp.Stream.is_blocking reports which behavior a stream uses. Interoperability code can check borrowed streams before releasing temporary Warp resources, then either synchronize a non-blocking stream on the host or make a blocking Warp stream wait on it. See Warp's non-blocking stream guidance for both patterns (#1618).repr() output such as wp.array4d[wp.uint32]. Generated API documentation uses the same form that appears in annotations (#1628).@wp.kernel(module="unique") declarations from factory functions are roughly 2x faster in microbenchmarks, whether their captured values produce identical or specialized kernels (#1486).Two new distributed Jacobi solver examples expand on example_jacobi_mpi.py, which uses CUDA-aware mpi4py directly for nearest-neighbor halo exchange. The new variants keep MPI for process setup but move the halo transfers to GPU communication libraries:
example_jacobi_nccl.py uses NCCL, NVIDIA's topology-aware inter-GPU communication library, through nccl4py, with MPI retained for rank setup (#1576).example_jacobi_nvshmem.py uses NVSHMEM, NVIDIA's symmetric-memory library for GPU clusters, to allocate Warp arrays on the symmetric heap and exchange halo rows through host-initiated nvshmem4py transfers (CUDA 12 package, CUDA 13 package), with MPI retained for bootstrap (#1582).pip does not check the installed driver and may install a wheel whose CUDA backend cannot run on an older driver.This repeats the tentative plan announced in Warp v1.15.0. See NVIDIA's CUDA minor-version compatibility table for driver requirements.
wp.vec3(...) or wp.mat22(...) when launching kernels or assigning struct fields (#1022).wp.Texture.copy_from_array() and wp.Texture.copy_to_array() are scheduled for removal in Warp 1.17. Use wp.Texture.copy_from() and wp.Texture.copy_to() instead (#1238).warp.jax_experimental namespace will remain available through Warp 1.17. Its removal, including the legacy custom-call jax_kernel() and graph-cache getter and setter APIs, is now scheduled for Warp 1.18. Migrate to the top-level wp.jax_kernel() and wp.jax_callable() APIs (#1370).warp.config.verbose and warp.config.quiet are scheduled for removal in Warp 1.18. Use warp.config.log_level = warp.LOG_DEBUG and warp.config.log_level = warp.LOG_WARNING, respectively. warp.config.verbose_warnings is unaffected (#1315).wp.HashGridQueryH and wp.HashGridQueryD are scheduled for removal in Warp 1.18. Use wp.HashGridQuery in public type annotations (#1452).max_verts, max_tris, and device arguments and compatibility attributes on wp.MarchingCubes are scheduled for removal in Warp 1.19. This includes the max_verts and max_tris arguments to wp.MarchingCubes.resize() and the id and runtime attributes. Output arrays are dynamically sized, and extraction uses the input field's device (#1594).masked=True form of warp.sparse topology-changing operations is scheduled for removal in Warp 1.19. Use topology="masked" with bsr_set_from_triplets(), bsr_assign(), bsr_set_transpose(), bsr_axpy(), and bsr_mm() (#1537).warp.fem.Nanogrid.from_environment_voxels() and warp.fem.AdaptiveNanogrid.from_environment_voxels() is deprecated. It is scheduled for removal in Warp 1.19. Pass flat points, cell_levels where applicable, point_envs, and env_count instead (#1606).warp.sparse.BsrMatrix.copy_nnz_async() is deprecated. It is scheduled for removal in Warp 1.20. Use warp.sparse.BsrMatrix.notify_nnz_changed() after updating the nonzero count (#987).We also thank the following contributors from outside the core Warp development team:
warp.fem.lookup() with wp.float64 warp.fem.Grid2D and warp.fem.Grid3D geometries (#1660) and updating the FEM examples for current Matplotlib releases.wp.float16 parameters (#1593).For a complete list of changes, see the full changelog.
wp.Volume.allocate_by_tiles(..., rebuildable=True),
wp.Volume.allocate_by_voxels(..., rebuildable=True), and wp.Volume.rebuild(). Support fixed capacities, optional
point masks, CPU execution, CUDA graph-capturable allocation and rebuilding, and in-place refreshes of rebuildable
warp.fem.Nanogrid topologies. See warp/examples/fem/example_apic_fluid.py for a captured grid-rebuild workflow
(GH-1606).wp.HashGrid for multi-environment workloads via the optional groups
argument to wp.HashGrid.build() and an optional trailing group argument to wp.hash_grid_query() that restricts
traversal to points sharing the requested group ID. Use wp.HashGrid.reserve(num_points, with_groups=True) to record
grouped rebuilds inside a CUDA graph without a warm-up build (GH-1579).t[2:6, :], t[:, ::2], t[::-1, :]),
dimension-collapsing integer indices with negative-index support (t[5, :], t[-1, :]), and slice assignment
(t[0:4, :] = src). Also add wp.tile_slice_indexed(), which gathers elements along a single axis
using a 1D integer index tile (t[indices, :]) (GH-1176).wp.jax_kernel() and wp.jax_callable() FFI calls, with automatic dispatch between CPU and CUDA
based on the device selected by JAX (GH-1661).block_dim argument to wp.jax_kernel() for selecting the CUDA thread-block size, including for
tile kernels and their generated adjoint launches (GH-1436).wp.utils.array_sum() and wp.utils.array_inner()
on CPU and CUDA (GH-1663).wp.HashGrid.build(), wp.Bvh.refit(), and
wp.Bvh.rebuild(). These captures are replay-only and cannot be saved with wp.capture_save()
(GH-1664, GH-1665).wp.quat_twist_angle_signed() to recover the signed rotational coordinate represented by a quaternion
(GH-1631).wp.cuda_profiler_start(), wp.cuda_profiler_stop(), and the wp.ScopedCudaProfiler context manager to
control CUDA profiler data collection from Python (equivalent to cuProfilerStart/cuProfilerStop) for an
external profiler's capture range. They accept an optional device argument and act on that device's current CUDA
context (GH-1596).wp.tile_lower_solve() with vector and matrix right-hand sides,
enabling it in backward passes (GH-1378).wp.Stream.is_blocking to report whether a CUDA stream is blocking
(GH-1618).warp.jax_experimental namespace, its legacy custom-call jax_kernel(),
and its graph-cache getter and setter APIs from Warp 1.16 to Warp 1.18, allowing more time to migrate.warp.fem.Nanogrid.from_environment_voxels() and
warp.fem.AdaptiveNanogrid.from_environment_voxels(); pass flat points, cell_levels where applicable,
point_envs, and env_count instead (GH-1606).warp.sparse.BsrMatrix.copy_nnz_async(); use warp.sparse.BsrMatrix.notify_nnz_changed() instead
(GH-987).@wp.kernel(module="unique") declarations from factory functions by roughly 2× in microbenchmarks,
whether the resulting kernels are identical or specialized with different captured values
(GH-1486).repr() of array type annotations in their subscript form (such as wp.array4d[wp.uint32]) so that it matches
the written annotation and can be evaluated back to the same annotation, instead of the previous
wp.array(dtype=..., ndim=...) form. Warp array return types now read more naturally in Sphinx-generated API
documentation (GH-1628).if/else branches
(GH-1574).deterministic mode,
which could execute consumed-return counter atomics twice or crash the process
(GH-1637).range() loop is reused by a later while or dynamic
for loop. Assignments to the reused index now carry across iterations, including when it is first re-declared
(e.g., i = int(0)), instead of being silently discarded (GH-1534).@wp.func functions when passing a Warp function to a wp.Function parameter
(GH-1648).wp.ref[T] parameters (e.g., x, y = a, b) so they update the caller's storage
instead of raising a type error (GH-1581).warp.optim.Adam.set_params() and warp.optim.SGD.set_params() to preserve compatible optimizer state.
This covers repeated calls with wp.float16 Adam parameters and parameters moved between devices
(GH-1593, GH-1615).@wp.func helper is used by both backward-enabled and
backward-disabled kernels in a module.@wp.func_grad or @wp.func_replay function needs more shared
memory than the corresponding forward function (GH-1646).wp.autograd.jacobian(),
wp.autograd.jacobian_fd(), and wp.autograd.gradcheck() (GH-1672).wp.tile_matmul() returning incorrect results on the MathDx path when an operand is a strided tile, such as a
wp.tile_view() into a wider tile (GH-1667).wp.tile_load_indexed() reading out of bounds for negative gather indices. Negative indices now yield zero,
matching indices past the end of the axis and allowing -1 to serve as a padding sentinel
(GH-1653).wp.func_grad() raising KeyError: 'output_arch' when defining a custom gradient for a function that uses
wp.tile_matmul() or other MathDx tile built-ins (GH-1668).wp.sparse.bsr_set_transpose() with topology="padded" when the
destination has insufficient row capacity (GH-1630).warp.fem.make_space_partition() for multi-environment spaces with
environment_first=True and max_node_count set (GH-1607).wp.capture_if() or wp.capture_while() body graph (GH-1641)..wrp files containing duplicate memory records or truncated initial data,
preventing out-of-bounds memory access.wp.bvh_get_group_root() and wp.mesh_get_group_root() returning roots that let queries traverse later groups
when group IDs are sparse (GH-1612).warp.fem.lookup() with wp.float64 geometries
(GH-1660).make_node_coords_in_element() returning None in Warp FEM kernels for SquareNedelecFirstKindShapeFunctions
and SquareRaviartThomasShapeFunctions (GH-1685).wp.from_dlpack() rejecting standards-conformant 8-bit Boolean tensors
(GH-1619).wp.quat_twist_angle() losing precision for small wp.float32 rotations
(GH-1631).wp.config.kernel_cache_dir or WARP_CACHE_PATH.cpu_compiler_flags.
Compatible headers are reused, while differing flags receive separate headers,
avoiding repeated full compilation and Clang target-feature errors
(GH-1658, GH-1649).nccl4py) and NVSHMEM (nvshmem4py)
for nearest-neighbor halo exchange
(GH-1576, GH-1582).wp.tile_matmul() supports only real floating-point tiles: remove the unsupported vec2h, vec2f, and
vec2d complex types from its documentation, and report a targeted error when complex tiles are passed
(GH-1682).Important This is an experimental feature. The API may change without a formal deprecation cycle.
Warp v1.15 adds opt-in deterministic execution for supported atomic patterns, broader replay for CPU and saved CUDA graphs, CUDA managed-memory allocation, and mipmapped textures. It also extends Warp functions with in-place references and function arguments, adds struct support to tile operations, and lets sparse matrices reserve BSR row capacity.
Simulation, validation, and regression-test workloads can now trade some performance for reproducible ordering of common atomic updates. wp.DeterministicMode.RUN_TO_RUN produces bit-exact repeated results on the same GPU architecture, while wp.DeterministicMode.GPU_TO_GPU uses a more conservative path intended to preserve results across GPU architectures (#1443).
Note
Deterministic mode changes the execution algorithm and uses temporary storage, so its performance impact depends on the atomic pattern and contention. Highly contended accumulations may run faster because deterministic reduction replaces a serialized atomic hotspot with parallel sort-and-reduce work. Sparse accumulations usually slow down because normal atomics have little contention, while deterministic mode still pays for recording and sorting. Counter-based slot allocation also requires counting, scanning, and replaying the kernel to assign stable slots, so it usually slows down. The more conservative GPU_TO_GPU mode can also be slower than RUN_TO_RUN. Benchmark affected kernels before enabling deterministic mode broadly.
The following kernel compacts even thread IDs into an output array. Since each thread uses the return value from wp.atomic_add() as its slot, deterministic mode assigns the same output ordering across repeated launches.
import warp as wp
wp.config.deterministic = wp.DeterministicMode.RUN_TO_RUN
@wp.kernel
def compact_even(count: wp.array[wp.int32], output: wp.array[wp.int32]):
tid = wp.tid()
if tid % 2 == 0:
slot = wp.atomic_add(count, 0, 1)
output[slot] = tid
num_threads = 4_096
count = wp.zeros(1, dtype=wp.int32, device="cuda:0")
output = wp.empty(num_threads // 2, dtype=wp.int32, device="cuda:0")
runs = []
for _ in range(5):
count.zero_()
wp.launch(compact_even, dim=num_threads, inputs=[count], outputs=[output], device="cuda:0")
runs.append(output.numpy())
reference = runs[0].tobytes()
if any(result.tobytes() != reference for result in runs[1:]):
raise RuntimeError("Deterministic launches produced different results")
print(runs[0][:8]) # [ 0 2 4 6 8 10 12 14]Key capabilities:
wp.config.deterministic before defining or importing kernel modules. Use wp.set_module_options() to override an existing module, or module="unique" with module_options for intentional per-kernel isolation.wp.atomic_add(), wp.atomic_sub(), wp.atomic_min(), and wp.atomic_max() updates, sorts them by destination and thread, and combines the updates for each destination with a deterministic reduction instead of applying them in CUDA's scheduler-dependent atomic order. It handles += and -= updates the same way, while counter-based allocation uses count, scan, and replay passes to assign stable output slots.wp.config.deterministic_max_records before defining or importing affected modules. The value is the maximum number of records one thread can emit for one atomic target, such as an output array. The default 0 uses Warp's code-generated estimate. For example, use wp.config.deterministic_max_records = 8 when each thread can execute that atomic at most eight times. A bound that is too small raises an overflow error outside CUDA graph capture. Capture and replay cannot report overflows, so size the bound conservatively. Larger bounds allocate more temporary storage and can increase captured-launch work.wp.atomic_cas(), wp.atomic_exch(), and tile atomics are not rewritten. Deterministic CUDA kernels cannot currently run inside conditional body graphs or APIC serialization. See the deterministic execution guide for the full limits.Important
This is an experimental feature. The API may change without a formal deprecation cycle.
CPU graphs can now replay fills, scans, sorts, run-length encoding, sparse topology updates, adjoint launches, reusable recorded launches, and conditional graph nodes. Saved CUDA .wrp graphs can now serialize and rebuild supported helper operations, conditional nodes, adjoint launches, and reusable record_cmd=True launches (#1431).
The captured scan below reads the current contents of src each time the graph launches. Changing the input produces a new prefix sum without recapturing the graph.
import numpy as np
import warp as wp
src = wp.array(np.ones(8, dtype=np.int32), device="cpu")
prefix = wp.zeros(8, dtype=wp.int32, device="cpu")
with wp.ScopedCapture(device="cpu") as capture:
wp.utils.array_scan(src, prefix, inclusive=True)
wp.capture_launch(capture.graph)
print(prefix.numpy()) # [1 2 3 4 5 6 7 8]
src.fill_(2) # Change the input without recapturing the graph.
wp.capture_launch(capture.graph)
print(prefix.numpy()) # [ 2 4 6 8 10 12 14 16]Notable replay limitations:
int32, float32, int64, and float64 scalar and vector arrays, but not negative strides. Non-contiguous copies and fills, along with the host-return form of wp.utils.runlength_encode(), are unsupported.wp.utils.runlength_encode(), and CUDA wp.sparse.bsr_set_from_triplets() are unsupported. Stream events and texture array copies are not recorded, and only wp.Mesh object handles can be serialized.See the runtime guide sections on CPU Graphs and Saving and Loading Graphs for the complete lists of recorded operations and limitations.
Applications can now allocate CUDA Unified Memory explicitly. These allocations remain CUDA arrays in Warp, while the CUDA driver manages physical page placement and may migrate pages between host memory and CUDA devices as they are accessed. On supported systems, CPU kernels can access them directly. wp.CudaManagedAllocator() works with the existing allocator scopes, and array.memory_kind identifies the allocation as wp.MemoryKind.CUDA_MANAGED (#1523).
import warp as wp
@wp.kernel
def increment(data: wp.array[float]):
i = wp.tid()
data[i] += 1.0
managed = wp.CudaManagedAllocator()
with wp.ScopedAllocator("cuda:0", managed):
data = wp.array([1.0, 2.0, 3.0], dtype=float, device="cuda:0")
if data.memory_kind is not wp.MemoryKind.CUDA_MANAGED:
raise RuntimeError("Expected a CUDA managed-memory allocation")
cpu_data = data if wp.can_access("cpu", data) else data.to("cpu")
wp.launch(increment, dim=data.size, inputs=[cpu_data], device="cpu")
print(cpu_data.numpy()) # [2. 3. 4.]Direct CPU access depends on the system's managed-memory capabilities, so portable code should check wp.can_access(). Allocate managed arrays before CUDA graph capture and reuse them inside the graph; Warp does not currently support allocating managed arrays during capture. Managed arrays cannot be exported through CUDA IPC because CUDA does not provide IPC handles for cudaMallocManaged() allocations.
wp.refUser-defined Warp functions previously received scalar, vector, matrix, quaternion, transform, and struct arguments by value. A helper that changed one of these values had to return it and rely on its caller to assign it back. Functions can now declare wp.ref[T] parameters to update caller-owned storage directly. This removes the return-and-reassign step from reusable helpers, especially when they update several values (#1277).
import warp as wp
@wp.func
def add_in_place(value: wp.ref[float], delta: float):
value += delta
@wp.kernel(enable_backward=False)
def add_ten(values: wp.array[float]):
i = wp.tid()
add_in_place(values[i], 10.0)
values = wp.array([1.0, 2.0, 3.0], dtype=float, device="cpu")
wp.launch(add_ten, dim=values.size, inputs=[values], device="cpu")
print(values.numpy()) # [11. 12. 13.]When calling a function with a wp.ref[T] parameter, pass writable storage such as a local variable, array element, or struct field. Literals, arithmetic results, function return values, and compile-time constants are rejected because there is no original storage to update. wp.address_of(expr) exposes a raw pointer to the same kinds of writable expressions for C++/CUDA snippets; use array.ptr for the base pointer of an entire array. Warp does not automatically differentiate @wp.func helpers that use wp.ref[T], so declare the calling kernel with enable_backward=False. For differentiable code, define the helper with @wp.func_native using a C++/CUDA snippet and provide its derivative in adj_snippet.
User-defined Warp functions can now accept another Warp function as an argument, so one helper can apply different operations without defining a separate wrapper for each one. Annotate the parameter with wp.Function and pass a user-defined @wp.func or a simple builtin such as wp.sin or wp.sqrt (#1424).
import warp as wp
@wp.func
def square(x: float):
return x * x
@wp.func
def apply(function: wp.Function, x: float):
return function(x)
@wp.kernel
def evaluate(output: wp.array[float]):
output[0] = apply(square, 3.0)
output[1] = apply(wp.sqrt, 9.0)
output = wp.empty(2, dtype=float, device="cpu")
wp.launch(evaluate, dim=1, outputs=[output], device="cpu")
print(output.numpy()) # [9. 3.]The target is chosen when Warp generates the kernel, not while it runs. Warp compiles a specialized version of the helper for each distinct target, so using many target combinations can increase compilation work. Defaults and keyword arguments are supported. Arbitrary Python callables and builtins such as wp.printf are not valid targets. A helper with a wp.Function parameter cannot also use @wp.func_grad or @wp.func_replay.
Performance-sensitive CUDA kernels can now opt out of Warp's grid-stride loop. @wp.kernel(grid_stride=False) uses a lean launch path with one thread per work item, reducing loop-control overhead and potentially lowering register pressure for lightweight or bandwidth-bound kernels (#1270).
Use the "default_grid_stride": False module option or set wp.config.default_grid_stride = False to opt in more broadly. Grid-stride launches remain the default. Lean kernels support launches of any practical size, but they cannot use max_blocks to cap the block count. Passing max_blocks > 0 raises an error.
@wp.kernel(cluster_dim=...) and wp.get_cuda_max_cluster_dim() expose CUDA Thread Block Clusters to native functions on cluster-capable GPUs (#1401).wp.utils.array_scan() now accepts 64-bit scalar and vector types. wp.utils.radix_sort_pairs() accepts 32-bit and 64-bit signed, unsigned, and floating-point keys with 4-byte or 8-byte values (#1538).Important
The texture API is experimental and subject to change without a formal deprecation cycle. Pass texture constructor options by keyword when possible.
Textures can now store a full level-of-detail chain and sample a chosen resolution from Warp kernels. Set num_mip_levels=0 to generate the full chain, or request a fixed number of levels, then pass lod to wp.texture_sample() (#1409).
import numpy as np
import warp as wp
@wp.kernel
def sample_levels(texture: wp.Texture2D, output: wp.array[float]):
level = wp.tid()
output[level] = wp.texture_sample(
texture,
wp.vec2(0.125, 0.125),
dtype=float,
lod=float(level),
)
pixels = np.arange(16, dtype=np.float32).reshape(4, 4)
texture = wp.Texture2D(data=pixels, num_mip_levels=0, device="cpu")
num_levels = texture.num_mip_levels
output = wp.empty(num_levels, dtype=float, device="cpu")
wp.launch(sample_levels, dim=num_levels, inputs=[texture], outputs=[output], device="cpu")
print(output.numpy()) # [0. 2.5 7.5]Each output element samples the same normalized coordinate at LOD 0, 1, and 2, so [0.0, 2.5, 7.5] shows how the sampled value changes across the mip chain.
Warp builds lower levels from construction data with a 2x box filter. Mipmapped textures are immutable after construction, cannot wrap an external CUDA array, and cannot use surface access. CUDA mipmapped textures also require normalized coordinates.
Important
This is an experimental feature. The API may change without a formal deprecation cycle.
Warp 1.15 expands the experimental cuBQL BVH backend with direct wp.Bvh construction and broader query support. Select it with constructor="cubql" for wp.Bvh or bvh_constructor="cubql" for wp.Mesh. The resulting structures support AABB, closest-point, furthest-point, and ray queries through Warp's existing query APIs (#1467).
Grouped BVHs and meshes are not supported, and wp.Mesh winding-number queries remain unavailable. A cuBQL BVH also cannot switch to or from another constructor during an in-place rebuild.
Tiles can now contain Warp struct values, so tile kernels no longer need a separate array for each struct field. wp.tile_map() can process struct elements and broadcast non-tile struct arguments. Struct tiles support field-wise addition and subtraction, reductions, atomic addition, gradients, and use as wp.tile_sort() value payloads (#573).
import warp as wp
TILE_SIZE = wp.constant(4)
@wp.struct
class Particle:
mass: float
velocity: wp.vec3
@wp.func
def double_mass(particle: Particle) -> Particle:
particle.mass *= 2.0
return particle
@wp.kernel
def update(input_particles: wp.array[Particle], output_particles: wp.array[Particle]):
tile = wp.tile_load(input_particles, shape=TILE_SIZE)
wp.tile_store(output_particles, wp.tile_map(double_mass, tile))
particles = []
for mass in (1.0, 2.0, 3.0, 4.0):
particle = Particle()
particle.mass = mass
particles.append(particle)
input_particles = wp.array(particles, dtype=Particle, device="cuda:0")
output_particles = wp.empty_like(input_particles)
wp.launch_tiled(
update,
dim=1,
inputs=[input_particles],
outputs=[output_particles],
block_dim=32,
device="cuda:0",
)
print(output_particles.numpy()["mass"]) # [2. 4. 6. 8.]wp.tile_sort() treats structs as value payloads rather than keys, and sorting is forward-only for struct values. Struct tiles do not define a canonical wp.tile_ones() value, bitwise in-place operations, or ordering reductions.
Sparse assembly can now reserve a bounded number of blocks per row, update topology without reallocating the whole matrix, and compact candidate entries in place. warp.sparse.bsr_zeros(row_capacity=...) creates padded storage, bsr_compress() sorts and coalesces active entries, and matrix status reports whether an operation exceeded its row capacity (#1537).
import warp as wp
from warp.sparse import BSR_STATUS_SUCCESS, bsr_compress, bsr_set_from_triplets, bsr_zeros
A = bsr_zeros(2, 3, float, device="cpu", row_capacity=2)
bsr_set_from_triplets(
A,
rows=wp.array([0, 0, 1, 1], dtype=int, device="cpu"),
columns=wp.array([0, 2, 1, 1], dtype=int, device="cpu"),
values=wp.array([2.0, 3.0, 4.0, -1.0], dtype=float, device="cpu"),
topology="padded",
)
print(A.row_counts.numpy())
print(A.status_sync() == BSR_STATUS_SUCCESS)
bsr_compress(A, inplace=True, topology="compact")
print(A.nnz_sync())Output:
[2 1]
True
3
The inplace=True compression path supports float32 and float64 scalar matrices and is not differentiable. The default out-of-place path remains differentiable. warp.fem.integrate() and warp.fem.interpolate() can use in-place compression during matrix assembly to reduce peak memory use.
Warp's source tree now includes ready-to-use CMake configuration for parallel and incremental local builds of the native warp and warp-clang libraries (#1495). This optional developer workflow does not build or replace Warp's Python wheels, which remain the primary way to install Warp. See the CMake build documentation.
Source builds can instrument native libraries and JIT-compiled CPU kernels with AddressSanitizer (#1387):
uv run build_lib.py --sanitize=addressRun the target kernel in release mode so Warp's debug bounds assertion does not fire before AddressSanitizer. Linux processes must preload the AddressSanitizer runtime before importing Warp. Standard pip installations are unchanged.
wp.copy() validates offsets, counts, and element sizesWarp now rejects non-integer or negative offsets, a negative count, and contiguous copies whose source and destination element sizes differ. Warp v1.14 treated a negative count as a request to copy the whole array, which could hide mistakes (#1584).
import warp as wp
source = wp.array([1, 2, 3, 4], dtype=wp.int32, device="cpu")
destination = wp.zeros_like(source)
# Warp v1.14 treated a negative count as "copy all".
# Warp v1.15 rejects it before any pointer arithmetic.
try:
wp.copy(destination, source, count=-1)
except RuntimeError as error:
print(error)
# Migration: pass the intended non-negative element count explicitly.
wp.copy(destination, source, count=source.size)
print(destination.numpy())Output:
Copy count must be non-negative, got -1
[1 2 3 4]
Warp now stops initialization with RuntimeError when the Python package version does not match warp.dll, warp-clang.dll, or their platform equivalents. Warp previously warned and continued, which could load an incompatible native ABI. Rebuild or reinstall Warp so the native libraries match the Python package (#1508).
We are tentatively planning to build Warp's PyPI wheels with a CUDA 13.x Toolkit starting with Warp 1.17, currently targeted for September 2026. CUDA Toolkit 13.0 was released in August 2025, more than a year before that target, and we recommend moving to an R580-series or newer NVIDIA driver in advance. If this plan proceeds, R580 will be the minimum driver branch for using Warp's CUDA backend from those wheels. pip does not check the installed driver and can install the wheels on systems with older drivers, but the CUDA backend will not work on those systems. The driver must support the CUDA 13.x family, but it does not need to match the exact CUDA 13.x Toolkit version used to build Warp. See NVIDIA's CUDA minor-version compatibility table.
CUDA 12 support will continue beyond Warp 1.17. After the PyPI wheels move to CUDA 13.x, users who need CUDA 12 can build Warp from source or install a pre-built CUDA 12 wheel from GitHub Releases.
warp.fem arguments have been removed. Pass a warp.fem.Quadrature or warp.fem.GeometryDomain through at instead of the old quadrature or domain arguments to warp.fem.interpolate(). Use space_topology for partitions and space_topology or space_partition for restrictions.warp.render.UsdRenderer.update_body_transforms() method has been removed. It referenced attributes that UsdRenderer never defined and could not update USD transforms. Use the supported shape-rendering methods to author time-varying USD transforms instead.wp.MarchingCubes sizing arguments and compatibility attributes are deprecated and scheduled for removal in Warp 1.19. The max_verts and max_tris values are retained only for compatibility and do not limit the dynamically sized output arrays. Likewise, device does not select the extraction device. Extraction always uses the input field's device. id no longer identifies a native resource, while runtime is not used during extraction. Neither has a replacement. Stop passing max_verts, max_tris, or device to the constructor, stop passing max_verts or max_tris to resize(), and remove accesses to these compatibility attributes (#1594).masked=True in warp.sparse topology-changing operations is deprecated. Pass topology="masked" instead. The old argument will be removed in a future feature release under the standard deprecation timeline (#1537).We also thank the following contributors from outside the core Warp development team:
_warp_fastcall crash, which informed the removal of fastcall support (#1578).For a complete list of changes, see the full changelog.
wp.config.deterministic or per module
via the "deterministic" module option, using wp.DeterministicMode.RUN_TO_RUN for bit-exact repeatable
results on the same GPU, or wp.DeterministicMode.GPU_TO_GPU to also target consistency across GPU
architectures. It covers the two common atomic patterns: accumulating values into array elements, and using
an atomic counter to hand out unique slot indices. It extends to the reductions Warp generates for backward
passes and custom adjoints. Counters used for slot allocation cannot be replayed automatically during a
backward pass, so Warp raises an error rather than risk misplaced gradients; supply a custom replay function
to differentiate through them. Raise wp.config.deterministic_max_records if a kernel needs a larger
per-thread record limit (GH-1443).wp.ref[T] parameter annotation for @wp.func and @wp.func_native, enabling explicit pass-by-reference
semantics for scalar, vector, matrix, quaternion, and struct types. Mutations to a wp.ref[T] parameter are visible
in the caller's storage without a return value. Add wp.address_of(expr) -> wp.uint64 for obtaining a raw pointer
to an addressable expression for use with native snippets, and allow @wp.func_native ref-parameter functions to
provide manual adjoints with adj_snippet (GH-1277).wp.Function-typed parameters in user-defined @wp.func functions, allowing user-defined Warp
functions and simple built-in Warp functions such as wp.sin() and wp.min() to be used as function targets from
kernels or other functions, including through defaults (GH-1424).wp.Texture1D, wp.Texture2D, and wp.Texture3D
via the new num_mip_levels and mip_filter_mode constructor parameters, and allow wp.texture_sample() to accept
an optional trailing lod argument for controlling sampled detail level
(GH-1409).wp.CudaManagedAllocator() for explicit CUDA managed-memory arrays. CPU kernels can use managed arrays as an
opt-in path to read and write CUDA managed-memory allocations through Unified Memory on systems where CUDA reports
compatible managed-memory access, while Warp CUDA arrays backed by non-managed memory still need explicit CPU
copies. Use wp.array.memory_kind to inspect whether an array is backed by host, pinned host, CUDA device, CUDA
mempool, or CUDA managed memory. Preallocated managed arrays work in CUDA graph captures, but capture-time
allocations remain unsupported (GH-1523).wp.launch(..., record_cmd=True) from current inputs. Also replay CUDA conditional graph nodes created with
wp.capture_if() and wp.capture_while(), as well as adjoint launches
(GH-1431).@wp.kernel(grid_stride=False) (and the "default_grid_stride" module option / wp.config.default_grid_stride
for a whole module or process) to compile a kernel without the grid-stride loop. Removing the loop lowers per-thread
overhead and register pressure, which can speed up launch-bound kernels; it still handles launches of any size, but
the block count can no longer be capped, so launching such a kernel with max_blocks > 0 raises an error pointing
back at the grid-stride default. The default launch is unchanged
(GH-1270).wp.tile_map() to process struct elements and broadcast
non-tile struct arguments; support field-wise addition, subtraction, reductions, atomic addition, and gradients
through compatible data-movement and additive operations; and allow wp.tile_sort() to reorder structs as value
payloads in the forward pass (GH-573).constructor="cubql" support to wp.Bvh and allow wp.Mesh(..., bvh_constructor="cubql") to
use AABB, point, furthest-point, and ray queries when Warp is built with cuBQL. Grouped BVHs/meshes and mesh
winding-number queries remain unsupported (GH-1467).warp.sparse.BsrMatrix, including padded topology policies for topology-changing
operations, warp.sparse.bsr_zeros(row_capacity=...) for reserving capacity, warp.sparse.bsr_compress() for
in-place compaction, and warp.sparse.BSR_STATUS_SUCCESS and warp.sparse.BSR_STATUS_ROW_CAPACITY_EXCEEDED
for status checks (GH-1537).warp.fem.integrate() and warp.fem.interpolate() to assemble matrices in place via
warp.sparse.bsr_compress() by passing bsr_options={"construction": "row_compress"}, reducing peak
memory usage (GH-1537).wp.utils.array_scan() to 64-bit scalar and vector types, and extend wp.utils.radix_sort_pairs() to 32- and
64-bit signed, unsigned, and floating-point keys with 4- or 8-byte values
(GH-1538).cluster_dim option to @wp.kernel and a wp.get_cuda_max_cluster_dim() query helper for using CUDA Thread
Block Clusters from @wp.func_native code (GH-1401).warp and warp-clang with Packman-managed dependencies and support for
parallel and incremental developer builds (GH-1495).build_lib.py --sanitize=address. These builds report out-of-bounds wp.array accesses as heap-buffer-overflow.
Standard pip installations are unaffected (GH-1387).warp.fem APIs: quadrature and domain from warp.fem.interpolate(), and
space from warp.fem.make_space_partition() and warp.fem.make_space_restriction(). Pass a
warp.fem.Quadrature or warp.fem.GeometryDomain to at when interpolating. Use space_topology for partitions
and either space_topology or space_partition for restrictions.warp.render.UsdRenderer.update_body_transforms() method. Use supported
warp.render.UsdRenderer shape-rendering methods to author time-varying USD transforms instead.max_verts, max_tris, and device arguments to wp.MarchingCubes, the max_verts and
max_tris arguments to wp.MarchingCubes.resize(), and the corresponding compatibility attributes, id, and
runtime. They are scheduled for removal in Warp 1.19. Output arrays are dynamically sized, and extraction uses the
input field's device. id and runtime have no replacement (GH-1594).masked=True arguments in warp.sparse topology-changing operations; use topology="masked" instead
(GH-1537).src_offset, dest_offset, and count arguments to wp.copy(), raising a clear error
for non-integer values, negative offsets, and a negative count (previously a negative count silently copied the
entire array). Also reject copies between contiguous arrays whose element sizes differ, matching the existing
behavior for non-contiguous arrays (GH-1584).wp.init() when the Warp Python package version does not match the loaded native
libraries (warp.dll, warp-clang.dll, and platform equivalents), instead of warning and continuing. Treat missing
or unreadable native version symbols the same way (GH-1508).Python.h and compatibility with free-threaded Python by
reverting the wp.float16 fast-call optimization added in Warp 1.14. Python-scope wp.float16 conversions again use
the ctypes path, with higher per-call overhead (GH-1339).@wp.kernel(module="unique") kernels created by factory
functions, by avoiding repeated source inspection and redundant scans of each kernel's Python code. This reduces
import time without changing native compilation or execution (GH-1486).wp.mesh_query_ray() and wp.mesh_query_ray_anyhit() BVH traversal performance by visiting the nearer child
first at each inner node, enabling earlier tightening of the closest-hit bound and more aggressive subtree pruning;
reusing loaded node payloads across traversal stack operations; and using a faster AABB intersection path for
non-parallel rays. These traversal changes use additional per-thread temporary storage
(GH-1529).wp.init() segfault while loading warp-clang.so on some Linux systems caused by aggressive symbol stripping
(GH-1554).RuntimeError instead of only logging native CUDA stderr and
continuing with stale outputs or gradients (GH-1535).wp.constant(wp.float64(...)), when the constant is the leading operand of a multiplication or division expression
(GH-1540).wp.copy() when src_offset and dest_offset differ
(GH-1533).wp.copy() to honor src_offset, dest_offset, and count for one-dimensional non-contiguous arrays such as
strided slices (GH-1533).@wp.struct plain-array field with None or a non-gradient array to release stale gradient
references (GH-1520).wp.mesh_query_ray() and related functions silently missing intersections for axis-aligned rays whose origin lies
on a BVH node's AABB slab boundary. Handle parallel slabs explicitly instead of relying on fixed epsilon inflation
that can round away at larger wp.float32 coordinate scales (GH-1530).wp.force_load(max_workers > 1) that could cause
modules sharing a @wp.func to fail to compile or raise a KeyError during module loading
(GH-1474, GH-1532).module="unique" kernels, capture-time allocations, untracked host buffers,
coexisting CPU and CUDA captures, and warp.fem partition node counts. Report capture state correctly and reject
operations that cannot replay safely
(GH-1431,
GH-1559, GH-1407).wp.load_module() and wp.ScopedCapture(force_module_load=True) after CPU launches so CUDA graph capture
precompiles the correct CUDA kernel variant. On CUDA drivers older than 12.3 this previously raised
CUDA_ERROR_STREAM_CAPTURE_UNSUPPORTED; on newer drivers it silently recompiled inside the capture window
(GH-564).wp.array.fill_() calls on wp.bool, wp.int8, and wp.uint8 arrays
(GH-1525).f = my_func, f = module.my_func,
assigning a wp.static(...) function result to f, and wp.grad(f) where f is such a local now compile and call
correctly instead of failing during code generation. Rebinding a function-valued local to a different function or
to a non-function value raises a clear error (GH-1423).wp.tile_dot() compilation failure for scalar wp.float64 and wp.float16 tiles, which previously narrowed
the result to float and mismatched the inferred result type (GH-1563).@wp.func(...) decorators, such as @wp.func(module="unique"), in directly executed scripts to
register successfully instead of raising AttributeError during decoration
(GH-1544).@wp.overload kernel stubs defined in nested scopes to register correctly instead of raising
IndentationError (GH-1557).wp.array[...]-style subscript annotations to allow for PEP 604 unions (for example, wp.array[float] | float)
on Python methods (GH-1548).wp.Tape.record_scope_end() to raise a clear error for unmatched scope ends and preserve nested non-empty tape
visualization scopes (GH-1515).None return annotations to raise a clear error instead of silently ignoring the annotation.
Write results to output arguments and omit the return annotation or use -> None
(GH-1471).wp.grad() to a non-gradient value to raise a clear code-generation
error instead of an internal AttributeError (GH-1487).block_dim=1, along with
portable workarounds (GH-1580).…so application warning filters can suppress Warp deprecation warnings again. wp.config.verbose and wp.config.quiet still work during the deprecation w…
Warp v1.14 expands serialized CPU capture support: captured graphs can now include backward launches, tiled kernels, richer launch arguments, and .wrp files with arrays nested inside structs or wp.indexedarray arguments. This release also adds multi-environment FEM support for batched simulations, reusable and batched linear solvers, pluggable logging, portable tile FFT and solver fallbacks, stable JAX integration APIs, and relaxed CPU/GPU array access for Heterogeneous Memory Management (HMM) and Address Translation Services (ATS) systems.
Building on the initial API Capture serialization support in Warp v1.13, Warp v1.14 primarily broadens the set of CPU graph patterns that can be saved and replayed. CPU captures can now include forward execution and reverse-mode passes from wp.Tape().backward(), wp.launch_tiled() kernels, and scalar parameters of any size (#1431). The shared .wrp serialization format also now supports @wp.struct arguments that contain arrays and wp.indexedarray arguments that carry data, gradient, and index buffers.
Important
Upgrade impact for APIC users:
.wrp files saved by Warp v1.13. Warp v1.14 writes APIC format version 10 and rejects the previous format.APICState* and APICGraph*. Ownership and destroy calls are unchanged. See the APIC migration diff.Saved APIC graphs can still be consumed from standalone C++ through the C API declared in warp/native/apic.h. Native replay behavior is unchanged apart from the explicit pointer spelling for APIC handles.
Key capabilities:
wp.Tape().backward() are recorded into the CPU APIC stream and replayed from live captures or loaded graphs.wp.indexedarray arguments, and handles inside serialized launch value blobs.Known limitations:
wp.utils.array_scan() is still not recorded into CPU APIC and raises NotImplementedError in CPU capture.array.fill_() operations on CPU are not recorded.wp.Mesh handles, but wp.Volume and wp.Bvh handles are not yet supported..wrp graphs requires the warp-clang backend and the companion _modules/ directory with the recorded CPU kernel objects.warp.femwarp.fem can now represent many independent simulation environments inside one geometry and one solve setup (#1407). Colocated Grid2D and Grid3D geometries expose an env_count, sparse Nanogrid and AdaptiveNanogrid geometries can pack per-environment voxels into one NanoVDB volume, and unstructured meshes can carry per-cell environment metadata through cell_env and env_count.
This feature changes two positional call signatures. See the FEM migration diff if your code passes requires_grad, device, or temporary_store positionally.
import warp as wp
import warp.fem as fem
geo = fem.Grid3D(res=(8, 8, 8), bounds_lo=wp.vec3(0.0), bounds_hi=wp.vec3(1.0), env_count=4)
pressure_space = fem.make_polynomial_space(geo, degree=0, discontinuous=True)
partition = fem.make_space_partition(space_topology=pressure_space.topology, environment_first=True)
# For scalar pressure spaces, these node offsets can be passed directly to
# warp.optim.linear.LinearOperator(batch_offsets=...).
pressure_batch_offsets = partition.env_offsetsEnvironment-aware lookup keeps colocated environments from interacting accidentally. When a geometry has more than one environment, pass an environment index to fem.lookup() and pass env_indices to fem.PicQuadrature when particles are binned from world-space positions. The new warp/examples/fem/example_apic_fluid_multi_env.py example uses these APIs to run colocated APIC fluid environments with environment-aware particle quadrature and batched pressure solves.
Key capabilities:
Nanogrid.from_environment_voxels() and AdaptiveNanogrid.from_environment_voxels() build packed sparse grids with per-cell cell_env metadata and hidden offsets for packed grids.cell_env and env_count so grouped BVH lookup only traverses the requested environment.make_space_partition(..., environment_first=True) exposes env_offsets that line up with the new linear-solver batch_offsets support for scalar spaces.environment_first=True does not support halo nodes for partitions that do not cover a whole geometry. Mesh environment indices are lookup and partition metadata, so callers must still provide disconnected mesh topology for independent mesh environments.The iterative solvers in warp.optim.linear can now preallocate their temporary buffers and reuse them across compatible solves (#1391). Passing run=False to cg(), cr(), bicgstab(), or gmres() returns a solver object that can be called repeatedly with replacement operands that have the same shape, dtype, device, and batch layout.
import numpy as np
import warp as wp
from warp.optim.linear import aslinearoperator, cg
diag = wp.array([2.0, 2.0, 5.0, 5.0], dtype=float, device="cpu")
b = wp.array([2.0, 4.0, 10.0, 15.0], dtype=float, device="cpu")
x = wp.zeros_like(b)
offsets = wp.array(np.array([0, 2, 4], dtype=np.int32), dtype=int, device="cpu")
A = aslinearoperator(diag, batch_offsets=offsets)
state = cg(A, b, x, maxiter=100, run=False)
state() # solve the original system
b2 = wp.array([4.0, 2.0, 15.0, 10.0], dtype=float, device="cpu")
x2 = wp.zeros_like(b2)
state(b=b2, x=x2) # reuse temporary buffers for a compatible systembatch_offsets is independent of solver-state reuse. It partitions a LinearOperator into scalar degree-of-freedom intervals that are solved as independent subproblems in one solver launch sequence, with convergence checked per batch and reported through the worst residual. For one-shot solves, call the solver directly as before.
Warp's Python-side diagnostics now flow through a configurable logger (#1315, #1434). Install a custom logger with wp.set_logger(), scope it with wp.ScopedLogger, and control verbosity through wp.config.log_level or wp.ScopedLogLevel.
wp.config.log_level accepts the standard Warp log-level constants:
wp.LOG_DEBUG: most verbose, including code-generation details and module loads.wp.LOG_INFO: the default, including the init banner and compile timings.wp.LOG_WARNING: warnings and errors only.wp.LOG_ERROR: errors only.import warp as wp
wp.config.log_level = wp.LOG_WARNING
class ListLogger:
def __init__(self):
self.records = []
def debug(self, message):
self.records.append(("debug", message))
def info(self, message):
self.records.append(("info", message))
def warning(self, message, category=None, stacklevel=1):
self.records.append(("warning", message))
def error(self, message):
self.records.append(("error", message))
logger = ListLogger()
with wp.ScopedLogger(logger), wp.ScopedLogLevel(wp.LOG_INFO):
wp.get_logger().info("captured by application logger")
print(logger.records) # [('info', 'captured by application logger')]The default Warp logger now routes warnings emitted by Warp's Python code through Python's warnings.warn(), so application warning filters can suppress Warp deprecation warnings again. wp.config.verbose and wp.config.quiet still work during the deprecation window, but they now emit one-time DeprecationWarnings and map to wp.config.log_level.
Warp no longer enforces the old rule that every Warp array passed to a kernel must be allocated on the launch device. wp.config.launch_array_access_mode now defaults to wp.config.LaunchArrayAccessMode.RELAXED, so CUDA kernels can receive CPU arrays on hardware where CUDA reports pageable CPU memory as GPU-accessible (#1461). The main targets are Linux Heterogeneous Memory Management (HMM) systems and NVIDIA Address Translation Services (ATS) platforms such as GH200, GB200, DGX Spark / GB10, and Jetson Thor, where ordinary CPU allocations can be GPU-accessible without an explicit .to(device) copy. Use wp.can_access() for concrete arrays and Device.can_access() for coarse device checks when code needs to choose a direct launch or an explicit copy at runtime.
import warp as wp
@wp.kernel
def increment(data: wp.array[float]):
i = wp.tid()
data[i] = data[i] + 1.0
device = wp.get_device("cuda:0")
cpu_data = wp.empty(1024, dtype=float, device="cpu")
if wp.can_access(device, cpu_data):
wp.launch(increment, dim=cpu_data.size, inputs=[cpu_data], device=device)
else:
gpu_data = cpu_data.to(device)
wp.launch(increment, dim=gpu_data.size, inputs=[gpu_data], device=device)Choose the validation mode based on how much pre-launch checking you want:
wp.config.LaunchArrayAccessMode.RELAXED is the default. It checks type, dtype, and dimensions, but does not reject cross-device array arguments before launch. If a GPU kernel dereferences CPU memory that the device cannot access, the failure surfaces as a CUDA runtime error instead of Warp's previous Python same-device error.wp.config.LaunchArrayAccessMode.CHECKED raises a Python error before launch when Warp can prove that an array is not accessible from the launch device. It warns and proceeds for custom or externally wrapped allocations whose provenance Warp cannot verify.wp.config.LaunchArrayAccessMode.STRICT restores the old same-device rule and requires every Warp array argument to be allocated on the launch device.Most users can keep the default. Use wp.config.LaunchArrayAccessMode.CHECKED when diagnosing mixed-device launches, and use wp.config.LaunchArrayAccessMode.STRICT in tests or libraries that depend on pre-launch same-device validation.
CUDA graph capture now exposes CUDA's stream-capture mode through the capture_mode keyword on wp.ScopedCapture and wp.capture_begin() (#1410). Choose the mode based on how strictly CUDA should reject capture-unsafe runtime calls while capture is active:
wp.CaptureMode.THREAD_LOCAL is the default and preserves Warp's historical behavior. Capture-unsafe runtime calls from the capturing thread invalidate the capture, while other threads are unaffected.wp.CaptureMode.GLOBAL is the strictest mode. Capture-unsafe runtime calls from any thread invalidate the capture.wp.CaptureMode.RELAXED tolerates capture-unsafe runtime calls. Use it when composing with libraries that may lazily initialize CUDA contexts or allocators during capture.import warp as wp
@wp.kernel
def add_one(x: wp.array[float]):
i = wp.tid()
x[i] = x[i] + 1.0
x = wp.zeros(1024, dtype=float, device="cuda:0")
with wp.ScopedCapture(device="cuda:0", capture_mode=wp.CaptureMode.RELAXED) as capture:
wp.launch(add_one, dim=x.size, inputs=[x], device="cuda:0")
wp.capture_launch(capture.graph)The function form uses the same capture_mode keyword, for example wp.capture_begin(device="cuda:0", capture_mode=wp.CaptureMode.RELAXED).
wp.tile_fft() and wp.tile_ifft() now run on CPU and on GPU builds that do not include libmathdx (#1396). CPU supports any power-of-two FFT length and non-power-of-two lengths up to 4096 elements. The GPU fallback, selected automatically when libmathdx is unavailable or explicitly with enable_mathdx_fft=False, supports power-of-two FFT lengths divisible by block_dim.
import warp as wp
@wp.kernel
def filter_signal(x: wp.array2d[wp.vec2f], y: wp.array2d[wp.vec2f]):
t = wp.tile_load(x, shape=(64, 64))
wp.tile_fft(t)
# Apply a spectral filter here.
wp.tile_ifft(t)
wp.tile_store(y, t)GPU scalar fallbacks are also available for wp.tile_cholesky(), wp.tile_cholesky_solve(), wp.tile_lower_solve(), and wp.tile_upper_solve(), including in-place variants and the wp.tile_cholesky() adjoint (#1402). Select them with wp.config.enable_mathdx_solver=False or module_options={"enable_mathdx_solver": False}. The fallback avoids a libmathdx dependency and reduces compile cost, at the expense of runtime performance. One fallback limitation is that differentiated wp.tile_cholesky() allocates extra per-block scratch storage, so large GPU tiles can exceed the device's shared-memory budget. If you hit that limit, reduce the Cholesky tile size or dtype, or use a libmathdx-enabled build with wp.config.enable_mathdx_solver=True.
wp.tile_empty()wp.tile_empty() allocates an uninitialized register or shared-memory tile for kernels that overwrite every element before the first read (#1312). Use it instead of wp.tile_zeros() for full overwrites, because skipping zero-fill work can improve performance. Keep wp.tile_zeros() for accumulators or partial writes where the initial zeros are part of correctness.
Several tile fixes improve correctness in edge cases:
wp.tile_load(), wp.tile_store(), and indexed tile operations now use 64-bit byte offsets, fixing overflows on arrays larger than 2 GiB (#1422). This correctness fix widens address arithmetic in the tile load/store paths, so tile-heavy kernels may see slightly higher address-calculation overhead.wp.tile_matmul() kernels and their adjoints (#1439, #1440).wp.tile_matmul() now rejects wp.bfloat16 output tiles. It also rejects wp.bfloat16 input tiles when backward compilation is enabled, because the backend cannot use bfloat16 accumulators for those paths (#1427).The JAX integration has been promoted from warp.jax_experimental into Warp's stable public API (#1370). New code should import warp.jax_kernel, warp.jax_callable, warp.clear_jax_callable_graph_cache, warp.JaxCallableGraphMode, and warp.JaxModulePreloadMode directly from the top-level warp namespace.
- from warp.jax_experimental import GraphMode, ModulePreloadMode, clear_jax_callable_graph_cache
- from warp.jax_experimental import jax_callable, jax_kernel
+ from warp import JaxCallableGraphMode, JaxModulePreloadMode, clear_jax_callable_graph_cache
+ from warp import jax_callable, jax_kernelAs part of the promotion, warp.jax_experimental is now a deprecated compatibility namespace and will be removed in Warp 1.16. The warp.jax_experimental.get_jax_callable_default_graph_cache_max() and warp.jax_experimental.set_jax_callable_default_graph_cache_max() helpers are also deprecated. Pass graph_cache_max to warp.jax_callable() or update the returned callable's graph_cache_max attribute instead. Top-level warp.jax_callable() defaults to graph_cache_max=32. Pass graph_cache_max=None for an unlimited graph cache.
Differentiable warp.jax_kernel() wrappers now accept launch_dims together with enable_backward=True (#1380). The dimensions are fixed at wrapper construction and reused for both the forward and adjoint launches, which is useful when the input array includes batch or channel dimensions that are not part of the kernel's wp.tid() iteration space.
# API shape example. Requires JAX installed.
import warp as wp
from warp import jax_kernel
@wp.kernel
def spatial_update(x: wp.array4d[float], out: wp.array4d[float]):
i, j, k = wp.tid()
out[0, i, j, k] = x[0, i, j, k] * 2.0
jax_update = jax_kernel(
spatial_update,
num_outputs=1,
launch_dims=(16, 16, 16),
enable_backward=True,
)When enable_backward=True, launch_dims cannot be overridden per call, and output_dims remains unsupported.
Source builds gain two new build options. build_lib.py --use-dynamic-cuda links Warp's native library against shared CUDA libraries instead of embedding them statically, for deployments that already provide the matching CUDA shared libraries at runtime (#1334). build_lib.py --sanitize=address builds the native libraries with AddressSanitizer instrumentation, a compiler/runtime memory-error detector for native out-of-bounds accesses, use-after-free, double-free, and similar bugs (#1387). Use it for debugging source builds when you want a failing test or repro to report the invalid memory access closer to where it happens.
uv run build_lib.py --use-dynamic-cuda
uv run build_lib.py --sanitize=address --mode debugOn Linux, install Python development headers if a source build fails with a missing Python.h, for example python3-dev or libpython3-dev on Debian and Ubuntu (#1339). Warp now compiles a small Python C API extension for faster wp.float16 conversion to and from Python float. The extension uses CPython's vectorcall protocol, the public fast-call convention introduced by PEP 590. Linux and macOS debug builds now use -Og -g, which keeps debug information while preserving useful compiler diagnostics such as uninitialized-value analysis (#1414).
Floating-point wp.min(), wp.max(), wp.clamp(), wp.atomic_min(), and wp.atomic_max() now use NaN-as-missing semantics matching C fmin() and fmax() (#1376). When exactly one operand is NaN, the non-NaN operand wins. Vector reductions and wp.argmin() / wp.argmax() skip NaN slots. Adjoint and atomic variants route gradients to the operand chosen by the forward pass.
import numpy as np
import warp as wp
@wp.kernel
def math_kernel(values: wp.array[float], out: wp.array[float]):
out[0] = wp.min(values[0], 2.0)
out[1] = wp.max(values[0], -3.0)
out[2] = wp.copysign(3.0, values[1])
values = wp.array([np.nan, -0.0], dtype=float, device="cpu")
out = wp.zeros(3, dtype=float, device="cpu")
wp.launch(math_kernel, dim=1, inputs=[values], outputs=[out], device="cpu")
print(out.numpy()) # [ 2. -3. -3.]Other math and autodiff fixes are grouped by affected path:
wp.copysign() is now available in kernels (#1444).wp.curlnoise() in 2D, 3D, and 4D, so differentiable curl-noise force fields no longer produce zero gradients (#1012).wp.array, such as arr[i].y = rhs, m[i][r, c] = rhs, transform .p / .q writes, and scalar or composite struct-field writes, now propagate gradients (#583, #248, #1174).wp.closest_point_edge_edge() is more reliable for near-parallel float32 segments (#1437). Well-conditioned inputs keep the same output, while near-parallel cases now use a more stable closest-point computation and a bounded analytic adjoint.wp.array.fill_() now passes fill values up to 3968 bytes inline through kernel arguments on both non-contiguous and contiguous array fill paths (#1412). This fixes forked-stream capture failures for scalar, vector/matrix, and many struct fills. Larger fill values still use the previous temporary-storage fallback and are not guaranteed to compose in the same forked-stream capture scenario.module="unique" kernel declarations that depend on other Warp functions no longer retain stale Python module references (#1462), and autodiff metadata is now cleaned up for non-differentiable builtins (#988, #1466).The component-write fix means the example below now gives x.grad == [2, 2, 2] because the adjoint crosses the stored out[i].y component instead of stopping at the component assignment.
import numpy as np
import warp as wp
@wp.kernel
def write_y(x: wp.array[float], out: wp.array[wp.vec3]):
i = wp.tid()
out[i].y = 2.0 * x[i]
@wp.kernel
def sum_y(out: wp.array[wp.vec3], loss: wp.array[float]):
i = wp.tid()
wp.atomic_add(loss, 0, out[i].y)
x = wp.array([1.0, 2.0, 3.0], dtype=float, device="cpu", requires_grad=True)
out = wp.zeros(x.size, dtype=wp.vec3, device="cpu", requires_grad=True)
loss = wp.zeros(1, dtype=float, device="cpu", requires_grad=True)
with wp.Tape() as tape:
wp.launch(write_y, dim=x.size, inputs=[x, out], device="cpu")
wp.launch(sum_y, dim=x.size, inputs=[out, loss], device="cpu")
tape.backward(loss)
np.testing.assert_allclose(x.grad.numpy(), [2.0, 2.0, 2.0])Warp v1.14 writes APIC format version 10. .wrp files captured by Warp v1.13 used the older format and must be recaptured. Native C/C++ code that includes apic.h must update APIC handles from the old typedef form, which hid the pointer, to explicit pointers. Ownership and destroy calls are unchanged. The migration is a source-level spelling change:
- APICState state = wp_apic_create_state();
+ APICState* state = wp_apic_create_state();
wp_apic_destroy_state(state);
- APICGraph graph = wp_apic_load_graph(context, "simulation", APIC_DEVICE_CUDA);
+ APICGraph* graph = wp_apic_load_graph(context, "simulation", APIC_DEVICE_CUDA);
wp_apic_destroy_graph(graph);warp.fem.PicQuadrature() now accepts env_indices before requires_grad, and warp.fem.make_space_partition() now accepts environment_first before the keyword-only device and temporary_store. Pass affected arguments by keyword to preserve behavior:
- q = fem.PicQuadrature(domain, particles, measures, True)
+ q = fem.PicQuadrature(domain, particles, measures, requires_grad=True)
- p = fem.make_space_partition(space_topology, geometry_partition, True, -1, device)
+ p = fem.make_space_partition(
+ space_topology=space_topology,
+ geometry_partition=geometry_partition,
+ with_halo=True,
+ max_node_count=-1,
+ device=device,
+ )Generated docs and public stubs now expose wp.HashGridQuery as the single query type. Runtime aliases wp.HashGridQueryH and wp.HashGridQueryD still warn and forward during the deprecation window, but function annotations should migrate to avoid IDE or type-checker failures. Use unparameterized wp.HashGridQuery for the default wp.float32 query type, and parameterize it when the query uses another coordinate precision:
- query: wp.HashGridQueryH
+ query: wp.HashGridQuery[wp.float16]
- query: wp.HashGridQueryD
+ query: wp.HashGridQuery[wp.float64]This is a source-typing change only. Runtime query objects and aliases remain compatible during the deprecation window.
Python.h (#1339)Linux source builds now compile a small Python C API extension for faster wp.float16 conversions, so the build needs Python.h. If a source build fails with a missing Python.h, install your distribution's Python development package before rebuilding Warp. The extension uses CPython's vectorcall protocol, the public fast-call convention introduced by PEP 590.
warp.jax_experimental is deprecated. Import jax_kernel, jax_callable, clear_jax_callable_graph_cache, JaxCallableGraphMode, and JaxModulePreloadMode from top-level warp instead. The graph-cache default helpers are also deprecated. See JAX integration graduates to stable API for the import diff and graph-cache migration. The deprecated namespace will be removed in Warp 1.16 (#1370).warp.config.verbose and warp.config.quiet are deprecated. Use warp.config.log_level = wp.LOG_DEBUG for verbose diagnostics and warp.config.log_level = wp.LOG_WARNING to suppress the init banner. See Pluggable logging for accepted log-level values. The legacy flags will be removed in a future feature release per the standard deprecation timeline (#1315).wp.HashGridQueryH and wp.HashGridQueryD are deprecated. Use wp.HashGridQuery[wp.float16] or wp.HashGridQuery[wp.float64] in function annotations that need explicit query precision. Use plain wp.HashGridQuery for the default wp.float32 query type. See the HashGrid query annotation migration. Runtime aliases remain available during the deprecation window, but public stubs now expose only wp.HashGridQuery (#1452).quadrature and domain arguments of warp.fem.interpolate(), plus the space argument of warp.fem.make_space_restriction() and warp.fem.make_space_partition(), remain available for this release and are now scheduled for removal in Warp 1.15.We also thank the following contributors from outside the core Warp development team:
launch_dims with differentiable warp.jax_kernel() wrappers (#1380).For a complete list of changes, see the full changelog.
@wp.struct / wp.indexedarray launch arguments that reference Warp array data,
gradient buffers, or wp.indexedarray indices
(GH-1431).warp.fem: colocated Grid2D/Grid3D and packed Nanogrid/AdaptiveNanogrid
geometries with environment-aware lookup and PicQuadrature environment indices, per-cell environment metadata for
unstructured FEM meshes with grouped-BVH lookup and nonconforming field evaluation, and environment-first space
partitions for batched solves plus a multi-environment APIC fluid example
(GH-1407).debug(), info(), warning(), and
error() methods to wp.set_logger(), or scope it temporarily with wp.ScopedLogger. wp.Logger documents
this interface. wp.config.log_level controls the global verbosity threshold and can be scoped temporarily with
wp.ScopedLogLevel. All Python-side diagnostic output emitted by Warp now routes through the configured logger
(GH-1315, GH-1434).warp.optim.linear.cg(), warp.optim.linear.cr(),
warp.optim.linear.bicgstab(), and warp.optim.linear.gmres(). Passing run=False returns a state object
that preallocates solver temporary buffers and can be reused for repeated solves with matching shape, dtype, and
device, plus the same batch layout for batched operators. This avoids per-call allocation overhead and enables use
in CUDA subgraphs (GH-1391).warp.optim.linear solvers: a warp.optim.linear.LinearOperator built with
batch_offsets partitions the DOF vector into independent subproblems that are all solved in a single launch
sequence, with per-batch convergence checks (GH-1391).wp.tile_fft() and wp.tile_ifft() (and their adjoints) on CPU and on GPU builds without libmathdx.
CPU supports any power-of-two FFT length and non-power-of-two FFT lengths up to 4096 elements. The GPU fallback
used without libmathdx, or selected with wp.config.enable_mathdx_fft=False or module option
"enable_mathdx_fft": False, requires a power-of-two FFT length divisible by block_dim, trading runtime
performance for faster kernel compile times (GH-1396).wp.tile_cholesky(), wp.tile_cholesky_solve(),
wp.tile_lower_solve(), and wp.tile_upper_solve() (and their in-place variants) and for the
wp.tile_cholesky() adjoint on GPU builds without libmathdx. The fallback can also be selected on GPU builds
with libmathdx through wp.config.enable_mathdx_solver=False or module option "enable_mathdx_solver": False
(GH-1402).wp.copysign() built-in function to copy one floating-point value's sign onto another value's magnitude,
matching C copysign() (GH-1444).wp.tile_empty() to allocate an uninitialized register or shared-memory tile when every element will be
overwritten before it is read, avoiding the initialization cost of wp.tile_zeros(). Use wp.tile_zeros() for
accumulators or partial writes (GH-1312).wp.curlnoise() in 2D, 3D, and 4D, so gradients now propagate through curl-noise force fields.
Previously, differentiating through wp.curlnoise() produced zero gradients
(GH-1012).wp.CaptureMode, wp.ScopedCapture, and wp.capture_begin(). Users can
select wp.CaptureMode.GLOBAL, wp.CaptureMode.THREAD_LOCAL (the default), or wp.CaptureMode.RELAXED to
control how CUDA handles capture-unsafe runtime calls during graph capture
(GH-1410).build_lib.py --use-dynamic-cuda for source builds that should link Warp's native library against shared
CUDA libraries instead of embedding them statically. The corresponding shared libraries must be present at runtime
(GH-1334).build_lib.py --sanitize=<name> for source builds of Warp's native libraries, enabling AddressSanitizer
instrumentation for warp.dll and warp-clang.dll
(GH-1387).warp.jax_experimental in favor of top-level warp JAX APIs. Migrate
warp.jax_experimental.jax_kernel() to warp.jax_kernel(), warp.jax_experimental.jax_callable() to
warp.jax_callable(), and warp.jax_experimental.GraphMode to warp.JaxCallableGraphMode.
warp.jax_experimental.ModulePreloadMode is also available as warp.JaxModulePreloadMode. The
get_jax_callable_default_graph_cache_max() and set_jax_callable_default_graph_cache_max() helpers are also
deprecated. Pass graph_cache_max to warp.jax_callable() or update the returned callable's graph_cache_max
attribute instead. warp.jax_callable() defaults to graph_cache_max=32. Pass graph_cache_max=None for an
unlimited graph cache. warp.jax_experimental will be removed in Warp 1.16
(GH-1370).warp.config.verbose and warp.config.quiet in favor of warp.config.log_level. Reading or setting
either flag continues to work during the deprecation window and now emits a one-time DeprecationWarning.
warp.config.verbose_warnings is unaffected and continues to control whether warning output includes the source
location (GH-1315).wp.HashGridQueryH and wp.HashGridQueryD. Use wp.HashGridQuery instead. The legacy runtime aliases
remain available during the deprecation period for compatibility, but public stubs now expose only
wp.HashGridQuery, so type annotations should migrate (GH-1452).quadrature and domain arguments of
warp.fem.interpolate(), and the space argument of warp.fem.make_space_restriction() and
warp.fem.make_space_partition().wp.indexedarray launch parameters and @wp.struct
arguments containing arrays in .wrp files. Files saved by Warp 1.13 use the old .wrp format and must be
recaptured. Native APIC C/C++ API users must update code that includes apic.h to declare APIC handles as
explicit C/C++ pointers, such as APICState* and APICGraph*, instead of the old pointer-typedef style
(GH-1431).warp.fem.PicQuadrature() and warp.fem.make_space_partition() positional signatures for
multi-environment FEM support. PicQuadrature() now accepts env_indices before requires_grad, while
make_space_partition() accepts environment_first before the now-keyword-only device and temporary_store.
Pass existing requires_grad, device, and temporary_store arguments by keyword to preserve behavior
(GH-1407).wp.can_access() and wp.config.LaunchArrayAccessMode. Use wp.config.launch_array_access_mode to select relaxed,
strict, or checked pre-launch diagnostics on mixed-device array launches
(GH-1461).wp.tid() overhead in generated kernels by specializing launch-bounds handling for the number of indices a
kernel reads. Existing behavior is preserved when a kernel unpacks fewer wp.tid() indices than the launch provides:
extra trailing launch dimensions still run additional threads that reuse the same returned leading indices
(GH-1270).wp.HashGrid query API around wp.HashGridQuery.
Query objects returned by wp.hash_grid_query() now appear as wp.HashGridQuery in generated docs and stubs
regardless of the grid coordinate precision. Public stubs no longer expose wp.HashGridQueryH or
wp.HashGridQueryD, though the runtime aliases remain available during the deprecation period
(GH-1452).launch_dims to be specified when enable_backward=True in warp.jax_kernel() and
warp.jax_experimental.jax_kernel() (GH-1380).wp.tile_cholesky() by replacing the previous thread-0
scalar adjoint with a cooperative shared-memory implementation that mirrors the libmathdx adjoint and distributes work
across the CUDA block. This keeps the fallback usable when libmathdx is unavailable or enable_mathdx_solver=False.
The tradeoff is two per-block scratch tiles (__shared__ T[n*n], 16 KiB at n=32 and 64 KiB at n=64 for
float64), so large differentiated Cholesky tiles can exceed shared-memory limits that previously allowed them
(GH-1402).wp.float16 values to and from Python float by using a native
Python C API fast-call path. This optimization is why Linux source builds now require Python.h
(GH-1339).wp.map() / wp.tile_map() from attempting backward passes for non-differentiable mapped
callables such as integer bitwise operations (GH-988).Python.h) when building from source on Linux. Warp now compiles
a native fast-call extension that uses the Python C API for wp.float16 conversions, bypassing the previous
ctypes path to lower per-conversion overhead. On Linux, install the distribution package that provides
Python.h (for example, python3-dev / libpython3-dev on Debian or Ubuntu) if the build reports missing
headers (GH-1339).-Og -g for the native library instead of
-O0 -g -fkeep-inline-functions. This enables debug-friendly optimizations, restores -Wuninitialized dataflow
analysis, and avoids leaking unused inline function bodies from <Python.h>
(GH-1414).wp.min(), wp.max(), wp.clamp(), wp.atomic_min(), and wp.atomic_max() on float arguments to use
NaN-as-missing semantics matching C fmin() / fmax(): the operation returns the non-NaN operand when exactly one
is NaN, and NaN only when both are NaN. Vector reductions and wp.argmin() / wp.argmax() skip NaN slots. Adjoint
and atomic variants are updated consistently to route gradients to whichever operand the forward picked. Integer
overloads are unchanged (GH-1376)..wrp graphs so output gradients and retained gradients are preserved and used
the same way as in a normal wp.Tape().backward() call. Older .wrp format versions now fail fast instead of
being misparsed (GH-1431).wp_apic_end_recording(NULL), asynchronous parameter
updates, and vector and small scalar parameters. wp.utils.array_scan() during CPU APIC capture now raises
NotImplementedError instead of producing stale replay output
(GH-1431).warnings.filterwarnings() (GH-1315).arr[i].y = rhs,
arr[i].field = rhs (scalar struct field), m[i][r, c] = rhs, t[i].p = v, t[i].q = q, and
state[i].position = rhs (composite-valued struct field) now correctly propagate gradients instead of silently
dropping them (GH-583,
GH-248, GH-1174).wp.sign(), wp.volume_sample_i(), and the 4-D wp.atomic_exch() overload
(GH-1466).wp.closest_point_edge_edge() at near-parallel configurations. Forward output is unchanged for
well-conditioned inputs and is now well-defined (and gradient-stable) for near-parallel edges, where it could
previously return geometrically valid but unstable barycentric weights
(GH-1437).wp.tile_load(), wp.tile_store(), and indexed tile operations for arrays larger
than 2 GiB. The fix uses 64-bit address arithmetic for tile global-memory accesses, which may slightly increase
address-calculation overhead in tile-heavy kernels (GH-1422).wp.tile_matmul() kernels, and correct the corresponding adjoints
(GH-1439, GH-1440).wp.map() and wp.utils.create_warp_function() failing with IndentationError for lambda expressions split
across multiple parenthesized lines, such as formatter-wrapped expressions
(GH-1351).wp.bfloat16 out tiles in wp.tile_matmul(). The underlying libmathdx GEMM backend requires out to be
wp.float16, wp.float32, or wp.float64. When the backward pass is enabled, wp.bfloat16 a and b tiles are
also rejected because the backward pass uses them as accumulators for adjA and adjB. Set enable_backward=False
on the kernel's module if gradients are not needed (GH-1427).wp.array.fill_() performing temporary device allocations for common fill values, which broke CUDA graph
composability. Oversized fill values still fall back to the previous staging path
(GH-1412).wp.ScopedMemoryTracker.report() that could corrupt the report when called concurrently from
multiple threads (GH-1415).module="unique" kernels that depend on other
Warp functions (GH-1462).+avx10.1-256 "invalid feature combination" warning emitted on every CPU kernel compile on Intel
Granite Rapids hosts (GH-1426).…) instead of the accumulate_dtype . See Breaking changes for the migration.
Warp v1.13 introduces experimental graph capture serialization with CPU replay, letting captured simulations roundtrip through a portable .wrp file and load from standalone C++ on either GPU or CPU. It also adds an experimental cuBQL BVH backend for wp.Mesh that accelerates ray-heavy mesh queries, the wp.bfloat16 scalar type, a pluggable CUDA allocator interface with built-in RAPIDS Memory Manager (RMM) integration, scoped memory tracking with C++-layer call-site attribution, and a batch of new tile primitives (tile_dot, tile_axpy, tile_stack, scatter helpers).
Important
This is an experimental feature. The API may change without a formal deprecation cycle.
Warp v1.13 introduces a portable serialized-graph format. Operations recorded during wp.capture_begin(apic=True) / wp.capture_end() can be saved to a .wrp file with wp.capture_save() and replayed from either Python or standalone C++ via wp.capture_load(), enabling cross-process and cross-language graph reuse (#1349). CPU graph capture is also new in this release: the same wp.Graph object now replays on CPU through wp.capture_launch(), and the underlying APIC operation log is what gets serialized. A new wp.handle (a uint64 alias) carries wp.Mesh handles across save and load so kernels can keep referencing meshes after deserialization.
import warp as wp
with wp.ScopedDevice("cpu"):
a = wp.zeros(64, dtype=float)
b = wp.zeros(64, dtype=float)
wp.capture_begin(apic=True)
wp.copy(b, a)
graph = wp.capture_end()
wp.capture_save(graph, "demo", inputs={"a": a}, outputs={"b": b})
# Later (in the same process or a fresh one): replay from disk.
with wp.ScopedDevice("cpu"):
loaded = wp.capture_load("demo")
loaded.set_param("a", wp.array([1.0] * 64, dtype=float))
wp.capture_launch(loaded)Loading and replaying from standalone C++ (CPU device shown). The full example also walks the _modules/ directory, loads each .o via wp_load_obj, resolves kernel symbols, and registers them with wp_apic_register_loaded_cpu_kernel before the first replay. The snippet below elides that boilerplate:
#include "apic.h"
#include "warp.h"
wp_init(nullptr);
APICGraph graph = wp_apic_load_graph(nullptr, "demo.wrp", 1); // 1 = CPU device
// (Walk demo_modules/, load each .o, and register kernels. See linked example.)
wp_apic_set_param(graph, "a", a_buffer, a_size);
wp_apic_cpu_replay_graph(graph); // For CUDA: cudaGraphLaunch(wp_apic_get_cuda_graph_exec(graph), stream)
wp_apic_get_param(graph, "b", b_buffer, b_size);
wp_apic_destroy_graph(graph);See warp/examples/cpp/02_apic_visualization (CUDA replay) and warp/examples/cpp/03_apic_visualization_cpu (CPU replay) for end-to-end demos with OpenGL visualization.
What gets written:
demo.wrp # operation byte stream + region snapshots + metadata
demo_modules/
<hash>.cubin / .meta # one per CUDA kernel module, arch-pinned
<hash>.o # one per CPU kernel module (CPU capture)
Key capabilities:
wp.capture_save(graph, path, inputs=..., outputs=...) registers named bindings so the consumer side can swap in fresh inputs and read outputs by name without touching the graph topology.wp.capture_load() + wp.capture_launch() support replay on both CPU and CUDA. Loaded graphs expose set_param, get_param, and get_param_ptr for each registered binding, plus params and is_loaded properties on wp.Graph.wp.handle scalar type and wp.Mesh remap let kernels accept mesh handles whose underlying objects are reconstructed on load. APIC walks @wp.struct fields recursively to find handle pointers and remap them.Stability and known gaps:
API Capture is experimental, and we plan to keep adding capabilities and closing gaps over future releases (tracker: #1388). For now, regenerate .wrp artifacts when upgrading Warp. The current operation set, handle types, and platform constraints are documented in the Graphs section of the user guide.
wp.MeshImportant
This is an experimental feature. The API may change without a formal deprecation cycle.
wp.Mesh now accepts bvh_constructor="cubql" to build its acceleration structure with cuBQL, an Apache 2.0-licensed header-only CUDA library for fast BVH construction and traversal (#1286). For ray-heavy workloads on dense static meshes, where the existing SAH builder's exhaustive construction dominates setup time and where ray traversal sits on the simulation hot path, cuBQL typically delivers faster ray queries alongside consistently lower build times than the SAH, median, and LBVH builders. As one specific data point, a Warp-based renderer benchmark on an RTX 4090 (Franka Emika Panda visual mesh, 8192 parallel worlds) saw simulation time drop from 1.41 s to 0.98 s after switching the constructor with no other changes. Speedups depend heavily on mesh size, query mix, and how much of the frame the mesh queries occupy, so benchmark on your own scene before relying on a particular win.
The "cubql" backend currently only routes wp.mesh_query_ray() through cuBQL's traversal kernels. Extending it to point queries, AABB queries, grouped queries, and winding-number support is future work. Today, passing groups=... or support_winding_number=True to a cuBQL wp.Mesh raises a RuntimeError at construction. Calling wp.mesh_query_point_* or wp.mesh_query_aabb_* against a cuBQL mesh silently returns no results. Stick with the default SAH/median/LBVH builders for kernels that mix query types or aren't ray-bound.
import warp as wp
mesh = wp.Mesh(
points=points, # wp.array of wp.vec3
indices=tri_indices, # wp.array of wp.int32, shape (num_tris * 3,)
bvh_constructor="cubql",
)CUDA device-memory allocations can now be routed through any object implementing the wp.Allocator protocol via wp.set_cuda_allocator(), wp.set_device_allocator(), or scoped with wp.ScopedAllocator (#781). The built-in wp.utils.AllocatorRmm delegates to RAPIDS Memory Manager so Warp can share a memory pool with PyTorch, CuPy, or any other RMM-aware framework, eliminating duplicate caching on GPUs that train and simulate in the same process.
import rmm
import warp as wp
rmm.reinitialize(pool_allocator=True, initial_pool_size=2**30)
wp.set_cuda_allocator(wp.utils.AllocatorRmm())
a = wp.zeros(1024, dtype=wp.float32, device="cuda:0") # served from the RMM poolwp.utils.AllocatorRmm requires the rmm package on Linux (pip install rmm-cu12).
wp.ScopedMemoryTracker and the wp.config.track_memory flag enable allocation tracking with call-site attribution and per-category reports across GPU, host, and pinned-host memory (#1269). Tracking is implemented in the C++ native layer by intercepting all wp_alloc_* / wp_free_* calls, so internal allocations from BVH, hash-grid, mesh, volume, and sparse subsystems show up alongside Python-originated arrays, labeled with their subsystem (e.g. (native:bvh), (native:hashgrid), (native:mesh)).
import warp as wp
@wp.kernel
def fill(x: wp.array[float]):
i = wp.tid()
x[i] = float(i)
with wp.ScopedMemoryTracker("training step"):
a = wp.zeros(1_000_000, dtype=wp.float32, device="cuda:0")
b = wp.zeros(2_000_000, dtype=wp.float32, device="cuda:0")
wp.launch(fill, dim=a.size, inputs=[a])
wp.launch(fill, dim=b.size, inputs=[b])Output:
Allocation Tracking Report
Total allocations:
cuda:0 2 (11.44 MB)
Peak usage:
cuda:0 11.44 MB
Live allocations:
(none)
Nest trackers to form hierarchical scopes (e.g. "simulation/collision"), and pass report_func=... to redirect the report to a logger or test-assert callback instead of stdout. For unscoped tracking across an entire process, set wp.config.track_memory = True before wp.init() and call wp.print_memory_report() at any time.
wp.bfloat16 scalar data typewp.bfloat16 joins wp.float16, wp.float32, and wp.float64 as a first-class Warp scalar type (#1332). It supports array allocation, kernel execution, autodiff, DLPack, PyTorch (wp.from_torch(t) for torch.bfloat16 tensors), JAX, and optional NumPy interop via the ml_dtypes package. wp.tile_matmul() and atomic operations also accept bfloat16 tiles, so transformer-style mixed-precision kernels can stay in Warp end to end.
import warp as wp
@wp.kernel
def axpy(a: wp.array[wp.bfloat16], b: wp.array[wp.bfloat16], out: wp.array[wp.bfloat16]):
i = wp.tid()
out[i] = a[i] * wp.bfloat16(2.0) + b[i]
a = wp.array([1.0, 2.0, 3.0], dtype=wp.bfloat16)
b = wp.array([0.5, 0.5, 0.5], dtype=wp.bfloat16)
out = wp.zeros_like(a)
wp.launch(axpy, dim=a.size, inputs=[a, b], outputs=[out])
print(out.numpy()) # [2.5 4.5 6.5]Note
Long chains of bfloat16 math accumulate rounding error at bfloat16 precision because each intermediate result is quantized back to 16 bits before feeding into the next op. This is true whether the underlying op runs on native bf16 hardware (Warp dispatches +, -, * to PTX bf16 instructions on sm_80 and newer GPUs) or goes through a float32 round-trip (/, comparisons, and math built-ins like wp.sqrt / wp.exp, matching how PyTorch, JAX, and CuPy handle those ops). For precision-sensitive code, cast to wp.float32 at the start of the chain and back to wp.bfloat16 only at the boundary so intermediates stay in float32.
wp.Texture1D, wp.Texture2D, and wp.Texture3D accept an externally allocated CUDA array via the new cuda_array= parameter, sharing memory without an extra copy (#1238). For OpenGL interop, the new wp.GLTextureResource class registers an OpenGL texture for use as a Warp texture, so rendering pipelines can sample from textures Warp writes (and vice versa) without going through host memory. Texture objects also gain copy_from() and copy_to() methods for transfers between textures, host arrays, and device arrays. The previous copy_from_array() / copy_to_array() methods are now deprecated.
import warp as wp
# Wrap an existing CUDA array (cuArrayHandle as Python int) in a Warp texture.
tex = wp.Texture2D(
cuda_array=external_cuda_array,
width=1024,
height=1024,
channels=4,
dtype=wp.float32,
filter_mode=wp.Texture.FILTER_LINEAR,
)
# Round-trip into a Warp array without intermediate host buffers.
out = wp.zeros((1024, 1024, 4), dtype=wp.float32, device="cuda:0")
tex.copy_to(out)This release adds several tile primitives that shorten common kernel patterns:
wp.tile_dot(a, b) (#1364) computes the dot product of two same-shape tiles, returning a single-element tile of the underlying scalar type. For tiles of vectors or matrices, each element pair is fully contracted (e.g. wp.dot(a[i], b[i]) for tiles of vec3f).wp.tile_axpy(alpha, src, dest) (#1363) performs the fused in-place update dest += alpha * src without allocating an intermediate scaled tile.wp.tile_scatter_add(a, i, value, has_value, atomic=True) (#1342) is a cooperative scatter-add into a shared-memory tile. Set atomic=False for faster writes when indices are guaranteed unique across threads.wp.tile_scatter_masked() (#1298) writes per-thread values into a shared-memory tile with cooperative synchronization, so each lane can publish a value at its own index without manual barriers.wp.tile_query_valid() (#1335) exposes a cleaner loop condition for tile BVH and mesh AABB queries that avoids the wp.tile_max() reduction overhead the previous pattern required.import warp as wp
@wp.kernel
def reduce_dot():
a = wp.tile_ones(dtype=float, shape=64)
b = wp.tile_ones(dtype=float, shape=64) * 2.0
d = wp.tile_dot(a, b)
print(d)
wp.launch_tiled(reduce_dot, dim=[1], inputs=[], block_dim=64)Output:
[128] = tile(shape=(1), storage=register)
wp.tile_stack() allocates a cooperative thread-block stack in shared memory, with wp.tile_stack_push(), wp.tile_stack_pop(), wp.tile_stack_clear(), and wp.tile_stack_count() operations (#1287). The stack lets all threads in a block contribute or consume entries without manual atomic-counter management. wp.tile_stack_push() accepts a has_value flag so threads can opt out of pushing while staying in the cooperative call, and wp.tile_stack_pop() returns a (value, slot) pair with slot == -1 for threads that did not get an element (for example, when the stack runs dry).
import warp as wp
BLOCK = 8
CAP = wp.constant(8)
@wp.kernel
def compact(data: wp.array[int], out: wp.array[int], out_count: wp.array[int]):
_i, j = wp.tid()
s = wp.tile_stack(capacity=CAP, dtype=int)
val = data[j]
wp.tile_stack_push(s, val, val > 5)
if j == 0:
out_count[0] = wp.tile_stack_count(s)
result, slot = wp.tile_stack_pop(s)
if slot != -1:
out[slot] = resultwp.tile_cholesky(), wp.tile_cholesky_inplace(), wp.tile_cholesky_solve(), and wp.tile_cholesky_solve_inplace() accept a new fill_mode parameter ("lower" or "upper") for upper Cholesky factorization and solve (#1318). The default lower path reads its input by columns, which hits a known slowdown at power-of-2 tile sizes. The new fill_mode="upper" path reads by rows and avoids that cliff, running roughly 1.2x to 1.7x faster on power-of-2 tiles in microbenchmarks on an RTX PRO 6000 Blackwell. Out-of-place tile_cholesky factorization also gains a backward pass, so it can be used as part of a differentiable solve (#1316).
import warp as wp
TILE = 64
@wp.kernel
def factor(A: wp.array2d[float]):
a = wp.tile_load(A, shape=(TILE, TILE))
L = wp.tile_cholesky(a, fill_mode="upper")
wp.tile_store(A, L)wp.tile_fft() and wp.tile_ifft() now accept tiles of rank N >= 2 instead of being limited to 2-D, computing the FFT along the last dimension and treating any leading dimensions as independent batches (#1317). A tile of shape (B1, B2, N) or (B1, B2, B3, N) no longer needs to be reshaped to a 2-D (B, N) tile before transforming. The other constraints carried over from the 2-D path are unchanged: the input must be a register-storage tile of wp.vec2f or wp.vec2d (interpreted as complex pairs), and the FFT length (the last dimension) must be at least 2 * block_dim.
wp.tile_load() and wp.tile_store() get two new fast paths without any kernel changes (#1236):
wp.mat33, wp.mat44, or wp.mat66 get a new shared-memory copy path that replaces the previous per-element loop, restoring the bandwidth that the old path lost on multi-byte elements. The bigger the element, the bigger the speedup.A new aligned parameter on both functions skips runtime alignment checks when the caller can guarantee 16-byte alignment, contiguity, and in-bounds access:
- wp.tile_load(arr, shape=N, offset=i * N)
+ wp.tile_load(arr, shape=N, offset=i * N, aligned=True)@wp.func_native@wp.func_native snippets now receive tile parameters by reference instead of by value, matching the behavior of @wp.func (#1362). Native snippets can therefore modify shared tiles in place rather than receiving a local copy that gets discarded on return.
warp.fem double-precision supportwarp.fem gains end-to-end wp.float64 support (#418). Precision is selected on the geometry (e.g. scalar_type=wp.float64 on grid constructors) and propagates automatically to function spaces, quadrature, fields, and integration kernels. Existing wp.float32 setups are unchanged.
import warp as wp
from warp.fem import Grid2D
geo = Grid2D(res=(64, 64), bounds_lo=wp.vec2(0.0), bounds_hi=wp.vec2(1.0), scalar_type=wp.float64)
# spaces, quadrature, integrate(...) calls built on `geo` now use float64 throughout.The default output_dtype of warp.fem.integrate() now follows the geometry's scalar type (wp.float32 or wp.float64) instead of the accumulate_dtype. See Breaking changes for the migration.
CPU kernel compile times in multi-module workloads drop substantially via precompiled-header support, controlled by the existing warp.config.use_precompiled_headers setting (#595). Repeated module compiles in a session no longer re-parse the bundled headers on every call.
CPU kernels now compile with -march=native by default, so generated code automatically picks up the wider SIMD instruction sets (AVX2, AVX-512, NEON variants) available on the build host instead of targeting a generic x86-64 / aarch64 baseline (#1308). Existing kernels do not need any change to benefit.
The follow-on implication is that AOT-compiled CPU modules and shared CPU kernel caches are now host-specific. wp.compile_aot_module() emits a warning when it produces a CPU module under the new default, because loading that module on a host with a narrower CPU feature set would fail with an illegal-instruction crash. Set wp.config.cpu_compiler_flags = "" before compiling to opt back into a portable baseline build. Cached CPU .o filenames also include a short host-ISA hash (e.g. module.cpu1a2b3c4d.o), so heterogeneous CI runners and shared kernel caches no longer cross-load incompatible binaries.
@wp.kernel@wp.kernel(module="unique", module_options={...}) accepts a dict of module-level compilation options inline (#1250), avoiding the previous pattern of toggling warp.config.* globally before defining the kernel. One useful application is disabling cuBLASDx for wp.tile_matmul during development to skip its slow LTO compile step:
import warp as wp
M, N, K = 16, 16, 16
@wp.kernel(module="unique", module_options={"enable_mathdx_gemm": False})
def matmul(a: wp.array2d[float], b: wp.array2d[float], c: wp.array2d[float]):
ta = wp.tile_load(a, shape=(M, K))
tb = wp.tile_load(b, shape=(K, N))
tc = wp.tile_zeros(dtype=float, shape=(M, N))
wp.tile_matmul(ta, tb, tc)
wp.tile_store(c, tc)On a clean kernel cache the kernel above compiles in roughly a second, instead of the tens of seconds the default cuBLASDx-backed path can take.
-O2, CUDA defaults remain -O3) (#1310).None) and its explicit default, avoiding spurious recompilation (#1307).wp.get_suggested_block_size(kernel) queries the CUDA driver's occupancy API for a launch configuration that maximizes per-SM occupancy (#1270). Returns (block_size, min_grid_size). Pass block_size to wp.launch(..., block_dim=...).and / or (#1329): chained boolean operators in kernels now use Python semantics, so guards like if arr and arr[i] == 0 no longer crash from eagerly evaluating the right-hand side.wp.vec3d(), wp.mat22h(), wp.quatd(), etc. now accept Python scalar literals directly without an explicit cast, preserving precision.wp.indexedarray fields in @wp.struct (#1327): structs may now hold wp.indexedarray fields, with assignment, device transfer, and NumPy structured-value support working the same way as wp.array fields.wp.Volume anisotropic voxels (#1193): wp.Volume.load_from_numpy() and wp.Volume.allocate() accept a 3-element sequence for voxel_size, enabling anisotropic voxel spacing.tape.backward() zeroes intermediate gradient buffers (#1062)tape.backward() now zeroes the gradient buffer of any array written to in the forward pass, matching PyTorch's behavior for intermediate gradients. This is a correctness fix: previously, calling tape.backward() multiple times accumulated stale .grad from the prior call into the new pass.
# v1.12 behavior: out.grad still holds the upstream seed after backward()
out = wp.zeros(N, dtype=float, requires_grad=True)
with wp.Tape() as tape:
wp.launch(forward_kernel, dim=N, inputs=[x], outputs=[out])
out.grad = wp.ones_like(out)
tape.backward()
# v1.12: out.grad == ones (still the seed)
# v1.13: out.grad == zeros (cleared during backward)
# v1.13 migration: pass fresh upstream gradients each backward call
with wp.Tape() as tape:
wp.launch(forward_kernel, dim=N, inputs=[x], outputs=[out])
tape.backward(grads={out: wp.ones_like(out)})
# Or opt out of the zeroing for arrays you want to inspect after backward.
# Only safe when each element is written at most once per forward pass.
out = wp.array(shape=N, dtype=float, requires_grad=True, retain_grad=True)Backward passes of kernels with many per-element array writes (e.g. matrix component assignments) may be slower because of the additional zeroing. See the differentiability guide for the full rationale.
warp.fem.integrate() output_dtype default change (#418)The default output_dtype of warp.fem.integrate() now follows the geometry's scalar type instead of accumulate_dtype (which itself defaults to wp.float64). For wp.float32 geometries this changes the default output from wp.float64 to wp.float32.
- result = warp.fem.integrate(form, fields=fields) # v1.12: float64
+ result = warp.fem.integrate(form, fields=fields, output_dtype=wp.float64) # explicitWarp now aims to publish a feature release every month. This replaces the previous schedule that alternated monthly between feature and bugfix releases. Bugfix releases are no longer regularly scheduled. They are issued ad hoc, only when an important issue cannot wait for the next feature release, and only against the most recent feature release line. See the Compatibility & Support page for the full versioning, deprecation, and support policy.
warp.torch, warp.context, etc.), used warp.mat / warp.vec, referenced warp.context.Devicelike, or relied on the Module.foo -> Module._foo proxy now raises rather than warning (#1352). Use the curated public API in the warp namespace. The API documentation is the source of truth for what is public.wp.isfinite(), wp.isnan(), and wp.isinf() no longer accept integer types. These functions now accept floating-point arguments only, finalizing the deprecation announced in v1.11 (#847). Drop integer call sites or wrap the operand in a float cast.copy_from_array() / copy_to_array() deprecated. Use the new copy_from() / copy_to() methods instead (#1238). Calling the deprecated names emits a DeprecationWarning. They will be removed in a future feature release per the standard deprecation timeline.We also thank the following contributors from outside the core Warp development team:
wp.indexedarray field support in @wp.struct, with assignment, device transfer, and NumPy structured-value handling (#1327).wp.Volume.load_from_numpy() and wp.Volume.allocate(), including the shared _normalize_voxel_size() validation helper and tests (#1193).For a complete list of changes, see the full changelog.
wp.capture_begin()/wp.capture_end() can be serialized to .wrp files via
wp.capture_save() and loaded for execution from Python or standalone C++ via
wp.capture_load(). CPU graph capture supports replay through wp.capture_launch(). Added
wp.handle scalar type for mesh-handle serialization
(GH-1349).wp.Mesh, selectable via bvh_constructor="cubql".
Currently only supports wp.mesh_query_ray(). Point queries, AABB queries, grouped queries,
and winding number queries are not yet supported
(GH-1286).wp.Texture1D, wp.Texture2D, or wp.Texture3D via the new cuda_array= parameter,
or register an OpenGL texture for CUDA-OpenGL interop via the new wp.GLTextureResource
class (GH-1238).wp.float64) support to warp.fem.
Precision is selected via the geometry (e.g. scalar_type=wp.float64 on grid constructors)
and propagated automatically to function spaces, quadrature, fields, and integration kernels
(GH-418).wp.Allocator protocol and pass it to wp.set_cuda_allocator(),
wp.set_device_allocator(), or scope it with wp.ScopedAllocator. Built-in
wp.utils.AllocatorRmm routes allocations through RAPIDS Memory Manager (RMM) for
pool-sharing with PyTorch, CuPy, or other RMM-aware frameworks
(GH-781).wp.ScopedMemoryTracker context manager and wp.config.track_memory flag for tracking
memory allocations with call-site attribution, scope grouping, and per-category reports (GPU,
host, pinned host). Tracking is implemented in the C++ native layer, capturing both
Python-originated and C++ internal allocations with descriptive labels (e.g. (native:bvh),
(native:hashgrid), (native:mesh), (native:volume), (native:sparse))
(GH-1269).wp.bfloat16 scalar data type with array allocation, kernel execution, autodiff, DLPack, PyTorch, JAX, and
optional ml_dtypes NumPy interop (GH-1332).wp.get_suggested_block_size() to query a suggested CUDA launch configuration for a kernel
based on occupancy (GH-1270).wp.tile_stack() cooperative thread-block stack for tile kernels, with wp.tile_stack_push(),
wp.tile_stack_pop(), wp.tile_stack_clear(), and wp.tile_stack_count() operations
(GH-1287).wp.tile_scatter_masked() for per-thread writes into a shared-memory tile, with
cooperative synchronization (GH-1298).wp.tile_scatter_add() for per-thread cooperative adds into shared-memory tiles,
with an atomic parameter (default True) that can be set to False for faster writes
when indices are guaranteed unique across threads
(GH-1342).wp.tile_query_valid() for tile BVH and mesh AABB queries, providing a cleaner loop condition that avoids the
wp.tile_max() reduction overhead (GH-1335).wp.tile_axpy(alpha, src, dest) for the fused in-place update
dest += alpha * src on tiles. Avoids allocating an intermediate scaled tile
(GH-1363).wp.tile_dot(a, b) to compute the dot product of two tiles of matching
shape and dtype, returning a single-element tile of the tile's scalar type.
For tiles of vectors or matrices, each element pair is fully contracted (e.g.
wp.dot(a[i], b[i]) for tiles of vec3f). Replaces the longer
wp.tile_sum(wp.tile_map(wp.tensordot, a, b)) pattern
(GH-1364).fill_mode parameter to wp.tile_cholesky(), wp.tile_cholesky_inplace(),
wp.tile_cholesky_solve(), and wp.tile_cholesky_solve_inplace() for upper Cholesky factorization and solve,
improving memory access patterns and eliminating shared-memory bank conflicts at power-of-2 tile sizes
(GH-1318).tile_cholesky factorization
(GH-1316).wp.tile_fft() and wp.tile_ifft() to operate on N-D tiles (N >= 2),
computing the FFT along the last dimension with all leading dimensions treated as independent batches
(GH-1317).wp.tile_load() and wp.tile_store() in two new cases without kernel changes:
3D and 4D aligned tiles (extending the existing optimization that was 2D-only) and tiles of
large element types like wp.mat33, wp.mat44, or wp.mat66. Add an aligned parameter
to skip runtime alignment checks when the caller guarantees 16-byte alignment, contiguity,
and in-bounds access (GH-1236).copy_from() and copy_to() methods to texture objects for copying between textures,
host arrays, and device arrays (GH-1238).module_options dict parameter to @wp.kernel for inline module-level compilation options
on "unique" modules (GH-1250).wp.indexedarray fields in @wp.struct (assignment, device transfer, and NumPy structured values)
(GH-1327).warp.config.cpu_compiler_flags to control CPU kernel compilation targeting.
The default (None) now detects host CPU features at runtime instead of compiling for a generic target
(GH-1308).warp.torch, warp.context, etc.), used warp.mat/warp.vec,
referenced warp.context.Devicelike, or relied on the Module.foo → Module._foo proxy
will now raise rather than emit a deprecation warning. Use the curated public API in
the warp namespace instead
(GH-1352).wp.isfinite(), wp.isnan(), and wp.isinf()
(GH-847).copy_from_array() and copy_to_array() on texture objects,
use copy_from() and copy_to() instead (GH-1238).tape.backward() to zero the gradient buffer of arrays written to in
the forward pass, matching PyTorch's behavior for intermediate gradients. Code that
inspects .grad on such arrays after backward will see zeros. Allocate the array with
wp.array(..., retain_grad=True) to preserve gradients for inspection, but only when each
element is written at most once in the forward pass. Setting retain_grad=True disables
the zeroing that prevents double-counting across multiple writes. When calling
tape.backward() multiple times, pass fresh upstream gradients via the grads=
argument. Backward passes of kernels with many per-element array writes (e.g. matrix
component assignments) may be slower due to the additional zeroing. See the
differentiability guide
for details (GH-1062).and/or operators in kernels, matching
Python semantics. Previously all operands were eagerly evaluated, so guards like
if arr and arr[i] == 0 could crash
(GH-1329).warp.fem.integrate() output_dtype default to the geometry's scalar type
(wp.float32 or wp.float64) instead of accumulate_dtype (which itself defaults to
wp.float64). Pass output_dtype=wp.float64 explicitly to restore the previous behavior
(GH-418).warp.config.optimization_level to the full CPU compilation pipeline.
The default CPU optimization level remains -O2. CUDA kernels default to -O3
(GH-1310).warp.config.use_precompiled_headers setting
(GH-595).None
and its explicit default (e.g. warp.config.optimization_level = None vs = 3). Both now
resolve to the same module hash (GH-1307).@wp.func_native by reference, matching @wp.func behavior.
Previously tile parameters were passed by value, preventing native snippets from modifying
shared tiles in-place (GH-1362).wp.vec3d(), wp.mat22h(), wp.quatd())
to accept scalar literals directly, preserving precision without explicit casts
(GH-1297).wp.Volume.load_from_numpy() and wp.Volume.allocate() to accept a 3-element sequence for voxel_size,
enabling anisotropic voxel spacing (GH-1193).warp.config.legacy_cpu_linker = True to opt
back in to the previous linker
(GH-1346).libdevice.10.bc, which caused all CUDA kernels compiled via the
Clang/LLVM toolchain (llvm_cuda=True) to fail when no system CUDA toolkit was available.wp.copy() with non-contiguous sources
(GH-1384).llvm_cuda or use_precompiled_headers are changed between runs
(GH-903).ValueError: Cell is empty crash during eager module hashing when a kernel closure references a variable
assigned later in the enclosing scope (GH-913).wp.transform_set_translation() silently operating on a temporary copy
instead of modifying the original when called with values from wp.transform_identity(),
wp.quat_identity(), or generic types created via wp.types.transformation(),
wp.types.vector(), etc. Kernel-scope usage was unaffected
(GH-1336).wp.constant(wp.int32(IntEnum_value)) emitting the symbolic enum name instead of the
integer value in generated C++/CUDA code on Python 3.10, causing compilation failures
(newton#2363).wp.Graph class returned by wp.capture_end() and wp.capture_load(),
including the APIC parameter-binding methods (set_param, get_param, get_param_ptr)
and the params and is_loaded properties.The following deprecations will be finalized in Warp 1.13.0 :
Warp v1.12.1 is a bugfix release following v1.12.0. For a complete list of changes, see the changelog.
Tile Correctness: Fixed several tile kernel issues, including kernel dispatch using incorrect block_dim across devices (#1254), wp.tile_matmul() and wp.tile_fft() ignoring module-level enable_backward (#1320), and @wp.func with tile parameters failing to compile with shared-memory tiles (#1313). Tile parameters in @wp.func are now passed by reference for both register and shared storage, matching Python's semantics for mutable objects.
Silent Correctness Bugs: Fixed compile-time constants silently losing precision when passed to 64-bit scalar constructors like wp.float64() (#485), wp.HashGrid neighbor queries missing results for negative coordinates (#1256), and augmented assignments with subscript or attribute targets (e.g., s.field += expr, arr[i] *= expr) double-evaluating the target expression (#1233).
Type System and Tooling: Fixed struct field assignments converting Warp scalar types to plain Python types (#1288), wp.array[dtype] not being recognized by mypy (#1278), and array annotation repr() displaying raw internal class paths (#1341).
Kit Extensions Removed: The Omniverse Kit extensions have been removed from this repository (#1296).
New Examples: Added a differentiable 2-D Navier-Stokes optimization example and three warp.fem examples for Taylor-Green vortex, Kelvin-Helmholtz instability, and shallow water equations.
The following deprecations will be finalized in Warp 1.13.0:
wp.isfinite(), wp.isnan(), and wp.isinf() will no longer accept integer types. These functions only produce meaningful results for floating-point inputs and have been emitting deprecation warnings for integer arguments since 1.12.0.warp._src directly (e.g., from warp._src.math import ...) will stop working without warning. Please migrate to the public warp.* API.Thanks to @thomasbbrunner for fixing boolean vector element assignment (#1302).
@wp.func functions by reference for both register and shared storage,
matching Python's semantics for mutable objects. Previously, register tiles were passed by value
(GH-1313).wp.float64(), wp.int64(), wp.uint64()) (GH-485).wp.HashGrid neighbor queries silently missing results when coordinates are negative
(GH-1256).s.field += expr, arr[i] *= expr),
causing side effects in target indices or the right-hand side to trigger multiple times
(GH-1233).wp.tile_matmul() and wp.tile_fft() ignoring the module-level enable_backward flag
(GH-1320).@wp.func with tile parameters failing to compile when called with shared-memory tiles
(GH-1313).block_dim when the same kernel is launched on different devices,
which could cause out-of-bounds shared memory access and memory corruption in tile kernels
(GH-1254).wp.tile_argmin(), wp.tile_argmax(), and related tile operations crashing in debug mode
when block_dim exceeds the tile element count (GH-1133).wp.tile_map() with wp.tile_store() failing for custom vector and matrix types
created via wp.types.vector() or wp.types.matrix() (GH-1311).@wp.func_native return type resolution for wp.types.vector(), wp.types.matrix(),
and other complex type annotations (GH-1300).set_module_options(), get_module_options(), and load_module() crashing with AttributeError
when called from code executed via runpy.run_module() (e.g., python -m package.module)
(GH-1274).Texture creation crashing with AttributeError when wp.init() has not been called
(GH-1272).wp.float32, wp.int32)
to plain Python types, causing subsequent reads to return float or int instead of the original type
(GH-1288).v[i] = True) raising an error
(GH-1302).wp.sign() returning the wrong type for custom vector types.wp.array[dtype] subscript syntax not being recognized by mypy,
which reported "array" expects no type arguments (GH-1278).repr() displaying raw internal class paths for dtype
(e.g., wp.array(dtype=<class 'warp._src.types.uint32'>, ndim=4)) instead of
clean type names (e.g., wp.array(dtype=wp.uint32, ndim=4)) (GH-1341).example_navier_stokes_perturbation.py)
for optimal initial perturbation, complementing the solver in example_fft_poisson_navier_stokes_2d.py.warp.fem examples for Taylor-Green vortex (example_taylor_green.py),
Kelvin-Helmholtz instability (example_kelvin_helmholtz.py), and shallow water equations (example_shallow_water.py).warp._src.lang leaking into published documentation page titles, URLs,
and search engine results for built-in functions (GH-1275).Experimental. This API may change without a formal deprecation cycle.
Warp v1.12 adds experimental hardware-accelerated texture sampling on CUDA GPUs, extends tile programming with element-wise arithmetic operators and differentiable FFT, and broadens JAX interoperability with jax.vmap support. This release also introduces subscript-style type hints for better IDE integration, new quaternion and approximate-math builtins, B-spline shape functions in warp.fem, and a collection of utility and diagnostics APIs.
Experimental. This API may change without a formal deprecation cycle.
Warp v1.12 introduces wp.Texture1D, wp.Texture2D, and wp.Texture3D classes that leverage CUDA texture memory for hardware-accelerated interpolation directly inside Warp kernels. On GPU, texture reads are routed through dedicated texture units that perform filtered lookups in a single instruction, making them ideal for rendering, volume sampling, signed-distance-field queries, and simulation lookup tables. On CPU, a software fallback provides identical semantics so the same kernel code runs on both devices.
import warp as wp
import numpy as np
wp.init()
# 64x64 single-channel height map
data = np.random.rand(64, 64).astype(np.float32)
# Create a 2D texture with bilinear filtering
tex = wp.Texture2D(data, filter_mode=wp.Texture.FILTER_LINEAR)
@wp.kernel
def sample_texture(tex: wp.Texture2D, coords: wp.array[wp.vec2f], out: wp.array[float]):
i = wp.tid()
# Coordinates are in [0, 1]; bilinear interpolation is automatic
out[i] = wp.texture_sample(tex, coords[i], dtype=float)
coords = wp.array(np.random.rand(1024, 2).astype(np.float32), dtype=wp.vec2f)
result = wp.zeros(1024, dtype=float)
wp.launch(sample_texture, dim=1024, inputs=[tex, coords, result])
print(f"Sampled {result.shape[0]} points, range: [{result.numpy().min():.4f}, {result.numpy().max():.4f}]")
# Example output: Sampled 1024 points, range: [0.0069, 0.9793]Key capabilities:
wp.Texture1D, wp.Texture2D, wp.Texture3D) with matching wp.texture_sample() overloads that accept scalar, vec2f, or vec3f coordinates.FILTER_POINT for nearest-neighbor sampling and FILTER_LINEAR for bilinear (2D) or trilinear (3D) interpolation.ADDRESS_WRAP, ADDRESS_CLAMP, ADDRESS_MIRROR, and ADDRESS_BORDER control how out-of-range texture coordinates are handled, configurable per axis.copy_from_array() and copy_to_array() methods to transfer data between wp.array objects and texture memory. A cuda_surface property exposes the CUDA surface handle for advanced interop.When annotating kernel parameters with call-syntax forms like wp.array(dtype=float), static type checkers such as Pyright and Pylance flag these as errors because the expressions look like constructor calls rather than type annotations. Warp v1.12 adds subscript-style alternatives that are recognized as valid generic aliases (#1216):
# Before (flagged as error by Pyright/Pylance):
@wp.kernel
def my_kernel(a: wp.array(dtype=float), b: wp.array2d(dtype=wp.vec3)):
...
# After (clean subscript syntax):
@wp.kernel
def my_kernel(a: wp.array[float], b: wp.array2d[wp.vec3]):
...The subscript syntax is supported for all array dimensionalities (wp.array[dtype] through wp.array4d[dtype]) as well as wp.tile[dtype] for tile-typed arguments.
Warp's static type checking compatibility is being improved incrementally, and you may encounter other Pyright/Pylance diagnostics that are not yet resolved. If you run into type checking issues, please report them as sub-issues of #549.
The new wp.print_diagnostics() function displays a comprehensive snapshot of the Warp build and runtime environment (software versions, CUDA information, build flags, and available devices) in a single call (#1221). Two companion helpers, wp.get_cuda_toolkit_version() and wp.get_cuda_driver_version(), return the CUDA toolkit and driver versions as integer tuples (#1172). Together these are useful for debugging environment issues, capturing context in CI logs, and providing system information when filing bug reports.
Warp v1.12 adds quaternion and spatial transformation helpers: wp.quat_from_euler(), wp.quat_to_euler(), wp.transform_twist(), and wp.transform_wrench() (#1237). The Euler conversion functions accept axis indices (0 = X, 1 = Y, 2 = Z) so you can specify arbitrary rotation-order conventions such as ZYX or XYZ, making them suitable for robotics and animation pipelines:
euler = wp.vec3(0.0, wp.PI / 4.0, 0.0)
q = wp.quat_from_euler(euler, 2, 1, 0) # ZYX convention
print(q) # [0.0, 0.3826834559440613, 0.0, 0.9238795042037964]wp.div_approx() and wp.inverse_approx() expose GPU hardware fast-math instructions (div.approx.f32 and rcp.approx.ftz.f64) for approximate floating-point division and reciprocal, offering higher throughput at reduced precision (#1199). Only floating-point types are supported. On CPU, both functions fall back to exact arithmetic so the same kernel code runs correctly on either device.
The internal marching cubes lookup tables are now exposed as public class attributes on wp.MarchingCubes: CUBE_CORNER_OFFSETS, EDGE_TO_CORNERS, CASE_TO_TRI_RANGE, and TRI_LOCAL_INDICES (#1151). These tables enable custom marching cubes implementations for advanced use cases such as sparse volume extraction or procedural mesh generation without having to duplicate the standard lookup data.
wp.utils.graph_coloring_assign(), wp.utils.graph_coloring_balance(), and wp.graph_coloring_get_groups() are now part of the public API (#1145). These graph coloring utilities were originally introduced in warp.sim in v1.5.0 for use with VBDIntegrator and were removed along with the warp.sim module in v1.10.0. They are now re-introduced as standalone functions in wp.utils, independent of any physics module. They partition a graph into independent color groups, which is useful for parallel constraint solving, conflict-free mesh updates, and other tasks that require concurrent writes to non-adjacent elements.
Tiles now support native Python * and / operators for element-wise multiplication and division, including broadcast between tiles and scalar constants (#1006, #1009). The supported forms are tile * tile, tile * constant, constant * tile for multiplication, and tile / tile, tile / constant, constant / tile for division. All combinations are differentiable and work with scalar, vector, and matrix element types.
import warp as wp
TILE_SIZE = wp.constant(64)
@wp.kernel
def scale_and_normalize(
a: wp.array[float],
b: wp.array[float],
out: wp.array[float],
):
i = wp.tid()
ta = wp.tile_load(a, shape=TILE_SIZE, offset=i * TILE_SIZE)
tb = wp.tile_load(b, shape=TILE_SIZE, offset=i * TILE_SIZE)
product = ta * tb # element-wise multiply
scaled = product * 0.5 # broadcast scalar multiply
result = scaled / tb # element-wise divide
wp.tile_store(out, result, offset=i * TILE_SIZE)
N = 256
a = wp.ones(N, dtype=float)
b = wp.full(N, value=2.0, dtype=float)
out = wp.zeros(N, dtype=float)
wp.launch_tiled(scale_and_normalize, dim=[N // 64], inputs=[a, b, out], block_dim=64)
print(out.numpy()[:8]) # [0.5 0.5 0.5 0.5 0.5 0.5 0.5 0.5]wp.tile_from_thread()wp.tile_from_thread() broadcasts a scalar or vector value held by a single thread to a shared tile visible to all threads in the block (#1178). This is useful when one thread computes a value (e.g., a reduction result or a loop-invariant parameter) that the entire block needs to use in subsequent tile operations. The function accepts a thread_idx argument to specify which thread's value is broadcast, and supports both "shared" and "register" storage modes.
wp.tile_fft() and wp.tile_ifft() now support reverse-mode automatic differentiation when recorded on a wp.Tape() (#1138). Warp automatically provides the correct gradient implementations for both transforms, so gradients propagate seamlessly through frequency-domain operations. This enables end-to-end gradient computation through pipelines that mix spatial and spectral steps, which is useful for differentiable signal processing, spectral methods, and PDE solvers.
Setting wp.config.enable_mathdx_gemm = False (or passing "enable_mathdx_gemm": False as a module option) disables cuBLASDx for wp.tile_matmul(), falling back to an optimized scalar GEMM implementation (#1228). This avoids the slow link-time optimization (LTO) step required by libmathdx during development iteration, while keeping libmathdx available for operations that have no scalar fallback, such as Cholesky factorization and FFT. The scalar fallback may be slower than cuBLASDx depending on tile sizes, data types, and block_dim, so this option is primarily intended for faster compile–edit–run cycles during development rather than production use.
Shared-memory tile loads and stores via wp.tile_load() / wp.tile_store() have been accelerated for non-power-of-two tile sizes (#1239). The improvement is most pronounced when source arrays fit within the GPU L2 cache, reducing overhead for tile-based kernels that operate on irregularly shaped data blocks.
jax.vmap supportjax.vmap() can now be used with Warp kernels and callables exposed through jax_kernel() and jax_callable() (#859). This enables vectorized mapping over batched inputs, letting JAX automatically handle batching of Warp kernel invocations without manual loop constructs. The vmap_method parameter controls how batching is implemented. "broadcast_all" broadcasts all arrays to include the batch dimension, while "sequential" iterates over the batch dimension one element at a time.
import warp as wp
import jax
import jax.numpy as jp
from warp.jax_experimental.ffi import jax_kernel
@wp.kernel
def add_kernel(a: wp.array[float], b: wp.array[float], out: wp.array[float]):
i = wp.tid()
out[i] = a[i] + b[i]
# Wrap the Warp kernel for JAX
jax_add = jax_kernel(add_kernel, vmap_method="broadcast_all")
# Batched inputs: 3 arrays of 4 elements each
a = jp.arange(12, dtype=jp.float32).reshape((3, 4))
b = jp.ones((3, 4), dtype=jp.float32)
# Use jax.vmap to apply the kernel over the batch dimension
(result,) = jax.jit(jax.vmap(jax_add))(a, b)
print(result)
# [[ 1. 2. 3. 4.]
# [ 5. 6. 7. 8.]
# [ 9. 10. 11. 12.]]has_side_effect flagThe new has_side_effect flag on jax_kernel() and jax_callable() ensures that JAX does not optimize away Warp FFI calls whose outputs are unused downstream (#1240). This is important for kernels that write to global state or perform I/O as side effects. Without the flag, JAX's dead-code elimination may silently skip their execution. Setting has_side_effect=True marks the call as effectful, forcing JAX to always execute it.
warp.fem enhancementsWarp v1.12 adds B-spline basis functions to warp.fem with SquareBSplineShapeFunctions (2D) and CubeBSplineShapeFunctions (3D), accessible via ElementBasis.BSPLINE (#1208). B-spline bases provide higher continuity across element boundaries compared to standard Lagrange bases, which is beneficial for applications requiring smooth solutions such as thin-shell mechanics or isogeometric analysis. Degrees 1 through 3 are supported on Grid2D, Grid3D, and Nanogrid geometries.
cells() operator now accepts traced fields, returning the underlying cell-level field for evaluation at cell-space samples (e.g., from lookup()).PicQuadrature particles can now span multiple cells by passing a tuple of 2D arrays (cell_indices, coords, particle_fraction) to specify per-particle cell contributions.quadrature and domain arguments of interpolate(), and the space argument of make_space_restriction and make_space_partition (scheduled for removal in 1.14).wp.compile_aot_module() to produce PTX/CUBIN during Docker image builds where no GPU is available (#1085).wp.compile_aot_module() now skips recompilation when the output binary already exists and wp.config.cache_kernels is enabled (#1246).--no-cuda flag has been added to build_lib.py for explicit CPU-only builds (#1223).wp.config.cuda_arch_suffix setting appends architecture-specific suffixes to the --gpu-architecture flag passed to NVRTC (#1065).Device.max_shared_memory_per_block exposes the maximum shared memory per block for CUDA devices (#1243).wp.HashGrid now supports wp.float16 and wp.float64 coordinate types (#1007, #1168).wp.int32) can now be used when indexing vectors and matrices in kernels (#1209).wp.config.enable_tiles_in_stack_memory to True (#1032).The new example_fft_poisson_navier_stokes_2d example demonstrates 2-D incompressible turbulence using a vorticity-streamfunction formulation. It uses tile-based fast Fourier transforms (FFT) to solve the Poisson equation on a periodic domain and advances vorticity transport with strong-stability-preserving Runge-Kutta (SSP-RK3) time integration, showcasing how Warp's tile FFT primitives can be applied to spectral PDE solvers.
Python 3.9 reached end-of-life on October 31, 2025 and no longer receives security updates. Warp v1.12 deprecates Python 3.9 support. A DeprecationWarning is emitted both at runtime (when import warp runs under Python 3.9) and at build time (when compiling the native library with a Python 3.9 interpreter). Support for Python 3.9 will be removed entirely in Warp 1.13. To continue receiving Warp updates beyond v1.12.1, please migrate to Python 3.10 or newer.
Implicit conversion of scalar values to composite types (vectors, matrices, etc.) when launching kernels or assigning to struct fields is now deprecated. Use explicit constructors such as wp.vec3(...) or wp.mat22(...) (#1022). Constructing matrices from row vectors via wp.matrix() has been removed after being deprecated in v1.9 (#1179); use wp.matrix_from_rows() or wp.matrix_from_cols() instead.
The internal API cleanup begun in Warp v1.11, which added deprecation warnings and forwarding calls for internal symbols accessed through the public warp namespace, will be finalized in Warp v1.13. At that point, the deprecation messages and forwarding shims will be removed, and code that still accesses deprecated internal APIs will break. If you have been seeing deprecation warnings about internal symbol access, please update your code before v1.13 is released. If you need help migrating, feel free to ask on GitHub Discussions.
We also thank the following contributors from outside the core Warp development team:
warp.fem temporaries not being released promptly due to reference cycles (#1075).distutils import from the OpenGL rendering example (#1205).For a complete list of changes, see the full changelog.
wp.Texture1D, wp.Texture2D, and wp.Texture3D classes
for hardware-accelerated texture sampling on CUDA devices,
with wp.texture_sample() for linear/bilinear/trilinear interpolation in kernels.
Includes CUDA interop APIs for array↔texture copies and surface handle access
(GH-1122).wp.config.enable_mathdx_gemm and "enable_mathdx_gemm" module option
to disable libmathdx (cuBLASDx) for wp.tile_matmul(), falling back to an optimized scalar GEMM.
Avoids slow LTO compilation during development
while keeping libmathdx available for Cholesky/FFT (GH-1228).wp.array[float], wp.array2d[float],
wp.tile[float]) as alternatives to the call-syntax forms that static type checkers like Pyright flag as errors
(GH-1216).wp.tile_from_thread(), which broadcasts a value from a particular thread to all threads in the block.
(GH-1178).tile * tile, tile * constant, and constant * tile syntax
for element-wise and broadcast multiplication (GH-1006).tile / tile, tile / constant, and constant / tile syntax
for element-wise and broadcast division (GH-1009).wp.tile_fft() and wp.tile_ifft() calls recorded on the tape
(GH-1138).jax.vmap() with jax_kernel() and jax_callable() foreign function interface (FFI) calls
(GH-859).has_side_effect flag to jax_kernel() and jax_callable()
to ensure FFI calls are always executed by JAX (GH-1240).wp.print_diagnostics() to display a comprehensive snapshot of the Warp build and runtime environment,
including software versions, CUDA info, build flags, and devices (GH-1221).wp.get_cuda_toolkit_version() and wp.get_cuda_driver_version() to query CUDA versions
(GH-1172).wp.config.cuda_arch_suffix setting to append architecture-specific ("a")
or family-specific ("f") suffixes to the --gpu-architecture flag passed to NVRTC
(GH-1065).Device.max_shared_memory_per_block attribute exposing the opt-in maximum shared memory per block in bytes
for CUDA devices (GH-1243).wp.float16 and wp.float64 support for wp.HashGrid (GH-1007,
GH-1168).wp.quat_from_euler(), wp.quat_to_euler(),
wp.transform_twist(), etc.) (GH-1237).wp.div_approx() and wp.inverse_approx() built-ins for approximate PTX intrinsics
(div.approx.f32, rcp.approx.ftz.f64) on GPU. Only floating-point types are supported;
falls back to exact arithmetic on CPU (GH-1199).wp.MarchingCubes: CUBE_CORNER_OFFSETS,
EDGE_TO_CORNERS, CASE_TO_TRI_RANGE, and TRI_LOCAL_INDICES. These enable custom marching cubes implementations
for advanced use cases like sparse volume extraction (GH-1151).wp.utils.graph_coloring_assign(), wp.utils.graph_coloring_balance(), and wp.graph_coloring_get_groups()
to the public API for graph coloring (GH-1145).warp.fem: Add B-spline shape functions with SquareBSplineShapeFunctions (2D)
and CubeBSplineShapeFunctions (3D), supporting degrees 1-3. Use via ElementBasis.BSPLINE.
Supported on Grid2D, Grid3D, and Nanogrid geometries (GH-1208).--no-cuda flag to build_lib.py for explicit CPU-only builds,
skipping CUDA toolkit detection and .cu compilation (GH-1223).wp.matrix() at the Python and kernel scopes
(GH-1179).warp.fem: Remove the deprecated .array attribute from TemporaryStore temporaries,
which are now wp.array objects directly.DeprecationWarning is now emitted at runtime and build time
when using Python 3.9. Support will be removed in Warp 1.13.wp.vec3(...) or wp.mat22(...) (GH-1022).warp.fem: Add deprecation warnings for the quadrature and domain arguments of interpolate(),
and the space argument of make_space_restriction and make_space_partition (scheduled for removal in 1.14).wp.float16, wp.float64, wp.int8, etc.)
from built-in functions and vector/matrix indexing instead of Python native types
for non-native scalar types.
Native scalar types (wp.int32, wp.float32, wp.bool) continue to return Python int, float,
and bool. Set wp.config.legacy_scalar_return_types = True to restore the previous behavior
(GH-905).wp.config.enable_tiles_in_stack_memory to True,
allocating shared tile storage on the stack for all CPU architectures
(GH-1032).wp.compile_aot_module()
to produce PTX/CUBIN during Docker image builds where no GPU is available
(GH-1085).wp.compile_aot_module() when the output binary already exists
and wp.config.cache_kernels is enabled, matching the caching behavior of Module.load()
(GH-1246).wp.int32) when indexing vectors and matrices in kernels
(GH-1209).wp.synchronize_stream() to be called on CPU devices without raising an exception
(GH-1225).wp.tile_load() / wp.tile_store() for non-power-of-two tile sizes,
particularly when source arrays fit within L2 cache (GH-1239).warp.fem: Allow the cells() operator to accept traced fields,
returning the underlying cell-level field for evaluation at cell-space samples (e.g., from lookup()).warp.fem: Allow PicQuadrature particles to span multiple cells
by passing a tuple of 2D arrays (cell_indices, coords, particle_fraction)
to specify per-particle cell contributions.Vector and Matrix generic type parameter order to dtype-first: Vector[Scalar, Length]
(was Vector[Length, Scalar]) and Matrix[Scalar, Rows, Cols] (was Matrix[Rows, Cols, Scalar]).
This only affects type stubs (.pyi) and internal TypeVar annotations
(GH-1216).wp.foo.bar.tid() resolved to wp.tid()) and raise an error instead
(GH-1198).wp.tile_assign() support for assigning register tiles (e.g., wp.tile_map() outputs) into shared tile views,
and fix reverse-mode gradients for overwritten destinations (GH-1232).module="unique" kernels incorrectly reusing cached kernels
when wp.static() expressions are deferred (e.g., referencing loop variables),
causing wrong kernel execution (GH-1211).@wp.func losing parameter type information in Pyright/Pylance
(GH-1219).x += expr, x *= expr, etc.) on scalar variables evaluating the RHS expression twice,
generating redundant loads and arithmetic in compiled kernels (GH-1230).jax_kernel() and jax_callable().BsrMatrix.notify_nnz_changed sometimes reading a stale non-zero count
from the device offsets array,
which could cause the matrix to report an incorrect nnz
and under-allocate internal buffers.warp.fem: Fix reference cycles in borrow_temporary() that prevented GPU memory
from being reclaimed until garbage collection,
causing unbounded memory growth in long-running simulations
(GH-1075).warp.fem: Fix uninitialized memory accesses in point-based function spaces with variable nodes per element.warp/examples/core/example_fft_poisson_navier_stokes_2d.py)
demonstrating a vorticity-streamfunction solver with tile-based FFT Poisson solve and SSP-RK3 timestepping.wp.array[wp.vec3]
instead of wp.array(dtype=wp.vec3)) in kernel and function signatures
(GH-1216).--num_frames becomes --num-frames)
(GH-1213).The following feature is deprecated and will be removed in v1.12 :
Warp v1.11.1 is a bugfix release following v1.11.0. For a complete list of changes, see the changelog.
This is primarily a bugfix release with no major new features. Key fixes include:
Tile Operations: Fixed wp.tile_matmul() sometimes producing NaN results when using the c = wp.tile_matmul(a, b) form due to reading uninitialized output memory (#1180). Also fixed tile multiplication with scalar constants when one operand is a vector or matrix type (#1175), and enabled scalar, vector, and matrix arguments in wp.tile_map() (#1136).
Code Generation: Fixed wp.static() incorrectly resolving loop variables to same-named global Python variables when used for static loop unrolling in kernels (e.g., wp.static(i) inside for i in range(n) would use a global i if one existed, instead of the loop iteration value) (#1139). Also fixed a segfault in conditional expressions (ternary if/else) when one branch accesses an array element and the other branch is taken.
CUDA Graphs: Fixed CUDA graphs with multiple temporary allocations using more memory than necessary. Previously, memory freed during graph capture wasn't properly sequenced for reuse by later allocations, causing memory to accumulate (e.g., three sequential 1GB allocations would consume 3GB instead of reusing the same 1GB).
Developer Experience: Fixed @wp.func decorated functions showing generic _Wrapped types in Pyright/Pylance instead of their actual signatures on Python 3.10+. Also fixed multiple issues with IDE autocomplete stubs that caused type checker errors in mypy and Pyright, including incorrect @overload usage, shadowed bool type references, and missing Literal[] syntax for integer type parameters.
Documentation: Added missing docstrings across the API, including built-in functions, type constructors, constants, and warp.fem symbols (#1159). Added example_particle_repulsion.py demonstrating how to use wp.grad() (#1137).
The following feature is deprecated and will be removed in v1.12:
Constructing matrices from column vectors via wp.matrix(): The ability to construct matrices by passing column vectors to wp.matrix() has been deprecated at both Python and kernel scopes. Use wp.matrix_from_cols() instead.
Constructing matrices from vectors via wp.matrix(): The ability to construct matrices by passing vectors to wp.matrix() has been deprecated at both Python and kernel scopes. The replacement depends on the scope because the behavior was inconsistent:
# Kernel scope (vectors were interpreted as columns)
# Deprecated (will be removed in v1.12):
m = wp.mat33(col0, col1, col2)
# Use instead:
m = wp.matrix_from_cols(col0, col1, col2)
# Python scope (vectors were interpreted as rows)
# Deprecated (will be removed in v1.12):
m = wp.mat22(wp.vec2(1, 2), wp.vec2(3, 4))
# Use instead:
m = wp.matrix_from_rows(wp.vec2(1, 2), wp.vec2(3, 4))We thank the following contributors:
build_lib.py CLI flags to use kebab-case for consistency (e.g., --cuda_path to --cuda-path,
--llvm_source_path to --llvm-source-path). Rename --libmathdx to --use-libmathdx for clarity.wp.tile_map() (GH-1136).wp.tile_matmul() reading from uninitialized output tile when using c = wp.tile_matmul(a, b)
(GH-1180).wp.static() incorrectly capturing global Python variables instead of loop variables when used inside for-loops
in kernels (GH-1139).if/else) when one branch accesses an array element
and the other branch is taken (GH-1094).wp.autograd.gradcheck() and wp.autograd.gradcheck_tape() support for kernels involving arrays with data types
other than single-precision floats (GH-1113).@wp.func decorated functions showing generic _Wrapped types in Pyright/Pylance instead of their actual
signatures on Python 3.10+ (GH-1163).zeros()). The stub generator now detects conflicts and
generates merged @overload definitions where appropriate (GH-1156).verbose flag in wp.capture_debug_dot_print() (GH-1202).--llvm-path build option to use existing LLVM installation when building warp-clang library instead of
downloading from packman.bsr_get_diag() not zeroing the output buffer when provided, causing missing diagonal blocks to
retain stale values instead of being set to zero (GH-1170).example_particle_repulsion.py to optimization examples, demonstrating how to use wp.grad()
(GH-1137).warp.fem symbols (GH-1159).Your coding agent can read these notes before it upgrades. Set up the MCP server →