NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3905 most downloaded on PyPI
A Python framework for high-performance simulation and graphics programming
Last release 1 months ago
31 Aug 2026
Ships on a steady schedule
a new release about every 4 weeks
Nearly every release is documented
notes for 52 of 52 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
52 releases · first in 2022
One column per quarter.
Symbols and namespaces intended for internal use now emit deprecation warnings when accessed and will be removed in v1.13 (nominally May 2026). The co…
Warp v1.11 introduces group-aware spatial queries for multi-world workloads, provides new options for managing JIT compilation overhead, and expands differentiation capabilities with wp.grad(). This release also includes expanded tile operations, the unpack operator in kernels, C++ integration examples, and a major API cleanup clarifying public versus internal interfaces.
Warp v1.11 introduces group-aware construction and queries for wp.Bvh and wp.Mesh data structures, enabling efficient spatial queries across multiple independent environments. This feature allows you to build a single acceleration structure containing geometry from multiple worlds or scenes, then query each world independently without traversing primitives from other worlds.
When constructing a BVH or Mesh, assign each primitive to a group using the groups parameter. Warp builds isolated sub-trees for each group within a unified structure:
# Build a BVH in Python containing multiple worlds
lowers = wp.array(...) # Shape bounds for all worlds
uppers = wp.array(...)
world_ids = wp.array([0, 0, 1, 1, 2, 2, ...], dtype=int)
bvh = wp.Bvh(lowers, uppers, groups=world_ids)
@wp.kernel
def raycast_world(
bvh_id: wp.uint64,
world_id: int,
ray_origin: wp.vec3,
ray_dir: wp.vec3
):
# Get the root node for this world's sub-tree
root = wp.bvh_get_group_root(bvh_id, world_id)
# Query only intersects geometry from this world
query = wp.bvh_query_ray(bvh_id, ray_origin, ray_dir, root)
# Process hits
shape_idx = int(0)
while wp.bvh_query_next(query, shape_idx):
# Handle intersection with shape_idx
pass
# Launch kernel to query world 2
wp.launch(raycast_world, dim=1, inputs=[bvh.id, 2, origin, direction])This example shows a single-world query for clarity. For production use, launch multiple threads in parallel, each querying its assigned world from arrays of world IDs and ray parameters. See Newton's raytrace implementation for a real-world example of parallel multi-world raycasting.
groups array during construction to organize primitives into isolated sub-treesroot parameter to limit traversal to a specific groupwp.bvh_get_group_root() and wp.mesh_get_group_root() retrieve sub-tree roots for each groupThanks to @StafaH for implementing this feature.
Warp v1.11 adds several new query functions and improvements for spatial queries:
wp.mesh_query_ray_anyhit(): Fast any-hit query that returns immediately upon finding any intersection, useful for shadow ray calculations in renderingwp.mesh_query_ray_count_intersections(): Counts all ray-triangle intersections along a ray pathwp.mesh_query_point_sign_parity(): Point-in-mesh query using perturbed ray casting with majority voting for improved robustness in challenging casesmax_dist parameter: wp.bvh_query_next() now accepts a maximum distance to filter intersections, useful for early ray terminationwp.bvh_query_aabb_tiled(), wp.bvh_query_ray_tiled(), wp.mesh_query_aabb_tiled(), etc.)wp.grad() directly evaluates the gradient of a Warp function at specific input values, computing gradients inline during the forward pass. This is useful for computing forces from energy functions or when implementing custom adjoints that need to call auto-generated gradients of subfunctions, avoiding the need to manually code the entire adjoint chain. This contrasts with wp.Tape(), which records an entire computation graph for reverse-mode automatic differentiation across multiple kernel launches. This feature was implemented in response to community feedback (#125).
import warp as wp
import numpy as np
k = 1.0
@wp.func
def compute_energy(x: float):
return 0.5 * k * x * x
@wp.kernel
def compute_force(x: wp.array(dtype=float), U: wp.array(dtype=float), F: wp.array(dtype=float)):
i = wp.tid()
U[i] = compute_energy(x[i])
F[i] = -wp.grad(compute_energy)(x[i])
N = 5
x = wp.array(np.arange(N, dtype=np.float32), dtype=float)
U = wp.zeros_like(x)
F = wp.zeros_like(x)
wp.launch(compute_force, N, inputs=[x], outputs=[U, F])
print(U.numpy()) # Energy: [0. 0.5 2. 4.5 8. ]
print(F.numpy()) # Force: [ 0. -1. -2. -3. -4.]wp.tile_map() supports n-ary maps (up to n=8)User-defined functions that accept up to 8 arguments may now be used as tile mapping functions. An equivalent number of tiles must be passed to wp.tile_map(). For example:
@wp.func
def weighted_sum(a: float, b: float, c: float):
return 0.5 * a + 0.3 * b + 0.2 * c
@wp.kernel
def compute():
a = wp.tile_arange(0.0, 1.0, 0.1, dtype=float)
b = wp.tile_ones(shape=10, dtype=float)
c = wp.tile_arange(1.0, 2.0, 0.1, dtype=float)
s = wp.tile_map(weighted_sum, a, b, c)
print(s)
wp.launch_tiled(compute, dim=[1], inputs=[], block_dim=16)wp.tile_randf() and wp.tile_randi() have been introduced to generate tiles of random floats and ints, respectively. These functions accept optional lower and upper bound arguments to control the range of generated values. This snippet generates 4x4 tensors of random floats using 2x2 tiles:
TILE_M, TILE_N = 2, 2
M, N = 2, 2
seed = 42
@wp.kernel
def rand_kernel(seed: int, x: wp.array2d(dtype=float)):
i, j = wp.tid()
rng = wp.rand_init(seed, i * TILE_M + j)
t = wp.tile_randf(shape=(TILE_M, TILE_N), rng=rng)
wp.tile_store(x, t, offset=(i * TILE_M, j * TILE_N))
x = wp.zeros(shape=(M * TILE_M, N * TILE_N), dtype=float)
wp.launch_tiled(rand_kernel, dim=[M, N], inputs=[seed, x], block_dim=32)
print(x.numpy())wp.tile_matmul()Optional alpha and beta scaling arguments have been added to wp.tile_matmul() builtins.
| Previous Behavior | Updated Behavior |
|---|---|
out = A * B + out |
out = alpha * A * B + beta * out |
out = A * B |
out = alpha * A * B |
wp.tile_cholesky_inplace(), wp.tile_cholesky_solve_inplace(), wp.tile_lower_solve_inplace(), and wp.tile_upper_solve_inplace() give the same results as their non-inplace counterparts, but overwrite input memory rather than allocate additional output memory, thereby halving shared memory usage. This is particularly beneficial in memory-constrained kernels where shared memory is limited. A standard example using Cholesky decomposition and the Cholesky solver looks like:
@wp.kernel()
def tile_math_cholesky_inplace(
gA: wp.array2d(dtype=wp.float64),
gy: wp.array1d(dtype=wp.float64),
):
i, j = wp.tid()
# Load A & y
a = wp.tile_load(gA, shape=(TILE_M, TILE_M), storage="shared")
y = wp.tile_load(gy, shape=TILE_M, storage="shared")
# Compute L st LL^T = A inplace
wp.tile_cholesky_inplace(a)
# Solve for y in LL^T x = y inplace
wp.tile_cholesky_solve_inplace(a, y)
# Store L & y
wp.tile_store(gA, a)
wp.tile_store(gy, y)Warp v1.11 brings three changes that aim to reduce the time to compile and load modules:
The CUDA C++ files that are generated from the Python modules all include the same set of header files. Warp now leverages NVRTC precompiled headers to cache the result of parsing these headers and reuse it for subsequent modules.
The first module that gets compiled incurs a 50 ms overhead to create the precompiled header, but every subsequent module in the same Python session gains 50-500 ms in compile time, with larger modules seeing the greatest benefit. The precompiled header is stored in a temporary directory and cached for the lifetime of the Python process. Each new Python process must recreate the precompiled header, as PCH files cannot be shared across processes due to internal memory layout requirements.
This feature is enabled by default, but can be disabled using wp.config.use_precompiled_headers=False.
Note for source builds: Precompiled headers require building Warp against CUDA Toolkit 12.8 or newer. Users installing from PyPI automatically have this feature because the Warp libraries on PyPI are now built against CUDA Toolkit 12.9.1.
For more details, see the NVRTC PCH documentation.
By default, the CUDA Runtime Compiler performs a high level of optimizations on GPU kernels, favoring runtime performance at the cost of longer compilation times. Warp v1.11 introduces the wp.config.optimization_level setting to control this tradeoff. When set to None (the default), Warp uses level 3, which corresponds to maximum runtime optimization.
This setting controls GPU kernel compilation and accepts values from 0 to 3:
The setting can be configured globally via wp.config.optimization_level or per-module using wp.set_module_options({"optimization_level": 2}).
This setting is available when Warp is built against CUDA Toolkit 12.9 or newer. PyPI wheels include this support.
Modules can now be compiled and loaded in parallel across multiple threads for both CPU and GPU. To benefit from parallel compilation, set wp.config.load_module_max_workers to a positive integer (default is 0, which disables parallelization) and explicitly load modules using wp.load_module() or wp.force_load(). You can also pass a max_workers argument directly to these functions to override the config setting. When modules are lazily compiled on-demand at wp.launch(), they are compiled one at a time and do not benefit from parallelization.
Parallel compilation can significantly reduce startup time when working with many modules. The most time-consuming step being parallelized is the two-stage compilation process: translating Python code to CUDA/C++ source code, then JIT-compiling to binary libraries using NVRTC for GPU or LLVM for CPU.
wp.load_module() requires recursive=True to enable parallel compilation. This loads the specified module along with all its submodules. The following example loads Newton's inverse kinematics module and its submodules:
import warp as wp
import newton._src.sim.ik
wp.config.load_module_max_workers = 4
wp.load_module(newton._src.sim.ik, device=wp.get_device("cuda:0"), recursive=True)With parallel compilation (wp.config.load_module_max_workers=4), this takes about 3 seconds for CUDA (versus 4.5 seconds serially) and 3.7 seconds for CPU (versus 10.5 seconds serially).
wp.force_load() provides lower-level control by accepting an explicit list of modules to compile, unlike wp.load_module() which operates on a module hierarchy. Warning: Without explicit modules and device arguments, wp.force_load() compiles all imported modules for all available devices, which can take much longer than not using it at all. For example, import newton followed by wp.force_load(device="cuda:0") will compile over 100 modules.
The following example shows selective compilation using a manually-built module list. This requires more setup but provides fine-grained control:
import warp as wp
wp.config.load_module_max_workers = 4
module_list = [
wp.get_module(m)
for m in [
"newton._src.sim.articulation",
"newton._src.sim.ik.ik_common",
"newton._src.sim.ik.ik_lm_optimizer",
"newton._src.sim.ik.ik_objectives",
"newton._src.sim.ik.ik_solver",
"newton._src.geometry.inertia",
"newton._src.viewer.gl.opengl",
"newton._src.viewer.kernels",
"newton._src.solvers.mujoco.kernels",
]
]
wp.force_load(device=wp.get_device("cpu"), modules=module_list)The above snippet takes about 9-10 seconds with parallel compilation (wp.config.load_module_max_workers=4), compared to about 24 seconds with serial loading (wp.config.load_module_max_workers=0), yielding a roughly 2.5x speedup in this case.
Parallel compilation is most effective when compilation time is distributed evenly across modules. Gains will be limited if a single module dominates the total compilation time.
Advanced optimization: For applications with many kernels in a single large module file, consider splitting them into separate submodules across multiple files. This enables parallel compilation of the submodules, trading some code organization complexity for faster compilation times.
Warp now supports Python's unpack operator (*) inside kernel function calls, enabling you to expand vectors, matrices, quaternions, and 1D array slices into individual arguments. This feature brings familiar Python idioms into Warp kernels and simplifies common patterns like constructing larger vectors from smaller ones, or copying array values into a vector.
The unpack operator works on composite types to expand their components:
@wp.kernel
def compute(
arr: wp.array(dtype=float),
):
# Unpack a 1D array slice into a vector.
v1 = wp.vec3(*arr[:3])
wp.expect_eq(v1, wp.vec3(1.0, 2.0, 3.0))
# Unpack a vector into function arguments.
v2 = wp.vec2(1.0, 2.0)
x2 = wp.max(*v2)
wp.expect_eq(x2, 2.0)
# Build larger vectors by unpacking smaller ones.
v3 = wp.vec3(1.0, 2.0, 3.0)
v4 = wp.vec4(*v3, 4.0)
wp.expect_eq(v4, wp.vec4(1.0, 2.0, 3.0, 4.0))
arr = wp.array((1, 2, 3, 4, 5, 6, 7, 8, 9), dtype=float)
wp.launch(compute, dim=1, inputs=(arr,))Important: When unpacking arrays, slice bounds must be compile-time constants and non-negative. The upper bound is required since the array length is not known at compile time.
Warp v1.11 introduces C++ integration examples demonstrating how to deploy Warp-compiled kernels in standalone C++ applications without runtime Python dependencies. Two approaches are demonstrated:
00_cubin_launch): Load Warp-generated CUBIN files at runtime using CUDA Driver API01_source_include): Statically include Warp-generated CUDA source (including forward and adjoint kernels) in C++ projectsThese examples are available at warp/examples/cpp/ with full build support for Make and CMake on Linux and Windows.
Community feedback: We're gathering input on Warp's AOT and C++ interoperability roadmap through a survey on GitHub Discussions. If you work with native workflows, deployment in minimal-Python environments, or interoperability with other CUDA/C++ libraries, your feedback will help shape future development in these areas.
Warp v1.11 refines the boundary between public and internal APIs alongside a major documentation reorganization. Symbols and namespaces intended for internal use now emit deprecation warnings when accessed and will be removed in v1.13 (nominally May 2026). The complete public API is now clearly documented in the restructured API Reference and Language Reference sections.
What this means for your code:
wp.config.verbose_warnings = True.warp package (e.g., wp.context) is deprecated. If your code relies on internal APIs like wp.context.runtime, you can access them via wp._src.context.runtime, but be aware these are not part of the public API and may change or be removed without notice.wp.DeviceLike instead of the deprecated wp.Devicelike (note the capitalization).If you depend on functionality that's no longer accessible and believe it should be part of the public API, please open a feature request on GitHub. Note: We're aware that some functionality (such as graph coloring and color balancing) currently lacks a public API and requires accessing internal modules. We're tracking these gaps in issues like #1145.
Warp 1.11.0 drops support for Python 3.8, which reached end-of-life in October 2024. Python 3.9 is now the minimum supported version.
PyPI wheels are now built with CUDA Toolkit 12.9.1 (up from 12.8.0 in previous releases). This enables new optimizations and features, including the wp.config.optimization_level setting for controlling kernel compilation.
For users building Warp from source, CUDA Toolkit 12.9.1 or newer is recommended for full GPU support.
We also thank the following contributors from outside the core Warp development team:
wp.Bvh and wp.Mesh, a new built-in function forwp.mesh_query_ray_anyhit(), support for the max_dist argument to wp.bvh_query_next(), and improving the BVH SAHalpha/beta scaling parameters towp.tile_matmul(), reducing shared memory usage and enabling operation fusion in tile kernels.For a curated list of all changes in this release, please see the v1.11.0 section in CHANGELOG.md.
wp.grad() to get a Warp function's gradient in functions, kernels, and custom gradients
(GH-125).wp.Bvh and wp.Mesh to support multi-environment workloads,
including groups constructor argument, optional root query parameter for group-restricted traversal,
and helper functions wp.bvh_get_group_root() and wp.mesh_get_group_root() to retrieve subtree roots
(GH-1074, GH-1097).wp.bvh_query_aabb_tiled(), wp.bvh_query_ray_tiled(), wp.bvh_query_next_tiled(),
wp.mesh_query_aabb_tiled(), and wp.mesh_query_aabb_next_tiled(). Aliases with tile_* prefix
(e.g., wp.tile_bvh_query_ray()) are also available (GH-1005).wp.mesh_query_ray_anyhit() for ray any-hit queries with group-aware support
(GH-1097).max_dist argument to wp.bvh_query_next() to control the maximum distance to check for intersections in ray
queries (GH-1052).wp.mesh_query_ray_count_intersections() to count all intersections with triangles along a ray
(GH-938).wp.mesh_query_point_sign_parity() for mesh point queries with parity-based inside/outside determination using
ray–triangle intersection counting (GH-938).alpha and beta scaling parameters to wp.tile_matmul() to reduce shared memory usage and fuse operations,
supporting C = alpha * A * B + beta * C and C = alpha * A * B
(GH-1023).wp.tile_cholesky_inplace(), wp.tile_cholesky_solve_inplace(), wp.tile_lower_solve_inplace(), and
wp.tile_upper_solve_inplace() (GH-1025).wp.int64 and wp.uint64 key types to wp.tile_sort()
(GH-1089).wp.tile_scan_max_inclusive() and wp.tile_scan_min_inclusive() for cumulative maximum and minimum operations
across tiles (GH-1090).wp.tile_randi() and wp.tile_randf() to support generating tiles of random numbers
(GH-1010).launch_bounds parameter to @wp.kernel decorator to specify CUDA __launch_bounds__ attribute for controlling
thread block occupancy. Can be an integer for maxThreadsPerBlock or a tuple of 1-2 integers for
(maxThreadsPerBlock, minBlocksPerMultiprocessor) (GH-1049).wp.types.is_vector(), wp.types.type_is_tile()).
See the warp.types documentation for the complete list.*) in kernels to expand vectors, matrices, quaternions,
and 1D array slices into individual arguments, enabling syntax like wp.vec4(*v3, 1.0), wp.max(*v),
or wp.vec3i(*arr[:3]) (GH-1083).wp.config.optimization_level to control kernel compilation optimization level, allowing trade-offs between
kernel compile time and runtime performance. Can be configured globally or per-module. Currently only affects
kernels compiled for GPU devices (GH-1084).max_workers parameter in wp.load_module() and
wp.force_load(). Defaults to serial loading; set wp.config.load_module_max_workers to enable parallel loading
(GH-1086).wp.config.use_precompiled_headers=False if needed (GH-595).GraphMode.WARP_STAGED and GraphMode.WARP_STAGED_EX) for JAX FFI to avoid
re-capturing graphs when input/output buffer addresses change between calls
(GH-1039).graph_compatible argument from warp.jax_experimental.jax_callable() (deprecated since 1.8.1).
Use graph_mode instead (GH-982).wp.isfinite(), wp.isnan(), and wp.isinf(). These functions will only
accept floating-point arguments in a future release (GH-847).warp.render.UsdRenderer.update_body_transforms() which references non-existent attributes and is
non-functional.wp.tile_map() to support user-defined functions with up to 8 input tiles, beyond the previous unary and binary
operations (GH-1077).wp.map() performance by caching kernels for repeated calls with the same function and similar input arguments
(GH-1108).array.view()
(GH-1112).wp.config.verbose is enabled or when compilation errors occur
(GH-1129).warp.fem.interpolate() API to decouple interpolation location from storage definition, allowing more flexible
usage patterns (GH-1091).wp.config.enable_tiles_in_stack_memory to use architecture-aware defaults. When None (the default),
automatically enables stack memory on aarch64 platforms and disables it on other architectures.OverflowError when slicing arrays with np.int32 strides
(GH-1120).RecursionError when calling repr() on vector instances (e.g., repr(wp.vec2i(42)))
(GH-1124).array.list() returning NumPy scalar types instead of Python native types (e.g., numpy.int64 instead of int)
(GH-1099).wp.tile_zeros() with a struct type containing an array field
(GH-1128).struct types
(GH-1135).warp.optim.SGD causing wrong momentum and weight decay behavior
(GH-1143).warp.render.OpenGLRenderer window incorrectly appearing when pyglet.options["headless"] is True but the
headless constructor parameter is not provided.The following feature is deprecated and will be removed in v1.11 (planned for January 2026):
Warp v1.10.1 is a bugfix release following v1.10.0. For a complete list of changes, see the changelog.
This is primarily a bugfix release with no major new features. Key fixes include:
module="unique": Fixed kernels using @wp.kernel(module="unique") to properly reuse existing module objects when the kernel is defined multiple times, avoiding unnecessary module creation overhead.wp.zeros() inside kernels, including .ptr access, indexing for subarrays, and accepting single integers for the shape parameter.@wp.func_grad) from compiling when used with nested function calls.wp.fem.TemporaryStore during tape capture and resolved reference cycles in wp.fem.Temporary and wp.fem.ShapeBasisSpace.The following feature is deprecated and will be removed in v1.11 (planned for January 2026):
graph_compatible parameter in jax_callable(): The boolean graph_compatible flag has been deprecated in favor of the new graph_mode parameter which accepts GraphMode enum values. Use GraphMode.JAX, GraphMode.WARP, or GraphMode.NONE instead.
# Deprecated (v1.10.1, will be removed in v1.11)
callable = wp.jax_experimental.jax_callable(func, graph_compatible=True)
# Use instead
from warp.jax_experimental import GraphMode
callable = wp.jax_experimental.jax_callable(func, graph_mode=GraphMode.JAX)
We thank the following contributors:
wp.static() expressions that prevented certain code patterns from compiling correctly.module="unique" kernels to properly reuse existing module objects when defined multiple times,
avoiding unnecessary module creation overhead (GH-995).wp.static() expressions that use the loop variable.
These loops are now always unrolled regardless of max_unroll settings to
ensure loop variables are available as compile-time constants within static
expressions (GH-560).wp.compile_aot_module() to detect generic kernels without overloads and generic kernels with
multiple overloads when strip_hash=True (GH-919).wp.tile_load_indexed() when indices tile has been reshaped or transformed
(GH-1008).wp.zeros() in kernels):
@wp.func_grad) when used with nested function calls
(GH-967).wp.fem.TemporaryStore during tape capture for automatic differentiation
(GH-1021).wp.fem.Temporary and wp.fem.ShapeBasisSpace
(GH-1076).wp.fem.lookup() and related functionality
(GH-1072).wp.tile() (GH-1042).Breaking change: As part of this optimization, support for passing lists, tuples, and other non-Warp array arguments to built-in functions has been re…
Warp v1.10 expands JAX integration with automatic differentiation support and multi-device jax.pmap() compatibility. The tile programming model has been enhanced with axis-specific reductions, component-level indexing, and convenience functions for creating tiles.
Performance has been significantly improved in several areas: BVH operations now support in-place rebuilding for CUDA graphs and configurable leaf sizes, built-in function calls from Python are up to 70× faster, and additional sparse matrix and FEM operations can now be captured in CUDA graphs.
Additional usability improvements include negative indexing and slicing for arrays, atomic bitwise operations, and new built-in functions including error functions and type casting.
Important: This release removes the warp.sim module (deprecated since v1.8), which has been superseded by the Newton physics engine. See the Announcements section below for migration guidance and other upcoming changes.
For a complete list of changes, see the full changelog.
Warp now supports experimental automatic differentiation with JAX, allowing kernels to participate in JAX automatic differentiation workflows. This feature is contributed by @mehdiataei and builds on earlier work by @jaro-sevcik. It enables computing gradients through Warp kernels using jax.grad() by passing enable_backward=True to jax_kernel().
Key capabilities include:
wp.vec2 or wp.mat22 are fully supportedjax.pmap() for distributed forward and backward passes across multiple GPUsimport jax
from warp.jax_experimental import jax_kernel
@wp.kernel
def my_kernel(a: wp.array(dtype=float), out: wp.array(dtype=float)):
i = wp.tid()
out[i] = a[i] ** 2.0
# Enable automatic differentiation
jax_func = jax_kernel(my_kernel, num_outputs=1, enable_backward=True)
# Compute gradients through the kernel
grad_fn = jax.grad(lambda a: jax.numpy.sum(jax_func(a)[0]))
gradient = grad_fn(input_array) # gradient: [2*a[0], 2*a[1], ...]
This feature is experimental and has some current limitations. See the JAX Automatic Differentiation documentation for complete examples, usage details, and limitations.
jax.pmap()Warp now properly supports jax.pmap() and jax.shard_map() for multi-device parallel execution, thanks to fixes contributed by @chaserileyroberts. Previously, device targeting issues prevented Warp callables from working correctly within these JAX primitives—JAX would invoke callbacks from multiple threads targeting different devices, but Warp would always execute on the default device. The fix ensures proper device coordination by extracting device ordinals from XLA FFI and adding thread synchronization for concurrent callbacks, enabling efficient data-parallel workflows across multiple GPUs.
A new wp.Bvh.rebuild() method enables rebuilding BVH hierarchies in-place without allocating new memory. This complements the existing refit() method and is particularly useful when primitive distributions change significantly.
CUDA graph capture: Unlike creating a new BVH, rebuild() reuses existing buffers, making it safe to capture in CUDA graphs. Previously captured graphs that include queries on the BVH remain valid after rebuilding, enabling high-performance repeated updates without graph re-capture overhead.
Construction algorithms: On CUDA devices, in-place rebuild supports "lbvh" only. On CPU, "sah" and "median" are supported. Defaults are chosen automatically based on the device.
The tile programming model has been enhanced with new capabilities to make tile-based computations more expressive and convenient:
The tile-reduction functions wp.tile_reduce() and wp.tile_sum() now support an optional axis parameter, enabling reductions along a specific dimension of a tile rather than reducing the entire tile to a single value. This enhancement brings NumPy-like axis semantics to tile operations.
@wp.kernel
def tile_reduce_axis(x: wp.array2d(dtype=float), y: wp.array(dtype=float)):
a = wp.tile_load(x, shape=(4, 8), storage="shared")
# Sum along axis 0, reducing shape from (4, 8) to (8,)
b = wp.tile_sum(a, axis=0)
wp.tile_store(y, b)
x = wp.array(np.arange(32).reshape(4, 8), dtype=float)
# x = [[ 0. 1. 2. 3. 4. 5. 6. 7.]
# [ 8. 9. 10. 11. 12. 13. 14. 15.]
# [16. 17. 18. 19. 20. 21. 22. 23.]
# [24. 25. 26. 27. 28. 29. 30. 31.]]
y = wp.zeros(8, dtype=float)
wp.launch_tiled(tile_reduce_axis, dim=(1,), inputs=[x], outputs=[y], block_dim=32)
# y = [48. 52. 56. 60. 64. 68. 72. 76.] (column sums)
Tiles of composite types (vectors, matrices, quaternions) now support component-level indexing and assignment. You can directly index into individual components using extended indexing syntax:
tile[i][1] extracts the second component of a vector at position itile[i][1, 1] accesses the element at row 1, column 1 of a matrix at position iThis provides more convenient and expressive syntax for working with structured data in tiles.
The new wp.tile_full() function provides a convenient way to create tiles initialized with a constant value, similar to NumPy's np.full():
# Create an 8x8 tile filled with 3.14
tile = wp.tile_full(shape=(8, 8), value=3.14, dtype=float)
The new example_tile_mcgp.py example demonstrates tile-based Monte Carlo methods by implementing a walk-on-spheres algorithm for solving Laplace's equation on volumetric domains.
Calling Warp built-in functions from Python scope (e.g., wp.normalize(), wp.transform_identity(), matrix arithmetic like mat * mat) is now significantly faster thanks to optimizations in overload resolution. Previously, each function call would iterate through all overloads, attempt argument binding, and pack parameters into C types until finding a match. Now, Warp caches the resolved overload and parameter packing strategy based on argument types using @functools.lru_cache, eliminating redundant resolution overhead on subsequent calls.
In microbenchmarks, repeated wp.mat44 multiplication at Python scope is up to 70× faster (~570 μs → ~8 μs), while operations like wp.transform_identity() see 3-4× speedups (~100 μs → ~30 μs). The magnitude of improvement varies by operation complexity, with greater gains for operations requiring more expensive overload resolution.
Breaking change: As part of this optimization, support for passing lists, tuples, and other non-Warp array arguments to built-in functions has been removed. Calls like wp.normalize([1.0, 2.0, 3.0]) must now be written as wp.normalize(wp.vec3(1.0, 2.0, 3.0)). This simplifies the function call path and removes expensive sequence-flattening logic that was incompatible with efficient caching.
wp.Bvh and wp.Mesh now expose tunable leaf_size and bvh_leaf_size parameters, respectively, allowing users to control the number of primitives stored in each leaf node for performance optimization. The optimal leaf size depends on the query workload:
Behavior change: The default leaf_size for wp.Bvh has changed from 4 (hardcoded) to 1, optimizing for intersection queries which are more common. wp.Mesh retains a default bvh_leaf_size of 4 as a compromise between intersection and closest-point query performance. Users performing primarily closest-point queries may benefit from explicitly setting larger leaf sizes.
Sparse matrix operations in warp.sparse can now be captured in CUDA graphs for allocation-free execution. Operations like bsr_axpy(), bsr_assign(), and bsr_set_transpose() preserve matrix topology when using masked=True, while bsr_mm() adds a new max_new_nnz parameter that allows specifying an upper bound on new non-zero blocks for flexible graph capture when sparsity patterns vary within known bounds.
Building warp.fem geometry and function space partitions can now be captured in CUDA graphs by specifying upper bounds on partition sizes: max_cell_count and max_side_count for ExplicitGeometryPartition, and max_node_count for make_space_partition(). Additionally, building fields and restrictions is now synchronization-free by default.
Warp arrays now support negative indexing and improved slicing behavior, making array manipulation more intuitive and consistent with NumPy conventions.
Negative indexing: Access elements from the end of an array using negative indices:
@wp.kernel
def use_negative_indexing(arr: wp.array(dtype=float)):
last = arr[-1] # Last element
second_last = arr[-2] # Second-to-last element
Enhanced array slicing: Arrays now support more flexible slicing operations within kernels, including stride-based access patterns. This works with both regular arrays and tile operations:
@wp.kernel
def tile_load_strided(input: wp.array2d(dtype=float), output: wp.array2d(dtype=float)):
# Load every other element from a 16x16 region into an 8x8 tile
tile = wp.tile_load(input[::2, ::2], shape=(8, 8))
wp.tile_store(output, tile)
input = wp.array(np.arange(256).reshape(16, 16), dtype=float)
output = wp.zeros((8, 8), dtype=float)
wp.launch_tiled(tile_load_strided, dim=(1,), inputs=[input, output], block_dim=32)
# output contains every other element from input:
# [[ 0. 2. 4. 6. 8. 10. 12. 14.]
# [ 32. 34. 36. 38. 40. 42. 44. 46.]
# [ 64. 66. 68. 70. 72. 74. 76. 78.]
# ...
# [224. 226. 228. 230. 232. 234. 236. 238.]]
wp.erf(), wp.erfc(), wp.erfinv(), and wp.erfcinv() for error function computationswp.cast() to reinterpret values as different types while preserving bit patterns (e.g., reinterpreting float bits as int)wp.atomic_and(), wp.atomic_or(), and wp.atomic_xor() for thread-safe bitwise operations on integers, contributed by @j3soonwp.sparse.bsr_row_index() and wp.sparse.bsr_block_index() as kernel-level functions to efficiently determine which row a given block belongs to without manually searching through the compressed offset arrayFixed segmentation faults when running tile-based kernels on AArch64 CPUs, affecting platforms including NVIDIA Jetson (Thor, Orin), DGX Spark, Grace Hopper, and Grace Blackwell systems. The fix uses stack memory allocation instead of static memory to work around limitations in LLVM's JIT compiler.
This change is enabled by default on all CPU architectures and can be disabled if needed via wp.config.enable_tiles_in_stack_memory = False. If you encounter issues that are resolved by disabling this setting, please report them on our GitHub Issues page.
Note: This primarily affects CPU execution of tile operations, which is less common in Warp workflows but useful for debugging or scenarios in which GPU memory transfer overhead outweighs compute benefits.
Warp now performs runtime version checking to detect mismatches between the Python package and native libraries (e.g., warp.dll, warp.so). This helps diagnose issues in which multiple Warp installations on the same system may cause the wrong native libraries to be loaded. When a mismatch is detected, a warning is issued but execution continues. If you see such warnings, ensure you're loading Warp from the expected installation location and that your environment doesn't have conflicting Warp versions.
warp.sim moduleThe warp.sim module has been removed in this release. This module was formally deprecated in Warp v1.8 (July 2025) and has been superseded by the Newton physics engine, an independent package managed as a Linux Foundation project with a redesigned API focused on robotics and robot learning.
Migration: Users relying on warp.sim should migrate to Newton. For guidance on transitioning from warp.sim to Newton, please consult the Newton migration guide. The original deprecation announcement and community discussion can be found in GitHub Discussion #735.
Questions and discussions about Newton should be directed to the Newton Discussions section. Existing issues in the Warp repository concerning warp.sim will be closed.
The default implementation of jax_kernel() is now based on JAX's Foreign Function Interface (FFI), which is required for JAX version 0.8 and newer. Most users should not need to change their code, as the FFI-based version has been available since Warp 1.7 and provides better performance through CUDA graph capture. The previous custom call implementation is still available as wp.jax_experimental.custom_call.jax_kernel() for users on older JAX versions, but it is deprecated and will not work with JAX version 0.8 or later.
_src folderAs part of ongoing efforts to clarify Warp's public API surface, internal implementation code has been reorganized into a warp._src subpackage. This change helps distinguish between public APIs that users should rely on versus internal implementation details that may change without notice.
What this means for users:
warp.context, warp.types, and warp.fem remain accessible at their current paths through compatibility shims.warp._src paths in error messages and stack traces (e.g., warp._src.context instead of warp.context).warp._src.* (acknowledging the use of private APIs).This reorganization is the first step in a multi-phase effort to establish a stable public API. If you encounter any issues introduced by this reorganization, please report them on our GitHub Issues page.
The following features will be removed in v1.11 (planned for January 2026):
wp.mat22(wp.vec2(1, 2), wp.vec2(3, 4))). Use wp.matrix_from_rows() or wp.matrix_from_cols() instead. This deprecation was originally announced in v1.9 with a planned removal in v1.10, but has been extended one release cycle. While kernel-scope usage had been emitting deprecation warnings since v1.9, it was discovered that Python-scope usage lacked proper warnings. Starting in v1.10, both contexts now emit deprecation warnings.graph_compatible parameter in jax_callable(): The boolean graph_compatible parameter has been deprecated in favor of the new graph_mode parameter which accepts GraphMode enum values (GraphMode.JAX, GraphMode.WARP, or GraphMode.NONE).We also thank the following contributors from outside the core Warp development team:
struct() and overload() decoratorswp.Bvh.rebuild() method that rebuilds the hierarchy without allocating new memory and can be
captured in CUDA graphs (GH-826).wp.atomic_and() (&=), wp.atomic_or() (|=), and wp.atomic_xor() (^=) for
scalar types, along with bitwise operations for vector, matrix, and tile types
(GH-886).wp.array() type
(GH-504).wp.cast() to reinterpret a value as a different type while preserving its bit pattern
(GH-789).wp.erf(), wp.erfc(), wp.erfinv(), and wp.erfcinv()
(GH-910).wp.tile_full(), which fills a tile with a constant value (GH-973).wp.tile_reduce() and wp.tile_sum()
(GH-835).tile[i][1] for
extracting vector components, tile[i][1, 1] for matrix elements)
(GH-941).warp/examples/tile/example_tile_mcgp.py, demonstrating how to implement a Monte Carlo Laplace solver.psutil package)
(GH-985).wp.get_cuda_supported_archs() to query supported CUDA compute architectures for compilation targets
(GH-964).bsr_row_index() and bsr_block_index() to warp.sparse
(GH-895).wp.transform() when constructing with individual scalars
(GH-1011).wp.intersect_tri_tri() (GH-1015).jax.pmap() (GH-976).jax_kernel(enable_backward=True)
(GH-912, GH-515).warp.sim module and related examples. This module has been superseded by the Newton library, a separate
package with a new API. For migration guidance, see the
Newton migration guide and the original GitHub announcement
(GH-735).wp.normalize([1.0, 2.0, 3.0])
should be wp.normalize(wp.vec3(1.0, 2.0, 3.0))).RuntimeError directing them to use Warp 1.9.x or earlier
(GH-1016).wp.select() (deprecated since 1.7). Use wp.where(cond, value_if_true, value_if_false) instead.wp.matrix(pos, quat, scale) built-in function. Use wp.transform_compose() instead
(GH-980).wp.mat22(wp.vec2(1, 2), wp.vec2(3, 4))
should become wp.matrix_from_rows(wp.vec2(1, 2), wp.vec2(3, 4)))
(GH-981).jax_kernel() to be wp.jax_experimental.ffi.jax_kernel().
The previous version is still available as wp.jax_experimental.custom_call.jax_kernel(), but it is not supported
with JAX v0.8 and newer (GH-974).RuntimeError from wp.load_module() when attempting to load a module that does not contain
any Warp kernels, functions, or structs (GH-920).leaf_size parameter to wp.Bvh and bvh_leaf_size to wp.Mesh to control the number of primitives per leaf
for performance tuning. The default is now 1 for wp.Bvh and 4 for wp.Mesh, changed from a hardcoded value of
4 (GH-994).warp.sparse operations with masked=True consistent with bsr_mm() by preserving result matrix topology,
enabling CUDA subgraph capture for bsr_axpy(), bsr_assign() and bsr_set_transpose()
(GH-987).max_new_nnz argument to wp.sparse.bsr_mm() providing a synchronization-free path without further assumptions
about non-zero topology.warp.fem geometry and function space partitions is now possible in CUDA graphs by passing an explicit
upper-bound for the number of cells and nodes to ExplicitGeometryPartition and make_space_partition.
Building fields and field restrictions is now synchronization-free by default
(GH-1021).q argument in wp.transform() to the identity quaternion at the kernel scope
(GH-923).wp.bvh_query_aabb(), wp.mesh_query_aabb() and wp.bvh_query_ray().
This fixes a performance regression introduced in Warp 1.6.0 (GH-758).wp.config.enable_tiles_in_stack_memory (enabled by default)
(GH-957).scalar * array
now work correctly (previously only array * scalar worked) (GH-892).wp.atomic_add() failing to accumulate wp.int64 values (GH-977).wp.map()
(GH-984).wp.transform() constructor at the Python scope
(GH-975).struct() and overload() decorators
(GH-971).TypeError and AttributeError exceptions during Python interpreter shutdown when Warp objects are being
cleaned up, as these can be safely ignored during process termination
(GH-1048).The following features have been deprecated in prior releases and will be removed in v1.10 (early November):
Warp 1.9.1 is a bugfix release that follows our recent feature update. For a full list of changes, see the changelog.
wp.mesh_query_aabb() and wp.mesh_query_aabb_next(), added a caveat concerning the use of __cuda_array_interface__ on a system with multiple GPUs, and fixed the labeling of built-in functions that were incorrectly labeled as differentiable.arr[i:i]) are now handled correctly at the Python scope, returning an empty array instead of raising an error.wp.copy() and wp.where() now work with tiles and compute correct gradients (adjoints).TypeError that occurred when using modern tuple type hints (e.g., tuple[int, int]) with @wp.func-decorated functions on Python 3.9 and 3.10.The following features have been deprecated in prior releases and will be removed in v1.10 (early November):
warp.sim - Use the Newton engine.wp.matrix() from column vectors - Use wp.matrix_from_rows() or wp.matrix_from_cols() instead.wp.select() - Use wp.where() instead (node: different argument order).wp.matrix(pos, quat, scale) - Use wp.transform_compose() instead.We thank the following contributors for their valuable contributions to this release:
wp.copy() and wp.where() to work correctly with tile arguments (#777).wp.map() (#953).IntFlag limitations in Warp kernels
(GH-917).arr[i:i] that previously failed with indexing errors
(GH-958).TypeError: Unrecognized type 'tuple[...]' with tuple type annotations on Python 3.10
(GH-959).wp.copy(), wp.select(), and wp.where() with tiles (GH-777).#line directives being emitted for wp.map() calls during code generation
(GH-953).wp.jax_experimental.ffi.jax_kernel().The following features have been deprecated in prior releases and will be removed in v1.10 (early November):
Warp 1.9 ships with a rewritten marching cubes implementation, compatibility with the CUDA 13 toolkit, and new functions for ahead-of-time module compilation. The programming model has also been enhanced with more flexible indexing for composite types, direct IntEnum support, and the ability to initialize local arrays in kernels.
A fully differentiable wp.MarchingCubes implementation, contributed by @mikacuy and @nmwsharp, has been added. This version is written entirely in Warp, replacing the previous native CUDA C++ implementation and enabling it to run on both CPU and GPU devices. The implementation also addresses a long-standing off-by-one bug (#324). For more details, see the updated documentation.
We have added wp.compile_aot_module() and wp.load_aot_module() for more flexible ahead-of-time (AOT) compilation.
These functions include a strip_hash=True argument, which removes the unique hashes from compiled module and function
names. This change makes it possible to distribute pre-compiled modules without shipping the original Python source code.
See the documentation on ahead-of-time compilation workflows for more details. In future releases, we plan to continue to expand Warp's support for ahead-of-time workflows.
CUDA Toolkit 13.0 was released in early August.
PyPI Distribution: Warp wheels on PyPI and NVIDIA PyPI will continue to be built with CUDA 12.8 to provide a transition period for users upgrading their CUDA drivers.
CUDA 13.0 Compatibility: Users requiring Warp compiled against CUDA 13.x have two options:
Driver Compatibility: CUDA 12.8 Warp wheels can run on systems with CUDA 13.x drivers thanks to CUDA's backward compatibility.
The iterative linear solvers in warp.optim.linear (CG, BiCGSTAB, GMRES) are now fully compatible with CUDA graph capture. This adds support for device-side convergence checking via wp.capture_while(), enabling full CUDA graph capture when check_every=0. Users can now choose between traditional host-side convergence checks or fully graph-capturable device-side termination.
warp.sparse now supports arbitrary-sized blocks and can leverage tile-based computations for certain matrix types. The system automatically chooses between tiled and non-tiled execution using heuristics based on matrix characteristics (block sizes, sparsity patterns, and workload dimensions). Note that the heuristic for choosing between tiled and non-tiled variants is still being refined, and that it can be manually overridden by providing the tile_size parameter to bsr_mm or bsr_mv.
warp.fem.integrate now leverages tile-based computations for quadrature point accumulation, with automatic tile size selection based on workload characteristics. The system automatically chooses between tiled and non-tiled execution to optimize performance based on the integration problem size and complexity.
We have enhanced the support for slice operations and negative indexing across all composite types (vectors, matrices, quaternions, and transforms).
m = wp.matrix_from_rows(
wp.vec3(1.0, 2.0, 3.0),
wp.vec3(4.0, 5.0, 6.0),
wp.vec3(7.0, 8.0, 9.0),
)
subm = m[:-1, 1:]
print(subm)
# [[2.0, 3.0],
# [5.0, 6.0]]
IntEnum and IntFlag inside kernelsIt is now possible to directly reference IntEnum and IntFlag values inside Warp functions and kernels. Previously, workarounds involving wp.static() were required.
from enum import IntEnum
class JointType(IntEnum):
PRISMATIC = 0
REVOLUTE = 1
BALL = 2
@wp.kernel
def count_revolute_joints(
joint_types: wp.array(dtype=JointType),
counter: wp.array(dtype=int)
):
tid = wp.tid()
joint = joint_types[tid]
# No longer requires wp.static(JointType.REVOLUTE.value)
if joint == JointType.REVOLUTE:
wp.atomic_add(counter, 0, 1)
wp.array() views inside kernelsThis enhancement allows kernels to create array views by accessing the ptr attribute of an array.
@wp.kernel
def kernel_array_from_ptr(arr_orig: wp.array2d(dtype=wp.float32)):
arr = wp.array(ptr=arr_orig.ptr, shape=(2, 3), dtype=wp.float32)
arr[0, 0] = 1.0
arr[0, 1] = 2.0
arr[0, 2] = 3.0
Additionally, these in-kernel views now support dynamic shapes and struct types.
It is now possible to allocate local arrays of a fixed size in kernels using wp.zeros(). The resulting arrays are allocated in registers, providing fast access and avoiding global memory overhead.
Previously, developers needed to create vectors to achieve a similar capability, e.g. v = wp.vector(length=8, dtype=float), but this came with various limitations.
@wp.kernel
def kernel_with_local_array():
local_arr = wp.zeros(8, dtype=wp.float32) # Allocated in registers
# ... use local_arr
Warp now provides three new indexed tile operations that enable more flexible memory access patterns beyond simple contiguous tile operations. These functions allow you to load, store, and perform atomic operations on tiles using custom index mappings along specified axes.
wp.tile_load_indexed() - Load tiles with custom index mapping along a specified axiswp.tile_store_indexed() - Store tiles with custom index mapping along a specified axiswp.tile_atomic_add_indexed() - Perform atomic additions with custom index mapping along a specified axisx = wp.array(
[
[0.77395605, 0.43887844, 0.85859792, 0.69736803],
[0.09417735, 0.97562235, 0.7611397, 0.78606431],
[0.12811363, 0.45038594, 0.37079802, 0.92676499],
],
dtype=float,
)
indices = wp.array([0, 2], dtype=int)
@wp.kernel
def indexed_data_lookup(data: wp.array2d(dtype=float), indices: wp.array(dtype=int)):
# [0 2] = tile(shape=(2), storage=shared)
indices_tile = wp.tile_load(indices, shape=(2,))
# [[0.773956 0.438878 0.858598 0.697368]
# [0.128114 0.450386 0.370798 0.926765]] = tile(shape=(2,4), storage=register)
data_rows_tile = wp.tile_load_indexed(data, indices_tile, axis=0, shape=(2, 4))
print(data_rows_tile)
# [[0.773956 0.858598]
# [0.0941774 0.76114]
# [0.128114 0.370798]] = tile(shape=(3,2), storage=register)
data_columns_tile = wp.tile_load_indexed(data, indices_tile, axis=1, shape=(3, 2))
wp.launch_tiled(indexed_data_lookup, dim=1, inputs=[x, indices], block_dim=2)
Warp now properly supports writing to individual matrix elements stored within struct fields. Previously, operations like struct.matrix[1, 2] = value would result in a compile-time error.
@wp.struct
class MatStruct:
m: wp.mat44
@wp.kernel
def kernel_nested_mat(out: wp.array(dtype=MatStruct)):
s = MatStruct()
s.m[1, 2] = 3.0 # This now works correctly (no longer raises a WarpCodegenError)
s.m[2][2] = 5.0 # This has also been fixed (used to silently fail)
out[0] = s
Early testing on NVIDIA Jetson Thor indicates that launching CPU kernels may sometimes result in segmentation faults. GPU kernel launches are unaffected. We believe this can be resolved by building Warp from source against LLVM/Clang version 18 or newer.
The following features have been deprecated in prior releases and will be removed in v1.10 (early November):
warp.sim - Use the Newton engine.wp.matrix() from column vectors - Use wp.matrix_from_rows() or wp.matrix_from_cols() instead.wp.select() - Use wp.where() instead (note: different argument order).wp.matrix(pos, quat, scale) - Use wp.transform_compose() instead.We thank the following contributors for their valuable contributions to this release:
warp.jax_experimental.ffi.jax_callable() with a function annotated with the -> None return type (#893).strip_hash=True option for the new ahead-of-time compilation functions (#661).For a curated list of all changes in this release, please see the v1.9.0 section in CHANGELOG.md.
wp.MarchingCubes.extract_surface_marching_cubes() to extract a triangular mesh from a 3D scalar field
(GH-788).wp.compile_aot_module() and wp.load_aot_module() to support basic ahead-of-time compilation workflows
(docs,
GH-766).wp.matrix()/wp.vector()/wp.quaternion() types
(GH-899).array.ptr
(GH-819).wp.array() constructor inside kernels
(GH-853).wp.zeros()
(GH-794).IntEnum and IntFlag inside Warp kernels (GH-529).bounds_check option to wp.tile_load(), wp.tile_store(), and wp.tile_atomic_add()
for performance optimization. When set to False, boundary checks are disabled for memory-aligned tiles,
improving performance (defaults to True for safety) (GH-797).wp.tile_index_load(), wp.tile_index_store(), and wp.tile_index_atomic_add()
(GH-684, GH-796).block_dim argument to wp.load_module() and wp.force_load().examples/core/example_render_opengl.py example with comprehensive ImGui usage examples
(GH-833).DeprecationWarning with this information.wp_ prefix for all exported, C-style symbols to prevent name conflicts
(GH-792).wp.MarchingCubes in pure Warp for cross-platform support and differentiability
(GH-788).wp.array from a pointer inside a kernel.wp.tile_map() to return a different type than their input arguments
(GH-732).wp.map()
(GH-732).warp.sparse to efficiently process sparse matrices with arbitrarily sized blocks and leverage tiled
computations when beneficial (GH-838).warp.fem.integrate for quadrature-point accumulation
(GH-854).wp.breakpoint() in CUDA kernels on Linux systems. This feature is not supported on Windows due to
CUDA-GDB Linux-target-only support (GH-795).__init__.py using typing re-export conventions to improve
static type checker support (GH-864).Failed to compile LTO)
(GH-608, GH-911).libffi
library bug (GH-356).str() and repr() implementations missing for scalar types at the Python scope
(GH-863).wp.array parameters.wp.load_module() when loading modules created with @wp.kernel(module="unique").scikit-image convention
(GH-324).UnboundLocalError when applying warp.jax_experimental.ffi.jax_callable to a function annotated with the
None return type (GH-893).warp.fem discrete fields.warp.fem integrands.warp.fem.#line directives for Python↔CUDA source correlation not being emitted by default when a module is compiled in
debug mode (GH-901).However, to support the adoption of Warp by the MuJoCo MJX physics engine, it also includes new features and deprecations limited to the jax_experimen…
This patch release primarily contains bug fixes as expected.
However, to support the adoption of Warp by the MuJoCo MJX physics engine, it also includes new features and deprecations limited to the jax_experimental module. We are flagging this deviation from our standard versioning practices to ensure clarity. Normal versioning practices will resume with the next release.
graph_compatible boolean flag in jax_callable() in favor of the new graph_mode argument with GraphMode enum (#848).wp.indexedarray() (#468).jax_callable() using Warp via the new graph_mode parameter (GraphMode.WARP), enabling capture of graphs with conditional nodes that cannot be used as subgraphs in a JAX capture (#848).tape.zero() to correctly reset gradient arrays in nested structs (#807).div(scalar, vec), div(scalar, mat), and div(scalar, quat), and other miscellaneous issues with adjoints (#831).wp.config.mode were not being picked up after module initialization (#856).wp.tile_sort() (#836).wp.tile_min() and wp.tile_argmin() to return correct values for large tiles with low occupancy (#725).wp.tile_sum() when using shared tiles (#822).cuDeviceGetUuid caused by using an incorrect version (#851).wp.sparse.bsr_from_triplets() ignored the prune_numerical_zeros=False setting (#832).wp.sim.VBDIntegrator with handle_self_contact=False (#862).OpenGLRenderer where meshes with different scale attributes were incorrectly instanced, causing them all to be rendered with the same scale OpenGLRenderer (#828).Remove wp.mlp() (deprecated in v1.6.0). Use tile primitives instead.
wp.map() to map a function over arrays and add math operators for Warp arrays (docs, #694).wp.capture_if() and wp.capture_while() (docs, #597).wp.capture_debug_dot_print() to write a DOT file describing the structure of a captured CUDA graph (#746).Device.sm_count property to get the number of streaming multiprocessors on a CUDA device (#584).wp.block_dim() to query the number of threads in the current block inside a kernel (#695).wp.atomic_cas() and wp.atomic_exch() built-ins for atomic compare-and-swap and exchange operations (#767).wp.config.compile_time_trace setting or the module-level "compile_time_trace" option. When used, JSON files in the Trace Event format will be written in the kernel cache, which can be opened in a viewer like chrome://tracing/ (docs, #609).wp.svd3() and wp.quat_to_axis_angle() (#503).wp.func functions (#682).wp.tile_squeeze() to remove axes of length one (#662).wp.tile_reshape() to reshape a tile (#663).wp.tile_astype() to return a new tile with the same data but different data type. (#683).wp.tile_cholesky_solve() (#773).wp.tile_scan_inclusive() and wp.tile_scan_exclusive() for performing inclusive and exclusive scans over tiles (#731).wp.transform_compose() and wp.transform_decompose() for converting between transforms and 4x4 matrices with 3D scale information (#576).wp.transform syntax operations for loading and storing (#710).as_spheres parameter to UsdRenderer.render_points() in order to choose whether to render the points as USD spheres using a point instancer or as simple USD points (#634).wp.sim.VBDIntegrator.rebuild_bvh() to rebuild the BVH used for detecting self-contacts.wp.sim.VBDIntegrator collisions, with strength is controlled by Model.soft_contact_kd.wp.fem.lookup() operator across geometries and add filtering parameters (#618).warp.fem: fem/example_elastic_shape_optimization.py and fem/example_darcy_ls_optimization.py (#698).py.typed marker file (per PEP 561) to the package to formally support static type checking by downstream users (#780).wp.mlp() (deprecated in v1.6.0). Use tile primitives instead.wp.autograd.plot_kernel_jacobians() (deprecated in v1.4.0). Use wp.autograd.jacobian_plot() instead.length and owner keyword arguments from wp.array() constructor (deprecated in v1.6.0). Use the shape and deleter keywords instead.kernel keyword argument from wp.autograd.jacobian() and wp.autograd.jacobian_fd() (deprecated in v1.6.0). Use the function keyword argument instead.outputs keyword argument from wp.autograd.jacobian_plot() (deprecated in v1.6.0).warp.sim module (planned for removal in v1.10). It will be superseded by the upcoming Newton library, a separate package with a new API. Migrating will require code changes; a future guide will be provided (current draft). See the GitHub announcement for details (#735).wp.matrix(pos, quat, scale) built-in function. Use wp.transform_compose() instead (#576).len() where possible.wp.types.type_length() to wp.types.type_size().wp.tile_cholesky_solve() input parameters to align with its docstring (#726).wp.tile_upper_solve() and wp.tile_lower_solve() to use libmathdx 0.2.1 TRSM solver (#773).wp.tile_matmul() if enable_backward is disabled (#644).preserve_type=True when tiling a value across the block with wp.Tile() (#772).wp.sparse.bsr_[set_]from_triplets differentiable with respect to the input triplet values (#760).warp.fem operators: node_count, node_index, element_coordinates, element_closest_point.wp.sim.VBDIntegrator rigid-body-contact handling to use only the shape's friction coefficient, rather than averaging the shape's and the cloth's coefficients.wp.assign_copy() hidden built-in to the kernel scope.inputs and outputs arguments in the Kernel documentation.wp.launch() by avoiding costly native API calls (#774).@wp.func-decorated functions from the Python scope (#521).Formal parameter space overflowed error during wp.sim.VBDIntegrator kernel compilation for the backward pass in CUDA 11 Warp builds. This was resolved by decoupling collision and elasticity evaluations into separate kernels, increasing parallelism and speeding up the solver (#442).UsdRenderer.render_points() not supporting multiple colors (#634).wp.fem module regarding the orientation of 2D geometry side normals (#629).Add missing adjoint method for tile assign operations (GH-680).
assign operations (GH-680).+= and -= invoke wp.atomic_add() and wp.atomic_sub(), respectively
(GH-505).wp.struct (now throws RuntimeError)
(GH-656).__cuda_array_interface__ (GH-624,
GH-670).omni.warp towards omni.warp.core
(GH-702).wp.Volume allocation
(GH-611).wp.tile_atomic_add() (GH-681).wp.svd2() with duplicate singular values and improved accuracy
(GH-679).OpenGLRenderer.update_shape_instance() not having color buffers created for the shape instances.wp.render.OpenGLRenderer (GH-704).ModelBuilder.collapse_fixed_joints()
(GH-631).UsdRenderer.render_points() erroring out when passed 4 points or less
(GH-708).wp.atomic_*() built-ins not working with some types (GH-733).Improve handling of deprecated JAX features (#613).
mpi4py in warp/examples/distributed/example_jacobi_mpi.py (#475).repr() for Warp types, including adding repr() for wp.array.framesPerSecond for time sampling instead of timeCodesPerSecond to avoid playback speed issues in some viewers (#617).Model.rigid_contact_tids are now -1 at non-active contact indices which allows to retrieve the vertex index of a mesh collision, see test_collision.py (#623).DeformedGeometry from wp.fem.Trimesh3D geometries (#614).lookup operator for wp.fem.Trimesh3D (#618).dtype parameter missing for wp.quaternion().dtype comparison when using the wp.matrix()/wp.vector()/wp.quaternion() constructors with literal values and an explicit dtype argument (#651).wp.sim.collide() (#459).wp.sim.ModelBuilder adds springs with -1 as vertex indices (#621).show_joints not working with wp.sim.render.SimRenderer set to render to USD (#510).OgnParticlesFromMesh node not being computed correctly.atol and rtol arguments to wp.autograd.gradcheck() and wp.autograd.gradcheck_tape() (#508).Deprecate constructing a matrix from vectors using wp.matrix().
#line directives in CUDA-C code. This setting is controlled by wp.config.line_directives and is True by default. (docs, #437)vec4f grid construction in wp.Volume.allocate_by_tiles().wp.svd2() (#436).wp.randu() for random uint32 generation.wp.matrix_from_cols() and wp.matrix_from_rows() (#278).wp.transform_from_matrix() to obtain a transform from a 4x4 matrix (#211).wp.where() to select between two arguments conditionally using a more intuitive argument order (cond, value_if_true, value_if_false) (#469).wp.get_mempool_used_mem_current() and wp.get_mempool_used_mem_high() to query the respective current and high-water mark memory pool allocator usage (#446 ).Stream.is_complete and Event.is_complete properties to query completion status (#435).wp.clear_lto_cache() to clear the LTO cache (#507).warp/examples/optim/example_fluid_checkpoint.py.wp.sim.VBDIntegrator.wp.matmul() functionality (including batched version). Users should use tile primitives for matrix multiplication operations instead.wp.matrix().wp.select() in favor of wp.where(). Users should update their code to use wp.where(cond, value_if_true, value_if_false) instead of wp.select(cond, value_if_false, value_if_true).wp.sim.Control no longer has a model attribute (#487).wp.sim.Control.reset() is deprecated and now only zeros-out the controls (previously restored controls to initial model state). Use wp.sim.Control.clear() instead.v[0] = x) now compile and run faster in the backward pass. Note: For correct gradient computation, each component should only be assigned once.@wp.kernel has now an optional module argument that allows passing a wp.context.Module to the kernel, or, if set to "unique" let Warp create a new unique module just for this kernel. The default behavior to use the current module is unchanged.wp.tile_reduce() on tiles with struct data types.wp.tile_broadcast() to support broadcasting to 1D, 3D, and 4D shapes (in addition to existing 2D support).wp.fem.integrate() and wp.fem.interpolate() may now perform parallel evaluation of quadrature points within elements.wp.fem.interpolate() can now build Jacobian sparse matrices of interpolated functions with respect to a trial field.wp.sparse routines (bsr_set_from_triplets, bsr_assign, bsr_axpy, bsr_mm) now accept a masked flag to discard any non-zero not already present in the destination matrix.wp.sparse.bsr_assign() no longer requires source and destination block shapes to evenly divide each other.wp.expect_near() to support all vectors and quaternions.wp.quat_from_matrix() to support 4x4 matrices.OgnClothSimulate node to use the VBD integrator (#512).globalScale parameter from the OgnClothSimulate node.edge_indices when adding a ModelBuilder to another (#557).Update project license from *NVIDIA Software License* to *Apache License, Version 2.0* (see LICENSE.md).
Document wp.Launch objects (docs, #428).
wp.Launch objects (docs, #428).wp.tile_load().wp.array() not initializing from arrays defining a CUDA array interface when the target device is CPU (#523).wp.Launch objects not storing and replaying adjoint kernel launches (#449).wp.config.verify_autograd_array_access failing to detect overwrites in generic Warp functions (#493).OpenGLRenderer app (#488).wp.sim.VBDIntegrator with CUDA graphs when handle_self_contact is enabled (#441).wp.collide.TriMeshCollisionDetector.target_ke, target_kd, and mode parameters (#454).ModelBuilder.add_builder() to use correct offsets for ModelBuilder.joint_parent and ModelBuilder.joint_child (#432)wp.randi() documentation to show correct output range of [-2^31, 2^31).Emit deprecation warnings for the use of the owner and length keywords in the wp.array initializer.
wp.tile_cholesky(), tile_cholesky_solve()
and tile_diag_add() (preview APIs are subject to change).wp.sim.VDBIntegrator by passing handle_self_contact=True.
See warp/examples/sim/example_cloth_self_contact.py for a usage example.wp.norm_l1(), wp.norm_l2(), wp.norm_huber(), wp.norm_pseudo_huber(), and wp.smooth_normalize()
for vector types to a new wp.math module.wp.sim.SemiImplicitIntegrator and wp.sim.FeatherstoneIntegrator now have an optional friction_smoothing
constructor argument (defaults to 1.0) that controls softness of the friction norm computation.assert statements in kernels (docs).
Assertions can only be triggered in "debug" mode (GH-366).ipc_handle() method to get an IPC handle for a wp.Event or a wp.array,
and call wp.from_ipc_handle() or wp.event_from_ipc_handle() in another process to open the handle
(docs).wp.set_module_options({"fuse_fp": False})
(GH-379).wp.set_module_options({"lineinfo": True}).wp.struct objects by defining wp.func functions
(GH-392).wp.len() to retrieve the number of elements for vectors, quaternions, matrices, and arrays
(GH-389).warp/examples/optim/example_softbody_properties.py as an optimization example for soft-body properties
(GH-419).warp/examples/tile/example_tile_walker.py, which reworks the existing example_walker.py
to use Warp's tile API for matrix multiplication.warp/examples/tile/example_tile_nbody.py as an example of an N-body simulation using Warp tile primitives.wp.tile_load() and wp.tile_store() indexing behavior so that indices are now specified in
terms of array elements instead of tile multiples.shape and offset parameters as tuples,
e.g.: wp.tile_load(array, shape=(m,n), offset=(i,j)).wp.Bvh constructor now supports various construction algorithms via the constructor argument, including
"sah" (Surface Area Heuristics), "median", and "lbvh" (docs)wp.Bvh and wp.Mesh.enable_backward set to False (GH-332).+= and -= operations compile and run faster in the backward pass
(GH-332).module_codegen (GH-431).block_dim.wp.autograd.gradcheck_tape() now has additional optional arguments reverse_launches and skip_to_launch_index.wp.autograd.gradcheck(), wp.autograd.jacobian(), and wp.autograd.jacobian_fd() now also accept
arbitrary Python functions that have Warp arrays as inputs and outputs.update_vbo_transforms kernel launches in the OpenGL renderer are no longer recorded onto the tape.enable_backward is set to False.owner and length keywords in the wp.array initializer.wp.mlp(), wp.matmul(), and wp.batched_matmul().
Use tile primitives instead.wp.Tape.zero() zeroes gradients passed via the grads parameter in wp.Tape.backward()
(GH-407).wp.array() not respecting the target dtype and shape when the given data is an another array with a CUDA interface
(GH-363).ImportError exception being thrown during interpreter shutdown on Windows when using the OpenGL renderer
(GH-412).AttributeError crash in the OpenGL renderer when moving the camera (GH-426).wp.sim.ModelBuilder default parameters (GH-429).wp.tile_extract() when the block dimension is smaller than the tile size.wp.autograd.jacobian() where in some cases gradients were not zeroed-out properly.wp.autograd.jacobian_plot().len() operator returning the total size of a matrix instead of its first dimension.wp.sim.SemiImplicitIntegrator and
wp.sim.FeatherstoneIntegrator (GH-349).up_axis, color in OpenGLRenderer (GH-448).Add PyTorch basics and custom operators notebooks to the notebooks directory.
notebooks directory.wp.launch_tiled() not returning a Launch object when passed record_cmd=True.wp.func when called from Python's runtime
(GH-386).wp.atomic_add(), wp.atomic_sub(),
wp.atomic_max(), or wp.atomic_min() as being written to (GH-378)..meta files into Warp kernel cache on Windows.Support for cooperative tile-based primitives using cuBLASDx and cuFFTDx, please see the tile documentation for details.
reversed() built-in for iterators (GH-311)..nvdb files with the save_to_nvdb method.wp.fem.Trimesh3D and wp.fem.Quadmesh3D geometry types for 3D surfaces with new example_distortion_energy example."add" option to wp.fem.integrate() for accumulating integration result to existing output."assembly" option to wp.fem.integrate() for selecting between more memory-efficient or more
computationally efficient integration algorithms.curl and div operators, respectively.wp.sim.ModelBuilder now includes methods to color particles for use with wp.sim.VBDIntegrator(),
users should call builder.color() before finalizing assets.wp.sim.Model.particle_radius
array (docs), replacing the previous
hard-coded value of 0.01 (GH-329).particle_radius parameter to wp.sim.ModelBuilder.add_cloth_mesh() and wp.sim.ModelBuilder.add_cloth_grid()
to set a uniform radius for the added particles.wp.array attributes (GH-364).notebooks directory.wp.Int, wp.Float, and wp.Scalar generic annotation types to the public API.wp.fem.cells(), wp.fem.to_inner_cell(), wp.fem.to_outer_cell() operators.wp.randn() samples a normal distribution of mean 0 and variance 1.wp.printf() built-in.place setting of paddle backend.wp.expect_neq() overloads missing for scalar types.wp.kernel or a wp.func object is annotated to return a None value..nvdb files.wp.printf() erroring out when no variadic arguments are passed (GH-333).Make the output of wp.print() in backward kernels consistent for all supported data types.
wp.print() in backward kernels consistent for all supported data types.1.3.0).dtype values.Texture Write node, used in the Mandelbrot Omniverse sample, sometimes erroring out in multi-GPU environments.Re-introduced the wp.rand*(), wp.sample*(), and wp.poisson() onto the Python scope to revert a breaking change.
iter_reverse() not working as expected for ranges with steps other than 1 (GH-311).wp.sparse.BsrMatrix object is reused for storing matrices of different shapes.wp.fem.utils.symmetric_eigenvalues_qr.ModelBuilder.add_builder(builder) to correctly update articulation_start and thereby articulation_count when builder contains more than one articulation.wp.rand*(), wp.sample*(), and wp.poisson() onto the Python scope to revert a breaking change.Support for a new wp.static(expr) function that allows arbitrary Python expressions to be evaluated at the time of function/kernel definition (docs).
wp.static(expr) function that allows arbitrary Python expressions to be evaluated at the time of
function/kernel definition (docs).warp.fem (docs).wp.kernel and wp.func objects from within closures.wp.func.jax_kernel() (GH-310).wp.mod() for vector types (GH-282).% to Python's runtime scalar and vector types.atomic_add, atomic_max, and atomic_min (GH-284).q.w).omni.warp extension.warp.sim.VBDIntegrator now supports body-particle collision.wp.sim.Model.edge_indices now includes boundary edges.wp.rand*(), wp.sample*(), and wp.poisson() from the Python scope.wp.Mesh.points is now a property instead of a raw data member, its reference can be changed after the mesh is initialized.if/else/elif statements with constant conditions are resolved at compile time with no branches being inserted in the generated code.warp.fem.wp.func erroring out when defining a Tuple as a return type hint (GH-302).+=, -=) adjoints to compute gradients correctly in the backwards passv[1] = x--num_tiles 1 in example_render_opengl.py (GH-306).FeatherstoneIntegrator when bodies and particles collide.FeatherstoneIntegrator where eval_rigid_jacobian could give incorrect results or reach an infinite
loop when the body and joint indices were not in the same order. Added Model.joint_ancestor to fix the indexing
from a joint to its parent joint in the articulation.add_edges() called from ModelBuilder.add_cloth_mesh() (GH-319).compute-sanitizer initcheck tool when using wp.Mesh.stream argument in array.__dlpack__().while statements.wp.Tape context.kDLBool instead of kDLUInt for DLPack interop of Booleans.Fix an aliasing issue with zero-copy array initialization from NumPy introduced in Warp 1.3.0.
wp.Volume.load_from_numpy() behavior when bg_value is a sequence of values.wp.svd3 with fp64 numbers (GH-281).wp.bvh_query_ray() where the direction instead of the reciprocal direction was used
(GH-288).wp.sim.collide.triangle_closest_point_barycentric() where the returned barycentric coordinates may be
incorrect when the closest point lies on an edge.np.int32.input_output_mask argument to autograd.jacobian and
autograd.jacobian_fd (GH-289).ModelBuilder.collapse_fixed_joints() to correctly update the body centers of mass and the
ModelBuilder.articulation_start array.wp.fem.ExplicitQuadrature (regression from 1.3.0).wp.bvh_query_aabb() returns parts that overlap the bounding volume.wp.synchronize() from PyTorch autograd function exampleTape.check_kernel_array_access() and Tape.reset_array_read_flags() are now private methods.Warp Core improvements
(compiled), loaded from the cache (cached), or was unable to be
loaded (error).wp.config.verbose = True now also prints out a message upon the entry to a wp.ScopedTimer.wp.clear_kernel_cache() to the public API. This is equivalent to wp.build.clear_kernel_cache().wp.config variables.wp.matmul() CPU fallback to use dtype explicitly in np.matmul() callfrom __future__ import annotations (GH-256).wp.launch() directly via __cuda_array_interface__ and __array_interface__, up to 2.5x faster conversion from PyTorchreturn_ctype argument to wp.from_torch()wp.abs() and wp.sign() for vector typeswp.float16(1.23) * wp.float16(2.34))wp.copy(), wp.clone(), and array.assign() differentiability__new__() methods for all class __del__() methods to handle when a class instance is created but not instantiated before garbage collectionwp.quat_t suffix: wp.BVHQuery, wp.HashGridQuery, wp.MeshQueryAABB, wp.MeshQueryPoint, and wp.MeshQueryRaywp.array(ptr=...) to allow initializing arrays from pointer addresses inside of kernels (GH-206)warp.autograd improvements:
warp.autograd module with utility functions gradcheck(), jacobian(), and jacobian_fd() for debugging kernel Jacobians (docs)wp.config.verify_autograd_array_access is true in-place operations on arrays on the Tape that could break gradient computation will be detected (docs)@wp.func_replay functions and native snippets would not trigger module recompilationwarp.sim improvements:
self.rigid_mesh_contact_max is zero (default behavior).mask argument to wp.sim.eval_fk() now accepts both integer and boolean arrays to mask articulations.ModelBuilder.joint_act in ModelBuilder.collapse_fixed_joints() (affected floating-base systems)ModelBuilder.plot_articulation() to visualize the articulation tree of a rigid-body mechanism__new__() method (missing instance return and *args parameter)upaxis variable in ModelBuilder and the rendering thereof in OpenGLRendererwarp.sparse improvements:
bsr_from_triplets(), bsr_axpy(), etc.) can now be captured in CUDA graphs; exact number of non-zeros can be optionally requested asynchronously.bsr_assign() now supports changing block shape (including CSR/BSR conversions)A += 0.5 * B, y = x @ Cwarp.fem new features and fixes:
wp.fem.lookup() operator now supports wp.fem.Tetmesh and wp.fem.Trimesh2D geometrieswp.fem.Subdomain), free-slip boundary conditionswp.fem.UniformField, wp.fem.ImplicitField and wp.fem.NonconformingFieldstreamlines, magnetostatics and nonconforming_contact examples, updated mixed_elasticity to use a nonlinear modelwp.fem.PicQuadrature w.r.t. positions and measuresFix accuracy of 3x3 SVD wp.svd3 with fp64 numbers (GH-281).
wp.svd3 with fp64 numbers (GH-281).wp.bvh_query_ray() where the direction instead of the reciprocal direction was used
(GH-288).wp.sim.collide.triangle_closest_point_barycentric() where the returned barycentric coordinates may be
incorrect when the closest point lies on an edge.np.int32.input_output_mask argument to autograd.jacobian and
autograd.jacobian_fd (GH-289).ModelBuilder.collapse_fixed_joints() to correctly update the body centers of mass and the
ModelBuilder.articulation_start array.wp.fem.ExplicitQuadrature (regression from 1.3.0).wp.bvh_query_aabb() returns parts that overlap the bounding volume.wp.synchronize() from PyTorch autograd function exampleTape.check_kernel_array_access() and Tape.reset_array_read_flags() are now private methods.(compiled), loaded from the cache (cached), or was unable to be
loaded (error).wp.config.verbose = True now also prints out a message upon the entry to a wp.ScopedTimer.wp.clear_kernel_cache() to the public API. This is equivalent to wp.build.clear_kernel_cache().wp.config variables.wp.matmul() CPU fallback to use dtype explicitly in np.matmul() callfrom __future__ import annotations (GH-256).wp.launch() directly via __cuda_array_interface__ and __array_interface__, up to 2.5x faster conversion from PyTorchreturn_ctype argument to wp.from_torch()wp.abs() and wp.sign() for vector typeswp.float16(1.23) * wp.float16(2.34))wp.copy(), wp.clone(), and array.assign() differentiability__new__() methods for all class __del__() methods to handle when a class instance is created but not instantiated before garbage collectionwp.quat_t suffix: wp.BVHQuery, wp.HashGridQuery, wp.MeshQueryAABB, wp.MeshQueryPoint, and wp.MeshQueryRaywp.array(ptr=...) to allow initializing arrays from pointer addresses inside of kernels (GH-206)Remove wp.synchronize() from PyTorch autograd function example
wp.synchronize() from PyTorch autograd function exampleTape.check_kernel_array_access() and Tape.reset_array_read_flags() are now private methods.Warp Core improvements
(compiled), loaded from the cache (cached), or was unable to be
loaded (error).wp.config.verbose = True now also prints out a message upon the entry to a wp.ScopedTimer.wp.clear_kernel_cache() to the public API. This is equivalent to wp.build.clear_kernel_cache().wp.config variables.wp.matmul() CPU fallback to use dtype explicitly in np.matmul() callfrom __future__ import annotations (GH-256).wp.launch() directly via __cuda_array_interface__ and __array_interface__, up to 2.5x faster conversion from PyTorchreturn_ctype argument to wp.from_torch()wp.abs() and wp.sign() for vector typeswp.float16(1.23) * wp.float16(2.34))wp.copy(), wp.clone(), and array.assign() differentiability__new__() methods for all class __del__() methods to handle when a class instance is created but not instantiated before garbage collectionwp.quat_t suffix: wp.BVHQuery, wp.HashGridQuery, wp.MeshQueryAABB, wp.MeshQueryPoint, and wp.MeshQueryRaywp.array(ptr=...) to allow initializing arrays from pointer addresses inside of kernels (GH-206)warp.autograd improvements:
warp.autograd module with utility functions gradcheck(), jacobian(), and jacobian_fd() for debugging kernel Jacobians (docs)wp.config.verify_autograd_array_access is true in-place operations on arrays on the Tape that could break gradient computation will be detected (docs)@wp.func_replay functions and native snippets would not trigger module recompilationwarp.sim improvements:
self.rigid_mesh_contact_max is zero (default behavior).mask argument to wp.sim.eval_fk() now accepts both integer and boolean arrays to mask articulations.ModelBuilder.joint_act in ModelBuilder.collapse_fixed_joints() (affected floating-base systems)ModelBuilder.plot_articulation() to visualize the articulation tree of a rigid-body mechanism__new__() method (missing instance return and *args parameter)upaxis variable in ModelBuilder and the rendering thereof in OpenGLRendererwarp.sparse improvements:
bsr_from_triplets(), bsr_axpy(), etc.) can now be captured in CUDA graphs; exact number of non-zeros can be optionally requested asynchronously.bsr_assign() now supports changing block shape (including CSR/BSR conversions)A += 0.5 * B, y = x @ Cwarp.fem new features and fixes:
wp.fem.lookup() operator now supports wp.fem.Tetmesh and wp.fem.Trimesh2D geometrieswp.fem.Subdomain), free-slip boundary conditionswp.fem.UniformField, wp.fem.ImplicitField and wp.fem.NonconformingFieldstreamlines, magnetostatics and nonconforming_contact examples, updated mixed_elasticity to use a nonlinear modelwp.fem.PicQuadrature w.r.t. positions and measuresUpdate to CUDA 12.x by default (requires NVIDIA driver 525 or newer), please see README.md for commands to install CUDA 11.x binaries for older driver
Warp Core improvements
(compiled), loaded from the cache (cached), or was unable to be
loaded (error).wp.config.verbose = True now also prints out a message upon the entry to a wp.ScopedTimer.wp.clear_kernel_cache() to the public API. This is equivalent to wp.build.clear_kernel_cache().wp.config variables.wp.matmul() CPU fallback to use dtype explicitly in np.matmul() callfrom __future__ import annotations (GH-256).wp.launch() directly via __cuda_array_interface__ and __array_interface__, up to 2.5x faster conversion from PyTorchreturn_ctype argument to wp.from_torch()wp.abs() and wp.sign() for vector typeswp.float16(1.23) * wp.float16(2.34))wp.copy(), wp.clone(), and array.assign() differentiability__new__() methods for all class __del__() methods to handle when a class instance is created but not instantiated before garbage collectionwp.quat_t suffix: wp.BVHQuery, wp.HashGridQuery, wp.MeshQueryAABB, wp.MeshQueryPoint, and wp.MeshQueryRaywp.array(ptr=...) to allow initializing arrays from pointer addresses inside of kernels (GH-206)warp.autograd improvements:
warp.autograd module with utility functions gradcheck(), jacobian(), and jacobian_fd() for debugging kernel Jacobians (docs)wp.config.verify_autograd_array_access is true in-place operations on arrays on the Tape that could break gradient computation will be detected (docs)@wp.func_replay functions and native snippets would not trigger module recompilationwarp.sim improvements:
self.rigid_mesh_contact_max is zero (default behavior).mask argument to wp.sim.eval_fk() now accepts both integer and boolean arrays to mask articulations.ModelBuilder.joint_act in ModelBuilder.collapse_fixed_joints() (affected floating-base systems)ModelBuilder.plot_articulation() to visualize the articulation tree of a rigid-body mechanism__new__() method (missing instance return and *args parameter)upaxis variable in ModelBuilder and the rendering thereof in OpenGLRendererwarp.sparse improvements:
bsr_from_triplets(), bsr_axpy(), etc.) can now be captured in CUDA graphs; exact number of non-zeros can be optionally requested asynchronously.bsr_assign() now supports changing block shape (including CSR/BSR conversions)A += 0.5 * B, y = x @ Cwarp.fem new features and fixes:
wp.fem.lookup() operator now supports wp.fem.Tetmesh and wp.fem.Trimesh2D geometrieswp.fem.Subdomain), free-slip boundary conditionswp.fem.UniformField, wp.fem.ImplicitField and wp.fem.NonconformingFieldstreamlines, magnetostatics and nonconforming_contact examples, updated mixed_elasticity to use a nonlinear modelwp.fem.PicQuadrature w.r.t. positions and measuresFix Warp not being initialized when constructing arrays with wp.array()
wp.array()wp.is_mempool_access_supported() not resolving the provided device arguments to wp.context.Devicewp.NAN or wp.nan.wp.isnan(), wp.isinf(), and wp.isfinite() for scalars, vectors, matrices, etc.wp.constant() variables declared in a Warp program.wp.MarchingCubes on field dimensions and sizewp.Mesh BVH (GH-225)sm_75 (from sm_70), enabling Turing ISA featureswarp.Volume):
wp.volume_lookup_index(), wp.volume_sample_index() and generic wp.volume_sample()/wp.volume_lookup()/wp.volume_store() kernel-level functionswarp.fem can now work directly on NanoVDB grids using warp.fem.Nanogridwp.volume_sample_v() and wp.volume_store_*() adjointswp.volume_store() from overwriting grid background valueswarp.femwp.render.OpenGLRenderer via pyglet.options["headless"] = Truewp.render.RegisteredGLBuffer can fall back to CPU-bound copying if CUDA/OpenGL interop is not availablewp.sparse.bsr_mm() by ~5x on benchmark problemsjoint_act arrayswp.sim.FeatherstoneIntegrator()--msvc_path in build scriptswp.copy() params to record dest and src offset parameters on wp.Tape()wp.randn() to ensure return values are finitebool types in generic kernelswp.copy(), wp.clone(), and array.assign() differentiability__new__() methods for all class __del__() methods to
handle when a class instance is created but not instantiated before garbage collection.wp.quatFix Warp not being initialized when constructing arrays with wp.array()
wp.array()wp.is_mempool_access_supported() not resolving the provided device arguments to wp.context.Devicewp.NAN or wp.nan.wp.isnan(), wp.isinf(), and wp.isfinite() for scalars, vectors, matrices, etc.wp.constant() variables declared in a Warp program.wp.MarchingCubes on field dimensions and sizewp.Mesh BVH (GH-225)sm_75 (from sm_70), enabling Turing ISA featureswarp.Volume):
wp.volume_lookup_index(), wp.volume_sample_index() and generic wp.volume_sample()/wp.volume_lookup()/wp.volume_store() kernel-level functionswarp.fem can now work directly on NanoVDB grids using warp.fem.Nanogridwp.volume_sample_v() and wp.volume_store_*() adjointswp.volume_store() from overwriting grid background valueswarp.femwp.render.OpenGLRenderer via pyglet.options["headless"] = Truewp.render.RegisteredGLBuffer can fall back to CPU-bound copying if CUDA/OpenGL interop is not availablewp.sparse.bsr_mm() by ~5x on benchmark problemsjoint_act arrayswp.sim.FeatherstoneIntegrator()--msvc_path in build scriptswp.copy() params to record dest and src offset parameters on wp.Tape()wp.randn() to ensure return values are finitebool types in generic kernelsAdd a not-a-number floating-point constant that can be used as wp.NAN or wp.nan.
wp.NAN or wp.nan.wp.isnan(), wp.isinf(), and wp.isfinite() for scalars, vectors, matrices, etc.wp.constant() variables declared in a Warp program.wp.MarchingCubes on field dimensions and sizewp.Mesh BVH (GH-225)sm_75 (from sm_70), enabling Turing ISA featureswarp.Volume):
wp.volume_lookup_index(), wp.volume_sample_index() and generic wp.volume_sample()/wp.volume_lookup()/wp.volume_store() kernel-level functionswarp.fem can now work directly on NanoVDB grids using warp.fem.Nanogridwp.volume_sample_v() and wp.volume_store_*() adjointswp.volume_store() from overwriting grid background valueswarp.femwp.render.OpenGLRenderer via pyglet.options["headless"] = Truewp.render.RegisteredGLBuffer can fall back to CPU-bound copying if CUDA/OpenGL interop is not availablewp.sparse.bsr_mm() by ~5x on benchmark problemsjoint_act arrayswp.sim.FeatherstoneIntegrator()--msvc_path in build scriptswp.copy() params to record dest and src offset parameters on wp.Tape()wp.randn() to ensure return values are finitebool types in generic kernelswp.init() is no longer required to be called explicitly and will be performed on first call to the APIomni.warp.core's startup timeSupport returning a value from @wp.func_native CUDA functions using type hints
@wp.func_native CUDA functions using type hintswp.sim.FeatherstoneIntegratorwp.sim.collide()wp.ScopedTimer()wp.Tape.visualize()__cuda_array_interface__ attributestruct.to(device) to migrate struct arrayswp.sim.Modelwp.HashGrid.build()Ruff for formatting and lintingwp.launch()wp.optim.linear.cr()import warp by eliminating raising any exceptionsa[i][j] vs. a[i, j]support_level entry to the configuration file of the extensionsMake examples runnable from any location
README.md examples USD locationexample_graph_capture.py descriptionDocument Device total_memory and free_memory
total_memory and free_memorypython -m warp.examples.browse for browsing the examples folderexamples/optim/example_walker.py samplewp.synchronize_event() for blocking the host thread until a recorded event completesstdout captureAdd FeatherstoneIntegrator which provides more stable simulation of articulated rigid body dynamics in generalized coordinates (State.joint_q and Stat
FeatherstoneIntegrator which provides more stable simulation of articulated rigid body dynamics in generalized coordinates (State.joint_q and State.joint_qd)warp.sim.Control struct to store control inputs for simulations (optional, by default the Model control inputs are used as before); integrators now have a different simulation signature: integrator.simulate(model: Model, state_in: State, state_out: State, dt: float, control: Control)joint_act can now behave in 3 modes: with joint_axis_mode set to JOINT_MODE_FORCE it behaves as a force/torque, with JOINT_MODE_VELOCITY it behaves as a velocity target, and with JOINT_MODE_POSITION it behaves as a position target; joint_target has been removedModel.shape_materials.ka which controls the contact distance at which the adhesive force is appliedwp.ScopedCaptureenable_backward warning for callablesAdd examples assets to the wheel packages
wp.config.quiet = TrueDeprecation of owner argument - use deleter to transfer ownership
examples directory under warp/python -m warp.tests --helptorch.autograd.function example + docsexample_graph_captureverify_fp causing compiler errors and support CPU kernelsmatmul to be called in CUDA graph capturewp.launch to support tuple argstest_femassert_np_equal when NaN's and tolerance are involvedwarp.config.max_unroll, fix custom gradient unrolling@wp.func_native(snippet, replay_snippet=replay_snippet)CUDA_PATH environment variable or --cuda_path build option are not usedwp.ones() to efficiently create one-initialized arrayswp.config.graph_capture_module_load_default to wp.config.enable_graph_capture_module_load_by_defaultwp.config.enable_mempools_at_init to enable pooled allocators during Warp initialization (if supported)wp.is_mempool_supported() - check if a device supports pooled allocatorswp.is_mempool_enabled(), wp.set_mempool_enabled() - enable or disable pooled allocators per devicewp.set_mempool_release_threshold(), wp.get_mempool_release_threshold() - configure memory pool release thresholdwp.is_peer_access_supported() - check if the memory of a device can be accessed by a peer devicewp.is_peer_access_enabled(), wp.set_peer_access_enabled() - manage peer access for memory allocated using default CUDA allocatorswp.is_mempool_access_supported() - check if the memory pool of a device can be accessed by a peer devicewp.is_mempool_access_enabled(), wp.set_mempool_access_enabled() - manage access for memory allocated using pooled CUDA allocatorswp.ScopedStream can synchronize with the previous stream on entry and/or exit (only sync on entry by default)wp.copy(), wp.launch(), wp.capture_launch())deleter argument when constructing arrays
owner argument - use deleter to transfer ownershipwp.zeros(), wp.full(), and more)wp.matmul() to always use the correct CUDA contextNoise Deform are deterministic across different Kit sessionsUpdate the license to *NVIDIA Software License*, allowing commercial use (see LICENSE.md)
LICENSE.md)CONTRIBUTING.md guidelines (for NVIDIA employees)snippet and adj_snippet strings to fix cachingbuild_docs.py on Windows.py extension to warp/tests/walkthrough_debugwp.bool usage in vector and matrix typesenable_backward setting is set to False upon calling wp.Tape.backward()wp.from_torch()Noise Deform node for OmniGraph that deforms points using a perlin/curl noiseShow deprecation warnings only once
pip install warp-lang.PACKAGING.md and resembling that of Python itself:
release-0.11 branch.public branch, previously used to merge releases into and corresponding with the GitHub main branch, is retired.try/finallywp.optim.linear.cg, wp.optim.linear.bicgstab, wp.optim.linear.gmres, and wp.optim.linear.LinearOperatorfloat(wp.float32(1.23)) or ctypes.c_float(wp.float32(1.23))wp.infexamples/example_diffsim_mass_spring_cage.py-s, --suite option for only running tests belonging to the given suiteswp.render.OpenGLRendererwp.mesh_query_ray()Remove deprecated wp.ScopedCudaGuard(), please use wp.ScopedDevice() instead
warp.array.reshape() to handle -1 dimensionswp.sim.create_soft_body_contacts()warp.from_torch(), warp.to_torch() plus documentationwp.mesh_query_point_nosign() - closest point query with no sign determinationwp.mesh_query_point_sign_normal() - closest point query with sign from angle-weighted normalwp.mesh_query_point_sign_winding_number() - closest point query with fast winding number sign determinationwarp.sparse module:
wp.sparse.BsrMatrixwp.sparse.bsr_zeros(), wp.sparse.bsr_set_from_triplets() for constructionwp.sparse.bsr_mm(), wp.sparse_bsr_mv() for matrix-matrix and matrix-vector products respectivelywp.utils.array_scan() - prefix sum (inclusive or exlusive)wp.utils.array_sum() - sum across arraywp.utils.radix_sort_pairs() - in-place radix sort (key,value) pairs@wp.func functions from Python (outside of kernel scope)wp.Launch object that can be replayed with low overhead, use wp.launch(..., record_cmd=True) to generate a command objectwp.struct kernel arguments, up to 20x faster launches for kernels with large structs or number of paramsomni.warp.nodes module
omni.warp.nodes.mesh_create_bundle(), omni.warp.nodes.mesh_get_points(), etc.wp.array:
wp.empty() no longer zeroes-out memory and returns an uninitialized array, as intendedarray.zero_() and array.fill_() work with non-contiguous arraysarray.fill_() can now take lists or other sequences when filling arrays of vectors or matrices, e.g. arr.fill_([[1, 2], [3, 4]])array.fill_() now works with arrays of structs (pass a struct instance)wp.copy() gracefully handles copying between non-contiguous arrays on different deviceswp.full() and wp.full_like(), e.g., a = wp.full(shape, value)device argument to wp.empty_like(), wp.zeros_like(), wp.full_like(), and wp.clone()indexedarray methods .zero_(), .fill_(), and .assign()indexedarray methods .numpy() and .list()array.list() to work with arrays of any Warp data typearray.list() synchronization issue with CUDA arraysarray.numpy() called on an array of structs returns a structured NumPy array with named fieldsError: No module named 'omni.warp.core' when running some Kit configurations (e.g.: stubgen)wp.struct instance address being included in module content hashwp.BVH.refit() when executed on the CPUwp.struct constructorwp.float16 vectors and matrices in Pythonwp.float16 members in structswp.ScopedCudaGuard(), please use wp.ScopedDevice() insteadwp.simwp.array.reshape() to handle -1 dimensionswp.sim.create_soft_body_contacts()wp.from_torch(), wp.to_torch() plus documentationDeprecate wp.Model.soft_contact_distance which is now replaced by wp.Model.particle_radius
wp.verbose if using gradients)wp.breakpoint(), and test_debug.py@wp.func functionspass, continue, and break statements__sincos_stret symbol for macOSwp.Mesh.points, and other cases where arrays are passed to native functions@ operator as an alias for wp.matmul()ModelBuilder.add_particle has a new radius argument, Model.particle_radius is now a Warp arrayModel.particle_flags Warp array, introduce PARTICLE_FLAG_ACTIVE to define whether a particle is being simulated and participates in contact dynamics&, |, ~, <<, >>cpu devicesomni.warp into omni.warp.core for Omniverse applications that want to use the Warp Python module with minimal additional dependencies@wp.struct registration during hot reloadunot() operator so kernel writers can use if not array: syntaxwp.hash_grid_point_id() now returns -1 if the wp.HashGrid has not been reserved beforewp.Model.soft_contact_distance which is now replaced by wp.Model.particle_radiusAdd Texture Write node for updating dynamic RTX textures from Warp kernels / nodes
Texture Write node for updating dynamic RTX textures from Warp kernels / nodeswp.load_module() to pre-load specific modules (pass recursive=True to load recursively)wp.poisson() for sampling Poisson distributionswp.sim.parse_usd()--standalone build optionwp.ScopedTimer()matrix.get_row(), matrix.set_row()wp.struct types within kernelsslice = array[indices] will now generate a sparse slice of array datadef compute(param: Any):with wp.ScopedDevice("cuda") as device: syntax (same for wp.ScopedStream(), wp.Tape())wp.vector(), and wp.matrix()I = wp.identity(n=3, dtype=float)wp.pos())wp.constant variables to be used directly in Python without having to use .val memberwp.struct typeswp.struct from functions--quick build for faster local dev. iteration (uses a reduced set of SASS arches)requires_grad parameter to wp.from_torch() to override gradient allocationwp.Tape()wp.matmul() with tape backward passwp.matmul()wp.launch(), up to 3x faster launches in common caseswp.randf() conversion to float to reduce bias for uniform samplingwp.func and wp.constant types from inside Python closureswp.constant variables can now be treated as their true type, accessing the underlying value through constant.val is no longer supportedwp.sim.model.ground_plane is now a wp.array to support gradient, users should call builder.set_ground_plane() to create the groundwp.sim capsule, cones, and cylinders are now aligned with the default USD up-axisReduce test time for vec/math types
Updated wp.from_torch() to support more data types
wp.from_torch() to support more data typeswp.from_torch() to automatically determine the target Warp data type if not specifiedwp.from_torch() to support non-contiguous tensors with arbitrary strideswp.matmul() and wp.matmul_batched()mat33 types, see wp.qr3(), and wp.eig3()wp.config.quiet = True)@wp.kernel(enable_backward=False)imp package with importlibwp.quat_slerp())Fix strides computation in array_t constructor, fixes a bug with accessing mesh indices through mesh.indices[]
Add smoothed particle hydrodynamics (SPH) example, see example_sph.py
example_sph.pyarray.shape inside kernels, e.g.: width = arr.shape[0]wp.Bvh and bvh_query_ray(), bvh_query_aabb() functionsspatial_vector, spatial_matrix typeswp.lerp() and wp.smoothstep() builtinswp.optim module with implementation of the Adam optimizer for float and vector typeswp.length_sq(), wp.trace() for vector / matrix types respectivelywp.quat_rpy(), wp.determinant()wp.atomic_min(), wp.atomic_max() operatorswarp.sim.model.add_cloth_mesh()wp.Volume.allocate(), and wp.Volume.allocate_by_tiles()wp.volume_store_i(), wp.volume_store_f(), wp.volume_store_v()example_jacobian_ik.pywp.Tape.zero() support for wp.struct typeswp.Mesh object accept both 1d and 2d arrays of face vertex indicesimportlib.reload()wp.constants() not invalidating kernels.ptx versions are presentUpdate all samples to use GPU interop path by default
wp.sample_unit_cube(), wp.sample_unit_disk(), etcwp.lower_bound() for searching sorted arrayswp.set_module_options({"enable_backward": False}), True by default__file__ attributewp.func() definitionswp.config.mode == "debug", this enables bounds checking on CUDA kernels in debug moderuntime = None errors when reloading the Warp modulewp.config.ptx_target_arch__init__.py, defer them to individual functions that need themwp.mesh_query_point()wp.HashGrid memory leak when creating/destroying grids@wp.struct referenceswp.is_cuda_available() to check availabilitywp.render classFix for marching cubes reallocation after initialization
wp.closest_point_edge_edge() builtinwp.sim.ModelBuilder.add_cloth_mesh()wp.float16wp.array.strides@wp.struct decoratorwp.config.verify_fpwp.MarchingCubes classr = m33[i]wp.log2() and wp.log10() builtinswp.sim.ModelBuilder objects to improve env. creation performance for RLwp.Tape.zero()requires_grad=Truewp.mat22 constructor adjointwp.closest_point_edge_edge() builtinwp.sim.ModelBuilder.add_cloth_mesh()wp.set_device(), wp.get_device(), wp.ScopedDevice"cuda:0", "cuda:1"), see wp.get_cuda_devices(), wp.get_cuda_device_count(), wp.get_cuda_device()"cuda"wp.map_cuda_device(), wp.unmap_cuda_device()wp.device_from_torch(), wp.device_to_torch()wp.get_preferred_device())wp.ScopedCudaGuard is deprecated, use wp.ScopedDevice insteadwp.synchronize() now synchronizes all devices; for finer-grained control, use wp.synchronize_device()"cuda" now refers to the current CUDA context, rather than a specific device like "cuda:0" or "cuda:1"NumPy implementations of many builtin methods have been moved to warp.utils and will be deprecated
wp.constant changes not updating module hashwp.array.grad, users should take care to always call wp.Tape.zero() to clear gradients between different invocations of wp.Tape.backward()wp.array.fill_() to set all entries to a scalar value (4-byte values only currently)capture option has been removed, users can now capture tapes inside existing CUDA graphs (e.g.: inside Torch)requires_grad=True at creation timefrom import * inside Warp initializationwp.curlnoise()wp.from_torch() to correctly preserve shape@wp.func functions based on argument typewp.copy() for detailswp.mesh_get_point(), wp.mesh_get_index(), wp.mesh_eval_face_normal()wp.quat_identity() now call the Warp native implementation directly and will return a wp.quat object instead of NumPy arraywarp.utils and will be deprecated@wp.func functions should not be namespaced when called, e.g.: previously wp.myfunc() would work even if myfunc() was not a builtinwp.rpy2quat(), please use wp.quat_rpy() insteadFix for unrolling loops with negative bounds
hash_grid_build_device() not found when lib is compiled without CUDA supportwarp.dll not found on some Windows installations@wp.func functionsx = array[i,j,k] syntax to address a 3-dimensional arraylaunch(kernel, dim=(i,j,k), ... and i,j,k = wp.tid() to obtain thread indiceswp.config.mode = "debug" to enablewp.mlp()wp.volume_sample_i(), wp.volume_sample_f(), and wp.volume_sample_vec()wp.Volume.load_from_nvdb()wp.intersect_tri_tri()wp.ScopedTimer()wp.inverse(), wp.quat_to_matrix(), wp.quat_from_matrix()--fast-math) to kernel compilation by defaultwarp.torch import by default (if PyTorch is installed)wp.volume_sample_world() is now replaced by wp.volume_sample_f/i/vec() which operate in index (local) space. Users should use wp.volume_world_to_index() to transform points from world space to index space before sampling.wp.mlp() expects multi-dimensional arrays instead of one-dimensional arrays for inference, all other semantics remain the same as earlier versions of this API.wp.array.length member has been removed, please use wp.array.shape to access array dimensions, or use wp.array.size to get total element countdense_gemm(), dense_chol(), etc methods as experimental until we revisit themYour coding agent can read these notes before it upgrades. Set up the MCP server →