NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #366 most downloaded on PyPI
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Last release 4 days ago
30 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 49 of 51 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
51 releases · first in 2018
Nothing published for this version
Nothing published for this version
The named tensor feature (a long-deprecated prototype) has been fully removed to reduce overhead and code bloat. All associated Python and C++ APIs ar…
<table> <tr><td><strong>FlexAttention</strong> lands on Apple Silicon (MPS), with up to ~12x speedup over SDPA on sparse patterns, and gains a deterministic backward path on CUDA for reproducible gradient computation.</td></tr> <tr><td><strong>CuTeDSL "Native DSL" backend</strong> gives Inductor a second high-performance code path (alongside Triton) for key GPU operations, with faster compilation. [Prototype]</td></tr> <tr><td><strong><code>nn.LinearCrossEntropyLoss</code></strong> combines the final prediction and loss computation to cut peak GPU memory by up to 4x for large-vocabulary language model training.</td></tr> <tr><td><strong>torchcomms</strong>, a new communications backend for PyTorch Distributed, improves fault tolerance, scalability, and debuggability for large-cluster training.</td></tr> <tr><td><strong>FSDP2</strong> now overlaps reduce-scatter and all-gather communications via a dedicated process group (opt-in), increasing distributed training throughput.</td></tr> <tr><td><strong>Python 3.15 wheel support</strong> for PyTorch on Linux via the pytorch repository index, including builds compatible with free-threaded 3.15t.</td></tr> <tr><td><strong>Broader platform support</strong>: ROCm gains AOTriton 0.12b with native HIP CMake, Arm adds Armv9-A <code>torch.compile</code> targeting, and Intel XPU exposes new device telemetry APIs.</td></tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
torch.compile on CPU in environments without a GPURunning a torch==2.13.0+rocm7.2 wheel in an environment where no GPU is available (torch.cuda.is_available() is False) breaks torch.compile on the CPU path: the first compile raises RuntimeError: Can't detect vectorized ISA for CPU (#189194). This is a regression from torch==2.12.1+rocm7.2, which compiles CPU code fine (detecting e.g. VecAVX2) in the same setup. The 2.13 ROCm wheel appears to rely on something present in the ROCm builder image to detect the CPU vectorized ISA, so it works when run on a ROCm image but fails on a plain CPU-only image.
Workaround: run the +rocm wheel on a ROCm image, or install a standard CPU/CUDA build for GPU-less environments.
Stop building CPython 3.13t (free-threaded) binaries (#182951)
Upstream pypa/manylinux removed CPython 3.13t (free-threaded) on 2026-05-07, because 3.13t
was experimental and has been superseded by the now-non-experimental CPython 3.14t. As a result,
PyTorch 2.13 no longer ships cp313t wheels (Linux, Triton, and related artifacts). Users on the
free-threaded interpreter should move to Python 3.14t.
PyTorch 2.12:
# cp313t (free-threaded 3.13) wheels were available
python3.13t -m pip install torch
PyTorch 2.13:
# Use free-threaded Python 3.14t instead
python3.14t -m pip install torch
Bare PyObject is no longer allowed in operator schemas (#184209)
Bare PyObject was accidentally accepted in operator schema strings in
PyTorch 2.12. This was undocumented and is now rejected, since torch.compile
does not support arbitrary PyObject inputs to custom ops. If
you parse or register a schema with a bare PyObject argument or return type,
you will now get a schema parse error.
PyTorch 2.12:
>>> from torch._C import parse_schema
>>> parse_schema("foo(PyObject x) -> ()") # accepted
PyTorch 2.13:
>>> from torch._C import parse_schema
>>> parse_schema("foo(PyObject x) -> ()") # raises a schema parse error
Remove Bazel build support (#180883)
The Bazel build was never broadly adopted and still depended on the antiquated Bazel 6,
while the wider ecosystem has since moved to Bazel 9. All Bazel build files and CI jobs have
been removed. Users building PyTorch with Bazel should migrate to the supported CMake/pip install
build flow.
PyTorch 2.12:
# Build PyTorch with Bazel
bazel build //:torch
PyTorch 2.13:
# Bazel build files have been removed; build from source with pip instead
pip install --no-build-isolation -e .
StorageImpl's built-in copy-on-write (COW) materialization is replaced by a pluggable materializer hook (#179063)
StorageImpl no longer knows about COW directly. Its internal COW entry points
StorageImpl::is_cow(), StorageImpl::maybe_materialize_cow(), and the friend
cow::materialize_cow_storage() have been removed in favor of a single pluggable
MaterializeFn hook (void(*)(StorageImpl*)) that a backend registers to run once,
on the first mutable data-pointer access. COW is now just one consumer of this hook
(c10::impl::cow::materialize_cow), and all COW behavior (lazy clone, refcounted
shared data, copy-on-write) is unchanged. This also gives accelerator backends and
eager-mode graph compilers a zero-fast-path-cost place to commit deferred allocations
or materialize symbolic buffers on first mutation.
This is a C++-only change. It affects out-of-tree backends/extensions that called the
removed StorageImpl COW symbols directly; they will fail to compile against 2.13
with errors such as no member named 'is_cow' in 'c10::StorageImpl'. Migrate to the
new hook API (set_materializer() / has_materializer() / clear_materializer()).
PyTorch 2.12:
// Detect a COW storage and force it to materialize.
if (storage.is_cow()) {
storage.maybe_materialize_cow();
}
PyTorch 2.13:
// Register a one-shot materializer; it runs on the next mutable-data access
// and then clears itself. COW registers c10::impl::cow::materialize_cow this way.
storage.set_materializer(&my_backend_materialize); // void(StorageImpl*)
// `has_materializer()` replaces `is_cow()` for "is a deferred materialization pending?"
if (storage.has_materializer()) { /* ... */ }
Convert shared_ptr<Node> to intrusive_ptr<Node> (#181139). This changes the signature of Tensor::grad_fn. Accesses to Tensor.grad_fn() should change from std::shared_ptr<Node> to c10::intrusive_ptr<Node>. Similarly, construction of a C++ autograd function should change:
PyTorch 2.12:
std::shared_ptr<CustomCppNode> node(new CustomCppNode(), torch::autograd::deleteNode);
PyTorch 2.13:
auto node = c10::make_intrusive<CustomCppNode>();
The minimum supported NCCL version when building from source is now 2.23 (#186292)
PyTorch now requires NCCL >= 2.23 at compile time, and the preprocessor/runtime gates that guarded NCCL features introduced in 2.23 or earlier have been removed. Users who build PyTorch from source against a system NCCL older than 2.23 will hit compile errors against the dropped gates. Upgrade the NCCL installation to >= 2.23 to build. The prebuilt PyTorch wheels already bundle a compatible NCCL, so pip/conda users are unaffected.
Remove named tensors (#173895)
The named tensor feature (a long-deprecated prototype) has been fully removed to reduce overhead and code bloat. All associated Python and C++ APIs are gone, including Tensor.names, Tensor.rename(), Tensor.refine_names(), Tensor.align_to(), Tensor.align_as(), torch.align_tensors(), the names= keyword on factory functions (e.g. torch.zeros, torch.empty, torch.ones), and the C++ Dimname / DimnameList APIs. Code that previously relied on named dimensions must track dimension order positionally and avoid usage of any of these now-removed APIs or op overloads.
The onednn::qconv2d_pointwise.binary and .binary_tensor operators no longer alias their input but rather return fresh tensors. Previously these ops mutated the qaccum input buffer and returned it directly, violating the PyTorch invariant that custom operator outputs must not alias inputs. This silently bypassed aliasing checks via the old -> Tensor(a!) schema and would become a hard error in a future PyTorch version (as mentioned in #182063), so the schema and implementation were corrected to return a fresh output. Most users are unaffected, only code that calls these ops directly and relies on the in-place mutation of qaccum must now read the returned tensor instead. (#177171)
Custom operators that return an output aliasing one of their inputs are deprecated (#182063)
When a custom operator returns an output that is the same tensor as (or otherwise aliases) one of its inputs under torch.compile, PyTorch now emits a UserWarning stating that this is deprecated and will become an error in a future version of PyTorch. Previously the warning stated the change would land in PyTorch 2.12; that timeline has been pushed back. To update your code, return a clone of the offending output instead of the input, or refactor the operator so it does not return the aliased tensor.
Deprecated:
@torch.library.custom_op("mylib::foo", mutates_args=())
def foo(x: torch.Tensor) -> torch.Tensor:
return x # output aliases the input -- deprecated
Updated:
@torch.library.custom_op("mylib::foo", mutates_args=())
def foo(x: torch.Tensor) -> torch.Tensor:
return x.clone() # return a clone instead
Creating tensors with the quantized dtypes quint8, qint8, and qint32 is now deprecated and emits a warning. This covers both Python and C++ call sites; see #184982 for migration guidance (#184984)
PyTorch 2.12:
>>> x = torch.quantize_per_tensor(torch.randn(3), 0.1, 0, torch.quint8)
PyTorch 2.13:
>>> x = torch.quantize_per_tensor(torch.randn(3), 0.1, 0, torch.quint8)
UserWarning: Creating tensors with quantized dtypes (quint8, qint8, qint32) is deprecated
Rename distributed collective ops to the _single naming scheme and deprecate the old names (#186123, #186124, #186125, #186134, #186135, #186144)
To align the public torch.distributed collective APIs with the naming used by torchcomms' TorchCommBackend, all_gather_into_tensor is renamed to all_gather_single and reduce_scatter_tensor to reduce_scatter_single. The previous names continue to work as thin wrappers that delegate to the new functions, but now emit a FutureWarning.
PyTorch 2.12:
dist.all_gather_into_tensor(output, input)
dist.reduce_scatter_tensor(output, input)
PyTorch 2.13:
dist.all_gather_single(output, input)
dist.reduce_scatter_single(output, input)
torch.Tag.inplace and torch.Tag.out, that let an operator declare how it writes its result: inplace means it mutates a tensor in place, and out means it writes into a caller-provided output tensor. Native PyTorch operators are tagged automatically, and custom operators defined with torch.library can opt in by adding the tag. To be tagged inplace, an operator must take the tensor it mutates as its first positional argument (declared as Tensor(a!), and the only mutable argument) and return that same tensor. Tagging a custom operator this way improves its behavior under torch.compile: inplace ops now go through auto_functionalize, so the reinplacing pass can analyze clones and skip unnecessary copies, and both inplace and out ops get their fake/meta kernels generated for free. See the Python custom operators tutorial for how to author and tag custom operators. (#181100, #181099, #184199, #184200, #184201, #184202, #184203, #180851, #180852)const_data_ptr() Python binding to torch.Tensor for read-only data pointer access (#180382)abbr property to torch.dtype that returns a dtype's short string abbreviation (e.g. torch.float32.abbr returns "f32") (#177296)Functions (#182206)rearrange in the torch.func namespace for einops-style tensor reshaping (#173183)nn.LinearCrossEntropyLoss, a fused linear-projection plus cross-entropy loss module that avoids materializing the full logits tensor (#181573, #185852, #172286, #186113)torch.autograd.graph.region_activation_memory_budget (#185979)dict to torch.autograd.grad and torch.autograd.backward (#178140)Add a registration API for symmetric memory arguments (lib.register_symm_mem_args()), letting operators (including out-of-tree ops) declare which arguments require symmetric-memory allocation (#173513)
Remove NCCLSymmetricMemory's explicit dependency on ProcessGroupNCCL, enabling symmetric memory to work with out-of-tree backends such as torchcomms (#184260)
Support accessing the ReduceOp.PREMUL_SUM factor from Python when implementing process group backends in Python (#185863)
Expose the NCCL 2.30 maxP2pPeers config binding (#181686)
Add rocSHMEM Triton integration for symmetric memory on ROCm (#178658)
Support passing extra keyword arguments to the loss function in pipeline schedules via a new loss_kwargs parameter to step(), enabling loss functions that require arguments beyond (output, target) (such as chunked cross-entropy needing token counts for scaling) (#181057)
FSDPModule.set_separate_reduce_scatter_group to give reduce-scatter its own NCCL communicator, enabling opt-in overlap of all-gather and reduce-scatter (#186335)set_reduce_scatter_max_input_buffers to keep multiple reduce-scatter input buffers in flight, so backward compute no longer stalls waiting to recycle a single reduce-scatter buffer (#186000)torch.compiler.set_default_backend to override the default torch.compile backend globally, so out-of-tree backend authors don't need to pass backend= at every call site (following the pattern of torch.set_default_dtype/torch.set_default_device). Explicit backend= arguments still take precedence (#178944)torch.compile(f, isolate_recompiles=True) to give each torch.compile call its own isolated cache bucket, preventing cross-compile interference in cache lookups and recompile-limit checks when multiple torch.compile calls target the same function (#178351)register_multi_grad_hook support to @leaf_function, allowing a backward hook to fire once per backward pass when all requires_grad inputs have their gradients computed (#179609)FlexAttention template (chosen when query length is 1) with a new configurable PARTITION_SIZE kernel option (#159835)decomp_comms) that eliminates all_gather for Gram-matrix optimizer patterns (Muon/Shampoo) under FSDP, gated by config.aten_distributed_optimizations.allow_comms_decompositions, yielding 1.25-1.95x training speedups (#184370)TORCH_TARGET_VERSION, so shims introduced in newer releases are only exposed when the target version supports them (#181916)torch._inductor.aoti_compile_and_package / aoti_load_package API, including packaging and loading of the multiple .so files emitted per kernel (#182251)torch_exception_get_what, torch_exception_get_what_without_backtrace, and STABLE_TORCH_ERROR_CODE_CHECK) so extensions built against the stable ABI can retrieve the original error message across the C API boundary (target version 2.13+) (#180135)aoti_torch_stream_native_handle and torch::stable::accelerator::Stream::nativeHandle(), gated behind TORCH_FEATURE_VERSION >= 2.13, for retrieving a native stream handle from the stable ABI (#183930)CUDAGraph.get_graph_data() for graph topology introspection (#183165)torch.distributions.Dirichlet on MPS by adding _sample_dirichlet and _dirichlet_grad Metal implementations (#185458, #185854)grid_sampler_2d backward support on MPS (#179756)grid_sampler_3d backward support on MPS (#179388)lcm support on MPS via a new Metal kernel (#186279)c10/metal/reduction_utils.h (#180708) and a complex->bool specialization (#185938)enable_language(HIP) (#180485)torch.xpu.* (#181082, #183427, #183428, #183429, #183430, #183431)scaled_mm on XPU (#173630, #176043)torch.load (#170592)Storage.pin_memory / Storage.is_pinned device-agnostic (#186223)op_overloads to OpOverloadPacket to enumerate an operator's overloads (#182993)num_splits in FlashAttention-2 and bump the flash-attention submodule (#179760)linear_bias in linear_cross_entropy on the reference and chunked paths (#185129, #185276)SequentialLR wrong learning rate initialization when milestones contain 0 (#185986)torch.nextafter (#148820)torch.autograd.enforce_grad_layout_policy to control the memory layout policy for accumulated gradients (#180552)When TorchComms is enabled, route new_group through split_group for subgroup creation, raising NotImplementedError for arguments split_group cannot honor (e.g. use_local_synchronization=True, sort_ranks=False) instead of silently falling back (#185416)
Delegate dist.new_group to custom process group subclasses (#184262)
Surface started-work metadata in NCCL watchdog timeouts (#183656)
Add a health check endpoint to the distributed debug server (#179326)
Make the DeviceMesh non-overlapping check stricter (#172343)
Allow elastic_launch/launch_agent to accept a pre-created torchelastic health check server, so it can be started before rendezvous (#180543)
Add an overlap_pp_comm flag to pipeline schedules (default True) that, when set to False, defers each pipeline RECV op to immediately before the compute op that consumes it, using rank-parity P2P ordering to avoid deadlock (helps platforms such as AMD ROCm where a pending RECV blocks unrelated compute) (#178815)
.out, inplace, functional, and foreach), expanding strategy coverage to hundreds of additional ops (#185386)scatter, upsample/interpolation backward, anti-aliased upsample, batch norm backward, and aten.detach_.default (#186149, #180311, #184626, #182743, #181876)torch.func.jvp) on models wrapped with fully_shard or replicate, including with mixed precision (#182732)Half and BFloat16 dispatch support for torch.trace on CPU (#184874)torch.linalg.lu (#185344).events() output (#180275)split_module now supports torch.Size crossing graph split boundaries by decomposing size() calls into per-dimension sym_size nodes, and builds submodules lazily for faster inference graph splitting (#179839)
CapabilityBasedPartitioner can now opt out of horizontal fusion via skip_horizontal_fusion=True, partitioning only through direct data dependencies (#184904)
Enable rewriting of FX traces containing complex tensors during compilation (#169832)
divmod (#185655)einops 0.8.2 (#185619), record_function as a decorator (#184703), inference_mode retracing helpers (#185066), mark_dirty in the autograd Function HOP (#184267), warn_only deterministic toggles (#180373), and the _maybe_view_chunk_cat functional collective (#180389)__setitem__/__delitem__) on more container types in Dynamo via sq_ass_item/mp_ass_subscript slots (#182862, #182996)torch.accelerator.device_index and torch.xpu.device in the device context manager (#181846, #181847)torch.compile: accept tl.constexpr values as kernel arguments (#181783) and handle capture_triton as a no-op during tracing (#183555)SeqSpec for list/tuple specs with better walk-spec errors (#185327), add ObjectSpec (#182764), pipe dynamic spec through torch.compile (#184501), and revisit guarding in mark_dynamic APIs (#181469)torch.compile device mismatch errors with a dedicated FakeTensorDeviceMismatchError and actionable guidance to place inputs, parameters, and buffers on the same device (#185412).any()/.all() (#180406), clearer torch._check tensor predicate errors (#185777), user-friendly reasons for skipped frames (#183596), carets in stack traces (#182393), and reporting why a symbol was created dynamically in symbolic_shapes logs (#168331)pin_memory for torch.ones, torch.zeros and torch.full (#174595)pointer_range_32 optimization (#179604)x_scale/x_zp and ReinterpretView strides (#181090)aot_inductor.autotune_per_kernel_alloc config to allocate-run-delete tensors per kernel during AOTI autotuning, avoiding OOM on large models (#181176)AsyncCompile future waits with the compile_worker_wait_timeout setting (#181293)a100_default_flex_config entries for head_dim=192 (#181835)aten.multinomial to avoid graph breaks (#182423)_scaled_mm_v2 (#182527)combo_kernel_autotune_grouping by default (#182567)cudagraph_partition_memory_budget config for partition reordering (#183569)bincount, unique variants, and AMP scale ops so they compile without graph breaks (#183590)addmm when the bias is a narrowing dtype cast (fp32->bf16/fp16) to preserve bias precision in XPU AMP training (#183680).grad buffer is invalidated on a later run (gradient accumulation) (#184003)index_add-style atomic scatter mutations into Triton template epilogues behind a config flag (#184179)cpp.march Inductor config knob so AOTInductor cpp-only builds can override or suppress the default CPU architecture flag (#184297)rsqrt, exp2, log2, log10, tan, acos, asin, atan, atan2, floor, logical_xor) (#184538)fake_mode argument to standalone_compile (with dynamic_shapes="from_example_inputs") so it can reuse the caller's FakeTensorMode/ShapeEnv instead of always creating a fresh one (#184776)signbit on unsigned integer dtypes (#185985)keep_static_cubin_raw config to retain cubin bytes in cached kernels so caches restored on another machine avoid recompilation (#186404)BatchLinearLHSFusion's matcher to also match the inlined torch._C._nn.linear form so the (opt-in) fusion can fire on Dynamo-inlined linear (#186632)torch._C._nn.linear operations are grouped into a single batched kernel (#180477)detach() method calls (#180513)update_constant_buffer (#181114)shim.h (#178120)cudaMemcpy for AOTI constant loading to reduce peak memory usage (#184823)proxy_executor error messages (#180884)cpp_wrapper (#182089)AOTIModelPackageLoader (#182149)torch.export save/load (#181676)torch.export-able (#179686)UpdateConstantBufferFromCpu for host-to-device copy (#181637)merge_view_inputs error messages, so non-differentiable view input mutation errors identify the specific offending inputs (#180424)_transformer_encoder_layer_fwd so it traces under torch.compile (#183916)torch.compile on AArch64 (#184555)c10d collectives in standalone compile (#181836)_foreach_max (#173483)adaptive_max_pool2d and adaptive_max_pool3d decompositions for ONNX export (#184396)set_python_module on torch::Library (#182720)== overloads for HeaderOnlyArrayRef (#185379)torch::stable::Generator (#186423)c10::layout typecaster for torch.layout (#179607)def_static (#175644)setup.py to CMake (#177641)BinaryDivFloorKernel.cu (#179260)opmath_t in i1 and i1e CUDA kernels (#183778)resize_ with address hint (#178215)bfloat16 in _embedding_bag_per_sample_weights_backward on CUDA (#185889)parsePerProcessMemoryFraction's return type with other parsers (#185139)cuda-bindings version is too old (#185990)torch.cuda.current_solver_handle for cuSOLVER handle sharing (#176705)cudnn_frontend submodule to 1.24 (#185554)bernoulli (#182210), native_dropout (#182232), uniform/normal/randint (#182386), randperm (#182528), comparison ops eq/ne/lt/le/gt/ge (#183019), bitwise ops (#182839), scatter/gather (#184028), copy-cast (#184740), gelu/gelu_backward (#181451), replication pad (#183065), embedding backward (#185119), trace (#183627), count_nonzero (#180725), amax/amin/aminmax/all/any (#180752), and cumsum/cumprod (#185609), native_group_norm (#183830, native_group_norm_backward #184437), topk and kthvalue Metal kernels (#184106), single-block and multi-block sort, including a stable sort path (#180714, #182242, #181736)histc (#178624)out variants of unary ops (#184743)is_causal support (#181855), head dim 256 (#181852), float mask support (#183458), and a clearer error for causal + attn mask (#181856).contiguous() calls: Col2Im (#181949), Im2Col (#182709), Repeat (#182718), LossOps (#182714), and HistogramKernel (#181951)NDHWC+DHWIO fast path for Conv3d on channels_last_3d (#184612)torch.mps.empty_cache() (#181485)stride > 0 in pool ops (#184875), F.fold (#182067), im2col (#183593), and bernoulli probabilities (#185065)dropout_p (#184126)cub::DeviceHistogram hipify mappings (#180433)head_dim != head_dim_v, use_deterministic_algorithms, gfx1100 and gfx1151 promoted out of experimental, partial FAv3 support on gfx950 (#184288)last_level_cache_size and is_integrated_gpu to XPU device properties (#184499, #182624)_fused_adagrad_ (#185577)torch.xpu.device in Dynamo device management (#181847)bmm_outer_product Triton override on XPU (#180441)vmap batching rule for torch.unbind_copy (#178035)vmap batching rule for Tensor.view(dtype) (aten::view.dtype) (#180728)mat1/mat2 layouts in sspaddmm and clarify error messages (#179037)opt_dtype validation to torch.nanmean() for consistent error handling (#172809)torch.nansum integer output dtype through nan_to_num + sum for correct results (#183808)CUDAStream::stream() (#184237)logspace/linspace ref tests with upstream XFAIL state (#178734)DataLoader file descriptor leak from atexit cleanup (#176607)stride/padding/kernel_size length in slow_conv3d (#181063)math_channel_shuffle (#181029)delta type in nn.HuberLoss constructor (#184012)torch/serialization.py so it matches the true values (#180959)layer_norm on CUDA for tensors with more than 2^32 elements (#181600)NestedTensor inputs in flex_attention with a clear error instead of an unclear backend failure (#183516)reflection_pad1d backward CUDA launch for large batches (#185024)lp_pool infinity norm handling (#183997)torch.autograd.enforce_grad_layout_policy decorator state leak (#183868)Fix NCCLComm::abort() to use the correct deregister API for window-registered handles (#181626)
Fix FakeProcessGroup all_gather on tensors that require grad (#181790)
Fix gather and allgather_coalesced on FakeProcessGroup to copy input to output (#182364)
Fix the scatter and reduce_scatter family on FakeProcessGroup to copy input to output (#182365)
Fix all_to_all on FakeProcessGroup and validate splits (#182366)
Fix conflict between broadcast_buffers and init_sync in DDP (#178054)
Fix gather on non-destination ranks for the TorchComms backend (#178533)
Fix TCPStore compilation with Clang 20 (#185785)
Fix NCCL symmetric memory mismatch by using an allocation-time counter instead of address for block ordering (#183489)
Fix a symbol lookup issue with the symmetric memory __init__ (#186416)
Fix the value returned by Work.exception() so the exception can be inspected from Python instead of being unusable (#184697)
Fix false assertion errors in the flight recorder when using the ncclx, gloo, rccl, rcclx, mccl, or hccl backends (#179753)
Fix a failure when creating a subgroup on a fake backend via new_group, which has no underlying communicator to split (#186172)
Fix torch.compile of the _c10d_functional all_gather_tensor_out and reduce_scatter_tensor_out ops, which previously failed functionalization with "Found a custom (non-ATen) operator whose output has alias annotations" (#183597)
Fix split_group on multi-backend process groups (e.g. init_process_group(backend="cpu:gloo,cuda:nccl")) to split only the relevant backend instead of every backend, avoiding spurious warnings, extra rendezvous overhead, and inconsistent process-group shapes (#182057)
Fix FakeProcessGroup to reject rank >= world_size at construction time, which previously failed silently and only surfaced later when collectives indexed past world_size (#182363)
Fix the torchelastic agent hanging indefinitely (and never exiting) when workers become stuck in an uninterruptible (D-state) process that SIGKILL cannot reap; the final proc.join()/proc.wait() in _close is now bounded by a timeout and the unkillable PID is logged (#185414)
Fix pipelining producing incorrect results or cryptic runtime errors when a PipelineScheduleMulti topology communicates between non-adjacent stages (e.g. skip connections); this is now detected at initialization and raises a clear RuntimeError (#179293)
Fix RuntimeError: only Tensors of floating point dtype can require gradients when building a pipeline for models with non-float intermediates (such as Hugging Face transformer models) (#183582)
Make LocalTensorMode transparent to torch.compile so compilation proceeds as if the debugging mode were not active (#182667)
Fix AssertionError in elastic c10d rendezvous when a node's rank changes across rendezvous rounds (e.g. a node becomes rank 0 after a peer leaves) (#182375)
Fix shared-weight gradient double-counting in zero-bubble pipeline schedules (#181365)
Fix None gradient handling in pipeline backward send/recv (#182182)
Fix a pipelining crash when split_module interleaves get_attr nodes with placeholder nodes (#182644)
FSDP.optim_state_dict and FSDP.optim_state_dict_to_load (#181261)DefaultStager crash when reused (#183424)squeeze leaving a DTensor's spec and local tensor out of sync (spec shape not matching the local tensor) by preventing squeeze from redistributing when strict_view is set (#175798).view() failures (e.g. in ColwiseParallel/RowwiseParallel) after an uneven Shard(dim>0)->Replicate redistribute by making the local tensor contiguous (#184443)make_fx, caused by sharding propagation recording dead global-shape shadow nodes into the graph (#185865)upsample backward crashes when the input is sharded (#182595)Partial placement being lost during the autograd layout invariant (#180511)OpSpec.mesh crash when specs contain None entries (#181541)redistribute(backward_dtype=...) ignoring the backward dtype (#182032)_StridedShard flag conflict during gradient accumulation (#183517)FakeTensor device hint in sharding propagation (#183970)group_norm scalar adjuster crash when weight=None (#184819)to_local() dropping the _is_param marker that nn.Parameter sets on custom tensors (#184422)pad_tensor/unpad_tensor creating unnecessary guards on symbolic pad sizes during tracing (#180887)torch.compile crash caused by an unbound inner symbol at the root tracer (#181797)DTensorSpec refcount leak in OpSchema._recompute_comparison_key (#181792)post_accumulate_grad_hook results under CPUOffloadPolicy (#180666)Tensor/DTensor error in chunk_cat (#183040)IndexError for modules called with no forward inputs by preserving empty args/kwargs, matching FSDP1 behavior (#183943)DTensor with a stale fp32 dtype, causing sharding propagation to fail in eager and torch.compile (#183805)fully_shard([norm, head]) group for chunked loss, where the model forward runs norm only and head is called standalone per chunk; unshard/reshard previously relied on _modules_to_run_forward and produced wrong results for this pattern (#180428)clip_grad_norm results with multiple data-parallel shard axes (e.g. dp_shard + cp passed via DataParallelMeshDims), which exposed separate Shard axes with the wrong float32 reduction order and an incorrect Shard instead of _StridedShard; the axes are now flattened into a single shard axis in the sharding spec (#183629)cast_forward_inputs is enabled, by casting forward inputs during recompute (#182580)_StridedShard (#186126)torch.linalg.ldl_solve CPU kernel (#181032)dict update (#185428), defaultdict inplace union (#185429), frozenset copy identity (#185430), sequence search (#185431), iand on bool constants (#184503), sequence * SymNode spurious graph break (#185260), and torch.Size tensor shape handling (#184613)float/bool + SymNode (#183362), PendingUnbackedSymbolNotFound for 0-d tensor Scalar args (#182660), GuardOnDataDependentSymNode on sparse tensors (#179616), and creating symbolic tensors from foreign fake tensors (#181794)CUDAStream/Event tp_dealloc overrides (#183403) and Dynamo dict guard cleanup (#183753)DeviceMesh is constructed inside torch.compile (#177201)torch.compile crash when an unsupported type is passed to a tensor method inside try/except (#182106)wrap_inline for exec'd Python functions (#181531)torch.compile (#183337)torch.full validation for nn.Parameter fill values in Dynamo (#183915)BlockMask placeholders (#184611)CudagraphsBackend.__call__ (#182989)def forward(self, ..., self, ...) SyntaxError in dynamo_graph_capture_for_export (#185314)context_fn clobbered by DDPOptimizer's propagate_metadata (#179496).grad reads for new in-graph parameters (#184972)manual_seed in Dynamo so compiled random calls stay reproducible (#185761)to_dense no-op (#184586)IndexError in compile mode matching eager mode (#184856)fullgraph=True and a non-default stance is set (#183623)is_compiling flag for the whole torch.compile session (#184614)max_autotune BMM correctness with dynamic OpenMP threads (#169128)LayerNorm by guarding the Welford variance computation (#173989)NameError from Python name mangling for user-defined Triton kernels whose names start with double underscores (#176100)pointer_range_32 to user-defined Triton kernels on ROCm to avoid compilation crashes (#178541)e8m0_rceil_log2 pattern failing to register on any CUDA device (#178698)cpp_wrapper backward compilation failing with a CudaKernelParamCache assertion when set via torch.compile options (#178847)FxGraphCachePickler crash on unpicklable pybind11 extension types during cache key computation (#178853)argmax/argmin indices in the CPP Tile2D vectorized kernel for transposed inputs (#179525)mutation_outputs for the intermediate buffer and add epilogue source to the cache key (#179803)tl.atomic_add for constant-index stores (#179833)fmod/remainder on CPU (#179923)dislike_padding conflicting with contiguous storage layout (#180197)check_bounds forward-reference C++ compile error in the CPU backend (#180212)waves_per_eu, matrix_instr_nonkdim, kpack) when combo kernels rewrite per-subkernel configs on ROCm (#180277)torch.compile crash on cumsum with broadcast input when the scan dimension is 129 or larger (#180369)while_loop codegen for the minimal-arrayref interface (wrap ArrayRefTensor inputs before assign) (#180370)addmm CUTLASS codegen and relax addmm tolerances (#180432)as_strided (#180581)NoValidChoicesError for scaled_mm/scaled_grouped_mm with use_fast_accum=True on Blackwell (#180586)UnicodeDecodeError on Windows when using icx-cl as the CXX compiler by decoding subprocess output with errors='replace' (#180853)tl.float64 in Triton scalar-shape-math codegen on XPU devices without fp64 support (e.g. Intel Arc A770) (#180854)adaptive_avg_pool2d is fused with downstream flatten + reduction for non-power-of-2 channel sizes (#180898)num_warps on AMD RDNA (wave32) GPUs so it is no longer incorrectly halved; only CDNA (warp_size 64) is halved (#181112)index_expr/kindexpr error when using flex_attention with bias/index expressions (#181207)slice lowering early-return path (#181219)fallback_random is enabled (#181245)NotImplementedError: View) when a flatten()-produced view is passed as the conv2d bias (#181363)bfloat16) when running under another (e.g. float16) (#181564)_to_copy decomposition so float64->float32 transfers to XPU devices without fp64 support no longer generate illegal fp64 buffers/crash (#181607)AutoHeuristic initialization from crashing in non-CUDA (e.g. XPU) environments (#181745)AttributeError crash in remove_noop_ops when a node's meta val is a SymInt/SymFloat (auto dynamic shapes) instead of a Tensor (#181752)ConcatKernel channels_last contiguous check to handle symbolic shapes, avoiding a data-dependent (DDE) LoweringException (#181845)tl.load/tl.store being incorrectly eliminated (and not allocated) when epilogue fusion is enabled (#181868)_CollectiveKernel.create_inplace (#181930)print_performance in generated benchmark_compiled_module so timing works on non-CUDA devices (e.g. XPU) (#181957)CompiledFxGraph._original_gm via GraphPickler to fix AOTAutograd/precompile cache serialization crashes for graphs with HOPs that have lifted buffers (e.g. flex_attention with a causal BlockMask) (#182088)maybe_realize in FlexAttention (#182610)AutoHeuristic crash in non-CUDA environments (e.g. XPU and CPU-only builds) (#182614)emulate_precision_casts on CPU C++ codegen so explicit/emulated fp32->fp16->fp32 precision barriers are no longer optimized away (#182882)count_flops_fx) for HigherOrderOperator targets such as flex_attention during overlap/bucketing (#182992)index_add/index_copy/index_reduce instead of silently succeeding under torch.compile (#183007)eager_input_vals propagation (#183334)is_linear_add_bias crash when the bias node argument is a Python float (#183514)torch.compile crash on cumprod backward by disallowing fusion of split-scan nodes with reductions (#183653)argmax logical index for transposed reductions (#183655)while_loop backward expanded gradient strides (#183658)torch.xpu stream APIs instead of CUDA-only calls (#183693)SymInts in associative_scan lowering (#183706)set_.source_Tensor lowering when the source is a view (#183724)select_algorithm input storage check (#183791)log for strict Inductor numerics (#183844)div_ to match eager type-promotion semantics (#183859)handle_synced_deallocation for XPU (#183865)hipModuleLoadData in StaticCudaLauncher on ROCm (#183926)flex_attention autotune logging to use size hints for symbolic dims, avoiding lowering failures under dynamic shapes (#183933)tuned_addmm with max-autotune when the bias is an unrealized view (e.g. a transposed expression) (#183973)gelu backward in opmath dtype for improved numerical accuracy (#183985)float8_e4m3fn dequantization on pre-sm89 CUDA devices to match eager bit patterns (#184008)bucketize/searchsorted on sliced or zero-stride lookup tensors (#184043)needs_fixed_stride_order so fallback custom ops receive tensor-list inputs in the recorded stride order (#184098)asinh overflow (#184105)RReLU (#184136)torch.jagged layout to match eager and meta behavior (#184146)out_dtype for scaled_mm by propagating it to MMKernelInputs (#184168)frexp exponent for non-finite inputs (#184176)convolution_backward lowering so dilated conv2d backward compiles under dynamic shapes (#184224)sort in combo kernels (#184227)as_strided storage offsets for graph input views (#184232)convolution_backward lowering so the Triton backward conv template no longer references undefined sympy symbols under dynamic shapes (#184255)index_put (#184263)aten.topk on ROCm/HIP to avoid a rocPRIM memory fault on cudagraph-tree replays (MI350 / ROCm 7) (#184265)gather bounds for negative indices (#184287)ExternKernelCaller) choices (#184493)INT_MIN / -1 matches eager instead of trapping (#184497)cumsum/cummax/cummin/cumprod) on unsigned integer dtypes (#184514)constant_offset alignment in TMACompatibilityChecker (#184564)acc_type for fp8 dtypes in mm autotuning (use triton_type and accumulate in fp32) (#184591)torch.cond subgraph buffer reuse by scoping EfficientPeakEstimate per subgraph during AOT codegen (#184623)addcmul_ and addcdiv_ under torch.compile (#184629)ZeroTensor view with symbolic sizes (#184651)cpp_micro_gemm that broke AOTI compilation of AVX512 VNNI/WoQ-Int4 kernels (#184693)constant_args and kwargs in ExternKernel.get_read_writes() to prevent index producers from being incorrectly eliminated (#184751)Py_None reference counting in the C++ wrapper (#184869)torch.compile runs inside a multiprocessing child process (#185070)torch.compile(dynamic=True) with torch.func.grad) (#185080)get_stride_order for assorted A/B layouts (#185437)NameError in combo kernels when TMA is enabled by handling DelayReplaceLine load expressions (#185514)cat lowering assertion error with 1-D statically-empty tensors (#185549)wrap_triton when TRITON_INTERPRET=1 is set (#185597)wait_stream reordering by registering it among the synchronization ops that preserve control dependencies (#185627)standalone_compile cache artifact save for grad-enabled graphs whose backward is lowered lazily (#185635)emulate_precision_casts to match eager (#185847)arange) in tensor value computation (#185853)view.dtype lowering when falling back (#185879)FakeTensorUpdater handling of higher-order ops and their subgraphs (#185962)get_stride for PermuteView to avoid a NotImplementedError (#185992)SymFloat under dynamic shapes in overlap scheduling (#186054)torch.bincount so downstream codegen uses the correct dtype (#186077)scatter_upon_const_tensor rewrite for low-precision const tensors (#186481)tl.broadcast_to(False, ...) by using tl.full instead (#186621)diagonal_scatter backward under torch.compile (#185146)SymInt crash in the overlap scheduler's collective/compute node benchmarking (#186065)torch.cat axis handling in Inductor pre-grad fusion (#183995)fp32->bf16->fp32 casts being dropped (#180575)atan numerics in Inductor (#183984)flip on 0-d tensors in the prims.rev lowering (#184104)softmax decomposition for symbolic empty dims (#184454)CppWrapper due to false-positive caching (#178147)cpp_wrapper (#182825)pointer_to_optional_list (#183764)c10::make_scope_exit to avoid exception leaks (#184520)AOTInductorModelContainer::run() during concurrent constant folding (#181941)CUmodule handles to prevent GPU code object leaks (#184860)cpp_wrapper temporary arrays (#179846)cpp_wrapper to fix MSVC compilation (#180120)float('inf')/float('-inf') kernel args (#180297)cond subgraph arrayref dispatch with generic lambda (#180558)cpp_wrapper while loop carried mutations (#183657)TORCHINDUCTOR_CACHE_DIR (#185723)IndexError during decomposition by also excluding lifted tensor constants and custom objects when identifying user-input placeholders (#181179)torch.export save/load (#181263)Min/Max of scaled symbolic terms (e.g. Min(128*s, 512*s) reduces to 128*s) so export no longer rejects valid branch guards (#185092)AssertionError on later eager buffer assignment (#184956)torch.export.load GIL contention during tensor deserialization (#175983)torch.compile crash with batched matmul in inference_mode (#181913)pixel_shuffle, pdist, and reflection/replication padding ops (#183814)max_pool2d/3d_with_indices (#179104)MaxUnpool output sizes in meta/decomposition kernels (#184706)conv2d kernel size in meta and symbolic-shape kernels (#180448)addmv decomposition dtype validation (#184140)fill_ meta value-tensor dimensionality validation (#179363)_weight_int8pack_mm meta inner-dims and scales validation (#179364)torch.empty(..., out=...) shape validation under torch.compile (#182349)torch.compile wrong output shape for norm() with a negative dim (#182405)frac decomposition signed-zero handling (#183640)pad_sequence mixed-dtype padding decomposition (#184173)istft fake tensor length padding (#184532)unfold_backward decomposition for overlapping windows (#183996)Tensor decomposition in Inductor (#184134)aten.hardtanh meta semantics for export (#185298)_fused_dropout decomposition at keep-probability zero (#184979)addmm decomposition crash with out_dtype under FakeTensorMode (#179634)torch.split decomposition for empty dim with nonzero split_size (#181493)torch.distributions.Gamma under torch.compile (#174090)miopen_batch_norm meta save_mean/save_var dtype (#179365)torch.compile (#179837)lp_pool2d compilation (#184000)hardtanh_backward decomposition (#185840)torch.sigmoid() in silu_backward decomposition (#185041)index_copy decomposition shape checks (#184338)zero (#185360)non_overlapping_and_dense (#186785)_cslt_sparse_mm meta registration for hipSPARSELt (#181609)FakeTensor (#183397)ProxyTensor (#183398)mix_order_reduction over-fusion via load count check (#179494)torch.compile crash from aten.lift functionalization on an already-functionalized tensor (e.g. randint followed by lift) (#185805)_foreach_sub under compile (#184421)CastLike handling logic from OpRecorder (#182197)_rotary_embedding_23_fake_impl stride drift for 3D and 4D inputs (#184854)invoke_subgraph export with lifted tensor constants (#182230)mode (#186428)-Winconsistent-dllimport warning in tensor_numpy and tensor_new headers on Windows (#183703)ValueError/NotImplementedError on Windows (#175340)nvrtcCompileProgram changing locale in CUDA < 12.6.2 (#180569)total_weight before accumulating in nll_loss2d (#182082)dtype promotion in max/min kernel (#181505)torch.cuda.ExternalStream(0) to wrap the NULL stream (#183258)native_group_norm in eager (#183946)multinomial SIGSEGV (#180493)F.linear on M5+ (#181466)uint32 offset overflow in scatter/gather kernels for strided views crossing 2^32 elements (#182054) and move col2im offset/stride to long to avoid overflow corruption (#185664)relu (#183571), softshrink (#183710), hardsigmoid (#183939), cholesky (#184588), and fast::tanh overflow (#186286); return NaN for std/var on empty input (#184510)layer_norm_backward silent correctness bug for frozen inputs (#183893)_amp_foreach_non_finite_check_and_unscale_ zeroing fp16/bf16 grads (#184286) and stop ignoring grad scale and found_inf (#186360)copy_kernel_mps (#184403)sort returning out-of-bounds indices for bool/int-max/NaN inputs (#184620)fill_ on byte-dtype views with misaligned storage offset (#183790)precise::sincos (#184749)deviceCount() consistent with Python to fix at::manual_seed() (#164571)scale not being cached (#184122)Support TheRock wheel distribution in _find_rocm_home (#180723)
Fix warpMergeSortTopK padding sentinel for integer dtypes (#182212)
Guard ck_group_gemm on USE_ROCM_CK_GEMM (#182615)
Fix large arange launch (#182657)
Fix triu/tril for 64-bit indexing for large matrices (#179717)
Drop dead CUDA/ROCm version gates from tests and helpers (#184879)
Fix LayerNorm backward kernel for AMD Strix Halo GPUs (#183864)
Decline CuteDSL scatter_add on ROCm (#185678)
For HSTU, fix CK flash-attn GQA seqlen_q==1 garbage output (#186434)
Inductor fixes:
pointer_range_32 optimization (#179604)maybe_hipify_code_wrapper for bare-token inputs (#183725)StaticCudaLauncher (#183926)lookup_device_info is now case-insensitive (#182284)Windows
CMAKE_HIP_FLAGS (#183856)dllimport (#183690, #183324, #183282, #183694)CMAKE_HIP_FLAGS (#183365)USE_ROCM_CK_SDPA on Windows (#183962)CurrentWorkStream (#179140)addmm shape handling and addmv_out stride preservation on XPU (#180985, #178498)XPUPluggableAllocator registration (#183865, #179392)logcumsumexp with complex inputs on XPU (#174492)SyclExtension Windows builds for oneAPI 2025.3 and later (#170701)getGlobalIdxFromDevice(-1) handling on XPU (#181361)MemoryReadAdapter::read (#181193)broadcast_shapes op missing in selective builds (#180860)binary_cross_entropy SymInt error with dynamic shapes by registering aten::broadcast_shapes as a TorchScript builtin (#180583)policy_fn during recompute (#176455)Speed up store-based metadata exchange on TCPStore by using multiGet and a server-side barrier, reducing network round trips from 2*(world_size-1) to 1 (#182132)
Coalesce the NCCL buffer and signal pad into a single symmetric-memory allocation so window registration runs only once (#183344)
Fuse slice-cat TP collective patterns (#184911)
GraphModule reconstruction in CSEPass when no common subexpressions were eliminated (#185479)repr in Dynamo ID_MATCH guard text (#184796)ao::offload ops to avoid per-tensor cudaHostAlloc overhead (gated by the pinned_memory_pool() context manager) (#186162)auto_functionalize, improving FP8 KV-cache performance (#173177)basic_gnn_sage fp32 single-thread performance regression (#177958)_FastCudaLauncher, a vectorcall C extension for pre-bound kernel launch that reduces per-launch overhead (#180507)sympy.gcd on very wide shape expressions, cutting some backward compiles from over 50 minutes to about 6 (#181275)num_stages=4) matmul config that speeds up large Hopper matmul shapes by ~1.3x and up (#181413)nested_compile_region (invoke_subgraph) subgraphs created with options, which were previously skipped (#181834)can_fuse_vertical, reducing kernel count (#183521)grid_sampler_2d lowering on CUDA/XPU when sizes fit, avoiding unnecessary int64 arithmetic (#184269)evict_first for coalesced last-use loads in persistent reductions (#184395)log in reused CPU pointwise so T5-style softmax bias inputs are materialized once (#184473)_safe_softmax SDPA math path back into a native scaled_dot_product_attention call (new SFDP patterns 29/30) (#185574)forward_inference prop kind only for channels-last/MKLDNN-layout, fixing a ~2x dense-contiguous slowdown (#185997)Note truncated.
One column per quarter.
This release is meant to fix the following regressions and silent correctness issues:
This release is meant to fix the following regressions and silent correctness issues:
Fix SyclExtension Windows build for oneAPI 2025.3+ breaking change
<table> <tr><td><strong>Batched linalg.eigh on CUDA</strong> is up to 100x faster due to updated cuSolver backend selection.</td></tr> <tr><td>New <strong>torch.accelerator.Graph</strong> API unifies graph capture and replay across CUDA, XPU, and out-of-tree backends.</td></tr> <tr><td><strong>torch.export.save</strong> now supports Microscaling (MX) quantization formats, enabling full export of aggressively compressed models.</td></tr> <tr><td><strong>Adagrad</strong> now supports <code>fused=True</code>, joining Adam, AdamW, and SGD with a single-kernel optimizer implementation.</td></tr> <tr><td><strong>torch.cond</strong> control flow can now be captured and replayed inside CUDA Graphs.</td></tr> <tr><td><strong>ROCm</strong> users gain expandable memory segments, rocSHMEM symmetric memory collectives, and FlexAttention pipelining.</td></tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
Strengthened SVE compile checks in FindARM.cmake, which may reject previously accepted but incorrect SVE configurations (#176646)
Source builds that enable SVE now validate the compiler configuration more strictly. If a build previously passed with an incomplete or mismatched SVE setup, it may now fail during CMake configuration instead of later in compilation. Update the compiler/toolchain flags so they accurately describe the target SVE support, or disable SVE for that build.
Updated the minimum CUDA version required to build PyTorch from source to CUDA 12.6 (#178925)
Building PyTorch from source with CUDA versions older than 12.6 is no longer supported. Users building custom binaries should install CUDA 12.6 or newer and make sure CUDA_HOME points to that installation.
Version 2.11:
CUDA_HOME=/usr/local/cuda-12.4 python setup.py develop
Version 2.12:
CUDA_HOME=/usr/local/cuda-12.6 python setup.py develop
Enforced a C++20 minimum in CMake build files (#178662)
Source builds now require a compiler and build configuration that support C++20. If you maintain custom build scripts or downstream extensions that build PyTorch from source, update the compiler and remove assumptions that PyTorch can be built as C++17.
torch.distributed.nn.functional ops now raise RuntimeError under torch.compile (#177342)
All ops in torch.distributed.nn.functional (e.g., broadcast, all_reduce, all_gather, reduce_scatter, all_to_all_single) now raise RuntimeError when called inside torch.compile. Users should migrate to the functional collectives API in torch.distributed._functional_collectives.
Version 2.11:
@torch.compile
def my_func(x):
return torch.distributed.nn.functional.all_reduce(x, op=ReduceOp.SUM)
Version 2.12:
@torch.compile
def my_func(x):
return torch.distributed._functional_collectives.all_reduce(x, reduceOp="sum", group=group)
torchrun now defaults to an OS-assigned free port for single-node training instead of port 29500 (#175699)
When running torchrun --nproc-per-node=N script.py without specifying --master-port or --standalone, the default behavior now automatically uses an OS-assigned free port via the c10d rendezvous backend. This eliminates "Address already in use" errors when running multiple training jobs concurrently. Multi-node training, explicit --master-port, PET_MASTER_PORT env var, and --standalone are unchanged.
Version 2.11:
# Used static rendezvous on port 29500 by default
torchrun --nproc-per-node=4 train.py
Version 2.12:
# Uses OS-assigned free port by default
torchrun --nproc-per-node=4 train.py
# To explicitly use a fixed port:
torchrun --nproc-per-node=4 --master-port=29500 train.py
All MPS tensors are now allocated in unified memory (#175818)
Previously, MPS tensors could be allocated in either device-only or unified memory. Now all MPS tensors use unified memory unconditionally. This simplifies memory management and enables CPU access to MPS tensor data without explicit copies. Code that relied on device-only memory placement may observe different performance characteristics.
The max_autotune layout-constraint deferral introduced in 2.11 is now opt-in (#175330)
In 2.11, Inductor deferred layout freezing for max_autotune templates to expose more fusion opportunities. This caused a regional-inductor failure mode, so the default in 2.12 reverts to immediate layout freezing. Users who relied on the deferred behavior for fusion opportunities should opt in explicitly via torch._inductor.config.max_autotune_defer_layout_freezing or TORCHINDUCTOR_MAX_AUTOTUNE_DEFER_LAYOUT_FREEZING=1.
Version 2.11:
# Deferred layout freezing was the default
torch.compile(model, mode="max-autotune")
Version 2.12:
import torch._inductor.config as cfg
cfg.max_autotune_defer_layout_freezing = True
# or set TORCHINDUCTOR_MAX_AUTOTUNE_DEFER_LAYOUT_FREEZING=1
torch.compile(model, mode="max-autotune")
Deprecate CUDA 12.8 builds in favor of CUDA 13.0 (#179072)
CUDA 12.8 binaries have been removed from the PyTorch binary build matrix. CUDA 13.0 is now the stable default and CUDA 12.6 remains available for users on older drivers. Users explicitly pinning the cu128 index URL will need to switch to cu130 (recommended) or cu126.
Version 2.11:
pip install torch --index-url https://download.pytorch.org/whl/cu128
Version 2.12:
# Use CUDA 13.0 (default on PyPI):
pip install torch
# Or explicitly:
pip install torch --index-url https://download.pytorch.org/whl/cu130
# Older driver fallback:
pip install torch --index-url https://download.pytorch.org/whl/cu126
Compatibility with CMake < 3.10 will be removed in a future release (#166259)
Source builds against CMake versions older than 3.10 now emit a deprecation warning. A future release will require CMake 3.10 or newer; please upgrade CMake before then.
Several CUDA linear algebra operators no longer use the MAGMA backend and now dispatch to cuSolver or cuBLAS unconditionally:
torch.linalg.eigh now dispatches to cuSolver (#174619)torch.linalg.lu_solve now dispatches to cuSolver/cuBLAS (#174248)torch.linalg.cholesky_inverse now dispatches to cuSolver (#174681)torch.linalg.cholesky_solve now dispatches to cuSolver (#174769)User code calling these APIs does not need to change. The practical impact is for users who depended on MAGMA-specific numerical behavior, performance characteristics, or debugging. Those calls now use the cuSolver/cuBLAS implementations on CUDA.
Compiling through FSDP2 hooks without graph breaks is no longer supported (#174863, #174906). If you use compiled autograd with FSDP2, update your code to allow graph breaks around FSDP2 hooks or disable compiled autograd for the FSDP2 training step.
Version 2.11:
with torch._dynamo.config.patch(compiled_autograd=True):
compiled_model = torch.compile(fsdp_model, fullgraph=True)
loss = compiled_model(input).sum()
loss.backward()
Version 2.12:
# Either run FSDP2 backward without fullgraph.
compiled_model = torch.compile(fsdp_model, fullgraph=False)
loss = compiled_model(input).sum()
loss.backward()
# Or apply compile before applying FSDP.
compiled_model_pre_fsdp = torch.compile(model, fullgraph=True)
compiled_model = fully_shard(compiled_model_pre_fsdp, ...)
loss = compiled_model(input).sum()
loss.backward()
Profiler's metadata_json field is now deprecated; use event_metadata instead (#179417)
Version 2.11:
metadata = event.metadata_json
Version 2.12:
metadata = event.event_metadata
torch.compile(fullgraph=True) now warns when a call runs no compiled code; will error in 2.13 (#181940)
Previously fullgraph=True was only validated once Dynamo actually compiled and ran the function. If Dynamo was bypassed at call time (e.g. under a user-defined TorchDispatchMode), the annotation silently had no effect. 2.12 emits a warning; 2.13 will raise. For graph-break errors without fullgraph's stronger guarantees, use torch._dynamo.error_on_graph_break.
Version 2.12:
from torch.utils._python_dispatch import TorchDispatchMode
class LoggingMode(TorchDispatchMode):
def __torch_dispatch__(self, func, types, args=(), kwargs=None):
return func(*args, **(kwargs or {}))
@torch.compile(fullgraph=True)
def model(x):
return x.sin() + 1
# A user-defined TorchDispatchMode is active, so Dynamo skips the frame
# and no compiled code runs — emits a warning in 2.12, will raise in 2.13.
with LoggingMode(): # Remove this to fix warning
model(torch.randn(3, 4))
The inline_inbuilt_nn_modules Dynamo config is deprecated (#177489, #178205)
Inlining of in-built nn.Module instances is now the default; setting the flag emits a deprecation warning and it will be removed in a future release.
Version 2.11:
import torch._dynamo.config as cfg
cfg.inline_inbuilt_nn_modules = True # was a tunable knob
Version 2.12:
# No action needed — inlining is on by default.
# Remove any explicit references to torch._dynamo.config.inline_inbuilt_nn_modules.
Added a deprecation framework to the torch.compile config module so individual options can be marked deprecated (#169837)
torch.accelerator.Graph as a unified frontend Graph interface (#171285)_foreach_clone operator, with a fast path for CUDA utilizing _foreach_copy_ (#177421)Store::barrier API and TCPStore client BARRIER support, reducing synchronization round trips compared to the existing ADD+WAIT pattern (#174920)suspend(), resume(), and memory_stats() APIs for managing communicator memory lifecycle (#176300)all_to_all support in the Gloo backend (#165435)reduce_scatter_offset to symmetric memory, supporting variable-sized block reductions with NVLink multicast or LSA fallback (#177791)batch_isend_irecv to work under torch.compile (#161213)torch.distributed.symmetric_memory.is_symm_mem_tensor() API to check if a tensor is a symmetric memory tensor (#178947)NanCheck to a standalone op (torch.ops.c10d.check_for_nan) usable outside of ProcessGroupNCCL (#174990)grad_placements parameter to DTensor.from_local(), allowing explicit control over gradient placements in the backward pass (#175867)fully_shard with DTensors on a full SPMD mesh via DataParallelMeshDims (#176334)--shutdown-timeout to torchrun for controlling the SIGTERM-to-SIGKILL timeout during worker shutdown (#172596)CPUBlas brgemm API for fp8 (e4m3 & e5m2) GEMM, backed by oneDNN (#172548)torch.cond with CUDA graphs, using conditional graph nodes (CUDA 12.4+) so data-dependent control flow can be captured entirely inside a single CUDA graph. Works with the eager and cudagraphs torch.compile backends (no Inductor support yet). (#168912)linalg_qr for MPS (#172536)cholesky_solve support on MPS (#176703)index_reduce on MPS (#174936)torch.distributions.Gamma (forward + backward) on MPS (#179228)mvlgamma on MPS (#178914)nonzero_static implementation on MPS (#179589) (from miscategorized)torch.accelerator.Graph on XPU (#176421)memory_clock_rate and memory_bus_width to XPU device properties (#171967)split_group API when TorchComms is used as a backend for TorchTitan on XPU (#178236)torch._dynamo.aot_compile public, with aot_eager and inductor backend support and docs (#179917, #180008)recompile_limit keyword argument to torch.compile to override the per-function recompile cap without touching global config (#177936)torch._dynamo.mark_unbacked for communicating value ranges to the symbolic shape system (#176313)bdb, a pdb-style debugger for stepping through nested frames during Dynamo tracing (n, u, d, r, bt), plus a user-callable breakpoint() that auto-starts it (#174626, #174746, #175200)torch.compile. Inductor now codegens stream context managers (enter/exit) and record_stream calls in the wrapper, enabling user streams to flow through compiled regions with proper synchronization, scheduler integration, and cross-stream dependency tracking (#165390, #165391, #165504, #165505, #174223, #176700, #177694)ao::offload, ao::reload, and ao::wait ops for asynchronous activation offloading. These ops encapsulate async CPU offloading stream management following the same async 2-op pattern as c10d functional collectives, reducing IR size from 7 nodes (offload) and 5 nodes (reload) down to 2 nodes each (#177621)relu()), parsing the user kernel source via AST and inlining the epilogue into the tl.store expression (#173662).out overloads, Inductor automatically lowers single-output and multi-output functional ops to their .out variants as ExternKernelOut, enabling memory planner buffer reuse (#175116, #176117)max_autotune now extends to combo kernels. The autotuning pipeline generates and benchmarks per-sub-kernel block-size phase configs, with chained sequential autotuning and per-sub-kernel reduction hints (#177715, #178936, #179317)mm and addmm for max-autotune, enabling persistent kernels on hardware without TMA (#177781, #179095)torch.float8_e5m2 dtype, including registration for FP8 GEMM autotuning (#171176)max-autotune-gemm, allowing CUTLASS-style GEMM templates to target Intel GPUs (#161938, #161939)sort, median, and mode operations (#178525)conv1d (#175280)at::vec::convert for the Inductor C++ x86 backend (#172309)disable_welford_reduction config flag to opt out of Welford reduction in codegen (#175778)float8_e8m0fnu and float4_e2m1fn_x2) to the AOTInductor C shim layer, enabling MXFP4 quantization (e.g., for AMD MI350) (#176496)tuple_return option to split_module that wraps submodule outputs in a tuple (#179007)ignore_raw_node option to GraphPickler (#176939)_merge_overlapping_fusions() method to FxNetSplitter which detects and merges overlapping fusion groups (#177099)float8_e8m0fnu dtype (#176270)torch.uint32 and torch.uint64 dtypes (#179434)List[List[float]]) (#178081)uniform and normal sampling on CPU to improve fp16/bf16 results (#175988)requires_grad to Optional[bool] in torch.asarray (#170897)narrow_copy derivative (#175609)grid_sample (#177487)torch.aminmax (#175215)num_splits in varlen attention to allow disabling split_kv (#176905)AutoNamingMode support in Selective Activation Checkpointing (#175348)torch.utils.checkpoint to no longer use autograd.Function for saving inputs (#174327)_int_mm unsigned int8 × signed int8 (u8s8) support on CPU (#168226)bias argument to nn normalization methods (LayerNorm, GroupNorm, RMSNorm, etc.) (#176573)MultiMarginLoss error message for inconsistent target size (#174072)enable_gqa flag to varlen_attn (#179468)eps=0 in batch_norm during eval mode (#175508)trunc_normal_ initialization (#176240)clone operator for semi-structured sparse tensors (#174991)alg_id (#178659)cpp_extension and cpp_builder to C++20 (#176659)at::Tag header-only changes and add a library.def override for tags (#181608)timeout parameter to torch.distributed.barrier() (#174974)reduce_scatter_tensor_coalesced support to ProcessGroupWrapper (#168961)batched_grad_copy option to reduce per-parameter kernel launches to 2 kernels per bucket (#176638)BucketCapacityConfig dataclass (#175217)ChildFailedError exitcode output for better debugging (#175254)dist.broadcast for FP8 tensors on GPUs older than SM90 (#175884)__torch_function__ handlers for distributed functions (#176376)split_group API for TorchComms on XPU (#178236)ncclx and gloo to FlightRecorder trace analyzer backend allowlist (#180268)Implement missing methods in ProcessGroupWrapper (#178779)_StridedSharding for full nn.Linear(DTensor) compatibility (#166483)is_pinned() support (#177235)print() HOP support (#175222)run_dtensor_rng_op compatible with compile_on_one_rank (#177447)_StridedShard through Replicate (#179059)Split(Flatten) sharding propagation (#179632)view_groups (#174629)index_select, index, index_fill, index_reduce, roll, fft, constant_pad_nd, squeeze.dims, interpolate, linalg ops, LayerNorm/RMSNorm FW/BW, foreach/fused ops, and einsum linearity (#176037, #176038, #178456, #175463, #175656, #173563, #176991, #176955, #179173, #177186, #177187, #176150, #174830)clip_grad_norm to match the documented behavior (#173641)ModuleList/ModuleDict subclasses that implement forward() (#175033)fully_shard (#173580)shard_mesh and shard_mesh_from_root handling (#174107)offset_t operators to be __host__ __device__ in SortStable.cu (#175997)avg_pool3d backward shape-check variables in CUDA (#178893)per_process_memory_fraction + throw_on_cudamalloc_oom (#179473) (#179473)enable_annotations kwarg to torch.cuda.graph` (#179867)ReduceLogicKernel (#176132)abs complex overflow/underflow on MPS (#174346)index_fill_ to native Metal (#175822)histogram to float/bfloat types on MPS (#176913)unfold_backward to torch.complex64 on MPS (#177274)scatter, gather, repeat, cumsum, logcumsumexp, cumprod, and nn.functional.linear on MPS (#177794, #178198, #178328, #178411, #178436, #178799)lerp, eye, relu, silu, fill_, xlogy, norm to native Metal kernels (#177093, #178683, #178866, #179071, #176101, #177749, #177328)DeviceCapability for MPS backend (#178180)enable_gqa parameter to SDPA MPS meta registration (#181550)addmv, addmm, and baddbmm on XPU (#174590)addcdiv lowering for XPU (#176163)bmm_outer_product Triton override for XPU (#180816)IntelGPUError in Inductor (#169167)enum.Enum iteration, nn.Module.__getattribute__, _enter_autocast/_exit_autocast and other context managers, next() on itertools.count, itertools.takewhile, bool(OrderedDict), NamedTuple.__eq__(tuple), numpy ndarray.flat, and locals()/vars() (#175176, #175527, #173877, #176521, #178818, #177876, #175394, #176729, #175787, #179595)nb_index/nb_bool/nb_float slots so Dynamo can trace operator.index(tensor), bool(...), and float(...); graph-break on torch.Generator methods (#178921, #178931, #179114, #180198, #178519)cond supports aliases and mutations under no_grad, autogradable leaf modules support pytree outputs, nonstrict_trace accepts nn.Module inputs, and invoke_subgraph supports subgraph reuse (#172836, #172152, #175010, #172372, #176644)torch.cuda.stream, sync barriers via a dependency HOP, triton.set_allocator inside torch.compile, and reuse of tracked objects for Triton prune_configs_by (#177610, #168894, #177470, #177874)OUT_DTYPE, ACC_TYPE, and INDEX_DTYPE codegen flow in Triton templates (#179453)addcdiv lowering for CUDA parity with eager and matching _foreach_addcdiv to _foreach_addcmul (#174912, #175309, #175310, #175839, #176237)lerp decompositions for bitwise parity with eager (#176804)torch.cat and avoided duplicate computation in cat/pad when inputs have multiple consumers (#175729)ExternKernelOut for output buffer reuse, and added symm_mem planning for graph inputs and fallback regions (#174856, #175449)pad_mm AutoHeuristics in deterministic mode (#176186, #179826)NotImplementedError when return_aux=AuxRequest(max_scores=True) is requested with BACKEND='FLASH' instead of failing later with an opaque error (#177434)allow_tf32 to fp32_precision to avoid divergence with the new TF32 API (#176098)prims.scalar_tensor and aten.arange.start_step (#179017, #179028)convert_element_type lowering to emulate PyTorch eager numerics (#176781)kpack Triton compile options on ROCm (#173179)aten.index_add (#179486)tile_k from nvMatmulHeuristics matching (#176845)aten._grouped_mm to AOTInductor fallback ops, enabling cpp-wrapper mode for grouped_mm (#177307)AOTIPythonKernelHolder, allowing a single compiled kernel to serve multiple input shapes (#176018)native_layer_norm, aminmax) (#176019)Optional[List[T]] arguments in cpp wrapper (#174460)_scaled_dot_product_attention_math_for_mps enable_gqa (#181549)get_source_partitioner to parse nn_module_stack metadata for improved source-based graph partitioning (#175788)split_module now uses _make_graph_module to support lazy recompile (#177907)fuser_utils.topo_sort to produce a stable ordering (#175378)DynamicInt __pow__ and __rpow__ methods (#179868)scaled_mm_v2 CPU implementation (#176266)PYTORCH_RELEASES_CODE_CC dict (#182369)torch.isclose broadcast failure with equal_nan=True (#175244)torch.trace backward for non-square matrices (#175068)layer_norm computes 3rd order derivatives (#176234)_wrap_sync_node to replace deps in output node's nested args (#178471)setup_context cache (#179475)addmv backward pass failure (#165777)linalg.det backward for 0-dimensional inputs (#177498)cholesky(upper=True) on macOS for matrices larger than block size (#179154)trunc_normal_ low precision issue when used with half-precision dtypes (#174997)nll_loss meta function to prevent invalid input types (#175151)Conv3d.reset_parameters for channels_last format (#175990)MSELoss failing to compute gradients when inputs have different dtypes (#175743)GroupNorm backward correctness bug on AMD wavefront-64 GPUs (#178872)nn.functional.pad compile crash with deterministic mode and replication padding (#177166)torch.bmm(COO, Dense) memory misalignment on CUDA (#175347)TORCH_BUILD_VERSION not updating when version.txt changes (#176167)_CoalescingManager not passing Opts to allgather_into_tensor_coalesced() (#175379)_CoalescingManager to raise exception when ops in the coalesced list are not the same type (#175573)getenv/setenv race condition causing segfault during NCCL initialization with heartbeat thread (#167523)AsyncCollectiveTensor inputs leaking into compiled regions, causing RuntimeError or silent data corruption in TP + compile workflows (#179849)USE_RCOM typo to USE_ROCM in intra_node_comm.cpp (#175078)NCCLPeerAllocInfo destructor to properly deregister windows and free resources (#177459)Metadata.storage_meta regression from dataclasses.replace() (#178001)FrameSummary._code on Python 3.13+ (#177754)__torch_dispatch__ bypass (#177878)_StridedShard not in safe globals for checkpoint loading (#178560)stack dim normalization (#174640)view_as_complex with P(max)/P(min) placements (#173935)get_mesh_from_args when first arg is not a tensor (#169265)tp_conv rejecting batch-dim-only sharding for valid configs (#176448)compute_local_stride for unevenly-sharded tensors (#177174)scaled_mm sharding strategy (#177234)propagate_shape_and_sharding (#177973)index_put sharding strategy for None indices (#179217)NestedRedistribute backward dtype handling (#179495)topk, sort, min, etc.) (#178668)InputDim.__eq__ type guard to prevent int comparison bugs (#178599)DeviceMesh.__getitem__ by disabling proxy tensor handling (#176007)_unshard() passing a CUDA stream where an event was expected (#170525)sync_module_states broadcast order for buffers with meta-device initialization (#178569)BlockMask as an argument (#179215)split_with_sizes_copy() missing dim argument (#169173)async_op=True profiler trace issues (#182100)stage_backward_weight with multi-output intermediate in pipeline parallelism (#175705)transpose_mxn specialization for AArch64 SVE, providing a deterministic vectorized BF16 transpose implementation independent of fixed vector widths (#174097)test/inductor/test_fp8.py hang on sm89 (#177573)AdaptiveMaxPooling2d.cu (#179261)ComplexTransform const kTransformB in fpA_intB_gemm.h (#179271)LayoutB in fpA_intB_gemm.h (#179269)AvgPool for channels_last + offset inputs (#175235)linalg_solve to return pivots (#175284)lu_solve for broadcasted bias (#175332)addmm/mm to return zero-filled matrix when an input is empty (#175905)index_reduce atomic misalignment for sub-32-bit types (#176009)masked_fill for non-contiguous outputs (#176171)layer_norm with noncontiguous bias (#176238)solve_triangular for noncontiguous inputs (#176335)histogram/histogramdd with noncontiguous weight (#175906)getStridedMPSNDArray (#176648)bmm on MPS (#176771)torch.cdouble tensor on MPS (#176985)_copy_from_and_resize logic (#177606)ops.masked variable name collisions in Metal codegen (#178304)self.add_(other, alpha) RuntimeErrors with type promotion (#178724)BatchNorm with mixed input/weight dtypes (#178775)getMPSScalar construction for uint64 (#179230)mm with stride-0 inputs on macOS < 26.4 (#180236)masked_scatter side-effect and aligned behavior with CPU (#175622)lgamma/digamma/polygamma noncontiguous behavior (#175603)masked_scatter to preserve scalar tensor shape (#174381)channels_last tensor handling (#181107)B > 1 (#181886)_get_amdsmi_device_index to return devices in correct order (#178398)torch.compile graph break inside torch.autocast('xpu') causing dtype mismatch (#180309)conv2d incorrect results and alignment errors for non-64-byte-aligned tensors on XPU (#177956)nn.Embedding module failures on XPU (#178987)_scaled_dot_product_fused_attention_overrideable to preserve query layout (#178986)DeviceOpOverrides registered incorrectly on XPU (#178959)SyclExtension Windows build for oneAPI 2025.3+ breaking change (#170701)__defaults__, guard tensor-method fallthrough against unknown methods, and closure-hash in CodeId so factory functions don't reuse stale graphs (#177103, #177191, #178420, #177737, #173512)@property setters bypassed by torch.compile, AttributeError swallowed by try/except on tensor attributes, torch.Size dict lookups with tensor-backed keys, graph break on enum members with class values, detach_ autograd metadata, allow_in_graph crash inside compiled functions, preserve original exception in GuardOnDataDependentSymNode, and einops 0.6.1 backwards patch (#176624, #175611, #177313, #177439, #177875, #178524, #176016, #177165)contextlib.contextmanager init, and parent-stack corruption in step_graph_break (#176906, #177090, #177195, #177408)torch.compiler.is_exporting() returning True during torch.compile, activation-checkpoint metadata loss through custom autograd.Function, mixed-dtype bmm/matmul, vjp_fn under torch.compile, _extract_distributed_info crash on FX-Node group_name, nested_compile_region cache keyed on fn.__code__, and reverting allow_in_graph deprecation warn-spam (#176499, #177396, #177696, #173883, #178108, #179148, #178340)cuda_stream pointer extraction for generic torch.Stream (#181019)fullgraph=True fallback to eager (#181940)no_grad views of differentiable intermediates (#175673)floordiv Inductor lowering for mixed signedness (Triton workaround) (#175168)Sm100CollectiveEpilogue on SM100 (#175305)aten.resize on overlapping-stride views (#176651)ConstructorMoverPass replacing CPU placeholder in graph output and creating mixed-device pointwise ops (#176164, #177646)VecMask::from for scalar masks in CPU codegen (#178148)triton_main_loop_scaled_mm template to use correct scale recipe (#178005)cpp_wrapper lazy compile stale state across fresh_cache resets (#178162)grid_y overflow for large batch dims and i32 overflow in template kernel signature for large storage offsets (#178617, #179333)remove_no_ops incorrectly eliminating ops on mutated values (#174938)name is in buffer_read_counts before access (#171245)SyntaxError when Triton kernel has docstrings (#176796)fallback_random dropout stride mismatch (#177077)AssertionError in ForeachKernelSchedulerNode loop reordering after fusion (#176849)floordiv (#177926)persistent_mm_template selection (#178178)max_autotune must include pointwise configs with max_autotune_pointwise (#177995)num_warps when max_autotune is enabled on HIP (#178023)nn.Dropout accuracy discrepancies between Triton and torch implementations (#178843)floor_divide with zero divisor (#178016)randn_like inconsistency between eager and compile with fallback_random=True (#177994)MetalScheduling constructor in MPSInductor (#179646)argmax/argmin returning incorrect indices for boolean tensors on CUDA (#174076)sign to match PyTorch NaN semantics (#176579, #176814)UnicodeDecodeError in Triton depthwise conv template (#176484)torch._check divisibility propagation to Triton tt.divisibility (#175755)block_ptr store dtype for inplace-mutated buffers (#177860)constexpr (#172354)'bool' object not callable) (#176090)_split_iteration_ranges silently dropping dimensions (#177673)sym_sum lowering to accept varargs (#178661)bias_addmm for AMD (#178929)isinf() to Float8_e4m3fn to fix nan_asserts crash with fp8 inputs (#160641)Grid2DWithYZOverflow (#178878)torch.cond (#179457)torch.compile performance regression for cumprod backward (#170388)lazy_triton_compile.h in the XPU cpp_wrapper header (#180815)run_single_threaded (#174998)int64_t type declaration for kernel numel variables (#176922)get_source_partitioner (#175935)make_fx handling of value typesNote truncated.
Fixed a ZipSlip directory traversal vulnerability in torch.hub that could allow malicious zip files to extract files outside the target directory. tor…
<table> <tr> <td> Added Support for <strong>Differentiable Collectives</strong> for Distributed Training </td> </tr> <tr> <td> <strong>FlexAttention</strong> now has a <strong>FlashAttention-4</strong> backend on <strong>Hopper</strong> and <strong>Blackwell</strong> GPUs </td> </tr> <tr> <td> <strong>MPS (Apple Silicon)</strong> Comprehensive Operator Expansion </td> </tr> <tr> <td> Added <strong>RNN/LSTM</strong> GPU Export Support </td> </tr> <tr> <td> Added <strong>XPU Graph</strong> Support </td> </tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
Starting with PyTorch 2.11, the CUDA 12.8 and 12.9 pre-built binaries no longer include support for Volta GPUs (compute capability 7.0, e.g. V100). This change was necessary to enable updating to CuDNN 9.15.1, which is incompatible with Volta.
Users with Volta GPUs who need CUDA 12.8+ should use the CUDA 12.6 builds, which continue to include Volta support. Alternatively, build PyTorch from source with Volta included in TORCH_CUDA_ARCH_LIST.
Version 2.10:
# CUDA 12.8 builds supported Volta (SM 7.0)
pip install torch --index-url https://download.pytorch.org/whl/cu128
# Works on V100
Version 2.11:
# CUDA 12.8 builds no longer support Volta
# For V100 users, use CUDA 12.6 builds instead:
pip install torch --index-url https://download.pytorch.org/whl/cu126
Starting with PyTorch 2.11, pip install torch on PyPI installs CUDA 13.0 wheels by default for both Linux x86_64 and Linux aarch64. Previously, PyPI wheels shipped with CUDA 12.x and only Linux x86_64 CUDA wheels were available on PyPI. Users whose systems have only CUDA 12.x drivers installed may encounter errors when running pip install torch without specifying an index URL.
Additionally, CUDA 13.0 only supports Turing (SM 7.5) and newer GPU architectures on Linux x86_64. Maxwell and Pascal GPUs are no longer supported under CUDA 13.0. Users with these older GPUs should use the CUDA 12.6 builds instead.
CUDA 12.6 and 12.8 binaries remain available via download.pytorch.org.
Version 2.10:
# PyPI wheel used CUDA 12.x
pip install torch
Version 2.11:
# PyPI wheel now uses CUDA 13.0
pip install torch
# To get CUDA 12.8 wheels instead:
pip install torch --index-url https://download.pytorch.org/whl/cu128
# To get CUDA 12.6 wheels (includes Maxwell/Pascal/Volta support):
pip install torch --index-url https://download.pytorch.org/whl/cu126
torch.hub.list(), torch.hub.load(), and torch.hub.help() now default the trust_repo parameter to "check" instead of None. The trust_repo=None option has been removed. (#174101)Previously, passing trust_repo=None (or relying on the default) would silently download and run code from untrusted repositories with only a warning. Now, the default "check" behavior will prompt the user for explicit confirmation before running code from repositories not on the trusted list.
Users who were explicitly passing trust_repo=None must update their code. Users who were already passing trust_repo=True, trust_repo=False, or trust_repo="check" are not affected.
Version 2.10:
# Default trust_repo=None — downloads with a warning
torch.hub.load("user/repo", "model")
# Explicit None — same behavior
torch.hub.load("user/repo", "model", trust_repo=None)
Version 2.11:
# Default trust_repo="check" — prompts for confirmation if repo is not trusted
torch.hub.load("user/repo", "model")
# To skip the prompt, explicitly trust the repo
torch.hub.load("user/repo", "model", trust_repo=True)
varlen_attn via window_size, making optional arguments keyword-only (#172238)The signature of torch.nn.attention.varlen_attn has changed: a * (keyword-only separator) has been inserted before the optional arguments. Previously, optional arguments like is_causal, return_aux, and scale could be passed positionally; they must now be passed as keyword arguments. A new window_size keyword argument has also been added.
# Before (2.10)
output = varlen_attn(query, key, value, cu_seq_q, cu_seq_k, max_q, max_k, True, None, 1.0)
# After (2.11) — pass as keyword argument
output = varlen_attn(query, key, value, cu_seq_q, cu_seq_k, max_q, max_k, window_size=(-1, 0), return_aux=None, scale=1.0)
is_causal flag from varlen_attn (#172245)The is_causal parameter has been removed from torch.nn.attention.varlen_attn. Causal attention is now expressed through the window_size parameter: use window_size=(-1, 0) for causal masking, or window_size=(W, 0) for causal attention with a sliding window of size W. The default window_size=(-1, -1) corresponds to full (non-causal) attention.
# Before (2.10)
output = varlen_attn(query, key, value, cu_seq_q, cu_seq_k, max_q, max_k, is_causal=True)
# After (2.11) — use window_size instead
output = varlen_attn(query, key, value, cu_seq_q, cu_seq_k, max_q, max_k, window_size=(-1, 0))
DebugInfoWriter now honors $XDG_CACHE_HOME for its cache directory in C++ code, consistent with the Python side. Previously it always used ~/.cache/torch. (#168232)This avoids issues where $HOME is not set or not writable. Users who relied on ~/.cache/torch being used regardless of $XDG_CACHE_HOME may see debug info written to a different location.
Version 2.10:
# C++ DebugInfoWriter always wrote to ~/.cache/torch
Version 2.11:
# C++ DebugInfoWriter now respects $XDG_CACHE_HOME/torch (same as Python code)
# Falls back to ~/.cache/torch if $XDG_CACHE_HOME is not set
DeviceMesh now stores a process group registry (_pg_registry) directly, enabling torch.compile to trace through get_group(). (#172272)This may break code that skips init_process_group, loads a saved DTensor (constructing a DeviceMesh with no PGs), and later creates PGs separately — during torch.compile runtime the PG lookup will fail. Users should ensure process groups are initialized before constructing the DeviceMesh.
Version 2.10:
# PGs resolved via global _resolve_process_group at runtime
mesh = DeviceMesh(...) # PGs could be created later
Version 2.11:
# PGs now stored on DeviceMesh._pg_registry; must exist at mesh creation
dist.init_process_group(...) # Must be called before creating mesh
mesh = DeviceMesh(...)
DTensor.to_local() backward now converts Partial placements to Replicate by default when grad_placements is not provided. (#173454)Previously, calling to_local() on a Partial DTensor would preserve the Partial placement in the backward gradient, which could produce incorrect gradients when combined with from_local(). Now, the backward pass automatically maps Partial forward placements to Replicate gradient placements, matching the behavior of from_local().
Users who relied on the previous behavior (where to_local() backward preserved Partial gradients) may see different gradient values. To ensure correctness, explicitly pass grad_placements to to_local().
Version 2.10:
# Partial placement preserved in backward — could produce incorrect gradients
local_tensor = partial_dtensor.to_local()
Version 2.11:
# Partial → Replicate in backward by default (correct behavior)
local_tensor = partial_dtensor.to_local()
# Or explicitly specify grad_placements for full control:
local_tensor = partial_dtensor.to_local(grad_placements=[Replicate()])
_PhiloxState.seed and _PhiloxState.offset now return torch.Tensor instead of int (#173876)The DTensor RNG internal _PhiloxState class changed its seed and offset properties to return tensors instead of Python ints, and the setters now expect tensors. This makes the RNG state compatible with PT2 tracing (the previous .item() calls were not fake-tensor friendly).
Code that directly reads _PhiloxState.seed or _PhiloxState.offset and treats them as ints will break. Call .item() to get the int value. When setting, wrap the value in a tensor.
Version 2.10:
from torch.distributed.tensor._random import _PhiloxState
philox = _PhiloxState(state)
seed: int = philox.seed # returned int
philox.offset = 42 # accepted int
Version 2.11:
from torch.distributed.tensor._random import _PhiloxState
philox = _PhiloxState(state)
seed: int = philox.seed.item() # now returns Tensor; call .item() for int
philox.offset = torch.tensor([42], dtype=torch.int64) # must pass Tensor
When caffe2 and PyTorch were separate projects, the ROCm support strategies were different. For caffe2, all files and classes would be renamed following the pattern of CUDA to HIP, Cuda to Hip, cuda to hip, and so on. PyTorch did not rename classes, but would create new files following the same renaming pattern (e.g., aten/src/ATen/cuda/CUDABlas.h to aten/src/ATen/hip/HIPBlas.h). As a consequence, caffe2 had a distinct device backend named "HIP" (renamed from "CUDA") while ROCm PyTorch masquerades as the "cuda" device (torch.empty(1, device="cuda")). Once caffe2 and PyTorch projects were merged, this caused a mismatch between caffe2 expecting to use a "HIP" device while PyTorch expecting a "cuda" device. To alleviate this mismatch, "Masquerading" classes were created under aten/src/ATen/hip/impl.
Hipify v2 (#174087, #174300, #174388, #174499, #175098) makes the following changes:
torch.export.export_for_training has been removed (#171714)export_for_training was previously available as a separate API for exporting models while preserving training semantics. This function has been removed. Users should use torch.export.export instead, which returns the same graph as the previous export_for_training.
fallback option from torch.onnx.export (#173189)The fallback parameter has been removed from torch.onnx.export(). Previously, when fallback=True, the exporter would automatically fall back to the legacy TorchScript-based exporter if the dynamo exporter failed. This fallback was removed because it was overly complicated, required different inputs, produced different models, and hid errors from the new exporter.
Migration: Remove fallback=True (or fallback=False) from your torch.onnx.export() calls. If you need fallback behavior, implement it explicitly in your own code by catching exceptions and calling the legacy exporter separately.
# Before
torch.onnx.export(model, args, "model.onnx", dynamo=True, fallback=True)
# After
torch.onnx.export(model, args, "model.onnx", dynamo=True)
The custom_translation_table parameter in torch.onnx.export() no longer accepts a list of functions for each torch op. Previously, users could pass a list of overloaded ONNX functions (e.g., one for float tensors, another for bool tensors), and the dispatcher would automatically select the correct overload based on input types. This complex type-matching logic has been removed because torchlib no longer uses overloads for the same opset version.
The type of custom_translation_table changed from dict[Callable, Callable | Sequence[Callable]] to dict[Callable, Callable]. Passing a Sequence as a value now raises a TypeError.
Migration: Provide a single function per operator instead of a list of overloads. If you need type-dependent behavior, handle it inside the single function.
# Before
custom_translation_table = {
torch.ops.aten.logical_and.default: [custom_impl_float, custom_impl_bool],
}
# After
custom_translation_table = {
torch.ops.aten.logical_and.default: custom_impl,
}
torch.ao.quantization.pt2e and torch.ao.quantization.quantizer) has been removed from PyTorch and migrated to torchao. (#169151)The following modules and classes have been removed:
torch.ao.quantization.pt2e (including DuplicateDQPass, PortNodeMetaForQDQ, export utils, graph utils, numeric debugger, lowering utilities)torch.ao.quantization.quantizer (including ComposableQuantizer, EmbeddingQuantizer, X86InductorQuantizer, XPUInductorQuantizer, XNNPACKQuantizer, QuantizationSpec, QuantizationAnnotation, QuantizationConfig, etc.)Users relying on the PT2E quantization flow should migrate to the torchao package, which now hosts these APIs.
Version 2.10:
from torch.ao.quantization.pt2e import prepare_pt2e, convert_pt2e
from torch.ao.quantization.quantizer.x86_inductor_quantizer import X86InductorQuantizer
Version 2.11:
# Install torchao: pip install torchao
from torchao.quantization.pt2e import prepare_pt2e, convert_pt2e
from torchao.quantization.pt2e.quantizer.x86_inductor_quantizer import X86InductorQuantizer
The MAGMA backend for linear algebra operations is now deprecated and will be removed in a future release. Setting torch.backends.cuda.preferred_linalg_library("magma") or retrieving a previously-set MAGMA preference will now issue a deprecation warning. cuSOLVER remains the default backend. (#172823)
If you see any errors when using cuSOLVER that did not occur with MAGMA, please file an issue on GitHub. To silence the warning, stop explicitly selecting the MAGMA backend:
Version 2.10:
# No warning
torch.backends.cuda.preferred_linalg_library("magma")
Version 2.11:
# Issues a deprecation warning — remove this call to use the default cuSOLVER backend
torch.backends.cuda.preferred_linalg_library("magma")
torch.linalg.svd no longer dispatches to MAGMA. The MAGMA backend is deprecated and cuSOLVER is now used unconditionally, providing significant speedups (2x–400x depending on matrix size and batch dimensions). (#172824)
Previously, setting torch.backends.cuda.preferred_linalg_library("magma") would route SVD through MAGMA. This setting is now ignored for SVD, and cuSOLVER is always used.
Version 2.10:
torch.backends.cuda.preferred_linalg_library("magma")
U, S, Vh = torch.linalg.svd(x) # Uses MAGMA
Version 2.11:
# MAGMA preference is ignored; cuSOLVER is always used
U, S, Vh = torch.linalg.svd(x) # Uses cuSOLVER
torch.linalg.solve_triangular and torch.triangular_solve no longer dispatch to MAGMA on CUDA. cuBLAS is now used unconditionally, providing speedups of 2x–24x for most matrix sizes (small matrices may see minor regressions of ~0.6x). (#174109)
Version 2.10:
torch.backends.cuda.preferred_linalg_library("magma")
torch.linalg.solve_triangular(A, B, upper=False) # Uses MAGMA
Version 2.11:
# MAGMA preference is ignored; cuBLAS is always used
torch.linalg.solve_triangular(A, B, upper=False) # Uses cuBLAS
torch.linalg.lstsq no longer dispatches to MAGMA. cuSOLVER/cuBLAS are now used unconditionally, providing speedups of 1.7x–620x depending on matrix size and batch dimensions. (#174779)
Version 2.10:
torch.backends.cuda.preferred_linalg_library("magma")
result = torch.linalg.lstsq(A, B) # Uses MAGMA
Version 2.11:
# MAGMA preference is ignored; cuSOLVER/cuBLAS is always used
result = torch.linalg.lstsq(A, B) # Uses cuSOLVER/cuBLAS
torch.distributed.symmetric_memory.enable_symm_mem_for_group is deprecated. The store can be retrieved directly via ProcessGroup.getStore() in C++, making this call unnecessary. (#172163)Version 2.10:
from torch.distributed.symmetric_memory import enable_symm_mem_for_group
enable_symm_mem_for_group(group)
Version 2.11:
# No longer needed — store is accessed directly from the ProcessGroup
Added native_handle property to torch.Stream, providing a unified way to retrieve the backend-specific opaque stream handle (e.g., cudaStream_t for CUDA, sycl::queue* for XPU). This is useful for passing stream handles to third-party libraries such as Triton. (#171040)
stream = torch.accelerator.current_stream()
handle = stream.native_handle # backend-specific stream handle
Function.clear_saved_tensors_on_access class attribute to automatically free saved tensors after they are accessed (#173833)activate_flash_attention_impl (#169866)scale for softmax to varlen attn (#171199)start_method option to torch.distributed.debug.start_debug_server to select the multiprocessing start method (fork, spawn, or forkserver), enabling CUDA-safe server startup (#173196)torch.distributed.debug (#174808)torch.distributed.all_gather) now automatically work with FakeTensorMode — meta implementations are registered at import torch time (#162119)SymmetricMemory as a torch class for use in op definitions (#174019)torchcomms _BackendWrapper shim layer in c10d (#174202)import torch
x = torch.rand(10, 1, 10, device='mps')
y = x[:, [1]]
torch.mps.synchronize() # will raise index out of bounds error
clock_rate, memory_clock_rate, memory_bus_width, memory_per_block, shared_memory_per_block. (#170572)TORCH_USE_HIP_DSA. (#172679)torch.compile now supports tracing through contextlib.ExitStack and contextlib.suppress context managers, allowing code that uses these patterns to be compiled without graph breaks (#146506, #147990)torch._dynamo.config.ignore_logging_functions config to skip arbitrary logging callables during tracing without causing graph breaks. Add functions to this set to have Dynamo treat them as no-ops during compilation (#168913)TORCH_DYNAMO_AUTOMATIC_DYNAMIC_SHAPES=0 environment variable to globally disable automatic dynamic shapes without modifying Python code (#172334)TORCH_COMPILE_OVERRIDE_BACKENDS environment variable for per-graph backend override, enabling binary search to find problematic compiled graphs. Supports filter syntax like ">10:eager" or "0-5:aot_eager;6-10:inductor" (#172411)torch._dynamo.decorators.leaf_function, which allows annotating functions as leaf operations that Dynamo and AOTAutograd will not trace into (#170471)register_hook on non-leaf tensors would fail under torch.compile (#172126)torch.cond dispatch generation (#167617)ldexp lowering with libdevice.ldexp (CUDA) and std::ldexp (CPU) codegen (#171721)pin_memory for torch.empty (#172578)triton_meta to TritonTemplate maybe_append_choice API for custom template development (#174292)max-autotune-gemm, which overlaps autotuning with lowering/scheduling in a subprocess to reduce compilation overhead (#170407)(BlockWise128x128, BlockWise1x128) scaling support in Inductor Triton templates (#170748)torch.export (#174720)ExportableModule wrapper for ONNX export (#170810)InputObserver to infer dynamic shapes for export (#172838)InputObserver for multimodal LLM export (#174964)torch.linalg._powsum and torch._foreach_powsum as fused kernels that compute sum(abs(x)**ord) (equivalent to vector_norm without the root extraction) (#172685)torch.load now produces clearer error messages when encountering miniz errors from PyTorchStreamReader, explicitly indicating that the checkpoint file is likely corrupt (#170244)torch.load(map_location='meta') no longer reads storage data from the filesystem, improving performance when loading checkpoints onto the meta device (#170619)check_out_variant and to_out_variant utilities for custom operator out variant validation. check_out_variant verifies that a custom op's out variant is compatible with Inductor's out_variant pass, and to_out_variant converts an OpOverload to its out variant. (#174473)remove_duplicate parameter to nn.Module.modules() function (#174383)nn.attention.flex_attention (#171744)Float8_e8m0fnu and Float4_e2m1fn_x2 dtypes to stable ABI (#173669)torch::stable::Tensor::layout() (#174735)context_parallel_shard more general (#170200)get_offset for symmetric memory (#172044)ProcessGroupNCCL: workaround for reduce_scatter with world_size=1 (#170922)ProcessGroupWrapper (#171920)pdb only when user calls breakpoint() in torch.distributed (#171818)ProcessGroupNCCL: use lowest rank as split color (#173687)accscalar_t for interpolation accumulators (#170661)index_fill backward pass (#174238)baddbmm and addbmm to integer and complex types (#170895)torch.special.erfcx (scaled complementary error function) (#172910)addmm behavior now takes into account preferred BLAS backend instead of forcing hipblaslt. (#174350)torch.view_as_real and torch.view_as_complex now support sparse tensors (#164964)torch.xpu._dump_snapshot API (#170186)torch.xpu._record_memory_history API (#169559)torch.xpu.memory_snapshot (#169442)local_mem_size to XPU device properties (#172314)torch.accelerator.get_device_capability on XPU (#170747)aot_inductor.emit_multi_arch_kernel on XPU (#171432)decompose_k choice for XPU (#170541)skip_actions flag to filter out specific events (#168183).post_process_timeout_s field to prevent post processing from blocking
further execution (#173957).fullgraph=True now recursively disables dynamo on compiled code to prevent unintentional re-invocation of torch.compile (#173080)Enum.__contains__ and constants (#173223)kwargs=True (#172519)object type in dynamo tracing (#171457)allow_in_graph (#173611)LoweringException occurs, making debugging easier (#171846)120a for .ptx to .fatbin compilation and cpp codegen (#174162, #172263)cvt_e8m0_rceil prim with PTX lowering for SM100+ GPUs (#172497)launch_cooperative_grid flag for cooperative reduction kernels (#167800)torch.float8_e5m2, enabling mixed FP8 (e4m3fn x e5m2) matrix multiplication (#171167)scaled_dot_product_flash_attention.low_p overload (#172622)record_function with _RecordFunctionFast in CompiledFxGraph for reduced profiling overhead (#163976)mutated_inputs, allowing more flexible template usage (#170721)combo_kernels_pointwise_only config option to exclude reduction kernels from combo kernel fusion (#174894)torch.fx.symbolic_trace now supports tracing HigherOrderOperators that do not take callable arguments (#173839)hint_int to size_hint, support size_hint in user code. (#171944)_disable_torch_fn_metadata_mode option to make_fx and aot_export_joint_with_descriptors (#172087)from_node provenance information is now preserved when serializing exported programs (#171726)expm1 for computing quantized ELU, improving numerical stability (#173968)spin lint command now supports pass-through arguments to lintrunner, including --take, --skip, and --tee-json flags, giving developers more control over which linters run (#169373)cpp_kernel_name to public API to match AOTI shim gen; add mm_type_out to AOTI fallback kernel (#174489)torch.load with FakeTensorMode or skip_data context would compute incorrect storage sizes (#170618)torch.ops.aten.index.Tensor to properly raise an IndexError when called with an empty indices list, instead of producing undefined behavior (#174009)torch.autograd.gradcheck when fast_mode=True (#166386)torch.view_as_complex() not working on the memory layout produced by .contiguous() after .transpose() (#169780)torch.bucketize crash during torch.export when test_elements is a scalar (#170751)MaxUnpool crash when input tensors are small (#169359)__getitem__ in Subset subclasses (#163961)max_q and max_k for varlen_attn (#173681)SymInt, MemoryFormat, ScalarType, Layout) correctly (#174734)std::string::find method in c10d (#170057)_set_pg_timeout not working for Gloo backend (#167052)flex_input_fn argument unwrapping issue (#170201)_unshard() passing Stream instead of Event (#170525)ProcessGroupGloo CUDA tensor stream handling with futures (#170812)split_with_sizes_copy() missing dim argument (#169173)wait_tensor (#171614)fully_shard arg typehint inconsistency (#171574)clip_grad_norm to align with documentation (FSDP) (#173641)ProcessGroupWrapper missing method forwarding (#173599)Partial(max/min) reduce op type on torch.max/torch.min output DTensors (#170203)Partial DTensors with different reduce ops (#170209)ctc_loss_gpu_template on SM12+ (#172447)tanh implementation (#172406)orgqr race condition on MPS (#174143)torch.sparse.spdiags crashing with zero-dimension shapes (#174052)GradTrackingTensor.tolist() not working on MPS (macOS) device tensors when used inside torch.func.grad or other function transforms (#171317)torch.xpu.memory_allocated / torch.xpu.memory_reserved reporting incorrect memory sizes (#171453)FakeTensorMode after compile (#171209), fixed CUDA memory usage for CPU-only compile (#163841)OrderedSet, set, and frozenset with activation checkpointing (#169535, #170291)MATCH_MAPPING, MATCH_KEYS, and MATCH_SEQUENCE opcodes for Python pattern matching (match/case) support (#173085, #173086, #173087)share_memory_ compile failure (#171162)defaultdict default factory and union functionality (#168028)MutableMapping subclasses (#173184)is_inference in the cudagraphs torch.compile backend (#174713)persistent_reduction heuristics, FusedMixOrderReductions grouping, and scalar store broadcast shape mismatch (#168939, #169509, #169721, #172658)torch.cond stride mismatch when subgraph outputs are padded (#169963)view_copy decomposition to handle dtype size changes (#171442)aten.isin with scalar inputs (#171272)aten.uniform (#170794)upsample_nearest (#171151)allow_tf32 usage from inductor internals (#173731)num_stages (#170071)select_scatter dtype assertion error (#171311)aten.pow lowering (#170960)silu output match eager mode in inductor (#171723)bucketize crash in compile mode (#171595)OverflowError when truncating infinity numbers (#166636)narrow_copy decomposition (#173043)remove_noop_ops (#172160)next_power_of_2 (#174330)_pdist_forward and _pdist_backward (#170959)libdevice.pow type mismatch in Triton codegen (#173685)exclusive_scan_decoupled_lookback_64 (#171153)RuntimeError in FxGraphCachePickler.dumps() for unpickleable pybind11 objects (#173577)var_ranges (#173607)n_spills None before threshold comparison in Triton heuristics (#169940)identify_mutated_tensors() (#170808)should_use_persistent_reduction (#172534)use_compute_types parameter name mismatch in CPP backend to_dtype() that caused TypeError when enabling emulate_precision_casts=True on CPU (#168125)torch.segment_reduce (#172981)torch.isin compile shape for scalar test_elements (#172531)torch.fx.symbolic_trace to_folder with torch.nn.Sequential modules (#169279)Node.type pickling in torch.fx (#169172)torch.export.save and torch.export.load (#171954)torch.export (#172805)MaxPool3d that could produce negative values (#171790)out_channels in quantized ConvTranspose modules (#171628)_foreach_copy_ producing incorrect results when destination tensors have mixed dtypes (#173531)_foreach_max returning incorrect results when tensors contain negative values, by initializing output with lowest() instead of zeros (#173241)_foreach_norm now properly raises an error when computing infinity norm on empty tensors, instead of returning undefined results (#173238)_foreach_norm to support symbolic tensors with dynamic shapes (#174026)__slots__ to pytree TreeSpec dataclasses, reducing memory usage and improving attribute access speed (#172172)_get_param_to_fqns from O(N^2) to O(N) in FSDP (#174675)grouped_mm a bit (#170802)atan2 to native MPS Metal kernel (#173405)pow_tensor_scalar and reciprocal to Metal shaders (#170077)log_normal and geometric distributions as Metal shaders (#174189)grid_sampler_2d to Metal (#174343)int_mm performance on Intel GPU when mat2 is non-contiguous (#169555)inspect.signature, var_getattr, attr source construction, and higher-order ops; fast paths for bind_args, GET_ITER on tuples, and tree_map on namedtuples; lazy variable tracker optimizations (#170100, #169959, #173582, #174437, #174438, #174141, #174020, #174130, #174901, #174598)dynamo_timed calls, used _has_symbolic_sizes_strides more pervasively, overlapped template fusion with async_compile, and optimized Triton template heuristics (#169370, #169661, #169662, #170444, #170565, #172249, #174408)avg_pool2d as a reduction for improved performance (#167228)_foreach_addcmul and _foreach_addcdiv when tensor1 is a 0D (scalar) tensor (#172731)torch.unique behavior when using the dim parameter (#171608)torch.as_tensor signature to document that dtype and device are keyword-only arguments (#173073)torch.tensordot documentation (#173893)torch.Tensor.permute to clarify it accepts variadic arguments unlike torch.permute (#170689)torch.autograd.set_multithreading_enabled docs (#170204)torch::stable namespace (#170912)torch.sparse.mm and torch.sparse.addmm (#174039)tensor_attributes.rst with additional torch.float4_e2m1fn_x2 dtype documentation (#170448)torch.hub that could allow malicious zip files to extract files outside the target directory. torch.hub now validates all extracted paths and raises a ValueError if an archive attempts to write outside the expected folder. (#171754)USE_NCCL=0 build failure in nccl_dev_cap.hpp (#171694)from_local/to_local, and backward passes (#171576)-march=armv8-a+sve+bf16 on Debian 13) by broadening the GCC version workaround for SVE compilation (#174647)at::numeric_limits (#171111)_local_scalar_dense_mps to DispatchV2 (#172967)spv to zebin (#167972)[compile_id] prefix style (#170217, #170218, #174110)fx_graph_runnable improvements: fixed single quotes in envvars, missing symints, nested triton kernels, and global constexpr support (#173704, #174035, #174038, #174533)GraphModule.print_readable() improvements: new additional_meta argument for displaying additional node metadata (#173734), long annotations are now truncated for readability (#173119), and fix trailing whitespace with inner graphs (#172644)GraphPickler improvements: support for custom ignored metadata field keys (#172587), a debug_dumps method for debugging pickle failures (#173675), respecting __getstate__ for GraphModule serialization (#173810), and automatic fallback to dill if available (#173801)invoke_subgraph nodes now point to the original model code for easier debugging (#170927)fx_graph_runnable (#173932)Removed deprecated imports for torch.utils.data.datapipes.iter.grouping (#163438). from torch.utils.data.datapipes.iter.grouping import SHARDING_PRIOR…
<table> <tr> <td> <strong>Python 3.14</strong> support for <code>torch.compile()</code>. Python 3.14t (freethreaded build) is experimentally supported as well. </td> <tr> <tr> <td> Reduced kernel launch overhead with <strong>combo-kernels</strong> horizontal fusion in torchinductor </td> <tr> <tr> <td> A new <strong>varlen_attn()</strong> op providing support for ragged and packed sequences </td> <tr> <tr> <td> Efficient eigenvalue decompositions with <strong>DnXgeev</strong> </td> <tr> <tr> <td> <code>torch.compile()</code> now respects <strong>use_deterministic_mode</strong> </td> <tr> <tr> <td> <strong>DebugMode</strong> for tracking dispatched calls and debugging numerical divergence - This makes it simpler to track down subtle numerical bugs. </td> <tr> <tr> <td> <strong>Intel GPUs support:</strong> Expand PyTorch support to the latest Panther Lake on Windows and Linux by enabling FP8 (core ops and scaled matmul) and complex MatMul support, and extending SYCL support in the C++ Extension API for Windows custom ops. </td> <tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
data_source argument from Sampler (#163134). This is a no-op, unless you have a custom sampler that uses this argument. Please update your custom sampler accordingly.from torch.utils.data.datapipes.iter.grouping import SHARDING_PRIORITIES, ShardingFilterIterDataPipe is no longer supported. Please import from torch.utils.data.datapipes.iter.sharding instead.nn.attention.flex_attention (#161734)fallback=False is now the default in torch.onnx.export (#162726)dynamo=True option without fallback. This is the recommended way to use the ONNX exporter. To preserve 2.9 behavior, manually set fallback=True in the torch.onnx.export call.We decided to deprecate an existing behavior which goes against the PyTorch design principle (explicit over implicit) for device mesh slicing of flattened dim.
import torch
from torch.distributed.device_mesh import
device_type = (
acc.type
if (acc := torch.accelerator.current_accelerator(check_available=True))
else "cpu"
)
mesh_shape = (2, 2, 2)
mesh_3d = init_device_mesh(
device_type, mesh_shape, mesh_dim_names=("dp", "cp", "tp")
)
mesh_3d["dp", "cp"]._flatten()
mesh_3["dp_cp"] # This comes with no warning
import torch
from torch.distributed.device_mesh import
device_type = (
acc.type
if (acc := torch.accelerator.current_accelerator(check_available=True))
else "cpu"
)
mesh_shape = (2, 2, 2)
mesh_3d = init_device_mesh(
device_type, mesh_shape, mesh_dim_names=("dp", "cp", "tp")
)
mesh_3d["dp", "cp"]._flatten()
mesh_3["dp_cp"] # This will come with a warning because it implicitly change the state of the original mesh. We will eventually remove this behavior in future release. User should do the bookkeeping of flattened mesh explicitly.
from/to to torch::stable::detail (#164956)torch.jit is not guaranteed to work in Python 3.14. Deprecation warnings have been added to user-facing torch.jit API (#167669).torch.jit should be replaced with torch.compile or torch.export.
dynamic_axes option in torch.onnx.export is deprecated (#165769)Users should supply the dynamic_shapes argument instead. See https://docs.pytorch.org/docs/stable/export.html#expressing-dynamism for more documentation.
export_memory_timeline method (#168036)The export_memory_timeline method in torch.profiler is being deprecated in favor of the newer memory snapshot API (torch.cuda.memory._record_memory_history and torch.cuda.memory._export_memory_snapshot). This change adds the deprecated decorator from typing_extensions and updates the docstring to guide users to the recommended alternative.
torch.utils.checkpoint.checkpoint (#166536)ComplexTensor subclass (#167621)LocalTensor:
LocalTensor is a powerful debugging and simulation tool in PyTorch's distributed tensor ecosystem. It allows you to simulate distributed tensor computations across multiple SPMD (Single Program, Multiple Data) ranks on a single process. This is incredibly valuable for: 1) debugging distributed code without spinning up multiple processes; 2) understanding DTensor behavior by inspecting per-rank tensor states; 3) testing DTensor operations with uneven sharding across ranks; 4) rapid prototyping of distributed algorithms. Note that LocalTensor is designed for debugging purposes only. It has significant overhead and is not suitable for production distributed training.LocalTensor is a torch.Tensor subclass that internally holds a mapping from rank IDs to local tensor shards. When you perform a PyTorch operation on a LocalTensor, the operation is applied independently to each local shard, mimicking distributed computation (LocalTensor simulates collective operations locally without actual network communication.). LocalTensorMode is the context manager that enables LocalTensor dispatch. It intercepts PyTorch operations and routes them appropriately. The @maybe_run_for_local_tensor decorator is essential for handling rank-specific logic when implementing distributed code.LocalTensor, users import from torch.distributed._local_tensor, initialize a fake process group, and wrap their distributed code in a LocalTensorMode context. Within this context, DTensor operations automatically produce LocalTensors.c10d:
shrink_group implementation to expose ncclCommShrink API (#164518)torch.compile now fully works in Python 3.14 (#167384)skip_fwd_side_effects_in_bwd_under_checkpoint) to allow eager and compile activation-checkpointing divergence for side-effects (#165775)torch._higher_order_ops.print for enabling printing without graph breaks or reordering (#167571)Added node metadata annotation API
Disable preservation of node metadata when enable=False (#164772)
Annotation should be mapped across submod (#165202)
Annotate bw nodes before eliminate dead code (#165782)
Add logging for debugging annotation (#165797)
Override metadata on regenerated node in functional mode (#166200)
Skip copying custom meta for gradient accumulation nodes; tag with is_gradient_acc=True (#167572)
Add metadata hook for all nodes created in runtime_assert pass (#169497)
Update gm.print_readable to include Annotation (#165397)
Add annotation to assertion nodes in export (#167171)
Add debug mode to print meta in fx graphs (#165874)
torch.compiler.config.force_disable_caches as a public API. (#166699)nn.functional.scaled_mm (#164142)nn.functional.scaled_grouped_mm (#165154)nn.attention.varlen_attn (#164502, #164504)nn.functional.grouped_mm (#168298)torch.onnx.testing with a testing utility assert_onnx_program (#162495)RecordFunctionFast (#162661)Add _scaled_mm_v2 API (#164141)
Add scaled_grouped_mm_v2 and python API (#165154)
Add embedding_bag_byte_prepack_with_rowwise_min_max and embedding_bag_{2/4}bit_prepack_with_rowwise_min_max (#162924)
Add MXFP4 support for _scaled_grouped_mm_v2 via. FBGEMM kernels (#166530)
scaled_mm and scaled_mm_v2 for Intel GPU (#166056)_weight_int8pack_mm for Intel GPU (#160938)torch.xpu.get_per_process_memory_fraction for Intel GPU (#165511)torch.xpu.set_per_process_memory_fraction for Intel GPU (#165510)torch.xpu.is_tf32_supported for Intel GPU (#163141)torch.xpu.can_device_access_peer for Intel GPU (#162705)torch.accelerator.get_memory_info for Intel GPU (#162564)torch.compile(backend="aot_eager") backend, it should now give bitwise equivalent results in eager. Previously it sometimes would not due to extra compile-only decompositions running (#165910)torch._check over torch._check_is_size (#164889,TORCH_CHECK_{COND} behavior to be non-fatal (#167004)TypeTraits, TypeList, Metaprogramming, DeviceType, MemoryFormat, Layout, version.h, and CppTypeToScalarType to torch::headeronly (#167386, #163999, #168034, #165153, #164381, #167610)libfmt submodule version to 12.0.0 (#163441)torch.cuda.rng_set_state and torch.cuda.rng_get_state work in CUDA graph capture. (#162505)header_code argument (#163165)launch_logcumsumexp_cuda_kernel (#164567)per_process_memory_fraction option to PYTORCH_CUDA_ALLOC_CONF (#161035)c10d
Context Parallel
parallelize_module (#162542)flex_cp_forward custom op for FlexAttention CP (#163185)_templated_ring_attention to the backward compatility stub (#166991)_LoadBalancer classes, and load-balance interface to Context Parallel APIs with process-time based Round-Robin load-balance (#161062, #163617)DeviceMesh
_unflatten on top of CuTe layout bookkeeping (#161224, #165521)_rank for use with non-global PGs (#162439)FullyShardDataParallel (FSDP1 and FSDP2)
DTensor
_foreach_pow, logsumexp and masked_fill_.Scalar to sharding propagation list. (#162895, #163879, #169668)SymmetricMemory
multimem_one_shot_reduce_out (#164517)multi_root_tile_reduce (#162243, #164757)symm_mem_sync Triton kernel to torch.ops.symm_mem (#168917)set_signal_pad_size API for SymmetricMemory (#169156)Pipeline Parallelism
PipelineScheduleRuntime (#164777)OVERLAP_F_B in schedule (#161072)torchelastic
Turn on capture_scalar_outputs and capture_dynamic_output_shape_ops when fullgraph=True (#163121, #163123)
Improved tracing for dict key hashing (#169204)
Tracing support for torch.cuda.stream (#166472)
Improved tracing of torch.autograd.Functions (#166788)
Miscellaneous smaller tracing support additions:
Extend collections.defaultdict support with *args, **kwargs and custom default_factory (#166793)
Support for bitwise xor (#166065)
Support repr on user-defined objects (#167372)
Support new typing union syntax X | Y (#166599)
torch._inductor.config.combo_kernels (#162442) (#166274) (#162759) (#167781) (#168127) (#168946) (#168109) (#164918)qconv_pointwise.tensor and qconv2d_pointwise.binary_tensor quantized operations. (#166608)out_dtype argument for matmul operations. (#163393)fast_tanhf on ROCm. (#162052)fallback_embedding_bag_byte_unpack. (#163803)pad_mm and decompose_mm_pass pass on Intel GPU. (#166618) (#166613)native_layer_norm_backward meta function. (#159830)assume_32bit_indexing inductor config option. (#167784)embedding_bag operator (#163012, #163931, #163281)embedding_bag and index_select ops (#166615, #168930, #166468)share_memory_ (#162272)scaled_dot_product_attention (#166318)BlockMask in nn.functional.flex_attention (#164702)tofile() in ONNX IR tensors for more efficient ONNX model serialization (#165195)Adam, AdamW work with nonzero-dim Tensor betas (#149939)user_metadata display to memory visualizer (#165939)torch.library and custom ops to support view functions (#164520)generator arg to rand*_like APIs (#166160)half and bf16 support for fused_moving_avg_obs_fake_quant (#162620, #164175)bf16 support for fake_quantize_learnable_per_channel_affine (#165098)bf16 support for backward of torch._fake_quantize_learnable_per_tensor_affine (#165362)NVFP4 two-level scaling to scaled_mm (#165774)fp8_input/fp8_weight/bf16_bias and bf16_output for fp8 qconv in CPU (#167611)torch.float4_e2m1fn_x2 dtype support equality comparisons (#169575)copy_ support for torch.float4_e2m1fn_x2 dtype (#169595)--nproc-per-node torchrun option for Intel GPU (#159474)torch.nonzero_static crash on CUDA when the input is a empty tensor (#162578)C10_CUDA_CHECK error messages (#162808)parallel_cat (#165023)const_cast in CUDA memcpy call (#168165)torch.cuda._compile_kernel (#162647)torch.unique_consecutive crash on CUDA (#162950)parallel_cat (#165446)torch.kthvalue deterministic on CUDA (#165762)reduce_kernel (#162995)Tensor.__dlpack__(stream=None) support during CUDA Graph capture (#163242)out_dtype overload for addmm/baddbmm (#167931)c10d
Context Parallel
DistributedDataParallel: (DDP)
DistributedStateDict
DTensor
foreach_max op (#169667)FullyShardDataParallel (FSDP1 and FSDP2)
Pipeline Parallelism:
SymmetricMemory
cProfile usage with torch.compile in Python 3.12+ (#170013)repeat_interleave out-of-bounds indices, divmod error, alpha/beta handling). (#165368) (#163482) (#167317)argmin/argmax returning incorrect indices for non-contiguous tensors. (#165983)static_input_indices subclass remapping under training. (#167127)torch.cond HOP device handling in inductor. (#167354)CppTile2DKernel for FP8 datatype. (#167451).cpu() correctness issue. (#168281)as_strided with .view(dtype) inputs, symbolic shapes in FlexAttention, sym_size_/sym_stride). (#163319) (#168383) (#167565)copy_ for scalar in inductor. (#164167)searchsorted for non-dense tensors. (#165064)FailedMatch format string. (#169611)welford_reduce. (#162709)max_numwarps in coordinate descent tuner. (#159146)group_m config. (#169514)groups != 1. (#163094)get_raw_stream undefined error. (#163707)weight_norm backward. (#167667)aot_compile typing. (#168320)median/nanmedian/mv, dot (#162846, #166561), #165237)fill and cat operation (#164108, #165373, #166556, #164416)torch.compile bugfixes (#169648, #163021, #162776, #163452)score_mod in nn.functional.flex_attention (#163677)nn.Module.load_state_dict for singleton tensor (#166335)torch.onnx.ops)SWALR.state_dict and load_state_dict to serialize properly with weights_only=True (#163122)LBFGS wolfe max iteration (#161488)ProfilerState typo ('Disable' → 'Disabled') and expose PRIVATEUSE1 in ActiveProfilerType (#169166)__builtin_amdgcn_rcpf for gfx90a (#166454)output_padding on Intel GPU (#169176)convolution_overrideable on Intel GPU(#166839)radix_sort_pairs to improve torch.sort performance (#167094)pinned_reserve_segment_size_mb to speed up pinned memory allocation (#164501)torch.EmbeddingBag performance (#167834)empty_permuted decomposition. (#169731)inference_mode hint message to use eval with inference. (#163619)tlparse (#171339).
tlparse is a compilation report tool that processes TORCH_TRACE logs to generate interactive HTML reports showing how your model was compiled.
When reporting bugs to PyTorch developers, we encourage you to attach the trace log or tlparse output to provide critical debugging information to help us bisect the issue.tlparse (#171339) (#162975). tlparse is a compilation report tool that processes TORCH_TRACE logs to generate interactive HTML reports showing how your model was compiled. When reporting bugs to PyTorch developers, we encourage you to attach the trace log or tlparse output to provide critical debugging information to help us bisect the issue.Update CTCLoss docs float32 input required for CuDNN (#162042)
Update LPPool docs to clarify ceil_mode padding semantics when ceil_mode=True (#163186)
FunctionEvent (#167688)Note truncated.
This release is meant to fix the following issues (regressions / silent correctness):
This release is meant to fix the following issues (regressions / silent correctness):
Significant Memory Regression in F.conv3d with bfloat16 Inputs in PyTorch 2.9.0 (#166643) This release provides work around this issue. If you are impacted please install nvidia-cudnn package version 9.15+ from pypi. (#166480) (#167111)
Fix Inductor bug when compiling Gemma (#165601) Fix InternalTorchDynamoError in bytecode_transformation (#166036) Fix silent correctness error_on_graph_break bug where non-empty checkpoint results in unwanted graph break resumption (#166586) Improve performance by avoiding recompilation with mark_static_address with cudagraphs (#162208) Improve performance by caching get_free_symbol_uses in torch inductor (#166338) Fix fix registration design for inductor graph partition for vLLM (#166458) (#165815) (#165514) Fix warning spamming in torch.compile (#166993) Fix exception related to uninitialized tracer_output variable (#163169) Fix crash in torch.bmm and torch.compile with PyTorch release 2.9.0 (#166457)
Fix warning spamming on new APIs to control TF32 behavior (#166956) Fix distributed crash with non-contiguous gather inputs (#166181) Fix indexing on large tensor causes invalid configuration argument (#166974) Fix numeric issue in CUDNN_ATTENTION (#166912) (#166570) Fix symmetric memory issue with fused_scaled_matmul_reduce_scatter (#165086) Improve libtorch stable ABI documentation (#163899) Fix image display on pypi project description section (#166404)
This upgrade is doing the same BC-breaking changes as the DLPack release. Objects in torch.utils.dlpack have been updated to reflect these changes, su…
<table> <tr> <td><strong>Unstable (API-Unstable)</strong></td> </tr> <tr> <td>Updates to the stable libtorch ABI for third-party C++/CUDA extensions</td> </tr> <tr> <td>Symmetric memory that enables easy programming of multi-GPU kernels</td> </tr> <tr> <td>The ability to arbitrarily toggle error or resume on graph breaks in torch.compile</td> </tr> <tr> <td>Expanded wheel variant support to include ROCm, XPU and CUDA 13</td> </tr> <tr> <td>FlexAttention enablement on Intel GPUs</td> </tr> <tr> <td>Flash decoding optimization based on FlexAttention on X86 CPU</td> </tr> <tr> <td>ARM Platform improvements and optimizations</td> </tr> <tr> <td>Enablement of Linux aarch64 binary wheel builds across all supported CUDA versions</td> </tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
The minimum version of Python required for PyTorch 2.9.0 is 3.10. We also have 3.14 and 3.14t available as preview with this release.
This is a reminder that outputs of PyTorch custom operators (that are registered using the torch.library or TORCH_LIBRARY APIs) are not allowed to return Tensors that share storage with input tensors. The violation of this condition leads to undefined behavior: sometimes the result will be correct, sometimes it will be garbage.
After #163227, custom operators that violated this condition that previously returned correct results under torch.compile may now return silently incorrect results under torch.compile. Because this is changing the behavior of undefined behavior, we do not consider this to be a bug, but we are still documenting it in this section as a "potentially unexpected behavior change".
This is one of the conditions checked for by torch.library.opcheck and is mentioned in The Custom Operators Manual
Outputs of PyTorch custom operators are not allowed to return Tensors that share storage with input tensors
For example, the following two custom operators are not valid custom operators:
@torch.library.custom_op("mylib::foo", mutates_args=())
def foo(x: torch.Tensor) -> torch.Tensor:
# the result of `foo` must not directly be an input to foo.
return x
@torch.library.custom_op("mylib::bar", mutates_args=())
def bar(x: torch.Tensor) -> torch.Tensor:
# the result of bar must not be a view of an input of bar
return x.view(-1)
The easiest workaround is to add an extra .clone() to the outputs:
@torch.library.custom_op("mylib::foo", mutates_args=())
def foo(x: torch.Tensor) -> torch.Tensor:
return x.clone()
@torch.library.custom_op("mylib::bar", mutates_args=())
def bar(x: torch.Tensor) -> torch.Tensor:
return x.view(-1).clone()
A common way to get into this situation is for a user to want to create a custom operator that sometimes mutates the input in-place and sometimes returns a new Tensor, like in the following example.
@torch.library.custom_op("mylib::baz", mutates_args=["x"])
def baz(x: torch.Tensor) -> torch.Tensor:
if inplace:
x.sin_()
return x
else:
return x.sin()
This dynamism is not supported and leads to undefined behavior. The workaround is to split the custom operator into two custom operators, one that always mutates the input in-place, and another that always returns a new Tensor.
@torch.library.custom_op("mylib::baz_outplace", mutates_args=())
def baz_outplace(x: torch.Tensor) -> torch.Tensor:
return x.sin()
@torch.library.custom_op("mylib::baz_inplace", mutates_args=["x"])
def baz_inplace(x: torch.Tensor) -> torch.Tensor:
x.sin_()
def baz(x):
if inplace:
baz_inplace(x)
return x
else:
return baz_outplace(x)
PyTorch MPS is only supported on MacOS-14 or later. If you need to use MPS on MacOS Ventura, please avoid updating to Python-3.9 or above
This upgrade is doing the same BC-breaking changes as the DLPack release. Objects in torch.utils.dlpack have been updated to reflect these changes, such as DLDeviceType.
See the PR for details on the exact changes and how to update your code.
torch.cat (#158249)torch.cat now raises ValueError, IndexError or TypeError where appropriate instead of the generic RuntimeError. If you code was catching these errors, you can update to catch the new error type.
dynamo=True for ONNX exporter (#159646, #162726)Previously torch.onnx.export(...) used the legacy TorchScript exporter if no arguments were provied. The ONNX exporter now uses the newer torch.export.export pipeline by default (dynamo=True). This change improves graph fidelity and future-proofs exports, but may surface graph capture errors that were previously masked or handled differently.
Previously in torch 2.8.0:
# API calls the legacy exporter with dynamo=False
torch.onnx.export(...)
Now in torch 2.9.0:
# To preserve the original behavior
torch.onnx.export(..., dynamo=False)
# Export onnx model through torch.export.export
torch.onnx.export(...)
Recommendation: first try the new default; only fall back if you hit blocking issues and report them upstream. Long term solution: fix the root cause instead of relying on fallback or TorchScript exporter.
To enable runtime asserts, use export(..., prefer_deferred_runtime_asserts_over_guards=True). Also kills the allow_complex_guards_as_runtime_asserts flag, merging it into the former option.
Additionally, exported_program.module() will generate a call to a _guards_fn submodule that will run additional checks on inputs. Users who do not want this behavior can either remove this call in the graph, or do exported_program.module(check_guards=False) to avoid the generation.
Opset 20 enables newer operator definitions. If your tooling or downstream runtime only supports opset 18, pin it explicitly. For the latest ONNX operators, you can experiment with opset 23.
Previously in torch 2.8.0:
# opset_version=18
torch.onnx.export(...)
Now in torch 2.9.0:
# To preserve the original behavior
torch.onnx.export(..., opset_version=18)
# New: opset_version=20
torch.onnx.export(...)
# Use the latest supported opset: opset_version=23
torch.onnx.export(..., opset_version=23)
draft_export in exporter API (#161454, #162225)Remove implicit draft tracing from the default exporter path, achieving clearer behaviour and faster failures.
The expensive torch.export.draft_export diagnostic path is no longer auto-invoked (which could take hours on large models). You can still opt in for deep diagnostics:
Previously in torch 2.8.0:
# If both torch.export.export(..., strict=False) and
# torch.export.export(..., strict=True) fail to capture
# the model graph, torch.export.draft_export(...) will be triggered,
# and uses real tensor to trace/export the model.
#
# Inside export_to_onnx.py:
# ... torch.onnx.export(..., dynamo=True)
python export_to_onnx.py
Now in torch 2.9.0:
# To trigger torch.export.draft_export once
# torch.export.export strict=False/True both
# fail:
TORCH_ONNX_ENABLE_DRAFT_EXPORT=True python export_to_onnx.py
torch.onnx.dynamo_export and the onnxrt torch compile backend (#158130, #158258)torch.onnx.dynamo_export is removed. Please use torch.onnx.export instead.
The experimental ONNX Runtime compile backend (torch.compile(backend="onnxrt")) is no longer supported.
torch.onnx.enable_fake_mode (#161222)The dynamo=True mode uses FakeTensors by default which is memory efficient.
Deprecated members in torch.onnx.verification are removed. Previously private torch.onnx.symbolic_opsets* functions will no longer be accessible. Consider making a copy of the source code if you need to access any private functions for compatibility with the TorchScript based exporter.
torch.onnx.symbolic_caffe2 (#157102)Support for caffe2 in the ONNX exporter has ended and is removed.
/d2implyavx512upperregs flag that slows build (#159431)Re-introduced AVX512 optimizations for Windows VS2022 builds, may cause issues with specific versions of VS2022, see #145702
ScalarType to shim conversion and stable::Tensor.scalar_type (#160557)Before, user extensions could only in abstract pass around obfuscated dtypes appearing as int32_ts. Now, users can confidently use torch::headeronly::ScalarType in their extensions for major scalar types. This PR enables ABI stability by adding a translation layer through the shim, so that even if the ScalarType enum values change in the future, user extensions need not fear.
This change adds ScalarType support for user extensions and is only narrowly BC breaking for unpopular dtypes: quint*s, qint*s, Bits*, dummy_uint*s, dummy_int*s, Float8_e8m0fnu, and Float4_e2m1fn_x2 in the use case where an extension retrieves a Tensor dtype of the above and passes it into aoti_torch_call_dispatcher.
pin_memory_device param in torch.utils.data.DataLoader (#158323)We move enabling pin_memory back inside BaseDataLoaderIter. This is required for StatefulDataloader which leveraged BaseDataLoaderIter direclty rather than the Dataloader class init
torch.export.export_for_training API in favor of equivalent torch.export.export API (#158203)torch.export.export_for_training exists because we couldn't migrate internal usages of export to the final IR. Now that we have completed the migration, we deprecated and deleted this API.
__torch_function__ handler to be triggered by elements within a list (#160256)torch.hash_tensor reduction function (#154149)is_fx_symbolic_tracing flag (#161385)Introduces an API for annotating dynamic integer inputs & attributes for torch.compile, by wrapping plain ints with DynamicInt().
DynamicInt objects also work in eager mode, acting as their underlying values when passed as scalar inputs.
a = DynamicInt(4)
y = a + 2 # DynamicInt(6)
z = torch.ones(a) # torch.ones(4)
fn = torch.compile(torch.ones)
fn(a) # compiled fn takes a dynamic integer input
fn(2) # returns torch.ones(2) without recompiling
torch/csrc/stable/ops.h: amax, narrow, new_empty + new_zeros dtype variant, pad, (#159328, #158974, #159508, #161597, #160214)torch::stable::Tensor() default constructor, is_cpu, and get_device_index(#159507, #160212, #160143)torch::stable::accelerator with support for DeviceGuard and Stream (#159679, #160453)torch/headeronly: c10 Macros, STD_TORCH_CHECK, ScalarTypes (like BFloat16 and Half, #158035, #158365, #157912, #158377, #159302, #159414, #159412, #159415, #159411, #159911)TORCH_ERROR_CODE_CHECK in headeronly from AOTI (#159604)wheel from build requirements (#158027)TORCH_STABLE_ONLY is defined in TensorBase.h (#161658)torch/csrc/stable (#158160)zero_() and empty_like(t) to torch/csrc/stable/ops.h (#158866)Add support for CUDA 13.0 in CI/CD builds. Enable CUDA compression mode for binary size reduction for CUDA 13.0 builds (#160956, #161073, #161257, #161663, #161316, #160201, #160770, #161013, #161916, #162268, #162322, #162383, #161833)
Enable CUDA 12.6, 12.8 and 13.0 support for Linux ARM64 CD builds (#162364, #160720, #159481)
Add support for Python 3.14 in CI/CD builds (#156889, #157559, #159261, #159869, #160593, #160788, #161255, #159725)
Enable NVSHMEM integration (#151261, #153010, #154538, #155506, #156685, #158938, #161321, #160778, #159907, #160465)
cudnn_batch_norm_out kernel to replace the autogen approach (#123020)avg_pool3d, max_unpool1d/2d/3d, max_pool3d, max_pool3d bwd pass, and avg_pool3d bwd pass for MPS (#158877,#159789, #156467, #157498, #159089)FlexAttention on Intel GPU (#143553)torch.load under FakeTensorMode by reducing random reads (#157931)torch.utils.benchmark.utils.timer accelerator agnostic (#157131)register_buffer with Tensor-like objects (#159455)NLLLoss (#161412)requires_grad=True to a scalar (#160389)SequentialLR deprecation warning about invoking step(epoch) (#149392)torch.nn.Upsample mode="trilinear" backward (#154239)ProcessGroupNCCL (#156748)ProcessGroupNCCL (#156790)ProcessGroupGloo (#158128)TCPStore (#159165)all_gather (#149913)unsafe_get_ptr for dist.ProcessGroupNCCL.NCCLConfig (#161136)send/recv_object_list (#160342)ProcessGroupGloo (#156633)work.isStarted (#160398)device_mesh argument constraint in local_map (#157049)DTensor slice (#157953)histc op (#158298)propagate_tensor_meta function that skips cache if _are_we_tracing (#161334)local_map as a decorator (#161353)init_device_mesh (#159371)all_gather and reduce_scatter comms (#155189)set_allocate_memory_from_process_group if used together with custom comm hooks (#157487)reduceOpSum when world size is 1 (#157529)allgather when world size is 1 (#160135)post_reduce_stream.record_event() on hsdp+cpuoffload (#160481)parallelize_module API to support more cases (#157182)torchrun (#149334)eval() API to schedule (#157795)OVERLAP_F_B computation type (#158978)DualPipeV schedule (#159591)is_impure() (#151524, #157981)self.module_stack in ModuleStackTracer (#159956)node_name_match to subgraph rewriter (#157574)lists (e.g. #153969)sets (e.g. #153150)dicts (e.g. #154794)iter (e.g. #156371)itertools (e.g. #159693)collections (e.g. #159365)collections.NamedTuple (#159367)dataclasses.dataclass (#159529)TorchDispatchMode to ignore torch.compile internals (#161648){non-inf, > int32_max} upper bound is provided (#159433)None & ellipsis slicing/select in non-strict (#157821)triton_kernel_wrapper_functional HOP (#161314)while_loop HOP subgraphs (#158467)aot_export_joint_with_descriptors and aot_compile_joint_with_descriptors (#158715)prepare_aot_module_simplified for use in next PR (#158319)aot_export_joint_with_descriptors (#159814)aten.add.Scalar (#161332)aten.expand_copy decomp (#161688)aten.linalg_vector_norm (#155111)aten.complex (#160894)HistogramObserver (#156457)bias=None for fbgemm_linear_fp16_weight CPU op (#158535)wrapped_fbgemm_linear_fp16_weight for Sigmoid (#160451)log_softmax() support (#159662)vector.reserve() consistently for non-inplace foreach operations (#161128)has_integral_tensor() (#161042)torch.tensor warning in ONNX symbolic_opset10 export (#158835)AllocatorConfig to be device-agnostic via new AcceleratorAllocatorConfig (#149601, #150312)Scalar::isUnsigned() method (#159877)ModelRunner from nativert as public (#159989)torch.binomial enforcing float inputs (#157658)Dependencies.cmake (#159702)libtorch without NVSHMEM (#160910)repeat_interleave kernel (#157996)kernelLaunchCheck to print help string (#158896)Wunused-but-set-variable (#159276)libc++ (#161101)shifted_chebyshev_polynomial_[tuvw], igamma/igammac,grid_sampler_3d, native_dropout/native_dropout_backward (#157488, #161927, #160541, #162108)index_put to complex types (#160159)addmm to integral types (#160270)kthvalue (#161817)logcumsumexp metal kernel (#156858)dlpack integration (#158888)avg_pool2d to use Metal kernel when ceil_mode=True (#161011)composable_kernel (CK) backend user interface to improve user experience (#152951)rocSOLVER for Cholesky inversion. (#157154)torch.backends.miopen.immediate to toggle MIOpen Immediate Mode instead of relying on deterministic=True and benchmark=False (#158951)reshape_ or unexpectedly change memory formats (#161687)device_id to Intel GPU properties to distinguish iGPUs with identical names (#156481)torch.utils.cpp_extension.load_inline to override gencode (#156850)max_width computation in Tensor printing (#126859)pin_memory error message on CPU-only systems (#159994)F.embedding DTensor-aware (#162117)torch.autograd.Function memory leak due to torch.utils.checkpiont early stopping (#161171)torch.autograd.graph.GradientEdge for torch.autograd.Function (#160098)setGroupName and setGroupDesc in group_split and merge_remote_group (#159429)batch_isend_irecv with 2D tensor views by making P2P tensors dense (#163719)allgather/reducescatter inputs (#163712)DDPOptimizer and donated buffers (#160745)OpSchema equality check (#161231)grouped_mm strategy for invalid stride cases (#158245)F.one_hot in DTensor (#162307)ShardingPropagation cache if compiling (#156868)pin_memory (#157147)NO_SHARD correctly by flattening tensors before copying (#154369)fsdp_pre_all_gather (#160817)set_reduce_scatter_divide_factor errors and MixedPrecisionPolicy (#155964)no_grad() (#159293)eval() (#159475)import torch if compiled without TensorPipe (#159461)torch.distributed.elastic.multiprocessing.start_processes() (#160396)split_module with symint (#160093)getattr_recursive with ModuleList (#161204)torch.Tensor (#162224)torch.compiler.reset() (#156527)FlexAttention (#163677)return_lse warning message in FlexAttention (#163578)FlexAttention head broadcast (#163426)constant_pad_nd (#159878)MutationOutput Buffer (#162020)NaN behavior (#159308)dtype consistency (#160851)FallbackKernel alias function to avoid incorrect aliasing for custom ops (#163227)load_constants (#161887)gen_aoti_c_shim (#159904)aoti_torch_as_strided (#162118)wait_tensor returned tensor (#159502)all_reduce (#159818)from_node provenance in unlift pass (#157943)NaN serialization (#155359)nn_module_stack for assert_tensor_metadata nodes (#159625)move_to_device_pass (#159992, #160528, #162301)dynamo.disables in unflattening (#161306).contiguous() when saving weights to raw bytes to preserve original storage size of tensor (#163587)NaN in fp8 output of CPU qlinear and qconv ops (#160957)choose_qparams_optimized (#161966)chunk_size should always be int64_t for Foreach functors (#156872)rotary_embedding_23 implementation (#162865)None as output (#160200)dynamo=True (#161056)index_put_ usage (#161263)TORCH_CUDA_ARCH_LIST to only log when running on non-distributed or on rank 0 (#162764)torch.utils.cpp_extension parser for clang version 20.1.7+libcxx (#157666)MakeTensor::computeStorageSize() calculation (#158690)AllocatorConfig (#159629)BUILD_BUNDLEPTXAS=1 to allow compile on newer GPUs(#163988)torch.backends.cuda.matmul.fp32_precision (#161102)cudaErrorNotSupported (#162412)__syncthreads in MultiMarginLoss backward (#158994)kernel size != 1 for cuDNN 9.8+ (#163581)MKLDNN path (#162168)tensor dim > INT_MAX (#158824)exponential_ for MPS (#159386)ConvTranspose3d (#160345)conv3d autocast bias dtype mismatch (#160423)torch.var on scalar (#160889)index_add for complex + int64, int64 input + zerodim index (#160926, #161511)constant_pad_nd_mps bug when pad is empty (#161149)index_select for scalar_types (#161206)index_copy for scalars and index_copy for strided indices (#161267, #161333)NaNs if SDPA is called with all values masked from query (#157727)scaled_dot_product_attention using MPS (#163598)fillBuffer into 4Gb slices to avoid regression on MacOS 26 (#164108)hip:0 device error (#161221)cpp_extension.py on Windows (#159790)logsumexp needs scaling back to natural base (#156903)memcpy_and_sync inline function (#158165)cpp_extension compatibility with intel-deep-learning-essentials-2025.2 (#161012)ErrorReport::CallStack thread-safe (#160386)RemoveProfileNodesAndSpecializeTypes handling for Tensor? that is resolved to None (#161538)addmm to improve Newton–Schulz orthogonalization in Muon (#161379)AveragedModel.update_parameters() (#157705)dict tag optimization for faster guard evaluation (#159183)GEMM template (#159127, #161148)fmod.Scalar and scale_gradient (#160654, #160454)torch.full for 1-byte types (#158874)argmax/argmin (#159524)max_pool3d (#157875)max_pool3d impl (#157874)max_pool2d to Metal for stride != 1 (#157876)hipblaslt is used by default on gfx908 for ROCm >= 6.3 (#159092)tensor.item() (#158486)fbgemm_gpu genai sources for grouped GEMM support (#160676)torch.lobpcg, torch.clone, torch.matmul, torch.max, torch.gather, torch.Tensor.scatter_, torch.empty_like, torch.randint, torch.mul, torch.min, torch.max. torch.sort, torch.full_like, torch.histogramdd, torch.hamming_window (#156139, #157007, #161424, #156153, #157929, #157920, #158050, #158731, #160312, #161539, #162051, #158275, #152682)torch.set_float32_matmul_precision docs (#158191)torch.nn.utils.clip_grads_with_norm_ to reflect clamping behavior (#158200)torch.gradient (#159130)torch.segment_reduce docs (#154352)torch.is_floating_point and torch.is_complex docs (#161951)padding for avg_poolnd (#159142)CrossEntropyLoss docs with example of incorrect target specification (#155649)scaled_dot_product_attention example (#161613)torch.optim.adam.Adam, properly (#158483, #158669, #160194)zero_grad description (#161239)torch.inference_mode docs and error message (#161164)device_id (#159389)distributed_c10d.py (#162158)TupleStrategy (#158132)redistribute_costs (#158495)torch/ (torch/fx/, #156604)torch/fx/experimental/migrate_gradual_types/util.py (#157236)symbolic_multi_out (#160702)onnx.md to simplify deprecated entities (#159312)fallback=False by default (#162622, #162726)waitcounter for watchdog and heartbeat monitoring thread (#157480)torch.distributed.breakpoint set a long timeout (#158481)check_rng_sync util (#160283)FlightRecorder support for ProcessGroupXCCL (#158568)early_stop kwarg to torch.utils.checkpoint (#160781)register_foward_pre_hook not supported on ScriptModule error (#156904)__eq__ function to NodeSource (#158170)__hash__ function to NodeSource (#158322)tlparse error (#158469)expanded_def option for FX printing, render descriptor, update tests (#158708)co_lnotab in favor of co_linetable (#159227)CompiledFxGraph (#159311)allow_tf32 in tl.dot(..., allow_tf32=...), use tl.dot(..., input_precision=...) (#160711)jit_post_compile_hook within Inductor Triton Kernel compile path (#161443)torch.backends.cuda.matmul.allow_tf32 in Inductor cache key (#159480)guard_size_oblivious on data dependent errors (#160510)torch.utils.data samplers benchmark script (#156974)torch.utils.data.Dataloader benchmark script (#159432)setup.py develop with pip install -e for development builds (#155998, #156027, #156710) (#156709)This is a BC-breaking change for the build system interface. Downstream projects that previously got NVTX3 through cmake/public/cuda.cmake (i.e.. call…
<table> <tr> <td><strong>Unstable</strong></td> </tr> <tr> <td>torch::stable::Tensor</td> </tr> <tr> <td>High-performance quantized LLM inference on Intel CPUs with native PyTorch</td> </tr> <tr> <td>Experimental Wheel Variant Support</td> </tr> <tr> <td>Inductor CUTLASS backend support</td> </tr> <tr> <td>Inductor Graph Partition for CUDAGraph</td> </tr> <tr> <td>Control Flow Operator Library</td> </tr> <tr> <td>HuggingFace SafeTensors support in PyTorch Distributed Checkpointing</td> </tr> <tr> <td>SYCL support in PyTorch CPP Extension API</td> </tr> <tr> <td>A16W4 on XPU Device</td> </tr> <tr> <td>Hierarchical compilation with torch.compile</td> </tr> <tr> <td>Intel GPU distributed backend (XCCL) support</td> </tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
Due to a bug introduced in CUDA 12.9.1, we are unable to complete full Windows wheel builds with this
version, as compilation of torch.segment_reduce() crashes the build. Thus, we provide a wheel
without torch.segment_reduce() included in order to sidestep the issue. If you need support
for torch.segment_reduce(), please utilize a different version.
Due to binary size limitations, support for sm50 - sm60 architectures with CUDA 12.8 and 12.9 has been dropped for the 2.8.0 release. If you need support for these architectures, please utilize CUDA 12.6 instead.
NotImplementedError instead of RuntimeError (#155470)Please update exception handling logic to reflect this.
In 2.7.0
try:
torch.nn.Hardshrink()(torch.randint(0, 5, (10,)))
except RuntimeError:
...
In 2.8.0
try:
torch.nn.Hardshrink()(torch.randint(0, 5, (10,)))
except NotImplementedError:
...
autograd.Function (#153094)In 2.8.0, if a custom autograd.Function mutates a view of a leaf requiring grad,
it now properly raises an error. Previously, it would silently leak memory.
class Func(torch.autograd.Function):
@staticmethod
def forward(ctx, inp):
inp.add_(1)
ctx.mark_dirty(inp)
return inp
@staticmethod
def backward(ctx, gO):
pass
a = torch.tensor([1.0, 2.0], requires_grad=True)
b = a.view_as(a)
Func.apply(b)
Output:
Version 2.7.0
Runs without error, but leaks memory
Version 2.8.0
RuntimeError: a view of a leaf Variable that requires grad is being used in an in-place operation
tensordot when called with a requires_grad=True tensor (#150270)Please avoid passing an out tensor with requires_grad=True as gradients cannot be
computed for this tensor.
In 2.7.0
a = torch.empty((4, 2), requires_grad=True)
b = torch.empty((2, 4), requires_grad=True)
c = torch.empty((2, 2), requires_grad=True)
# does not error, but gradients for c cannot be computed
torch.tensordot(a, b, dims=([1], [0]), out=c)
In 2.8.0
a = torch.empty((4, 2), requires_grad=True)
b = torch.empty((2, 4), requires_grad=True)
c = torch.empty((2, 2), requires_grad=True)
torch.tensordot(a, b, dims=([1], [0]), out=c)
# RuntimeError: tensordot(): the 'out' tensor was specified and requires gradients, and
# its shape does not match the expected result. Either remove the 'out' argument, ensure
# it does not require gradients, or make sure its shape matches the expected output.
mark_dynamic applied now correctly errors (#152661)Prior to 2.8, it was possible for a guard on a symbolic shape to be incorrectly
omitted if the symbolic shape evaluation was previously tested with guards
suppressed (this often happens within the compiler itself). This has been fixed
in 2.8 and usually will just silently "do the right thing" and add the correct
guard. However, if the new guard causes a tensor marked with mark_dynamic to become
specialized, this can result in an error. One workaround is to use
maybe_mark_dynamic instead of mark_dynamic.
See the discussion in issue #157921 for more context.
Version 2.7.0
import torch
embed = torch.randn(2, 8192)
x = torch.zeros(8192)
torch._dynamo.mark_dynamic(x, 0)
@torch.compile
def f(embedding_indices, x):
added_tokens_mask = torch.where(x > 10000, 1, 0)
ei = torch.narrow(embedding_indices, 1, 0, x.size(0))
return ei.clone()
f(embed, x)
Version 2.8.0
import torch
embed = torch.randn(2, 8192)
x = torch.zeros(8192)
torch._dynamo.maybe_mark_dynamic(x, 0)
@torch.compile
def f(embedding_indices, x):
added_tokens_mask = torch.where(x > 10000, 1, 0)
ei = torch.narrow(embedding_indices, 1, 0, x.size(0))
return ei.clone()
f(embed, x)
torch.compile have been renamed or removedenable_cpp_framelocals_guard_eval has changed to no longer have any effect (#151008).rocm.n_max_profiling_configs is deprecated (#152341).
Instead, use ck-tile based configs rocm.ck_max_profiling_configs and
rocm.ck_tile_max_profiling_configs.autotune_fallback_to_aten is deprecated (#154331).
Inductor will no longer silently fall back to ATen. Please add "ATEN" to
max_autotune_gemm_backends for the old behavior.use_mixed_mm and mixed_mm_choice are deprecated (#152071). Inductor now supports prologue fusion, so there is no need for
special cases now.descriptive_names = False is deprecated (#151481). Please use one of the other available
options: "torch", "original_aten", or "inductor_node".custom_op_default_layout_constraint has moved from inductor config to functorch config (#148104). Please reference it via
torch._functorch.config.custom_op_default_layout_constraint instead of
torch._inductor.config.custom_op_default_layout_constraint.emit_current_arch_binary is deprecated (#155768).aot_inductor.embed_cubin has been renamed to aot_inductor.embed_kernel_binary (#154412).aot_inductor.compile_wrapper_with_O0 has been renamed to compile_wrapper_opt_level (#148714).HigherOrderOperators (e.g. cond), which will explicitly error out if alias/mutation among inputs and outputs is unsupported (#148953, #146658).For affected HigherOrderOperators, add .clone() to aliased outputs to address this.
Version 2.7.0
import torch
@torch.compile(backend="eager")
def fn(x):
return torch.cond(x.sum() > 0, lambda x: x, lambda x: x + 1, [x])
fn(torch.ones(3))
Version 2.8.0
import torch
@torch.compile(backend="eager")
def fn(x):
return torch.cond(x.sum() > 0, lambda x: x.clone(), lambda x: x + 1, [x])
fn(torch.ones(3))
guard_or_x and definitely_x have been consolidated (#152463)We removed definitely_true / definitely_false and associated APIs, replacing them with
guard_or_true / guard_or_false, which offer similar functionality and can be used to
achieve the same effect. Please migrate to the latter.
Version 2.7.0
from torch.fx.experimental.symbolic_shapes import definitely_false, definitely_true
...
if definitely_true(x):
...
if definitely_false(y):
...
Version 2.8.0
from torch.fx.experimental.symbolic_shapes import guard_or_false, guard_or_true
...
if guard_or_false(x):
...
# alternatively: if guard_or_false(torch.sym_not(y))
if not guard_or_true(y):
...
torch.export.export_for_inference has been removed in favor of torch.export.export_for_training().run_decompositions() (#149078)Version 2.7.0
import torch
...
exported_program = torch.export.export_for_inference(mod, args, kwargs)
Version 2.8.0
import torch
...
exported_program = torch.export.export_for_training(
mod, args, kwargs
).run_decompositions(decomp_table=decomp_table)
strict=False in torch.export.export and export_for_training (#148790, #150941)This differs from the previous release default of strict=True. To revert to the old default
behavior, please explicitly pass strict=True.
Version 2.7.0
import torch
# default behavior is strict=True
torch.export.export(...)
torch.export.export_for_training(...)
Version 2.8.0
import torch
# strict=True must be explicitly passed to get the old behavior
torch.export.export(..., strict=True)
torch.export.export_for_training(..., strict=True)
torch.onnx.export is now 18 (#156023)When dynamo=False, the default ONNX opset version has been updated from 17 to 18. Users can set opset_version to explicitly select an opset version.
Version 2.7
# opset_version=17
torch.onnx.export(...)
Version 2.8
# To preserve the original behavior
torch.onnx.export(..., opset_version=17)
# New: opset_version=18
torch.onnx.export(...)
JitTraceConvertStrategy has been removed (#152556)Support for JIT traced and scripted modules in the ONNX exporter when dynamo=True has been removed. You are encouraged to export an nn.Module directly, or create an ExportedProgram using torch.export before exporting to ONNX.
onnxscript>=0.3.1 is required for the dynamo=True option (#157017)You must upgrade onnxscript to version 0.3.1 or higher for it to be compatible with PyTorch 2.8.
torch/types.h include from Dispatcher.h (#149557)This can cause build errors in C++ code that implicitly relies on this include (e.g. very old versions of torchvision).
Note that Dispatcher.h does not belong as an include from torch/types.h and was only present as a
short-term hack to appease torchvision. If you run into torchvision build errors, please
update to a more recent version of torchvision to resolve this.
DLPack to 1.0 (#145000)As part of the upgrade, some of the DLDeviceType enum values have been renamed. Please switch
to the new names.
Version 2.7.0
from torch.utils.dlpack import DLDeviceType
d1 = DLDeviceType.kDLGPU
d2 = DLDeviceType.kDLCPUPinned
...
Version 2.8.0
from torch.utils.dlpack import DLDeviceType
d1 = DLDeviceType.kDLCUDA # formerly kDLGPU
d2 = DLDeviceType.kDLCUDAHost # formerly kDLCPUPinned
...
cmake/public/cuda.cmake to cmake/Dependencies.cmake (#151583)This is a BC-breaking change for the build system interface. Downstream projects that previously got NVTX3 through cmake/public/cuda.cmake
(i.e.. calling find_package(TORCH REQUIRED)) will now need to explicitly configure NVTX3 support in the library itself (i.e. use USE_SYSTEM_NVTX=1).
The change is to fix the broken behavior where downstream projects couldn't find NVTX3 anyway due to the PROJECT_SOURCE_DIR mismatch.
Version 2.7.0:
-DUSE_SYSTEM_NVTX would be able to find NVTX3 and torch::nvtx3 via PyTorch's cmake/public/cuda.cmake logic.-DUSE_SYSTEM_NVTX would encounter build errors with CUDA 12.8 or above.Version 2.8.0:
-DUSE_SYSTEM_NVTX will not be able to find NVTX3 or torch::nvtx3 via PyTorch's cmake/public/cuda.cmake. The downstream project now needs to explicitly find NVTX3 and torch::nvtx3 by implementing the same logic in PyTorch's cmake/Dependences.cmake.-DUSE_SYSTEM_NVTX will proceed building without NVTX unless another part of the build process re-enables NVTX.PyTorch 2.8 is the last release that will support GPU acceleration on MacOS Ventura. In the next release (2.9), MacOS Sonoma (released in Sept. 2023) or above will be required to use the MPS backend.
torch.ao.quantization is deprecated and will be removed in 2.10 (#153892)To migrate:
torch.ao.quantization.quantize, torch.ao.quantization.quantize_dynamic)
torchao eager mode quantize_.torchao PT2E quantization.torch.ao.quantization.quantize_fx.prepare_fx, torch.ao.quantization.quantize_fx.convert_fx): use torchao PT2E quantization (torchao.quantization.quantize_pt2e.prepare_pt2e, torchao.quantization.quantize_pt2e.convert_pt2e).Note that PT2E quantization has been migrated to torchao (https://github.com/pytorch/ao/tree/main/torchao/quantization/pt2e). See https://github.com/pytorch/ao/issues/2259 and https://docs.pytorch.org/ao/main/quick_start.html#pytorch-2-export-quantization for more details.
dynamo=False (current default) option for torch.onnx.export is deprecated (#152478, #155580)The default will be dynamo=True starting from PyTorch 2.9. You are encouraged to migrate to use the dynamo=True option in torch.onnx.export. This flag makes torch.export.export the default export path, replacing TorchScript.
To maintain the old behavior, set dynamo=False explicitly. You are encouraged to also experiment with the fallback=True option that will make the exporter fall back to the dynamo=False path if there are errors.
nested_compile_region (#156449)guard_filter_fn (#150936)dont_skip_tracing decorator to skip over most Dynamo skipfiles rules (#150586)draft-export, an export variant designed to consistently produce a graph and generate a debugging report of issues encountered during tracing (#152637, #153219, #149465, #153627, #154190, #155744, #150876, #150948, #151051, #151065, #150809, #151797)TorchBind objects (#150196, #154265)aot_inductor.model_name_for_generated_files for specifying model name (#154129)MPSInductor: torch.compile for Apple GPUs (#150121, #149342, #151449, #151754, #149687, #149180, #149221, #153598, #152788, #153787, #152214, #151152, #155891, #154578, #151272, #151288, #153997, #151871, #153362, #156566, #150661, #153582)Added new strategy draft_export (#147529, docs) to provide debugging information upon data-dependent / constraint errors when obtaining an ExportedProgram with torch.onnx.export
Added support for symbolic operators in the dynamo=True export path (#148905, #149678, #150038, docs). Two operators torch.onnx.ops.symbolic and torch.onnx.ops.symbolic_multi_out are defined to allow you to create symbolic ONNX operators directly in your PyTorch models. You can use them in a forward method:
def forward(self, x: torch.Tensor) -> torch.Tensor:
# Optionally use is_in_onnx_export to control the behavior during onnx export
if torch.onnx.is_in_onnx_export():
# Create a symbolic ONNX operator with the name "CustomOp" in the "custom_domain" domain.
# The output tensor will have the specified dtype and shape
return torch.onnx.ops.symbolic(
"custom_domain::CustomOp",
(x,),
dict(attr_key="attr_value"),
dtype=x.dtype,
shape=x.shape,
version=1,
)
else:
return x
torch.float4_e2m1fn_x2 dtype (#148791)TORCH_CUDA_ARCH_LIST (#152715, #155314)bicubic mode for torch::nn::functional::grid_sample (#150817)no_implicit_headers mode for load_inline() on custom CUDA extensions (#149480)TCPStore with clone and queuing features (#150966, #151045, #150969, #151485)getDefaultBackend more fault tolerant without relying on exceptions (#149152)masterListenFd in TCPStoreLibUvBackend (#150215)TORCH_NCCL_USE_TENSOR_REGISTER_ALLOCATOR_HOOK (#150682)global_rank when group_rank is used (#151373)ProcessGroupNCCL via an unsafe API (#152496)needs_contiguous_strides tag in functional collective (#153399, #153523)split_group to work with non-nccl backends (#152175)new_subgroups() by using new_subgroups_by_enumeration() (#153843)ProcessGroupNCCL (#153990)c10::Half for gloo (#153862)get_process_group_ranks() to accept group=None (#154902)init_process_group support index-only device id (#156214)ProcessGroup (#151723)reduce_scatter and ReduceOp::AVG in ProcessGroupGloo (#149781, #149869)ProcessGroupNCCL (#152706)ibverbs backend in gloo and enabled gloo CUDA when used with a backend that supports GPUDirect (#153015, #153425, #153406)use_python_reducer to C++ reducer (#152735)
DistributedStateDict (DSD)write_size in planner write items (#149699)StridedShard support uneven sharding (#150490)torch.cumsum (#151071)DTensor redistribute fwd/bwd datatype conversion to enable SimpleFSDP mixed precision training (#150740)torch.distributed.tensor.debug.visualize_sharding (#152027)PrivateUse1 backend in FSDP collectives and device type to pre forward hook (#147260, #149487)set_reshard_after_forward (#149103)reshard_after_forward=True for root model and kept root unsharded when not specifying reshard_after_forward (#154704, #155319)all_reduce_event only if it's not CPU device (#150316)get_pipeline_order() for Gpipe and 1F1B (#155935)ShardedTensor and recalculated metadata from all_gather (#152583)ParallelStyle PrepareModuleInputOutput (#150372)__torch_function__, and namedtuple subclasses (#153150, #149792, #153982)reason field to torch.compiler.disable (#150341)lru_cache warnings for functions in the top-level torch namespace (#157718)aot_inductor.custom_ops_to_c_shims and aot_inductor.custom_op_libs: allow for specifying custom op C shim (#153968)max_fusion_buffer_group_pairwise_attempts: limits fusions to specified node distance (#154688)cuda.cutlass_enabled_ops: controls CUTLASS operation selection (#155770)triton.cudagraph_capture_sizes: allows specifying certain shapes for which to capture CUDAGraphs; skips CUDAGraphs for other shapes (#156551)use_static_cuda_launcher: enables launching compiled triton statically to improve cold start times (#148890)assume_unaligned_fallback_output: allows inductor to track unaligned outputs (#150777)cuda.cutlass_tma_only: controls whether or not to only use TMA-compatible kernels in CUTLASS (#152815)static_launch_user_defined_triton_kernels: enables statically launching user defined triton kernels (#153725)precompilation_timeout_seconds: controls the timeout on precompilation (#153788)disable_decompose_k: disables new DecomposeK GEMM Kernels (#154421)min_num_split: sets the minimum number of splits in a split reduction (#155941)max_autotune_flex_search_space: allows specifying the size of the search space for flex attention autotuning (#156307)LOG_AUTOTUNE_RESULTS for autotune log (#156254)min, max, math.pow) (#151348)pytree.register_dataclass (#147752)jit.scripted functions in export (#155180)num_runners to AOTIModelPackageLoader (#149364)== (#150611)normalize_function (#143689)graph_code_verbose_log artifact for FX passes (#153775)fx.passes.split_module to normalize input names (#157793)cross (#154999)torch.special operations as well as index_copy, hardshrink, rsub, col2im, and isin (#149174, #149203 #149123, #149368, #149378, #149563, #149687, #149705, #149783, #149407/#149680, #150279, #151754, #153786, #154326, #155304, #156263, #155382, #154010, #149816, #152282, #156090, #150060, #151600, #155002, #154671)index_put with half precision floats (#151869)ConvTranspose3D with FP32 and complex (#154696)log1p and sigmoid with int64 (#151791)weight_norm on CPU (#148878)dynamo=True (#149901, #154596)Attention-23 and RotaryEmbedding-23 as native PyTorch ops (#156431, #156367, #154745)torch.scan (#154513)group_norm support from opset 21 (#152138)asdict method to VerificationInfo class (#151024)dynamic_shapes behavior to use torch.export.dim.DYNAMIC (#153065)sym_float, sym_not, sym_min, sym_max (#153200, #152111, #152196)TensorLR variant for fused Adagrad on CPU (#153078)lr_lambda type check in MultiplicativeLR (#151973)torch.AcceleratorError (#152023)Size.__radd__() (#152554)get_default_device() to also respect torch.device context manager (#148621)mul / add / add_relu and batch_norm2d), qconv1d-relu fusion, and lowering pass (#151112, #152411, #152811, #150751, #149708)torch.fused_moving_avg_obs_fake_quant on CUDA (#153699)cpp_extension (#152432)mm/bmm/addmm (#153262)PrivateUse1 extension (#149374)torch.Tensor.scatter_add_ (#150543), torch.matrix_exp (#155202)embed_cubin and multi_arch_kernel_binary options in AOTI for Intel GPU (#154514, #153924)UserDefineClass (#155787)CMake-4.x (#150203)gcc-12+ (#150847)/permissive- flag (#149035)torch.norm for scalar input (#144073)log_softmax reduced-precision fp kernel (#156379)torch.backends.cuda.matmul.allow_fp16_accumulation crash when using cuBLASLt (#153083)AsyncMM on Blackwell (#153519)torch.cuda.MemPool for multithreaded use-cases (#153356)sum() on a default-constructed gamma / beta in layer_norm (#156600)empty_cache under mempool context (#158180)all_to_all (#149485)group input argument in new_subgroups() (#152765, #153798)broadcast_object util function (#155912)DDPOptimizer issue on static tensor index (#155746)local_map with multi-threading (#149070)new_local_tensor in redistribute be None case (#152303)TensorPipe (#154382)gather when a local tensor on certain ranks has zero elements (#150914)dict(mapping_proxy), and the FlexAttention HOP (#157754, #157515, #157519)lru_cache method (#158689, #157308)TORCH_LOGS argument is passed (#151678)aten.is_nonzero (#149637), torch.bincount() (#152497), aten.div (#150874) slicing (#150104), and attn_mask (#158618), aten.to (#153972), scalar tensor construction (#154661)dynamic_shapes spec for kwargs (#148772, #149528, #150103)functools.partial (#153408), and higher order ops (#149295)None inputs (#150515), math module (#154643), call_torchbind (#155647), and enums (#154821)update_constant_buffer issue (#149243)model_package_loader (#152334)AOTIModel if they don't exist (#152692)ConstantFolding (#153152)dot and gemv (#152676)torch.lobpcg to compute same largest eigenvalue as scipy and np.linalg.eig (#152789)ReducedPrecisionGemV (#150949)2**32+ element inputs, binary ops with inputs with different dtypes, ops with complex scalar inputs, cholesky decomp, floor_divide type promotion, index_kernel with large inputs, lerp with complex inputs, logit with half/bfloat16 inputs, SDPA memory leak, torch.special.entr, tri[ul], matrix inversion with N>1024, and where with non-contiguous cond (#152479, #155183, #149233, #151176, #151282, #158239, #152371, #149974, #158237, #146754, #158867, #155184, #152204)load_state_dict behavior for nn.LazyLinear (#147599)onnx_program callable (#151121)lr_scheduler unexpectedly calls step() when init argument last_epoch > -1 (#149312)CosineAnnealingWarmRestarts resetting T_cur (#151289)MixtureSameFamily distribution (#151317)Wishart or Uniform distribution modifies constraints on the first (#154361)torch::utils::tensor_to_numpy symbol (#154178)torch.[con]cat[enate] to avoid crashing on empty inputs (#155460)torch.tensor and torch.ops.aten.scalar_tensor behavior (#158655)ScaledGEMM (#149677)ScaledGEMM (#152403)torch.is_vulkan_available() on Mac (#155595)offset > 0 (#154495)torch.xpu.is_bf16_supported to correctly report presence of Intel GPU (#152317)ELU(0) with the cheaper definition (#155765)cat and index_select (#150233, #152380, #151715)SubsetRandomSampler by iterating over list instead of tensor (#149126)cpp.use_small_dequant_buffer to use a small dequant buffer for WOQ int4 GEMM (#156395)torch.dot with float16/bfloat16 (#152799)LayerNorm, mm / bmm, sum / prod reductions, arithmetic ops,
binary kernels, SDPA, linear, and cumsum / cumprod (#152010, #150541, #150566, #147644, #149730, #152781, #152210, #157494)torch.tensordot when contracting to a scalar (#145936)softmax, NLLLoss, in-place sum, max pooling backward / reductions on NHWC
inputs, max pooling, multi-dimensional reductions, and non-vectorized elementwise kernels (#149076, #149779, #149548, #151230, #152267, #154522, #154619, #155806, #153184)HipSparseLT to further accelerate semi-structured (e.g. 2:4) sparsity (#150578)addmm, baddmm to reduce oneDNN integration overhead on Intel GPU (#153051)ctx.save_for_backward is important in note about extending autograd (#153005)torch.autograd.graph.saved_tensors_hooks to avoid refcycle (#153049)torch.amin and torch.amax (#155071)NCCLConfig with QOS variable (#151821)get_default_backend_for_device (#158236)ignored_params docstring and added unit tests (#149074)Dims and ExportGraphSignature (#156262, #156244)torch.linalg.norm()'s ord argument of +2 & -2 (#155148)nn.RNN, nn.functional loss functions, interpolate saturate cast behavior, ConvTranspose2d stride / output_size arguments, and register_full_backward_hook (#155123, #153620, #148436, #151304, #150819, #150609, #151785)nn.Sequential and nn.LazyModuleMixin (#147304, #150596)nn.modules.padding and AvgPoolND (#155618, #152680)LRSchedulers (#149189)CosineAnnealingLR to accurately reflect its recursive learning rate schedule (#152936)Adafactor documentation (#145209)load_state_dict hint doc about invoke order work with lr_scheduler (#149942)torch.Library's kind have no default value to be consistent with the code (#149390)requires_grad=True in tensor.to() (#150913)cdist param description (#151178)Example: and not Example:: in docs (#153978)as_strided() docs (#149146)keepdim param optional description (#151197)torch.trapezoid docs (#151190)out_dtype arg for torch GEMM operations (#151704)torch.min(), torch.max(), torch.all(), and torch.any() (#152658)torch.triu_indices, torch.tril_indices dtype description (#150749)torch.equal description (#149618)get_default_qat_qconfig in prepare_qat_fx docs (#155100)nccl_version and thread name/id, for flight record in PGNCCL (#150356, #150513, #151048, #152648, #155142, #155754)new_subgroups() for Non-Divisible World Sizes (#154124)get_backend() with more details (#141796)FlatParamHandle (#151336)rpc_init to CPython (#154325)torch.distributed.run option to provide destination for event logging (#155268)TracingContext (#149294)detect_attr_assignment (#151824)AOTInductor runtime API for Intel GPU (#153929)stable::Tensor is_contiguous API (#156228)lr_scheduler.py (#151219)step() with default value (#153367)setup-python from for Mac tests (#155698)This release is meant to fix the following issues (regressions / silent correctness):
This release is meant to fix the following issues (regressions / silent correctness):
Fix Excessive cudagraph re-recording for HF LLM models (#152287) Fix torch.compile on some HuggingFace models (#151154) Fix crash due to Exception raised inside torch.autocast (#152503) Improve Error logging in torch.compile (#149831) Mark mutable custom operators as cacheable in torch.compile (#151194) Implement workaround for a graph break with older version einops (#153925) Fix an issue with tensor.view(dtype).copy_(...) (#151598)
Fix assertion error due to inductor permuting inputs to flex attention (#151959) Fix performance regression on nanogpt speedrun (#152641)
Fix extra CUDA context created by barrier (#149144) Fix an issue related to Distributed Fused Adam in Rocm/APEX when using nccl_ub feature (#150010) Add a workaround random hang in non-blocking API mode in NCCL 2.26 (#154055)
Fix MacOS compilation error with Clang 17 (#151316) Fix binary kernels produce incorrect results when one of the tensor arguments is from a wrapped scalar on MPS devices (#152997)
Improve PyTorch Wheel size due to introduction of addition of 128 bit vectorization (#148320) (#152396) Fix fmsub function definition (#152075) Fix Floating point exception in torch.mkldnn_max_pool2d (#151848) Fix abnormal inference output with XPU:1 device (#153067) Fix Illegal Instruction Caused by grid_sample on Windows (#152613) Fix ONNX decomposition does not preserve custom CompositeImplicitAutograd ops (#151826) Fix error with dynamic linking of libgomp library (#150084) Fix segfault in profiler with Python 3.13 (#153848)
Users should move to use the dynamo=True option on torch.onnx.export as torch.onnx.dynamo_export is now deprecated. Leverage the `dynamic_shapes` argu…
<table> <tr> <td><strong>Beta</strong> </td> <td><strong>Prototype</strong> </td> </tr> <tr> <td>Torch.Compile support for Torch Function Modes </td> <td>NVIDIA Blackwell Architecture Support </td> </tr> <tr> <td>Mega Cache </td> <td>PyTorch Native Context Parallel </td> </tr> <tr> <td> </td> <td>Enhancing Intel GPU Acceleration </td> </tr> <tr> <td> </td> <td>FlexAttention LLM <span style="text-decoration:underline;">first token processing</span> on X86 CPUs </td> </tr> <tr> <td> </td> <td>FlexAttention LLM <span style="text-decoration:underline;">throughput mode optimization</span> on X86 CPUs </td> </tr> <tr> <td> </td> <td>Foreach Map </td> </tr> <tr> <td> </td> <td>Flex Attention for Inference </td> </tr> <tr> <td> </td> <td>Prologue Fusion Support in Inductor </td> </tr> </table>
For more details about these highlighted features, you can look at the release blogpost. Below are the full release notes for this release.
Some users with 12.2 CUDA driver (535 version) report seeing "CUDA driver error: invalid argument" during NCCL or Symmetric Memory initialization. This issue is currently under investigation, see #150852. If you use PyTorch from source, a known workaround is to rebuild PyTorch with CUDA 12.2 toolkit. Otherwise, you can try upgrading the CUDA driver on your system.
py_limited_api=True is now built with -DPy_LIMITED_API (#145764)We formally began respecting the py_limited_api=True kwarg in 2.6 and stopped linking libtorch_python.so when the flag was specified, as libtorch_python.so does not guarantee using APIs from from the stable Python limited API. In 2.7, we go further by specifying the -DPy_LIMITED_API flag which will enforce that the extension is buildable with the limited API. As a result of this enforcement, custom extensions that set py_limited_api=True but do not abide by the limited API may fail to build. For an example, see #152243.
This is strictly better behavior as it is sketchy to claim CPython agnosticism without enforcing with the flag. If you run into this issue, please ensure that the extension you are building does not use any APIs which are outside of the Python limited API, e.g., pybind.
torch.Tensor.new_tensor() to be on the given Tensor's device by default (#144958)This function was always creating the new Tensor on the "cpu" device and will now use the same device as the current Tensor object. This behavior is now consistent with other .new_* methods.
With Migration to manylinux_2_28 (AlmaLinux 8 based), we can no longer support OS distros with glibc2_26. These include popular Amazon Linux 2 and CentOS 7. (#143423, #146200, #148028, #148135, #148195, #148129)
torch.onnx.dynamo_export now uses the ExportedProgram logic path (#137296)Users using the torch.onnx.dynamo_export API may see some ExportOptions become
unsupported due to an internal switch to use torch.onnx.export(..., dynamo=True): diagnostic_options, fake_context and onnx_registry are removed/ignored by ExportOptions. Only dynamic_shapes is retained.
Users should move to use the dynamo=True option on torch.onnx.export as
torch.onnx.dynamo_export is now deprecated. Leverage the dynamic_shapes argument in torch.onnx.export for specifying dynamic shapes on the model.
Version 2.6.0
torch.onnx.dynamo_export(model, *args, **kwargs)
Version 2.7.0
torch.onnx.export(model, args, kwargs=kwargs, dynamo=True)
LRScheduler.print_lr() along with the verbose kwarg to the LRScheduler constructor. (#147301)Both APIs have been deprecated since 2.2. Please use LRScheduler.get_last_lr() to access the learning rate instead.print_lr and verbose were confusing, not properly documented and were little used, as described in #99270, so we deprecated them in 2.2. Now, we complete the deprecation by removing them completely. To access and print the learning rate of a LRScheduler:
Version 2.6.0
optim = ...
lrsched = torch.optim.lr_scheduler.ReduceLROnPlateau(optim, verbose=True)
// lrsched will internally call print_lr() and print the learning rate
Version 2.7.0
optim = ...
lrsched = torch.optim.lr_scheduler.ReduceLROnPlateau(optim)
print(lrsched.get_last_lr())
Previously, the symbols in libtorch_python.so were exposed with default visibility. We have transitioned to being more intentional about what we expose as public symbols for our python API in C++. After #142214, public symbols will be marked explicitly while everything else will be hidden. Some extensions using private symbols will see linker failures with this change.
torch.export.export instead of capture_pre_autograd_graph to export the model for pytorch 2 export quantization (#139505)capture_pre_autograd_graph was a temporary API in torch.export. Since now we have a better longer term API: export available, we can deprecate it.
Version 2.6.0
from torch._export import capture_pre_autograd_graph
from torch.ao.quantization.quantize_pt2e import prepare_pt2e
from torch.ao.quantization.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer,
get_symmetric_quantization_config,
)
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config()
)
m = capture_pre_autograd_graph(m, *example_inputs)
m = prepare_pt2e(m, quantizer)
Version 2.7.0
from torch.export import export
from torch.ao.quantization.quantize_pt2e import prepare_pt2e
# please get xnnpack quantizer from executorch (https://github.com/pytorch/executorch/)
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer,
get_symmetric_quantization_config,
)
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config()
)
m = export(m, *example_inputs)
m = prepare_pt2e(m, quantizer)
torch.fx.passes.graph_transform_observer.GraphTransformObserver to enable Node Level provenance tracking (#144277)We now track a mapping between the nodes in the pre-grad and post-grad graph. See the issue for an example frontend to visualize the transformations. To update your GraphTransformObserver subclasses, instead of overriding on_node_creation and on_node_erase, there are new functions get_node_creation_hook, get_node_erase_hook, get_node_replace_hook and get_deepcopy_hook. These are registered on the GraphModule member of the GraphTransformObserver upon entry and exit of a with block
Version 2.6.0
class MyPrintObserver(GraphTransformObserver):
def on_node_creation(self, node: torch.fx.Node):
print(node)
Version 2.7.0
class MyPrintObserver(GraphTransformObserver):
def get_node_creation_hook(self):
def hook(node: torch.fx.Node):
print(node)
return hook
torch.ao.quantization.pt2e.graph_utils.get_control_flow_submodules is no longer public (#141612)We are planning to make all functions under torch.ao.quantization.pt2e.graph_utils private. This update marks get_control_flow_submodules as a private API. If you have to or want to continue using get_control_flow_submodules, please make a private call by using _get_control_flow_submodules.
Example: Version 2.6:
>>> from torch.ao.quantization.pt2e.graph_utils import get_control_flow_submodules
Version 2.7:
>>> from torch.ao.quantization.pt2e.graph_utils import get_control_flow_submodules
ImportError: cannot import name 'get_control_flow_submodules' from 'torch.ao.quantization.pt2e.graph_utils'
>>> from torch.ao.quantization.pt2e.graph_utils import _get_control_flow_submodules # Note: Use _get_control_flow_submodules for private access
torch.onnx.dynamo_export is deprecated (#146425, #146639, #146923)Users should use the dynamo=True option on torch.onnx.export.
Version 2.6.0
torch.onnx.dynamo_export(model, *args, **kwargs)
Version 2.7.0
torch.onnx.export(model, args, kwargs=kwargs, dynamo=True)
XNNPACKQuantizer is deprecated in PyTorch and moved to ExecuTorch, please use it from executorch.backends.xnnpack.quantizer.xnnpack_quantizer instead of torch.ao.quantization.quantizer.xnnpack_quantizer. (#144940)XNNPACKQuantizer is a quantizer for xnnpack that was added into pytorch/pytorch for initial development. However, as it is not related to our core quantization workflow, we have moved it to ExecuTorch instead. Please use it from executorch.backends.xnnpack.quantizer.xnnpack_quantizer instead of torch.ao.quantization.quantizer.xnnpack_quantizer.
Version 2.6.0
from torch._export import capture_pre_autograd_graph
from torch.ao.quantization.quantize_pt2e import prepare_pt2e
from torch.ao.quantization.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer,
get_symmetric_quantization_config,
)
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config()
)
m = capture_pre_autograd_graph(m, *example_inputs)
m = prepare_pt2e(m, quantizer)
Version 2.7.0
# we also updated the export call
from torch.export import export
from torch.ao.quantization.quantize_pt2e import prepare_pt2e
# please get xnnpack quantizer from executorch (https://github.com/pytorch/executorch/)
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer,
get_symmetric_quantization_config,
)
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config()
)
m = export(m, *example_inputs)
m = prepare_pt2e(m, quantizer)
torch.utils.serialization.config namespace for all serialization related configurations (#143324)torch.serialization.config.save.use_pinned_memory_for_d2h to speed up torch.save when passed gpu devices (#143342)torch.utils.serialization.config.load.calculate_storage_offsets to reduce random reads and significantly improve performance for storage with bad random access performance (#143880)__torch_function__ handler on dtype arguments, similar to subclass objects (#145085)torch.nn.functional.scaled_dot_product_attention over the sequence dimension. We implemented
Ring Attention (#131351) and an AllGather-based approach (#132820) where the all-gather is issued before the first local SDPA
and the subsequent local SDPAs will have to wait until the all-gather completes, and offered a user API (#142093) to select the desired approach. The implementation
currently supports three SDPA kernels: SDPBackend.FLASH_ATTENTION, SDPBackend.EFFICIENT_ATTENTION, and SDPBackend.CUDNN_ATTENTION (#148537). We also
verified that our Context Parallel implementation is compatible with other parallelisms and torch.compile.torch.compile (#145270)torch.cuda.gds APIs public (#147120)torch.compile on Windows Platform for XPU (#147637, #144316, #149511)torch.utils.cpp_extension APIs (#132945)contextlib.contextmanager in Dynamo (#136033)nonstrict_trace escape hatch to apply non-strict tracing to difficult-to-compile code (#146367)list subclasses (#146819)head_dim for FlexAttention (#133495).num_warps and num_stages (#139639).ConfigFuzzer: a new debugging tool designed to fuzz Torch compile configurations. Given a test function, it will identify combinations of configs that throw errors during compilation and execution (#139736) (#145565).TORCHINDUCTOR_PROLOGUE_FUSION enables this feature (#147008).TORCHINDUCTOR_CUTLASS_INSTANTIATION_LEVEL. Consult config.py for information (#146230).cuda.cutlass_max_profiling_swizzle_options (#146088).package_cpp_only is specified in AOTI (#143352).graph_partition functions. Set the graph_partition in inductor config to enable (#147038).experimentalConfig (#143659)torch.ops.aten._dyn_quant_matmul_4bit, while the weights, scaled and optional bias are packed in torch.ops.aten._dyn_quant_pack_4bit_weight. To use it on your model you can quantize it using the following example that leverages torchao:from torchao.dtypes import PlainLayout
from torchao.experimental.packed_linear_int8_dynamic_activation_intx_weight_layout import (
PackedLinearInt8DynamicActivationIntxWeightLayout,
)
from torchao.experimental.quant_api import (
int8_dynamic_activation_intx_weight,
)
from torchao.quantization.granularity import (
PerGroup,
PerRow,
)
from torchao.quantization.quant_api import quantize_
from torchao.quantization.quant_primitives import MappingType
my_model = Model()
quantize_(
my_model,
int8_dynamic_activation_intx_weight(
weight_dtype=torch.int4,
granularity=PerGroup(32), # PerRow() is also supported
has_weight_zeros=True, # Should be True
weight_mapping_type=MappingType.SYMMETRIC_NO_CLIPPING_ERR # MappingType.SYMMETRIC can also be used but increases error
layout=PackedLinearInt8DynamicActivationIntxWeightLayout(target="aten"),
),
)
torch.onnx.verification.verify_onnx_program (#148396, #148706, #148730, #148707)A new verification API torch.onnx.verification.verify_onnx_program can now be used to verify numerical accuracy of the exported ONNX model. Users can use the compare_intermediates option to identify any operator that causes numerical discrepancies in intermediate tensors. It is possible to use a tool like model-explorer to visualize the verification results.
dynamic_shapes (#146321)torch.onnx.export(dynamo=True) now optimizes the output model by default (#146187)torch.addcmul (#143264)-DPy_LIMITED_API flag for py_limited_api=True cpp_extensions (#145764)torch.jit.load (#143403)torch.save configurable (#147788)with statement on torch.Stream (#140138)torch.autograd.graph.GradientEdge as torch.autograd.backward outputs #144744residuals of torch.linalg.lstsq #148526reflection_pad2d_backward (#136241)in_order is False (#142324)device argument. device and pin_memory_device are discouraged and will be deprecated in the future. (#131858)torch.cum{min,max}. (#143920)chunk() backward on batch dim (#144584)*_like factory functions for NJT (#144889)matmul with NJTs via backward support and composition with dense tensors (#144587, #146405)strict kwarg to nn.Module.set_submodule and fix bug for non dot-delineated strings (#143455)reflection_pad1d, reflection_pad2d and reflection_pad3d (#141670)isAcceleratorExcluded (#144959)abort and shutdown by adding both to Backend and ProcessGroup objects (#148798)new_group instead of split_group on non-CUDA device (#141469)call_guard in pybind object init of c10d (#143598)getDefaultBackend more fault tolerant (#148596)init_sync option to control collectives during initialization (#142824)reduce_dtype in lazy init (#143297)aten.amin/amax to linear_reduction_strategy (#143747)src_data_rank to distribute_tensor API (#143883)_scaled_mm (#143760)aten.view.dtype op support (#144404)shard_dim_alltoall to use alltoall_single (#148868)_shard_tensor to use src_data_rank=None (#144171)aten.minimum (#145816)src_data_rank kwarg in TP API (#144005)etcd_rendezvous publicly importable (#145396)generate_stage_to_rank_mapping utility (#146193)stage_index_to_group_rank from schedule (#146217)brgemm (#143384)sharedMemPerMultiprocessor device property to python (#143119)cudaDeviceProps to python (#143226)index >= 0 of cuda device (#140791)get_stream_from_external API for CUDA backend (#143799)angle, entr, spherical_bessel_j0,xlog1py, sinc,round.decimals, linalg.det, cholesky.ex, bilineard2d_aa,linalg.solve, zeta, cholesky, fused_rms_norm, lu_unpack, lu_factor_ex, slogdet and logdet (#143449, #147948, #146818, #147687, #146539, #147266, #146279, #146799, #145526, #146531, #146465, #145701, #145301, #146681, #144651, #145341, #146771, #147914)angle and atan2 for long type, torch.special.sinc to complex, torch.mm / torch.bmm to integral types (#149017, #146648, #145809, #147526)torch.accelerator.synchronize() on MPS (#143171)gamma, zeta, sinc, spherical_bessel_j0, entr (#145341, #146465, #146539, #147650, #148128)convolution_backward output layout between fake tensor and real output tensor (#146880)torch.xpu.get_device_properties API error message (#144379)nested_layer_norm support for XPU (#148593)is_big_gpu() check in Inductor (#143491)dict subclasses (#143548)trace_rules.py skipfiles (#145856)transformers ModelOutput) (#143567)buffer.copy_(int) (#141161)torch.inference_mode (#147925)topk (#147017)interpolate(antialias=True) backward (#141198)nonzer_static (#146006)_compute_symbolic_stride() (#138844)torch._check (#144471)backed_size_oblivious config (#148696)mark_unbacked strict mode (#147333, #147342)Several operator decomps received improvements/bugfixes:
torch._refs.tensor (#143461)torch._refs.mean (#147188)linspace (#147997)addmv (#143792)
New meta tensor implementations for a few pytorch operators:nonzero (#144727)silu, sigmoid, _softmax, embedding (#147862)
New fake tensor implementation for a few pytorch operators:unique_consecutive (#145649)
Several general FakeTensor improvementsUntypedStorage.from_buffer(buf) to return meta storage under FakeTensorMode (#146642)meta_tensor.to(device='cpu') under fake_mode (#146729)swizzle (#147223).INDUCTOR_CPP_ENABLE_FLOATING_POINT_CONTRACT_FLAG will be passed to ffp-contract (#143450).TORCHINDUCTOR_WORKER_START to one of "subprocess", "fork", or "spawn" (#144491).do_bench (#133058).emulate_precision_casts: TORCHINDUCTOR_EMULATE_PRECISION_CASTS (#145948).TORCHINDUCTOR_CUTLASS_ALLOWLIST and TORCHINDUCTOR_CUTLASS_DENYLIST (#148161).TORCHINDUCTOR_SCALAR_ASSERTS (#146462).layout_optimization and comprehensive_padding (#148450).AOT_INDUCTOR_COMPILE_WRAPPER_WITH_O0=1 (#144866).global_scratch arg, fix cpp_wrapper (#148051, #149973)._int_mm in AOTI (#144571).torch.ops.aten._assert_tensor_metadata.default for AOTI (#145028).aot_compile and aoti_compile_and_package (#148506)._weight_int4pack_mm_cpu_tensor (#149031)"+export" logging to de/serialization process (#145283)builtins.getattr with serializable higher-order-op for tensor subclasses (#145772)Dim.AUTO) for TorchScript -> Export Converter (#138273)ShapesCollection (#147534)expression_created logging in draft export (#146859)report as return output for draft export, attached as ep._report (#147558)builtin bitshift ops in verifier (#145802)aoti_call_delegate higher-order-op for eager-mode runnability (#145630)FlatArgsAdapter (#146107)upsample_bilinear2d.vec, nearest2d.vec from default decomposition table (#147153)keep_original_weights in _lower_to_native_backend (#141049)torch.float8_e8m0fnu dtype to PyTorch (#147466)dynamic_axes to dynamic_shapes with torch.export.Dim.AUTO (#143158)torch.cdist into onnx and support 'compute_mode' (#144213)LegacyDynamoStrategy (#145442)torch.onnx call site (#147165)dynamic_shapes renaming (#147407)run_decomposition() (#148617)torch export to get dynamic_shapes for JIT convert strategy (#148627)torch.export.Dim.AUTO in dynamo_export (#144356)verify=True (#148619)torch.lerp type promotion (#141117)torch.Tensor when both slots and python gc are used (#143203)torch.bfloat16 support for __cuda_array_interface__. (#143042)torch.Tensor incorrect. (#145530)torch.load under FakeTensorMode to create FakeTensor with correct devices (for plain Tensors) (#147786)torch.acos, torch.asin, torch.atan, torch.exp, torch.sigmoid, torch.div, for torch.complex datatypes on CPU (#134838, #140358, #140391, #140375, #144749)torch.autograd.graph.allow_mutation_on_saved_tensors for inplace foreach ops #145520hardswish backward (#143899)batchnorm backward on CPU (#147353)torch.compile + ddp + non-reentrant AC pack hook firing count (#144271)eigh (#146456)min / max backward() for non-ragged reductions (#144583)frexp() to handle both outputs (#144585)fill.Scalar for contiguous inputs (#144586)#pragma diagnostic pop in VecLib (#148354)CudaEventCache for dangling references (#144496)all-reduce input contiguous in distributed.nn.all_reduce (#144267)Alltoallv specialization for PyTorch generic all_to_all (#145045)torch.device (#146290)dist.init_process_group on windows (#148266)isend and irecv (#148462)strict=False case for DDP (#143038)buffer_dtype without root parameters (#143989)torch.distributed._functional_collectives.AsyncCollectiveTensor for aten.to. (#134661)_scaled_dot_product_flash_attention sharding (#148125)all-reduce (#148761)asinh codegen (#142360)Conv/Linear + broadcast add fusion (#141759)PYTORCH_NO_CUDA_MEMORY_CACHING has effect only when value is 1 (#145905)complex128 scan (#143401)Int64 indexing fix for UpSampleNearest3D (#144865)avg_pool2d backward for SM 10.0 to prevent runtime crash (#145669)f8f8bf16 rowwise scaled matmul to SM 9.0 (precedes #148421 adding of kernel) (#145728)Upsample2D (#141923)_preload_cuda_deps (#149808)sm100, sm120 (#150640)gather_out in MPS backend (#135543)torch.add(x,y, alpha=2) crash (#143949)nllnd_loss_backward crash with different dtypes (#144170)_scaled_dot_product_attention_math_mps (#146623)cholesky_ex for empty inputs (#147159)torch.chalf (#148285)unary_kernel_strided logic (#148512)lstm_mps causing leaked memory (#145503)c10::metal::log_gamma correctness on M4 (#145740)enable_gqa crash on MPS (#149147)tril op not handling infs correctly (#149866)min/max reductions over large dims (#149004)masked/where for inf values (#144500)addmm op (#144519)torch.layer_norm invalid configuration when input is large tensor (#144007)isnan integer overload errors on MicroSoft STL (#146605)torch.utils._content_store to accelerate XPU tensor hashing for tensor serialization (#147785)OffsetBasedRNGTracker to unbreak torch.distributed (#148360)torch.backends.mkldnn.flags() CM should not warn (#150358)step() for iterative on-demand tracking behind environment variable (#144494)torch.profiler (#144237)torch._functorch import (#149683)torch.compile calls was ignored (#145131).nanj in cpp wrapper CPU (#144064).fractional_max_pool lowering in Inductor (#144395).associative_scan (#143048).torch.polygamma(n) when n == 0 (#144058).avg_pool that was causing 0 rounding (#144059).avg_pool with uint to match eager (#144313).torch.logit decomposition (#145576).benchmark_harness isn't generated, but is called in some cases (#145532).SVE256 features were run on SVE128 systems (#146207).mm_template (#146293).cpp_wrapper (#145527).AOTIModelPackageLoaderPybind::boxed_run (#146100).None and equal_to_1 arguments issue in Triton kernel generated by AOTI (#148102)AOTIModelPackageLoader() constructor defaults (#149082)get_source_partitions when weights are tied (#142446)and_ operator (#145506)unbacked_bindings (#144894)unbacked_bindings keys (#145777)unbacked_bindings in deserialization (#145882)nn_module_stack (#145901).requires_grad field (#146351)math.trunc ops for serialization (#146715)math.inf and NaN as strings (#146490)lazy_trace_handler bug in draft export logging (#146106)._modules corner case for nn_module_stack metadata in strict-mode (#142823)all_reduce, all_gather, all_gather_into_tensor, all_to_all_single, reduce_scatter_tensor) in non-strict mode (#147133, #147417)stack_trace field optional in insert_custom_op_guards pass (#146438)ScriptModules and ScriptObjects for TorchBind (#147399)export_tracepoint (#148709)transpose_ (#149057)rename_dynamic_shapes_with_model_inputs (#146002)dynamic_shapes string cases (#148025)clones throughout codebase (#148159)ALLOC_BUFFER_SIZE from 4000 to 4096 to be a power of 2 for TCPStore (#145759)ExpandableSegments object and the GIL in WorkNCCL destructor (#148805)fmadds (#144486)sort (#142391)prop_kind to forward_inference when grad is not needed for mkldnn_convolution_pointwise (#142855)add and max (#144065)conv src zp mask (#149473)PYTORCH_NO_CUDA_MEMORY_CACHING has effect only when value is 1 (#145905)complex128 scan (#143401)Int64 indexing fix for UpSampleNearest3D (#144865)avg_pool2d backward for SM 10.0 to prevent runtime crash (#145669)f8f8bf16 rowwise scaled matmul to SM 9.0 (precedes #148421 adding of kernel) (#145728)Upsample2D (#141923)masked_fill_scalar as shader (#147369)bilineard2d as shader (#145581)_load_dwordx4 ISA for BFloat16 and Half (#141397)persistent_tma (#142101)._weight_int4pack_mm_for_cpu (#146756).SmoothQuant int8 linear pattern (#142036).max_pool2d_with_indices as a reduction for large window sizes (#147876).Graph.nodes.__iter__ (#144631)map_aggregate(immutable_dict) (#147691)input in torch.addbmm() (#146664)torch.cat type promotion documentation (#141339)torch.topk indices stability when duplicate values (#143736)torch.diagonal documentation (#144214)torch.{min,max} documentation (#146725)torch.max optional args dim, keepdim description (#147177)torch.bucketize documentaion (#148400)torch.autograd.gradcheck #144287CrossEntropyLoss doc (#145444)set_requires_gradient_sync and no_sync (#148715)wait() (#143305)from_group API and add a 2D test (#146364)__create_chunk_list__ in the doc (#144100)set_optimizer_state_dict (#148918)record (#146968)suppress_errors on compiler error (#146553)emulate_precision_casts (#145579).inductor.config.descriptive_names = False is no longer a suggested option (#145523).dims() and suggested fixes parser: #142510torch.export.load(): #141490__add__ and __mul__ hints to torch.Size (#144322)fully_shard so that the return value can be chained with typing enabled (#147489)add_, addcmul, arange, baddbmm, bmm, clamp, div, div_, gelu, index_add, logical_and, mul_, sub_, topk, where} to operator benchmark (#145625)torch._dynamo.optimize() with torch.compile() (#142451)TORCH_LOGs under the name "autotuning" (#147222).set by OrderedSet: only use OrderedSet in the Inductor codebase (#138466).GPU_TYPE (#143634).qlinear (#143903).TORCH_COMPILE_DEBUG to TORCH_LOGS. For example, TORCH_LOGS="+ir_pre_fusion" (#147248).deepcopy for AOTICompiledModel (#145423)fx.node.map_arg() and .map_aggregate() generic (#146248)Your coding agent can read these notes before it upgrades. Set up the MCP server →