NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #358 most downloaded on PyPI
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Last release 19 days ago
02 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 49 of 50 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
50 releases · first in 2018
However, this change can potentially be backward incompatible since there may be small numerical differences between the results computed with the for…
We are excited to announce the release of PyTorch® 2.0 (release note) which we highlighted during the PyTorch Conference on 12/2/22! PyTorch 2.0 offers the same eager-mode development and user experience, while fundamentally changing and supercharging how PyTorch operates at compiler level under the hood with faster performance and support for Dynamic Shapes and Distributed.
This next-generation release includes a Stable version of Accelerated Transformers (formerly called Better Transformers); Beta includes torch.compile as the main API for PyTorch 2.0, the scaled_dot_product_attention function as part of torch.nn.functional, the MPS backend, functorch APIs in the torch.func module; and other Beta/Prototype improvements across various inferences, performance and training optimization features on GPUs and CPUs. For a comprehensive introduction and technical overview of torch.compile, please visit the 2.0 Get Started page.
Along with 2.0, we are also releasing a series of beta updates to the PyTorch domain libraries, including those that are in-tree, and separate libraries including TorchAudio, TorchVision, and TorchText. An update for TorchX is also being released as it moves to community supported mode. More details can be found in this library blog.
This release is composed of over 4,541 commits and 428 contributors since 1.13.1. We want to sincerely thank our dedicated community for your contributions. As always, we encourage you to try these out and report any issues as we improve 2.0 and the overall 2-series this year.
Summary:
<table> <tr> <td> <strong>Stable</strong> </td> <td><strong>Beta</strong> </td> <td><strong>Prototype</strong> </td> <td><strong>Platform Changes</strong> </td> </tr> <tr> <td>Accelerated PT 2 Transformers </td> <td>torch.compile </td> <td>DTensor </td> <td>CUDA support for 11.7 & 11.8 (deprecating CUDA 11.6) </td> </tr> <tr> <td> </td> <td>PyTorch MPS Backend </td> <td>TensorParallel </td> <td>Python 3.8 (deprecating Python 3.7) </td> </tr> <tr> <td> </td> <td>Scaled dot product attention </td> <td>2D Parallel </td> <td>AWS Graviton3 </td> </tr> <tr> <td> </td> <td>Functorch </td> <td rowspan="2" >Torch.compile (dynamic=True) </td> <td> </td> </tr> <tr> <td> </td> <td>Dispatchable Collectives </td> <td> </td> </tr> <tr> <td> </td> <td>torch.set_default_device and torch.device as context manager </td> <td> </td> <td> </td> </tr> <tr> <td> </td> <td>X86 quantization backend </td> <td> </td> <td> </td> </tr> <tr> <td> </td> <td>GNN inference and training performance </td> <td> </td> <td> </td> </tr> </table>
*To see a full list of public 2.0, 1.13 and 1.12 feature submissions click here
Previously the minimum supported version of Python for PyTorch was 3.7. This PR updates the minimum version to require 3.8 in order to install PyTorch. See Hardware / Software Support for more information.
This PR updates the minimum CUDA version to 11.0. See the getting-started for installation or building from source for more information.
None instead of zeros by default in torch.optim.*.zero_grad() and torch.nn.Module.zero_grad() (#92731)This changes the default behavior of zero_grad() to zero out the grads by setting them to None instead of zero tensors. In other words, the set_to_none kwarg is now True by default instead of False. Setting grads to None reduces peak memory usage and increases performance. This will break code that directly accesses data or does computation on the grads after calling zero_grad() as they will now be None. To revert to the old behavior, pass in zero_grad(set_to_none=False).
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> import torch
>>> from torch import nn
>>> module = nn.Linear(2,22)
>>> i = torch.randn(2, 2, requires_grad=True)
>>> module(i).sum().backward()
>>> module.zero_grad()
>>> module.weight.grad == None
False
>>> module.weight.grad.data
tensor([[0., 0.],
[0., 0.]])
>>> module.weight.grad + 1.0
tensor([[1., 1.],
[1., 1.]])
</td> <td>
>>> import torch
>>> from torch import nn
>>> module = nn.Linear(5, 5)
>>> i = torch.randn(2, 5, requires_grad=True)
>>> module(i).sum().backward()
>>> module.zero_grad()
>>> module.weight.grad == None
True
>>> module.weight.grad.data
AttributeError: 'NoneType' object has no attribute 'data'
>>> module.weight.grad + 1.0
TypeError: unsupported operand type(s) for +:
'NoneType' and 'float'
</td> </tr> </table>
torch.tensor and nn.Parameter to serialize all their attributes (#88913)Any attribute stored on torch.tensor and torch.nn.Parameter will now be serialized. This aligns the serialization behavior of torch.nn.Parameter, torch.Tensor and other tensor subclasses
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
# torch.Tensor behavior
>>> a = torch.Tensor()
>>> a.foo = 'hey'
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>> print(a.foo)
hey
>>> print(b.foo)
AttributeError: 'Tensor' object has no attribute 'foo'
# torch.nn.Parameter behavior
>>> a = nn.Parameter()
>>> a.foo = 'hey'
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>> print(a.foo)
hey
>>> print(b.foo)
AttributeError: 'Parameter' object has no attribute 'foo'
# torch.Tensor subclass behavior
>>> class MyTensor(torch.Tensor):
... pass
>>> a = MyTensor()
>>> a.foo = 'hey'
>>> print(a.foo)
hey
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>>print(b.foo)
hey
</td> <td>
# torch.Tensor behavior
a = torch.Tensor()
a.foo = 'hey'
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>> print(a.foo)
hey
>>> print(b.foo)
hey
# torch.nn.Parameter behavior
>>> a = nn.Parameter()
>>> a.foo = 'hey'
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>> print(a.foo)
hey
>>> print(b.foo)
hey
# torch.Tensor subclass behavior
>>> class MyTensor(torch.Tensor):
... pass
>>> a = MyTensor()
>>> a.foo = 'hey'
>>> print(a.foo)
hey
>>> buffer = io.BytesIO()
>>> torch.save(a, buffer)
>>> buffer.seek(0)
>>> b = torch.load(buffer)
>>>print(b.foo)
hey
</td> </tr> </table>
If you have an attribute that you don't want to be serialized you should not store it as an attribute on tensor or Parameter but instead it is recommended to use torch.utils.weak.WeakTensorKeyDictionary
>>> foo_dict = weak.WeakTensorKeyDictionary()
>>> foo_dict[a] = 'hey'
>>> print(foo_dict[a])
hey
{Adadelta, Adagrad, Adam, Adamax, AdamW, ASGD, NAdam, RAdam, RMSProp, RProp, SGD} default to faster foreach implementation when on CUDA + differentiable=FalseWhen applicable, this changes the default behavior of step() and anything that calls into adadelta(...), adagrad(...), adam(...), adamax(...), adamw(...), asgd(...), nadam(...), radam(...), rmsprop(...), rprop(...), sgd(...) directly to use the foreach implementation instead of the for-loop for better performance. However, this change can potentially be backward incompatible since there may be small numerical differences between the results computed with the foreach implementation and the previous default. The foreach implementation will be the default only if the following conditions are met.
foreach, fused, or differentiable),torch.jit.is_scripting is False.When these conditions are satisfied, the implementation used will match the implementation used when one passes foreach=True. The user defined flag for foreach will NOT be overwritten in order to preserve user selections. For more details, check the documentation. There should be no significant differences between the results returned by these optimizers. To revert to the old behavior, say, for adam, pass in adam(..., foreach=False, ...) or initialize Adam with Adam(..., foreach=False, ...).
Pull Requests: #92306, #92716, #92723,#92724, #92726, #92727, #92728, #92715, #91896, #92730, #90865, #93184, #92181, #92923, #95415, #95818, #95811
torch.nn.utils.stateless.functional_call now respects tied weights (#90477)Assume a module has two tied weights, x and x_tied. Previously, invoking functional_call(module, parameters_and_buffers, args, kwargs=None, *, strict=False) with a parameter dictionary of only one of the tied weights would result in the other one(s) not being updated.
We’ve changed the behavior so that providing one of the tied weights in the parameter dictionary will update all other tied weights. If you would like the behavior in previous versions of PyTorch, please set tie_weights=False.
Please also see the related deprecation section "torch.nn.stateless.functional_call in favor of torch.func.functional_call".
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> class Foo(nn.Module):
... def __init__(self):
... super().__init__()
... self.x = nn.Parameter(torch.zeros([]))
... self.x_tied = self.x
...
... def forward(self, inp):
... return self.x + self.x_tied
>>> foo = Foo()
>>> params = {'x': torch.ones([])}
>>> result = functional_call(foo, params, torch.randn([]))
>>> print(result)
1.0
</td> <td>
>>> class Foo(nn.Module):
... def __init__(self):
... super().__init__()
... self.x = nn.Parameter(torch.zeros([]))
... self.x_tied = self.x
...
... def forward(self, inp):
... return self.x + self.x_tied
>>> foo = Foo()
>>> params = {'x': torch.ones([])}
>>> result = functional_call(foo,
... params,
... torch.randn([]),
... tie_weights=False)
>>> print(result)
1.0
</td> </tr> </table>
return_complex to be passed explicitly to torch.stft for real input (#86724)torch.stft takes an optional return_complex parameter that indicates whether the output should be a floating point tensor or a complex tensor. return_complex previously defaulted to False for real input tensors. This PR removes the default and makes return_complex a required argument for real inputs. However, complex inputs will continue to default to return_complex=True.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> a = torch.rand(1024)
>>> _ = torch.stft(a, n_fft=128)
</td> <td>
>>> t = torch.rand(1024)
>>> _ = torch.stft(t, n_fft=128, return_complex=False)
</td> </tr> </table>
torch.istft to be complex valuedtorch.istft no longer supports input in the form of real tensors
with shape (..., 2) to mimic complex tensors. Instead, convert
inputs to a complex tensor first before calling torch.istft.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> t = torch.rand(65, 33, 2)
>>> _ = torch.istft(t, n_fft=128, length=1024)
</td> <td>
>>> t = torch.rand(65, 33, 2)
>>> _ = torch.istft(t, n_fft=128, length=1024)
RuntimeError: istft requires a complex-valued input
tensor matching the output from stft with return_complex=True.
>>> t_complex = torch.view_as_complex(t)
>>> _ = torch.istft(t_complex, n_fft=128, length=1024)
</td> </tr> </table>
We now disable the costly component verification of torch.sparse_coo/csr/csc/bsr/bsc/compressed_tensor by default. The user can use the new check_invariants flag or torch.sparse.check_sparse_tensor_invariants to locally enable component verification. This allows users to constrain these costly checks to specific regions of their code and enables better overall performance. Previously users had no access to public constructors that disable these checks.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> i = [[0, 1, 1],
[2, 0, 5]]
>>> v = [3, 4, 5]
>>> s = torch.sparse_coo_tensor(i, v, (2, 3))
RuntimeError: size is inconsistent with
indices: for dim 1, size is 3 but found index 5
</td> <td>
>>> i = [[0, 1, 1],
[2, 0, 5]]
>>> v = [3, 4, 5]
>>> s = torch.sparse_coo_tensor(i,
... v,
... (2, 3),
... check_invariants=True)
RuntimeError: size is inconsistent with indices: for
dim 1, size is 3 but found index 5
>>> with torch.sparse.check_sparse_tensor_invariants():
... s = torch.sparse_coo_tensor(i, v, (2, 3))
...
RuntimeError: size is inconsistent with indices: for
dim 1, size is 3 but found index 5
</td> </tr> </table>
torch.testingHistorically, torch.testing exposed a lot of private and undocumented functionality publicly. The 2.0 release completes the deprecation cycle for the following items and removes them:
rand and randn (#87970)get_all_device_types (#87971)make_non_contiguous (#87973).grad() (#85849)This is a bug fix. Per the docs, hooks registered to Tensor should fire any time gradients are computed w.r.t. to that tensor. This change corrects the behavior to be consistent with the documentation. See documentation for more details about backward hooks execution..
2.0
a = torch.tensor(1., requires_grad=True)
b = a.clone()
b.register_hook(hook) # the hook registered here didn't fire before!
torch.autograd.grad(b.clone(), inputs=(b,))
grad_fn post-hooks can always observe the modifications to gradient by any grad_fn pre-hooks or hooks registered to Tensor, even if this is a leaf tensor (#85849)This corrects the behavior of hooks to be consistent with the documentation in the case where the tensor is a leaf tensor, i.e. the node is a grad accumulator node. See documentation for more details about backward hooks execution.
2.0
def hook(grad):
# updates grad
return grad * 3
def hook2(grad_input, grad_output):
# Before this change, grad_output would NOT see the x3
print(grad_output)
a = torch.tensor(1., requires_grad=True)
b = a.clone()
acc_grad = b.grad_fn.next_functions[0][0]
acc_grad.register_hook(hook2)
b.register_hook(hook)
torch.autograd.backward(b.clone(), inputs=(a,)) # hook fire
params_with_grad (#87480)In FSDP, we used to have an API params_with_grad for users to get parameters which have gradients from the FSDP module. We decided not to expose this helper because it is not a common paradigm.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
m = FullyShardedDataParallel(module)
m.params_with_grad()
</td> <td>
m = FullyShardedDataParallel(module)
m.params_with_grad() # Runtime error thrown
# For work-around, users can still do
[p for p in self.parameters() if p.grad is not None]
</td> </tr> </table>
Users could previously import both public and non-public symbols:
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
from torch.distributed.fsdp.fully_sharded_data_parallel import *
ShardingStrategy.FULL_SHARD # Non-public API
FullyShardedDataParallel(module) # public API
</td> <td>
from torch.distributed.fsdp.fully_sharded_data_parallel import *
ShardingStrategy.FULL_SHARD # Non-public API, this will fail now
Fully`Sharded`DataParallel(module) # public API
...
# Users can instead
from torch.distributed.fsdp.fully_sharded_data_parallel import (
FullyShardedDataParallel,
ShardingStrategy,
)
FullyShardedDataParallel(module, sharding_strategy=ShardingStrategy.FULL_SHARD)
</td> </tr> </table>
auto_wrap_policy related APIs were changed in (#88450).<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
lambda_auto_wrap_policy(m, unwrapped_params=...)
transformer_auto_wrap_policy(m, unwrapped_params=...)
size_based_auto_wrap_policy(m, unwrapped_params=...)
</td> <td>
lambda_auto_wrap_policy(m, nonwrapped_numel=...)
transformer_auto_wrap_policy(m, nonwrapped_numel=...)
size_based_auto_wrap_policy(m, nonwrapped_numel=...)
</td> </tr> </table>
alltoall signature to be consistent with other c10d APIs (#90569)The keyword argument names have been changed.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
alltoall(output=..., input=...)
</td> <td>
alltoall(output_tensors=..., input_tensors=...)
</td> </tr> </table>
This commit removes the following unused functions from both the torch.quantization and the torch.ao.quantization namespaces:
graph_pretty_strget_per_tensor_qparamsquantize_nodeget_qconv_opcreate_qparam_nodesnode_return_type_is_intis_get_tensor_info_nodetorch.ao.quantization.backend_config.BackendConfig accept inputs in the right order (#90698)The existing BackendConfig fusion pattern uses a "reversed nested tuple" format that is unintuitive.
This pattern format also complicates the signatures of the user specified "fuser methods", which needed to accept arguments in reverse nested order to match
the patterns:
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
import torch as nn
import torch.ao.nn.intrinsic as nni
from torch.ao.quantization.backend_config import (
BackendPatternConfig
)
def fuse_linear_relu(is_qat, relu, bn_conv):
(bn, conv) = bn_conv
return nni.ConvBnReLU2d(conv, bn, relu)
config = (
BackendPatternConfig((nn.ReLU, (nn.BatchNorm2d, nn.Conv2d)))
.set_dtype_configs(...)
.set_fuser_method(fuse_conv_bn_relu)
.set_fused_module(nni.ConvBnReLU2d)
)
backend_config.configs # returns Dict[Pattern, BackendPatternConfig]
</td> <td>
def fuse_linear_relu(is_qat, conv, bn, relu):
return nni.ConvBnReLU2d(conv, bn, relu)
config = (
BackendPatternConfig((nn.Conv2d, nn.BatchNorm2d, nn.ReLU))
.set_dtype_configs(...)
.set_fuser_method(fuse_conv_bn_relu)
.set_fused_module(nni.ConvBnReLU2d)
)
# Or for backward-compatibility
def fuse_linear_relu(is_qat, relu, bn_conv):
(bn, conv) = bn_conv
return nni.ConvBnReLU2d(conv, bn, relu)
config = (
BackendPatternConfig()
._set_pattern_complex_format((nn.ReLU, (nn.BatchNorm2d, nn.Conv2d)))
.set_dtype_configs(...)
.set_fuser_method(fuse_conv_bn_relu)
.set_fused_module(nni.ConvBnReLU2d)
)
backend_config.configs # returns List[BackendPatternConfig]
</td> </tr> </table>
If users were using any of the AO private APIs then these would have to be accessed with a preceding _ to conform with the guidelines.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
get_observer_dict()
</td> <td>
_get_observer_dict()
</td> </tr> </table>
Pull Requests: (#86029, #87515, #87516, #87517, #87518, #87519, #88392, #88394, #88396, #88397, #87521, #88395, #87883, #88399, #88398, #86022, #86023, #86024, #86025, #86026, #86027, #86028, #86030, #86031, #86032, #86033, #86034, #86037, #90315, #88391, #90554, #87520)
This commit removes overwrite_output_observer and overwrite_output_fake_quantize overwrite observer settings in the BackendConfig. Instead, we represent the observer constraints for
fixed qparams ops through the existing DTypeWithConstraints mechanism. Note that, however, to be consistent with other DTypeWithConstraints checks, we no longer throw an error if an incorrect observer is specified, but simply ignore the offending QConfig and log a warning instead. This is the BC-breaking part of the change.
1.13
from torch.ao.quantization.qconfig import default_qconfig
from torch.ao.quantization.quantize_fx import prepare_fx
model = ModelWithFixedQParamsOps()
qconfig_mapping = QConfigMapping().set_global(default_qconfig)
example_inputs = ...
prepare_fx(model, qconfig_mapping, example_inputs)
Before this commit, running the above leads to an exception because the wrong observers are used for fixed qparams ops. After this commit, the above will only encounter a warning,and the fixed qparams ops will not be quantized. In both cases, switching to get_default_qconfig_mapping will cause the fixed qparams ops to be quantized.
torch.ao.quantization.quantization_patterns and torch.ao.quantization.fusion_patterns(#89872)The following classes under the torch.ao.quantization.fx.quantization_patterns namespace are migrated to the torch.ao.quantization.fx.quantize_handler
namespace:
QuantizeHandlerBinaryOpQuantizeHandlerCatQuantizeHandlerConvReluQuantizeHandlerLinearReLUQuantizeHandlerBatchNormQuantizeHandlerEmbeddingQuantizeHandlerRNNDynamicQuantizeHandlerDefaultNodeQuantizeHandlerFixedQParamsOpQuantizeHandlerCopyNodeQuantizeHandlerGeneralTensorShapeOpQuantizeHandlerCustomModuleQuantizeHandlerStandaloneModuleQuantizeHandlerThe following classes under the torch.ao.quantization.fx.fusion_patterns namespace are migrated to the torch.ao.quantization.fx.fuse_handler namespace:
DefaultFuseHandlerFuseHandlertorch.ao.quantization.fx.backend_config_utils namespace(#89810)The following APIs that were mistakenly public under the torch.ao.quantization.fx.backend_config_utils namespace are removed in this commit.
get_quantize_handler_clsget_fusion_pattern_to_fuse_handler_clsget_native_quant_patternsget_pattern_to_quantize_handlers<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
from torch.ao.quantization.fx.backend_config_utils import (
get_quantize_handler_cls,
get_fusion_pattern_to_fuse_handler_cls,
get_native_quant_patterns,
get_pattern_to_quantize_handlers,
)
all_quant_patterns = get_native_quant_patterns()
</td> <td>
from torch.ao.quantization.fx.quantization_patterns import (
_get_quantize_handler_cls,
_get_pattern_to_quantize_handlers,
)
from torch.ao.quantization.fx.fusion_patterns import (
_get_fusion_pattern_to_fuse_handler_cls,
)
from torch.ao.quantization.backend_config import (
get_native_backend_config,
)
all_quant_patterns = _get_pattern_to_quantize_handlers(
get_native_backend_config()
)
</td> </tr> </table>
These operators are primarily used by the functionalization pass, used in AOTAutograd. Previously, they would always return contiguous tensors. Now, they return a tensor with the same striding as their first argument.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> x = torch.ones(2, 2, 2)
>>> base = x[:, :, 1]
>>> base.stride()
(4, 2)
>>> x = torch.zeros(2, 2, 2)
>>> base = x[:, :, 1]
>>> base.stride()
(4, 2)
>>> torch.diagonal_scatter(base, torch.ones(2)).stride()
# returns a tensor with same strides as base.
(4, 2)
</td> <td>
>>> x = torch.ones(2, 2, 2)
>>> base = x[:, :, 1]
>>> base.stride()
(4, 2)
>>> x = torch.zeros(2, 2, 2)
>>> base = x[:, :, 1]
>>> base.stride()
(4, 2)
>>> torch.diagonal_scatter(base, torch.ones(2)).stride()
# returns a contiguous tensor
(2, 1)
</td> </tr> </table>
The Deprecated monkey patches to torch.Graph, torch.Block and torch.Node are removed
Monkey patches to the classes torch.Graph, torch.Block and torch.Node from torch.onnx have been removed. This means the methods torch.Graph.op(), torch..Graph.at(), torch.Block.op(), torch.Graph.constant(), and torch.Node.__getitem__ are no longer available.
Users creating custom symbolic functions for the torch.onnx exporter can continue to assume the g.op() interface for creating an operator in the graph, which is now exposed via the GraphContext class. Users should not assume any other methods from the GraphContext class other than those defined natively by torch.Graph and .op().
Code change to existing symbolic functions is not expected with this change.
This removes boolean value of full_check parameter in TORCH API check_onnx_proto, and forces full_check with warning messages if it fails.
Also, the API didn’t check on types in the graph even with full_check=True previously. With the change, a warning message will show if the graph contains type error.
torch::deploy has been migrated to over to MultiPy. Ongoing development will continue in this repository.
lazy::View (#87822)The view and aliasing infrastructure in lazy tensor core has been deprecated in favor of functionalization.
c10::fromIntArrayRef to c10::fromIntArrayRefSlow and changed call sites (#86235)The function has been renamed to more accurately reflect its performance characteristics.
We’re excited to announce that, as the final step of upstreaming and integrating functorch into PyTorch, the functorch APIs are now available in the torch.func module. Our function transform APIs are identical to before, but we have changed how the interaction with NN modules work.
We’ve deprecated functorch._ function transforms (e.g. vmap, grad, jvp) in favor of their identical torch.func._ counterparts (#92279).
PyTorch has consolidated on torch.func.functional_call as the NN module functional API. Please migrate from functorch.{make_functional, make_functional_with_buffers} to it. For more details see this Guide
Please migrate from functorch.combine_state_for_ensemble to torch.func.stack_module_state. For more details see this Guide
We are no longer supporting functorch.compile (also known as AOTAutograd) as a frontend for compilation in PyTorch; we have integrated AOTAutograd into PyTorch’s compilation story. If you are a user, please use torch.compile() instead.
Typed storages have been removed from the C++ side and torch.UntypedStorage is used in place. The use of torch.TypedStorage and all of its subclasses is now deprecated.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
tensor.storage()
torch.TypedStorage(...)
</td> <td>
tensor.untyped_storage()
torch.UntypedStorage(...)
</td> </tr> </table>
If you need to access individual elements in a storage as a particular dtype, you can simply create a tensor to view it:
torch.tensor(storage, dtype=...)
tensor.mT,tensor.T,tensor.mH,tensor.H on 0D-tensors (#92143)<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
>>> a = torch.tensor(10)
>>> a.T
>>> a.H
</td> <td>
>>> a = torch.tensor(10)
>>> a.T
UserWarning: Tensor.T is deprecated on 0-D tensors.
This function is the identity in these cases.
>>> a.H
UserWarning: Tensor.H is deprecated on 0-D tensors.
Consider using x.conj().
</td> </tr> </table>
Decorating classes with torch.no_grad is now deprecated. You should be decorating its functions or methods instead. To preserve the current behavior of class decoration, you can directly decorate the __init__ method and nothing else.
<table> <tr> <th>1.13</th> <th>2.0</th> </tr> <tr> <td>
@torch.no_grad()
class Blah():
pass
</td> <td>
class Blah():
@torch.no_grad()
def __init__(self):
pass
</td> </tr> </table>
In continuation with the deprecation process from release 1.12 the tensor overload for this function has been removed. This function was not used in the bindings of Pytorch and should not impact users of torch.norm.
functional.{tanh, sigmoid} functions (#86905)Both these ops are heavily used and so will not be removed. Deprecation warnings have been removed.
We’ve moved torch.nn.stateless.functional_call under the torch.func module to reflect how it is useful for working with nn.Modules in a functional style. As of PyTorch 2.0, torch.func.functional_call is a drop-in replacement for torch.nn.stateless.functional_call and we will remove torch.nn.utils.stateless.functional_call in a future version of PyTorch. However, please note that we did change the default behavior of torch.nn.stateless.functional_call in PyTorch 2.0 (see “torch.nn.utils.stateless.functional_call now respects tied weights” under BC-breaking notes).
Removed the Python 2 and 3 compatibility library six and future and torch._six. 2.0
# from torch._six import string_classes
str
# from torch._six import int_classes
int
# from torch._six import inf, nan
from torch import inf, nan
# torch._six.string_classes
str
Users must use PyTorch 1.x versions to use Caffe2 ONNX exporter. This capability will be completely removed from PyTorch 2.x series.
torch.nn.functional.scaled_dot_product_attention() to allow writing fast Transformer-like functions and use it to speed up nn.Transformer() ( #91362, #91066, #90413, #87312, #94008, #89470, #90776, #92189)Module.register_{buffer,module,parameter} functions (#86148, #87369)Module.full_backward_pre_hook (#86700)Module.state_dict_pre_hook (#90435)Module.call_super_init: bool flag that can be used to ensure Module initialization is properly calling parent’s __init__ (#91819)functorch support for torch.autograd.Function: one is now able to apply function transformations (e.g. vmap, grad, jvp) over torch.autograd.Function. (#92023, #91452, #91222, #90037, #90077, #90966, #89860, #91211, #92030)Logcumsumexp for complex dtypes for CUDA (build-time optimized) (#94310)set_to_none flag for C++ optim endpoint (#92989)tensor.to() for NestedTensor backend (#87146)gelu and relu operators (#94776)torch.neg operator (#88131)sharded_state_dict is added as well. (#87987, #88698, #89256, #89398, #89399, #89501, #89503, #89537, #89542, #89873, #89964, #90212, #91036, #91092, #91209, #91269, #92553, #92705, #92829, #92869, #92933, #94379, #94501)use_orig_params=True in the FSDP constructor (#84911)init_process_group API which changes backend to an optional argument. For users, this feature will allow for code that runs on both GPU and CPU machines without having to change the backend specification. The dispatchability feature will also allow users to perform both GPU and CPU collectives using the same ProcessGroup, as PyTorch will automatically find an appropriate backend for the tensor type (as of PyTorch 2.0, the default is NCCL for CUDA and Gloo for CPU). Existing backend specifications by users will be honored and will not require change (#83679, #83735, #83810, #83859, #83876, #83916, #84423, #86166, #86368, #86407, #86408, #86409, #88351, #88846, #88889, #88903, #89317, #89318, #89505, #89813, #88330, #91257, #91172)torch.nn.functional.group_norm(#91190), torch.var_mean (#91190), torch.nansym(#93845), torch.frac(#86625), torch.signbit(#87214), torch.exp1m(#87147), torch.cumsum(#88319), torch.trace(#87910), torch.nn.Hardswish (#87952),torch.inverse(#90428), torch.floor_divide(#91126), unfold(#91266), bincount(#91267), nonzero(#91616), norm_dtypeandcdist(#91643), uniqueandunique_consecutive(#88532), nan_to_num(#91110), torch.linalg.cross(#91642), randperm(#91708), triangular_solve(#94345), grid_sampler2d(#94273), remainder(#92139), addr(#94538), fmod(#94722), repeat_interleave (#88649),sortandargSort(#94697),range (#91075)torch.mps.{get_rng_state, set_rng_state, synchronize, manual_seed, seed} (#94417)mps device for torch.Generator (#91348)torch.int64 support for unary ops (#86615)torch._foreach_lerp (#87562),fused adamw (#88015)_foreach_addc(div/mul)(_).Tensor (#88157)clamp_min clamp_max (#91384)adamw (#88015)Executor and Compiler classes which compiles the XNNPACK graph and preps for execution (#88779, #88778, #88780, #89090)torch.sparse.check_sparse_tensor_invariants context manager that allows users to opt into more checks at runtime for better debugging. (#92094)check_invariants flag to torch.sparse_coo/csr/csc/bsr/bsc/compressed_tensor to allow users to verify components at construction time. (#92094)reduce flag for CPU to torch.sparse.mm with support for sum, mean, amax, amin (#83727){Adadelta, Adagrad, Adamax, AdamW, ASGD, NAdam, RAdam, RProp} differentiable (#86096, #86258, #86183)torch.abs (#87414)torch.select for height and width dimensions (#94612)requires_backend_transfers flag of a model is set to false, then input tensors do not to be transferred to the GPU (via tensor.gpu()) and output tensors do not to be transferred back to the CPU (via tensor.cpu()) since these transfers are inserted into the modelMobileOptimizer.VULKAN_AUTOMATIC_GPU_TRANSFER under torch.utils.mobile_optimizer to the optimization_blocklist argument of optimize_for_mobile (#92081)hipGraph support for pytorch mainline (#88202)any_chain() in operator support (#90949)example_kwarg_inputs argument (#81623, #94032)torch.squeeze to allow squeezing multiple dimensions at once (#89017)where to have cpu scalar args (#87022)torch.tensor.asarray (#90914)Tensor.set_ when dtypes mismatch(#88804)torch.max(#85926)torch.ormqr (#86800)setup_context (#89859, #92312)
forward should no longer take ctx as an input.torch.autograd.set_multithreading_enabled for disabling multithreading in the autograd engine (#86245)remove_duplicate flag to Module.named_buffers() method (#84984) and Module.named_parameters() (#88090)Module forward-pre and forward hooks (#89389)Transformer() fast path (#90783) and kernel selection (#90783)torch.bf16 for Embedding (#94163)freeze argument to Embedding() (#86769)torch.channels_last_3d support for SyncBatchNorm() (#88401)torch.bfloat16 support on CPU for functional.{mish,hardtanh,silu} (#82460)LayerNorm() (#81851, #88064), BatchNorm{1,2,3}d() (#84410), GroupNorm() (#89485, #81852, #88663, #92671, #92668)ModuleList() (#90452)torch.uint8 support for functional.interpolate() on CPU (#90771)functional.max_pool1d error checking consistent between CPU and CUDA (#90211)SyncBatchNorm() fallback to BatchNorm() when it is used in a non-distributed setting (#89706)GroupNorm() on XPU (#87680)is_causal kwarg to TransformerEncoder() layer (#90508)prepend argument to Module hooks to register a hook that will be called before the existing ones (#87370)None from apply_activation_checkpointing (#87871)checkpoint_sequential (#86331)PackedSequence support when device_ids is specified (#86614)BACKWARD_PRE for the backward_prefetch of FSDP (#88428)NO_SHARD in clip_grad_norm_ (#89137)BACKWARD_PRE and BACKWARD_POST in the post-backward assert (#89791)ModuleWrapPolicy.__repr__ (#89058)clip_grad_norm_ for low prec grads (#90028)ModuleWrapPolicy for simplicity in FSDP autowrap (#88450)use_orig_params=True, no_sync and mixed precision to work together (#91193)summon_full_params(with_grads=True) (#85738, #87314)keep_low_precision_grads support when CPU offloading (#86495)state_dict offload_to_cpu settings (#86211)set_state_dict_type API to setup state_dict_type without using context manager (#86243)use_orig_param for FSDP’s optim_state_dict (#89898, #89899, #89900)optim_state_dict and optim_state_dict_to_load for FSDP (#90798, #91343, #92744, #92118, #92991, #92992, #93285, #93318, #94109, #94129)torchrun and TorchElastic to take optional local_addr param to allow skip local IP lookup if specified (#88922)get_worker_info (#87017)torch.nn.functional.conv_transpose3d (#87967), torch.log1p (#89214,#90422), torch.lerp (#75584), torch.logcumsumexp for CPU (#93153)prims.clone (#86705), ndtr, ndtri, log_ndtr, erfcx (#86077), NLL loss (#81128), conv backward (#87047), xlogy and xlog1py (#77712), alpha_dropout (#87989)_adaptive_avg_pool2d_backward (#86359), (#87074), avg_pool2d and avg_pool2d_backward (#87043), scalar_tensor and argmax (#88590), topk (#88694), max_pool2d_with_indices_backward (#88743), grid_sampler_2d_backward (#88745), linalg_cholesky and linalg_cholesky_ex (#89430), aten._cdist_forward (#90042), aten.pixel_shuffle (#91605)linear (#86137, #86302), mm, log1p(#86301, #88155), to_sparse_*(#90281)sparse_dim, dense_dim (#86203, #86203), torch.sum(#86300, #92979), torch.sparse.sampled_addmm(#86401),frac, deg2rad, rad2deg, relu(#88153, #88156, #88442, #86749),conj()(#91695),to_sparse(#90718),sparse_mask` (#92248, #94829)sparse_mask (#91964)indices, values, (c)row_indices, (c)col_indices (#93149) and addmm (#94843)col2im opset 18 (#84594), mse_loss (#90717), aten::contains (#91660), src/index dynamic axes support for aten::scatter_add (#90090), aten::zero (#91731), Raise Unsupported for GridSample with volumetric 5D input (#92212)torch.onnx.export API (#83186)JitScalarType API (#87245)share_from_this to torch::jit::Graph (#87343)ONNX_ATEN_FALLBACK mode (#85736)INT64_MAX magic numbers (#88341)torch.fx compatible with Python-3.11 (#92895)getitem node before split_module (#88510)torch.nn.Linear (#89774), torch.nn.GELU (#86218)torch.bitwise_not (#87286), torch.nn.LayerNorm (#94212), many backward functions (#94343), torch.nn.functional.hardswish (#94342), torch.topk (#91884), torch.arange (#94485), torch.linal.inv (#94551),nn.Conv2d when inputs are on different devices (#86303)torch.nn.{Fold, UnFold} (#94491)k greater than 16 for torch.topk (#94639)@pytorch// in bazel build files which improves embedding usecases (#89660)USE_CUDA for bazel build (#92640)Literal, Protocol, and Final from standard library typing as of Python 3.8+ (#94490)amin/amax (#93091)torch.tensor.scatter (#88244)torch.tensor.index_select over scalar tensor (#94347)torch.tensor.where (#92849)torch.histc consistent between CPU and CUDA (#87832)linalg.solve (#91456), linalg.lstsq (#91460)Modules to work with stateless.functional_call() (#91111), better error messages (#87442),EmbeddingBag (#85433)Upsample and EmbeddingBag module printing (#93850)Conv3D CPU implementation (#94325)Upsample (#94290)functiona.pixel_{shuffle,unshuffle} to consistently return views or not (#86608)Conv3d() (#87527), Upsample() (#87901)Conv{1,2,3}d() (#86521), functional.adaptive_{avg, max}_pool() (#88906)Upsample() (#89252), MaxUnpool3d() (#94372)functional.grid_sample() loss of precision for torch.float16 inputs (#90427)functional.interpolate() bicubic interpolation to properly preserve memory format (#90470)make_functional.py (#91579)CUDA_VISIBLE_DEVICES into account for nvml calls (#94568)autocast_gpu_dtype in custom_fwd and custom_bwd for BFloat16 autocast (#88029)CUDA_VISIBLE_DEVICES into account for nvml calls (#94568)send, recv return type (#92152)backend_type for backend/PG plugin (#93129)isinstance with torch.distributed.ReduceOp (#87303, #88275)__eq__ for ReduceOp (#90088)use_orig_params=True for reentrant activation checkpointing by disabling the post-backward hooks (#87413)_lazy_init in case module changing after FSDP constructor (#87837)NO_SHARD by handling sharded and non-sharded parameters differently in FSDP.clip_grad_norm_ (#88955)ActivationWrapper directly to the inner wrapped module to fix state_dict issues (#87950)use_orig_params=True in FSDP (#91767, #92662)keep_low_precision_grads=True for use_orig_params=True (#90027)use_orig_params=True + no_sync (#90546)no_sync, use_orig_params=True, mixed precision, sharded (#92874)_mp_shard in record_stream (#91096)clip_grad_norm_ issues (#94835), (#86337)load_sharded_state_dict FQN mismatches for shared parameters (#86524)None edge case (#87308)state_dict transformations of modules with persistent buffers failure with mixed precision enabled (#93396)nn.Parameter usage for 2D and use_orig_params=True (#89782, #89845, #90562)_foreach_norm on some tensor sizes (#91844)_foreach_norm from autograd_not_implemented_fallback check (#93995)conj and neg_view (#88182)group["capturable"], not defaults["capturable"] in Adam(W) (#94149)FusedAdam(W) should take OptState into account before unscaling grads (#94060)torch.save (#88867)cat: fix striding (#89332)prelu: Fix prelu ref when a.ndim < 2 (#89809)huber_loss_backward fix (#86955)uniform fix (#90094)unfold_copy fix (#86371)aten.group_norm type promotion fix (#86607)torch.as_strided_scatter_backward memory initialization (#88342)aten.copy preserve strides (#89464)torch.mm: (#90763), (#90917), (#91094)mul when given CUDA CSR Tensor and scalar (#91239)torch.triangular_solve for CSR on CPU when unitriangular=True. (#93352)symint::sizes() instead of sizes() on convolution error messages. (#89549)torch.linspace result on CPU consistent with numpy (#89048)exponential_ few fixes (1) lambda > 0 (2) mkl kernel to continuous (3) better error log on dtype (#92891)cauchy_ few fixes (1) check gamma > 0 (2) better dtype error log (#93314)make_fx invocations isolated (opaque to higher make_fx invocations) by default (#93290)triu/tril operator export with diagonal input (#86843)aten::index_put(self, mask, v) export when rank(mask) < rank(self) (#92862)scatter_add with different static shape of src and index (#89787)_pad_circular export (#86984)ceil_mode and count_include_pad to align torch ceil_mode results in corner case (#87892)unconvertible_ops as per #89261 (#89299)Gather replacement in RNN peephole (#93120)cat operator for tensors with unknown rank (#94870)onnx::Max into standard Op for scalar type alignment (#88750)setType from user into InferredType and Reliable in ConstantValueMap (#88622)BUILD_CAFFE2=0 builds (#88504)torch.autograd.Function.symbolic method support (#94746)FindCommonAncestor in function_extraction (#86650)ScriptedModule (#86745)torch.median (#90326, #88807), torch.{std,var} correction argument (#91203), torch.index_select (#94117, #91064), torch.cumsum (#94119), torch.where (#86240), torch.nn.Embedding (#82809), torch.nn.Softplus (#88555), torch.nn.functional.pad (#89864), torch.max (#91520), padding functions (#91522), torch.nn.functional.upsample (#91669), pooling functions (#91519, #94348), torch.nn.{NLLLoss,SmoothL1Loss} (#94226), torch.nn.SoftPlus (#94256), torch.masked_fill (#94263), torch.fill_ (#94479), torch.median (#94489), torch.nonzero (#94442), torch.nn.BatchNorm (#94351), torch.{min,max} (#94386), torch.nn.GELU (#94529), torch.nn.LSTM (#94889), #95137),torch.nn.Conv2d(#95078),torch.nn.functional.bilinear(#94892),torch.copy\_ (#95272),torch.max_pool2d(#94963),torch.div (#95769)torch.bool for Unary ops (#91120), scatter ops (#94464),torch.float16 for torch.nan_to_num (#94220), torch.nn.HuberLoss (#94567)torch.int64 inputs for torch.dot (#94270), torch.floor_divide (#94488), torch.square (#94766),torch.int64 to torch.int32 for reduction ops and raise warning. (#94484)torch.nn.Conv3d (#94492),torch.float inputs by casting them to torch.float (#88542)torch.cat (#91786, #94662), torch.Conv2d (#91822, #94384), torch.nn.{ELU,ReLU,Hardswish} (#94664), torch.nn.BatchNorm (#94760), torch.nn.MaxPool2d (#94877)extern "C" block (#87853)benchmark_limit ignoring failed kernels in FIND (#91032)Note truncated.
One column per quarter.
This release is meant to fix the following issues (regressions / silent correctness):
This release is meant to fix the following issues (regressions / silent correctness):
The release tracker should contain all relevant pull requests related to this release as well as links to related issues
We are excited to announce the release of PyTorch 1.13! This includes stable versions of BetterTransformer. We deprecated CUDA 10.2 and 11.3 and compl…
We are excited to announce the release of PyTorch 1.13! This includes stable versions of BetterTransformer. We deprecated CUDA 10.2 and 11.3 and completed migration of CUDA 11.6 and 11.7. Beta includes improved support for Apple M1 chips and functorch, a library that offers composable vmap (vectorization) and autodiff transforms, being included in-tree with the PyTorch release. This release is composed of over 3,749 commits and 467 contributors since 1.12.1. We want to sincerely thank our dedicated community for your contributions.
Summary:
The BetterTransformer feature set supports fastpath execution for common Transformer models during Inference out-of-the-box, without the need to modify the model. Additional improvements include accelerated add+matmul linear algebra kernels for sizes commonly used in Transformer models and Nested Tensors is now enabled by default.
Timely deprecating older CUDA versions allows us to proceed with introducing the latest CUDA version as they are introduced by Nvidia®, and hence allows support for C++17 in PyTorch and new NVIDIA Open GPU Kernel Modules.
Previously, functorch was released out-of-tree in a separate package. After installing PyTorch, a user will be able to import functorch and use functorch without needing to install another package.
PyTorch is offering native builds for Apple® silicon machines that use Apple's new M1 chip as a beta feature, providing improved support across PyTorch's APIs.
| Stable | Beta | Prototype |
|---|---|---|
| <ul><li>Better Transformer</li><li>CUDA 10.2 and 11.3 CI/CD Deprecation </li></ul> | <ul><li>Enable Intel® VTune™ Profiler's Instrumentation and Tracing Technology APIs</li><li>Extend NNC to support channels last and bf16</li><li>Functorch now in PyTorch Core Library</li><li>Beta Support for M1 devices</li></ul> | <ul><li>Arm® Compute Library backend support for AWS Graviton</li><li> CUDA Sanitizer</li></ul> |
You can check the blogpost that shows the new features here.
Prior to 1.13, key_padding_mask could be set to uint8 or other integer dtypes in TransformerEncoder and MultiheadAttention, which might generate unexpected results. In this release, these dtypes are not allowed for the mask anymore. Please convert them to torch.bool before using.
1.12.1
>>> layer = nn.TransformerEncoderLayer(2, 4, 2)
>>> encoder = nn.TransformerEncoder(layer, 2)
>>> pad_mask = torch.tensor([[1, 1, 0, 0]], dtype=torch.uint8)
>>> inputs = torch.cat([torch.randn(1, 2, 2), torch.zeros(1, 2, 2)], dim=1)
# works before 1.13
>>> outputs = encoder(inputs, src_key_padding_mask=pad_mask)
1.13
>>> layer = nn.TransformerEncoderLayer(2, 4, 2)
>>> encoder = nn.TransformerEncoder(layer, 2)
>>> pad_mask = torch.tensor([[1, 1, 0, 0]], dtype=torch.bool)
>>> inputs = torch.cat([torch.randn(1, 2, 2), torch.zeros(1, 2, 2)], dim=1)
>>> outputs = encoder(inputs, src_key_padding_mask=pad_mask)
torch.floor_divide to perform floor division (#78411)Prior to 1.13, torch.floor_divide erroneously performed truncation division (i.e. truncated the quotients). In this release, it has been fixed to perform floor division. To replicate the old behavior, use torch.div with rounding_mode='trunc'.
1.12.1
>>> a = torch.tensor([4.0, -3.0])
>>> b = torch.tensor([2.0, 2.0])
>>> torch.floor_divide(a, b)
tensor([ 2., -1.])
1.13
>>> a = torch.tensor([4.0, -3.0])
>>> b = torch.tensor([2.0, 2.0])
>>> torch.floor_divide(a, b)
tensor([ 2., -2.])
# Old behavior can be replicated using torch.div with rounding_mode='trunc'
>>> torch.div(a, b, rounding_mode='trunc')
tensor([ 2., -1.])
torch.index_select on CPU to error that index is out of bounds when the source tensor is empty (#77881)Prior to 1.13, torch.index_select would return an appropriately sized tensor filled with random values on CPU if the source tensor was empty. In this release, we have fixed this bug so that it errors out. A consequence of this is that torch.nn.Embedding which utilizes index_select will error out rather than returning an empty tensor when embedding_dim=0 and input contains indices which are out of bounds. The old behavior cannot be reproduced with torch.nn.Embedding, however since an Embedding layer with embedding_dim=0 is a corner case this behavior is unlikely to be relied upon.
1.12.1
>>> t = torch.tensor([4], dtype=torch.long)
>>> embedding = torch.nn.Embedding(3, 0)
>>> embedding(t)
tensor([], size=(1, 0), grad_fn=<EmbeddingBackward0>)
1.13
>>> t = torch.tensor([4], dtype=torch.long)
>>> embedding = torch.nn.Embedding(3, 0)
>>> embedding(t)
RuntimeError: INDICES element is out of DATA bounds, id=4 axis_dim=3
Prior to this PR, overflows during tensor construction from scalars would not throw an error. In 1.13, such cases will error.
1.12.1
>>> torch.tensor(1000, dtype=torch.int8)
tensor(-24, dtype=torch.int8)
1.13
>>> torch.tensor(1000, dtype=torch.int8)
RuntimeError: value cannnot be converted to type int8 without overflow
Prior to 1.13, cpu_tensor[cuda_indices] was a valid program that would return a cpu tensor. The original use case for mixed device indexing was for non_cpu_tensor[cpu_indices], and allowing the opposite was unintentional (cpu_tensor[non_cpu_indices]). This behavior appears to be rarely used, and a refactor of our indexing kernels made it difficult to represent an op that takes in (cpu_tensor, non_cpu_tensor) and returns another cpu_tensor, so it is now an error.
To replicate the old behavior for base[indices], you can ensure that either indices lives on the CPU device, or base and indices both live on the same device.
1.12.1
>>> a = torch.tensor([1.0, 2.0, 3.0])
>>> b = torch.tensor([0, 2], device='cuda')
>>> a[b]
tensor([1., 3.])
1.13
>>> a = torch.tensor([1.0, 2.0, 3.0])
>>> b = torch.tensor([0, 2], device='cuda')
>>> a[b]
RuntimeError: indices should be either on cpu or on the same device as the indexed tensor (cpu)
# Old behavior can be replicated by moving b to CPU, or a to CUDA
>>> a[b.cpu()]
tensor([1., 3.])
>>> a.cuda()[b]
tensor([1., 3.], device='cuda:0')
torch.eig, torch.matrix_rank, torch.lstsq (#70982, #70981, #70980)The deprecation cycle for the above functions has been completed and they have been removed in the 1.13 release.
bias has the same dtype as input and weight for convolutions on CPU (#83686)To align with the implementation on other devices, the CPU implementation for convolutions was updated to enforce that the dtype of the bias matches the dtype of the input and weight.
1.12.1
# input and weight are dtype torch.int64
# bias is torch.float32
>>> out = torch.nn.functional.conv2d(input, weight, bias, ...)
1.13
# input and weight are dtype torch.int64
# bias is torch.float32
>>> with assertRaisesError():
>>> out = torch.nn.functional.conv2d(input, weight, bias, ...)
# Updated code to avoid the error
>>> out = torch.nn.functional.conv2d(input, weight, bias.to(input.dtype), ...)
.data of a tensor that requires_grad=True with an integer tensor (#78436)Setting the .data of a tensor that requires_grad with an integer tensor now raises an error.
1.12.1
>>> x = torch.randn(2, requires_grad=True)
>>> x.data = torch.randint(1, (2,))
>>> x
tensor([0, 0], requires_grad=True)
1.13
>>> x = torch.randn(2, requires_grad=True)
>>> x.data = torch.randint(1, (2,))
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
RuntimeError: data set to a tensor that requires gradients must be floating point or complex dtype
Prior to this change, C++ custom autograd Function considers tensors passed in TensorList to not be tensors for the purposes of recording the backward graph. After this change, custom Functions that receive TensorList must modify their backward functions to also compute gradients for these additional tensor inputs. Note that this behavior now differs from that of custom autograd Functions in Python.
1.12.1
struct MyFunction : public Function<MyFunction> {
static Variable forward(AutogradContext* ctx, at::Tensor t, at::TensorList tensors) {
return 2 * tensors[0] + 3 * t;
}
static variable_list backward(
AutogradContext* ctx,
variable_list grad_output) {
return {3 * grad_output[0]};
}
};
1.13
struct MyFunction : public Function<MyFunction> {
static Variable forward(AutogradContext* ctx, at::Tensor t, at::TensorList tensors) {
return 2 * tensors[0] + 3 * t;
}
static variable_list backward(
AutogradContext* ctx,
variable_list grad_output) {
return {3 * grad_output[0], 2 * grad_output[0]};
}
};
View operations registered as CompositeExplicitAutograd kernels are no longer allowed to return input tensors as-is. You must explicitly create a new tensor (e.g., using .alias()).
1.12.1
torch::Tensor view_op(const torch::Tensor& self) {
return self;
}
1.13
torch::Tensor view_op(const torch::Tensor& self) {
return self.alias();
}
torch.onnx.register_custom_op_symbolic now only registers the symbolic function at the specified opset version (#85636)This updates register_custom_op_symbolic's behavior to only register the symbolic function at a single version. This is more aligned with the semantics of the API signature. Previously the API registers a symbolic function to all versions up to the specified version. As a result of this change, users will need to register a symbolic function to the exact version when they want to override an existing symbolic function. Users are not affected if (1) an implementation does not exist for the op, or (2) the symbolic function is already registering to the exact version for export.
1.12.1
# Assuming an implemented symbolic function `custom_op_function`
torch.onnx.register_custom_op_symbolic("aten::foo", custom_op_function, 16)
1.13
# Assuming an implemented symbolic function `custom_op_function`
for opset in range(1, 17):
torch.onnx.register_custom_op_symbolic("aten::foo", custom_op_function, opset)
The update is done in regularly to ensure we are in sync with the onnx updates. Users can specify opset_version in torch.onnx.export to maintain opset version 13.
torch.onnx.symbolic_registry is removed (#84382)We removed the symbolic_registry module and hid it as an internal implementation detail. Users previously relying on the register_op function to register custom symbolic functions should move to use the torch.onnx.register_custom_op_symbolic API.
ScalarType and global variables in torch.onnx.symbolic_helper are removed (#82995)The ScalarType class in torch.onnx.symbolic_helper, along with the global variables cast_pytorch_to_onnx, pytorch_name_to_type, scalar_name_to_pytorch, scalar_type_to_onnx and scalar_type_to_pytorch_type are removed from the module. Users previously using these global variables for PyTorch JIT-ONNX type conversion in symbolic functions should move to use the torch.onnx.JitScalarType class.
1.12.1
# 1
torch.onnx.symbolic_helper.scalar_type_to_onnx[
symbolic_helper.scalar_type_to_pytorch_type.index(x.dtype)
].value
# 2
torch.onnx.symbolic_helper.scalar_name_to_pytorch[element_type] in cast_pytorch_to_onnx.keys()
# 3
torch.onnx.symbolic_helper.cast_pytorch_to_onnx["Long"]
# 4
torch.onnx.symbolic_helper.cast_pytorch_to_onnx[tensor.type().scalarType()]
1.13
# 1
torch.onnx.JitScalarType.from_dtype(x.dtype).onnx_type()
# 2
torch.onnx.JitScalarType.from_name(element_type).onnx_compatible()
# 3
torch.onnx.TensorProtoDataType.INT64
# 4
torch.onnx.JitScalarType.from_name(tensor.type().scalarType()).onnx_type()
We added a check to validate all dtype across all input tensors. Previously, users were allowed to pass in tensors with diferent dtypes for c10d collectives. Now, passing in tensors with different dtypes will throw a RuntimeError with the following message: “Invalid usage of tensors with different dtypes Found torch.float and torch.half”. Users can use tensor.to(dtype={some_dtype}) to fix this.
1.12.1
# users could pass inputs having different dtypes
>>> tensor = torch.ones(2, 2) * 7
>>> tensor_h = tensor.half()
>>> tensor_list = [torch.zeros(2, 2) for _ in range(4)] # Assume world_size = 4
# Both cases work.
>>> dist.all_gather(tensor_list, tensor)
>>> dist.all_gather(tensor_list, tensor_h)
...
1.13
# all inputs of c10d collectives need to have the same dtype
>>> tensor = torch.ones(2, 2) * 7
>>> tensor_h = tensor.half()
>>> tensor_list = [torch.zeros(2, 2) for _ in range(4)] # Assume world_size = 4
# Only allow same dtype for all input tensors.
>>> dist.all_gather(tensor_list, tensor) # RuntimeError thrown
...
We limit the usage of c10d APIs to public APIs, so if a user does a wildcard import and calls an internal API, it will fail. Please see the example below:
1.12.1
# users could import both public and non-public symbols:
from torch.distributed.distributed_c10d import *
>>> is_nccl_available() # public API
>>> _check_single_tensor(...) # Non-public API
...
1.13
# users can only import public symbols
from torch.distributed.distributed_c10d import *
is_nccl_available() # public API
_check_single_tensor(...) # Non-public API, this will fail now
...
Details of the changes and the updated tutorial can be found in the PyTorch tutorial PR #2099
1.12.1
// users use relative path to import C++ headers and Work resides in ProcessGroup class
#include <c10d/ProcessGroup.hpp>
#include <c10d/Store.hpp>
#include <c10d/Types.hpp>
#include <c10d/Utils.hpp>
...
class WorkDummy : public ProcessGroup::Work {
...
}
1.13
// users must use absolute path of import C++ files and Work is its own class
#include <torch/csrc/distributed/c10d/ProcessGroup.hpp>
#include <torch/csrc/distributed/c10d/Store.hpp>
#include <torch/csrc/distributed/c10d/Types.hpp>
#include <torch/csrc/distributed/c10d/Utils.hpp>
...
#include <torch/csrc/distributed/c10d/Work.hpp>
class WorkDummy : public Work {
...
}
example_args argument to prepare_fx and prepare_qat_fx (#249) (#77608)We added an additional required example_inputs argument to prepare_fx and prepare_qat_fx APIs, this can be used to do type inference to figure out the type information for each of the fx Node in the graph.
1.12.1
m = resnet18(...)
m = prepare_fx(m, qconfig_dict)
# or
m = prepare_qat_fx(m, qconfig_dict)
1.13
m = resnet18(...)
m = prepare_fx(m, qconfig_dict, example_inputs=(torch.randn(1, 3, 224, 224),))
# or
m = prepare_qat_fx(m, qconfig_dict, example_inputs=(torch.randn(1, 3, 224, 224),))
Previously, we automatically moved the model to CPU in torch.ao.quantization.fx.convert to work around the issue where certain functions called by convert expect CPU arguments. This commit pushes this responsibility to the caller since it is the user's decision of which device to use.
1.12.1
model = resnet18(...)
model = prepare_fx(model, qconfig_mapping, example_inputs)
# calibrate
model = convert_fx(model)
1.13
model = resnet18(...)
model.cpu() # if needed
model = prepare_fx(model, qconfig_mapping, example_inputs)
# calibrate
model = convert_fx(model)
is_reference flag of the torch.ao.quantize_fx.convert_fx function with the convert_to_reference function (#80091, #81326)This PR removes the is_reference flag from the existing convert_fx API and replaces it with a new convert_to_reference function. This separates (1) converting the prepared model to a reference model from (2) lowering the reference model to a quantized model, enabling users to call their custom lowering function for
custom backends.
1.12.1
from torch.ao.quantization.quantize_fx import (
prepare_fx,
convert_to_reference,
)
prepared = prepare_fx(model, ...)
reference = convert_to_reference(prepared, ...)
1.13
from torch.ao.quantization.quantize_fx import (
prepare_fx,
convert_to_reference_fx,
)
prepared = prepare_fx(model, ...)
reference = convert_to_reference_fx(prepared, ...)
This commit adds qconfigs with special observers for fixed qparams ops (operators whose corresponding quantized version has fixed quantized parameters for output) like sigmoid in get_default_qconfig_mapping and get_default_qat_qconfig_mapping. For correctness, we also require users to use these special observers if we detect these fixed qparams ops in prepare.
1.12.1 (fails after this PR):
from torch.ao.quantization.quantize_fx import prepare_fx
model = ModelWithFixedQParamsOps()
qconfig_mapping = QConfigMapping()
example_inputs = ...
prepare_fx(model, qconfig_mapping, example_inputs)
1.13
from torch.ao.quantization import get_default_qconfig_mapping
from torch.ao.quantization.quantize_fx import prepare_fx
model = ModelWithFixedQParamsOps()
qconfig_mapping = get_default_qconfig_mapping()
example_inputs = ...
prepare_fx(model, qconfig_mapping, example_inputs)
qconfig_dict with a typed QConfigMapping object (#78452, #79618)Previously, FX graph mode quantization configurations were specified through a dictionary of qconfigs. However, this API was not in line with other core APIs in PyTorch. This commit replaces this dictionary with a config object that users will create and pass to prepare and convert. This leads to better type safety and better user experience in notebook settings due to improved auto completion.
1.12.1 (deprecated)
from torch.ao.quantization.quantize_fx import prepare_fx
qconfig_dict = {
"": qconfig,
"object_type": [
(torch.nn.Linear, qconfig),
],
"module_name_regex": [
("foo.*bar", qconfig),
],
"module_name": [
("mod", qconfig),
],
}
prepare_fx(model, qconfig_dict)
1.13
from torch.ao.quantization import QConfigMapping
from torch.ao.quantization.quantize_fx import prepare_fx
qconfig_mapping = QConfigMapping()
.set_global(qconfig)
.set_object_type(torch.nn.Linear, qconfig)
.set_module_name_regex("foo.*bar", qconfig)
.set_module_name("mod", qconfig)
prepare_fx(model, qconfig_mapping)
*custom_config_dict with typed config objects (#79066)This commit replaces the following config dicts with python objects:
This leads to better type safety and better user experience in notebook settings due to improved auto completion. 1.12.1
from torch.ao.quantization.quantize_fx import prepare_fx, convert_fx
prepare_custom_config_dict = {
"float_to_observed_custom_module_class": {
"static": {
FloatClass: ObservedClass
}
},
"non_traceable_module_name": ["mod1", "mod2"],
"non_traceable_module_class": [class1, class2],
"input_quantized_idxs": [0, 1],
"output_quantized_idxs": [0],
"preserved_attributes": ["attr1", "attr2"],
}
convert_custom_config_dict = {
"observed_to_quantized_custom_module_class": {
"static": {
FloatClass: ObservedClass
}
},
"preserved_attributes": ["attr1", "attr2"],
}
model = prepare_fx(
model,
qconfig_mapping,
example_inputs,
prepare_custom_config_dict=prepare_custom_config_dict)
model(data)
model = convert_fx(model, convert_custom_config_dict=convert_custom_config_dict)
1.13
from torch.ao.quantization.fx.custom_config import (
PrepareCustomConfig,
ConvertCustomConfig,
)
from torch.ao.quantization.quantize_fx import prepare_fx, convert_fx
prepare_custom_config = PrepareCustomConfig() \
.set_float_to_observed_mapping(float_class, observed_class) \
.set_non_traceable_module_names(["mod1", "mod2"]) \
.set_non_traceable_module_classes([class1, class2]) \
.set_input_quantized_indexes([0, 1]) \
.set_output_quantized_indexes([0]) \
.set_preserved_attributes(["attr1", "attr2"])
convert_custom_config = ConvertCustomConfig() \
.set_observed_to_quantized_mapping(observed_class, quantized_class) \
.set_preserved_attributes(["attr1", "attr2"])
model = prepare_fx(
model,
qconfig_mapping,
example_inputs,
prepare_custom_config=prepare_custom_config)
model(data)
model = convert_fx(model, convert_custom_config=convert_custom_config)
remove_quant_dequant_pairs and fix tests (#84203)This PR removed some passes in convert_fx, and also fixes the way we quantize layer_norm operator, so the qconfig for layer_norm op needs to be updated as well.
1.12.1
import torch
from torch.ao.quantization.qconfig_mapping import QConfigMapping, QConfig
from torch.ao.quantization.observer import default_weight_observer
from torch.ao.quantization.backend_config import (
DTypeConfig,
ObservationType,
)
from torch.ao.quantization.quantize_fx import prepare_fx, convert_fx
qconfig = QConfig(activation=qconfig.activation, weight=default_weight_observer)
qconfig_mapping = QConfigMapping().set_object_type(torch.nn.LayerNorm, q_config) \
.set_object_type(torch.nn.functional.layer_norm, q_config)
# assuming mymodel contains a LayerNorm layer or torch.nn.functional.layer_norm
m = MyModel()
example_inputs = (torch.rand(3, 3),)
m = prepare_fx(m, qconfig_mapping, example_inputs)
1.13
import torch
from torch.ao.quantization.qconfig_mapping import QConfigMapping, QConfig
from torch.ao.quantization.observer import default_placeholder_observer
from torch.ao.quantization.backend_config import (
DTypeConfig,
ObservationType,
)
from torch.ao.quantization.quantize_fx import prepare_fx, convert_fx
qconfig = QConfig(activation=qconfig.activation, weight=default_placeholder_observer)
qconfig_mapping = QConfigMapping().set_object_type(torch.nn.LayerNorm, q_config) \
.set_object_type(torch.nn.functional.layer_norm, q_config)
# assuming mymodel contains a LayerNorm layer or torch.nn.functional.layer_norm
m = MyModel()
example_inputs = (torch.rand(3, 3),)
m = prepare_fx(m, qconfig_mapping, example_inputs)
Before this PR, the dtype attribute of observers was not clearly defined. It originally meant interface_dtype in the eager mode workflow, which is how the codebase before this PR is using it. In the new reference model spec, dtype attribute of an observer represents the dtype value which needs to be passed into a quantize function in the reference model spec. This PR aligns the codebase to this definition of dtype.
1.12.1
dynamic_quant_observer = PlaceholderObserver.with_args(
dtype=torch.float, compute_dtype=torch.quint8)
1.13
dynamic_quant_observer = PlaceholderObserver.with_args(
dtype=torch.quint8, compute_dtype=torch.quint8)
If an operator in ATen takes in a list of tensors, and is marked as “structured” in native_functions.yaml (example), then previously, TensorList was represented as at::TensorList, or c10::ArrayRef<at::Tensor>. Now, it is represented as a more efficient type: const ITensorListRef&.
1.12.1
at::Tensor cat_kernel(at::TensorList tensors,int64_t dim) {
...
}
TORCH_LIBRARY_IMPL(aten, dispatch_key, m) {
...
m.impl("cat", &cat_kernel);
}
1.13
at::Tensor cat_kernel(const at::ITensorListRef& tensors,int64_t dim) {
...
}
TORCH_LIBRARY_IMPL(aten, dispatch_key, m) {
...
m.impl("cat", &cat_kernel);
}
Prior to 1.13, the default for the dtype argument of torch.randint, torch.long, was set via manual python binding. However, in the C++ API, torch::randint would default to the global default data type, which is usually float. In 1.13 we changed the default for dtype in the C++ API to int64 in order to match the python API. To reproduce the old behavior, one can set the dtype argument.
1.12.1
torch::randint(/*low=*/0, /*high=*/10, {2, 3});
1.13
// assuming default dtype is float
torch::randint(/*low=*/0, /*high=*/10, {2, 3}, torch::kFloat);
dim=None for torch.{std, var, std_mean, var_mean} (#81845, #82765, #82912)Prior to 1.13, a C++ API call that has argument types torch::{std, var, std_mean, var_mean}(Tensor, OptionalIntArrayRef, int64_t, bool) used to resolve to the {std, var, std_mean, var_mean}.correction overload. In this release, it resolves to the {std, var, std_mean, var_mean}.dim overload. With the .correction overload, the third argument of type int64_t could be used to pass a correction δN other than 1. In order to call the {std, var, std_mean, var_mean}.correction overload in 1.13, the old int64_t argument can be wrapped in a c10::optional.
1.12.1
// using std as an example
int64_t correction = 2;
torch::std(t, /*dim=*/dim, /*correction=*/correction, /*keepdim=*/True);
1.13
// To replicate in 1.13 using std as an example
auto correction = c10::make_optional<int64_t>(2);
torch::std(t, /*dim=*/dim, /*correction=*/correction, /*keepdim=*/True);
We are deprecating the following APIs of c10d: *_coalesced APIs (#85959), *_multigpu APIs (#85961) and ProcessGroupRoundRobin (#85158)
We added warnings when users call c10d’s *_coalesced, *_multigpu and ProcessGroupRoundRobin APIs. Previously, users can use these APIs without any warnings but now they will see warnings like “torch.distributed.all_reduce_coalesced will be deprecated. If you must use it, please revisit our documentation later at https://pytorch.org/docs/master/distributed.html#collective-functions”. There are still workarounds for *_coalesced APIs but no workarounds will be provided for the other two.
1.12.1
# users could use the following APIs with no warnings:
all_reduce_coalesced(...)
all_gather_coalesced(...)
broadcast_multigpu(...)
all_reduce_multigpu(...)
reduce_multigpu(...)
all_gather_multigpu(...)
reduce_scatter_multigpu(...)
...
1.13
# users can still use these APIs but it will come with warnings:
all_reduce_coalesced(...)
# Warnings:
# torch.distributed.all_reduce_coalesced will be deprecated. If you must
# use it, please revisit our documentation later at
# https://pytorch.org/docs/master/distributed.html#collective-functions"
# Potential workaround:
reqs = []
with dist._coalescing_manager(group, reqs):
reqs.append(dist.all_reduce(tensor1, async_op=True))
reqs.append(dist.all_reduce(tensor2, async_op=True))
for req in reqs:
req.wait()
...
We are deprecating passing optim_input into the FSDP optimizer state checkpointing APIs. The user can simply not pass the optim_input argument, and all behavior is preserved. No fix is needed from users side for now.
1.12.1
# the user can use the following APIs with no warnings
full_optim_state_dict(...)
sharded_optim_state_dict(...)
shard_full_optim_state_dict(...)
flatten_sharded_optim_state_dict(...)
scatter_full_optim_state_dict(...)
rekey_optim_state_dict(...)
1.13
# users can still use these APIs, but they will come with warnings
# The `optim_input` argument is deprecated and will be removed after PyTorch 1.13.
# You may remove it from your code without changing its functionality.
The new operation has a cleaner API and better docs. The update rule is as follows:
1.12.1
LU2, pivots2, info = torch.lu(A, compute_pivots, get_infos=True)
LU1, pivots1, info = torch.lu(A, compute_pivots)
1.13
LU2, pivots2, info = torch.linalg.lu_factor_ex(A, compute_pivots)
LU1, pivots1 = torch.linalg.lu_factor(A, compute_pivots)
The new operation has a notation consistent with linalg.solve, and has an extra parameter adjoint=False. The update rule is as follows:
1.12.1
X = torch.lu_solve(B, LU, pivots)
1.13
X = linalg.lu_solve(LU, pivots, B)
torch._C.Graph, torch._C.Block and torch._C.Node are deprecated. (#83006)Deprecated methods include Graph.op(), Graph.constant(), Graph.at(), Block.op(), and Node.__getitem__(). Previously, these methods are patched into the classes above when users call torch.onnx.export() and are typically used in custom symbolic functions. Users can continue to expect g.op() and g.at() in symbolic functions to work. The g parameter has been substituted by the GraphContext object (#84728). The methods are now exposed by the GraphContext class with APIs unchanged. Users should not rely on the Graph.op(), Graph.constant(), Graph.at(), Block.op(), Node.__getitem__() methods when they are directly interacting with the C classes. Users should use only the op() and at() methods of the GraphContext object, as other fields in the class will change in future releases.
scatter_add on CUDA for all input sizes (#79466)torch.concatenate that aliases torch.cat (#85073)Tensor.is_cpu() that returns whether a tensor is on CPU (#78887)force kwarg to Tensor.numpy() that enables returning a numpy ndarray that does not share storage with the tensor (#78564)torch.special.{airy_ai, bessel_j0, bessel_j1, bessel_y0, bessel_y1, modified_bessel_i0, modified_bessel_i1, modified_bessel_k0, modified_bessel_k1, scaled_modified_bessel_k0, scaled_modified_bessel_k1, spherical_bessel_j0} (#78900), (#78901), (#78902), (#78912), (#78451)torch.special.{chebyshev_polynomial_t, chebyshev_polynomial_u, chebyshev_polynomial_v, chebyshev_polynomial_w, hermite_polynomial_h, hermite_polynomial_he, laguerre_polynomial_l, legendre_polynomial_p, shifted_chebyshev_polynomial_t, shifted_chebyshev_polynomial_u, shifted_chebyshev_polynomial_v, shifted_chebyshev_polynomial_w} (#78196), (#78293), (#78304), (#78366), (#78352), (#78357)weights_only option to torch.load that restricts load to state_dict only, enabling safe loading. This can also be set using the TORCH_FORCE_WEIGHTS_ONLY_LOAD environment variable (#86812)-Werror=unused-but-set-variable build flag (#79305)-Werror=type-limits in Bazel CPU build (#79139)-Werror=unused-variable in Bazel CPU build (#79156)-Wconstant-conversion to catch errors detected in #75400 (#80461)-Werror=non-virtual-dtor build flag (#81012)-Wunused-local-typedef build flag (#86154)torch.{index_select, index_add} (#79217), (#79897).torch.roll (#79970), torch.fft.{fftshift, ifftshift} (#79970), torch.{acos, acosh, asinh, atanh}, (#80030), torch.{cos, sinh, cosh, tanh} (#78718), torch.sqrt, rsqrt (#77490), torch.{triu, tril, diag, trace}(#78062).torch.where (#78665), torch.{where, pow, masked_fill, sgn, tan, angle}(#78665)torch.nn.ConvTranspose1d (#79694).pop function to nn.Sequential and nn.ModuleList (#81601)nn.Module (#80811)maximize kwarg for optim.SparseAdam (#80336), optim.ASGD (#81875), optim.Rprop (#81864), optim.RMSprop (#80326)differentiable kwarg optim.SGD (#80938), optim.Adam (#82205), optim.RMSprop (#83578)optim.Adam (#80279), optim.AdamW (#80280), optim.Adamax (#80319), optim.RMSprop (#83860), optim.Rprop (#83858),optim.{RMSprop, ASGD} (#83860), (#84472)optim.lr_scheduler.PolynomialLR (#82769)foreach maximum and minimum (#82523)linalg.lu_solve, linalg.solve_ex, linalg.vecdot, linalg.vander (#77634, #80073, #70542, #76303)torch.sparse.spdiags for easier creation of diagonal sparse matrices (#78439)torch.ops.nvprims namespace for nvFuser-specific prims (#82155)conv_transpose2d.input, convolution, convolution_backward (#77283, #83557, #80860)aten::_convolution when it is 2D conv in NNC (#84038)ProcessGroup::Work.wait() API to TorchScript (#83303)prim::PythonOp for Autograd Function Export (#74765)aten::index_add.out operator for MPS backend (#79935)aten::prelu operator for MPS backend (#82401)aten::bitwise-not operator native support for MPS backend (#83678)aten::tensor::index_put operator for MPS backend (#85672)aten::upsample_nearest1d operator for MPS backend (#81303)aten::bitwise_{and|or|xor} operators for MPS backend (#82307)aten::index.Tensor_out operator for MPS backend (#82507)aten::masked_select operator for MPS backend (#85818)aten::multinomial operator for MPS backend (#80760)torch.cumsum (#78554, #81107)torch.nn.LSTM (#78943, #79702)torch.nn.ReplicationPad2d (#79057, #79291)torch.nn.threshold (#78654, #79717)torch.nn.BatchNorm2d (#80510)torch.nn.LayerNorm (#80980)torch.nn.GLU (#80910, #81729)torch.select (#81771)torch.stack (#81064)torch.quantize_per_tensor (#81492)torch.dequantize (#81493)Upsample2D (#81720)Distributed Checkpointing (Prototyping)Distributed(c10d)all_gather (#83713) and uneven output support for reduce_scatter (#87010)ReduceOp (#84243)DistributedDataParallel
FullyShardedDataParallelsharded_optim_state_dict and flatten_sharded_optim_state_dict. (#77628)torch.distributed.elasticActivation Memory Management (Prototyping)torch.distributed.algorithms.checkpoint.checkpoint_wrapper to wrap nn.Modules with activation checkpointing or activation offloading to easily use and experiment with activation checkpoint techniques without modifying model code. This makes it simpler to leverage activation checkpointing to reduce memory footprint of your training applications and train larger models. (#83035, #78704, #78854, #79830, #80089, #84907, #84908, #85448, #85449)-Werror=all with a few exceptions in Bazel build for CUDA (#79306)float16 support for torch.{arange, linspace} (#80492)torch.index_reduce (#80464)stable kwarg to torch.argsort that controls the relative order of equivalent elements (#75162)torch.distributions.kl_divergence for two Bernoulli distributions (#79944)torch.{as_tensor, as_subclass} (#86105)torch.{addcmul, addcdiv} (#74234)bfloat16 support for torch.save with XLA/HPU tensors (#77534)TensorOption signatures for consistency with JIT schemas (#82241)torch.library.Library with PYTORCH_DISABLE_LIBRARY (#85190)dim=None for torch.{mean, sum, nanmean, nansum} (#81286), (#79881), (#82912)logsumexp to amp.autocast (#76330)const T& access to ListElementReference (#83177)stderr in torch.utils.cpp_extension (#82097)torch.utils.cpp_extension (#82860)__all__ to torch.utils.cpp_extension, torch.utils.hooks and torch.utils.show_pickle (#85331)torch.{amin, amax, nansum, nanmean} (#80082), torch.scatter_reduce (except reduction=prod) (#85000), torch.linalg.det (#79487), torch.{elu_, celu_, selu_} (#83080)nn.functional.{binary_cross_entropy} (#77852) , nn.functional.{embedding}(#79699), nn.functional.{mse_loss, softplus, l1_loss, smooth_l1_loss, prelu, hardswish} (#78740), nn.functional.{nll_loss, batch_norm, layer_norm, group_norm, cross_entropy, soft_min} (#84976) torch.{log_softmax, softmax}(#84976), torch.amin, amax, nansum (#80082)torch.linalg.det for real inputs (#80217)torch.utils.checkpoint with use_reentrant=False (#80987)torch.autograd.graph.disable_saved_tensors_hooks (#85971)ctx->needs_input_grad(idx) (#82544)check_nan flag to torch.autograd.detect_anomaly which enables users to run anomaly mode without nan checking (#83481).cu to improve compile times (#81193)append_cxx_flag_if_supported macro (#82883)groups argument validation for nn.Conv{1,2,3}d modules (#77919)nn.Module full backward hooks by removing reference cycles (#80139)kl_div at boundary and its general implementation (#80334)nn.AdaptiveAvgPool2d (#84061)groups argument validation for nn.Conv{1,2,3}d (#85248)nn.MaxUnpool{2,3}d (#78280)optim and nn (#80237)nn.Sequential: + (#81170), extend (#81179), insert (#81402), +=, * and *= (#81279),nn.MaxUnpool{1,2,3}d (#84766)nn.functional.kl_div on CUDA (#77676)fused kwarg to optim.Adam to enable a fused implementation on CUDA (#85739)functionalize() API that lives with functorch (#77129, #77126, #77125, #78199, #77132, #77713, #77714, #78819, #78820, #82008, #82009, #81702, #80416, #80418, #80251, #80526, #82326, #81454, #81471, #83542, #83701, #85975)__torch_dispatch__ subclasses and modes to override more tensor metadata: device/size/stride/dim (#77684, #77970, #78646, #78691)torch.library API, for registering python functions to the pytorch dispatcher:
torch.library (#77990)torch.library decorators return function, to allow for chaining (#78996)cholesky, linalg_qr, linalg_eigh and linalg_eighvalsh to structured kernels, giving them support with meta tensors (#79300, #79054, #79072)c10::FunctionSchema::operator<< to print native_functions.yaml syntax (#79645)x.detach().resize_(...) (#83590)torch.ops.ns.opname.overload accessor in __torch_dispatch__ (#85132)weights for WeightedRandomSampler (#78585)radom_split to accept percentages as lengths (#78877)functorch.jacfwd now accepts a randomness kwarg (#84220)vmap on a function with no Tensor inputs (#83016)Tensor.as_strided batching rule. This is a primitive used in forward-mode AD (among other things) and improves composability of vmap with other transforms (like jvp).functorch.functionalize: added support for in-place views on inputs (#83993)functorch.functionalize: moved this API out of the functorch.experimental namespace (#85742)linalg.cholesky, linalg.eigvals, linalg.eigvalsh, linalg.matrix_norm, linalg.matrix_power, linalg.norm, linalg.tensorinv, linalg.solve_triangular (#82177)linalg.solve (#82814)linalg.cross (#83759)linalg.matrix_rank (#83760)linalg.pinv (#83761)Tensor.fill_ (#84015)linalg.lstsq (#82325)linalg.lu_solve (#85175)driver= kwarg to torch.linalg.svd and svdvals. Add cusolver gesvdaStridedBatched driver to linalg.svd (#74521)torch.einsum (#86219)einsum (#84890)einsum to remediate MPS regression (#87135)einsum (#87199)sparse_dim and dense_dim for batched, hybrid CSR/CSC/BSR/BSC (#80565, #80901)mm, addmm, matmul and F.linear (#85551, #85308, #85379, #85307)permute (#79707)torch.nonzero and add(dense, CSR) (#79062)transpose (#82122)mul (#82962)empty_like (#82310)select (#82119)device_for_folded_attrs parameter and sets the requires_grad option for a folded tensor (#79067)ignore_parameters_and_buffers flag to FxGraphDrawer (#79982)is_fx_tracing flag in the FX tracer (#80255)__torch_dispatch__ (#82549)enable_tracing flag for ProxyTorchDispatchMode instead of modifying torch dispatch mode stack inner attributes (#82643)__deepcopy__ for fx.Tracer (#83130)torch._refs.var for nvFuser executor (#79517)where (tensor, python_scalar, tensor) type promotion (#80347)torch.jit.fuser() option for disabling all fusers (#81731)silu (#81724)prims.sign, refs.sign, squeeze, native_batch_norm, transpose) (#83167, #85562, #84629, #84117)masked_fill (#78368, #85108)index_put (#78384, #85685)LSTM and MultiHeadAttention (#79959, #79956, #79960, #83304, #85068)matmul (#83885)torch.tensor_split (#77437), torch.lerp (#78891), torch.movedim and torch.moveaxis (#78931), torch.scatter_add (#79103), torch.argsort (#80234), aten::native_dropout (#81743), aten::native_layer_norm (#81754), aten::convolution (#81815), aten::_log_softmax (#81804), aten::layer_norm for ONNX opset version 17 using LayerNormalization (#84293), nn.init.normal (#84149)aten::reshape, aten::reshape_as, aten::t, aten::transpose, aten::numpy_T, aten::expand, aten::expand_as, aten::embedding, aten::embedding_bag, aten::view, aten::select, aten::eq, aten::ne, aten::gt, aten::lt, aten::le, aten::ge, aten::elu, aten::selu, aten::hardtanh, aten::hardswish, aten::as_strided, quantized::sigmoid, quantized::layer_norm, quantized::group_norm, quantized::leaky_relu, quantized::instance_normnn.module (#82038), (#82039), (#82040)torch.onnx APIs now support runtime type checking when @beartype is present in the Python environment. A warning is emitted when a type mismatch is detected.TORCH_ONNX_EXPERIMENTAL_RUNTIME_TYPE_CHECK=ERRORS. To disable this behavior, set TORCH_ONNX_EXPERIMENTAL_RUNTIME_TYPE_CHECK=DISABLED which effectively makes it a no-op.torch.onnx submoduletorch.cuda.is_bf16_supported() returns True (#80410)complex32 for tan, atan, sin, asin (#77802),(#77606)logical_{or, xor} (#75947)get_current_stream (#78066)device_count (#85192)torch.cuda.device_count (#84878)__launch_bounds__ for torch.mode with CUDA 11.7 (#79710)cumsum (#75693)torch.{im2col,col2im} on CUDA (#84372)ReflectionPad (#84949)__all__ to torch.cuda (#85193)lerp on CPU (#84327)prelu op and module for quantized CPU backend (#73491)TensorLikePair (#77836)aten::softplus operator by adding RankedPlaceholder for graph nodes instead of constants (#81169)torch.addmm (#81519)dispatch1DJob for MPS native implementations (#82982)torch.adaptive_avgpool_2d for larger output sizes (#85726)torch.constant_pad_nd for 4D+ padding (#85991)Engine::evaluate_function event. (#77696)torch.nn.linear (#81773)relu in metal shader (#78544)max_pool2d, linear, conv2d FP32 operator tests for XNNPACK (#83131)Distributed(c10d)get_local_rank, get_global_rank and get_global_ranks (#82134, #84363)_all_gather_base with a public API all_gather_into_tensor (#85686)_reduce_scatter_base with a public API reduce_scatter_tensor (#85867)ncclGetLastError (#83724, #85825, #85850)ncclRemoteError (#85887)NCCL_ASYNC_ERROR_HANDLING=2 that does not crash the process (#84386)Distributed OptimizerDistributedDataParallelFullyShardedDataParalleloptim_input orders across ranks (#78599)sharded_state_dict logic to the post hook to avoid OOM (#82613)_init_from_local_tensor to create ShardedTensor to avoid communication overhead (#82911)state_dict on CPU cases (#85640)FSDPExtensions for TP support (#85039)torch.distributed.elastictrunk / linux-bionic-cuda10.2-py3.9-gcc7 / test (default from 2 -> 4 (#83424)dim out of range check for logcumsumexp on CUDA when the source tensor is empty(#78284)__init__.py for torch.utils.jit (#78629)gather with an empty index tensor when sparse_grad=True (#78698)torch.distributions.kl_divergence (#78432)end in the output of torch.arange for some inputs (#80758)torch.distributions.Transform to be pickle-able (#81707)self and mask are on the same device for torch.masked_fill (#82737)torch.utils.checkpoint (#82776)Tensor.__hash__ for Tensor subclasses (#83174)torch.cat for 0-dim tensors with different dtypes (#83391)torch.equal on CPU when inputs have different dtypes (#83350)torch.districutions.{HalfCauchy, HalfNormal} (#84322)tau is less than or equal to that of input in torch.ormqr (#85278)weights is a 1D tensor in torch.bincount (#85881)out arguments that have a large number of dims (#85294)torch.utils.dlpack strides to 1 where size of corresponding dimensions < 2 (#83158)torch.empty_strided that sizes has the same dimensionality as strides (#82422)torch.istft default output length to prevent trimming of last element (#80031)IListRefTag::Materialized to IListRefIterator destructor. (#85467)im2col by adding a check that pad_width and pad_height are non-negative (#85541)check_compiler_ok_for_platform on non-English locales in torch.utils.cpp_extension (#85891)torch.sgn which fixed forward-over-backward for torch.linalg.svd and other spectral decompositions, and torch.norm, torch.linalg.{norm, matrix_norm}(#80082)backward(inputs=) behavior in-place (#79996)create_graph=True and full backward hook registered (#82788)torch.stack to correctly handle implicit real->complex casting (#84993)torch.nn.functional.{leaky_relu, threshold} when inplace=True (#85634)torch.utils.checkpoint when use_reentrant=False (#81766)nn.functional.binary_cross_entropy_with_logits (#80083)norm(p=inf) (#78105)-Wno-unused-but-set-variable for clang 13.0.0 (#79666)torch._dl extension (#84361)torch.acosh for complex numbers (#80841).nn.Embedding ‘s max_norm argument when forward mode AD is used (#78560)nn.ChannelShuffle when given empty Tensors (#77029)nn.RReLU backward on CUDA (#80434)torch.nn.parallel.* APIs (#81476)nn.Conv2d fallback implementation for single channel inputs and channels last weight (#82392)nn.Conv{1,2,3}d for in_channels (#84302)nn.GeLU for empty inputs (#84926)nn.Conv2d on ARM-based machines (#85711)nn.ParameterList printing of Tensors on the “meta” device (#78529)nn.MaxPool3D on CUDA (#80748)nn.MaxPool1d (#85594)nn.Softmax for large input tensors (#84182)nn.RReLU (#84996)torch.nn.grad by calling into the c++ backward kernel directly (#81839)torch.nn.PixelShuffle for empty inputs (#86262)torch.nn.BatchNorm (#84410)optim.SGD maximize flag when momentum is involved (#81859)optim.lr_scheduler.CyclicLR (#85462)lr in optim.lr_scheduler.SequentialLR (#72856)seqlen > 1024 (#83639)num_head in TransformerEncoder to slow_path (#83483)__torch_function__ bug in getindex that causes an error not set exception (#78781)__torch_dispatch__ usage with inplace views (#79902)NoneType object has no attribute python_exit_status when DataLoader exits (#83985)functorch.grad: fixed silent correctness issue from calling a view operation on a captured tensor followed by an in-place operation (#85374)functorch.jacrev, functorch.jacfwd: fixed loud in-place errors when passing in inputs to the transforms and mutating them (#84914, #84915)functorch.vmap: Fixed support for in-place view operations (Tensor.unsqueeze_, Tensor.transpose_, Tensor.t_, Tensor.squeeze_) (#82899, #82903, #82972)functorch.vmap: added an error on incorrect weight shape to torch.nn.functional.prelu (#83106)functorch.vmap: fixed support for multinomial (#83838)functorch.vmap: fixed incorrect support for conv_transpose with groups > 1 (#84938)vmap x vjp x vjp composition for torch.nn.functional.prelu (#84939)cross to match unbatched behavior (#86926)linalg.cross (#83798)linalg.lstsq (#85357)linalg.lu_solve/torch.unpack to prevent bad memory usage on CPU (#85922)matrix_exp. (#81330)mul on tiny COO tensors (#80254)select if given all zero integer COO tensors(#82215)torch.sparse.sampled_addmm (#85194)function.__name__ rather than function.__code__.co_name (#84373)to_folder by adding custom_builtins to dump (#81433)import warnings (#82760)CalculatedNecessaryArgs to avoid underflow with schemas where all args have defaults. (#79331)map_location arg to torch.jit.load in torch.load (#78733)import_ir_module for pickle case to reduce memory usage (#80131)enumerate() (#80585)std::out_of_range when using NNC and ConstantChunk input shapes are unknown (#82698)sum_mean_dim symbolic shape fn (#83357)resize_ to avoid _MapBase::at runtime error (#81422)define_constant pybind signature to match std::complex scalar in NVFuser (#83684)torch.ScriptObject in torch::jit::as_object (#84398)torch.jit.trace check that was causing tracing to fail for MPS inputs (#84850)None to futures (#85304)fused_moving_avg_obs_fake_quant_* (#78148)ceil_mode of the avgpool op (#79028)Note truncated.
This release is meant to fix the following issues (regressions / silent correctness):
This release is meant to fix the following issues (regressions / silent correctness):
an input without a batch dimension, and dropout behavior was changed to drop along the first dimension. This was a silent breaking change.
We are excited to announce the release of PyTorch 1.12! This release is composed of over 3124 commits, 433 contributors. Along with 1.12, we are releasing beta versions of AWS S3 Integration, PyTorch Vision Models on Channels Last on CPU, Empowering PyTorch on Intel® Xeon® Scalable processors with Bfloat16 and FSDP API. We want to sincerely thank our dedicated community for your contributions.
Summary:
Updated type promotion for torch.clamp (#77035)
In 1.11, the ‘min’ and ‘max’ arguments in torch.clamp did not participate in type promotion, which made it inconsistent with minimum and maximum operations. In 1.12, the ‘min’ and ‘max’ arguments participate in type promotion.
1.11
>>> import torch
>>> a = torch.tensor([1., 2., 3., 4.], dtype=torch.float32)
>>> b = torch.tensor([2., 2., 2., 2.], dtype=torch.float64)
>>> c = torch.tensor([3., 3., 3., 3.], dtype=torch.float64)
>>> torch.clamp(a, b, c).dtype
torch.float32
1.12
>>> import torch
>>> a = torch.tensor([1., 2., 3., 4.], dtype=torch.float32)
>>> b = torch.tensor([2., 2., 2., 2.], dtype=torch.float64)
>>> c = torch.tensor([3., 3., 3., 3.], dtype=torch.float64)
>>> torch.clamp(a, b, c).dtype
torch.float64
Updates the type promotion rule such that given a complex scalar and real tensor, the value type of real tensor is preserved
1.11
>>> a = torch.randn((2, 2), dtype=torch.float)
>>> b = torch.tensor(1, dtype=torch.cdouble)
>>> (a + b).dtype
torch.complex128
1.12
>>> a = torch.randn((2, 2), dtype=torch.float)
>>> b = torch.tensor(1, dtype=torch.cdouble)
>>> (a + b).dtype
torch.complex64
PyTorch 1.12 makes the default math mode for fp32 matrix multiplications more precise and consistent across hardware. This may affect users on Ampere or later CUDA devices and TPUs. See the PyTorch blog for more details.
In 1.11.0, unlike scatter which takes a reduce kwarg or scatter_add, scatter_reduce was not an in-place function. That is, it did not allow the user to pass an output tensor which contains data that is reduced together with the scattered data. Instead, the scatter reduction took place on an output tensor initialized under the hood. Indices of the output that were not scattered to were filled with reduction inits (or 0 for options ‘amin’ and ‘amax’).
In 1.12.0, scatter_reduce (which is in beta) is in-place to align with the API of the related existing functions scatter/scatter_add. For this reason, the argument input in 1.11.0 has been renamed src in 1.12.0 and the new self argument now takes a destination tensor to be scattered onto. Since the destination tensor is no longer initialized under the hood, the output_size kwarg in 1.11.0 that allowed users to specify the size of the output at dimension dim has been removed. Further, in 1.12.0 we introduce an include_self kwarg which determines whether values in the self (destination) tensor are included in the reduction. Setting include_self=True could, for example, allow users to provide special reduction inits for the scatter_reduction operation. Otherwise, if include_self=False, indices scattered to are treated as if they were filled with reduction inits.
In the snippet below, we illustrate how the behavior of scatter_reduce in 1.11.0 can be achieved with the function released in 1.12.0.
Example:
>>> src = torch.arange(6, dtype=torch.float).reshape(3, 2)
>>> index = torch.tensor([[0, 2], [1, 1], [0, 0]])
>>> dim = 1
>>> output_size = 4
>>> reduce = "prod"
1.11
>>> torch.scatter_reduce(src, dim, index, reduce, output_size=output_size)
`tensor([[ 0., 1., 1., 1.],
[ 1., 6., 1., 1.],
[20., 1., 1., 1.]])`
1.12
>>> output_shape = list(src.shape)
>>> output_shape[dim] = output_size
# reduction init for prod is 1
# filling the output with 1 is only necessary if the user wants to preserve the behavior in 1.11
# where indices not scattered to are filled with reduction inits
>>> output = src.new_empty(output_shape).fill_(1)
>>> output.scatter_reduce_(dim, index, src, reduce)
`tensor([[ 0., 1., 1., 1.],
[ 1., 6., 1., 1.],
[20., 1., 1., 1.]])`
nn.GroupNorm: Report an error if num_channels is not divisible by num_groups (#74293)Previously, nn.GroupNorm would error out during the forward pass if num_channels is not divisible by num_groups. Now, the error is thrown for this case during module construction instead.
1.11
m = torch.nn.GroupNorm(3, 7)
m(...) # errors during forward pass
1.12
m = torch.nn.GroupNorm(3, 7) # errors during construction
nn.Dropout2d: Return to 1.10 behavior: perform 1D channel-wise dropout for 3D inputsIn PyTorch 1.10 and older, passing a 3D input to nn.Dropout2D resulted in 1D channel-wise dropout behavior; i.e. such inputs were interpreted as having shape (N, C, L) with N = batch size and C = # channels and channel-wise dropout was performed along the second dimension.
1.10
x = torch.randn(2, 3, 4)
m = nn.Dropout2d(p=0.5)
out = m(x) # input is assumed to be shape (N, C, L); dropout along the second dim.
With the introduction of no-batch-dim input support in 1.11, 3D inputs were reinterpreted as having shape (C, H, W); i.e. an input without a batch dimension, and dropout behavior was changed to drop along the first dimension. This was a silent breaking change.
1.11
x = torch.randn(2, 3, 4)
m = nn.Dropout2d(p=0.5)
out = m(x) # input is assumed to be shape (C, H, W); dropout along the first dim.
The breaking change in 1.11 resulted in a lack of support for 1D channel-wise dropout behavior, so Dropout2d in PyTorch 1.12 returns to 1.10 behavior with a warning to give some time to adapt before the no-batch-dim interpretation goes back into effect.
1.12
x = torch.randn(2, 3, 4)
m = nn.Dropout2d(p=0.5)
out = m(x) # input is assumed to be shape (N, C, L); dropout along the second dim.
# throws a warning suggesting nn.Dropout1d for 1D channel-wise dropout.
If you want 1D channel-wise dropout behavior, please switch to use of the newly-added nn.Dropout1d module instead of nn.Dropout2d. If you want no-batch-dim input behavior, please note that while this is not supported in 1.12, a future release will reinstate the interpretation of 3D inputs to nn.Dropout2d as those without a batch dimension.
F.cosine_similarity: Improve numerical stability (#31378)Previously, we first compute the inner product, then normalize. After this change, we first normalize, then compute inner product. This should be more numerically stable because it avoids losing precision in inner product for inputs with large norms. Because of this change, outputs may be different in some cases.
Functions in torch.ops.aten.{foo} no longer accept self as a kwarg
torch.ops.aten.{foo} objects are now instances of OpOverloadPacket (instead of a function) that have their __call__ method in Python, which means that you cannot pass self as a kwarg. You can pass it normally as a positional argument instead.
1.11
>>> torch.ops.aten.sin(self=torch.ones(2))
tensor([0.8415, 0.8415])
1.12
# this now fails
>>> torch.ops.aten.sin(self=torch.ones(2))
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: __call__() got multiple values for argument 'self'
# this works
>>> torch.ops.aten.sin(torch.ones(2))
tensor([0.8415, 0.8415])
torch_dispatch now traces individual op overloads instead of op overload packets (#72673)
torch.ops.aten.add actually corresponds to a bundle of functions from C++, corresponding to all over the overloads of add operator (specifically, add.Tensor, add.Scalar and add.out). Now, __torch_dispatch__ will directly take in an overload corresponding to a single aten function.
1.11
class MyTensor(torch.Tensor):
....
def __torch_dispatch__(cls, func, types, args=(), kwargs=None):
# Before, func refers to a "packet" of all overloads
# for a given operator, e.g. "add"
assert func == torch.ops.aten.add
1.12
class MyTensor(torch.Tensor):
....
def __torch_dispatch__(cls, func, types, args=(), kwargs=None):
# After, func refers to an individual operator overload,
# e.g. "add.Tensor"
assert func == torch.ops.aten.add.Tensor
# you can recover the old behavior with func.overloadpacket
assert func.overloadpacket == torch.ops.aten.add
The forward-backward correlation is no longer captured as to workaround a profile crash. This feature may be reenabled in a future release after the underlying issue is fixed.
with torch.profiler.profile() as p:
loss = model(inputs)
loss.backward() # Invoke autograd
# The exported chrome trace will not have forward-backward flow events. (arrows)
p.export_chrome_trace(...)
The minimum supported bytecode version is being bumped from 3 to 4. We no longer support version 3 bytecode models because the bytecode version was bumped from 3 to 4 more than half a year ago, and there was code in operator loading that performed differently on one operator on the global bytecode version 3.
If the model is generated before Oct 5, 2020, please use the following lines to update the model to the latest version:
1.12
import torch
from torch.jit.mobile import _get_model_bytecode_version
old_model_path = "old_model.ptl"
new_model_path = "new_model.ptl"
# Load full jit model
jit_model = torch.jit.load(old_model_path)
# Save model for mobile
jit_model._save_for_lite_interpreter(new_model_path)
# Verify the model can be loaded
mobile_m = _load_for_lite_interpreter(new_model_path)
# Get bytecode version from the new model
bytecode_version = _get_model_bytecode_version(new_model_path)
print(f"bytecode version is {bytecode_version}")
FullyShardedDataParallel's optional param name ‘fsdp_auto_wrap_policy’ (1.11) changed to ‘auto_wrap_policy’ (1.12). ‘default_auto_wrap_policy’ (1.11) is changed to ‘size_based_auto_wrap_policy’ (1.12).
In 1.11, when wrapping a model with FSDP instead of:
model = MyModel()
wrapped_model = FullyShardedDataParallel(
model,
**fsdp_auto_wrap_policy**=functools.partial(
default_auto_wrap_policy,
min_num_params=0, # wrap all modules
)
...
1.12
model = MyModel()
wrapped_model = FullyShardedDataParallel(
model,
**auto_wrap_policy**=functools.partial(
size_based_auto_wrap_policy,
min_num_params=0, # wrap all modules
)
...
TorchScript models created with PyTorch 1.5 or earlier and using the operators quantized::linear_prepack_legacy, linear_prepack_fp16_legacy, quantized::linear_unpack.legacy, or quantized::linear_unpack_fp16.legacy will no longer work and need to be re-exported. Please use PyTorch Quantization to quantize the Linear module instead.
Deprecated torch.testing.make_non_contiguous (#72705)
torch.testing.make_non_contiguous is being deprecated and will be removed in a future release. Depending on the use case there are different replacement options: If you are using make_non_contiguous in the PyTorch test suite, you can use torch.testing._internal.common_utils.noncontiguous_like
1.11
a = torch.randn(1, 2, 3)
torch.testing.make_non_contiguous(a)
1.12
a = torch.randn(1, 2, 3)
torch.testing._internal.common_utils.noncontiguous_like(a)
If you are using make_non_contiguous in combination with a creation function to create a noncontiguous tensor with random values, you can use make_tensor.
1.11
a = torch.randn(1, 2, 3)
torch.testing.make_non_contiguous(a)
1.12
torch.testing.make_tensor(..., noncontiguous=True)
If you are using make_non_contiguous with a specific tensor, you can use torch.repeat_interleave
1.11
a = torch.tensor([[1., 2.], [1., 2.]])
torch.testing.make_non_contiguous(a)
1.12
a = torch.tensor([[1., 2.], [1., 2.]])
torch.repeat_interleave(input, 2, dim=-1)[..., ::2]
torch.lu() is deprecated in favor of torch.linalg.lu_factor() and torch.linalg.lu_factor_ex(). torch.lu() will be removed in a future PyTorch release. If you were previously using get_infos=False (this is the default), you should use torch.linalg.lu_factor instead:
1.11
LU, pivots = torch.lu(A, compute_pivots)
1.12
LU, pivots = torch.linalg.lu_factor(A, compute_pivots)
If you were previously using get_infos=True you should use torch.linalg.lu_factor_ex:
1.11
LU, pivots, info = torch.lu(A, compute_pivots, get_infos=True)
1.12
LU, pivots, info = torch.linalg.lu_factor_ex(A, compute_pivots)
torch.lu_solve() is deprecated in favor of torch.linalg.lu_solve(). torch.lu_solve() will be removed in a future PyTorch release.
1.11
X = torch.lu_solve(B, LU, pivots)
1.12
X = torch.linalg.lu_solve(LU, pivots, B)
torch.solve which was deprecated in a previous release is now being removed. You should use torch.linalg.solve. instead. Note that torch.linalg.solve has its arguments reversed and does not return the LU factorization. To get the LU factorization see torch.lu, which can be used with torch.lu_solve or torch.lu_unpack.
1.11
X = torch.solve(B, A).solution
1.12
X = torch.linalg.solve(A, B)
nn.Module: Deprecate positional args for state_dict() (#72780)state_dict can currently be called in two ways: destination, prefix, and keep_vars can be passed as positional arguments, or as kwargs. The ability to do the former is being deprecated and will be removed in a future release. You should pass the arguments in as kwargs only.
Deprecated __torch_function__ as instance method for more functions (#74829)
__torch_function__ should be defined as a class method. Defining __torch_function__ as a plain method has already been previously deprecated for the functions handling __torch_function__ in Python. This change makes it so that that is also the case for functions that handle __torch_function__ in c++.
1.11
class Bad():
def __torch_function__(self, *args, **kwargs):
pass
t = Bad()
torch.abs(t)
1.12
class Good():
@classmethod
def __torch_function__(cls, *args, **kwargs):
pass
t = Good()
torch.abs(t)
torch.jit.quantized (#72690)Instead of using functions defined in torch.jit.quantized, please use PyTorch Quantization to dynamically quantize Linear/RNNCell/LSTMCell/GRUCell/LSTM modules. It’s both supported in Eager Mode Quantization and FX Graph Mode Quantization
1.11
>> torch.jit.quantized.QuantizedLSTMCell(...)
1.12
>> torch.jit.quantized.QuantizedLSTMCell(...)
"torch.jit.QuantizedLSTMCell is deprecated and will be removed in an upcoming
PyTorch release. Please use the torch.nn.quantized.dynamic.LSTMCell instead."
mps that can be used to leverage GPU acceleration on macOS platform with Apple Native Silicon (M1) or discrete AMD GPUs. (blogpost with details)torch.special.log_ndtr (#74795)torch.distributions.transforms.{SoftplusTransform,CumulativeDistributionTransform} (#52300, #72495)torch.testing to stable (#73348)maximize flag for optim.Adadelta(#75330)torch.complex32 to help computing with complex datatype with lower memory usage at the cost of lower precision. Note that this is an experimental feature (#78245) and the major focus in this release was to support operators under torch.fft on CUDA. Besides those operators we have added support and testing for following limited set of ops (NOTE: few operators are only supported on CUDA): Tensor.copy_, torch.complex, torch.testing.make_tensor, cat, Tensor.fill_, Tensor.item, torch.atleast_1d, torch.atleast_2d, torch.atleast_3d, torch.dsplit, torch.vsplit, torch.hsplit, torch.hstack, torch.dstack, torch.vstack, Tensor.conj, torch.add, torch.sub, torch.mul, torch.sub, torch.div, torch.view, torch.view_as, torch.real, torch.imag, torch.neg, Tensor.__getitem__, torch.sum, torch.prod, torch.abs, torch.sgn, torch.exp, torch.log, torch.eq, torch.masked_fill, torch.index_put, torch.rand, torch.randn, torch.full, torch.empty, torch.ones, torch.zeros, torch.block_diag, Tensor.chunk, Tensor.clone, Tensor.contiguous, torch.diag_embed, torch.diagonal, torch.as_strided, torch.column_stack, Tensor.T, Tensor.H, Tensor.mT, Tensor.mH, Tensor.narrow, torch.isfinite, torch.isinf, torch.isreal, torch.flatten, Tensor.chalf, torch.empty_like, torch.movedim ( #73847, #74667, #74854, #75010,#75156, #75311, #75498, #76132, #76158, #75592, #76615, #77179, #77339, #77446, #77483, #77479, #77192, #76724, #77404).torch.fft now support tensors with torch.complex32 dtype (CUDA only) (#74857).torch.complex32 tensor now participate in type-promotion (#76893)torch.chalf alias for torch.complex32 and Tensor.chalf method (#75320).torch.chalf tensors (#76614).torch.complex32, torch.complex64, torch.complex128)
torch.linalg.ldl_factor_ex and torch.linalg.ldl_solve (#69828)linalg.vander (#76303)linalg.lu (#67833)linalg.lu_solve (#72935)where and huber_loss (#77353)dot/group_norm/instance_norm/var_mean/index_reduce/matmul/bernoulli/adaptive_avg_pool (#77499) index_select/abs/min/max (#76916), reflection_pad2d (#77681), square (#77682), log_sigmoid_forward (#77739), several more ops (#77362)nn.Dropout1d: New module for 1D channel-wise dropout (#79545)nn.Module: Public API for stateless / functional module computation (#75834)nn.Module: Support for hooks that run after state dict loading (#76823, #77392)__torch_function__ and __torch_dispatch__
__torch_function__ mode, which allows you to override the meaning of all __torch_function__ overrideable functions within a dynamic scope. (#75154)enable_torch_dispatch_mode, which allows nesting of different __torch_dispatch__ modes. (#75965)__torch_dispatch__ (#73684)super().__torch_dispatch__ with arguments list (#74509, #74720)__torch_function__ fixes (#75484, #75110)__torch_function__ override protocol supporting to some factory functions (#75639)__torch_dispatch__. (#74357)__torch_dispatch__ (#72623, #74577)is_contiguous() to be overridden in __torch_dispatch__ (#77906)functorch. You can run it with functorch.experimental.functionalize(). Example usages can be found here. (#75913, #76083, #76084, #73442, #77285, #73441, #75302, #75818, #75819, #76125, #76318, #77358)torch.library API to allow users to override kernels for existing C++ ops through Python (#75905, #76892)from torch._decomp import register_decomposition, get_decompositions. (#76311, #76814)
ccol_indices and row_indices methods for CSC and BSC tensors. (#77503)to_sparse_csc with support for 2D Strided and 2D CSC input (#77521)to_sparse_bsr with support for 2D CSR input (#77366)index_reduce (#76997, #75981, #76296)sigmoid, exp, sqrt, rsqrt, log, log10, log2, addcmul, abs, addcdiv, sgn, neg , logical_and, angle(#73643, #73776, #73781, #74160, #74161, #74533, #74455, #74827, #74814, #74863, #75123, #76692)sigmoid and tanh (#76289, #74948)kaiser_window , prod (#73734, #75231)bfloat16, conv-bias-activation fusion (#60755)torch.cuda.is_current_stream_capturing (#77789)torch.nn.GRU) (#72692, #73599)torch.lerp) (#76544)ShardedTensor and tensor parallel
torch.tensor is being sharded across multiple GPUs or hosts and a high level APIs for users to specify how to shard, enabling basic tensor ops for ShardedTensor and enabling optimizer for ShardedTensor. In addition, we have added PartialTensor, ReplicatedTensor and checkpoint with ShardedTensor (#63997, #65511, #65671, #65855, #66012, #66351, #66464, #66603, #66604, #67057, #67188, #67199, #67799, #68021, #68096, #68607, #68771, #68786, #68806, #69226, #69493, #69569, #69725, #69874, #69945, #69946, #70145, #70228, #70266, #70331, #70476, #71445, #72062, #72130, #73309, #76360, #76477, #72733, #73392, #76199, #75374, #71624, #74040, #73529, #74941, #73703, #75712, #73873, #75991, #75844, #76824, #76897, #77185, #77191, #76758, #77209, #77214, #77367, #77580, #77626, #77800, #77707, #78056)
FlatParameter to track the information of a flat parameter (#69241)summon_full_params for FSDP. (#71225)no_sync() context manager (#72446)apply() (#72925)full_state_dict (#73324)clip_grad_norm for FSDP (#73405)no_sync() (#73535)full_optim_state_dict (#74215)reshard_flatten_tensor (#75192)scatter_full_optim_state_dict() (#75517)sharded_state_dict and load_sharded_state_dict (#77356)FullStateDictConfig to allow full state dict checkpoint with rank0 only and CPU offload (#75908)always_wrap policy (#73687)torch.jit.freeze (#74178)torch.jit.set_fusion_strategy is now a public API, allowing one to set if they want fusion based on static or dynamic tensor sizes (#72639)tensor.__getitem__() (#73952)torch.jit.save_jit_module_to_flatbuffer (#77870)torch.distributions.wishart.Wishart (#72993)mode property to torch.distributions.Distribution (#76690)foreach flag for torch.optim.{Adadelta, Adagrad, Adamax, Adam, ASGD, NAdam, RAdamSGD, Rmsprop, Rprop, AdamW} (#69980, #69981, #69982, #70295, #70481, #70229, #70230, #70231, #70482, #70483, #70484)torch.softmax and torch.log_softmax (#75833)torch.combinations (#70270)torch.autocast (#75250).set_(storage, offset, size, strides) (#77007)torch.return_types.* as pytree nodes (#75915)torch.return_type (#74199)torch module (#75801)NotImplementedError verbosity for torch.distributions.kl_divergence (#72845)torch.optim.Adagrad (#75968)optim.{Adagrad, Adam, Adamax, AdamW, RAdam}: Updated step in functional optimizers and pass state_steps instead of state (#71333)torch.lerp numerical precision by doing intermediate math in opmath_t (#76062)torch.finfo.tiny to torch.finfo.smallest_normal (#76292)col2im (#73719)stft (#73432)torch.{atan2, dist, logsumexp, log_softmax, norm, polar, put softmax} (#73741, #74205, #75027, #75326, #77421)torch.nn.functional.{cross_entropy, pairwise_dist, nll_loss, normalize} (#73741, #74205)torch.cholesky_inverse (#75033)torch.nn.functional.{embedding,prelu, bilinear, rrelu, logsigmoid} (#77421)torch.nn.BCELoss (#77755)Tensor.__rsub__ (#75326)torch.clamp when bounds are tensors (#74042)torch.nn.functional.{dropout, glu}(#75288, #77186)torch.nn.functional.{leaky_relu, glu, elu, selu, celu} (#75294, #77309, #75297)torch.{linalg.cholesky, cholesky} (#76032)torch.linalg.qr (#76115)torch.cholesky_inverse (#75033)torch.nn.functional.binary_cross_entropy wrt target (#77416)torch.nn.functional.batch_norm when running_{mean,var} have forward grad defined (#73655)torch.nn.functional.max_unpool (#68625)masked_softmax (#71502).pyi.in files exportable from torch/_C/ folder (#74962)pin_memory_device to Dataloader to pin Tensor to the corresponding GPU device (#65402)ForEach L1 and L2 norm by using OpMathType tensor for intermediate results (#68107)nn.init.orthogonal_ no-op for empty input (#75553)nn.{Conv1d, Conv2d, Conv3d}: Added support for complex datatypes (#75310, #75412, #75581)nn.Conv2d: Added bfloat16 support for mkl-dnn backend (#55864)nn.Conv2d: Added support for channels last memory format on CPU for mkl-dnn backend, naive algorithm, and dilated algorithm (#55584, #68101, #70665)nn.EmbeddingBag: Added half precision support on CPU (#74844)nn.FractionalMaxPool*d: Added support 0s in out_size (#73634)nn.Module: Changed to throw error for non-dict inputs to load_state_dict() (#77197)nn.{PixelShuffle, PixelUnshuffle}: Added support for channels last memory format (#50573)nn.PReLU: Enabled fp32/bfloat16 forward and backward for mkl-dnn backend (#60427)F.elu: Improve numerical precision by using opmath and expm1 (#77062)F.{hardshrink, hardsigmoid, hardswish, logsigmoid, smooth_l1_loss, softplus, softshrink}, nn.{BatchNorm, GLU, Upsample}: Add bfloat16 support on CPU (#62558, #63134, #77496, #61944, #76935)uru10x10_to_trt_eval script (#74707)scatter_reduce (#74606,#74607)to_sparse_csr (#77521)to_dense (#74486, #77521)to_sparse (#73642, #77521)sparse_csr_tensor (#74542)__str__ for CSC, BSR, and BSC tensors (#77530, #76650)addmm, addmv, mm (#77615)torch.sparse.sampled_addmm (#68084)torch.sparse.addmm and torch.sparse.mm (#76591)mul (#74266, #77177)sum (#74766)addmm, addmv, triangular_solve (#77255)torch.sparse.sampled_addmm on CUDA (#77243)torch.sparse.sampled_addmm on CPU (#76589)torch.select (#76228)Tensor.to (#76400)torch.empty (#77508)torch.clone (#77512)copy_ (#77605)torch.mm (#73686)torch.sparse.mm (#73075)addmm on CPU (#73076)bool support to coalesce and to_dense (#74495)half support to sparse_mask (#76862)coalesce (#73548)atomicAddNoRet() for all gfx targets. (#75451)ncclAllToAll for ROCm (#75128)num_threads for ROCm, Depthwise kernels, Embedding kernels, Normalization kernels, Softmax kernels, Tensor kernels, Index, Repeat and Sort kernels, Range and Multinomial Kernels (#69942, #72682, #72809, #73543, #73545, #73546, #73549, #73550)sort operator BF16 support (#72854)topk operator for bfloat16 dtype (#71913)init in torch.cuda (#72404)allow_tf32_cublas (#77114)torch.{nn.PReLU, nn.Upsample,nn.GLU, randperm, multinomial, poisson, nn.ELU, nn.SELU, nn.CELU, nn.LogSigmoid, nn.Hardsigmoid, nn.Hardshrink, nn.Softshrink, nn.Hardswish, nn.Softplus, nn.SmoothL1Loss, histc, atan2, logcumsumexp, diag, fmod, cumsum, cumprod, nn.utils.weight_norm , nn.BatchNorm2d} and allow autocast enabled (#63634, #58297, #61944, #63215 , #62546, #63134, #72694, #61897, #73845, #74410, #68725)
torch{norm,argmax,argmin, scatter, gather} performance on CPU (#64479, #64478)torch.nn.functional{log_softmax``, softmax} performance on CPU (#73953)conv_transpose3d (#76888)PReLU (#60427)torch.nn.init to list of functions overridable by __torch_function__ (#76014)torch.Tensor(#73850)torch.jit.trace now treats tensor.numel() as aten::numel, instead of a constant value (#74081)torch.jit.freeze to allow for folding arguments that can be promoted to floating point (eg integer tensor arguments) (#73278)torch.jit.save and torch.jit.load are now supported for meta tensors ( aka torch.Tensor(device="meta")) (#73435)Linear-Bn1d (#72431, #72796)qint32 quantization support (#72472)get_default_qconfig_dict&get_default_qat_qconfig_dict (#73528)Matmul Op (Naive Implementation) (#71783)Softmax Op (Naive Implementation) (#75415)Softmax Op (#75799)torch.matmul quantization (#72444)conv1d and its fusion variants in QAT (#74506)prepare_*fx from training/eval modes (#75401)default_affine_fixed_qparams_observer and default_symmetric_fixed_qparams_observer (#76637)opset_version to 13. The previous default was 9. To get the old behavior, just specify opset_version=9 when calling torch.onnx.export. Going forward we plan to update the default regularly to "latest as of 18 months ago". (#73898)torch.minimum with different dtype combinations (#76022)Expand shape inference (#72985)matmul shape inference (#72990)topk export with non-int64 k (#73761)numel tracing (#74081)onnx::ReduceProd (#74082)Squeeze and Unsqueeze (#73104)Torch.Package metadata (#74610)TORCH_DISTRIBUTED_DEBUG implementation (#73166)threading.Thread (#74462)verify_params_across_processes (#74113)nn.functional.all_gather/reducescatter/gather (#75276)ucc_lib available (#69564)summon_full_params when not sharded (#72572)all_gather stream in summon_full_params (#73314)unflatten_parameter in _summon_full_parameters (#72467)summon_full_params in get_full_params (#73242)state_dict (#73323)load_local_state_dict (#73325)summon_full_params a public method (#73116)fsdp_modules() (#73553)state_dict_type (#73716)summon_full_params (#73903)summon_full_params (#73904)_lazy_init() in rebuild full params (#74263)named_parameters() for clean names in summon_full_params() (#74333)summon_full_params context, similar to named_params in named_buffers (#74517)full_optim_state_dict (#74879)full_optim_state_dict (#74912)state_dict hooks for FlatParamsWrapper even if params_list is empty (#74860)apply_to_tensors support OrderedDict type (#75560)rank0_only to full_optim_state_dict() (#75516)summon_full_params a static method (#75423)apply_for_tensors (#76265)compute_device (#76664)_get_param_name_to_param to be faster(#76665)Optim state dict (#76671)ignored_modules (#76784)ignored_modules (#76994)FSDP.forward (#76899)_get_param_to_unflat_param_names() for shared params (#75430)Note truncated.
This argument used to default to 100 in PyTorch 1.10.2, but was deprecated (previously you would see a deprecation warning if you didn’t explicitly pa…
We are excited to announce the release of PyTorch 1.11. This release is composed of over 3,300 commits since 1.10, made by 434 contributors. Along with 1.11, we are releasing beta versions of TorchData and functorch. We want to sincerely thank our community for continuously improving PyTorch.
You can check the blogpost that shows the new features here.
deepcopy to correctly copy all attributes on Tensor objects (#65584)This change ensures that the deepcopy operation on Tensor properly copies all the attributes (and not just the plain Tensor properties).
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> a = torch.rand(2) a.foo = 3 torch.save(a, "bar") b = torch.load("bar") print(b.foo)
</pre></sub></td>
<td><sub><pre lang="python">
a = torch.rand(2) a.foo = 3 torch.save(a, "bar") b = torch.load("bar") print(b.foo)
</pre></sub></td>
</tr>
</table> </p>
steps argument is no longer optional in torch.linspace and torch.logspaceThis argument used to default to 100 in PyTorch 1.10.2, but was deprecated (previously you would see a deprecation warning if you didn’t explicitly pass in steps). In PyTorch 1.11, it is not longer optional.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a = torch.linspace(1, 10)
</pre></sub></td>
<td><sub><pre lang="python">
a = torch.linspace(1, 10, steps=100) </pre></sub></td> </tr> </table> </p>
torch.hub.import_module function that was mistakenly public (#67990)This function is not intended for public use.
If you have existing code that relies on it, you can find an equivalent function at torch.hub._import_module.
aten operators that they actually used (#68247, #68687, #68688, #68714, #68689, #68690, #68697, #68691, #68692, #68693, #69840)When you #include a header from the C++ frontend, you can no longer assume that every aten operators are transitively included. You can work around this by directly adding #include <ATen/ATen.h> in your file, which will maintain the old behavior of including every aten operators.
c10::List and c10::Dict move constructors have been removed (#69370)The semantics have changed from "make the moved-from List/Dict empty" to "keep the moved-from List/Dict unchanged"
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="cpp"> c10::List<string> list1({"3", "4"}); c10::List<string> list2(std::move(list1)); std::cout << list1.size() // 0 </pre></sub></td> <td><sub><pre lang="cpp"> c10::List<string> list1({"3", "4"}); c10::List<string> list2(std::move(list1)); // calls copy ctr std::cout << list1.size() // 2 </pre></sub></td> </tr> </table> </p>
THCeilDiv function and corresponding THC/THCDeviceUtils.cuh header (#65472)As part of cleaning up TH from the codebase, the THCeilDiv function has been removed. Instead, please use at::ceil_div, and include the corresponding ATen/ceil_div.h header
THCudaCheck (#66391)You can replace it with C10_CUDA_CHECK, which has been available since at least PyTorch 1.4, so just replacing is enough even if you support older versions
THCudaMalloc(), THCudaFree(), THCThrustAllocator.cuh (#65492)If your extension is using THCThrustAllocator.cuh, please replace it with ATen/cuda/ThrustAllocator.h and corresponding APIs (see examples in this PR).
This PR also removes THCudaMalloc/THCudaFree calls. Please use c10::cuda::CUDACachingAllocator::raw_alloc(size)/raw_delete(ptr), or, preferably, switch to c10:cuda::CUDaCachingAllocator::allocate which manages deallocation. Caching allocator APIs are available since PyTorch 1.2, so just replacing it is enough even if you support older versions of PyTorch.
libaot_compiler.so (#66227)Building aot_compiler.cpp as a separate library is not necessary, as it’s already included in libtorch.so.
You can update your build system to only dynamically link libtorch.so.
typing.Union type unsupported for mobile builds (#65556)typing.Union support was added for TorchScript in 1.10. It was removed specifically for mobile due to its lack of use and increase in binary size of PyTorch for Mobile builds.
torch.distributed.rpc: Final Removal of ProcessGroup RPC backend (#67363)ProcessGroup RPC backend is deprecated. In 1.10, it threw an error to help users update their code, and, in 1.11, it is removed completely.
The backend type “PROCESS_GROUP” is now deprecated, e.g.
torch.distributed.rpc.init_rpc("worker0", backend="PROCESS_GROUP", rank=0, world_size=1)
and should be replaced with:
torch.distributed.rpc.init_rpc("worker0", backend="TENSORPIPE", rank=0, world_size=1)
getitem in FX Graph Mode Quantization (#66647)getitem used to be quantized in FX Graph Mode Quantization, and it is no longer quantized. This won’t break any models but could result in a slight difference in numerics.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> from torch.ao.quantization.quantize_fx import convert_fx, prepare_fx class M(torch.nn.Module): def init(self): super().init() self.linear = torch.nn.Linear(5, 5) def forward(self, x): x = self.linear(x) y = torch.stack([x], 0) return y[0] m = M().eval() m = prepare_fx(m, {"": torch.ao.quantization.default_qconfig}) m = convert_fx(m) print(m)
</pre></sub></td>
<td><sub><pre lang="python">
from torch.ao.quantization.quantize_fx import convert_fx, prepare_fx class M(torch.nn.Module): def init(self): super().init() self.linear = torch.nn.Linear(5, 5) def forward(self, x): x = self.linear(x) y = torch.stack([x], 0) return y[0] m = M().eval() m = prepare_fx(m, {"": torch.ao.quantization.default_qconfig}) m = convert_fx(m) print(m)
zero_point=0, qscheme=torch.per_tensor_affine)
linear_input_zero_point_0, torch.quint8)
</pre></sub></td>
</tr>
</table> </p>
fuse_modules for PTQ fusion and fuse_modules_qat for QAT fusion (#69878, #71956)There are two types of fusion supported by fuse_modules api: PTQ and QAT fusion. Previously we relied on module.training to decide which mode user wanted, but this was a misuse of the training attribute since that is not the intended purpose. This PR removes the dependency on module.training and uses separate APIs to make the fusion requested by the user explicit.
Previously, fuse_module used to support both cases and distinguished PTQ/QAT fusion based on module.training, but now fuse_module only supports the PTQ fusion. So, in the case when user wants to do QAT fusion, they need to change the call to fuse_modules_qat, instead of using fuse_modules, otherwise, they would silently get unwanted fusion results (PTQ fusion), or if the model is in training mode, it might result in error.
Note: Currently it is still enforced that if the model is in eval mode, only PTQ fusion can be used; if the model is in training mode, then only QAT fusion can be used. In the future this constraint will be relaxed.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> import torch from torch.ao.quantization import fuse_modules class M(torch.nn.Module): def init(self): super().init() self.conv = torch.nn.Conv2d(3, 3, 3) self.bn = torch.nn.BatchNorm2d(3) def forward(self, x): return self.bn(self.conv(x)) m = M().train() m = fuse_modules(m, ["conv", "bn"]) print(type(m.conv)) m = M().eval() m = fuse_modules(m, ["conv", "bn"]) print(type(m.conv)) <class 'torch.nn.intrinsic.modules.fused.ConvBn2d'> <class 'torch.nn.modules.conv.Conv2d'> </pre></sub></td> <td><sub><pre lang="python"> import torch from torch.ao.quantization import fuse_modules class M(torch.nn.Module): def init(self): super().init() self.conv = torch.nn.Conv2d(3, 3, 3) self.bn = torch.nn.BatchNorm2d(3) def forward(self, x): return self.bn(self.conv(x)) m = M().train()
m = fuse_modules_qat(m, ["conv", "bn"]) print(type(m.conv)) m = M().eval() m = fuse_modules(m, ["conv", "bn"]) print(type(m.conv))
<class 'torch.nn.intrinsic.modules.fused.ConvBn2d'> <class 'torch.nn.modules.conv.Conv2d'> </pre></sub></td> </tr> </table> </p>
f arg from onnx.export_to_pretty_string (#69546)The arg has always been ignored. Simply remove it from your code.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.onnx.export_to_pretty_string(model, inputs, "file_name") </pre></sub></td> <td><sub><pre lang="python"> torch.onnx.export_to_pretty_string(model, inputs) </pre></sub></td> </tr> </table> </p>
use_external_data_format arg from onnx.export (#67809)The arg has been deprecated and ignored since 1.10. The external data format is now used automatically if and only if the exported file would exceed protocol buffer’s file size limit. Simply remove it from your code.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name, use_external_data_format=True) </pre></sub></td> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name) </pre></sub></td> </tr> </table> </p>
example_outputs arg from torch.onnx.export (#67809)The arg has been deprecated and ignored since 1.10. The provided model is instead executed once to produce example outputs. Simply remove it from your code.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name, exaple_outputs=(foo,)) </pre></sub></td> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name) </pre></sub></td> </tr> </table> </p>
enable_onnx_checker arg from onnx.export (#67276)The arg has been deprecated and ignored since 1.10. The ONNX checker is always enabled. If it fails, onnx.CheckerError will be raised. Users can catch and ignore that exception.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name, enable_onnx_checker=False) </pre></sub></td> <td><sub><pre lang="python"> try: torch.onnx.export(model, inputs, f_name) except torch.onnx.CheckerError: pass # ignore error </pre></sub></td> </tr> </table> </p>
onnx.utils.ONNXCheckerError to onnx.CheckerError (#66644)Previously the documentation was incorrect and stated ONNXCheckerError was in the onnx module, so this moves the class to the originally intended module and brings the code in line with the documentation. The new name is shorter and less redundant with the module name.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> except torch.onnx.utils.ONNXCheckerError: </pre></sub></td> <td><sub><pre lang="python"> except torch.onnx.CheckerError: </tr> </table> </p>
_retain_param_name arg from onnx.export (#67276)The arg has been deprecated and ignored since 1.10. Param names are now always retained. Simply remove it from your code. If you want to remove param names, you can do so by editing the exported ONNX model.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.onnx.export(model, inputs, f_name, _retain_param_name=True) </pre></sub></td> <td><sub><pre lang="python"> torch.onnx.export(model, inputs, f_name) </tr> </table> </p>
x.T on tensors of dimension other than 0 or 2 (#64180)x.T only accepts tensors with 0 or 2 dimensions. Calling x.T on tensors with a different number of dimensions has been deprecated.
<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> a = torch.ones(2, 3, 4) a.T.size()
</pre></sub></td>
<td><sub><pre lang="python">
a = torch.ones(2, 3, 4) a.T.size()
x.T on tensors of dimension other than 2x.mT to transpose batches of matrices or x.permute(*torch.arange(x.ndim - 1, -1, -1))</tr>
</table> </p>
torch.ao.quantization.QConfigDynamic is deprecated and going to be removed in next the release, please use torch.ao.quantization.QConfig instead (#69875, #69864)<p align="center"> <table align="center"> <tr><th>1.10.2</th><th>1.11.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> qconfig = torch.ao.quantization.QConfigDynamic(...) </pre></sub></td> <td><sub><pre lang="python"> qconfig = torch.ao.quantization.QConfig(...) </tr> </table> </p>
set_deterministic_debug_mode and get_deterministic_debug_mode (#67778, #66233)torch.fft.ifftn and torch.fft.hfftn (#63890)Wishart distribution to torch.distributions (#70377)torch and torch.linalg modules. PyTorch implements over 90% of the operators defined by the Python Array API, including the torch.from_dlpack operation for improved DLPack support (#60627)torch.testing from prototype to beta (#69668)torch.utils.checkpoint implementation that does not use reentrant autograd (can be toggled with the new use_reentrant flag) (#69508)batched_grad parameter to autograd.grad to allow batched gradient computation (#65564)ctx.save_for_forward function to autograd.Function (#71569)autograd.forward_ad.unpack_dual returns a named tuple instead of plain tuple (#68062, #68628)IS_LINUX and IS_MACOS global vars for cpp extensions building (#69093)linalg.matrix_exp operation (see the docs here) (#62715)linalg.cross operation (see the docs here) (#63285)linalg.diagonal operation, an alias for torch.diagonal (see the docs here) (#70599)linalg.lu_factor operation (see the docs here) (#66933)torch.nn.utils.rnn.{unpack_sequence,unpad_sequence} functions (#66550)torch.sparse.sampled_addmm for CSR Tensors on GPU (#68007)lcm, i0e, i1e, ndtri, efcx, digamma, trigamma, lgamma (#70663)exp2, erfc, erfinv and entr (#71295)polygamma (#71162)addmv_out (#61407)vulkan_perf_test benchmark binary to benchmark Vulkan ops under various input conditions. (#67230)bool and int dtypes in the copy kernel by default when using Tracing Based Selective Build (#69106, #69297)FullyShardedDataParallel
DistributedDataParallel
torch.jit.freeze() and torch.jit.optimize_for_inference on functions that are not forward (#68668, #69367)torch.jit.freeze to work on for sparse COO tensors (#69614)torch.jit.script(), torch.jit.freeze() and serialization for tensors in Compressed Sparse Row (CSR) format (#69555)torch.jit.fuser through the now public torch.jit.set_fusion_strategy . (#72937)torch.jit.set_fusion_strategy (#72036)torch.nn.functional.grid_sample 2d operator (#66879)torch.quantize_per_tensor_dynamic operator (#68004)torch.nn.Embedding and torch.nn.EmbeddingBag
torch.nn.Conv3d in qnnpack, for use in quantization
nn.Module calls as ONNX local functions (#66140, #67803)torch.Tensor.view(dtype): enable all dtype combinations (#66493)torch.diff by adding support for n greater than 1 (#67260)torch.movedim to handle scalar as no-op (#69537)cartesian_prod: fixed a warning in the docs example (#68753)max_unpool{}d operators (#67328)torch.distributions
torch.distributions (#71375)non-negative constraint in exponential distribution (allowing it to include zero). (#67184)kl divergence between normal and laplace distribution. (#68807)torch.Tensor.real for real-valued tensors (#71718)torch.logaddexp, torch.logaddexp2, torch.remainder: added BFloat16 support on CPU (#63621)torch.bucketize and searchsorted: added Half precision support (#67077)torch.slice_scatter,torch.select_scatter, torch.diagonal_scatter ops (#64430)torch.scatter_reduce a public API (#68580, #73125)hfftn (#66127)MaybeOwned<IValue> (#68157)set_to_none option for zero_grad() to C++ API (#68801)TORCH_CPP_LOG_LEVEL, that you can use to toggle the log level in the c10 library (#71746)torch.autograd.graph.saved_tensor_hooks (#70932)torch.{col2im,im2col} (#68199)torch.scatter_reduce (#71788)torch.{remainder,fmod} (#69908)strategy flag to autograd.functional.{Jacobian, Hessian} to enable vectorized computation (#67041, #66292)check_backward_ad flag to torch.autograd.gradcheck to be able to skip backward mode AD checks (#65040)native_functions.yaml in many core files (#64499, #66914, #64172, #64171, #66620, #66793, #66913, #66794, #64169, #64173, #64170, #67735)-Wno-unused-variable compliant (#66041)packaging in torch_version (#71345)Sequence and Mapping for utils.data.default_collate (#68779)num_samples to RandomSampler when replacement is False (#71568)utils.data.default_collate (#71065)ForEach L1 & L2 norm (#62646)linalg.matrix_rank (docs) and linalg.pinv (docs) operations now support specifying absolute and relative tolerances for better handling of singular values (#63102)channels_last support for ChannelShuffle (#50247)nn.{AdaptiveLogSoftmaxWithLoss, Bilinear, Conv*d, ConvTranspose*d, CrossEntropyLoss, CTCLoss, Fold, FractionalMaxPool3d, GaussianNLLLoss, GRU, GRUCell, InstanceNorm*d, LSTM, LSTMCell, MarginRankingLoss, MultiheadAttention, MultiLabelSoftMarginLoss, RNN, RNNCell, Transformer, TransformerDecoderLayer, TransformerEncoderLayer} (#69054, #69539, #70506, #71055, #70092, #64909, #69732, #69783, #70236, #65323, #71056, #64975, #67176, #70590, #65690, #70977, #70597, #70322, #69291)BFloat16 support on CPU to nn.{AdaptiveAvgPool2d, AdaptiveMaxPool2d, AvgPool2d, MaxPool2d} (#56902, #66929, #66927, #56903)maximize support to optim.{Adam, AdamW, SGD} (#68164, #70146, #67847, #68733, #71023)F.interpolate: Add nearest-exact mode to fix off-by-one error in nearest mode (#64501)F.interpolate: Added support for anti-aliasing to bilinear and bicubic algorithms (#70930, #68819, #65142, #69318)F.interpolate: Improved error message for invalid shapes (#66417)nn.Conv*d: Accepts 0-sized channel inputs (#66256)nn.LogSigmoid: Used log1p for improved precision (#66441)nn.Module: Added flag for removing duplicates from parameters (#71542)nn.Module: Added register_module alias for registering a sub-module (#65174)nn.ModuleList: Supported concatenation (#70887)nn.MultiheadAttention: Added flag to optionally average output attention weights across heads (#70055)nn.ParameterDict: Supported full set of dict methods (#69403)nn.{RNN, GRU}: Allowed hidden_size to be 0 (#70556)nn.Sequential: Added append method (#71326)nn.Upsample: Exposed recompute_scale_factor (#66419)nn.ZeroPad2d: Added extra_repr for printing purposes (#69206)optim.{ChainedScheduler, SequentialLR}: Added optimizer attribute (#67406, #69817)optim.swa_utils.AveragedModel: Added use_buffers flag for averaging buffers in addition to parameters (#65921, #71763)fx.Graph’s code generation function, including support for setting a breakpoint in the generated code (#67139)torch.triangular_solve, torch.addmv, torch.addmm, torch.add for all arguments on CPU (#62180, #61536, #65606, #64391)torch.triangular_solve, torch.addmv, torch.addmm, torch.add for all arguments on GPU (#61407, #61858, #63511, #63948)torch.empty, torch.resize_, torch.copy_, torch.randn_like, torch.clone (#63509, #63510, #68083, #70581)transpose (#70582)zeros_like (#68108)to_sparse (#66774)nanmedian result (#68591)igamma kernel instantiations (#70666)compare kernels by unifying them (#69111)bernoulli tensor tensor kernel instantiations (#70169)cub::FutureValue to simplify 64bit indexing split of cub scan (#66711)hascuSOLVER flag to Context (#69825)CUDACachingAllocator (#69174)masked_softmax perf for element_size is not 8 (#70271)TensorCompare.cu (#68835)interpolation (#72066)pow kernels for non-existent case (#70017)bmm and baddbmm (#66636)NNAPI
quantized::mul and quantized::convtranspose2d to converter (torch.backends._nnapi.prepare.convert_model_to_nnapi) (#63913, #63914)int32 and qint16 type in Torchscript expressions (#70197, #70621)CoreML
torch.distributed
TCPStore’s socket implementation (#68225)ncclAvg for reductions (#62835)NCCL comms in constructor (#65173, #66393)ProcessGroup and Work (#66338)c10d extension Backend class attr the same way as builtin ones (#66991)ProcessGroup trampoline (#67236)bfloat16 support for NCCL (#67843)c10d TCP store race condition with mutex (#68499)ncclUniqueId store broadcast error (#68597)FileStore destructor (#68603)ProcessGroupNCCL (#66745)ProcessGroupNCCL (#70029)gather_object on NCCL (#71623)allreduce_coalesced for ProcessGroupNCCL (#62140)deleteKey for FileStore (#69953)TSAN issue in TCPStore (#69590)DistributedDataParallel
torch.distributed.rpc
torch.distributed.autograd
torch.distributed.elastic
torch.distributed.run (#66179)CudaFusionGroup, and addition of a graph segmentation cache to the hierarchical caching system. (#63745, #65137, #63745, #65137)profile_ivalue to convert dynamic scalar into compile time constants in NVFuser. (e.g. reduction axes). (#63745, #65137)torch.jit.trace for tracing already JITted subgraphs(#59949)torch.jit.freeze now can preserve attributes of submodules - previously, it was only possible to prevent inlining of attributes of the top level module.(#66102)torch.jit.freeze now coalesces consecutive calls to torch.concat into a single call (#67000)None into an undefined Tensor(#67793)torch.jit.script now recognizes union of scalars as a JIT NumberType (#66591)torch.jit.optimize_for_inference, there is a new graph pass to precompute transposes for linear layers. (#65631, 68024)torch.jit.freeze, there is a new pass where we concat together multiple linear layers with same input Tensor (different weight/bias) (#63198, #68024)torch.Tensor.__rsub__ in normalize_ops JIT pass(#65014)torch.ao.FakeQuantize now supports fp32/fp16 zero_point. (#65836)torch.ops.quantized.add now supports broadcasting (#66049)torch.Tensor.dequantize now supports fp16 + cuda (#67234)torch.nn.GELU (#69968)torch.nn.quantized.functional.hardsigmoid supports an inplace flag (#65740)torch.nn.Linear + torch.nn.BatchNorm1d fusion for PTQ (#66484)torch.ao.quantization.quantize_fx.convert_fx to accept qconfig_dict to skip quantization (#66878)torch.nn.qat.dynamic.modules.Linear module (#67325)torch.nn.ConvTranspose{n}d + torch.nn.BatchNorm{n}d fusion support (#70022)torch.ao.quantization.prepare_qat with allow_list argument, to allow custom mapping and custom QAT module (#65119)torch.ao.quantization.default_replay_qconfig which allows observer reuse for torch.reshape in FX graph mode quantization (#69249)ir_version of the exported model based on opset_version. This increases the odds that the exported ONNX model will be usable. Before this change, we were setting the IR version to a hard-coded value which may be higher than what the model consumer supports. (#67803)torch.reciprocal to ONNX Reciprocal operator instead of Div(1, x) (#67271)beta!=1 in softplus (#66146)tensor.shape in tracing mode (#66142)instance_norm in training mode (#64375)aten::foo is allowed as well as “foo”). (#67810)prim namespace (#66139)OneHot, bool for Einsum (#66147)all_path function(#65602)multiply, subtract, divide} (#65937)torch.save when saving storages that view same data with different type (#66949)torch.save error if storages are unallocated (#68787)k out-of-bounds in torch.kthvalue (cpu kernel) (#68863)inference_mode decorator: with inference_mode(mode=False) used to ignore the mode argument and always set inference mode. (#68617)cdist_backward in the case when cdist inputs are not contiguous (#70016)cdist error message typo (#70178)scatter for empty indexes (#70662)torch.{unique, unique_consecutive} out of bound (#71540)torch.isin in the case when inputs are non-contiguous on CPU (#70659)hsplit vsplit dsplit crash when section is 0 (#69342)torch.gradient ignores dim argument when checking edge_order (#67926)TransformedDistribution.icdf should perform validation after applying the inverse transformation rather than before. (#71393)torch.all and torch.any internal assert error with requires_grad=True (#65714)torch.logsumexp type promotion: promote integral inputs to floating for(#63393)at::Tensor::print() linking error (#69615)torch.utils.checkpoint API (#71169)torch.nn.functional.conv_transpose3d backward when grad_out is non-contiguous (#67829)Tensor.copy_ forward AD to handle broadcasting (#69592)autograd.Function when non-Tensor argument precedes tensor argument (#71530)autograd.Function forward AD when forward is a no-op to no longer raise an internal error (#71531)pocketfft is not found and at_mkl is not enabled (#67909)_GLIBCXX_USE_CXX11_ABI (#72081)torch.autograd.gradcheck to generate valid inputs for forward AD computation for complex functions (#68001)torch.Tensor.copy_ transpose path for tensors with conjugate or negative bit set (#69026)torch.Tensor.copy_ behavior for the case when two conjugated or negated tensors of the same dtype (one or both of which are non-contiguous) are copied into each other (#68963)ProcessException picklable (#70118)pin_memory_thread (#71579)nn.AdaptiveAvgPool*d: Throws an error for negative output_size (#70488)nn.Conv1d: Fixed for 1D convolution on MKL-DNN backend (#68166)nn.CrossEntropyLoss: Fixed for usage of weight, ignore_index, and label_smoothing together (#69511)nn.Fold: Checked that block height and width are positive (#69048)nn.LayerNorm: Fixed incorrect result on CUDA when gamma or bias are missing (#69210)nn.LayerNorm: Avoided overflow by doing computation in float for half (#66920)nn.Module: Throws a proper error message from load_state_dict for non-tensor values (#70596)nn.ModuleList: Fixed incorrect return type in __getitem__ (#69083)nn.MultiheadAttention: Used query dtype for mask type (#68077)nn.NLLLoss: Fixed backward computation with negative weights (#64572)nn.{RNN, GRU}: Fixed RNN modules with input shapes containing-0 in CUDA (#71696)nn.utils.rnn.pad_sequence: Fix regression to support tuples for padding (#72436)optim._LrScheduler: Fixed print formatting (#68338)optim.ChainedScheduler: Fixed get_last_lr() (#69112)optim.CosineAnnealingWarmRestarts: Fixed ordering bug when last_epoch > 0 (#64758)optim.SequentialLR: Updated _last_lr on step (#70558)torch.layout as arg (#66048)concrete_args (#59569)GraphModule.delete_all_unused_submodules deletes submodules from called leaf modules (#66430)torch.fx.subgraph_rewriter.replace_pattern mechanism so that multiple one-liner instances of the pattern are captured correctly (#66442)to_folder not saving dtype (#69983)default_value arg to fx.Graph.placeholder and fix split_module (#71016)pinv_jvp and pinv_backward (#67948)_cudnn_impl functions (#70406)mem_get_info when querying on a device other than the current device (#69640)torch.utils.benchmark.Timer (#70050)OperatorHandle destructor, so that the symbol shows up in windows builds (#70033)torch.utils.tensorboard parsing JIT graph incorrectly (#65692)NNAPI (#70847)MTLCreateSystemDefaultDevice returns nil (#66859)irange not having a header included in Metal (#66877)type_parser. (#71341)flatbuffer_loader. (#71500)pytorch_jni_common (#71508)CoreML (#67737)torch.distributed
GLOO_SOCKET_IFNAME_ENV (#68933)DistributedDataParallel
torch.distributed.elastic
rdzv_handler.shutdown() on premature agent failures (#67749)CompilationUnit, resulting in memory leaks when class objects were in JIT graphs. (#65442)torch.jit.optimize_for_inference did not torch.jit.freeze a module when passed a a non-frozen module (#71436)torch.jit.freeze ed module ran the wrong graph (#68316)torch.split , resulting in invalid optimizations in various JIT optimization passes (#69745)torch.autocast together with autodiff (module.backwards()) in a JIT graph had the wrong number of arguments and would error out. (#67648)torch.jit.trace (#68242)torch.jit.freeze ops are converted to MKLDNN(#66628)pickle version.(#69807)torch.jit.script fails when comments in function has less indent than surrounding code (#70227)torch.jit.script) code (#69645)torch::jit::Function::call is only partially overridden in class torch::jit::GraphFunction (4bf1be898d)torch.nn.functional.interpolate for quantized tensors (#65570)torch.nn.Conv{n}d (#61647, #71426)torch.quantize_per_tensor on non floats more specific (#66050)torch.nn.Embedding conversion with unsupported dtype: make error message clearer (#66051)torch.nn.qat.EmbeddingBag from_float error message (#66989)Note truncated.
This release is meant to deploy additional fixes not included in 1.10.1 release:
This release is meant to deploy additional fixes not included in 1.10.1 release:
This release is meant to fix the following issues (regressions / silent correctness):
This release is meant to fix the following issues (regressions / silent correctness):
The release tracker should contain all relevant pull requests related to this release as well as links to related issues
This is the end of the deprecation cycle for both of these functions. You should be using torch.use_deterministic_algorithms andtorch.are_deterministi…
We are excited to announce the release of PyTorch 1.10. This release is composed of over 3,400 commits since 1.9, made by 426 contributors. We want to sincerely thank our community for continuously improving PyTorch.
PyTorch 1.10 updates are focused on improving training and performance of PyTorch, and developer usability. Highlights include:
torch.special, and nn.Module Parametrization, have moved from beta to stable.You can check the blogpost that shows the new features here.
torch.any/torch.all behavior changed slightly to be more consistent for zero-dimension, uint8 tensors. (#64642)These two functions match the behavior of NumPy, returning an output dtype of bool for all support dtypes, except for uint8 (in which case they return a 1 or a 0, but with uint8 dtype). In some cases with 0-dim tensor inputs, the returned uint8 value could mistakenly take on a value > 1. This has now been fixed.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.all(torch.tensor(42, dtype=torch.uint8)) tensor(1, dtype=torch.uint8) torch.all(torch.tensor(42, dtype=torch.uint8), dim=0) tensor(42, dtype=torch.uint8) # wrong, old behavior </pre></sub></td> <td><sub><pre lang="python"> torch.all(torch.tensor(42, dtype=torch.uint8)) tensor(1, dtype=torch.uint8) torch.all(torch.tensor(42, dtype=torch.uint8), dim=0) tensor(1, dtype=torch.uint8) # new, corrected and consistent behavior </pre></sub></td> </tr> </table> </p>
torch.{is,set}_deterministic (#62158)This is the end of the deprecation cycle for both of these functions. You should be using torch.use_deterministic_algorithms andtorch.are_deterministic_algorithms_enabled instead.
tensor.conj() now returns a view tensor that aliases the same memory and has conjugate bit set (#54987, #60522, #66082, #63602).This means that .conj() is now an O(1) operation and returns a tensor that views the same memory as tensor and has conjugate bit set. This notion of conjugate bit enables fusion of operations with conjugation which gives a lot of performance benefit for operations like matrix multiplication. All out-of-place operations will have the same behavior as before, but an in-place operation on a conjugated tensor will additionally modify the input tensor.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
import torch x = torch.tensor([1+2j]) y = x.conj() y.add_(2) print(x) tensor([1.+2.j]) </pre></sub></td> <td><sub><pre lang="python"> import torch x = torch.tensor([1+2j]) y = x.conj() y.add_(2) print(x) tensor([3.+2.j]) </pre></sub></td> </tr> </table> </p>
Note: You can verify if the conj bit is set by calling tensor.is_conj(). The conjugation can be resolved, i.e., you can obtain a new tensor that doesn’t share storage with the input tensor at any time by calling conjugated_tensor.clone() or conjugated_tensor.resolve_conj() .
Note that these conjugated tensors behave differently from the corresponding numpy arrays obtained from np.conj() when an in-place operation is performed on them (similar to the example shown above).
tensor.conj().neg() returns a view tensor that aliases the same memory as both tensor and tensor.conj() and has a negative bit set (#56058).conjugated_tensor.neg() continues to be an O(1) operation, but the returned tensor shares memory with both tensor and conjugated_tensor.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
x = torch.tensor([1+2j]) y = x.conj() z = y.imag z.add_(2) print(x) tensor([1.+2.j]) </pre></sub></td> <td><sub><pre lang="python"> x = torch.tensor([1+2j]) y = x.conj() z = y.imag print(z.is_neg()) True z.add_(2) print(x) tensor([1.-0.j]) </pre></sub></td> </tr> </table> </p>
tensor.numpy() now throws RuntimeError when called on a tensor with conjugate or negative bit set (#61925).Because the notion of conjugate bit and negative bit doesn’t exist outside of PyTorch, calling operations that return a Python object viewing the same memory as input like .numpy() would no longer work for tensors with conjugate or negative bit set.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
x = torch.tensor([1+2j]) y = x.conj().imag print(y.numpy()) [2.] </pre></sub></td> <td><sub><pre lang="python"> x = torch.tensor([1+2j]) y = x.conj().imag print(y.numpy()) RuntimeError: Can't call numpy() on Tensor that has negative bit set. Use tensor.resolve_neg().numpy() instead. </pre></sub></td> </tr> </table> </p>
TypeError instead of RuntimeError when assigning to a Tensor’s grad field with wrong type (#64876)Setting the .grad field with a non-None and non-Tensor object used to return a RuntimeError but it now properly returns a TypeError. If your code was catching this error, you should simply update it to catch a TypeError instead of a RuntimeError.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> try: # Assigning an int to a Tensor's grad field a.grad = 0 except RuntimeError as e: pass </pre></sub></td> <td><sub><pre lang="python"> try: a.grad = 0 except TypeError as e: pass </pre></sub></td> </tr> </table> </p>
autograd.grad are empty (#52016)Calling autograd.grad with an empty list of inputs used to do the same as backward. To reduce confusion, it now raises the expected error. If you were relying on this, you can simply update your code as follows:
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> grad = autograd.grad(out, tuple()) assert grad == tuple() </pre></sub></td> <td><sub><pre lang="python"> out.backward() </pre></sub></td> </tr> </table> </p>
autograd.gradcheck and autograd.gradgradcheck are now kwarg-only (#65290)These two functions now have a significant number of optional arguments controlling what they do (i.e., eps, atol, rtol, raise_exception, etc.). To improve readability, we made these arguments kwarg-only. If you are passing these arguments to autograd.gradcheck or autograd.gradgradcheck as positional arguments, you can update your code as follows:
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.autograd.gradcheck(fn, x, 1e-6) </pre></sub></td> <td><sub><pre lang="python"> torch.autograd.gradcheck(fn, x, eps=1e-6) </pre></sub></td> </tr> </table> </p>
detach_) now errors for views that return multiple outputs (#58285)This change is finishing the deprecation cycle for the inplace-over-view logic. In particular, a few things that were warning are updated:
* `detach_` will now raise an error when invoked on any view created by `split`, `split_with_sizes`, or `chunk`. You should use the non-inplace `detach` instead.
* The error message for when an in-place operation (that is not detach) is performed on a view created by `split`, `split_with_size`, and `chunk` has been changed from "This view is an output of a function..." to "This view is the output of a function...".
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> b = a.split(1)[0] b.detach_() </pre></sub></td> <td><sub><pre lang="python"> b = a.split(1)[0] c = b.detach() </pre></sub></td> </tr> </table> </p>
In-place on the unpacked SavedVariables used to be ignored. They are now properly detected which can lead to errors saying that a variable needed for backward was modified in-place. This is a valid error and the user should fix this by cloning the unpacked saved variable before using it.
No internal formula will trigger this, but it might be triggered by user custom autograd.Function if the backward modifies a saved Tensor inplace and you do multiple backwards. This used to silently return the wrong result and will now raise the expected error.
__torch_function__ handling checks (#63967)This fixes the has_torch_function*() checks throughout torch.nn.functional to correctly pass in optional tensor arguments; prior to this fix, handle_torch_function() was not called for these optional tensor arguments. Previously, passing a tensor-like object into a function that accepts an optional tensor might not trigger that object's __torch_function__. Now, the object's __torch_function__ will be triggered as expected.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> import torch import torch.nn.functional as F class TestTensor(object): def init(self, weight): self.weight = weight def torch_function(self, func, _, args=(), kwargs=None): print(func) print(func == F.group_norm)
features = TestTensor(torch.randn(3,3)) F.group_norm(features, 3)
features = torch.randn(3,3) weight = TestTensor(torch.randn(3)) F.group_norm(features, 3, weight=weight)
</pre></sub></td>
<td><sub><pre lang="python">
import torch import torch.nn.functional as F class TestTensor(object): def init(self, weight): self.weight = weight def torch_function(self, func, _, args=(), kwargs=None): print(func) print(func == F.group_norm)
features = TestTensor(torch.randn(3,3)) F.group_norm(features, 3)
features = torch.randn(3,3) weight = TestTensor(torch.randn(3)) F.group_norm(features, 3, weight=weight)
</pre></sub></td>
</tr>
</table> </p>
Calls to backward() or grad() synced only the calling thread's default stream with autograd leaf streams at the end of backward. This made the following weird pattern safe:
with torch.cuda.stream(s):
# imagine forward used many streams, so backward leaf nodes may run on many streams
loss.backward()# no sync
use grads
but a more benign-looking pattern was unsafe:
with torch.cuda.stream(s):
# imagine forward used a lot of streams, so backward leaf nodes may run on many streams
loss.backward()
# backward() syncs the default stream with all the leaf streams, but does not sync s with anything,
# so counterintuitively (even though we're in the same stream context as backward()!)
# it is NOT SAFE to use grads here, and there's no easy way to make it safe,
# unless you manually sync on all the streams you used in forward,
# or move "use grads" back to default stream outside the context.
use grads
Note: this change makes it so that backward() has same user-facing stream semantics as any cuda op.** In other words, the weird pattern is unsafe, and the benign-looking pattern is safe. Implementation-wise, this meant backward() should sync its calling thread's current stream, not default stream, with the leaf streams. This PR deletes syncs on the default stream.
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> with PackageExporter(buffer, verbose=False) as e: e.intern("") e.save_pickle("res", "mod1.pkl", mod1) e.save_pickle("res", "mod2.pkl", mod2) </pre></sub></td> <td><sub><pre lang="python"> with PackageExporter(buffer) as e: e.intern("") e.save_pickle("res", "mod1.pkl", mod1) e.save_pickle("res", "mod2.pkl", mod2) </pre></sub></td> </tr> </table> </p>
Previously the way we insert observers/fake_quants are specific to fbgemm/qnnpack backend, as we work on making FX Graph Mode Quantization extensible to custom backends, we are changing some behaviors for the fbgemm/qnnpack path as well. The above changes are adding extra observer/fake_quant to the output of some operators to make sure we model the quantized operator more accurately in quantization aware training, the comprehensive list of operators where the behavior changes are the following:
We will show an example with torch.nn.MaxPool2d:
class M(torch.nn.Module):
def __init__(self):
super().__init__()
self.maxpool2d = torch.nn.MaxPool2d(kernel_size=3)
def forward(self, x):
x = self.maxpool2d(x)
return x
m = M().eval()
m = prepare_fx(m, {"": torch.quantization.default_qconfig})
print(m.code)
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> def forward(self, x): x_activation_post_process_0 = self.x_activation_post_process_0(x); x = None maxpool2d = self.maxpool2d(x_activation_post_process_0); x_activation_post_process_0 = None return maxpool2d </pre></sub></td> <td><sub><pre lang="python"> def forward(self, x): x_activation_post_process_0 = self.x_activation_post_process_0(x); x = None maxpool2d = self.maxpool2d(x_activation_post_process_0); x_activation_post_process_0 = None maxpool2d_activation_post_process_0 = self.maxpool2d_activation_post_process_0(maxpool2d); maxpool2d = None return maxpool2d_activation_post_process_0 </pre></sub></td> </tr> </table> </p>
Note that self.maxpool2d_activation_post_process_0 and self.x_activation_post_process_0 will refer to the same observer/fake_quant instance, this is to simulate the numerics for the quantized maxpool implementation, where the output would reuse the quantization parameter of the input. Simple illustration with graph:
Before:
observer_0 - maxpool - ...
After:
observer_0 - maxpool - observer_0 (same observer instance as input observer) - ...
aten arg from torch.onnx.export(). (#62759)The new OperatorExportTypes.ONNX removes the need for an explicit aten argument. If Pytorch was built with -DPYTORCH_ONNX_CAFFE2_BUNDLE the a None value means OperatorExportTypes.ONNX_ATEN_FALLBACK
<p align="center"> <table align="center"> <tr><th>1.9.1</th><th>1.10.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.onnx.export(..., aten=True) </pre></sub></td> <td><sub><pre lang="python"> torch.onnx.export(..., operator_export_type=torch.onnx.OperatorExportTypes.ONNX_ATEN) </pre></sub></td> </tr> </table> </p>
__torch_function__ as a plain methods (#64843)The __torch_function__ function used to create Tensor like objects did not have any constraint whether it should be a method, class method or static method.
To make it compatible with newer features on Tensor-like objects, we are deprecating setting it as a plain method. You can define it as a class method to get the current class and scan the argument list if you need an object that is an instance of this class.
This API caused many issues and is not really necessary. The functionality (run model with bundled input) can be achieved by using get_all_bundled_inputs. For example:
1.9.1:
model.run_on_bundled_input(0)
1.10.0:
model(*model.get_all_bundled_inputs()[0])
torch.distributed.rpc: Removed ProcessGroup RPC backend (#62411 , #62985)ProcessGroup RPC backend has been deprecated and 1.9 was the last release which carried it. The default RPC backend is TensorPipe which is the recommended backend for RPC. Users who use torch.distributed.rpc.BackendType.PROCESS_GROUP will be given an error message to switch to torch.distributed.rpc.BackendType.TENSORPIPE.
enable_onnx_checker argument is removed. ONNX checker will now always run by default. Users can catch exceptions to ignore raised failures. strip_doc_string has been rolled into the verbose arg in torch.onnx.export(). _retain_param_name argument has been removed in torch.onnx.export() will default to True . There is no way to get the old behavior of _retain_param_name=False. Users should stop setting this arg.
1.9.1:
torch.onnx.export(..., enable_onnx_checker=False, strip_doc_string=False)
1.10.0:
try:
torch.onnx.export(verbose=True)
except torch.onnx.utils.ONNXCheckerError:
pass
ParallelTBB config/codepath is no longer actively tested by PyTorch CI and as result is subject to code/functionality degradation
torch.isin() (#53125), torch.bitwise_{left/right}_shift, __rlshift__, __rrshift__ (#59544), torch.Tensor.{__rand__, __ror__,__rxor__} (#59240), torch.aminmax (#62401), torch.new_ones (#58405)torch.cov (#58311), torch.frombuffer (#59077), torch.corrcoef (#60420), torch.nanmean (#62671), torch.cumulative_trapezoid (#61615)torch.optim:
torch.cpu.amp.autocast: enable new API for CPU autocast (#57386, #63534)BFloat16 support for torch.{cross, tril, triu, tril_indices, triu_indices, cumsum, cummax, cummin, median, kthvalue, nansum, nextafter, range, sinh, cosh, frexp, nan_to_num, sigmoid, sigmoid_backward, tanh_backward, addcmul, addcdiv, bucketize, bernoulli, dropout, fold, unfold, MaxPool2D, AdaptiveAvgPool2D, topk} on CPU (#62454, #63307, #55210, #60074, #61083, #61829, #55221, #61826, #55588, #56372, #62880, #55202, #59547)BFloat16 support for torch.{ceil, floor, frac, round, trunc, sort, topk, aminmax, cumsum, logcumsumexp, cumprod, cummin, cummax} on CUDA (#57910, #58196, #59977, #62767, #57904).torch.cuda.is_bf16_supported (#63798)torch.segment_reduce (#59951, #60018, #61141, #61266, #59521, #60379, #60379)torch.isclose (#61271)torch.trapezoid (#61475).torch.gradient support for second order central differences (edge_order=2) (#58165)torch.sigmoid: CUDA support and complex autograd support (#48647)torch.bilinear and torch.nn,MaxUnpool2d (#56322, #49984)autograd.Function to implement your own forward-mode-AD-supported operator.autograd.Function (#64061, #63434)torch.{acos, add, addbmm, addcdiv, addcmul, addmm, addmv, addr, angle, acosh, asinh, atanh, asin, atan, conj, baddbmm, bmm, cat, ceil, clamp, clamp_min, clamp_max, complex, copy_sign, cos, cosh, cross, cumprod, cumsum, cummax, cummin, deg2rad, div, dot, vdot, exp, exp2, expm1, expand, floor, frac, frexp, gather, hardswish, hstack, hypot, index_add_, index_copy_, index_put_, index_select, kthvalue, lerp, lgamma, digamma, polygamma, log, log10, log1p, log2, logaddexp, logaddexp2, xlogy, masked_fill_, masked_fill_, masked_scatter_, masked_select, max, maximum, fmax, mean, min, mininum, fmin, mm, mode, mul, lu, lu_solve, vstack} (#57768, #57863 #59711, #64742)torch.{mvlgamma, nan_to_num, permute, pow, reciprocal, remainder, repeat, round, rsqrt, sigmoid, logit, sign, sgn, sin, sinc, sinh, sqrt, squeeze, sub, sum, t, flip, roll, rot90, take, tan, tanh, trace, transpose, tril, triu, trunc, unfold, unsqueeze, view, zero_, hardshrink} (#59993)torch.special.{xlog1py, entr} (#59711, #59993)torch.linalg.{cholesky, cholesky_ex, eigh, inv, inv_ex, solve} (#62160, #64646, #62163, #62159)torch.functional.leak_relu (#59993)autograd.Function to use with the saved tensor hooks (#60551).is_inference() method (#58729)torch.lu_solve: Implement support for backward AD (#61681).nn.{ReflectionPad3d, LazyInstanceNorm*d} (#59791, #60837, #61308, #60982)nn.CrossEntropyLoss: Added support for class probability targets (#61044)nn.CrossEntropyLoss: Added support for label smoothing (#63122)nn.Module: Added support for arbitrary objects in state_dicts via get_extra_state() / set_extra_state() (#62976)nn.utils.skip_init(): Added function to skip module parameter / buffer initialization (#57555)#364,#368,#383, #422)at::meta:: namespace (#58570)cpu_kernel, cpu_kernel_vec and cpu_kernel_multiple_outputs (#58949)at::native::resize_bytes_cpu to resize Storage in ATen (#60324)transpose to PackedTensorAccessor (#61114)torch::linalg::qr as the C++ API (#60529)amin and amax to aten symbols (#61550)c10::optional to compare with different but comparable types (#62890)c10::util::check_env to check environment variable (#59052)torch.distributed.rpc.is_available() (#58887)torch.jit.script (#62420)torch::deploy C++ API (#62669)torch::deploy. (#63817)aten::{avgpool2d,softmax,to,div,flatten,detach,slice,log_softmax,conv2d_transpose} to NNAPI converter (#58538, #58539, #58540, #58541, #60885, #58543, #59364, #61378, #59529aten::{conv2d,linear,cat,flatten} converter accept flexible batch (#61021, #61022, 76c0f223d3, #61024)aten::{hardswish,tanh,clamp} for iOS Metal (#64588, #61383)DistributedDataParallel
torch.distributed
__fx_create_arg__ dunder method for controlling custom classes are handled as node args (#61780)autowrap_functions kwarg to Tracer (#62106)conv2d, BatchNorm2D, ReLU, maxpool2D, AdaptiveAvgPooling2D, flatten (#61093, #61012, #61150, #61188, #61239, #61265)get_attr operations in typechecker (#62682)remove_duplicate_output_args (#65134)torch.{linspace, new_ones, nn.LSTMCell, bernoulli, dot, nn.utils.spectral_norm,bernoulli, distributions.normal.Normal, roll} (#58854, #59255, #62757, #62765, #59536,#61560,#58697)torch.fft. operators on ARM-based platforms using pocket FFT (#60976, #62222, #63714)torch.einsum: added support for the “sublist” format (#56625)torch.linalg.det: added support for complex autograd (#58195)Tensor.to_sparse (#58413)max_pool2d, tanh, hardshrink, log_softmax, leaky_relu, softmax (#58806, #60695, #62870, #63193, #62239)torch.floor_divide deprecation warning (#64034)torch.nansum accuracy (#61082)torch.i0: now promote integer inputs to float (#52735)torch.kthvalue: added change to adjust output dim size for numpy compatibility (#59214)torch.scatter operation. (#57015)torch.testing.assert_close (#58926)torch.isclose upcast to most precise dtype within their category before the comparison (#60536)alpha to acc_type for torch.add and torch.sub (#60227)torch.cat shape check and removed unnecessary offending index information (#64556).torch.gather (#65006).float64 in tensorboard instead of float32 (#59435).use_strict_trace to tensorboard add_graph method (#63120).torch.hub (#62139)output_size to tensor.repeat_interleave(#58881)torch.isclose (#63571)torch.{testting.assert_close,is_close} consistent with numpy (#63841).backward() is called with create_graph=True (#59412)Tensor::grad() on a non-leaf Tensor in the C++ API (#59362)grad_output creation for .backward() and autograd.grad() (#59532)NotImplementedError for forward and backward-mode AD formulas that are not implemented (#59482, #59483)torch.relu for common use cases (#63089)autograd.backward() function inputs argument (#60521)requires_grad=True is passed to a non-differentiable function (#60610)binary_cross_entropy differentiable w.r.t. target (#59447)nn.{AdaptiveAvgPool*d, AdaptiveMaxPool*d, AvgPool*d, CosineEmbeddingLoss, Dropout, FractionalMaxPool2d, Linear, LPPool1d, MaxPool*d, MaxUnpool*d, NLLLoss, PairwiseDistance, ReflectionPad*d, ReplicationPad*d, TripletMarginLoss, ZeroPad*d}, most other loss modules, and all activation modules (#61264, #61847, #61860, #64590, #61911, #62490, #60992, #62190, #62206, #61984, #61310, #62651, #64882, #62183, #61060, #61262, #62729, #61300, #61461, #62726)nn.{AdaptiveAvgPool*d, AdaptiveMaxPool*d, Bilinear, FractionalMaxPool*d, LocalResponseNorm, MaxPool*d, MaxUnpool*d, TransformerDecoder, TransformerDecoderLayer, TransformerEncoder, TransformerEncoderLayer} (#62025, #62088, #47106, #62083, #62801, #64082, #62800)nn.AvgPool2d: Added channels_last support on CPU (#58725)nn.BatchNorm: Use resize_output and empty instead of empty_like to improve flexibility in output memory format choice (#63084)nn.Bilinear: Added support for non-contiguous tensor inputs (#38409)nn.GELU: Added support for fp32/bfloat16 in CPU path using mkldnn implementation (#58525)nn.GroupNorm: Improved numerical stability by using the Welford algorithm and cascade summation (#54921)nn.LayerNorm: Improved numerical stability by using the Welford algorithm and pairwise sums (#59987)nn.NLLLoss: Added support for target of dtype byte (#60308, #60650)nn.SmoothL1Loss: Added support for integral target within the backward pass (#61112)nn.Transformer: Added configurable pre/post LayerNorm placement (#60593, #61692)nn.{RNN, LSTM, GRU} (#60269)nn.{LeakyReLU, RReLU} (#61514)channels_last memory format in nn.{AdaptiveMaxPool2d, GroupNorm} (#48920, #49821)nn.{MultiheadAttention, Transformer, TransformerDecoderLayer, TransformerEncoderLayer} (#61355, #62342)profiler.profile argument with_flops when set to True to report total FLOPs rather than FLOP/s, and support more operators (#62779, #61895)#361, #404,#416,#421)#351)Subset to dataset (#59513)ConcatDataset must be Sized (#64114)IterableDataset to accept keyword-only arguments and abc class (#58450)DataLoader to accept non-integer Sampler as input(#63500)torch.scatter_add for 1D tensors (#58761)--torch_jit_enable_rethrow_caught_exception=true (#63348)torch.nn.ModuleList to support arbitrary step size (#58361)Tuple[()] annotation (#58340)torch.nn.Parameter type for Profile-Directed-Typing (#59249)torch.einsum (#59265)torch.jit.isinstance with multiple types (#60465)checkScriptRaisesRegex (#63901)optimize_for_mobile to preserve nodes’ debug information (#63106)torch::deploy (#58117)torch.utils.model_dump APIs:
quantized::linear (#58282) and quantized::embedding_bag_byte_prepack (#64081)qconfig_dict argument handling (#59605, #58566)torch.index_select on quantized tensors (#61406)DistributedDataParallel
NCCL_ASYNC_ERROR_HANDLING environment variable to control NCCL error handling (#59109)mul and copy_ instead of mul’s out= variant when gradient tensor requires grad in DDP (#63831)Tensor.set_ instead of directory assigning data in model averaging (#63895)torch.distributed
torch.distributed launcher (#59152)torch.distributed.optim.ZeroRedundancyOptimizer (#61370)torch.distributed.nn.RemoteModule
torch.distributed.elastic
torch.distributed.rpc
threading.Locks (#57943), torch.cuda.Event (#61354)torch.distributed.Store
torch.distributed.pipeline
WithDevice wrapper to specify device execution for a module. (#65190)torch.nn.Module constructor (#61334)torch.deploy for GraphModules with non-torch dependencies (#61680)torch.memory_format as a BaseArgumentType (#62593)__matmul__ to the magic methods for FX tracing (#64512)torch.{any, all, fmax, fmin, remainder, glu, argmax, argmin, avg_pool3d_backward, isposinf, isneginf, fmod, fmin, signbit, slow_conv_transpose2d, nll_loss_backward, cumprod, aminmax, addcmul, addcdiv, gather, hardshrink_backward, softshrink_backward, hardshrink, gelu, gelu_backward, avg_pool2d, avg_pool2d_backward, avg_pool3d, reflection_pad1d_backward, all, any, silu_backward, sgn, softplus, leaky_relu_backward, hardsigmoid_backward, elu_backward, eq, xlogy, ne, lt, gt, le, ge, sigmoid_backward, tanh_backward, logit_backward, bitwise_or, bitwise_xor, bitwise_and, nll_loss_forward, log_softmax, log_softmax_backward_data, prod, norm, sum.dim_IntList, clamp} (#64642, #58458,#58732, #61800, #60363, #60364, #59084, #60633, #60809, #60810, #57936, #55503, #62144, #61899, #62401, #62318, #62319, #63312, #58662, #58663, #58664, #58665, #58987, #59082, #59083, #59103, #60360, #60361, #58661, #58197, #58482, #58483, #58484, #58660, #60177, #60814, #60942, #60815, #60816, #60817, #60811, #60812, #60813, #61443, #57374, #62372, #62024, #62711, #61642, #61361)torch.utils.collect_env (#59632)CMAKE_PREFIX_PATH choice set by caller (#61904)torch.__version__ comparisons (#61556, #64565, #63848)Note truncated.
Deprecate use_env in torch.distributed.run #59409
.names() access in max_pool2d #60059torch.hub.load #62139log.warning in torch.distributed.run to print OMP_NUM_THREADS warning #63953torch.distribtued.run to set nproc_per_node to 1 by default #61552torch.distributed.elastic.utils.store #60807use_env in torch.distributed.run #59409torch.distributed.run #63910torch.distributed.launch and torch.distributed.run and clarify the documentation accordingly #61294torch.hub.load for TorchVision models #62072torch.mm to check input matrix sizes shapes #61394torch.distribtued.run warning message #61127Your coding agent can read these notes before it upgrades. Set up the MCP server →