NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #358 most downloaded on PyPI
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Last release 20 days ago
02 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 49 of 50 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
50 releases · first in 2018
…use rounding_mode='trunc' (and are also deprecated for that reason).
We are excited to announce the release of PyTorch 1.9. The release is composed of more than 3,400 commits since 1.8, made by 398 contributors. Highlights include:
We’d like to thank the community for their support and work on this latest release. We’d especially like to thank Quansight and Microsoft for their contributions.
You can find more details on all the highlighted features in the PyTorch 1.9 Release blogpost.
torch.divide with rounding_mode='floor' now returns infinity when a non-zero number is divided by zero (#56893).
This fixes the rounding_mode='floor' behavior to return the same non-finite values as other rounding modes when there is a division by zero. Previously it would always result in a NaN value, but a non-zero number divided by zero should return +/- infinity in IEEE floating point arithmetic. Note this does not effect torch.floor_divide or the floor division operator, which currently use rounding_mode='trunc' (and are also deprecated for that reason).<p align="center"> <table align="center"> <tr><th>1.8.1</th><th>1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a = torch.tensor([-1.0, 0.0, 1.0]) b = torch.tensor([0.0]) torch.divide(a, b, rounding_mode='floor') tensor([nan, nan, nan]) </pre></sub></td> <td><sub><pre lang="python"> a = torch.tensor([-1.0, 0.0, 1.0]) b = torch.tensor([0.0]) torch.divide(a, b, rounding_mode='floor') tensor([-inf, nan, inf]) </pre></sub></td> </tr> </table> </p>
Tensor.new no longer support passing both Tensor and device as inputs (#58108).
This fixes a bug in which 1-element integer tensors were misinterpreted as specifying tensor size, yielding an uninitialized tensor. As noted in the error message, use the new-style torch.tensor(...) or torch.as_tensor(...) to copy or alias an existing tensor. If you want to create an uninitialized tensor, use torch.empty(...).
<p align="center">
<table align="center">
<tr><th>1.8.1</th><th>1.9.0</th></tr>
<tr valign="top">
<td><sub><pre lang="python">a = torch.tensor([1]) torch.LongTensor(a, device='cpu') # uninitialized tensor([7022349217739848992]) a.new(a, device='cpu') tensor([4294967295]) # uninitialized </pre></sub></td> <td><sub><pre lang="python"> a = torch.tensor([1]) torch.LongTensor(a, device='cpu') RuntimeError: Legacy tensor constructor of the form torch.Tensor(tensor, device=device) is not supported. Use torch.tensor(...) or torch.as_tensor(...) instead. a.new(a, device='cpu') RuntimeError: Legacy tensor new of the form tensor.new(tensor, device=device) is not supported. Use torch.as_tensor(...) instead. </pre></sub></td> </tr> </table> </p>
torch.divide with rounding_mode='true' is replaced with rounding_mode=None (#51988).
torch.divide's undocumented rounding_mode='true' option has been removed, and instead rounding_mode=None should be passed to indicate no rounding should take place. This is equivalent to omitting the argument entirely.<p align="center"> <table align="center"> <tr><th>1.8.1</th><th>1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a, b = torch.full((2,), 4.2), torch.full((2,), 2) torch.divide(a, b, rounding_mode='true') tensor([2.1000, 2.1000]) </pre></sub></td> <td><sub><pre lang="python"> a, b = torch.full((2,), 4.2), torch.full((2,), 2) torch.divide(a, b, rounding_mode=None) # equivalent to torch.divide(a, b, rounding_mode='true') from the prior release tensor([2.1000, 2.1000]) </pre></sub></td> </tr> </table> </p>
import torch.tensor as tensor is no longer supported (#53424).
Instead, use from torch import tensor<p align="center"> <table align="center"> <tr><th>1.8.1</th><th>1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
import torch.tensor as tensor torch.tensor(1.) tensor(1.) </pre></sub></td> <td><sub><pre lang="python"> import torch.tensor as tensor ModuleNotFoundError: No module named 'torch.tensor' from torch import tensor tensor(1.) tensor(1.) </pre></sub></td> </tr> </table> </p>
numpy is no longer a required dependency
If you require numpy (and don't already have it installed) you will need to install it separately.torch.autograd.gradcheck.get_numerical_jacobian and torch.autograd.gradcheck.get_analytical_jacobian no longer support functions that return complex valued output as well as any other values of grad_out not equal to 1 (#55692).
This change is a part of a refactor of gradcheck’s internals. Note that gradcheck itself still supports functions with complex output. This new restriction only applies to calls to the two internal helper functions. As a workaround, you can wrap your functions to return either the real or imaginary component of its output before calling these functions. Additionally these internal helpers no longer accept any other value except 1 for grad_out for any input function. Note that these helper functions are also being deprecated in this release.1.8.1:
get_numerical_jacobian(torch.complex, (a, b), grad_out=2.0)
1.9.0:
def wrapped(fn):
def wrapper(*input):
return torch.real(fn(*input))
return wrapper
get_numerical_jacobian(wrapped(torch.complex), (a, b), grad_out=1.0)
torch.autograd.gradcheck now throws GradcheckError (#55656).
This change is a part of a refactor of gradcheck’s internals. All errors that are able to be silenced by raise_exception=False now raise GradcheckError (which inherits from RuntimeError). If you explicitly check that the type of the error is RuntimeError you'll need to update your code to check for GradcheckError instead. Otherwise if you use something like except or isinstance, no changes are necessary.1.8.1:
# An example of a situation that will now return GradcheckError instead of
# RuntimeError is when there is a jacobian mismatch, which can happen
# for example when you forget to specify float64 for your inputs.
try:
torch.autograd.gradcheck(torch.sin, (torch.ones(1, requires_grad=True),))
except RuntimeError as e:
assert type(e) is RuntimeError # explicitly check type -> NEEDS UPDATE
1.9.0:
try:
torch.autograd.gradcheck(torch.sin, (torch.ones(1, requires_grad=True),)
except RuntimeError as e:
# GradcheckError inherits from RuntimeError so you can still catch this
# with RuntimeError (No change necessary!)
# BUT, if you explicitly check type...
assert type(e) is torch.autograd.GradcheckError
nn.MultiheadAttention to now apply bias flag to both in and out projection layers (#52537).
In PyTorch 1.6, a regression was introduced that caused the bias flag of nn.MultiheadAttention only to apply to the input projection layer. This caused the output projection layer to always include a bias parameter, even with bias=False specified. The regression is now fixed in PyTorch 1.9, making the bias flag correctly apply to both the input and output projection layers. This fix is BC-breaking for the bias=False case as it will now result in no bias parameter for the output projection layer.<p align="center"> <table align="center"> <tr><th>v1.6 - v1.8.1:</th><th>pre 1.6 & 1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
mha = torch.nn.MultiheadAttention(4, 2, bias=False) print(mha.out_proj.bias) Parameter containing: tensor([0., 0., 0., 0.], requires_grad=True) </pre></sub></td> <td><sub><pre lang="python"> mha = torch.nn.MultiheadAttention(4, 2, bias=False) print(mha.out_proj.bias) None </pre></sub></td> </tr> </table> </p>
nn.Module to fire full backward hooks even when no input requires grad (#56693).
Prior to this release, full backward hooks were not fired when no input requires gradients. This has been changed so that full backward hooks will always fire during the backward pass, regardless of whether or not any input requires gradients. If you are using full backward hooks, be aware that they may fire more frequently than pre-1.9 due to this change.<p align="center"> <table align="center"> <tr><th>1.8.1:</th><th>1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
m = torch.nn.Linear(2, 3) def hook(mod, grad_input, grad_output): print('hook called:', grad_input, grad_output) m.register_full_backward_hook(hook) input_no_grad = torch.rand(1, 2, requires_grad=False) m(input_no_grad).sum().backward() input_grad = torch.rand(1, 2, requires_grad=True) m(input_grad).sum().backward() hook called: (tensor([[0.1478, 0.6517]]),) (tensor([[1., 1., 1.]]),) </pre></sub></td> <td><sub><pre lang="python"> m = torch.nn.Linear(2, 3) def hook(mod, grad_input, grad_output): print('hook called:', grad_input, grad_output) m.register_full_backward_hook(hook) input_no_grad = torch.rand(1, 2, requires_grad=False) m(input_no_grad).sum().backward() hook called: (None,) (tensor([[1., 1., 1.]]),) input_grad = torch.rand(1, 2, requires_grad=True) m(input_grad).sum().backward() hook called: (tensor([[0.1478, 0.6517]]),) (tensor([[1., 1., 1.]]),) </pre></sub></td> </tr> </table> </p>
DataLoader with num_workers > 0 will now set independent random seed for NumPy random functions on each worker by default. So, users now won’t be required to set random seed for NumPy using worker_init_fn to force NumPy random operations deterministic and independent across DataLoader workers. This PR won’t affect users who have already set random seed for NumPy random functions using worker_init_fn. # dataset returns numpy.random.randint(1, 10000)
ctx = mp.get_context('fork')
gen = torch.Generator().manual_seed(0)
dl = DataLoader(dataset, batch_size=2, num_workers=2, multiprocessing_context=ctx, generator=gen)
for epoch in range(2):
print("=" * 4, "Epoch", epoch, "=" * 4)
for batch in dl:
print(batch)
<p align="center"> <table align="center"> <tr><th>1.8.1:</th><th>1.9.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
========== Epoch 0 ========== tensor([[ 0, 340], [ 1, 7512]]) tensor([[ 2, 340], [ 3, 7512]]) ========== Epoch 1 ========== tensor([[ 0, 340], [ 1, 7512]]) tensor([[ 2, 340], [ 3, 7512]]) </pre></sub></td> <td><sub><pre lang="python">
DataLoader workers in each epoch.========== Epoch 0 ========== tensor([[ 0, 8715], [ 1, 5555]]) tensor([[ 2, 6379], [ 3, 1432]]) ========== Epoch 1 ========== tensor([[ 0, 1374], [ 1, 996]]) tensor([[ 2, 143], [ 3, 3507]]) </pre></sub></td> </tr> </table> </p>
A new attribute named type has been introduced for IterableDataset using the typing annotation at each class declaration. By adding this attribute, we are able to extend IterableDataset to have type inference and lazy initialization to incorporate the new DataLoader architecture. But, several BC-breaking restrictions are introduced due to this feature.
1.8.1:
# Users can use string to bypass the invalid type annotation without any error.
# And, incorrect type annotations attached to `__iter__` function are ignored.
1.9.0:
# The following scenario will now raise different Exceptions
# 1) The type annotation is required to be valid now. Previous workaround
# like using string to represent the invalid type annotation is not supported now.
# Raises Exception from the evaluation `eval("invalid_type", globals, locals)`
class DS(IterableDataset["invalid_type"]):
...
# Raises TypeError if the return type of __iter__ is not an Iterator
class DS(IterableDataset[str]):
def __iter__(self) -> str:
...
# Raise TypeError if the return type of __iter__ is of the form Iterator[X],
# but the argument type X is not a subtype of the IterableDataset.type attribute.
class DS(IterableDataset[str]):
def __iter__(self) -> Iterator[int]:
...
# IterableDatset now has a metaclass, which will conflict with
# existing user-defined metaclasses on IterableDatasets
class DS(IterableDataset[str], metaclass=MyMeta):
...
type(type(torch.tensor([1]))) now prints <class 'torch._C._TensorMeta'> (used to be <class 'type'>)resize_, resize_as_, resize_as_sparse_, sparse_resize_, and sparse_resize_and_clear_ has changed to return a const Tensor& instead of a Tensor&. This may break users’ TORCH_LIBRARY operators that called these functions but returned a non-const Tensor&. Ideally, users can change their operators to also consume and return const Tensor&, but simply casting the result of the changed function with const_cast<Tensor&> is also an option.1.8.1:
const at::Tensor a = at::randn({2, 2});
const at::Tensor b = at::ones({1, 4}, at::kInt);
at::Tensor& out = at::resize_as_(a, b); # success
1.9.0:
const at::Tensor b = at::ones({1, 4}, at::kInt);
at::Tensor& out = at::resize_as_(a, b);
# error: binding value of type 'const at::Tensor' to reference to type 'at::Tensor' drops 'const' qualifier
const at::Tensor& out = at::resize_as_(a, b); # Success
kron_out now throw an error when an undefined tensor is passed as input for out argument (#53218, #53640).
sum_out, nansum_out, prod_out, std_var_out have been changed to require users allocating result Tensor before calling these ops. The C++ API allocate_reduction_result has changed to resize_reduction_result to disallow allocating result Tensor in these reduction ops.c10::Error when executed. This code compiled and executed successfully in the prior release.at::Tensor out; # Undefined Tensor
const at::Tensor a = at::randn({2, 2});
at::IntArrayRef dim = {1};
at::sum_out(out, a, dim);
# c10::Error: Expected a Tensor of type Variable but found an undefined Tensor for argument #4 'out'
expand_inplace and expand_outplace now return c10::MaybeOwned<Tensor> instead of std::tuple<Tensor> (#55065, #55245).
The rationale for this change is to avoid unnecessary Tensor creation, thus improving performance. Functions in ExpandUtils return c10::MaybeOwned<Tensor> because expansion may not actually be needed, in which case we can improve efficiency by returning c10::MaybeOwned<Tensor>::borrowed(to_expand). However, this means that you need to be careful: the returned c10::MaybeOwned<Tensor> must not outlive the original Tensor object that to_expand referred to! The deleted rvalue reference overloads of these functions help with this by preventing trivial use of a temporary resulting from a function call, but it is still possible to make a mistake.@torch.jit.script. This new feature attempts to script the type, and falls back to the old behaviour of marking the class type attribute as "failed" if scripting fails. However, if the class definition does not have type annotations, the definition of the scripted class can different from users might expect (see code sample). If needed, users can explicitly disable the scripting of a class type attribute by adding its name to the __jit_ignored_attributes__ class attribute of the module being scripted.1.8.1:
class MyClass:
def __init__(self, a):
self.attr = a
class MyModule(torch.nn.Module):
def __init__(self):
self.attr = MyClass(4)
sm = torch.jit.script(MyModule())
1.9.0:
class MyClass:
def __init__(self, a):
self.attr = a
class MyModule(torch.nn.Module):
def __init__(self):
self.attr = MyClass(4)
# RuntimeError: Could not cast attribute 'attr' to type Tensor: Unable to cast Python instance of type <class 'int'> to C++ type 'at::Tensor'
sm = torch.jit.script(MyModule())
This error occurs because MyClass is automatically scripted, but self.attr is inferred to be a Tensor instead of an int because a is not annotated. To fix this, annotate a with the right type int, or mark attr as an attribute that should be ignored by the scripting process and not recursively processed:
class MyModule(torch.nn.Module):
__jit_ignored_attributes__ = ["attr"]
def __init__(self):
self.attr = MyClass(4)
torch.quantization.quantize_fx.convert_fx’s debug argument has been changed to is_reference (#52179).
<p align="center">
<table align="center">
<tr><th>1.8.1:</th><th>1.9.0</th></tr>
<tr valign="top">
<td><sub><pre lang="python">
import torch.quantization.quantize_fx as quantize_fxm = quantize_fx.convert_fx(m, debug=True) (Runs successfully) </pre></sub></td> <td><sub><pre lang="python"> m = quantize_fx.convert_fx(m, is_reference=True) # Runs successfully m = quantize_fx.convert_fx(m, debug=True) Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: convert_fx() got an unexpected keyword argument 'debug' </pre></sub></td> </tr> </table> </p>
torch.cat is now quantized to torch.cat instead of torch.ops.quantized.cat (#54924).
Previously, we produced torch.ops.quantize.cat which took inputs, dequantized them
and requantized them with new qparams. This behavior has been changed to produce torch.cat directly. torch.cat uses the same observer/fake_quant instance for all inputs and output, assumes all inputs are sharing the same qparam, and produces a quantized Tensor with
the same qparam as all inputs. Using torch.cat is expected to be more efficient since it does not introduce extra quant/dequant.
torch.cat was quantized to torch.ops.quantized.cat.torch.cat is quantized to torch.cat (torch.cat works on both floating point and quantized Tensor).DistributedDataParallel: Removed support for inter-process device replication in DDP (#54454, #54825, #54826, #55212, #55253, #55353).
DistributedDataParallel now errors out when users attempt to use it in single-process multi-device mode, where a module is replicated across more than one device in a single process. This mode had been previously deprecated and is now removed. Use cases should switch to spawning a single process for each device that is used in replication, which is the performant way to use DistributedDataParallel and supports a variety of newly developed features.1.8.1:
>>> # Assume the below is ran on 2 ranks in a distributed setting.
>>> rank_to_devices = { 0: [0, 1], 1: [2, 3] }
>>> # Each rank replicates model across 2 GPUs.
>>> model_ddp = torch.nn.parallel.DistributedDataParallel(
model,
device_ids=rank_to_devices[rank]
)
>>> # No error is raised, but below warning is produced.
>>> UserWarning: Single-Process Multi-GPU is not the recommended mode for DDP. In this mode, each DDP instance operates on multiple devices and creates multiple module replicas within one process. The overhead of scatter/gather and GIL contention in every forward pass can slow down training. Please consider using one DDP instance per device or per module replica by explicitly setting device_ids or CUDA_VISIBLE_DEVICES.
1.9.0:
>>> # Assume the below is ran on 2 ranks in a distributed setting.
>>> rank_to_devices = { 0: [0, 1], 1: [2, 3] }
>>> # Each rank replicates model across 2 GPUs.
>>> model_ddp = torch.nn.parallel.DistributedDataParallel(
model,
device_ids=rank_to_devices[rank]
)
>>> # Single process multi-GPU mode now produces an error on initialization.
>>> ValueError: device_ids can only be None or contain a single element.
torch.distributed.elastic: Replaced torch.distributed.launch with torch.distributed.elastic_launch (#56037, #56214).
* --logdir → —log_dir. The stdout and stderr log dir arg name and destination changed. The file destination changed from $logdir/node_{}_local_rank_{}_stdout to $log_dir/$rank/stdout.log. If users used the —logdir introduced in 1.8 pytorch version, they need to use —log_dir parameter now.1.8.1:
#!/bin/bash
# Assumes training script train.py exists.
python -m torch.distributed.launch --nproc_per_node=2 --nnodes=1 --node_rank=0 --master_addr="127.0.0.1" --master_port="29500" --logdir test_logdir train.py
# Logs are written to $logdir/node_{}_local_rank_{}_stdout
1.9.0:
#!/bin/bash
# Assumes training script train.py exists.
python -m torch.distributed.launch --nproc_per_node=2 --nnodes=1 --node_rank=0 --master_addr="127.0.0.1" --master_port="29500" --log_dir test_logdir train.py
# Logs are written to $log_dir/$rank/stdout.log
torch.floor_divide has been deprecated in favor of torch.div(..., rounding_mode=‘floor’) (#50281).
torch.floor_divide incorrectly divides then truncates (rounds towards zero) instead of dividing then flooring (rounds “down”). Use the rounding_mode argument of torch.div to indicate if you’d like to continue performing truncation division or floor division, instead, since torch.floor_divide will be removed in a future PyTorch release.torch.{cholesky, qr, symeig, chain_matmul, solve, eig, matrix_rank, lstsq} have been deprecated in favor of torch.linalg.{cholesky, qr, symeig, chain_matmul, solve, eig, matrix_rank, lstsq} (#57725,#57745, #57732,#53453, #57741, #57727, #57734, #57743).torch.norm has been deprecated in favor of the new linalg module norm functions: torch.linalg.vector_norm, torch.linalg.matrix_norm, and torch.linalg.norm (#57986).torch.det, torch.slogdet, torch.matrix_power, torch.inverse, and torch.pinverse to their linalg module counterparts (#57821).AutoNonVariableTypeMode to AutoDispatchBelowAutograd and added a warning. (#56422)
AutoNonVariableTypeMode is deprecated and will be removed in 1.10 release. For kernel implementations, please use AutoDispatchBelowAutograd instead. Check out more details on how to migrate your kernel here. If you are looking for a user-facing API to enable running your inference-only workload, please use c10::InferenceMode. Using AutoDispatchBelowAutogradMode in user code is under risk of producing silently wrong result for some edge cases.1.8.1:
{
at::AutoNonVariableTypeMode guard(true);
}
1.9.0:
{
c10::AutoDispatchBelowAutograd guard(true); // for kernel implementations
// c10::InferenceMode guard(true); --> consider inference mode if you are looking for a user-facing API
}
Function (#57357).
Instantiating a custom autograd function is now deprecated and will raise a warning. Users should call .apply() on the class itself because it is a static method.1.8.1:
# Instantiating custom function will raise a warning in 1.9
Func().apply
1.9.0:
# You should directly call the `apply` (classmethod) on the class
Func.apply
get_analytical_jacobian and get_numerical_jacobian (#54378, #54049).
torch.autograd.gradcheck.get_analytical_jacobian and torch.autograd.gradcheck.get_numerical_jacobian are internal-facing functions that are not a part of our public API. We’ve refactored some PyTorch internals to work without it and will
remove it in a future release. For gradient checking purposes, please use torch.autograd.gradcheck.linalg_ prefix from torch::linalg::linalg_det and torch::linalg::linalg_norm C++ API (#57464).
C++ code that used to call torch::linalg::{linalg_det, linalg_norm} should be updated to call torch::linalg::{det, norm}torch.distributed.rpc: Added a warning message to retire ProcessGroup RPC backend (#55616)
torch.{ceil, floor, frac, round, trunc, lerp, roll, diag, logaddexp, logaddexp2, nan_to_num, exp2, expm1, rsqrt, erfc, atan2, hypot} on CUDA (#57910, #57907, #57916, #57908, #58063, #57913, #57905).torch.pow() for torch.{float16, BFloat16} on CPU (#55280).torch.{index_select, argmax, argmin, min, max, amin, amax} for torch.{float16, BFloat16} (#53898, #52582, #51244, #52579).torch.dot for BFloat16 on CUDA (#57903).min and max arguments in torch.clamp (#52695, #56367).torch.special namespace similar to scipy.special (#52296).
alpha to torch.index_add (#54176).torch.assert_async (#53086)interpolation to torch.quantile (#49267).torch.{std, var, std_mean, var_mean} with a correction argument specifying the difference between the sample size and number of degrees of freedom.torch.{logit, rad2deg, deg2rad, polygamma} (#52028, #51853,#57462)stable (#51790).torch.linalg module, analogous to NumPy’s linalg module but with several additional functions, is stable! Added torch.linalg.{multi_dot, lstsq, vector_norm, matrix_norm, matrix_power, det, eig, eigvals, svdvals, cholesky_ex, inv_ex} (#51807, #49093, #51099, #57127, #52608, #53119, #52491, #56684, #56724, #58039).device=meta API (#53143)
torch.ones(2, device='meta') + torch.ones(1, 2, device='meta') will return a new meta tensor of size [1, 2] (performing broadcasting), without allocating memory or running an actual kernel.device=meta API is implemented for upsample_linear1d(#51917), upsample_bilinear2d and upsample_bicubic2d (#52012), upsample_nearest3d (#52065), sin(#52277), mul(#52692), pow(#53669), sub(#53679), div(#53680), copysign(#55040), atan2(#55130), sinh(#55538), acosh(#55540), cosh(#55563), cos (#55564), replication_padding1d (#55481), replication_padding3d (#55499), replication_pad1d_backward (#55537), fractional_max_pool2d (#55581), reflection_pad1d (#55531), replication_pad2d (#55511), addmv (#55746), all unary float functions (#56082), adaptive_max_pool2d(#56317), adaptive_max_pool3d (#56320), all non-float unary operators (and rsqrt) (#56151), adaptive_max_pool2d_backward (#56799), adaptive_max_pool3d_backward (#56800), neg(#57212), max_pool2d_with_indices(#56459), trunc (#57350), floor (#57587), sign (#57588), ceil (#57589), gcd (#57624), nextafter (#57625), igamma and igammac(#57626), hypot(#57627), lcm (#57628), logaddexp and logaddexp2 (#57629), maximum and minimum (#57630), topk (#57790), max_pool2d_with_indices_backward (#57797), threshold (#57810), addmm (#57417), heaviside (#57933), elu(#57619), softplus (#57620), leaky_relu (#57621), hardsigmoid (#57622), softshrink (#57623), silu (#58050), empty_strided (#53397), non-composite in-place operators (#54901)torch.{masked_fill, polar, cumsum, lerp, prod, rsub, unfold, symeig, index_copy} (#52483, #52488, #53240, #53689, #48125, #53702, #52999, #55085, #52203).torch.index_copy and torch.{take} and torch.Tensor.put_ on both CPU and CUDA (#52203, #53356).torch.{add, mul, sub, as_tensor} (#52881).cmath unary ops: cmath.{phase, log, log10, sqrt, exp, sin, cos, tan, asin, acos, atan, sinh, cosh, tanh, asinh, acosh, atanh} (#54089).cmath.{infj, nanj} (#54328).cmath.{isinf, isnan, isfinite, rect} (#54541).tensor.real/imag) (#54692).test_variant_consistency_jit_addmm for complex types (#54917, #57129).torch.{sparse_coo_tensor, coalesce, to_dense, to_sparse, sparse_add, sspaddmm, saddmm}.torch.Tensor.{cfloat, cdouble} functions (#58137).torch.{std, var} to return a real valued output tensor for complex inputs (#58066) .torch.nn modules: nn.LazyBatchNorm*d (#51862), nn.HuberLoss (#50553), nn.Mish (#58375).nn.Conv*d: Added padding='same' mode for non-strided convolutions (#45667).nn.EmbeddingBag: Added padding_idx support (#49237, #56065, #56618).MaxPool2d (#56361).skip_first parameter to the default schedule (#58025).gzip format support for chrome tracing (#56554).sequenceNr and fwdThreadId to the trace (#57182).fast_mode argument to autograd.gradcheck (#54480).torch.utils.checkpoint functions (#52422).FilterIterDataPipe (#51783).DataPipe at construct-time (#54066).DataPipe at runtime (#54544).issubtype for DataLoader type hints (#54299).ConcatDataPipe (#53301).DataLoader (#53271).ZipIterDataPipe (#53554).TransformsIterDataPipe (#52604).MapIterDataPipe (#51879).'max' reduction for torch.segment_reduce (#56704).torch::nn::functional::huber_loss (#50553).padding='same' mode to torch::conv{1,2,3}d (#45667).padding_idx argument to EmbeddingBag (#49237).pow on CPU (#56308).matmul for NNC lowering/unified dtypes (#56456).conv without bias (#57512).conv with dynamic shapes (#57514).t/transpose/permute/expand (#57426).GELU To NNC (#57753).GELU Backward (#58249).torch.type (#51904)dict() constructor (#51934).torch::deploy to manage multiple python interpreters in a single
process to deploy PyTorch models packaged with torch.package (#51754).torch.any (#52360).embedding_bag for SR (#52429).AliasDb in Python (#51336).DictConstruct (#54438)sliceHead/sliceTail APIs with short parameter list (#55115).CUDAFuture (#56516).MonkeyType (#57202).torch.jit.ignore as a context manager (#55172).hardswish/hardsigmoid on MKLDNN tensors (#55218).model_dump tool for model inspection (#56868)pow (#52374)TorchBind (#51253).optimize_for_inference API (#58193).aten::index_out (#51742).PYTORCH_TENSOREXPR_DONT_FUSE env variable to disable fusion on specified operators (#55650).torch.distributed.Store
torch.distributed.rpc
DistributedDataParallel
join context manager that enables throwing an error across all ranks when this flag is specified (#56755)torch.distributed.algorithms.default_hooks.fp16_compress_wrapper wrapper that can be combined with other communication hooks (#53808)torch.distributed
work.result API for MPI backend (#57168)work.result for ProcessGroupGloo::AsyncWork objects (#57565)work.get_future() API for ProcessGroupMPI and ProcessGroupGloo (#57818,#57214) torch.distributed.monitored_barrier API (Gloo-only) (#53773, #53787, #55009, #55010, #55197, #55265, #55989, #55990)options field to process group initialization APIs (#53662, #54090, #53663)TORCH_DISTRIBUTED_DEBUG environment variable (#52481)compareSet method for torch.distributed.{HashStore, FileStore} (#53803).torch.distributed.elastic module that upstreams pytorch/elastic
PeriodicTimer (#55919)DynamicRendezvousHandler and RendezvousBackend. (#55635)C10dRendezvousBackend. (#55636)EtcdRendezvousBackend. (#55637)torch.distributed.elastic.launchers.api, torch.distributed.elastic.metrics, torch.distributed.events, torch.distributed.rendezvous, torch.distributed.elastic.agent modules (#55471, #53870, #53574, #53760, #53172, #54343)torch.distribute.elastic.timer and torch.distributed.elastic.multiprocessing (#53574)torch.distributed.nn.RemoteModule: Enable RemoteModule to directly send GPU tensors over the wire on TensorPipe RPC backend if a device map is provided (#57288)torch.distributed.optim:
torch.fx.Node.format_node() (#51737).Transformer to normalize args/kwargs of torch.nn.functional calls into only kwargs (#51816).GraphModule (#52358).Graph.eliminate_dead_code (#52658).inspect.Signature instances for PyTorch operations (#53830).torch namespace operations (#53832).NormalizeArgs to work on torch namespace operations (#54236).optimize_for_inference for Intel CPUs (#53805, #58293).Node and switch shape-prop to use that (#54926).torch.randn to capture it during tracing (#54060).Node (#55887).concrete_args (#55888).hardswish and hardsigmoid activation functions (#53362).reflection_pad2d op (#53604).sigmoid activation function (#57867).TCP_TLS transport (#56442).index_put_ when accumulate=False (#55827).torch.index_add on CUDA (#56521).torch.index_copy on CPU (#56900).torch.Tensor.unflatten to be able to infer size value in sizes from -1 (#51955).out= input tensor for torch.tensordot (#56286).out and input tensors for torch.cat (#53004).torch.multinomial on CUDA (#53288).NotImplementedError (#53610).torch.multiprocessing (#53750).torch.int32 indices in torch.repeat_interleave (#55102).torch_shm_manager when running torch.multiprocessing (#57307, #57310).index_copy_cuda with index_put (#58144).torch.einsum (#56475).torch.orgqr (#51348) and torch.ormqr (#57316).torch.geqrf on both CPU and CUDA (#56249, #56251).torch.{linspace, logspace} to correctly infer complex type and return a complex tensor when the start and (or) end values are complex numbers, and the dtype value is None (#38875).inputs argument for .backward() (#53827).torch.orgqr (#52637), torch.segment_reduce (#56792).torch.gather for dim=1 (#55573).nn.init._calculate_fan_in_and_fan_out: Support usage with __torch_function__ (#53522).nn.Transformer / nn.MultiheadAttention: Add batch_first argument (#55285).nn.Transformer: Add layer_norm_eps arg (#54494).nn.AvgPool2d: Add channels_last support on CPU (#48918).clip_grad_norm_: Add error_if_nonfinite flag (#53843, #55169).Module.train: Raise nicer error when called with invalid modes (#58247).nn.Linear: Support 0 in_features (#56505).nn.EmbeddingBag: Support mix of int32 and int64 offsets/indices (#55189).xnnpack::linear: Handle 1D input (#54986).nn.Module: Add allow_duplicate flag to named_modules() (#54812).nn.Module: Add to_empty() function for moving to a device without copying storage (#56610).pad_sequence callable from C++ API (#57868).generate_state for NumPy seeding (#56797).LoadFilesFromDisk (#57056).DataPipe (#52858).MapIterDataPipe (#52856).torch.half dtype RNNs with MIOpen (#52475).hiprtc precompiler feature (#54350).hipfft and rocfft detection for ROCm build (#53408).torch.cuda.amp.GradScaler scale updates in-place for better composability with graph capture (#55562).USE_MAGMA build flag (#55994).torch.fx.map_arg require a callable (#51907).torch.fx.Tracer.create_arg (#51927).Transformer to accept output result that is not Proxy (#52473).TracerBase._find_user_frame private (#53654).GraphModule init (#53444).Interpreter (#54726).stride on Node during ShapeProp pass (#55108).memory_format on Node during ShapeProp pass (#55815).NamedTuple in ShapeProp (#55930).split_module (#56212).shape_prop handle targets with aggregate outputs (#56221).Node and not a pass (also augment tests to be exhaustive) (#55992).PythonKey out of tree (#57427).GraphDrawer when shape, type or stride are not present (#57845).get_attr directly in split_by_tags (#57844).args/kwargs in symbolic tracing(#57840).conv2ds (#56289).fuseLoops API to return bool flag and not throw any exceptions (#56353).unroll and flatten APIs which do not require return stmt pointer (#56420).Buf on mutation instead of creating a new one (#57513).flatten transformation to be in-place (#56629).with item is not an object (#52335).ModuleList non-literal indexing (#53410).mypy ignore annotation with particular rule specified (#51675).hardtanh(0,6) to the set of MKLDNN fusible ops for mobilenetv2 (#56203).has_bf16_support (#57408).out version for sum (#52225)torch.nn.Linear as aten::linear (#51897).is_tracing scriptable (#49853).Note truncated.
One column per quarter.
The `torch.profiler` submodule is now available. It leveraged the newly released kineto library for profiling. You can find more details in this blogp
torch.profilerThe torch.profiler submodule is now available. It leveraged the newly released kineto library for profiling.
You can find more details in this blogpost: https://pytorch.org/blog/introducing-pytorch-profiler-the-new-and-improved-performance-tool/
The torch.cuda.autocast package can now be used in conjunction with torch xla to provide easy mixed-precision training.
torch. submodule import more autocomplete-friendly (#52339)torch.{isinf,any,all} (#53529)torch.randperm for performance (#54537)torch.distributions validation checks (#53763)nn.Embedding (#53447)Ellipsis in TorchScript (#53766)24 (#54015)int8_t instead of char in {load,store}_scalar (#52616)torch.set_num_thread (#53871)pixel_{un}shuffle (#54178)ValueError (#53548)padding_idx (#53931)This PR removes the deprecated torch.{fft,rfft,ifft,irfft} and their corresponding methods on torch.Tensor. PyTorch programs using these functions mus…
We are excited to announce the availability of PyTorch 1.8. This release is composed of more than 3,000 commits since 1.7. It includes major updates and new features for compilation, code optimization, frontend APIs for scientific computing, and AMD ROCm support through binaries that are available via pytorch.org. It also provides improved features for large-scale training for pipeline and model parallelism, and gradient compression. A few of the highlights include:
torch.fx;torch.fft), Linear Algebra functions (torch.linalg), added support for autograd for complex tensors and updates to improve performance for calculating hessians and jacobians; andAlong with 1.8, we are also releasing major updates to PyTorch libraries including TorchCSPRNG, TorchVision, TorchText and TorchAudio. For more on the library releases, see the post here. As previously noted, features in PyTorch releases are classified as Stable, Beta and Prototype. You can learn more about the definitions in the post here.
You can find more details on all the highlighted features in the PyTorch 1.8 Release blogpost.
Inplace modulo in python %= was wrongfully done out of place for Tensors. This change fixes the behavior.
Previous code that was relying on this operation being done out of place should be updated to use the out of place version t = t % other instead of t %= other.
<p align="center"> <table align="center"> <tr><th>1.7.1</th><th>1.8.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a = torch.arange(0, 10) b = a b %= 3 print(a) tensor([0, 1, 2, 3, 4, 5, 6, 7, 8, 9]) print(b) tensor([0, 1, 2, 0, 1, 2, 0, 1, 2, 0]) </pre></sub></td> <td><sub><pre lang="python"> a = torch.arange(0, 10) b = a b %= 3 print(a) tensor([0, 1, 2, 0, 1, 2, 0, 1, 2, 0]) print(b) tensor([0, 1, 2, 0, 1, 2, 0, 1, 2, 0]) </pre></sub></td> </tr> </table> </p>
torch.clamp edge cases (#43288)For ease of exposition let a_min be the value of the "min" argument to clamp, and a_max be the value of the "max" argument to clamp.
This PR changes the behavior of torch.clamp to always compute min(max(a, a_min), a_max). torch.clamp currently computes this in its vectorized CPU implementation but uses different approaches for other backends.
These implementations are the same when a_min < a_max, but divergent when a_min > a_max. This divergence is easily triggered:
>>> t = torch.arange(200).to(torch.float)
>>> torch.clamp(t, 4, 2)[0]
tensor(2.)
>>> torch.clamp(t.cuda(), 4, 2)[0]
tensor(4., device='cuda:0')
>>> torch.clamp(torch.tensor(0), 4, 2)
tensor(4)
This PR makes the behavior consistent with NumPy's clip. C++'s std::clamp's behavior is undefined when a_min > a_max. Python has no standard clamp implementation.
.grad field (#50663)The deepcopy protocol will now properly copy the .grad field of Tensors when it exists.
The old behavior can be recovered by setting the .grad field to None after doing the deepcopy.
<p align="center"> <table align="center"> <tr><th>1.7.1</th><th>1.8.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
t.grad tensor([0.8883, 0.5765]) deepcopy(t).grad None </pre></sub></td> <td><sub><pre lang="python"> t.grad tensor([0.8883, 0.5765]) deepcopy(t).grad tensor([0.8883, 0.5765]) </pre></sub></td> </tr> </table> </p>
torch.fmod type promotion (#47323, #48278)1.7.1 Raises RuntimeError for integral tensor and floating-point tensor. The dtype of output is determined by the first input.
>>> x = torch.arange(start=1, end=6, dtype=torch.int32) # tensor([1, 2, 3, 4, 5])
>>> y = torch.arange(start=1.1, end=2.1, step=0.2, dtype=torch.float32) # tensor([1.1, 1.3, 1.5, 1.7, 1.9])
>>> torch.fmod(x, y)
RuntimeError: result type Float can't be cast to the desired output type Int
>>> z = torch.arange(start=0.2, end=1.1, step=0.2, dtype=torch.float64) # tensor([0.2, 0.4, 0.6, 0.8, 1.], dtype=torch.float64)
>>> torch.fmod(y, z).dtype
torch.float32
>>> torch.fmod(z, y).dtype
torch.float64
>>> torch.fmod(x, 1.2)
tensor([0, 0, 0, 0, 0], dtype=torch.int32)
1.8.0: Support integral tensor and floating-point tensor as inputs. The dtype of output is determined by both inputs.
>>> x = torch.arange(start=1, end=6, dtype=torch.int32) # tensor([1, 2, 3, 4, 5])
>>> y = torch.arange(start=1.1, end=2.1, step=0.2, dtype=torch.float32) # tensor([1.1, 1.3, 1.5, 1.7, 1.9])
>>> torch.fmod(x, y)
tensor([1.0000, 0.7000, 0.0000, 0.6000, 1.2000])
>>> z = torch.arange(start=0.2, end=1.1, step=0.2, dtype=torch.float64) # tensor([0.2, 0.4, 0.6, 0.8, 1.], dtype=torch.float64)
>>> torch.fmod(y, z).dtype
torch.float64
>>> torch.fmod(z, y).dtype
torch.float64
>>> torch.fmod(x, 1.2)
tensor([1.0000, 0.8000, 0.6000, 0.4000, 0.2000])
All the *_like factory functions will now generate the same striding as out of place operations would.
This means in particular that non-contiguous tensors will produce non-contiguous outputs.
If you require a contiguous output, you can pass the memory_format=torch.contiguous keyword argument to the factory function. Such factory functions include clone, to, float, cuda, *_like, zeros, rand{n}, etc.
torch.norm and torch.linalg.norm consistent for complex inputs (#48284)Previously, when given a complex input, torch.linalg.norm and torch.norm would return a complex output. torch.linalg.cond would sometimes return a complex output and sometimes return a real output when given a complex input, depending on its p argument. This PR changes this behavior to match numpy.linalg.norm and numpy.linalg.cond, so that a complex input will result in a real number type, consistent with NumPy.
torch.svd return V, not V.conj() for complex inputs (#51012)torch.svd added support for complex inputs in PyTorch 1.7, but was not documented as doing so. The complex V tensor returned was actually the complex conjugate of what's expected. This PR fixes the discrepancy.
Users that were already using the previous version of torch.svd with complex inputs can recover the previous behavior by taking the complex conjugate of the returned V.
torch.angle: properly handle pure real numbers (#49163)This PR updates PyTorch's torch.angle operator to be consistent with NumPy's. Previously torch.angle would return zero for all real inputs (including NaN). Now angle returns pi for negative real inputs, zero for non-negative real inputs, and propagates NaNs.
torch.distributions (#48743)This may slightly slow down some models. Concerned users may disable validation by using torch.distributions.Distribution.set_default_validate_args(False) or by disabling individual distribution validation via MyDistribution(..., validate_args=False).
This may cause new ValueErrors in models that rely on unsupported behavior, e.g. Categorical.log_prob() applied to continuous-valued tensors (only {0,1}-valued tensors are supported).
Such models should be fixed but the previous behavior can be recovered by disabling argument validation using the methods mentioned above.
Assigning to a sparse Tensor did not work properly and resulted in a no-op. The following code now properly raises an error:
>>> t = torch.rand(10).to_sparse()
>>> t[0] = 42
TypeError: Cannot assign to a sparse tensor
Tensors cannot be called with ArrayRef<Tensor> anymore (#49138)This PR changes the C++ API representation of lists of optional Tensors (e.g. in the Tensor::``index method) from ArrayRef<Tensor> to List<optional<Tensor>>. This change breaks backwards compatibility, since there is no implicit conversion from ArrayRef<Tensor> to List<optional<Tensor>>.
A common call pattern is tensor.index({indices_tensor}), where indices_tensor is a Tensor. This will continue to work because the {} initializer_list constructor for List<optional<Tensor>> can take Tensor elements that are implicitly converted to optional<Tensor>.
However, another common call pattern is tensor.index(indices_tensor), where previously the Tensor got implicitly converted to an ArrayRef<Tensor>. To implicitly convert Tensor -> optional<Tensor> -> List<optional<Tensor>> would chain two implicit conversions, which C++ doesn't allow. So those call sites should be rewritten to use the tensor.index({indices_tensor}) pattern.
After this fix, an error will properly be thrown to avoid wrong gradients when an in-place operation is performed on a view of a view, when in-place operation were not allowed on the first view.
This means that code that used to return wrong gradients in 1.7.1 (such as t.unbind()[0].select(0, 0).add_(1)) will now properly raise an error.
This PR removes the deprecated torch.{fft,rfft,ifft,irfft} and their corresponding methods on torch.Tensor. PyTorch programs using these functions must now update to use the torch.fft namespace.
torch.digamma : properly handle all inputs (#48302)This PR updates PyTorch's torch.digamma function to be consistent with SciPy's special.digamma function. This changes the result of the torch.digamma function on the nonpositive integers, where the gamma function is not defined. Since the gamma function is undefined at these points, the (typical) derivative of the logarithm of the gamma function is also undefined at these points, and for negative integers this PR updates torch.digamma to return NaN. For zero, however, it returns -inf to be consistent with SciPy.
Interestingly, SciPy made a similar change, which was noticed by at least one user: scipy/scipy#9663
SciPy's returning of negative infinity at zero is intentional: https://github.com/scipy/scipy/blob/59347ae8b86bcc92c339efe213128f64ab6df98c/scipy/special/cephes/psi.c#L163
This change is consistent with the C++ standard for the gamma function: https://en.cppreference.com/w/cpp/numeric/math/tgamma
torch.remainder type promotion (#48668)1.7.1: In the case where the second argument is a python number, the result is casted to the dtype of the first argument.
>>> torch.remainder(x, 1.2)
tensor([0, 0, 0, 0, 0], dtype=torch.int32)
1.8.0 In the case where the second argument is a python number, the dtype of result is determined by type promotion of both inputs.
>>> torch.remainder(x, 1.2)
tensor([1.0000, 0.8000, 0.6000, 0.4000, 0.2000])
The args input argument of the torch.onnx.export function is updated to better support optional arguments. An optional dictionary can be passed in addition as the last argument in the args tuple, specifying inputs with the corresponding named parameter. Note that this is backward breaking for cases where the last input is also of a dictionary type. In the new API, for such cases, it is mandatory to have an empty dictionary as the last argument in the args tuple.
More details can be found at: https://pytorch.org/docs/1.8.0/onnx.html?highlight=onnx#using-dictionaries-to-handle-named-arguments-as-model-inputs.
torch.quantization.quantize function #48537The run_args argument must now contain a list or tuple containing the positional arguments, even if there is only a single argument.
In particular, code like: qmodel = quantize(float_model, default_eval_fn, img_data) that was working in 1.7.1 will now raise the error: TypeError: default_eval_fn() takes 2 positional arguments but 3 were given.
You should update this code to provide the image in a list for example: qmodel = quantize(float_model, default_eval_fn, [img_data])
Starting with version 1.8.0, in the eager mode quantization flow, relu is not observed anymore as it is not needed.
In previous versions, quantized leaky_relu and sigmoid did not require observation and just inherited the quantization parameters from their input, but that does not work very well in eager mode quantization. Starting with version 1.8.0, they are observed operator so that they work better in eager mode quantization.
This update is BC-breaking because the values drawn by the engine will be different from the ones drawn in 1.7.1 even with the same seed.
<p align="center"> <table align="center"> <tr><th>1.7.1</th><th>1.8.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
from torch.quasirandom import SobolEngine eng = SobolEngine(1) eng.draw(3) tensor([[0.5000], [0.7500], [0.2500]]) </pre></sub></td> <td><sub><pre lang="python"> from torch.quasirandom import SobolEngine eng = SobolEngine(1) eng.draw(3) tensor([[0.0000], [0.5000], [0.7500]]) </pre></sub></td> </tr> </table> </p>
nn.Module backward hooks (#46163)Old style nn.Module backward hooks have been broken for a long time (they do not behave as advertised in the documentation). We now have new nn.Module.register_full_backward_hook that provide a fully working implementation of these hooks.
The old function should not be used and migrated to the new full version.
An example of this discrepancy is shown in the example below where a Linear layer takes as input a single Tensor of size 5 and returns a single Tensor of size 5 but old style hook would return two gradients with respect to the input for only one input.
1.7.1:
import torch
from torch import nn
mod = nn.Linear(5, 5)
def hook(mod, grad_inp, grad_out):
print(f"grad input size: " + " ".join(str(g.size()) for g in grad_inp))
print(f"grad output size: " + " ".join(str(g.size()) for g in grad_out))
mod.register_backward_hook(hook)
mod(torch.rand(5, requires_grad=True)).sum().backward()
>>> `grad input size: torch.Size([5]) torch.Size([5]) # One too many
>>> grad output size: torch.Size([5])`
1.8.0: Old style hooks are deprecated and will warn when providing wrong result.
import torch
from torch import nn
mod = nn.Linear(5, 5)
def hook(mod, grad_inp, grad_out):
print(f"grad input size: " + " ".join(str(g.size()) for g in grad_inp))
print(f"grad output size: " + " ".join(str(g.size()) for g in grad_out))
mod.register_backward_hook(hook)
mod(torch.rand(5, requires_grad=True)).sum().backward()
>>> grad input size: torch.Size([5]) torch.Size([5]) # One too many
>>> grad output size: torch.Size([5])
>>> `UserWarning: Using a non-full backward hook when the forward contains multiple
autograd Nodes is deprecated and will be removed in future versions. This hook
will be missing some grad_input.`
Full hooks should be used to get the proper result all the time and avoid warnings
mod.register_full_backward_hook(hook)
mod(torch.rand(5, requires_grad=True)).sum().backward()
>>> grad input size: torch.Size([5])
>>> grad output size: torch.Size([5])
torch.stft: Deprecate default value of the require_complex argument (#49022, #50102)Previously torch.stft took an optional return_complex parameter that indicated whether the output would be a real tensor or a complex tensor. return_complex has the default value of False. This default value is deprecated (meaning that this optional argument is becoming mandatory) and will be removed in future versions. You can pass this argument explicitly to avoid this deprecation.
torch.set_deterministic in favor of torch.use_deterministic_algorithms (#49904)This beta feature is being renamed for improved clarity. Users should migrate to use the new name.
torch.* linear algebra functions in favor of the torch.linalg.* variant for cholesky (#51460), slogdet (#51354), inverse (#51672), pinverse (#51671)All the linear algebra functions are being moved to the torch.linalg submodule that provided a compatible API with NumPy. These new functions have the same set of features as the torch. ones and should be used instead.
torch.nan_to_num (#44592), torch.tensor_split (#45168), torch.nanmedian (#45847), torch.ravel (#46098), torch.igamma (#46183), torch.igammac (#48171), torch.{column_stack,row_stack} (#46313), torch.kron (#45358), torch.copysign (#46396), Tensor.new_empty_strided (#47225), torch.{swapdims,swapaxes} (#46041), torch.tile (#47974), torch.float_power (#44937), torch.moveaxis (#48581), torch.inner (#46716), torch.msort (#48440), torch.sinc (#48740), torch.broadcast_to (#48997), torch.xlogy (#48777), torch.f{max,min} (#49312), torch.diff (#50569), torch.ldexp (#45370), torch.broadcast_shapes (#43935),torch.fft new features: 2D FFT functions (#45164), use new FFT operators in stft (#47601), helper functions (#44877), fuzzing benchmark (#47872)torch.linalg new features: linalg.tensorsolve (#46142), linalg.cholesky (#46083), linalg.tensorinv (#45969), linalg.{eigh,eigvalsh} (#45526), linalg.matrix_rank (#48206), linalg.solve (#48456), linalg.qr (#47764, #50046), linalg.svd (#45562), linalg.inv (#48261), linalg.pinv (#48399), linalg.slogdet (#49194), linalg.cond (#45832)torch.nn Modules: nn.PixelUnshuffle (#49334), nn.GaussianNLLLoss (#50886)torch.nn: new nn.LazyLinear (#44538), nn.LazyConv{1,2,3}d and nn.LazyConvTranspose{1,2,3}d (#47350)torch.nn.AdaptiveAvgPool2d (#48916)cpp_extensions (#47862)torch.futures.Future.add_done_callback (#45675)three_phase optional argument to torch.optim.lr_scheduler.OneCycleLR (#42715)bicubic option for the mode argument of torch.nn.functional.grid_sampler (#44780)torch.distributions: Kumaraswamy (#48285), LKJCholesky (#48798)torch.distributions.OneHotCategorical (#46610)torch.distributions: CorrCholeskyTransform (#48041)torch.distributions: independent (#50547, #50302)close method to torch.hub.tqdm mock (#46040)importance_scores keyword argument (#48378)torch.symeig (#45121), torch.pinverse (#45819), torch.det (#45980), torch.diagflat (#47564), torch.{addcmul, addcdiv} (#46639), torch.lu_solve (#48028), torch.matrix_exp (#48363), torch.eig (#49168), torch.{acosh, asinh, atanh} (#50387), torch.masked_scatter (#51281), torch.bmm and torch.baddbmm (#42553), torch.orgqr (#50502), torch.index_fill_ (#50578), torch.cholesky_inverse (#50269)torch.qr (#45032), torch.lu (#45898), torch.prod(#45980), torch.triangular_solve (#46916), torch.solve (#47045), torch.cholesky_solve (#47047), torch.mean (#47048), torch.svd (#45795), torch.inverse (#47595), torch.Tensor.index_put_ (#51148)torch.trace (#50380)torch.nn.DataParallel (#48686), torch.nn.L1Loss (#49912), Padding functions (#50594)torch.distributed.{all_reduce, all_gather} (#45879, #46270)torch.{atan, log, log10, log1p, log2, reciprocal, tan, pow, rsqrt, tanh, asinh, acosh} (#46275), torch.{cholesky, triangular_solve, mm, mv, ger} (#45737), torch.take(), torch.Tensor.fill_() (#46860), torch.matrix_exp (#48363), torch.{baddbmm, addbmm, addmm, addmv} (#50632), torch.qr (#48489), torch.svd and torch.pinverse (#47761), torch.sqrt (#49461), torch.diag (#51268), torch.trace (#51537), torch.exp (#47194), torch.mean (#47566), torch.addr (#50667), torch.{stack, gather, index_select}, torch.Tensor.index_add_(#49552), torch.{masked_scatter, masked_select} (#51281), torch.{addcmul, addcdiv} (#46639), torch.{acosh, asinh, atanh} (#50387), torch.solve (#47045), torch.cholesky_solve (#47047), torch.inverse (#47595)torch.nn.Module to complex dtypes (#44788)scalar.conj() (#46596)Tensor.copy_() for ComplexHalf tensors (#45339)inputs argument to autograd.backward() both in python and c++ (#46855, #47214)torch.autograd.gradcheck (#45732)vectorize flag to torch.autograd.functional.{jacobian, hessian} (#50915, #51638)torch.lu differentiable. (#46284)torch.no_grad (#49017)BufferedShuffleDataset (#45290)BatchIterDataPipe (#49186, #51880)SamplerIterDataPipe (#49363, #52104)BucketBatchIterDataPipe (#51126, #51880)MapIterDataPipe (#51488#51879)set_per_process_memory_fraction. (#48172)torch.cuda.can_device_access_peer (#50446)torch::nn::ModuleDict (#47707)torch::cuda::synchronize (#50072)torch::jit::freeze C++ api introduced (#52337, #52392)__setitem__ with dynamic shape (#45828)torch.jit.isinstance support for typed containers (#46062)+= style statements) (#44621)in operator with str (#47057)Type::{castRaw,expectRef} (#50061)Union[NoneType, T] as input type (#51605)send and recv in c10d NCCL backend (#44921, #44922)fairscale.nn.Pipe into PyTorch as torch.distributed.pipeline (#44090)--logdir option to log subprocess output to files in DDP launcher. (#33193)RRef.backward() for local RRefs. (#46568) and Owner RRefs. (#46641)"worker_name/device" (#46773)ScriptModule over RPC (#48293)torch.distributed.irecv(src=None, ...) as recv_anysource (#49383)alltoall_single in TorchScript (#48345)TensorPipeAgent (#44418)rref._get_type() (#50498)ZeroRedundancyOptimizer (#46750)Adam optimizer (#50624), sgd optimizer (#50618), Adadelta optimizer (#50623), RMSprop optimizer (#50619), l AdamW optimizer (#50620)set_exception API in torch.futures.Future (#50983)scatter_object_list API for c10d (#43930)Tracer.trace() just return a Graph (#45704)graph_copy examine existing values in val_map (#46104)GraphModule.to_folder (#47544)Node.all_input_nodes (#48270)Graph (#50878)inplace option for fuse_fx (#46953) and prepare_fx (#46954)torch.masked_{scatter,fill} (#45584)torch.{var,var_mean,std_mean} ops (#45678)torch.silu operator support for onnx (#51519)torch.hardswish symbolic in opset 9 (#48423)aten::is_floating point (#46442)torch.logical_{and,or,xor} torch op support in pytorch exporter (#50909)torch.binary_cross_entropy_with_logits op to ONNX opset version 12 (#50908)nn.Squeeze and nn.Unsqueeze (#50906)prim::data (#45747)torch.nonzero(*, as_tuple=True) export (#47421)torch.{cos,sin,tan} (#45733, #46706), torch.log{2,10} (#46810), torch.{a}tanh (#47064), torch.a{cos, tan} (#47005), torch.a{cosh, sinh} (#47152), torch.sqrt (#47293), torch.log1p (#48002). torch.erf{c} (#48472), torch.asin (#48461), torch.sigmoid (#47551), torch.sinh (#48644), torch.cosh (#48923), torch.exp{2, m1}(#48926), torch.reciprocal (#49102), torch.erfinv (#49155), torch.rsqrt (#47909), torch.exp (#50093), torch.lgamma (#50140)dtype argument to Tensor.view (#47951)out optional arguments to torch.{reshape,flatten} (#51249), torch.tensordot (#47278), torch.fft.* (#49335), torch.narrow_copy (#49502)nn.Embedding and nn.EmbeddingBag (#46758)torch.where (#47454), torch.mul and Tensor.__mul__ (#48637), torch.diag (#47455), torch.{all,any} (#44790), Tensor.to_dense (#50019)torch.cum{sum,prod}_ (#47651)torch.sqrt (#50088)dtype and ord arguments in torch.linalg.norm (#46637)torch.nn Module accept batch size of 0: nn.ReplicationPad (#39137), nn.Unfold (#40689), nn.PixelShuffle (#49187), nn.AvgPool{1,2,3}d (#50008), nn.MultiLabelMarginLoss and nn.MultiMarginLoss (#50007)utils.cpp_extensions Ensure default extra_compile_args are properly handled (#45956)torch.LongTensor legacy construction improved error message (#46147)torch.utils.checkpoint allow having Tensors that don’t require gradients (#45934)torch.nan_to_num: fix deprecated warnings (#46309)torch.nn.cpp (#46490), torch.nn.parallel.comm (#46736), torch.nn.modules.* (#46828, #45772, #46013, #49957, #49479, #49045, #49035, #49494, #48969), autograd functions from c++ (#46622), torch.distributed functions from c++ (#46623), torch.storage (#46876), torch._tensor_str (#48463, #48584), torch.nn.modules.pooling (#48412), common_nn (#48190), torch.lobpcg (#47680), torch.nn.functional (#50106), torch.overrides (#50824), torch.generate_torch_version (#51637), torch.distributions (#45689), torch.quantization.quantize_jit (#45548), torch.utils.tensorboard (#49834), torch.multiprocessing (#47756), torch.cuda (#47134), torch._C._distributed_rpc (#46624), torch.distributed.* (#47531, #47532, #47533, #47534), torch.nn.parallel._functions (#49687)torch.svd (#47440)torch.index_copy, torch.median on CUDA and torch.kthvalue on CUDA (#46942)torch.where (#49004), torch.matmul (#47873)torch.flip and torch.flip{lr, ud} (#49895)indices as a Tensor for torch.tensor_split (#49169)torc.nn.init.calculate_gain (#50664)torch.optim optimizers and refactor existing classes to use the functional version: SGD (#45597), Adadelta (#50409), RMSProp (#50410), AdamW (#50411)torch.fft.stft (#51128)torch.div (#51706, #52242)torch.optim.lr_scheduler.LambdaLR (#46813)torch.nn.Unflatten (#49838)torch.multiprocessing.spawntorch.fft submodule (#46004)data.TensorDataset: Add more specific error message (#46905)data.DistributedSampler: Additional validation (#48865)torch.sign for complex tensors (#43280)torch.{min,max} pointwise ops (#50465)torch::cuda::nccl::{send,recv} (#45926)torch.cuda.amp (#48154)torch.utils.checkpoint.checkpoint and torch.cuda.amp at the same time (#49757)DeviceCachingAllocator's error handling more defensive and a bit easier to read (#51158)WorkerInfo and name. (#46221)HashStore getNumKeys() (#46048) and deleteKey() (#46049)ScriptModule methods (#48339)NamedTuple data type to work with DDP (#44220)error_dict in PowerSGD state with bucket index (#48867)CUDAFuture remember and restore current device in callback (#48789)group.WORLD appropriately in process group initialization. (#48767)wait() to value() in some callbacks of PowerSGD communication hook (#49709)find_unused_parameters. (#49908)rref.get_type non-blocking. (#50977)balance and devices parameter from Pipe. (#48432)GradBucket for PowerSGD (#48757)FutureNCCL record streams w/ allocator in addCallback (#48496) and events in current stream (#48497)FutureNCCL callback (#48498)FutureNCCL inside markCompleted() (#48499)FutureNCCL's completed() disagreeing with wait() (#48503)FutureNCCL not recording DataPtrs with caching alloc in wait() (#48563)FutureNCCL (#48500)FutureNCCL (#48501)FutureNCCL (#48502)FutureNCCL's CUDA-specific parts from generic future logic (#48504)at::ivalue::Future (#48505)CUDAFuture from FutureNCCL (#48506)DataPtrs in CUDAFuture (#48788)Pipe to return an RRef. (#47829)dtype of the error tensor (#50985)start_PowerSGD_iter > 1 and add guidance on tuning PowerSGD configs. (#51427)warnings.warn only executes once inside TorchScript. (#45382)Resolver::resolveType (#47731)torch.device in recursive compilation of classes (#47734)__prepare_scriptable__ duck typing to allow replacing nn.Modules with scriptable preparations (#45645) (#49242)torch.quantization.observer (#45630), torch.quantization._numeric_suite (#46330), torch.quantization.stubs (#46475), quantization.fx.Quantizer (#48343), quantization.fx.Quantizer (#48350), quantization_mappings.py (#49179), fusion_patterns.py (#49606), torch/nn/quantized/modules (#49941), quantization-related files in torch/jit (#49939), fuser (#48844), quantization_patterns (#48851), observed_module.py (#49607), quantization (#49942)torch/quantization/fx/* (#48331)Note truncated.
*NOTE*: Conda installs for Python 3.9 will require the conda-forge channel, example: conda install -y -c pytorch -c conda-forge pytorch.
NOTE: Conda installs for Python 3.9 will require the conda-forge channel, example:
conda install -y -c pytorch -c conda-forge pytorch.
This upgrade fix regressions on Ampere cards introduced in cuDNN 8.0.4. It will improve performance for 3090 RTX cards, and may improve performance in other RTX-30 series card.
torch.sqrt: fix wrong output values for very large complex input (#48216)max_pool1d: fix for discontiguous inputs (#48219)collect_env: fix detection of DEBUG flag (#48319)collect_env: Fix to work when PyTorch is not installed (#48311)amp memory usage when running in no_grad() mode (#48936)nn.ParameterList and nn.ParameterDict: Remove spurious warnings (#48215)This change is backward incompatible because TorchScript now attempts to compile user code that creates custom exception messages instead of ignoring…
The PyTorch 1.7 release includes a number of new APIs including support for NumPy-Compatible FFT operations, profiling tools and major updates to both distributed data parallel (DDP) and remote procedure call (RPC) based distributed training. In addition, several features moved to stable including custom C++ Classes, the memory profiler, the creation of custom tensor-like objects, user async functions in RPC and a number of other features in torch.distributed such as Per-RPC timeout, DDP dynamic bucketing and RRef helper.
A few of the highlights include:
To reiterate, starting PyTorch 1.6, features are now classified as stable, beta and prototype. You can see the detailed announcement here. Note that the prototype features listed in this blog are available as part of this release.
FFT-related functionality is commonly used in a variety of scientific fields like signal processing. While PyTorch has historically supported a few FFT-related functions, the 1.7 release adds a new torch.fft module that implements FFT-related functions with the same API as NumPy.
This new module must be imported to be used in the 1.7 release, since its name conflicts with the historic (and now deprecated) torch.fft function.
Example usage:
>>> import torch.fft
>>> t = torch.arange(4)
>>> t
tensor([0, 1, 2, 3])
>>> torch.fft.fft(t)
tensor([ 6.+0.j, -2.+2.j, -2.+0.j, -2.-2.j])
>>> t = tensor([0.+1.j, 2.+3.j, 4.+5.j, 6.+7.j])
>>> torch.fft.fft(t)
tensor([12.+16.j, -8.+0.j, -4.-4.j, 0.-8.j])
Since PyTorch 1.5, we’ve continued to maintain parity between the python and C++ frontend APIs. This update allows developers to use the nn.transformer module abstraction from the C++ Frontend. And moreover, developers no longer need to save a module from python/JIT and load into C++ as it can now be used it in C++ directly.
Reproducibility (bit-for-bit determinism) may help identify errors when debugging or testing a program. To facilitate reproducibility, PyTorch 1.7 adds the torch.set_deterministic(bool) function that can direct PyTorch operators to select deterministic algorithms when available, and to throw a runtime error if an operation may result in nondeterministic behavior. By default, the flag this function controls is false and there is no change in behavior, meaning PyTorch may implement its operations nondeterministically by default.
More precisely, when this flag is true:
torch.backends.cudnn.deterministic = True is set.Note that this is necessary, but not sufficient, for determinism within a single run of a PyTorch program. Other sources of randomness like random number generators, unknown operations, or asynchronous or distributed computation may still cause nondeterministic behavior.
See the documentation for torch.set_deterministic(bool) for the list of affected operations.
Users can now see not only operator name/inputs in the profiler output table but also where the operator is in the code. The workflow requires very little change to take advantage of this capability. The user uses the autograd profiler as before but with optional new parameters: with_stack and group_by_stack_n. Caution: regular profiling runs should not use this feature as it adds significant overhead.
Torchelastic offers a strict superset of the current torch.distributed.launch CLI with the added features for fault-tolerance and elasticity. If the user is not be interested in fault-tolerance, they can get the exact functionality/behavior parity by setting max_restarts=0 with the added convenience of auto-assigned RANK and MASTER_ADDR|PORT (versus manually specified in torch.distributed.launch).
By bundling torchelastic in the same docker image as PyTorch, users can start experimenting with torchelastic right-away without having to separately install torchelastic. In addition to convenience, this work is a nice-to-have when adding support for elastic parameters in the existing Kubeflow’s distributed PyTorch operators.
PyTorch 1.7 introduces a new context manager to be used in conjunction with models trained using torch.nn.parallel.DistributedDataParallel to enable training with uneven dataset size across different processes. This feature enables greater flexibility when using DDP and prevents the user from having to manually ensure dataset sizes are the same across different process. With this context manager, DDP will handle uneven dataset sizes automatically, which can prevent errors or hangs at the end of training.
In the past, NCCL training runs would hang indefinitely due to stuck collectives, leading to a very unpleasant experience for users. This feature will abort stuck collectives and throw an exception/crash the process if a potential hang is detected. When used with something like torchelastic (which can recover the training process from the last checkpoint), users can have much greater reliability for distributed training. This feature is completely opt-in and sits behind an environment variable that needs to be explicitly set in order to enable this functionality (otherwise users will see the same behavior as before).
remote and rpc_synctorch.distributed.rpc.rpc_async has been available in TorchScript in prior releases. For PyTorch 1.7, this functionality will be extended the remaining two core RPC APIs, torch.distributed.rpc.rpc_sync and torch.distributed.rpc.remote. This will complete the major RPC APIs targeted for support in TorchScript, it allows users to use the existing python RPC APIs within TorchScript (in a script function or script method, which releases the python Global Interpreter Lock) and could possibly improve application performance in multithreaded environment.
PyTorch provides a broad set of optimizers for training algorithms, and these have been used repeatedly as part of the python API. However, users often want to use multithreaded training instead of multiprocess training as it provides better resource utilization and efficiency in the context of large scale distributed training (e.g. Distributed Model Parallel) or any RPC-based training application). Users couldn’t do this with with distributed optimizer before because we need to get rid of the python Global Interpreter Lock (GIL) limitation to achieve this.
In PyTorch 1.7, we are enabling the TorchScript support in distributed optimizer to remove the GIL, and make it possible to run optimizer in multithreaded applications. The new distributed optimizer has the exact same interface as before but it automatically converts optimizers within each worker into TorchScript to make each GIL free. This is done by leveraging a functional optimizer concept and allowing the distributed optimizer to convert the computational portion of the optimizer into TorchScript. This will help use cases like distributed model parallel training and improve performance using multithreading.
Currently, the only optimizer that supports automatic conversion with TorchScript is Adagrad and all other optimizers will still work as before without TorchScript support. We are working on expanding the coverage to all PyTorch optimizers and expect more to come in future releases. The usage to enable TorchScript support is automatic and exactly the same with existing python APIs, here is an example of how to use this:
import torch.distributed.autograd as dist_autograd
import torch.distributed.rpc as rpc
from torch import optim
from torch.distributed.optim import DistributedOptimizer
with dist_autograd.context() as context_id:
# Forward pass.
rref1 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 3))
rref2 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 1))
loss = rref1.to_here() + rref2.to_here()
# Backward pass.
dist_autograd.backward(context_id, [loss.sum()])
# Optimizer, pass in optim.Adagrad, DistributedOptimizer will
# automatically convert/compile it to TorchScript (GIL-free)
dist_optim = DistributedOptimizer(
optim.Adagrad,
[rref1, rref2],
lr=0.05,
)
dist_optim.step(context_id)
Support for using the PyTorch profiler in conjunction with the RPC framework was first introduced in PyTorch 1.6. In PyTorch 1.7, the following enhancements have been made:
rpc.functions.async_execution).User are now able to use familiar profiling tools such as with torch.autograd.profiler.profile() and with torch.autograd.profiler.record_function, and this works transparently with the RPC framework with full feature support, profiles asynchronous functions, and TorchScript functions.
PyTorch 1.7 brings prototype support for DistributedDataParallel and collective communications on the Windows platform. In this release, the support only covers Gloo-based ProcessGroup and FileStore.
To use this feature across multiple machines, please provide a file from a shared file system in init_process_group.
# initialize the process group
dist.init_process_group(
"gloo",
# multi-machine example:
# Shared files need six "/"
# init_method = `"file://////{machine}/{share_folder}/file"`
# Local file need three "/"
init_method="file:///{your local file path}",
rank=rank,
world_size=world_size
)
model = DistributedDataParallel(local_model, device_ids=[rank])
PyTorch Mobile supports both iOS and Android with binary packages available in Cocoapods and JCenter respectively. You can learn more about PyTorch-Mobile here.
On some mobile platforms, such as Pixel, we observed that memory is returned to the system more aggressively. This results in frequent page faults as PyTorch being a functional framework does not maintain state for the operators. Thus outputs are allocated dynamically on each execution of the op, for the most ops. To ameliorate performance penalties due to this, PyTorch 1.7 provides a simple caching allocator for CPU. The allocator caches allocations by tensor sizes and, is currently, available only via the PyTorch C++ API. The caching allocator itself is owned by client and thus the lifetime of the allocator is also maintained by client code. Such a client owned caching allocator can then be used with scoped guard, c10::WithCPUCachingAllocatorGuard, to enable the use of cached allocation within that scope.
Example usage:
#include <c10/mobile/CPUCachingAllocator.h>
.....
c10::CPUCachingAllocator caching_allocator;
// Owned by client code. Can be a member of some client class so as to tie the
// the lifetime of caching allocator to that of the class.
.....
{
c10::optional<c10::WithCPUCachingAllocatorGuard> caching_allocator_guard;
if (FLAGS_use_caching_allocator) {
caching_allocator_guard.emplace(&caching_allocator);
}
....
model.forward(..);
}
.....
NOTE: Caching allocator is only available on mobile builds, thus the use of caching allocator outside of mobile builds won’t be effective.
torch.conj now returns the input as-is for real Tensors (#43270)Previously, torch.conj and Tensor.conj were making a clone for Tensors of real dtype. It now returns the Tensor as-is to improve performance.
You can recover the original behavior by adding a .clone() for real Tensors.
Note that this behavior is different from numpy for which np.conj returns a new ndarray and ndarray.conj returns the ndarray as-is.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
t.is_complex() False t.conj() is t False
</pre></sub></td> <td><sub><pre lang="python"> t.is_complex() False t.conj() is t True t.conj().clone() is t False
</pre></sub></td> </tr> </table> </p>
torch.tensor, torch.as_tensor, and torch.sparse_coo_tensor now use the input Tensor’s device when it is not specified (#41984)This will change the device on which the Tensor is created and so the user can start seeing device mismatch errors.
It also means for sparse Tensors that both of the provided Tensors must be on the same device if the device is not specified.
You can recover the original behavior by passing the device argument.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
t.device device(type=‘cuda:0’)
tensor constructor
torch.tensor(t, dtype=torch.float32).device device(type=‘cpu’)
sparse constructor
torch.sparse_coo_tensor( torch.tensor(([0], [2]), device="cpu"), torch.tensor(([1.],), device="cuda"), size=(3, 3, 1)).device device(type='cuda', index=0)
</pre></sub></td> <td><sub><pre lang="python"> t.device device(type=‘cuda:0’)tensor constructor
torch.tensor(t, dtype=torch.float32).device device(type=‘cuda:0’)
Specify the device to get the same behavior as 1.6
torch.tensor(t, dtype=torch.float32, device='cpu').device device(type=‘cpu’)
sparse constructor
torch.sparse_coo_tensor( torch.tensor(([0], [2]), device="cpu"), torch.tensor(([1.],), device="cuda"), size=(3, 3, 1)).device RuntimeError: backend of indices (CPU) must match backend of values (CUDA)
Specify the device to get the same behavior as 1.6
torch.sparse_coo_tensor( torch.tensor(([0], [2]), device="cpu"), torch.tensor(([1.],), device="cuda"), size=(3, 3, 1), device="cuda:0").device device(type='cuda', index=0)
</pre></sub></td> </tr> </table> </p>
torch.nn.utils.pack_padded_sequence: remove hidden cross-device copy for lengths (#41984)In previous versions, when the lengths argument was a CUDA tensor, it would incorrectly be moved to the CPU silently.
This can lead to surprising performances and CPU/GPU sync when using CUDA so this has been removed.
You need to make sure that the provided lenghts is a CPU Tensor when it is provided as a Tensor.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
inp = torch.rand(10, 2, 3, device="cuda") lengths = torch.tensor([10, 7], device="cuda") torch.nn.utils.rnn.pack_padded_sequence(inp, lengths)
Implicitly move lengths to the CPU and runs fine
</pre></sub></td>
<td><sub><pre lang="python">
inp = torch.rand(10, 2, 3, device="cuda") lengths = torch.tensor([10, 7], device="cuda") torch.nn.utils.rnn.pack_padded_sequence(inp, lengths) RuntimeError: 'lengths' argument should be a 1D CPU int64 tensor, but got 1D cuda:0 Long tensor
Ensure the lenghts is already on the right device
lengths = lengths.cpu() torch.nn.utils.rnn.pack_padded_sequence(inp, lengths)
Runs fine with no implicit move across device
</pre></sub></td>
</tr>
</table> </p>
torch.norm handling of keepdim=True (#41956)Before this change, when calling torch.norm with keepdim=True and p='fro' or p=number, leaving all other optional arguments as their default values, the keepdim argument would be ignored. It is now properly respected.
Also, any time torch.norm was called with p='nuc' and keepdim=True, the result would have one fewer dimension than the input, and the dimensions could be out of order depending on which dimensions were being reduced. It is now properly keeping all the dimensions.
You can recover the original behavior by setting keepdim=False.
NOTE: this function is now deprecated (see below) and we recommend you use torch.linalg.norm, which follows NumPy’s conventions.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
t.size() torch.Size([4, 4]) t.norm(p=‘fro’, keepdim=True).size() torch.size([]) t.norm(p=3, keepdim=True).size() torch.size([]) t.norm(p=‘nuc’, keepdim=True).size() torch.size([1]) </pre></sub></td> <td><sub><pre lang="python"> t.size() torch.Size([4, 4]) t.norm(p=‘fro’, keepdim=True).size() torch.size([1, 1]) t.norm(p=3, keepdim=True).size() torch.size([1, 1]) t.norm(p=‘nuc’, keepdim=True).size() torch.size([1, 1])
</pre></sub></td> </tr> </table> </p>
torch.split and torch.chunk: Fix view tracking for the autograd (#41567)The autograd system is able to correctly handle modifications through views of Tensors by explicitly tracking known view operations. In prior releases, torch.split and torch.chunk were not marked as known view operations, which could lead to silently wrong gradients.
Note that since v1.5, inplace modification of views created by functions that return multiple views is deprecated. Such case is not properly handled by the autograd and can lead to internal errors or wrong gradients. So, as a side effect of this view fix, inplace modifications of the outputs of torch.split and torch.chunk will now raise a warning and can lead to internal errors or wrong gradients while they were previously silently computing wrong gradients.
If you see such a warning, you should replace the inplace operation with an out of place one.
You can recover the original behavior by using the new torch.unsafe_split and torch.unsafe_chunk. Note that these functions are only here to ease the transition and will also be removed in a future version.
torch.{argmin,argmax} now always return the first min/max index (#42004)torch.argmin (torch.argmax) now always returns the index of the first minimum (maximum) element. This choice is consistent with NumPy. Previously if there were multiple minima (maxima) the index returned could be the index of any of them.
You cannot recover the original behavior as it was platform dependent and not guaranteed. If your code was relying on a specific index for your specific platform, you should update it to work with the first index and this new code will work on all platforms.
torch.{min,max,median}: Update backward formula when doing full reduction (dim argument not provided) (#43519)When no dimension is specified, full reduction is performed and the gradient will now flow back evenly towards all the input that realized the output value. The old behavior was to propagate the gradient only for one of such input selected arbitrarily.
This should improve stability of training by gradient descent.
To recover the previous behavior, you can perform the reduction with the dim= argument. It will ensure that the gradient only flows back for the input whose index was returned.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a tensor([3, 2, 3]) a.max().backward() a.grad tensor([0, 0, 1])
</pre></sub></td> <td><sub><pre lang="python"> a tensor([3, 2, 3]) a.max().backward() a.grad tensor([0.5, 0, 0.5]) a.max(dim=0).max(dim=0).max(dim=0).backward() a.grad tensor([0, 0, 1])
</pre></sub></td> </tr> </table> </p>
nn.BCELoss size mismatch warning is now an error (#41426)This is the end of the deprecation cycle for this op to make sure it does not have different broadcasting semantic compared to numpy’s broadcasting semantic used everywhere else in PyTorch’s codebase. You need to make sure all inputs are the same size to avoid the error.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
bceloss = nn.BCELoss() a = torch.rand(25) b = torch.rand(25, 1) bceloss(a, b) UserWarning: Using a target size (torch.Size([25, 1])) that is different to the input size (torch.Size([25])) is deprecated. Please ensure they have the same size. tensor(1.0604)
</pre></sub></td> <td><sub><pre lang="python"> bceloss = nn.BCELoss() a = torch.rand(25) b = torch.rand(25, 1) bceloss(a, b) ValueError: Using a target size (torch.Size([25, 1])) that is different to the input size (torch.Size([25])) is deprecated. Please ensure they have the same size. b = b.reshape(25) bceloss(a, b) tensor(1.0604) </pre></sub></td> </tr> </table> </p>
autograd.Function stop materializing None output Tensors (#41490)To improve performance, the custom autograd.Function will not create a Tensor full of zeros when an input is differentiable but the user’s backward function returns None for it. This means that code for which the .backward() or autograd.grad() final result will now be None while it used to be a Tensor full of zeros.
You can recover the previous behavior by having your custom autograd.Function materialize the zero Tensor with torch.zeros_like(input) to replace the None output for the backward method.
import torch
# Custom Function that returns None for the gradient
class GetTwos(torch.autograd.Function):
@staticmethod
def forward(ctx, inp):
return inp.clone().fill_(2)
@staticmethod
def backward(ctx, grad_out):
# To recover the 1.6 behavior, replace the line below with `return torch.zeros_like(grad_out)`
return None
a = torch.rand(10, requires_grad=True)
b = GetTwos.apply(a)
b.sum().backward()
print(a.grad)
# In PyTorch 1.6 this will print
# tensor([0., 0., 0., 0., 0., 0., 0., 0., 0., 0.])
# In PyTorch 1.7 this will print
# None
We fixed a bug in the inplace detection code that was preventing the detection of some inplace operations for output that are not differentiable (like integer type Tensors). This can lead to code that used to run fine to throw the error “a Tensor that was needed for backward was modified in an inplace operation”. Such failure is true and the user code must be fixed to compute proper gradients. In general, this involves cloning the Tensor before modifying it inplace to make sure the backward pass can happen safely.
import torch
a = torch.rand(10, requires_grad=True)
with torch.no_grad():
a[2] = 10
b, ind = a.max(dim=0)
# ind is 2 here
with torch.no_grad():
t = torch.rand(10)
t[4] = 10
res = torch.max(t, dim=0, out=(torch.Tensor(), ind))
# ind becomes 4 here
# This backward runs in 1.6 but will fail in 1.7
b.sum().backward()
print(a.grad)
# tensor([0., 0., 0., 0., 1., 0., 0., 0., 0., 0.])
# The value is wrong is at index 4 while it should be at index 2
# The issue is avoided by not modifying ind inplace by replacing the line
# above with:
# res = torch.max(t, dim=0, out=(torch.Tensor(), ind.clone()))
__torch_functions__ for methods (#37091)Functions, slicing and Tensor methods will now properly preserve the subclass type when possible.
>>> class SubTensor(torch.Tensor):
... pass
>>> type(torch.add(SubTensor([0]), SubTensor([1]))).__name__
'SubTensor'
>>> type(torch.add(SubTensor([0]), torch.Tensor([1]))).__name__
'SubTensor'
The old behavior of “any operations on your subclass produces a torch.Tensor instead of the subclass” can be recovered by doing:
from torch._C import _disabled_torch_function_impl
class SubTensor(torch.Tensor):
__torch_function__ = _disabled_torch_function_impl
For all details on how to use this feature, please refer to the doc page for it.
tensor.__iter__: Use torch.unbind instead of a for loop (#40884)This improves performances significantly but it changes the behavior of in-place operations on the value returned by the iterator. This happens only if either the input Tensor or any argument of the in-place operation is a Tensor that requires gradients. And it will fail with "Output X of UnbindBackward is a view and is being modified inplace".
You can recover the previous behavior by manually slicing the Tensor: [t[i] for i in range(t.size(0))] as shown in the example below.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
x = torch.randn(5, 10, requires_grad=True) for i, v in enumerate(x): v.fill_(i) </pre></sub></td> <td><sub><pre lang="python"> x = torch.randn(5, 10, requires_grad=True) for i, v in enumerate([x[j] for j in range(x.size(0))]): v.fill_(i) </pre></sub></td> </tr> </table> </p>
It fixes silent correctness errors: something that used to be silently incorrect now errors out. Code that raises this error must be updated to avoid doing such op that was returning wrong results as shown in the example below:
>>> x = torch.randn(1, 3)
>>> # Create a tensor that has internal memory overlap
>>> y = x.expand(2, 3)
# In 1.6, this would not error out, but in 1.7, this errors out
>>> torch.nn.functional.elu(y, inplace=True)
RuntimeError: unsupported operation: more than one element of the written-to tensor refers to a single m
emory location. Please clone() the tensor before performing the operation.
# Here is the fix in 1.7
>>> torch.nn.functional.elu(y, inplace=False)
c++ API: Any external users of TensorIterator now always get the memory overlap check. The previous behavior can be recovered by setting set_check_mem_overlap(false) when creating the iterator.
@property of TorchScript classes and ScriptModules. Custom setters and getters are also supported. Custom deleters are not supported.Modules. If these properties use Python or Pytorch features that are not supported in Torchscript, scripting will fail.@torch.jit.unused to annotate problematic properties, the other is to update the implementation of the property so that the getter and setter are scriptable.torch.absolute_ has been removed, the Tensor method (Tensor.absolute_) should be used instead just like all other inplace ops.torch.ExtraFilesMap is an internal jit construct and should not be used.In 1.7, we are enabling a Profiling Executor and a new Tensor-Expressions-based (TE) Fuser. All compilations will now go through one (an adjustable setting) profiling run and one optimization run. For the profiling run, complete tensor shapes are recorded and used by the new Fuser. For the optimization run, the focus is on finding (in torch.jit.ScriptModules) and fusing element-wise operations over CUDA tensors into a single CUDA kernel.
The TE fuser is expected to deliver performance similar to the old fuser used in 1.6. It however unlocks more opportunities for performance improvements in future releases. In rare cases, performance of some models may degrade 5-10%. If you experience any regressions please report it on Github, so we can address them as soon as possible! For 1.7, we are providing an option for our users to revert back to the old fuser by calling torch._C._jit_set_profiling_executor(False) in Python and torch::jit::getExecutorMode()`` = false; in C++. For more information, please see “Graph Executor” section in our documentation.
torch.norm and torch.functional.norm are deprecated in favor of torch.linalg.norm (#44321)The new torch.linalg.norm has the same behavior as numpy.linalg.norm
Both deprecated functions had odd behaviors for matrix and vector norms. You should refer to the doc here to find the exact behavior they had and how to replicate it with the new API.
torch. namespace in favor of torch.fft. namespace (#44876)Please use torch.fft.foo as a drop-in replacement for torch.foo for the following functions: fft, ifft, rfft and irfft.
out= functions need to resize an output which is not 0-size (#42079)This behavior is dangerous and leads to an API that is hard to use. It is being deprecated to be able to fix that API in future versions. You should resize the output before-hand to avoid any issue in the future:
a = torch.rand(5)
b = torch.rand(25)
# This is deprecated
torch.add(a, a, out=b)
# This has the same behavior but will work in future versions
torch.add(a, a, out=b.resize_(0))
torch.optim: Warn for duplicate params in param group (#41597)Providing multiple times the same Parameter in a single param group is most likely due to user error and is being deprecated. Please open an issue if you have a valid use case that require this feature.
torch.linspace and torch.logspace: Not giving the step argument is deprecated (#43860)The default steps argument that has been used historically in PyTorch is not consistent with other libraries and so is being removed to avoid confusion.
For both functions, passing steps=100 keyword argument can be used to recover the original behavior.
<p align="center"> <table align="center"> <tr><th>1.6.0</th><th>1.7.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.linspace(0, 10).size() torch.Size([100])
</pre></sub></td> <td><sub><pre lang="python"> torch.linspace(0, 10).size() UserWarning: Not providing a value for linspace's steps is deprecated and will throw a runtime error in a future release. torch.Size([100]) torch.linspace(0, 10, steps=100).size() torch.Size([100]) </pre></sub></td> </tr> </table> </p>
ProcessGroup and ProcessGroup::Work APIs which will be retired soon. (#46366)New namespaces:
New operators:
torch.count_nonzero added (#39992)nn.SiLU activation added (#41034)torch.logit added (#41062)torch.gcd, torch.lcm added (#40651, #41552, #42254)torch.functional.atleast_{1d/2d/3d} added (#41317)torch.isreal added (#41298)nn.Unflatten added (#41564)torch.movedim added (#41480)torch.isposinf, torch.isneginf added (#41588)torch.signbit added (#41589)torch.absolute added (#42586)torch.clip alias added (#42770)torch.quantile added (#42755)torch.linalg.det and torch.outer alias added (#42802)torch.nansum added (#38628)torch.hypot added (#42291)torch.nextafter added (#42580)torch.hstack, torch.vstack, torch.dstack added (#42799)torch.arccosh alias added (#43107)Tensor.movedim as a method added (#43122)torch.matrix_exp added (#40161)torch.fix alias added (#43326)torch.arccos, torch.arcsin, torch.arctan aliases added (#43319)torch.negative alias added (#43400)torch.maximum, torch.minimum added (#42579)torch.arctanh, torch.arcsinh aliases added (#43762)torch.linalg.norm added (#42749, #43907)torch.amax, torch.amin added (#43819)torch.heaviside added (#42523)torch.i0 added (#43132)torch.not_equal, torch.greater, torch.greater_equal, torch.less, torch.less_equal aliases added (#43870)torch.exp2 added (#44184)torch.kaiser_window added (#44271)torch.nanquantile added (#44393)torch.multiply, torch.divide aliases added (#44463)nn.TripletMarginWithDistanceLoss added (#43680)torch.fft.fft, torch.fft.ifft, torch.fft.rfft, torch.fft.irfft, torch.fft.hfft, torch.fft.ihfft added (#43011)torch.fft.fftn, torch.fft.ifftn, torch.fft.rfftn, torch.fft.irfftn added (#44550)optim.functional.adagrad added (#44715)optim.functional.adam added (#44791)torch.complex, torch.polar added (#39617)Tensor.__complex__ added (#43844)torch.vdot added (#43004)API extension:
torch.full added support for bool and integer dtypes (#41912)torch.lt and torch.masked_select added support for half dtype (#43704)torch.div, torch.true_divide, torch.atan2 added support for integer to float type promotion in (#42359)unflatten added support for non-named dimensions (#42563)torch.polygamma added support for n >= 2 (#42499)torch.qr added backward support for wide input matrices (#42216)nn.Linear for MKLDNN added support for no-bias (#43703)torch.lerp added support for half dtype (#43541)torch.div to perform true division (end of deprecation cycle) (#42907)torch.scatter added support for reductions on CUDA (#41977)torch.pow (#44760), unary ops and activations (#44813, #44824, #44834), torch.i0 (#44750), softmax (#44837), div, addcdiv, addcmul, mean, var (#44758), layernorm (#45002),all pooling layers (#44836, #45151)), torch.logspace (CPU and CUDA) (#44675), random kernels on Windows (#44918), torch.addmm, torch.addmv (#44986), loss functions (#45011), batched gemm (#45167), nccl path (#38515), binary logical operators (#42485), torch.neg (#45240), Conv (non-cuDNN) (#45007), torch.abs (#44804), torch.erfinv (#43399), comparison ops (#44748)torch.asin, torch.neg added support for sparse Tensors (#44028)torch.softmax added support for CUDA (#42307)Tensor.{real,imag} added setter for these attributes (#39860)torch.{addmm,addmv} added support for complex on CUDA (#40431, #43827)torch.bmm added support for complex on CPU #42383,torch.{dot, vdot} added support for complex (#42745)torch.stft, torch.istft added support for complex (#43886)torch.cholesky added support for complex (#44895, #45267)torch.sgn added (to support complex) (#39955)autograd.Function (#41821)autograd.functional API (#43428)reset_grad API to remove gradient instead of setting them to zero (#44423)torch.autograd.gradcheck (#43877)@torch.no_grad() decorator (#44633)torch.lobpcg backward (#43002)torch.cuda.amp.GradScaler now supports sparse gradients (#36786)backends.cudnn.allow_tf32 flag to control it (#40737)torch.cuda.memory.list_gpu_processes to list running processes on a give GPU (#44616)atomicAdd() (#41538)nn::TransformerEncoderLayer added (#42633)nn::TransformerDecoderLayer added (#42717)nn::TransformerEncoder added (#43187)nn::TransformerDecoder added (#42886)nn::Transformer added (#44333)nn::Unflatten added (#42613)nn.ParameterList added (#41259)torch::cuda::manual_seed and torch::cuda::manual_seed_all added (#42638)mobile_optimized boolean flag to optimized model. (#45479)adaptive_avg_pool2d (#41220), mm (#41221), reshape (#41223), max_pool2d (#41379), add_ and relu_ (#41380), cat (#41434), add and mul (#42674) and avg_pool2d (#42675).torch.utils.optimize_for_vulkan (#44903)torch.set_deterministic and torch.is_deterministic: Raise error when the flag is set and a non-deterministic operation is used (#15359, #41377)CONTRIBUTING.md (#42635, #43294)prefetch_factor argument to control the number of batch loaded ahead of time(#41130)np.memmap objects (#39847)utils.cpp_extension (#41257, #43528)nn.ReflectionPad: Add support for 0-dim batch sizes. (#39231)torch.scatter: Add reductions for CPU (#36447)torch.iinfo, torch.finfo: Improve printing (#40488)torch.where: Add support for scalar input (#40336)torch.nonzero: Remove deprecation warning for as_tuple argument (#45413)torch.distributions.Categorical: Clamp logit to avoid -inf when calculating entropy (#41002)torch.futures.Future: Add done function to query the status of the future (#42013)nn.EmbeddingBag: Add support for incude_last_offset=True when reduction is mean or max (#42215)nn.AvgPooling{1,2,3}d: Ensure all cells are valid in ceil mode to avoid division by 0 (#41368)nn,[Adaptive]MaxPool{1,2,3}d: Handle edge case when input is filled with -inf (#40665)nn.Hardsigmoid, nn.Hardswish: Add inplace option (#42346)nn.MSELoss, nn.L1Loss, nn.SmoothL1Loss: Add support for target that requires gradients. (#44437, #44471, #44486)nn.Parameter{List,Dict}: Add warning when improperly used (with DataParallel or weight_norm) (#44405)nn.functional.smooth_l1: Add beta parameter (#44433)torch.cuda.nccl APIs. (#43247)@torch.no_grad (#41371)del to TorchScript classes (#44352)@torch.jit.unused syntax for ignoring properties (#45261)@torch.jit.unused on a @torch.no_grad decorated function (#41496)to_backend API now accepts wrapped modules (#43612)whitelist to allowlist (#41771, #41802)dequantize now supports list and tuple of tensors (#41079)register_activation_post_process_hook function (#42342)add/mul now support different variants (#42769)OP_LIST_TO_FUSER_METHOD is exposed to the user (#43286)quantize_jit can handle new upsample overloads (#43407)convert_jit can now take preserved_attrs argument (#44490)SyncBN: preserve qconfig if it exists (#45317)state_dict (#44846)conv parameters (#43524, #43086, #43651, #44671)In PyTorch 1.7, we have continued to add and improve PyTorch operator export to ONNX. We have enabled export of 10 new operators, and further enhanced and optimized export of 10+ torch operators to ONNX. We have also focused on improving export of TorchScript modules, in particular laying some groundwork required for better support in near future. We have also created an API (torch.onnx.utils._find_missing_ops_onnx_export) as a diagnostic tool (preview only) to get a list of operators in a model that are not supported or implemented by ONNX exporter. Support for export of torch.quantization.FakeQuantize has also been added to help enable some QAT workflows.
Add support to export more torch ops torch.view_as (#40496), fake quantize functions (#39738), embedding_bag (#41234, #44693), torch.eye (#41357), Tensor.as_strided (#41569), torch.tensor (#41872), addition between list of tensors (#41888), Tensor.__floordiv__ (#43022), torch.nn.KLDivLoss (#41858), Tensor.new_empty and Tensor.new_zeros (#43506)
Improves existing export logic and optimizing exported ONNX graph
torch.full_like (#40063)torch.where export, add support for ByteTensor (#42264)torch.scatter export, add support for src being scalar or different dtype (#42765, #43440)torch.where (#41544)torch.slice (#42935), torch.split (#43670), torch.repeat (#43430), torch.arange (#43777), len (#43824), torch.narrow (#44039), flatten (#40418), adaptive_pool (#46100)Update export to follow pytorch changes
torch.utils.collect_env: Collect more informations (python 32/64bit, clang version, CPU architecture, ROCm version) (#42887, #42961, #44106)torch.hub.load_local: Allow to load models from any local directory (#44204)import torch is called from the source root (#39995)--continue-through-error option to run_test.sh script (#41136)run_name and ``hparam_domain_discreteinadd_hparams` (#40660, #40720)with_source parameter to enable tracking source code (#43898)zero_grad, avoid using inpalce detach when it is not required (#41283)torch.div backward formula to improve numerical stability (#43627)torch.repeat as only input.dim() is needed in backward (#40766)device_count and cuda init error detection and messages (#42249)torch/*.py (#40235, #40873)Tensor attributes and methods: T and grad_fn (#40879), Tensor._version (#41125), ndim (#42909), nonzero (#43053), #40499)torch.serialization (#40862)torch.tensor (#45077)torch.Size (#40879)torch.futures (#41675)torch.random (#42234)torch.hub (#42252)collect_env.py (#43062)torch.utils (#39392, #42647, #42711, #42960, #43806, #44136, #44216)torch.nn (#43044, #44093, #43080, #42231, #40669)torch.sparse (#43108)torch.cuda.nvtx (#43443)torch.cuda.memory (#43444)torch.functional (#43446)torch.autograd (#44451, #46206)torch.quantization.fuse_modules (#43786)torch.nn.quantized (#43186, #44154, #43110)torch.testing._internal submodules (#44575, #44805, #44832, #44911, #44927, #44985, #44971, #45107, #45368, #45375)torch.backends.quantized ([#44794](httpsNote truncated.
The following are a list of BC-breaking changes to some of PyTorch’s internal components.
The PyTorch 1.6 release includes a number of new APIs, tools for performance improvement and profiling, as well as major updates to both distributed data parallel (DDP) and remote procedure call (RPC) based distributed training.
A few of the highlights include:
Additionally, from this release onward, features will be classified as Stable, Beta and Prototype. Prototype features are not included as part of the binary distribution and are instead available through either building from source, using nightlies or via compiler flag. You can learn more about what this change means in the post here.
AMP allows users to easily enable automatic mixed precision training enabling higher performance and memory savings of up to 50% on Tensor Core GPUs. Using the natively supported torch.cuda.amp API, AMP provides convenience methods for mixed precision, where some operations use the torch.float32 (float) datatype and other operations use torch.float16 (half). Some ops, like linear layers and convolutions, are much faster in float16. Other ops, like reductions, often require the dynamic range of float32. Mixed precision tries to match each op to its appropriate datatype.
PyTorch 1.6 introduces a new backend for the RPC module which leverages the TensorPipe library, a tensor-aware point-to-point communication primitive targeted at machine learning, intended to complement the current primitives for distributed training in PyTorch (Gloo, MPI, ...) which are collective and blocking. The pairwise and asynchronous nature of TensorPipe lends itself to new networking paradigms that go beyond data parallel: client-server approaches (e.g., parameter server for embeddings, actor-learner separation in Impala-style RL, ...) and model and pipeline parallel training (think GPipe), gossip SGD, etc.
# One-line change needed to opt in
torch.distributed.rpc.init_rpc(
...
backend=torch.distributed.rpc.BackendType.TENSORPIPE,
)
# No changes to the rest of the RPC API
torch.distributed.rpc.rpc_sync(...)
The torch.autograd.profiler API now includes a memory profiler that lets you inspect the tensor memory cost of different operators inside your CPU and GPU models.
Here is an example usage of the API:
import torch
import torchvision.models as models
import torch.autograd.profiler as profiler
model = models.resnet18()
inputs = torch.randn(5, 3, 224, 224)
with profiler.profile(profile_memory=True, record_shapes=True) as prof:
model(inputs)
# NOTE: some columns were removed for brevity
print(prof.key_averages().table(sort_by="self_cpu_memory_usage", row_limit=10))
# --------------------------- --------------- --------------- ---------------
# Name CPU Mem Self CPU Mem Number of Calls
# --------------------------- --------------- --------------- ---------------
# empty 94.79 Mb 94.79 Mb 123
# resize_ 11.48 Mb 11.48 Mb 2
# addmm 19.53 Kb 19.53 Kb 1
# empty_strided 4 b 4 b 1
# conv2d 47.37 Mb 0 b 20
# --------------------------- --------------- --------------- ---------------
PyTorch Distributed supports two powerful paradigms: DDP for full sync data parallel training of models and the RPC framework which allows for distributed model parallelism. Currently, these two features work independently and users can’t mix and match these to try out hybrid parallelism paradigms.
Starting PyTorch 1.6, we’ve enabled DDP and RPC to work together seamlessly so that users can combine these two techniques to achieve both data parallelism and model parallelism. An example is where users would like to place large embedding tables on parameter servers and use the RPC framework for embedding lookups, but store smaller dense parameters on trainers and use DDP to synchronize the dense parameters. Below is a simple code snippet.
// On each trainer
remote_emb = create_emb(on="ps", ...)
ddp_model = DDP(dense_model)
for data in batch:
with torch.distributed.autograd.context():
res = remote_emb(data)
loss = ddp_model(res)
torch.distributed.autograd.backward([loss])
RPC Asynchronous User Functions supports the ability to yield and resume on the server side when executing a user-defined function. Prior to this feature, when an callee processes a request, one RPC thread waits until the user function returns. If the user function contains IO (e.g., nested RPC) or signaling (e.g., waiting for another request to unblock), the corresponding RPC thread would sit idle waiting for these events. As a result, some applications have to use a very large number of threads and send additional RPC requests, which can potentially lead to performance degradation. To make a user function yield on such events, applications need to: 1) Decorate the function with the @rpc.functions.async_execution decorator; and 2) Let the function return a torch.futures.Future and install the resume logic as callbacks on the Future object. See below for an example:
@rpc.functions.async_execution
def async_add_chained(to, x, y, z):
return rpc.rpc_async(to, torch.add, args=(x, y)).then(
lambda fut: fut.wait() + z
)
ret = rpc.rpc_sync(
"worker1",
async_add_chained,
args=("worker2", torch.ones(2), 1, 1)
)
print(ret) # prints tensor([3., 3.])
This release adds support for a language-level construct as well as runtime support for coarse-grained parallelism in TorchScript code. This support is useful for situations such as running models in an ensemble in parallel, or running bidirectional components of recurrent nets in parallel, and allows the ability to unlock the computational power of parallel architectures (e.g. many-core CPUs) for task level parallelism.
Parallel execution of TorchScript programs is enabled through two primitives: torch.jit.fork and torch.jit.wait. In the below example, we parallelize execution of foo:
import torch
from typing import List
def foo(x):
return torch.neg(x)
@torch.jit.script
def example(x):
futures = [torch.jit.fork(foo, x) for _ in range(100)]
results = [torch.jit.wait(future) for future in futures]
return torch.sum(torch.stack(results))
print(example(torch.ones([])))
The minimum version of Python we support now is 3.6. Please upgrade your Python to match. If you use conda, instructions for setting up a new environment with Python >= 3.6 can be found here.
torch.div and torch.addcdiv integer floor division behavior (#38762, #38620)In 1.5.1 and older PyTorch releases torch.div , torch.addcdiv, and the / operator perform integer floor division. In 1.6 attempting to perform integer division throw a RuntimeError, and in 1.7 the behavior will change so that these operations always perform true division (consistent with Python and NumPy division).
To floor divide integer tensors, please use torch.floor_divide instead.
<p align="center"> <table align="center"> <tr><th>1.5.1</th><th>1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor(3) / torch.tensor(2) ../aten/src/ATen/native/BinaryOps.cpp:81: UserWarning: Integer division of tensors using div or / is deprecated, and in a future release div will perform true division as in Python 3. Use true_divide or floor_divide (// in Python) instead. tensor(1) </pre></sub></td> <td><sub><pre lang="python">
NB: the following is equivalent to
torch.floor_divide(torch.tensor(3), torch.tensor(2))
torch.tensor(3) // torch.tensor(2) tensor(1) </pre></sub></td> </tr> </table> </p>
The fix for torch.addcdiv is similar.
<p align="center"> <table align="center"> <tr><th>1.5.1</th><th>1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
input = torch.tensor(0) tensor = torch.tensor(1) other = torch.tensor(3) value = 1 torch.addcdiv(input, tensor, other, value=value) ../aten/src/ATen/native/PointwiseOps.cpp:81: UserWarning: Integer division with addcdiv is deprecated, and in a future release addcdiv will perform a true division of tensor1 and tensor2. The current addcdiv behavior can be replicated using floor_divide for integral inputs (self + value * tensor1 // tensor2) and division for float inputs (self + value * tensor1 / tensor2). The new addcdiv behavior can be implemented with true_divide (self + value * torch.true_divide(tensor1, tensor2). tensor(0) </pre></sub></td> <td><sub><pre lang="python"> input = torch.tensor(0) tensor = torch.tensor(1) other = torch.tensor(3) value = 1 (input + torch.floor_divide(value * tensor, other)) tensor(0) </pre></sub></td> </tr> </table> </p>
In previous versions of PyTorch, zero dimensional CUDA tensors could be moved across devices implicitly while performing binary pointwise operations (e.g. addition, subtraction, multiplication, division, and others). For example,
torch.tensor(5, device='cuda:0') + torch.tensor((1, 1), device='cuda:1')
would work, even though the tensors are on different CUDA devices. This is a frequent source of user confusion, however, and PyTorch generally does not move data across devices without it being explicit. This functionality is removed in PyTorch 1.6.
To perform binary pointwise operations on data of different devices, please cast the tensors to the correct device by using Tensor.to:
<p align="center"> <table align="center"> <tr><th>Version 1.5.1</th><th>Version 1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor(5, device='cuda:0') + torch.tensor((1, 1), device='cuda:1') torch.tensor([6, 6], device='cuda:1') </pre></sub></td> <td><sub><pre lang="python"> torch.tensor(5, device='cuda:0').to('cuda:1') + torch.tensor((1, 1), device='cuda:1') torch.tensor([6, 6], device='cuda:1') </tr> </table> </p>
In previous versions of PyTorch, we provided an installation option for Windows environments running CUDA 9.2. Starting from PyTorch 1.6.0, we are no longer providing those binaries. Please upgrade your CUDA version to 10.1 or 10.2 and install a PyTorch binary for one of those CUDA versions instead.
To check whether you are affected, please find your GPU in a table inthis link.
If you are using a Nvidia GPU with compute capability 6.1, you may notice a performance hit when using the release binaries (installed via pip or conda). We stopped building for CUDA compute capability 6.1 but PyTorch programs should still continue to work with those devices. If you do notice a performance hit, a workaround is to compile PyTorch from source.
If you are using a Nvidia GPU with compute capability 3.7 and relied on PTX, we have dropped support for that in our release binaries (installed via pip or conda). Potential workarounds are: install a previous version of PyTorch or to compile PyTorch from source.
In previous versions of PyTorch, when a bool tensor is constructed from a floating-point tensor, we would first convert the tensor to a long tensor, then to float tensor. This is not consistent with how bools are interpreted in Python, C++, and NumPy (just to name a few), which interpret 0 floating-point values as False and everything else as True.
If you were relying on the previous behavior, the following code will achieve the same effect.
<p align="center"> <table align="center"> <tr><th>Version 1.5.1</th><th>Version 1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor([-2, -1, -0.9, 0, 0.9, 1, 2], dtype=torch.bool) tensor([ True, True, False, False, False, True, True]) </pre></sub></td> <td><sub><pre lang="python"> torch.tensor([-2, -1, -0.9, 0, 0.9, 1, 2]).long().bool() tensor([ True, True, False, False, False, True, True]) </tr> </table> </p>
In PyTorch 1.6 bool and integral fill values given to torch.full must set the dtype our out keyword arguments. In prior versions of PyTorch these fill values would return float tensors by default, but in PyTorch 1.7 they will return a bool or long tensor, respectively. The documentation for torch.full has been updated to reflect this.
In previous versions of PyTorch, running .backward() in multiple threads causes them to be serialized in a specific order, resulting in no parallelism on CPU. In PyTorch 1.6.0, running .backward() in multiple threads no longer serializes the execution and instead autograd will run those in parallel.
This is BC-breaking for the following two use cases:
In more detail, in 1.6.0, when you run backward() or grad() via python, TorchScript or the C++ API in multiple threads on CPU, you should expect to see extra concurrency. For example, you can manually write multithreaded Hogwild training code like:
# Define a train function to be used in different threads
def train_fn(model, input):
# forward
y = model(input)
# backward
y.sum().backward()
# potential optimizer update
# define your model in python or in TorchScript
model = Model()
# User write their own threading code to drive the train_fn
threads = []
for _ in range(10):
# define or load the data
input = torch.ones(5, 5, requires_grad=True)
p = threading.Thread(target=train_fn, args=(model, input))
p.start()
threads.append(p)
for p in threads:
p.join()
Note when you use the same model and call backward() concurrently in multiple threads, model parameters are automatically shared across threads. The gradient accumulation might become non-deterministic as two backward calls might access and try to accumulate the same .grad attribute. Although we do proper locking to avoid data corruption, we don't guarantee the order in which the ops are executed, so non-determinism might arise, but this is an expected pattern in multithread training. You could use the functional API torch.autograd.grad() to calculate the gradients instead of backward() to avoid the non-determinism.
For thread safety:
.grads that match the weights' memory layout (#40358)In previous versions of PyTorch, autograd would yield contiguous gradients. Now, gradients have the same memory layout as their respective weights. This should result in silent performance improvements. Since PyTorch operators generally support non-contiguous tensors, this should have no functional effect on most PyTorch programs. A known exception is when accessing param.grad and performing an operation that requires a contiguous tensor, such as param.grad.view(-1). In this case, you will receive an error as follows:
RuntimeError: view size is not compatible with input tensor's size and stride (at least one dimension spans across two contiguous subspaces). Use .reshape(...) instead.
If a user wants to force accumulation into a grad with a particular layout, they can preset param.grad to a zeroed tensor with the desired strides or manually set grad to have the desired strides ( param.grad = param.grad.contiguous(desired format).)
See the below section on “Note: BC-breaking memory format changes” for more details.
In previous versions of PyTorch, performing a binary pointwise operation between a Contiguous and a Channels Last tensor produced a Channels Last. In PyTorch 1.6, this now returns a tensor with the layout of the first operand.
See the below section on“Note: BC-breaking memory format changes” for more details.
Operations that now return tensors in a different memory format generally should have no functional effect on most PyTorch programs because PyTorch operators generally support non-contiguous tensors.
The most common incompatibility with Python programs is with the view operator, which has specific stride requirements. If these requirements are no longer met as a result of this change, you will get an error message indicating that you should use reshape instead, i.e. "RuntimeError: view size is not compatible with input tensor's size and stride (at least one dimension spans across two contiguous subspaces). Use .reshape(...) instead."
Another possible exception incompatibility is if you have a (usually) C++ operator implementation that works directly on memory (i.e. calls data_ptr and relies on the strides being contiguous).
nn.functional.interpolate: recompute_scale_factor default behavior changed from True to False (#39453)In PyTorch 1.5.1 and older versions, nn.functional.interpolate(input, size, scale_factor, ..., recompute_scale_factor) has a default of recompute_scale_factor = True. In PyTorch 1.6, we’ve changed the default to recompute_scale_factor = False.
Depending on the precision of the scale_factor, this may result in an output tensor with different values than before. To retain the old behavior, simply change your code to use recompute_scale_factor = True.
More concretely, what recompute_scale_factor = True means is, if the user passes in a scale_factor:
scale_factor by dividing the output size by the input size and sending it to an internal helper function.scale_factor is used in the interpolate computation but in some cases is different from the scale_factor the user passed in.This behavior resulted in loss of precision so we deprecated it in PyTorch 1.5.0. In PyTorch 1.6 and onward, recompute_scale_factor has a default of False, which means that we pass it directly to an internal helper function.
out= arguments of pointwise and reduction functions no longer participate in type promotion (#39655)In PyTorch 1.5 passing the out= kwarg to some functions, like torch.add, could affect the computation. That is,
out = torch.add(a, b)
could produce a different result than
torch.add(a, b, out=out)
This is because previously the out argument participated in the type promotion rules. For greater consistency with NumPy, Python, and C++, in PyTorch 1.6 the out argument no longer participates in type promotion, and has no effect on the computation performed.
torch.quasirandom.SobolEngine(..., scramble=True, seed=None) to respect torch.manual_seed when a seed has not been provided (#36427)In previous versions of PyTorch, SobolEngine(..., scramble=True, seed=None) did not respect any calls to torch.manual_seed. The expected behavior for random number generation functions is to respect the seed set by torch.manual_seed, so we’ve changed SobolEngine to match.
If you were relying on the old behavior where SobolEngine ignores torch.manual_seed, please explicitly pass a different seed to SobolEngine:
<p align="center"> <table align="center"> <tr><th>Version 1.5.1</th><th>Version 1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.manual_seed(1337)
x1 = SobolEngine(dimension=1, scramble=True, seed=None).draw(3)</pre></sub></td> <td><sub><pre lang="python"> import time torch.manual_seed(1337)
ms_since_epoch = int(round(time.now() * 1000)) x1 = SobolEngine(dimension=1, scramble=True, seed=ms_since_epoch).draw(3) </pre></sub></td> </tr> </table> </p>
Tensor.random_(to, from): Enforce check that from and to are within the bounds of the Tensor’s dtype (#37507)In previous versions of PyTorch, to and from did not have to be within the bounds of the tensor’s dtype (this raised a warning). The behavior of random_ in that case can be unexpected. We are making this a hard error starting from PyTorch 1.6.0; please modify your code if you run into the error.
<p align="center"> <table align="center"> <tr><th>Version 1.5.1</th><th>Version 1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
tensor = torch.zeros(10, dtype=torch.uint8)
to for torch.uint8tensor.random_(0, 257) UserWarning: to - 1 is out of bounds for unsigned char. </pre></sub></td> <td><sub><pre lang="python"> tensor = torch.zeros(10, dtype=torch.uint8)
to for torch.uint8tensor.random_(0, 256) </pre></sub></td> </tr> </table> </p>
If you build PyTorch from source, we’ve dropped support for using CUDA < 9.2 (run nvcc --version to check your CUDA version). Users who install PyTorch packages via conda and/or pip are unaffected.
DataLoader’s __len__ changed to return number of batches when holding an IterableDataset (#38925)In previous versions of PyTorch, len(<instance of dataloader holding an IterableDataset>) would return the number of examples in the dataset. We’ve changed it to be the number of batches (e.g., the number of examples divided by the DataLoader’s batch_size) to be consistent with the computation of length when the DataLoader has a BatchedSampler.
torch.backends.cudnn.flags: deleted unused verbose flag (#39228)The verbose flag did nothing, so we deleted it. If you were passing a value to flags for verbose, please remove it.
RpcBackendOptions takes float instead of timedelta for timeout argument to stay consistent with timeout types in other TorchScriptable RPC APIs.
# v1.5
rpc.init_rpc(
"worker1",
rank=0,
world_size=2,
rpc_backend_options=rpc.ProcessGroupRpcBackendOptions(
num_send_recv_threads=16,
datetime.timedelta(seconds=20)
)
)
# v1.6
rpc.init_rpc(
"worker1",
rank=0,
world_size=2,
rpc_backend_options=rpc.ProcessGroupRpcBackendOptions(
num_send_recv_threads=16,
20 # seconds
)
)
We rolled back to the old fuser and the legacy executor in this release in order to recover some reported performance regressions. In future releases we plan to reach the same or better performance with a new redesigned executor and fuser.
In order to switch back to the executor used in the 1.5 release one could use the following API:
torch._C._jit_set_profiling_executor(True) before you call your model for the first time,#include <torch/csrc/jit/runtime/graph_executor.h> and set getExecutorMode() = true before you invoke your model for the first time.Note: this isn’t actually BC-breaking but we are listing it here because it is BC-Improving.
The PyTorch Team recommends saving and loading modules with the same version of PyTorch. Older versions of PyTorch may not support newer modules, and newer versions may have removed or modified older behavior. These changes are explicitly described in PyTorch’s release notes, and modules relying on functionality that has changed may need to be updated to continue working properly.
In this release, the historic behavior of torch.div and torch.full is preserved for models saved via torch.jit.save in previous versions of PyTorch. Modules saved with the current version of PyTorch will use the latest torch.div and torch.full behavior. See the notes above for the BC changes to those operators.
The following are a list of BC-breaking changes to some of PyTorch’s internal components.
KernelFunction::makeFromUnboxedFunctorFactory; use makeFromUnboxedFunctor directly instead (#35488)autograd.gradcheck and autograd.gradgradcheck: Added a new default-true argument check_undefined_grad (#39400)Internally, in the autograd engine, we use a special undefined Tensor value to represent zero-filled gradients and expect backward functions and user-defined torch.autograd.Functions to gracefully handle those values. When check_undefined_grad is True (the default for PyTorch 1.6+), gradcheck/gradgradcheck test that the operation in question supports undefined output gradients. This may cause a previously succeeding gradcheck to fail.
You can turn the check off by setting check_undefined_grad to False. As long as autograd does not error out due to an undefined gradient in your model, then everything should be fine.
<p align="center"> <table align="center"> <tr><th>Version 1.5.1</th><th>Version 1.6.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.autograd.gradcheck(my_custom_function, inputs) True </pre></sub></td> <td><sub><pre lang="python">
To keep the previous behavior
torch.autograd.gradcheck(my_custom_function, inputs, check_undefined_grad=False) True </pre></sub></td> </tr> </table> </p>
TensorIterator is an implementation detail for writing kernels that is exposed in our C++ API. We’ve modified how developers interact with TensorIterator, please see the Pull Request for more details.
torch._min and torch._max(#38440)torch._min and torch._max are undocumented and were intended to be an implementation detail; we expect very few users, if any at all, to be using it. We’ve deleted it in PyTorch 1.6.0. Please use torch.min/torch.max instead if you are using torch._min/torch._max.
torch.save serialization format (#39460, #39893, #40288, #40793)We have switched torch.save to use a zip file-based format by default rather than the old Pickle-based format. torch.load has retained the ability to load the old format, but use of the new format is recommended. The new format is:
__getstate__, __setstate__) functions on Modules that depended on serialized Tensor values were getting the wrong dataUsage is as follows:
m = MyMod()
torch.save(m.state_dict(), 'mymod.pt') # Saves a zipfile to mymod.pt
To use the old format, pass the flag _use_new_zipfile_serialization=False
m = MyMod()
torch.save(m.state_dict(), 'mymod.pt', _use_new_zipfile_serialization=False) # Saves pickle
Calling torch.nonzero(tensor, as_tuple=False) with one argument or Tensor.nonzero(as_tuple=False) with no arguments is deprecated and will be removed in a future version of PyTorch. Please specify the as_tuple argument.
New Utilities
torch.nn.Module (#38972)TORCH_SHOW_CPP_STACKTRACES=1 (#38127)torch.utils.show_pickle for showing pickle contents in saved models (#35168)New Operators
torch.logcumsumexp added (#36308)torch.logaddexp added (#38384)torch.rad2deg, torch.deg2rad added (#38852)torch.arccosh, torch.arcsinh, torch.arctanh added (#38388)torch.flip{lr, ud} added (#38599)torch.bucketize, torch.searchsorted added (#34577)torch.istft (Inverse Short Time Fourier Transform) added (#35569)torch.vander: added support for generating Vandermonde matrices (#36725)torch.block_diag added (#33449)nn.Hardswish, nn.functional.hardswish added (#34747)torch.nn.init.trunc_normal_ (truncated normal initializer) added (#32397)torch.optim.AveragedModel and torch.optim.SWALR for more details.(#35032)AdamW to C++ frontend (#40009)load_inline under ROCm (#35897)The PyTorch 1.6 release brings beta-level support for complex tensors. The UX is similar to existing PyTorch tensors and the new complex-specific functionality is compatible with NumPy’s complex arrays. In particular, you’ll be able to create and manipulate complex tensors, interop with previously existing code that represented complex tensors as tensors of size (..., 2), and more.
While this is an early version of this feature, and we expect it to improve over time, the overall goal is provide a NumPy compatible user experience that leverages PyTorch’s ability to run on accelerators and work with autograd to better support the scientific computing and ML communities.
Please find the full documentation here.
Python API:
torch.is_signed() for complex tensors. (#33773)torch.randn and torch.normal_ for complex tensors. (#34037, #35056)torch.full. (#34709)is_complex tensor attribute for complex numbers. (#34093)torch.rand for complex dtypes. (#34924, #35585)torch.copy_ , on cuda. (#35344)torch.from_numpy for complex dtypes. (#35531)torch.exp CPU implementation for complex tensors. (#35715)torch.masked_fill for complex tensors. (#36335)torch.abs to return float tensors for complex tensors. (#35871)torch.isfinite and torch.isinf for complex tensors. (#36648)torch.isclose for complex tensors. (#36456)torch.angle to return float tensors for complex tensors. (#36896)requires_grad for complex tensors. (#36932)torch.reciprocal for complex tensors on CUDA. (#36749)Complex Storage. (#35771)torch.addmv for complex tensors. (#37924, #40238)torch.tensor . (#38030)torch.pow for complex tensors on CUDA. (#36793)torch.pow .(#36793, #39117)torch.roll for complex tensors on CUDA. (#38664)torch.gather for complex tensors on CPU. (#36430)torch.tanh for complex tensors on CUDA. (#38786)torch.cumsum, torch.cumprod for complex tensors on CUDA. (#39063)real and imag views as tensor attributes. (#39033)torch.flip and torch.rot90 for complex tensors. (#37826)torch.view_as_real, torch.view_as_complex for complex tensors. (#39099)torch.tan for complex tensors on CUDA (#38400)torch.tanh backward function (#37791, #38786)C++ API:
at::tensor() and torch::tensor() for complex numbers (#39793)torch.distributed: Add all_to_all API to the MPI backend in the distributed module (#32361).torch.distributed: Add c10d dynamic loading mechanism to support 3rd-party c10d implementations (#28068).torch.nn.parallel.DistributedDataParallel: Add distributed data parallel benchmark tool (#35198).torch.nn.parallel.DistributedDataParallel and torch.distributed.rpc: allow DDP to work with RPC (#37998, #39916, #40130, #40139, #40495).torch.utils.mobile_optimizer.optimize_for_mobile to encapsulate several model optimizations appropriate for mobile models. (Note: currently broken on Windows.) (#35227) (#36357)PyTorch 1.6 has a new, pybind11-based operator registration API which replaces the torch::RegisterOperators() class.
Before:
static auto registry =
torch::RegisterOperators("my_ops::warp_perspective", &warp_perspective);
After:
TORCH_LIBRARY(my_ops, m) {
m.def("warp_perspective", warp_perspective);
}
You can read more about this API in the custom C++ operators tutorial or the reference documentation.
The new API was developed in PRs #35061, #35629, #35706, #36222, #36223, #36258, #36742, #37019. Internal code was ported to this API in #36799, #36800, #36389, #37834, #38014; you may find the code examples in these PRs helpful for your ports.
In PyTorch 1.6, we have added support for ONNX Opset 12. We have also enhanced export of torchvision models, such as FasterRCNN, MaskRCNN, and KeypointRCNN to support dynamic input image size. Export support for several new ops have also been added. A new operator export mode, ONNX_FALLTHROUGH, has been added to the export API that allows exporting the model with non-standard ONNX operators. For large (> 2 GB) model export (using external_data_format=True argument), we now support models with large tensor data in attributes (not just model parameters).
New ONNX operator support:
aten::max_pool2d to onnx jit pass (#34912)upsample_nearest2d, sigmoid and reshape as no scale in onnx (#36325)New quantization operators:
torch.distributed.rpc: Add TensorPipe RPC backend (#36197, #35483, #37839, #37918, #37919,#37850,#37851, #37852,#37980, #38052, #38265, #38266, #40162, #40389, #37910, #38448, #38818, #38819, #38926, #38931, #38930, #38933, #38934, #39010, #39011, #39397)torch.distributed.rpc: Support per-RPC timeouts for rpc_sync and rpc_async (#34650)torch.distributed.rpc.functions.async_execution: Add an @async_execution decorator to allow pause and resume executions in RPC target functions (#39216, #39267, #39485, #39486, #39758).torch.futures.Future:Expose a Future type to Python API (#39008, #37311, #39119, #39597, #39964, #39950)torch.distributed.rpc: Allow profiler to be enabled remotely with RPC (#38748, #40066)torch.distributed.rpc: Implement TorchScript-compatible RemoteModule API (#37139, #40173)torch.distributed.rpc.RRef: enable retrying RRef control messages on communication failures (#33636)torch.distributed.rpc: Let RPC use torch._C.Future instead of exposing a dedicated future type. No impact on user side (#35039)torch.distributed.autograd: Add profiler support for backward of the distributed autograd engine (#35261)torch.distributed.rpc.RRef: Add TorchScript support for RRef.local_value() (#35433)torch.distributed.rpc.WorkerInfo: Add TorchScript support for WorkerInfo (#35447)torch.distributed.rpc: Allow profiling RPC with TorchScript target functions (#36275)torch.distributed.rpc.RRef: Add RRef Python Helper to launch function on the remotely referenced object (#36619)torch.distributed.rpc: Add timeout argument to TorchScriptable rpc_async (#37884)torch.distributed.rpc: Enable RPC Server Global Profiler (#38847)torch.distributed.rpc: Implement timeout support for rpc.remote and RRef.to_here() (#38590)torch.distributed.rpc: Enable RRef timeout for TensorPipe (#39531)torch.distributed.rpc.WorkerInfo: Add WorkerInfo python __repr__ magic method (#40004)torch.add: Prevent unbounded growth while adding sparse tensors (#36030)torch.mv: enabled for sparse tensors (#21782)torch.bmm: enabled for sparse x dense tensor operations (#33430)torch.cat: improved error message (#38978)torch.masked_select: enabled bfloat16 support (#36859)torch.absolute: added as an alias for torch.abs (#36597)torch.device: improved error message to include xla as an acceptable device (#36446)torch.linspace, torch.logspace: improved precision (#35461)Tensor.true_divide method variant added (#34794)Tensor.isnan(), Tensor.isinf(), Tensor.isfinite() method variants added (#37942)Tensor.is_nonzero: improved error message (#38150)Tensor.cauchy_, Tensor.log_normal_, Tensor.exponential_: added support for bfloat16 (#38427)Tensor.as_subclass method added. (#34369)collect_env.py: improved to detect relevant conda-installed numpy and cudatoolkit (#35646)collect_env.py: made it more robust on Windows (#39136)torch.utils.data: Add generator= kwarg for DataLoader & random samplers (#39737)torch.utils.data.DataLoader: properly diagnose exceeding file descriptor limit (#34768)torch.utils.data.DataLoader: added repr for WorkerInfo (#39975)torch.utils.data.random_split: added option to pass a generator for determinism (#34043)torch.utils.data.IterableDataset: make the warning for when a DataLoader holds an IterableDataset clearer (#41185)torch.nn: Added support for non-persistent buffers that do not show up in a Module’s state dict (#37191)nn.Fold, nn.Unfold: added double backwards support (#36379)nn.MultiheadAttention: added support for bool/byte attn_mask tensor (#33763)nn.functional.upsample: enabled uint8 sampling support (#35029)nn.functional.kl_div: added option to accept target in log space (#34586)nn.functional.softmax: added support for sparse tensors (CPU) (#36305)nn.Softmin, nn.Softmax: improved repr (#39084)torch.cuda: Change DeprecationWarning to FutureWarning (#32142)torch.utils.cmake_prefix_path pointing to share/cmake folder (#38559)torch.hub: Added file_name argument to load_state_dict_from_url (#39749)torch.cuda.get_arch_list() and torch.cuda.get_gencode_flags() added. These return the architecture list and gencode flags PyTorch was compiled with. (#41212)torch.min, torch.max: significantly improved CUDA performance (#38440, #39029)torch.multinomial with replacement=False: significantly improved performance (#39742)torch.autograd: add type hints in-line (#38080)torch.finfo, torch.iinfo type annotations added (#38220)torch.cuda annotations inline (#40075)torch.cuda._CudaStreamBase and torch.cuda._CudaEventBase classes (#40256)torch.types.Device and stubbed all torch._C functions comprehensively (#38173)torch.nn modules type annotations inline (#38211)torch.autograd.anomaly_mode: fixed type hints stub (#39324)torch.backends.cudnn added type annotations (#38947)torch.channels_last, torch.preserve_format: added annotations (#39120)torch.topk: enabled support for BFloat16 type on ROCm. (#34849)torch.dot: enabled fp16 support on ROCm (#30431, #30432)torch.add: enabled support for BFloat16 type on ROCm for sparse tensors(#35978)torch.log: improved ROCm support (#40079)torch.pow, torch.exp, torch.erf: enabled support for BFloat16 type on ROCm (#40236)cpp_extensions on Windows (#35272)
Note: Above two PRs eliminate unnecessary compile warnings for windows build, make build log more readable.torch.distributed: Enhance error message for MPI unavailability. (#36781).torch.distributed: Expose torch.distributed.is_available() API (#37021).torch.utils.data: Only create torch.generator and seed in DistributedSampler when shuffling (#37604).ProcessGroup: Log incorrect device in ProcessGroupGloo (#38844).torch.utils.data: Improve DistributedSampler docs and add seed option (#39628).torch.cuda.comm.reduce: Avoid initializing unnecessary tensors in nccl.reduce (#39688).torch.nn.parallel.DistributedDataparallel: Remove obsolete warning message from DDP (#40190).distributions.Cauchy: Implemented kl divergence (#36477)distributions.Transform: Add a .with_cache() method (#36882)distributions.Binary: Implemented BTRS algorithm for fast/efficient binomial sampling (#36858)TORCH_FN for passing in compile time function pointers as regular function arguments rather than template arguments (#39823, #40110)__torch_function__ benchmarks (#36138)torch.autograd.profiler: Make RecordFunction callbacks thread local and modernize interface (#37491)torch.autograd.profiler: Make profiler thread local (#36291)torch.distributed.rpc.RRef: Throw an actionable error message on user call RRef.to_here() in TorchScript (#35369)torch.distributed.rpc.RRef: Handle exceptions returned via remote() calls (#35331)torch.distributed.rpc.RRef: Make RRef type hint mismatch exception message more actionable to users (#35943)torch.distributed.rpc:Allow abort RecvWork::wait() in ProcessGroupAgent::listenLoop (#36084)torch.distributed.autograd: Appropriately handle exceptions in autograd engine. (#36019)torch.distributed.autograd: Catch exception in distributed engine callbacks. (#36118)torch.distributed.autograd: Avoid some future callback self-captures. (#36502)torch.distributed.rpc: Propagate error from RPC retries to the original attempt (#35263)torch.distributed.autograd: Ensure future is complete when exiting Engine::mark_graph_task_completed() (#36856)torch.distributed.autograd: Trigger pre/post hooks of output function nodes under distributed autograd (#34501)torch.distributed.rpc: Supporting create an RPC gang of world size 1 (#32731)torch.distributed.autograd: Improve Error Message for Dist Autograd Context Cleanup Failure (#37255)torch.distributed.rpc: Guard against negative rpcTimeout being passed in to RpcBackendOptions (#38267)torch.distributed.rpc: Use infinite timeout for operations in ProcessGroup RPC backend (#38577)torch.distributed.rpc.WorkerInfo: Add stringify WorkerInfo (#39974)torch.distributed.rpc: Avoid using default process group in ProcessGroupAgent. (#39909)torch.distributed.rpc: Ignore expected errors in TensorPipe RPC backend (#39182)torch.distributed.rpc: Don't use separate heap allocation for metrics in TensorPipe RPC backend (#39183)torch.distributed.rpc: Bind to hostname's IP address instead of localhost in TensorPipe RPC backend (#39184)torch.distributed.rpc: Use PrefixStore to avoid conflicting keys in TensorPipe RPC backend (#39185)id function (#34975)str to float (#35352)if statements with statically determinable predicates (#35834)toBool (#35570)hardsigmoid, hardswish, and elu to make them scriptable (#35885)strict tracer flag to guard against risky behaviors (#36277)Dict as output when connecting script and tracing (#36265)dtype with torch.tensor when dtype is not specified (#36587)str to int (#36016)Tensor.tolist (#37465)code_with_constants method to module printing (#37586)del statements with variables as targets in TorchScript (#37608)init on custom C++ classes (#37474)@staticmethod access from self on modules (#37702)@torch.jit.unused to be used on TorchScript classes (#38522, #39336)%= operator in TorchScript (#38983)Tensor (#38527)Note truncated.
This most notably affects torch.argmin, torch.argmax, and torch.argsort. This change is BC-Breaking because previously one could obtain an integer-typ
This most notably affects torch.argmin, torch.argmax, and torch.argsort. This change is BC-Breaking because previously one could obtain an integer-type tensor that requires grad in 1.5.0. However, said tensors were not usable by autograd; calling .backward() on them resulted in an error, so most users are likely to not have been relying on this behavior.
<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
tensor = torch.randn(3, requires_grad=True) torch.argmax(tensor).requires_grad True <td><sub><pre lang="python"> tensor = torch.randn(3, requires_grad=True) torch.argmax(tensor).requires_grad False </pre></sub></td> </tr> </table> </p>
You may see error messages like the following when using the torch.multiprocessing package. This bug has primarily affected users with AMD CPUs.
`Error: mkl-service + Intel(R) MKL: MKL_THREADING_LAYER=INTEL is incompatible with libgomp.so.1 library.
Try to import numpy first or set the threading layer accordingly. Set MKL_SERVICE_FORCE_INTEL to force it.`
You can get rid of the error and the error message by setting the environment MKL_THREADING_LAYER=GNU. This can be done either by including the following in your python code:
import os
os.environ['MKL_THREADING_LAYER'] = 'GNU'
or by specifying the environment variable when running your script:
MKL_THREADING_LAYER=GNU python my_script.py
To learn more about what triggers this bug and other workarounds if the above isn’t working, please read this comment on the issue.
torch.multinomial: Fixed a bug where CUDA multinomial generated the same sequence over and over again with a shift of 4. (#38046)nn.Conv2d: Fixed a bug where circular padding applied padding across the wrong dimension (#37881)<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
circular = nn.Conv2d(6, 1, (3, 3), padding=(0, 1), padding_mode='circular') circular(torch.zeros(1, 6, 10, 10)).shape
torch.Size([1, 1, 10, 8]) </pre></sub></td> <td><sub><pre lang="python">
tensor = torch.randn(3, requires_grad=True) other = tensor + 1 output = nn.LeakyReLU(0, inplace=True)(other) output.sum().backward() torch.Size([1, 1, 8, 10]) </pre></sub></td> </tr> </table> </p>
torch.gather, torch.scatter: added checks for illegal input dtypes that caused silently incorrect behaviors (#38025, #38646)torch.argmin, torch.argmax: Fixed silently incorrect result for inputs with more than 2^32 elements (#39212)requires_grad=True flag. (#37355)<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.zeros(5, 14400, 14400, device='cuda').sum(0)RuntimeError: sub_iter.strides(0)[0] == 0 INTERNAL ASSERT FAILED at /pytorch/aten/src/ATen/native/cuda/Reduce.cuh:706, please report a bug to PyTorch.</pre></sub></td> <td><sub><pre lang="python"> torch.zeros(5, 14400, 14400, device='cuda').sum(0)
</pre></sub></td>
</tr>
</table> </p>
<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
pickle.dumps(torch.tanh)PicklingError: Can't pickle <class 'torch._C._VariableFunctions'>: it's not the same object as torch._C._VariableFunctions </pre></sub></td> <td><sub><pre lang="python"> pickle.dumps(torch.tanh)
</pre></sub></td>
</tr>
</table> </p>
nn.LeakyReLU: Fixed a bug where using autograd with in-place nn.LeakyReLu with a slope of 0 incorrectly errored out. (#37453, #37559)<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
tensor = torch.randn(3, requires_grad=True) other = tensor + 1 output = nn.LeakyReLU(0, inplace=True)(other) output.sum().backward() RuntimeError: In-place leakyReLu backward calculation is triggered with a non-positive slope which is not supported. This is caused by calling in-place forward function with a non-positive slope, please call out-of-place version instead. </pre></sub></td> <td><sub><pre lang="python"> tensor = torch.randn(3, requires_grad=True) other = tensor + 1 output = nn.LeakyReLU(0, inplace=True)(other) output.sum().backward()
</pre></sub></td>
</tr>
</table> </p>
torch.as_strided : Fixed crash when passed sizes and strides of different lengths. (#39301)nn.SyncBatchNorm.convert_sync_batchnorm: Fixed bug where it did not respect the devices of the original BatchNorm module, resulting in device mismatch errors (#39344)nn.utils.clip_grad_norm_: Fixed ability to operate on tensors on different devices (#38615)torch.min, torch.max: added check for illegal output dtypes (#38850)import torch error (#36941).This bug mainly affected users of ubuntu 16.04. We’re certain it affected the following configurations:
include_dirs for ahead-of-time compilation (#38264)nn.Conv1d, nn.Conv2d, nn.Conv3d: Fixed a bug where convolutions were using more memory than previous versions of PyTorch. (#38674)In 1.5.0, the in-place floor division magic method mistakingly performed the floor division out-of-place. We’ve fixed this in 1.5.1.
<p align="center"> <table align="center"> <tr><th>Version 1.5.0</th><th>Version 1.5.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
tensor = torch.ones(1) expected_data_ptr = tensor.data_ptr() tensor //= 1 tensor.data_ptr() == expected_data_ptr False <td><sub><pre lang="python"> tensor = torch.ones(1) expected_data_ptr = tensor.data_ptr() tensor //= 1 tensor.data_ptr() == expected_data_ptr True </pre></sub></td> </tr> </table> </p>
Weight quantization was done incorrectly for LSTMs, the statistics for all weights (across layers) were combined in the observer. This meant that weights for later layers in a LSTM would use sub-optimal scales impacting accuracy. The problem gets worse as the number of layers increases.
torch.unique (#38156)pow operator export (#39791)Your coding agent can read these notes before it upgrades. Set up the MCP server →