NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #358 most downloaded on PyPI
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Last release 21 days ago
02 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 49 of 50 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
50 releases · first in 2018
In previous versions of PyTorch, there were two ways to write autograd Functions. We deprecated one of them in 1.3.0 and dropped support for it entire…
This release includes several major new API additions and improvements. These include new APIs for autograd allowing for easy computation of hessians and jacobians, a significant update to the C++ frontend, ‘channels last’ memory format for more performant computer vision models, a stable release of the distributed RPC framework used for model parallel training, and a new API that allows for the creation of Custom C++ Classes that was inspired by PyBind. Additionally torch_xla 1.5 is now available and tested with the PyTorch 1.5 release providing a mature Cloud TPU experience.
The C++ frontend API is now at parity with Python and the features overall has been moved to ‘stable’. (previously tagged as experimental). Some of the major highlights include:
narrow / select / index_select / masked_select, which is clunky and error-prone compared to the Python API’s elegant tensor[:, 0, ..., mask] syntax. With the 1.5 release users can use tensor.index({Slice(), 0, "...", mask}) to achieve the same result.Channels Last memory format is an alternative way of ordering NCHW tensors in memory while preserving the NCHW semantic dimensions ordering. Channels Last tensors are ordered in memory in such a way that channels become the densest dimension (aka storing images pixel-per-pixel).
Channels Last memory format unlocks the ability to use performance efficient convolution algorithms and hardware (NVidia’s Tensor Cores, FBGEMM, QNNPACK). Additionally it was designed to automatically propagate through the operators, which allows easy switching between memory layouts.
Learn more here on how to write memory format aware operators.
This release adds a new API for binding custom C++ classes into TorchScript and Python simultaneously. This API is almost identical in syntax to pybind11. It allows users to expose their C++ class and its methods to the TorchScript type system and runtime system such that they can instantiate and manipulate arbitrary C++ objects from TorchScript and Python. An example C++ binding:
template <class T>
struct MyStackClass : torch::CustomClassHolder {
std::vector<T> stack_;
MyStackClass(std::vector<T> init) : stack_(std::move(init)) {}
void push(T x) {
stack_.push_back(x);
}
T pop() {
auto val = stack_.back();
stack_.pop_back();
return val;
}
};
static auto testStack =
torch::class_<MyStackClass<std::string>>("myclasses", "MyStackClass")
.def(torch::init<std::vector<std::string>>())
.def("push", &MyStackClass<std::string>::push)
.def("pop", &MyStackClass<std::string>::pop)
.def("size", [](const c10::intrusive_ptr<MyStackClass>& self) {
return self->stack_.size();
});
Which exposes a class you can use in Python and TorchScript like so:
@torch.jit.script
def do_stacks(s : torch.classes.myclasses.MyStackClass):
s2 = torch.classes.myclasses.MyStackClass(["hi", "mom"])
print(s2.pop()) # "mom"
s2.push("foobar")
return s2 # ["hi", "foobar"]
You can try it out in the tutorial here.
The torch.distributed.rpc package aims at supporting a wide range of distributed training paradigms that do not fit into DistributedDataParallel. Examples include parameter server training, distributed model parallelism, and distributed pipeline parallelism. Features in the torch.distributed.rpc package can be categorized into four main sets of APIs.
SGD, Adagrad, etc.) and a list of parameter RRefs, and its step() function automatically uses the local optimizer to update parameters on all distinct RRef owner workers.Learn more here.
torch_xla is a Python package that uses the XLA linear algebra compiler to accelerate the PyTorch deep learning framework on Cloud TPUs and Cloud TPU Pods. torch_xla aims to give PyTorch users the ability to do everything they can do on GPUs on Cloud TPUs as well while minimizing changes to the user experience. This release of torch_xla is aligned and tested with PyTorch 1.5 to reduce friction for developers and to provide a stable and mature PyTorch/XLA stack for training models using Cloud TPU hardware. You can try it for free in your browser on an 8-core Cloud TPU device with Google Colab, and you can use it at a much larger scale on Google Cloud.
See the full torch_xla release notes here and the full docs here.
PyTorch 1.5 brings new functions including jacobian, hessian, jvp, vjp, hvp and vhp to the torch.autograd.functional.* submodule. This feature builds on the current API and allow the user to easily perform these functions.
See the full docs here.
For PyTorch 1.5.0 we will no longer support Python 2, specifically version 2.7. Going forward support for Python will be limited to Python 3, specifically Python 3.5, 3.6, 3.7 and 3.8 (first enabled in PyTorch 1.4.0).
torch.nn.parallel.DistributedDataParallel does not work in Single-Process Multi-GPU mode.DistributedDataParallel (DDP) used to support two modes
module to all specified devices and trains on all module replicas. This mode is enabled when application passes in a device_ids argument that contains multiple devices. Or if device_ids is not presented, DDP will try to use all available devices.module without creating additional replicas. This mode is enabled when device_ids only contains a single device or if there is only one visible device (e.g., by setting CUDA_VISIBLE_DEVICES).A recent change (#33907) in torch.nn.parallel.replicate breaks DDP’s assumption on replicated modules and leads to failures in the SPMG mode. However, since SPMG is known to be slower due to GIL contention and additional overhead caused by scattering input and gathering output, we are planning to retire this mode in future releases and make MPSG the only supported mode in DDP. The code below shows an example of the recommended way to construct DDP.
import torch
from torch.nn.parallel import DistributedDataParallel as DDP
# use "cuda:1" as the target device
target_device = 1
local_model = torch.nn.Linear(2, 2).to(target_device)
ddp_model = DDP(local_model, device_ids=[target_device])
See #36268 for more discussion.
Tensor.exponential_(0) used to return Inf, now it incorrectly returns 0Previously in 1.4, x.exponential_(0) gives a tensor full of inf. On 1.5.0, it wrongly gives a tensor full of zeros.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.randn(3).exponential_(0) tensor([inf, inf, inf]) </pre></sub></td> <td><sub><pre lang="python"> torch.randn(3).exponential_(0)
tensor([0., 0., 0.]) </pre></sub></td> </tr> </table> </p>
See #36798 for more details
Tensor.clone, Tensor.to, Tensor.empty_like, and similar functions preserve stride information instead of returning contiguous tensorsclone, to, type, cuda, cpu, byte, char, double, bool, half, int, long, short, float, bfloat16, empty_like, full_like, ones_like, zeros_like, rand_like, randn_like, randint_like operators now propagate memory format (roughly, the strides) of the input tensor to the output tensor.
Since PyTorch operators generally support non-contiguous tensors, this should have no functional effect on most PyTorch programs.
The most common incompatibility with Python programs is with the view operator, which has specific stride requirements. If these requirements are no longer met as a result of this change, you will get an error message indicating that you should use reshape instead, i.e. "RuntimeError: view size is not compatible with input tensor's size and stride (at least one dimension spans across two contiguous subspaces). Use .reshape(...) instead."
Another possible exception incompatibility is if you have a (usually) C++ operator implementation that works directly on memory (i.e. calls data_ptr and relies on the strides being contiguous).
In the following example, we go through the implementation of a simple clone operation and see how it needs to change between versions.
# Version 1.4.0
Tensor simple_clone(const Tensor& input) {
TORCH_CHECK(input.dim() == 1);
auto output = at::empty_like(input);
auto input_stride = input.strides()[0];
auto* output_ptr = output.data_ptr<float>();
auto* input_ptr = input.data_ptr<float>();
// Before 1.5.0, the result of `empty_like` is always contiguous.
for (int64_t idx = 0; idx < input.size(); idx++) {
output[idx] = input[idx * input_stride]
}
}
# Version 1.5.0
Tensor simple_clone(const Tensor& input) {
TORCH_CHECK(input.dim() == 1);
// From 1.5.0 on, the result of `empty_like` may not be contiguous.
auto output = at::empty_like(input);
// As a result, we need to keep track of the output stride.
auto input_stride = input.strides()[0];
auto output_stride = output.strides()[0];
auto* output_ptr = output.data_ptr<float>();
auto* input_ptr = input.data_ptr<float>();
for (int64_t idx = 0; idx < input.size(); idx++) {
output[idx * output_stride] = input[idx * input_stride]
}
}
Please explicitly pass in the desired dtype when constructing tensors with NumPy float64 scalars to get the old behavior.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor(np.float64(0)) tensor(0.) </pre></sub></td> <td><sub><pre lang="python">
torch.tensor(np.float64(0), dtype=torch.get_default_dtype()) tensor(0.) </pre></sub></td> </tr> </table> </p>
This can cause your program to execute in torch.float64, potentially slowing down your program or can lead to errors for operators that don't support torch.float64 or mixed-dtypes.
numpy integer scalars are now treated as integers for the purposes of type promotion (#30486 (https://github.com/pytorch/pytorch/pull/30486))
Previously, in 1.4.0, they were mistakenly treated as floats (so for example, torch.ones(3) * np.int64(3) would return a float32 tensor. In 1.5.0, we’ve fixed that behavior; torch.ones(3) * np.int64(3) returns an int32 tensor.
This can cause your code to fail if you performed operations between PyTorch tensors and numpy scalars and then passed the result into an operation that does not support integral types or mixed types. To fix your code, please cast the resulting tensor to the desired dtype.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.ones(3) * np.int64(3) tensor([3., 3., 3.]) </pre></sub></td> <td><sub><pre lang="python"> (torch.ones(3) * np.int64(3)).float() tensor([3., 3., 3.]) </pre></sub></td> </tr> </table> </p>
Previously, in 1.4.0, they were mistakenly treated as floats (so for example, torch.ones(3) * np.int64(3) would return a float32 tensor. In 1.5.0, we’ve fixed that behavior; torch.ones(3) * np.int64(3) returns an int32 tensor.
This can cause your code to fail if you performed operations between PyTorch tensors and numpy scalars and then passed the result into an operation that does not support integral types or mixed types. To fix your code, please cast the resulting tensor to the desired dtype.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.ones(3) * np.int64(3) tensor([3., 3., 3.]) </pre></sub></td> <td><sub><pre lang="python"> (torch.ones(3) * np.int64(3)).float() tensor([3., 3., 3.]) </pre></sub></td> </tr> </table> </p>
torch.autograd.Function: dropped support for old-style Functions (#33956).In previous versions of PyTorch, there were two ways to write autograd Functions. We deprecated one of them in 1.3.0 and dropped support for it entirely in 1.5.0. Old-style autograd Functions will no longer work in user code.
These Functions be identified by not having staticmethod forward and backward functions (see the example below) Please see the current documentation for how to write new-style Functions.
# Version 1.4.0
class Exp(torch.autograd.Function):
def forward(self, i):
result = i.exp()
self.save_for_backward(result)
return result
def backward(self, grad_output):
result, = self.saved_tensors
return grad_output * result
Exp()(torch.tensor(1.))
# Version 1.5.0
class Exp(torch.autograd.Function):
@staticmethod
def forward(ctx, i):
result = i.exp()
ctx.save_for_backward(result)
return result
@staticmethod
def backward(ctx, grad_output):
result, = ctx.saved_tensors
return grad_output * result
Exp.apply(torch.tensor(1.))
torch.optim optimizers changed to fix in-place checks for the changes made by the optimizer (#33640, #34211)If this causes your code to fail, there are two possible reasons:
Reason 1: The value of that parameter was actually saved and used and we were computing incorrect gradients in previous versions of PyTorch. This would result in an error message mentioning incorrect version numbers. You should replace code that uses self.my_param by self.my_param.clone() to make sure the saved version is different from the one that is modified by the optimizer. For example:
Before 1.5.0, the following may have worked.
def model(input, target, param):
return `(input * param ** 2 - target).norm()`
param = torch.randn(2, requires_grad=True)
input = torch.randn(2)
target = torch.randn(2)
sgd = optim.SGD([param], lr=0.001)
loss = model(input, target, param)
loss.backward(retain_graph=True)
sgd.step()
loss.backward()
param.grad
If after upgrading to 1.5.0, the above fails due to a version counter error, then that means the gradient computed was incorrect. To remedy this, clone param before using it in the model:
def model(input, target, param):
return (input * param ** 2 - target).norm()
param = torch.randn(2, requires_grad=True)
input = torch.randn(2)
target = torch.randn(2)
sgd = optim.SGD([param], lr=0.001)
loss = model(input, target, param.clone())
loss.backward(retain_graph=True)
sgd.step()
loss.backward()
param.grad
Reason 2: You know what you're doing and change the values back to the right thing before the next backward. However, you're running into an error because the version counter cannot be decremented. Open an issue with your particular use case and we will help you to work around the version counter issue.
utils.cpp_extensions now use ninja as the default compilation backend (#32495)ninja enables parallel compilation of your C++ extension, greatly speeding up compilation. This change will not break most user code; if you do not have ninja installed, we fallback to the old distutils backend.
However, if you do have ninja installed, it is possible that this change will cause your C++ extension build to fail by oversubscribing your system with too many worker processes. There are two potential workarounds to this.
Method 1: If a previously succeeding python setup.py install now fails, try setting the MAX_JOBS environment variable.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="sh"> python setup.py install </pre></sub></td> <td><sub><pre lang="sh"> MAX_JOBS=2 python setup.py install </pre></sub></td> </tr> </table> </p>
Method 2: Switch back to the old distutils backend inside your setup.py
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> cmdclass={'clean': clean, 'build_ext': BuildExtension}, </pre></sub></td> <td><sub><pre lang="python"> cmdclass={'clean': clean, 'build_ext': BuildExtension.with_options(use_ninja=False)}, </pre></sub></td> </tr> </table> </p>
torch.optim.Adam, torch.optim.SGD changed to not modify gradients in-place (#30257)In previous versions of PyTorch, the Adam and SGD optimizers modified gradients (e.g. param.grad) in-place via in-place addition of params.grad += weight_decay * param. To make this consistent with the behavior of other optimizers and to prevent surprises about the behavior, we’ve changed them to stop modifying gradients in-place.
This should not have an effect on most PyTorch programs unless they relied on this behavior. The easiest way to replicate the old behavior is to create a custom optimizer that implements it.
torch.masked_select now always returns a 1D tensor (#29923)The behavior of torch.masked_select when both "self" and "mask" are 0-dimensional was changed. In previous versions of PyTorch, this would return a 0-dimensional tensor. Now, we return a 1-dimensional tensor to be consistent with other input sizes and our documentation.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.masked_select(torch.tensor(0), torch.tensor(True)) tensor(0) </pre></sub></td> <td><sub><pre lang="python"> torch.masked_select(torch.tensor(0), torch.tensor(True)) tensor([0]) </pre></sub></td> </tr> </table> </p>
torch.index_select on a 0-d tensor now returns a 0-d tensor. (#30790)In previous versions of PyTorch, the output of torch.index_select on a 0D input tensor produced a 1D tensor. This was inconsistent with our documentation on it, which stated "The returned tensor has the same number of dimensions as the original tensor (input)." Now, we return a 0D tensor.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.index_select(torch.tensor(5), 0, torch.tensor([0])) tensor([5]) </pre></sub></td> <td><sub><pre lang="python"> torch.index_select(torch.tensor(5), 0, torch.tensor([0])) tensor(5) </pre></sub></td> </tr> </table> </p>
nn.MultiLabelMarginLoss: 'none' reduction on 1D tensor now returns a 0D tensor (#30768)In previous versions of PyTorch, the output of nn.MultiLabelMarginLoss on 1D and 0D tensors incorrectly produced 1-D tensors. Now, those cases return a 0D tensor to be consistent with the 2-D tensor case.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
nn.MultiLabelMarginLoss(reduction='none')(torch.randn(3), torch.zeros(3, dtype=torch.long)) tensor([0.2959]) </pre></sub></td> <td><sub><pre lang="python"> nn.MultiLabelMarginLoss(reduction='none')(torch.randn(3), torch.zeros(3, dtype=torch.long)) tensor(0.2959) </pre></sub></td> </tr> </table> </p>
nn.MultiMarginLoss: ‘none' reduction on 1D target now returns a 1D tensor (#30826)In previous versions of PyTorch, the output of nn.MultiMarginLoss on a 1D target tensor produced a 0D output. We changed this to return a 1D target tensor to make it consistent with other input sizes which return an output that matches the target shape.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
nn.MultiMarginLoss(reduction='none')(torch.tensor([1.]), torch.tensor([0])) tensor(0.) </pre></sub></td> <td><sub><pre lang="python"> nn.MultiMarginLoss(reduction='none')(torch.tensor([1.]), torch.tensor([0])) tensor([0.]) </pre></sub></td> </tr> </table> </p>
Tensor.exponential_(lambda) no longer supports lambda < 0 (#32501)lambda, the rate parameter of the exponential distribution, mathematically should be greater than 0. We’ve disabled support lambda < 0 to be mathematically correct; most users will not have used a lambda less than zero.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> tensor = torch.empty(3).exponential_(-1.5) </pre></sub></td> <td><sub><pre lang="python">
</pre></sub></td>
</tr>
</table> </p>
nn.BCELoss, nn.functional.binary_cross_entropy no longer accept inputs with the same number of elements that are not broadcastable (#31365)Previously, we supported accepting inputs with the same number of elements. However, this behavior was deprecated and we removed it in 1.5.0. In order to replicate the old behavior, please explicitly reshape your input and target tensors to have the same shape.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
input = torch.rand(3, 3) target = torch.randn(9) torch.nn.functional.binary_cross_entropy(input, target) </pre></sub></td> <td><sub><pre lang="python"> input = torch.rand(3, 3) target = torch.randn(9) torch.nn.functional.binary_cross_entropy(input, target.reshape_as(input)) </pre></sub></td> </tr> </table> </p>
torch.normal out argument is now required to have the same size as the computed output (#32031)Previously, on CPU devices, torch.normal(mean, std, out=out) would resize out to the correct size. To be consistent with the CUDA implementation, we’ve changed it so that out must either already have the correct size, or be an empty tensor with size [0]. To work around this, please ensure that your out tensor has the correct size.
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.normal(torch.zeros(3), torch.ones(3), out=torch.randn(2)) tensor([ 0.0300, 0.7830, -1.3579]) </pre></sub></td> <td><sub><pre lang="python"> torch.normal(torch.zeros(3), torch.ones(3), out=torch.randn(2)) RuntimeError: inconsistent tensor, output size ([2]) is not the same as broadcasted mean and std size (3) </pre></sub></td> </tr> </table> </p>
Tensor.geometric_ no longer supports integral Tensors (#31878)Previously, on CPU devices, Tensor.geometric_ supported Tensors with integral dtype. Now, it only supports floating point. We removed support for this because it doesn’t make sense for geometric_ to operate on integral dtypes.
torch.floor_divide input positional argument name to self (#34552)Before PyTorch 1.5, torch.floor_divide took two positional arguments: torch.floor_divide(input, other). We’ve changed the name of the input argument to self; this will break code that called torch.floor_divide via keyword argument. For example:
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> torch.floor_divide(input=x, other=y) </pre></sub></td> <td><sub><pre lang="python">
torch.floor_divide(self=x, other=y) torch.floor_divide(x, y) </pre></sub></td> </tr> </table> </p>
RNNOutput, RNN / GRU forward method now returns std::tuple<Tensor, Tensor>, and LSTM forward method now returns std::tuple<Tensor, std::tuple<Tensor, Tensor>>, matching Python API.torch::optional<std::tuple<Tensor, Tensor>>, matching Python API.forward_with_packed_input method which accepts PackedSequence as input and optionally hidden state, matching the forward(PackedSequence, ...) variant in Python API.w_ih / w_hh / b_ih / b_hh. Instead, to access the weights and biases of the gates, users should do e.g. rnn->named_parameters()["weight_ih_l0"], which mirrors the Python API rnn.weight_ih_l0.RNNOptions
tanh() / relu() / activation are removed. Instead, nonlinearity is added which takes either torch::kTanh or torch::kReLUlayers is renamed to num_layerswith_bias is renamed to biasLSTMOptions
layers is renamed to num_layerswith_bias is renamed to biasGRUOptions
layers is renamed to num_layerswith_bias is renamed to biasUpsampleOptions and InterpolateFuncOptions:
size is changed from std::vector<int64_t> to c10::optional<std::vector<int64_t>>. If you want to pass a list of int64_t to this argument, you must pass it as std::vector<int64_t>.scale_factor is changed from std::vector<double> to c10::optional<std::vector<double>>. If you want to pass a list of double to this argument, you must pass it as std::vector<double>.torch::nn::functional::MultiLabelMarginLossFuncOptions is renamed to torch::nn::functional::MultilabelMarginLossFuncOptionstorch::nn::functional::MultiLabelSoftMarginLossFuncOptions is renamed to torch::nn::functional::MultilabelSoftMarginLossFuncOptionstorch::nn::BatchNorm is removed in favor of torch::nn::BatchNorm{1,2,3}dtorch::nn::FeatureDropout is removed in favor of torch::nn::Dropout{2,3}dtorch::nn::modules_ordered_dict is removed. User should do Sequential sequential({{"m1", MyModule(1)}, {"m2", MyModule(2)}}) instead.torch::nn::init::Nonlinearity is removed, in favor of these enums: torch::kLinear / torch::kConv1D / torch::kConv2D / torch::kConv3D / torch::kConvTranspose1D / torch::kConvTranspose2D / torch::kConvTranspose3D / torch::kSigmoid / torch::kTanh / torch::kReLU / torch::kLeakyReLUtorch::nn::init::FanMode is removed, in favor of these enums: torch::kFanIn / torch::kFanOutOptimizer::step now accepts closure function as optional input and returns a tensor, and LossClosureOptimizer is removed (#34790) (#34957). If you had a custom optimizer class defined as:struct MyOptimizer : Optimizer {
using Optimizer::Optimizer;
void step() override {...}
};
* you would need to update your optimizer class definition as follows:
struct MyOptimizer : Optimizer {
using Optimizer::Optimizer;
torch::Tensor step(LossClosure closure = nullptr) override {
...
// return `torch::Tensor()` if `closure` is nullptr
// (i.e. we are not computing the loss)
return torch::Tensor();
}
};
AdagradOptions, learning_rate is renamed to lr.Adagrad, sum_buffers and step_buffers are now removed, and parameter state should be accessed by calling the accessors on the parameter’s corresponding state object. For example:auto& param_state = static_cast<AdagradParamState&>(
*optimizer.state()[c10::guts::to_string(parameter.unsafeGetTensorImpl())]);
// Use the following to access parameter state:
//
// param_state.sum()
// param_state.step()
SGDOptions, learning_rate is renamed to lr.SGD, momentum_buffers is now removed, and parameter state should be accessed by calling the accessors on the parameter’s corresponding state object. For example:auto& param_state = static_cast<SGDParamState&>(
*optimizer.state()[c10::guts::to_string(parameter.unsafeGetTensorImpl())]);
// Use the following to access parameter state:
//
// param_state.momentum_buffer()
AdamOptions:
learning_rate is renamed to lrbeta1 and beta2 are replaced by a tuple betasAdam, step_buffers, exp_average_buffers, exp_average_sq_buffers and max_exp_average_sq_buffers are now removed, and parameter state should be accessed by calling the accessors on the parameter’s corresponding state object. For example:auto& param_state = static_cast<AdamParamState&>(
*optimizer.state()[c10::guts::to_string(parameter.unsafeGetTensorImpl())]);
// Use the following to access parameter state:
//
// param_state.step()
// param_state.exp_avg()
// param_state.exp_avg_sq()
// param_state.max_exp_avg_sq()
RMSpropOptions:
learning_rate is renamed to lrRMSprop, square_average_buffers, momentum_buffers and grad_average_buffers are now removed, and parameter state should be accessed by calling the accessors on the parameter’s corresponding state object. For example:auto& param_state = static_cast<RMSpropParamState&>(
*optimizer.state()[c10::guts::to_string(parameter.unsafeGetTensorImpl())]);
// Use the following to access parameter state:
//
// param_state.square_avg()
// param_state.momentum_buffer()
// param_state.grad_avg()
LBFGSOptions:
learning_rate is renamed to lrmax_eval‘s type is changed from int64_t to c10::optional<int64_t>tolerance_grads type is changed from float to doubletolerance_change type is changed from float to doublehistory_size type is changed from size_t to int64_tLBFGS, d, H_diag, prev_flat_grad, t, prev_loss, ro, al, old_dirs, old_stps, func_evals and state_n_iter are now removed, and parameter state should be accessed by calling the accessors on the parameter’s corresponding state object. For example:auto& param_state = static_cast<LBFGSParamState&>(
*optimizer.state()[c10::guts::to_string(parameter.unsafeGetTensorImpl())]);
// Use the following to access parameter state:
//
// param_state.d()
// param_state.H_diag()
// param_state.prev_flat_grad()
// param_state.t()
// param_state.prev_loss()
// param_state.ro()
// param_state.al()
// param_state.old_dirs()
// param_state.old_stps()
// param_state.func_evals()
// param_state.n_iter()
AutoGIL/AutoNoGIL in favor of pybind11::gil_scoped_* functions (#34301)If your code released or acquired the GIL via AutoNoGIL or AutoGIL, please change the invocations to pybind11::gil_scoped_release or pybind11::gil_scoped_release, respectively.
torch::tensor(floating-point values) will always produce tensor of default dtype, and torch::tensor(integer values) will always produce tensor of torch::kLong dtype, matching Python API behavior (#32367).torch::Tensor::base() is renamed to torch::Tensor::_base() , matching Python API. (#33316)The simple executor skips the number of fusion-related passes and analyses that are very time-consuming. Disabling these optimizations fixes pathologically long compilation times. The users that rely on GPU fusion to have their desired performance profile, should turn on the profiling executor. We provide C++ and python API to enable the profiling executor:
torch._C._jit_set_profiling_mode(True) before you call your model for the first time.#include <torch/csrc/jit/runtime/graph_executor.h> and set getProfilingMode() = true before you invoke your model for the first time.In eager mode quantization, one needs to manually insert quant and dequant stubs in a model to specify where activations are quantized. Having a qconfig_dict that specifies the quantization configuration for each module is not useful as one needs to manually modify the model with quant/dequant stubs. The new API makes it explicit that the model needs to be manually modified for quantization.
# previously qconfig_dict was an optional argument to prepare
def prepare(model, qconfig_dict=None, inplace=False):
# now replaced with
def prepare(model, inplace=False):
More specifically, callers must pass context_id to torch.distributed.autograd.backward() and torch.distributed.optim.step(). (#33711)
# Before
import torch.distributed.autograd as dist_autograd
import torch.distributed.rpc as rpc
from torch import optim
from torch.distributed.optim import DistributedOptimizer
with dist_autograd.context() as context_id:
# Forward pass.
rref1 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 3))
rref2 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 1))
loss = rref1.to_here() + rref2.to_here()
# Backward pass.
dist_autograd.backward([loss.sum()])
# Optimizer.
dist_optim = DistributedOptimizer(
optim.SGD,
[rref1, rref2],
lr=0.05,
)
# After
import torch.distributed.autograd as dist_autograd
import torch.distributed.rpc as rpc
from torch import optim
from torch.distributed.optim import DistributedOptimizer
with dist_autograd.context() as context_id:
# Forward pass.
rref1 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 3))
rref2 = rpc.remote("worker1", torch.add, args=(torch.ones(2), 1))
loss = rref1.to_here() + rref2.to_here()
# Backward pass.
dist_autograd.backward(context_id, [loss.sum()])
# Optimizer.
dist_optim = DistributedOptimizer(
optim.SGD,
[rref1, rref2],
lr=0.05,
)
dist_optim.step(context_id)
The motivation is to prevent potential invalid device errors when the number of devices on the sender and the receiver does not match. However applications, can always move CUDA tensors to CPU before sending (#33604).
<p align="center"> <table align="center"> <tr><th>Version 1.4.0</th><th>Version 1.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python"> import torch import torch.distributed.rpc as rpc rpc.init_rpc("worker0", rank=0, world_size=2) x = torch.zeros(2, device=0) ret = rpc.rpc_sync("worker1", torch.add, args=(x, 3)) rpc.shutdown() </pre></sub></td> <td><sub><pre lang="python"> import torch import torch.distributed.rpc as rpc rpc.init_rpc("worker0", rank=0, world_size=2) x = torch.zeros(2, device=0) ret = rpc.rpc_sync("worker1", torch.add, args=(x.cpu(), 3)) rpc.shutdown() </pre></sub></td> </tr> </table> </p>
__torch_function__ API Override Mechanism (#30730, #32194, #32799, #34240, #34303).We introduced __torch_function__, an API override mechanism for subclassing torch.Tensor in Python. This is useful for creating custom objects that implement the torch.* APIs. These currently support overriding most torch.*, and torch.nn.functional APIs; we’ve also planned future support for subclassing torch.Tensor (see tracking issue #22402).
torch.logical_and and torch.logical_or operations added (#30521).torch.square added (#30719).torch.bitwise_and added (#31104).torch.cummax, torch.cummin added (#32169, #32238, #32537, #33492).torch.floor_divide , Tensor.floor_divide added (#30493, #34552).torch.true_divide , Tensor.true_divide added, analogous to Python 's, and NumPy's (true) division (#34236, #34794)nn.functional.hardsigmoid added(#34545).torch.pca_lowrank, torch.svd_lowrank), torch.lobpcg for positive-defined generalized eigenvalue problem (#34721).distributions.von_mises added (#33418).distributions.mixture_same_family : Added support for mixture distributions (#22742, #33408).distributions.transforms.TanhTransform added(#19785).distributions.continuous_bernoulli added (#34619).isinf (#31099).at::Tensor::retain_grad API (#33349).c10d.Store using pybind11 trampoline class #30415.nn.RNN: Ensure MIOpen is called on same stream as operator (#30672)elementwise_kernel settings on ROCm (#32609).nn.BatchNorm{1,2,3}d: Use C10_WARP_SIZE to fix functionality on HIP vs CUDA for gradient computation (#33098).batch_norm (#32065).torch.pdist: improved precision by enabling double __shfl_down (#34103).nn.RNN: Check if weights need to be flattened (#34265).torch::nn::Sequential (#33027) (#33718)torch::nn::Sequential::push_back(AnyModule) methods public (#34208).Conv{1,2,3}d, padding_mode now accepts torch::kZeros / torch::kReflect / torch::kReplicate / torch::kCircular, matching Python API behavior. (#35023)F::interpolate and torch::nn::Upsample implementation to match Python API behavior (#35025) (#36274)std::vector<OptimizerParamGroup> as inputoptimizer.add_param_group(...) can be used to add parameter group to an existing optimizeroptimizer.state() should be used to access parameter stateat::Tensor::base() to _base(), matching Python API (#33316)distributions.independent: added explicit string representation (#33676).categorical.sample: Reduced memory overhead (#34900).distributions.MultivariateNormal: improved numeric stability and performance (#32092).In PyTorch 1.5, we have added support for 10 additional operators and also enhanced support for another set of 10+ existing operators. We have also added support for exporting large models (> 2GB) to ONNX. Additionally, we have made enhancements and optimizations to the export of ScriptModules and will continue to do that in the next release. We have also made improvements to the custom op export experience.
binary_test to benchmark binary ops (#31326).Tensor.copy_ operator (#31327).torch.diag (#32597).REGISTER_DISPATCH (#33682).FROM_SCHEMA for qadd, qmul, qclamp, qconcat (#33359).fake_quant_slice to TensorIterator (#33744).init_method (#30208).rpc_agent handlers with generic Future (#31224).rref.localValue() call (#31199).RpcAgent::getWorkerInfos() API to return all WorkInfos in the group (#30241).RRef.str() API to return a string representation of the RRef (#30609).get_metrics and get_debug_info to RPC agent (#30833).RpcBackendOptions Constructor API (#34081).RRefContext before graceful shutdown (#31893).rpc.remote for builtin operators (#34931).clearAndWaitForOutstandingRpcsAsync. (#32952).default_collate type hint added (#28935).Tensor.rsub, Tensor.rpow, Tensor.rtruediv, Tensor.map_ type hints were added (#30576).torch.optim: added more missing type hints (#31130).nn.functional.grid_sample, nn.functional.affine_grid: added missing align_corners annotation (#32492).torch.nn.Parameter constructor type hint was fixed (#32617).nn.MultiheadAttention, nn.Transformer: added type hints (#28396).torch.optim.LambdaLR constructor type hint was fixed (#33271).torch.optim: added missing default value for LRScheduler.step() (#32411).Tensor.type() more specific (#32353).torch.optim.optimizer.Optimizer type hints were fixed (#32900).optim.AdamW type hints were fixed (#34299).torch.utils.data.Sampler subclasses type hints were added (#33679).nn.Sequential, nn.ModuleList, nn.ParameterList, nn.ParameterDict type hints were fixed (#33686).Tensor.bfloat16() type hint was added (#33747).torch.bfloat16, nn.Module.training, Tensor.cuda, and 10s of other type hints added (#33762).torch.add type hint was fixed(#33935).Tensor.shape type hint was fixed (#34595).utils.data imports (#33543).Tensor.__radd__ type hint was fixed (#35231)autograd.detect_anomaly: added support for Sparse Tensors (#29803).autograd.detect_anomaly: Error messages now print the current Node name (#33875).autograd.profiler: added better error message when crashing while profiling multi-worker DataLoader (#31473).autograd.profiler Enable using torch.autograd.profiler.record_function as decorator (#30861).autograd.profiler Speed up export_chrome_trace by up to 4x (#30724).torch.autograd: added better error message when attempting to fork (#33885).torch.cuda.memory.caching_allocator_alloc, torch.cuda.memory.caching_allocator_delete exposed in Python API (#33860).torch.roll: added bool tensor support (#31194).torch.flip: added support for bool tensors (#31267).torch.equal: added support for bfloat16 CPU scalar types (#30817).torch.save, torch.load: added error message for minimum dill version support (#30985).torch.diagonal: added named tensor support(#30193).torch.linspace: added support for integral types on CPU (#32218).torch.eig: Added autograd support in the case where eigenvalues are real (#33090).torch.mvlgamma: improved error message (#32665).torch.no_grad, torch.enable_grad: added support for decorating generator functions (#31792).torch.narrow: added Tensor overload for start (#34317).Tensor.random_: enabled support for half on CPU (#34030).Tensor.grad: added warnings when accessing it if it won't be populated for known reasons (#30531).torch.cuda.comm.gather: improved error message (#27456).nn.functional.max_pool{1,2,3}d: added named tensor support (#31669).nn.Module.load_state_dict: Include the contents of the exception in error messages (#32693).nn.MultiheadAttention: add support for 3D attention mask (#31996).nn.MSELoss : Added performance warning for using CPU Half (#33021).nn.ModuleList, nn.ParameterDict, nn.ParameterDict: added more descriptive error messages when attempting to call these like Modules (#29991).nn.init.dirac_: Added groups option for compatibility with initializing group convolutions (#32825).torch.Tensor (#33615).hub_dir alongside TORCH_HOME env variable for storing hub models (#32844).Tensor.new is called on alternate layouts or dtypes (#31485).utils.checkpoint.checkpoint_sequential: Removed deprecated variadic arguments behavior (#25985).output_ratio for FractionalMaxPool{2,3}d module and fractional_max_pool{2,3}d functional should accept double as data type (#33304)AdaptiveAvgPool{2,3}d and AdaptiveMaxPool{2,3}d, output_size is changed to accept c10::nullopt in its elements, matching Python API behavior. (#35022)fractional_max_pool3d_with_indices implementation (#35024)namespace F = torch::nn::functional from torch/nn/modules/batchhnorm.h, so that people don't have to use F to alias torch::nn::functional if they don't want to (#30684)AutogradContext, get_dirty() is removed and get_and_bump_dirty() is added, and the latter always bumps the version counter of the returned tensors (#33068)using namespace torch::autograd from torch/csrc/api/include/torch/nn/modules/_functions.h (#34423)torch::tensor(floating-point values) will always produce tensor of default dtype, and torch::tensor(integer values) will always produce tensor of torch::kLong dtype, matching Python API behavior (#32367)torch::allclose to handle std::numeric_limits::lowest() for integral types (#32978)torch::empty_like to use merge_in to process TensorOptions (#33505)rank or world_size is specified in Process Group init_method URL (#32016).requires_grad for Parameter replica so it's not always set to True by default (#32356)allreduce results to input tensors (#32226)zero_grad is used in DataParallel (#33064)torch.stfttorch.lu,torch.lu_unpacktorch.cdisttorch.normtensor.tolist() compilation now supported, requires output type annotation (#33472)def foo(float_matrix, scalar_ten):
# type: (Tensor, Tensor) -> Tuple[List[List[float]], bool]
out1 : List[List[float]] = float_matrix.tolist()
out2 = torch.jit.annotate(bool, scalar_ten.tolist())
return out1, out2
torch.rand_like and other _like constructors no longer require additional arguments in TorchScriptnn.Module APIs added (#29495):
childrennamed_childrenmodulesnamed_modulesPackedSequence (#32955)index and type properties on Device (#32953)
device.indexdevice.typeTensor properties (#33906)
tensor.ndimtensor.Ttensor.nametensor.is_leaftorch.device to be usedlen on tuples containing different types #35768SELECTED_OP_LIST file path issue (#33942).gettimeofday on iOS (#30361).weight_norm export for dim=0 (#31015).copy_ with index as tensor input (#32801).rand_like as well (#33095).torch/onnx/symbolic_opset11.py (#31814).torch.mm export (#34794)aten::size for opset 11 (#35984)Note truncated.
One column per quarter.
torch::modules_ordered_dict is deprecated (28774).
The PyTorch v1.4.0 release is now available.
The release contains over 1,500 commits and a significant amount of effort in areas spanning existing areas like JIT, ONNX, Distributed, Performance and Eager Frontend Improvements and improvements to experimental areas like mobile and quantization. It also contains new experimental features including rpc-based model parallel distributed training and language bindings for the Java language (inference only).
PyTorch 1.4 is the last release that supports Python 2. For the C++ API, it is the last release that supports C++11: you should start migrating to Python 3 and building with C++14 to make the future transition from 1.4 to 1.5 easier.
Following the experimental release of PyTorch Mobile in the 1.3 release, PyTorch 1.4 adds additional mobile support including the ability to customize build scripts at a fine-grain level. This allows mobile developers to optimize library size by only including the operators used by their models and, in the process, reduce their on device footprint significantly. Initial results show that, for example, a customized MobileNetV2 is 40% to 50% smaller than the prebuilt PyTorch mobile library. Learn more about how to create your own custom builds, and please engage with the community on the PyTorch forums to provide any feedback you have.
With the scale of models, such as RoBERTa, continuing to increase into the billions of parameters, model parallel training has become ever more important to help researchers push the limits. This release provides a distributed RPC framework to support distributed model parallel training. It allows for running functions remotely and referencing remote objects without copying the real data around, and provides autograd and optimizer APIs to transparently run backwards and update parameters across RPC boundaries.
To learn more about the APIs and the design of this feature, see the links below:
For the full tutorials, see the links below:
As always, you can connect with community members and discuss more on the forums.
In addition to supporting Python and C++, this release adds experimental support for Java bindings. Based on the interface developed for Android in PyTorch Mobile, the new bindings allow you to invoke TorchScript models from any Java program. Note that the Java bindings are only available for Linux for this release, and for inference only. We expect support to expand in subsequent releases. See the code snippet below for how to use PyTorch within Java:
Learn more about how to use PyTorch from Java here, and see the full Javadocs API documentation here.
Pruning functionalities have been added to PyTorch in the nn.utils.prune module. This provides out-of-the-box support for common magnitude-based and random pruning techniques, both structured and unstructured, both layer-wise and global, and it also enables custom pruning from user-provided masks.
To prune a tensor, first select a pruning technique among those available in nn.utils.prune (or implement your own by subclassing BasePruningMethod).
from torch.nn.utils import prune
t = torch.rand(2, 5)
p = prune.L1Unstructured(amount=0.7)
pruned_tensor = p.prune(t)
To prune a module, select one of the pruning functions available in nn.utils.prune (or implement your own) and specify which module and which parameter within that module pruning should act on.
m = nn.Conv2d(3, 1, 2)
prune.ln_structured(module=m, name='weight', amount=5, n=2, dim=1)
Pruning reparametrizes the module by turning weight (in the example above) from a parameter to an attribute, and replacing it with a new parameter called weight_orig (i.e. appending "_orig" to the initial parameter name) that stores the unpruned version of the tensor. The pruning mask is stored as a buffer named weight_mask (i.e. appending "_mask" to the initial parameter name). Pruning is applied prior to each forward pass by recomputing weight through a multiplication with the updated mask using PyTorch's forward_pre_hooks.
Iterative pruning is seamlessly enabled by repeatedly calling pruning functions on the same parameter (this automatically handles the combination of successive masks by making use of a PruningContainer under the hood).
nn.utils.prune is easily extensible to support new pruning functions by subclassing the BasePruningMethod base class and implementing the compute_mask method with the instructions to compute the mask according to the logic of the new pruning technique.
torch.optim: It is no longer supported to use Scheduler.get_lr() to obtain the last computed learning rate. to get the last computed learning rate, call Scheduler.get_last_lr() instead. (26423)Learning rate schedulers are now “chainable,” as mentioned in the New Features section below. Scheduler.get_lr was sometimes used for monitoring purposes to obtain the current learning rate. But since Scheduler.get_lr is also used internally for computing new learning rates, this actually returns a value that is “one step ahead.” To get the last computed learning rate, use Scheduler.get_last_lr instead.
Note that optimizer.param_groups[0]['lr'] was in version 1.3.1 and remains in 1.4.0 a way of getting the current learning rate used in the optimizer.
Tensor.unfold on a 0-dimensional Tensor now properly returns a 1-dimensional Tensor.<p align="center"> <table align="center"> <tr><th>Version 1.3.1</th><th>Version 1.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor(5).unfold(dimension=0, size=1, step=1) tensor(5) </pre></sub></td> <td><sub><pre lang="python"> torch.tensor(5).unfold(dimension=0, size=1, step=1) tensor([5]) </pre></sub></td> </tr> </table> </p>
torch.symeig now return a 0-element eigenvectors tensor when eigenvectors=False (the default).<p align="center"> <table align="center"> <tr><th>Version 1.3.1</th><th>Version 1.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.symeig(torch.randn(3,3)).eigenvectors.shape torch.Size([3, 3]) </pre></sub></td> <td><sub><pre lang="python"> torch.symeig(torch.randn(3,3)).eigenvectors.shape torch.Size([0]) </pre></sub></td> </tr> </table> </p>
torch.jit.get_trace_graph private (it is now torch.jit._get_trace_graph) (29149)
traced_module.graph instead, like:@property on ScriptModules has been disabled (28395)
@property accesses were silently broken before, where we would evaluate the the get function once and store that as the attribute permanently. They properly error now; a workaround is to make your @property a regular method.torch::jit::RegisterOperators has been removed, use torch::RegisterOperators instead (28229). The usage and behavior should remain the same. torch.jit._register_* bindings from Python (e.g. torch.jit._register_attribute). These were private functions that were not intended to be used. (29499)This change simplifies our C++ API and matches previous changes we did at the python level that merged Tensors and Variables into a single type.
This change is unlikely to affect user code; the most likely exceptions are:
Argument-dependent lookup for torch::autograd may no longer work. This can break because Variable is now defined as an alias for Tensor (using Variable = Tensor;). In this case, you must explicitly qualify the calls to torch::autograd functions.
Because Variable and Tensor are now the same type, code which assumes that they are different types (e.g., for the purposes of templating, or std::enable_if checks) will not work until you delete the (now) redundant overload/specialization.
Some operators may trace differently. If this happens, please file a bug. The most likely situations are:
aten::empty)"there is no observable dependence" with the inputs)torch::nn::LinearOptions are renamed to match the Python API. (27382)in -> in_featuresout -> out_featureswith_bias -> biastorch::nn::Conv{1,2,3}dOptions are renamed to match the Python API. (28917) (29838)input_channels -> in_channelsoutput_channels -> out_channelswith_bias -> biastorch::nn::Conv{1,2,3}dOptions no longer has the transposed argument. (31005)transposed originally set to true in torch::nn::Conv{1,2,3}dOptions, they should migrate their code to use torch::nn::ConvTranspose{1,2,3}d layers instead.torch::nn layers and functionals are changed to have torch::KEnumNAME syntax. (27942, 26837)torch::Reduction::Mean. Now, torch::Reduction::Mean has been renamed to the shorter torch::kMean.torch::tensor constructor is improved to match Python API behavior. (28523) (29632) (29066)torch::tensor({{1}, {2}}) produced a tensor of sizes {2}. Now, it produces a tensor of sizes {2, 1}.torch::tensor(1.1) produced a 1-dim tensor. Now it produces a 0-dim tensor.torch::tensor with a double (e.g. torch::tensor(1.1)) or a (nested) braced-init-list of doubles (e.g. torch::tensor({{1.1, 2.2}}) produces a tensor with dtype torch::kDouble. Now it produces a tensor with dtype torch::get_default_dtype().torch::tensor with an integer type (e.g. torch::tensor(1)) or a (nested) braced-init-list of integer types (e.g. torch::tensor({{1, 2}})) produces a tensor with the same dtype. Now it always produces a tensor of dtype torch::kLong (aka. int64_t).TensorOptions without a dtype set to the torch::tensor constructor, it always produces a tensor of dtype torch::get_default_dtype(). Now it produces a tensor of different dtypes based on the dtype of the braced-init-list and the default dtype.std::initializer_list (NOT braced-init-list) to torch::tensor will no longer compile, and the user should pass the equivalent braced-init-list to torch::tensor instead. For example, write torch::tensor({1.1, 1.2}) instead of torch::tensor(std::initializer_list<double>({1.1, 1.2})).forward function now take Tensor instead of Tensor& as input. (28501)torch::nn layers affected: ELU / SELU / Hardtanh / LeakyReLU / ReLU / ReLU6 / RReLU / CELU
This change ensures that the above layers can be used in a torch::nn::Sequential module. If your C++ model uses any of the above layers, you must recompile your C++ code with the new libtorch binary.
Learning rate schedulers (torch.optim.lr_scheduler) now support “chaining.” This means that two schedulers can be defined and stepped one after the other to compound their effect, see example below. Previously, the schedulers would overwrite each other.
>>> import torch
>>> from torch.optim import SGD
>>> from torch.optim.lr_scheduler import ExponentialLR, StepLR
>>>
>>> model = [torch.nn.Parameter(torch.randn(2, 2, requires_grad=True))]
>>> optimizer = SGD(model, 0.1)
>>>
>>> scheduler1 = ExponentialLR(optimizer, gamma=0.9)
>>> scheduler2 = StepLR(optimizer, step_size=3, gamma=0.1)
>>>
>>> for epoch in range(4):
>>> print(epoch, scheduler2.get_last_lr()[0])
>>>
>>> optimizer.step()
>>> scheduler1.step()
>>> scheduler2.step()
0 0.1
1 0.09000000000000001
2 0.08100000000000002
3 0.00729000000000002
4 0.00656100000000002
allgather_coalesced API to ProcessGroup (28634,29059)abort API in ProcessGroupGloo Send/Recv Work (29928).--no_python flag to allow using a bash script wrapper in the launch command (29144).torch.distributed.rpc is a newly introduced package. It contains basic building blocks to run functions remotely in model training and inference, which will be useful for scenarios like distributed model parallel or implementing parameter server frameworks. More specifically, it contains four pillars: RPC, Remote Reference, Distributed Autograd, and Distributed Optimizer. Please refer to the documentation and the tutorial for more details.
rpc_sync and rpc_async for builtin operators and Python user functions (23228, 23569, 28392).remote and RRef for builtin operators and Python user functions (25169, 25499).remote and RRef with distributed autograd (28630, 28656).get_gradients() method to retrieve gradients from distributed autograd context. (28926).RRefs on local values and to-self remote calls (28948, 29634).shutdown to ProcessGroup agent (30330).script::Module: implement more of of the nn.Module API (28828)
attr() method to simplify attribute access.@staticmethod on ScriptModules (27163)OrderedDict is supported (26465)hasattr() (29332)clone_instance for ScriptModules (30168)torch.memory_format support to the TorchScript (28544)forward() is now allowed on container modules (28988)layout() in script (27100)ProcessGroupNCCL (27224).torch/csrc/cuda NCCL usage safe for NCCL 2.5 (29014).test_distributed for ROCm but only with NCCL backend (28814).rpc_sync and rpc_async APIs (26570).torch.distributed.{autograd,rpc} (27529).rpc_timeout to RpcAgent to make it reusable for other RpcAgent implementations. (29341).process_group_agent (29253).clean_shutdown=False. (29148).initializedContextIds_ map is cleaned up appropriately in distributed autograd engine. (29787).WorkerInfo (29958).RpcAgentOptions struct type to bundle arguments for different RpcAgents (29972).FutureMessages and throw exceptions in ProcessGroupAgent (29601).RRef leaks during shutdown (30217).torch.distrbuted.rpc (29276, 28030, 29971, 30160, 30050, 30069, 30179, 30218, 30240, 30243, 30259).PythonUDF{Call,Resp} (27530).std::shared_ptr for DistAutogradContext (29770).c10d::~NCCLUtils as noexcept (29118).std::stringstream cases for improved performance. (29351)torch::save() avoid zip compressing small header records. (28180)Node::print (27524)torch.addcdiv, torch.addcmul Added named tensor support (28975).torch.{ones,zeros,full,rand,randn}_like Added named tensor support (28981).torch.cdist Added named tensor support (29129).torch.equal Added named tensor support (29322).Tensor.align_to Fixed error message (27221).Tensor.align_to Make method-only. (27304).Tensor.align_to Accept partially named tensors (27308).torch.mean(Tensor, Dimname) Fixed autograd support (29199).Tensor.unflatten Fix when dim is a negative integer (#31208) (31432).In PyTorch 1.4, we have mainly focused on expanding the coverage for ONNX Opset 11, and enabling exporting torchvision models. Most of the torchvision models can be exported to ONNX (Opset 11, with fixed input size), including FasterRCNN, MaskRCNN, and KeypointRCNN. We have also enhanced export support for some tensor indexing scenarios, with more enhancements to come in the next release. In addition, 20+ new PyTorch operators are enabled in ONNX exporter.
torch.sort/torch.topk are supported in Opset 11 (25739)torch.size/torch.squeeze/torch.unsqueeze/torch.mm/torch.index_fill/torch.index_copy are supported in Opset 11 (27578)torch.masked_select/torch.masked_scatter are supported in Opset 11 (25949)torch.arange is supported in Opset 11 (26875)avg_pool, constant_pad_nd, reflection_pad, replication_pad Support enhanced in Opset 11 (28225)torch.hardtanh is supported in Opset 11 (30169)torch.remainder is enabled in exporter (24410)torch.unfold is enabled in exporter (24970)torch.slice/torch.select with negative index are enabled in exporter (25273, 26549)torch.ones/torch.ones_like/torch.zeros/torch.zeros_like/torch.full/torch.full_like with default dtype are enabled in exporter (27577)torch.unbind is enabled in exporter (27247)torch.nn.functional.interpolate export is enhanced (27179, 27566, 28560, 29489)torch.det is enabled in exporter (26958)torch.group_norm is enabled in exporter (27071)torch.meshgrid is enabled in exporter (26037)torch.randn/torch.randn_like are enabled in exporter (28470, 29354)torch.weight_norm enabled in exporter (28618)torch.scalar_tensor is enabled in exporter (28713)torch.logdet is enabled in exporter (29767)torch.batch_norm 2D with affine=False is enabled in exporter (29458)torch.bitshift is enabled in exporter (28210)Quantization updates correspond to a mix of bug-fixes and feature improvements, with feature improvements adding improved operator coverage and performance improvements. We have also made a lot of progress towards enabling graph mode quantization support.
torch.argmax/argmin Allow half type (28787).torch.cuda.memory_stats / memory_summary instrumentation for CUDA memory allocator (27361).torch.set_num_threads Allow calling multiple times with TBB (27190).torch.set_num_threads Allow calling multiple times in parallel native (27947).torch.logical_xor Allow non-bool tensors (27248).torch.promote_types Nicer error message. (27941).torch.batch_norm_elemt Add an out-variant (27621).torch.lerp Implement derivative with respect to weight (28219).torch.complex32 Add type promotion support (27929).torch.unique Support bool tensors (28374).torch.reshape Improve backward for viewable geometries (28901).torch.lu Generalized factorization (28608).torch.equal Add the intra-op parallelism (28810).torch.randint Accept generator=None (29748).torch.bfloat16 Enabled for cuda (27259).torch.multinomial Enable for torch.half (29266).nn.RNN Respect the current stream in cudnn (27026).nn.RNN Preserve nonlinearity attribute (28058).nn.Linear Support 0-batch size. (27211).nn.functional.binary_cross_entropy implement double backwards (26983).nn.AdaptiveAvgPool2d Add support for NHWC memory format (24396).nn.GELU Add GELU activation (28944).nn.LayerNorm Handle batch size of zero (28614).nn.BatchNorm Add NHWC support on cudnn (23861).nn.BatchNorm2d support torch.channels_last (28982).nn.BatchNorm2d Handle empty inputs (30035).nn.LayerNorm Enable the intra-op parallelism (28464).nn.utils.prune Add pruning functionality (24076).nn.Sequential Make iterable (28987).dtype.is_signed Ability to differentiate signed dtypes (29511).optim.lr_scheduler.MultiplicativeLR Add new multiplicative learning rate scheduler. (27254).cuda.comm.scatter, gather Add channel-last support (28077).at::parallel_for Choose number of OMP threads based on GRAIN_SIZE (26963).FileStore with concurrent accesses. (28812).nn.MultiheadAttention (26826).ProcessGroupAgent termination detection algorithm (26984).ProcessGroupAgent listener thread until contexts are initialized (28013).rpc_* / remote requests (29781).RRefContext singleton leaky, deal with module destruct order race. (30172).torch::nn::init::Nonlinearity and torch::nn::init::FanMode will be removed in 1.5.torch.arange dtypetoIValue dict iteration (26856)torch.kthvalue Fix CUDA shared memory out of bound access in findPattern (28989).
torch.save Fix source files not being saved (28965).
torch.load Fix OSError loading files larger than 2GB. (27125).
torch.linspace clearer error message for negative step sizes. (28274).
torch.histc Add range checks to avoid segfaults (27712).
torch.lu Fix thread local issue on cpu (28546).
torch.max_pool2d Limit tensor size to max CUDA grid size (28931).
torch.renorm Fix a memory leak in CUDA renorm. (29873).
torch.index_add Fix bug in atomicAdd on CUDA for some dtypes (29231).
torch.addmm Fix handling of empty tensors (28613).
nn.CTCLoss Fix incorrect gradient for large target sizes (27460).
nn.functional.ctc_loss Fix incorrect gradient on cudnn (27039).
nn.Embedding Incorrect gradient at padding_idx in cuda kernel. (27731).
nn.LayerNorm Fix an illegal memory access error (28196).
nn.Conv2d handle zero stride (28784).
nn.PoissonNLLLoss Fix incorrect result with full=True (28637).
nn.AvgPool2d fix an overflow for 2^31-1 sized inputs (30793).
nn.RNNBase Fix an issue with use of children of RNN third party device types (28562).
nn.Upsample Fix “invalid configuration argument” error (28927).
nn.Upsample Fix a CUDA launch config failure (29016).
optim.lr_scheduler.OneCycleLR Correctly handle div_factor parameter (28217).
PackedSequence.to Ensure all tensors are moved (27245).
EventList.total_average Fix a regression caused by missing iadd (27498).
Tensor.record_stream Ensure stream is recorded for shifted view tensors (27371).
torch.hub Handle branch names containing a slash. (27960).
Fix error handling in Magma kernels (29003).
Fix avx for c++14 (28207).
Fix illegal memory access thread safety issue in sparse CUDA (29426).
__cuda_array_interface__ Fix stride calculation (31450).
torch.optim: Scheduler.step(epoch) is now deprecated; use Scheduler.step() instead. (26432)For example:
>>> for epoch in range(10):
>>> optimizer.step()
>>> scheduler.step(epoch)
DeprecationWarning: The epoch parameter in `scheduler.step()` was not necessary and is being deprecated where possible. Please use `scheduler.step()` to step the scheduler. During the deprecation, if epoch is different from None, the closed form is used instead of the new chainable form, where available. Please open an issue if you are unable to replicate your use case: https://github.com/pytorch/pytorch/issues/new/choose.
warnings.warn(EPOCH_DEPRECATION_WARNING, DeprecationWarning)
Tensor::is_variable() has been deprecated. As noted in the Backwards Incompatible Changes section, the distinction between variable and non-variable has been eliminated, so this check is no longer meaningful. Generally, is_variable() will now return true except in some special circumstances (see 29653 for more details). (29653)torch::nn::modules_ordered_dict has been deprecated. It is generally no longer necessary and can just be removed. (28774)torch.jit.quantized API has been deprecated in favor of torch.quantization.quantize_dynamic (28766)A benchmark suite is available to easily measure the performance of operators with a range of input shapes. The generated benchmark data fully characterize the performance of operators in terms of execution time. For more details see README.md in the benchmarks/operator_benchmark directory.
torch.nn.functional.threshold, torch.nn.functional.layer_norm, torch.cdist Performance of threshold (CPU), layer norm (CUDA) and cdist operations was improved (27155,27634, 25799)torch.Tensor.fill_ Performance for half and bfloat16 types on CPU was improved (28397).torch.nn.MaxPool2d implementation for channels_last format was added (24872)tensor.numel devirtualized, improving performance (27294)>>> a = torch.tensor([[True, True], [False, True]])
<p align="center"> <table align="center"> <tr><th>Version 1.3.0</th><th>Version 1.3.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a = torch.tensor([[True, True], [False, True]])
a_transpose = a.t()
a_transpose == 0
</pre></sub></td>
<td><sub><pre lang="python">
a = torch.tensor([[True, True], [False, True]])
a_transpose = a.t()
a_transpose == 0 tensor([[False, True], [False, False]]) </pre></sub></td> </tr> </table> </p>
<p align="center"> <table align="center"> <tr><th>Version 1.3.0</th><th>Version 1.3.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
a = torch.ones(5, 2, dtype=torch.float) b = torch.zeros(5, dtype=torch.long) a[:, [1]] = b.unsqueeze(-1) a
</pre></sub></td>
<td><sub><pre lang="python">
a = torch.ones(5, 2, dtype=torch.float) b = torch.zeros(5, dtype=torch.long) a[:, [1]] = b.unsqueeze(-1) RuntimeError: expected dtype Float but got dtype Long </pre></sub></td> </tr> </table> </p>
x and y were of different dtypes. Mixed dtype operations of this form are currently disabled, as they were in version 1.2. (29078)<p align="center"> <table align="center"> <tr><th>Version 1.3.0</th><th>Version 1.3.1</th></tr> <tr valign="top"> <td><sub><pre lang="python">
x = torch.randn(2, 3) y = torch.randint(0, 10, (2, 3)) torch.where(x < 0, x, y) tensor(...)
</pre></sub></td>
<td><sub><pre lang="python">
x = torch.randn(2, 3) y = torch.randint(0, 10, (2, 3)) torch.where(x < 0, x, y) RuntimeError: expected scalar type Float but found Long </pre></sub></td> </tr> </table> </p>
torch.argmax: fix regression on CUDA that disabled support for torch.float16 inputs. (28915)Tensor.names. (28922)deepcopy for quantized tensors. (28612)nn.quantized.ReLU with inplace=True. (28710)torch.lgamma and torch.polygamma are now documented. (28964)Updated docs and added deprecation warnings to acknowledge a bool tensor (22261).
Previous versions of PyTorch supported a limited number of mixed dtype operations. These operations could result in loss of precision by, for example, truncating floating-point zero-dimensional tensors or Python numbers.
In Version 1.3, PyTorch supports NumPy-style type promotion (with slightly modified rules, see full documentation). These rules generally will retain precision and be less surprising to users.
<p align="center"> <table align="center"> <tr><th>Version 1.2</th><th>Version 1.3</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.tensor(1) + 2.5 tensor(3) torch.tensor([1]) + torch.tensor(2.5) tensor([3]) torch.tensor(True) + 5 tensor(True) </pre></sub></td> <td><sub><pre lang="python"> torch.tensor(1) + 2.5 tensor(3.5000) torch.tensor([1]) + torch.tensor(2.5) tensor([3.5000]) torch.tensor(True) + 5 tensor(6) </pre></sub></td> </tr> </table> </p>
<p align="center"> <table align="center"> <tr><th>Version 1.2</th><th>Version 1.3</th></tr> <tr valign="top"> <td><sub><pre lang="python">
int_tensor = torch.tensor(1) int_tensor.add_(1.5) tensor(2) bool_tensor = torch.tensor(True) bool_tensor.add_(5) tensor(True) </pre></sub></td> <td><sub><pre lang="python"> int_tensor = torch.tensor(1) int_tensor.add_(1.5) RuntimeError: result type Float cannot be cast to the desired output type Long bool_tensor = torch.tensor(True) bool_tensor.add_(5) RuntimeError: result type Long cannot be cast to the desired output type Bool </pre></sub></td> </tr> </table> </p>
These rules can be checked at runtime via torch.can_cast.
torch.flatten: 0-dimensional inputs now return a 1-dim tensor. (25406).<p align="center"> <table align="center"> <tr><th>Version 1.2</th><th>Version 1.3</th></tr> <tr valign="top"> <td><sub><pre lang="python">
torch.flatten(torch.tensor(0)) tensor(0) </pre></sub></td> <td><sub><pre lang="python"> torch.flatten(torch.tensor(0)) tensor([0]) </pre></sub></td> </tr> </table> </p>
nn.functional.affine_grid: when align_corners = True, changed the behavior of 2D affine transforms on 1D data and 3D affine transforms on 2D data (i.e., when one of the spatial dimensions has unit size).Previously, all grid points along a unit dimension were considered arbitrarily to be at -1, now they are considered to be at 0 (the center of the input image).
torch.gels: removed deprecated operator, use torch.lstsq instead. (26480).utils.data.DataLoader: made a number of Iterator attributes private (e.g. num_workers, pin_memory). (22273)Variable::backward will no longer implicitly create a gradient for non-1-element Variables. Previously, a gradient tensor of all 1s would be implicitly created . This behavior matches the Python API. (26150)auto x = torch::randn({5, 5}, torch::requires_grad());
auto y = x * x;
y.backward()
// ERROR: "grad can be implicitly created only for scalar outputs"
GRUOptions::bidirectional_) are now private, use the function variants (GRUOptions::bidirectional(...)) instead. (26419).In PyTorch 1.3, we are launching experimental support for mobile. Now you can run any TorchScript model directly without any conversion. Here are the full list of features in this release:
We decided not to create a new framework for mobile so that you can use the same APIs you are already familiar with to run the same TorchScript models on Android/iOS devices without any format conversion. This way you can have the shortest path from research ideas to production-ready mobile apps.
The tutorials, demo apps and download links for prebuilt libraries can be found at: https://pytorch.org/mobile/
This is an experimental release. We are working on other features like customized builds to make PyTorch smaller, faster and better for your specific use cases. Stay tuned and give us your feedback!
Named Tensors aim to make tensors easier to use by allowing users to associate explicit names with tensor dimensions. In most cases, operations that take dimension parameters will accept dimension names, avoiding the need to track dimensions by position. In addition, named tensors use names to automatically check that APIs are being used correctly at runtime, providing extra safety. Names can also be used to rearrange dimensions, for example, to support "broadcasting by name" rather than "broadcasting by position".
Create a named tensor by passing a names argument into most tensor factory function.
>>> tensor = torch.zeros(2, 3, names=('C', 'N'))
tensor([[0., 0., 0.],
[0., 0., 0.]], names=('C', 'N'))
Named tensors propagate names across operations.
>>> tensor.abs()
tensor([[0., 0., 0.],
[0., 0., 0.]], names=('C', 'N'))
Rearrange to a desired ordering by using align_to .
>>> tensor = tensor.align_to('N', 'C', 'H', 'W')
>>> tensor.names, tensor.shape
(('N', 'C', 'H', 'W'), torch.Size([3, 2, 1, 1]))
And more! Please see our documentation on named tensors.
PyTorch now supports quantization from the ground up, starting with support for quantized tensors. Convert a float tensor to a quantized tensor and back by:
x = torch.rand(10,1, dtype=torch.float32)
xq = torch.quantize_per_tensor(x, scale = 0.5, zero_point = 8, dtype=torch.quint8)
# xq is a quantized tensor with data represented as quint8
xdq = x.dequantize()
# convert back to floating point
We also support 8 bit quantized implementations of most common operators in CNNs, including:
We also support dynamic quantized operators, which take in floating point activations, but use quantized weights (in torch.nn.quantized.dynamic).
Quantization also requires support for methods to collect statistics from tensors and calculate quantization parameters (implementing interface torch.quantization.Observer). We support several methods to do so:
For quantization aware training, we support fake-quantization operators and modules to mimic quantization during training:
torch.fake_quantize_per_tensor_affine, torch.fake_quantize_per_channel_affinetorch.quantization.FakeQuantizeIn addition, we also support workflows in torch.quantization for:
All quantized operators are compatible with TorchScript.
For more details, see the documentation at: https://pytorch.org/docs/master/quantization.html
Arithmetic and comparison operations may now perform mixed-type operations that promote to a common dtype.
This below example was not allowed in version 1.2. In version 1.3, the same code returns a tensor with dtype=torch.float32.
>>> torch.tensor([1], dtype=torch.int) + torch.tensor([1], dtype=torch.float32)
See the full documentation for more details.
torch.result_type Provide function to determine result of mixed-type operations (26012).torch.can_cast Expose casting rules for type promotion (26805).torch.promote_types Expose promotion logic (26655).nn.functional.affine_grid / nn.functional.grid_sample: USING The Align_CORNER Default value is now deprecated, because it will be changed in 1.4 release.The align_corner parameter was added in this release; the behavior in the previous release was equivalent to setting the parameter to True. This is also the current default value but it will be changed to False from 1.4 release. Note that using the default will trigger a warning as demonstrated below; set the value explicitly to remove the warning.
>>> torch.nn.functional.affine_grid(torch.randn(1,2,3),
(1,3,2,2))
UserWarning: Default grid_sample and affine_grid behavior will be changed
to align_corners=False from 1.4.0.
See the documentation of grid_sample for details.
...
>>> torch.nn.functional.affine_grid(torch.randn(1,2,3),
(1,3,2,2),
align_corners=True)
# NO WARNING!
...
torch::Tensor::data<T>() in favor of torch::Tensor::data_ptr<T>() (24847, 24886).torch.utils.tensorboard supports 3D mesh and points plus hyperparameter logging. More details can be found in the documentation for SummaryWriter with add_mesh and add_hparams.
A simple example exercising both methods:
from torch.utils.tensorboard import SummaryWriter
vertices_tensor = torch.as_tensor([
[1, 1, 1],
[-1, -1, 1],
[1, -1, -1],
[-1, 1, -1],
], dtype=torch.float).unsqueeze(0)
colors_tensor = torch.as_tensor([
[255, 0, 0],
[0, 255, 0],
[0, 0, 255],
[255, 0, 255],
], dtype=torch.int).unsqueeze(0)
faces_tensor = torch.as_tensor([
[0, 2, 3],
[0, 3, 1],
[0, 1, 2],
[1, 3, 2],
], dtype=torch.int).unsqueeze(0)
with SummaryWriter() as w:
w.add_mesh('my_mesh', vertices=vertices_tensor, colors=colors_tensor, faces=faces_tensor)
for i in range(5):
w.add_hparams({'lr': 0.1*i, 'bsize': i},
{'hparam/accuracy': 10*i, 'hparam/loss': 10*i})
This release adds macOS support for torch.distributed with the Gloo backend. You can more easily switch from development (e.g. on macOS) to deployment (e.g. on Linux) without having to change a single line of code. The prebuilt binaries for macOS (stable and nightly) include support out of the box.
torch.distributed.all_reduce_coalesced Support allreduce of a list of same-device tensors (24949, 25470, 24876)torch.distributed.all_reduce Add bitwise reduction ops (BAND, BOR, BXOR) (26824)We now provide Libtorch binaries for building applications compatible with the C++11 ABI. The download links for libtorch binaries with C++11 ABI can be found in https://pytorch.org/ “QUICK START LOCALLY”.
not in support for TorchScript (23637).torch.jit.is_scripting() API (25955).x is not None unwrap the optional type of x (23949).+=) support to TorchScript (23639).grad and data attribute for tensor in TorchScript (23842).@ignore for TorchScript classes (23614).set_grad_enabled() into TorchScript (25350).in membership checks for lists (25796).tuple keyword (25474).__getitem__ to class types (25664).__setitem__ to class types (25750).min() and max() for lists to TorchScript (26351).We are on our way to better API parity between our Python and C++ frontends. Specifically, we made the following improvements:
torch::autograd::backward and torch::autograd::grad (24342)torch::autograd::Variable::register_hook (24393).torch::tensor (26210, 26890, 26756).
torch::tensor({{1, 2}, {3, 4}}) in C++ to construct the same tensor as torch.tensor([[1, 2], [3, 4]]) in Python. Some caveats are noted in this comment.torch::tensor (23337).torch::nn::Module::unregister_module function, for unregistering a submodule from a torch::nn::Module (26088).torch.distributed Detect and handle NCCL errors appropriately instead of blocking peers until timeout in ProcessGroupNCCL (25012, 25905)torch.distributed Make scatter/gather arguments optional (25575)torch.distributed.launch Add a -m flag to allow users to launch python modules (24910).torch.distributed Add function to get NCCL version for logging (26583).torch.distributed Add timeout parameter to connect function in TCPStore (26554).torch.distributed use timeout in connect function to prevent against infinite loop (26364).torch.nn.modules.batchnorm Allow SyncBatchNorm to run without DDP in inference mode (24815)torch.argmax/argmin Rewrite as TensorIterator reductions (26181).torch.erfinv Vectorize unary operator (26629).torch.sin/cos/tan Use intrinsics for trigonometric functions on CPU (26431).torch.qr Fix a regression (23591).nn.Conv Use Caffe2's implementation of grouped depthwise 3x3 convolutions (26556).nn.Conv Use parallel_for in DepthwiseConvKernel (26879).nn.Conv Change shape for conv and unary ops (25477).NoneType a subtype of Optional[T] (25361).In PyTorch 1.3, we have added support for exporting graphs with ONNX IR v4 semantics, and set it as default. We have achieved good initial coverage for ONNX Opset 11, which was released recently with ONNX 1.6. Further enhancement to Opset 11 coverage will follow in the next release. We have enabled export for about 20 new PyTorch operators. Also, we have focused on enabling the export for all models in torchvision. We have introduced some necessary groundwork for that in this release, e.g., accepting PyTorch models with inputs/outputs of Dict or String. We continue to work on torchvision models, such as FasterRCNN and MaskRCNN, to enable their export.
torch.det/logdet/slogdet Allowing batching (22909).torch.logical_not Add new operator (23839).torch.logical_xor Add new operator (23847).torch.symeig Improve the stability of gradient updates (23018).torch.eye Enable for bool and half (24148).torch.tril / triu Enable for bool and half (24163).torch.logical_not/xor support non-bool tensors. (23916, 23978).torch.index_select Implement indexing methods for sparse tensors (24937).torch.lu_solve Enable broadcasting of batch dimensions (24333).torch.cholesky Enable batches greater than 262140 (24438).torch.det Simplify generation of singular matrices to avoid numerical issue on PowerPC (25773).torch.erfinv In the CUDA implementation, use erfinv() for double to preserve accuracy (25337).torch.erfinv Add a float version of erfinv on CPU (26070).torch.cuda.stream Updates autograd engine to respect streams set in forward (8354).torch.backends.mkldnn.enabled Allow disabling MKLDNN at runtime (25459).torch.cholesky_solve Add derivative (26185).torch.cholesky_inverse Add derivative (26451).torch.polygamma Ensure that n is non-negative (26294).torch.pinverse Enable batching (26095).torch.digamma/trigamma Fix type mismatches on CUDA (25791).torch.where Enable for bool tensor on CUDA (26430).torch.load default encoding change to 'utf-8' (26421).torch.repeat_interleave Respect the current stream (26946).torch.bernoulli_ Implement for bool tensors (25076).torch.norm Fix nuclear norm with requires_grad=True (26303).torch.hub.download_url_to_file Make function public (26723).nn.modules.conv add padding_mode to repr (23996).nn.Transformer Extend to support BERT (gelu) (24181).nn.BatchNorm2d Add support for non-affine batch norm with float stats and half inputs (22750).nn.Parameter Fix type hints (25586).nn.CTCLoss Improve error message (26325).nn.Conv Allow batch size of 0 (26214).nn.LSTM/GRU enable double backward for non-cudnn (26660).optim.Adagrad Add epsilon argument (24980).optim.LBFGS Change default tolerance_grad to 1e-7 (25240).optim.lr_scheduler.OneCycleLR Add new 1cycle learning rate scheduler (25324).optimizer.step Fix type annotation (26930).bfloat16 Add support for sub, mul, and div on CPU (22851).bfloat16 Enabled comparison ops on CPU (24182).bfloat16 Enabled masked methods (24183).bfloat16 Enabled torch.mm and torch.mv (24224).bfloat16 Enable log_softmax and CrossEntropyLoss (24457).bfloat16 Enabled conv methods (26167).bfloat16 Enabled dtype on CUDA (26407).quasirandom.SobolEngine Use random seed if not specified (24884).utils.data.dataloader Add possible out of shared memory error message (25730).cuda.set_rng_state Add type hint (26200).~ and bitwise_not() when user tries to apply neg (-) on a bool tensor. (23621).autograd.grad Validate shapes of outputs (25349).iteration_ in SGD optimizer serialization (26906).torch::tensor Fix an ambiguous overload issues in constructor (26890).SummaryWriter.add_graph: Fix empty graph output in some cases (25599).SummaryWriter.make_video: Fix write_gif call to moviepy for newer lib (21218).step_size in LBFGS optimizer (25909).torch.jit.Attribute work when PYTORCH_ENABLED=0 (23851).nn.Module has not been initialized but you try to script it (24852).AliasAnalysisKind::PURE on MSVC (25375).NamedTuple types properly in Python (26443).optional (25965).is_optional check more robust (26312).torch.is_pinned pin_memory should not copy on already pinned tensors (23484).torch.cdist Fix incorrect gradients on CUDA non-batch tensors (22915).torch.from_numpy Fix failure on windows for int32 (25139).torch.tensor Fix memory leak creating a tensor from numpy (24267).torch.index Don't save self in index backward (25594).torch.bincount Fix int32 overflow on CUDA (25748).torch.bernoulli Fix the distribution sampler (26864).torch.pow Fix precision (25476).torch.cdist Fix gradient computation when first arg is 1xn (26254).torch.scatter_add_ Fix scatter CPU kernel when (input size, src size) > index size (25839).nn.ConvTranspose2d Fixed an error with float16 inputs and weights on CUDA. (23552).nn.CTCLoss Fix zero-length targets on CUDA (23298).nn.Conv2d Correct an overflow in an error message (25146).optim.Adam apply a small mathematical fix. (23737).dataloader Fix IndexError on shutdown if not all workers are started (23761).Tensor.repeat Fix crash on for 0 repeats (23766).torch.pin_memory only use one thread (25111).distributions.Uniform,HalfCauchy,Gamma Fix log_prob when value is a float (23017).intrusive_ptr.reset_() (24464).torch.hub: Fix SSL cert issue for hub in Python 2 (25042).Module.cuda Fix type hints (25018).batch_size=None or with namedtuple (26065).Vec256::abs() for floating point when applied on -0.0 (26422).torch.distributed Error phrasing in torch.distributed helper functions (25574)torch.distributions.negative_binomial clarified ambiguous doc string in NegativeBinomial (25923)trace_module to docs (24258).script and trace (24208).item() call in docs (25404).torch.record_stream Add documentation (24078).torch.fold Describe the relation between fold and unfold operations (24840).torch.argmax Fix incorrect doc (23775).torch.random add docs (23553).torch.empty_strided Add docs (23735).torch.bitwise_not Document for bool tensors (23800).torch.cdist Add documentation (25221).torch.where Update parameter names in doc (25554).torch.atan2 Clarify and correct the doc (26180).nn.functional.bilinear Added documentation (24951).nn.functional.upsample Fix align_corners doc (23707).nn.Transformer Fixed an error in the example (24837).optim.lr_scheduler.CosineAnnealingWarmRestarts Add documentation (25421).optim.SGD Updated with subscripts (23985).optim.RMSprop Highlighting in the doc that square root comes before adding epsilon (26735).autograd.detect_anomaly Add a warning (26615).…achieve feature parity with torch.uint8. See the Breaking Changes section for details about how this could affect existing programs. (21032, etc.)
We have just released PyTorch v1.2.0.
It has over 1,900 commits and contains a significant amount of effort in areas spanning JIT, ONNX, Distributed, as well as Performance and Eager Frontend Improvements.
Version 1.2 includes a new, easier-to-use API for converting nn.Modules into ScriptModules. A sample usage is:
class MyModule(torch.nn.Module):
...
# Construct an nn.Module instance
module = MyModule(args)
# Pass it to `torch.jit.script` to compile it into a ScriptModule.
my_torchscript_module = torch.jit.script(module)
torch.jit.script() will attempt to recursively compile the given nn.Module, including any submodules or methods called from forward(). See the migration guide for more info on what's changed and how to migrate.
In 1.2, TorchScript has significantly improved its support for Python language constructs and Python's standard library. Highlights include:
for..in loops, zip(), and enumerate().NamedTuples.math and string library support.See the detailed notes below for more information.
In PyTorch 1.2, working with Microsoft, we’ve added full support to export ONNX Opset versions 7(v1.2), 8(v1.3), 9(v1.4) and 10 (v1.5). We’ve have also enhanced the constant folding pass to support Opset 10, the latest available version of ONNX. Additionally, users now are able to register their own symbolic to export custom ops, and specify the dynamic dimensions of inputs during export. Here is a summary of the all of the major improvements:
Updated docs can be found here and also a refreshed tutorial using ONNXRuntime can be found here.
Read the documentation or simply type fromtorch.utils.tensorboardimport SummaryWriter to get started!
We include a standard nn.Transformer module, based on the paper “Attention is All You Need”. The nn.Transformer module relies entirely on an attention mechanism to draw global dependencies between input and output. The individual components of the nn.Transformer module are designed so they can be adopted independently. For example, the nn.TransformerEncoder can be used by itself, without the larger nn.Transformer. New APIs include:
nn.Transformernn.TransformerEncoder and nn.TransformerEncoderLayernn.TransformerDecoder and nn.TransformerDecoderLayerSee the Transformer Layers documentation for more info.
lt (<), le (<=), gt (>), ge (>=), eq (==), ne, (!=) ) return dtype has changed from torch.uint8 to torch.bool (21113)Version 1.1:
>>> torch.tensor([1, 2, 3]) < torch.tensor([3, 1, 2])
tensor([1, 0, 0], dtype=torch.uint8)
Version 1.2:
>>> torch.tensor([1, 2, 3]) < torch.tensor([3, 1, 2])
tensor([True, False, False])
For most programs, we don't expect that any changes will need to be made as a result of this change. There are a couple of possible exceptions listed below.
Mask Inversion
In prior versions of PyTorch, the idiomatic way to invert a mask was to call 1 - mask. This behavior is no longer supported; use the ~ or bitwise_not() operator instead.
Version 1.1:
>>> 1 - (torch.tensor([1, 2, 3]) < torch.tensor([3, 1, 2]))
tensor([0, 1, 1], dtype=torch.uint8)
Version 1.2:
>>> 1 - (torch.tensor([1, 2, 3]) < torch.tensor([3, 1, 2]))
RuntimeError: Subtraction, the `-` operator, with a bool tensor is not supported.
If you are trying to invert a mask, use the `~` or `bitwise_not()` operator instead.
>>> ~(torch.tensor([1, 2, 3]) < torch.tensor([3, 1, 2]))
tensor([False, True, True])
sum(Tensor) (python built-in) does not upcast dtype like torch.sum
Python's built-in sum returns results in the same dtype as the tensor itself, so it will not return the expected result if the value of the sum cannot be represented in the dtype of the tensor.
Version 1.1:
# value can be represented in result dtype
>>> sum(torch.tensor([1, 2, 3, 4, 5]) > 2)
tensor(3, dtype=torch.uint8)
# value can NOT be represented in result dtype
>>> sum(torch.ones((300,)) > 0)
tensor(44, dtype=torch.uint8)
# torch.sum properly upcasts result dtype
>>> torch.sum(torch.ones((300,)) > 0)
tensor(300)
Version 1.2:
# value cannot be represented in result dtype (now torch.bool)
>>> sum(torch.tensor([1, 2, 3, 4, 5]) > 2)
tensor(True)
# value cannot be represented in result dtype
>>> sum(torch.ones((300,)) > 0)
tensor(True)
# torch.sum properly upcasts result dtype
>>> torch.sum(torch.ones((300,)) > 0)
tensor(300)
TLDR: use torch.sum instead of the built-in sum. Note that the built-in sum() behavior will more closely resemble torch.sum in the next release.
Note also that masking via torch.uint8 Tensors is now deprecated, see the Deprecations section for more information.
__invert__ / ~: now calls torch.bitwise_not instead of 1 - tensor and is supported for all integral+Boolean dtypes instead of only torch.uint8. (22326)Version 1.1:
>>> ~torch.arange(8, dtype=torch.uint8)
tensor([ 1, 0, 255, 254, 253, 252, 251, 250], dtype=torch.uint8)
Version 1.2:
>>> ~torch.arange(8, dtype=torch.uint8)
tensor([255, 254, 253, 252, 251, 250, 249, 248], dtype=torch.uint8)
torch.tensor(bool) and torch.as_tensor(bool) now infer torch.bool dtype instead of torch.uint8. (19097)Version 1.1:
>>> torch.tensor([True, False])
tensor([1, 0], dtype=torch.uint8)
Version 1.2:
>>> torch.tensor([True, False])
tensor([ True, False])
nn.BatchNorm{1,2,3}D: gamma (weight) is now initialized to all 1s rather than randomly initialized from U(0, 1). (13774)Version 1.1:
>>> torch.nn.BatchNorm2d(5).weight
Parameter containing:
tensor([0.1635, 0.7512, 0.4130, 0.6875, 0.5496],
requires_grad=True)
Version 1.2:
>>> torch.nn.BatchNorm2d(5).weight
Parameter containing:
tensor([1., 1., 1., 1., 1.], requires_grad=True)
| Removed | Use Instead |
|---|---|
btrifact |
lu |
btrifact_with_info |
lu with get_infos=True |
btrisolve |
lu_solve |
btriunpack |
lu_unpack |
gesv |
solve |
pstrf |
cholesky |
potrf |
cholesky |
potri |
cholesky_inverse |
potrs |
cholesky_solve |
trtrs |
triangular_solve |
.data is no longer supported. (17072)>>> x = torch.randn(2,3)
>>> x.data = torch.sparse_coo_tensor((2, 3))
RuntimeError: Attempted to call `variable.set_data(tensor)`,
but `variable` and `tensor` have incompatible tensor type.
Version 1.1:
>>> i = torch.tensor([[0, 1]])
>>> v = torch.ones(2)
>>> s = torch.sparse_coo_tensor(i, v)
>>> i.resize_(1, 1)
>>> v.resize_(1)
>>> s.coalesce().indices().shape
torch.Size([1, 1])
>>> s.coalesce().values().shape
torch.Size([1])
Notice indices() and values() reflect the resized tensor shapes.
Version 1.2:
>>> i = torch.tensor([[0, 1]])
>>> v = torch.ones(2)
>>> s = torch.sparse_coo_tensor(i, v)
>>> i.resize_(1, 1)
>>> v.resize_(1)
>>> s.coalesce().indices().shape
torch.Size([1, 2])
>>> s.coalesce().values().shape
torch.Size([2])
Notice indices() and values() reflect the original tensor shapes.
.grad will no longer retain Python object identity. (17072)Version 1.1:
>>> m = torch.nn.Embedding(10, 3, sparse=True)
>>> m(torch.tensor([[1,2,4,5],[4,3,2,9]])).sum().backward()
>>> assert m.weight.grad.layout == torch.sparse_coo
>>> m_weight_grad_saved = m.weight.grad
# accumulate dense gradient into sparse .grad, change sparsity
>>> m.weight.sum().backward()
>>> assert m.weight.grad.layout == torch.strided
# m_weight_grad_saved still refers to the .grad of m's weight
# even though the sparsity has changed
>>> assert id(m_weight_grad_saved) == id (m.weight.grad)
Version 1.2:
>>> m = torch.nn.Embedding(10, 3, sparse=True)
>>> m(torch.tensor([[1,2,4,5],[4,3,2,9]])).sum().backward()
>>> assert m.weight.grad.layout == torch.sparse_coo
>>> m_weight_grad_saved = m.weight.grad
# accumulate dense gradient into sparse .grad, change sparsity
>>> m.weight.sum().backward()
>>> assert m.weight.grad.layout == torch.strided
# m_weight_grad_saved NO LONGER refers to the .grad of m's weight
>>> assert id(m_weight_grad_saved) == id (m.weight.grad)
AssertionError
nn.utils.convert_sync_batchnorm has been replaced with nn.SyncBatchNorm.convert_sync_batchnorm (18787)Example of new usage:
>>> # Network with nn.BatchNorm layer
>>> module = torch.nn.Sequential(
>>> torch.nn.Linear(20, 100),
>>> torch.nn.BatchNorm1d(100)
>>> ).cuda()
>>> # creating process group (optional)
>>> process_group = torch.distributed.new_group(process_ids)
>>> sync_bn_module = torch.nn.SyncBatchNorm.convert_sync_batchnorm(module, process_group)
torch.addcmul and torch.lerp operators enforce stronger shape requirements on the output tensor (out= keyword argument) and do not allow output tensor to be resized if it is also used as one of the inputs.Version 1.1:
>>> x=torch.zeros(1)
>>> torch.addcmul(x, x, torch.zeros(2,3), out=x)
tensor([[0., 0., 0.],
[0., 0., 0.]])
Version 1.2:
>>> x=torch.zeros(1)
>>> torch.addcmul(x, x, torch.zeros(2,3), out=x)
RuntimeError: output with shape [1] doesn't match the broadcast shape [2, 3]
If you run into this error, please ensure the out parameter is of the correct output shape (post-broadcasting).
PyTorch’s autograd system uses a version tracking mechanism to ensure that Tensors that are saved for backwards computations retain their correct values when the backward pass is computed (i.e. that they haven’t been updated in-place since they were saved). See In Place Correctness Checks in the docs for more information.
In PyTorch 1.2 we have enhanced the version tracking in a number of cases, which may flag issues that were not caught previously. There is now additional tracking through the Variable() constructor, the nn.Parameter() constructor, after setting .data, and via nn.Module._apply (internal API).
Track changes through Variable constructor:
>>> x = torch.ones(1, requires_grad=True)+1
>>> y = x*x
# do an in-place update through Variable constructor
>>> torch.autograd.Variable(x).add_(1)
>>> y.backward()
RuntimeError: one of the variables needed for gradient computation has been modified
by an inplace operation: [torch.FloatTensor [1]] is at version 1; expected version 0
instead.
Track changes on an nn.Parameter:
>>> x = torch.ones(1)
>>> p = torch.nn.Parameter(x)
>>> y = p * p
# do an in-place update on a saved Parameter
>>> x.add_(1)
>>> y.sum().backward()
RuntimeError: one of the variables needed for gradient computation has been modified
by an inplace operation: [torch.FloatTensor [1]] is at version 1; expected version 0
instead.
Track changes after setting .data:
>>> x = torch.zeros(1, requires_grad=True)+1
>>> y = x * x
>>> x.data = torch.zeros(1, requires_grad=True)+1
>>> x.add_(1)
>>> y.backward()
RuntimeError: one of the variables needed for gradient computation has been modified
by an inplace operation: [torch.FloatTensor [1]], which is output 0 of AddBackward0,
is at version 1; expected version 0 instead.
@ignoredtorch.jit.script now recursively compiles everything it finds in the original function, so if you had Python functions called from in your scripted function or module, you must now explicitly @ignore it. See the new API guide for more details.
Version 1.1
def my_unscriptable_python_fn():
# weird stuff
@torch.jit.script
def fn():
# This gets inserted as a Python call, and only errors on `save()`.
my_unscriptable_python_fn()
Version 1.2
@torch.jit.ignore # this needs to be added ...
def my_unscriptable_python_fn():
...
@torch.jit.script
def fn():
# ... or else recursive compilation will attempt to compile this call
my_unscriptable_python_fn()
NOTE: This is also a change to behavior of the @torch.jit.ignore decorator. In version 1.1, @ignore tells the compiler to omit compiling a function entirely, to mark Python functions that you know will not be called after export. In version 1.2 @ignore, tells the compiler to insert a call back to the Python interpreter instead of trying to compile the function.
To get the old behavior, use @torch.jit.ignore(drop_on_export=True) (@torch.jit.ignore with no arguments is equivalent to @torch.jit.ignore(drop_on_export=False)).
optimize for ScriptModules is now a context managerWhether optimization passes are run is now a thread-local flag. This better reflects how optimization actually happens in the JIT (i.e. it is decided at runtime, not compilation time).
Version 1.1
@torch.jit.script(optimize=False)
def fn(inputs):
...
fn(inputs)
Version 1.2
@torch.jit.script
def fn(inputs):
...
with @torch.jit.optimized_execution(False):
fn(inputs)
script::Module is now a reference typeTo better align with the PyTorch C++ API philosophy, script::Module and script::Method are now reference types. Our APIs have been updated to use script::Module instead of std::shared_ptr<script::Module>.
Version 1.1
using torch::jit::script::Module;
std::shared_ptr<Module> m = torch::jit::load("my_model.py");
m->forward(...);
Version 1.2
using torch::jit::script::Module;
Module m = torch::jit::load("my_model.py");
m.forward(...);
Version 1.1 API:
Tensor sum(IntArrayRef dim, bool keepdim=false) const;
Tensor sum(IntArrayRef dim, ScalarType dtype) const;
Version 1.2 API:
Tensor sum(IntArrayRef dim, bool keepdim=false,
c10::optional<ScalarType> dtype=c10::nullopt) const;
that is, to override dtype, keepdim must now be provided.
We have streamlined our conda and wheel binary distributions, so that it is easier than ever to install the version of PyTorch appropriate for your needs. The install instructions on https://pytorch.org/ have been updated, but if you have tooling to download and install PyTorch, here is a detailed description of the changes we made:
Wheels now have local version identifiers. Wheels that are for non-default CUDA configurations (the default CUDA version for this release is 10.0) now have local version identifiers like +cpu and +cu92. This means that, when installing, it is no longer necessary to specify a full wheel URL—just specify an appropriate version constraint like torch==1.2.0+cu92.
Version 1.1 (for Python 3.7 on Linux only):
pip install numpy
pip install https://download.pytorch.org/whl/cpu/torch-1.1.0-cp37-cp37m-linux_x86_64.whl
Version 1.2 (works for all versions of Python, and both Linux and Mac):
pip install torch==1.2.0+cpu -f https://download.pytorch.org/whl/torch_stable.html
CPU-only binaries on conda can be selected with the cpuonly feature. We’ve eliminated the pytorch-cpu conda package; instead, the cpu-only conda package can be enabled by installing the cpuonly metapackage. Similarly, there is no longer both a torchvision and torchvision-cpu package; the feature will ensure that the CPU version of torchvision is selected.
Version 1.1:
conda install -c pytorch pytorch-cpu
Version 1.2:
conda install -c pytorch pytorch cpuonly
Conda nightlies now live in the pytorch-nightly channel and no longer have “-nightly” in their name. We have added a new dedicated channel for nightlies called pytorch-nightly; all nightlies (pytorch, torchvision, torchaudio, etc.) will now be uploaded to this channel, but with the same name as their corresponding stable versions (unlike before, when we had a separate pytorch-nightly, torchvision-nightly, etc. packages.) This makes it more difficult to accidentally install a copy of the nightly and stable at the same time.
Version 1.1:
conda install -c pytorch pytorch-nightly
Version 1.2:
conda install -c pytorch-nightly pytorch
Wheel nightlies no longer have -nightly in their name. Similar to the changes we made in Conda, we no longer suffix wheel nightlies with “-nightly”, to make it harder to accidentally install a copy of nightly and stable at the same time.
Version 1.1:
pip install --pre torch_nightly -f https://download.pytorch.org/whl/nightly/torch_nightly.html
Version 1.2:
pip install --pre torch -f https://download.pytorch.org/whl/nightly/torch_nightly.html
torch.bool: added support for many operators (masking, comparison, arithmetic operators) to achieve feature parity with torch.uint8. See the Breaking Changes section for details about how this could affect existing programs. (21032, etc.)torch.sparse.HalfTensor: Added support for torch.float16 sparse Tensors on both CPU and CUDA. (19695)torch.bfloat16: Added basic creation and serialization support for Brain Floating Point Tensors. (21522, 21523, 21860, 22852)nn.Transformer: added implementation of Transformer from Attention is All You Need. (20170, 22588)nn.Embedding: support float16 embeddings on CUDA. (19695)nn.Flatten: added a Module that performs torch.flatten. (22245)nn.functional.gelu: Added support for Gaussian Error Linear Units. (20665, 21237)nn.Module hooks: add ability to replace input/output via forward_pre_hook and forward_hook. (22285)nn.Module: add requires_grad_() method for turning on/off requires_grad for Module parameters. (22576)Tensor.to_sparse: now supports autograd. (20458)Tensor.fill_diagonal_: operator to fill the main diagonal of a Tensor. (21892)torch.qr: supports autograd. (21274)torch.bitwise_not: add operator for boolean/integer types. Also have python ~ operator use this. (22283, 22320)torch.trapz: integrate using the trapezoid rule; equivalent to numpy.trapz. (21610)torch.var_mean / torch.std_mean: compute variance and mean at the same time.(18731)torch.utils.ThroughputBenchmark: benchmark utility for measuring the throughput of PyTorch operators. (20766).Logging: lightweight at-most-once logging to record operators that are used (c10::Logging). (20745)optim.AdamW: introduce AdamW optimizer from Decoupled Weight Decay Regularization. (21250)optim.LBFGS: added support for strong Wolfe line search. (8824)DistributedDataParallel: support CPU modules. (20236)DistributedDataParallel: support sparse tensors. (19146)DistributedDataParallel: support local gradient accumulation. (21736)IterableDataset: introduces a new type of Dataset designed for data read from a stream. (19228)SummaryWriter.flush: now supported. (20607)SummaryWriter.add_mesh: add support for 3D point clouds. (20413)List, Tuple, Dict, Tensor, String and you can also use zip(), enumerate(), and for...in. (21801, 22006, 21990, 21985)in membership checks. (21527)math support. (20979, 19707, 21151, 21131, 21129, 21130, 21512, 21126, 21127, 21128)NamedTuple. (21428)dict methods. (21979)sorted() keyword for lists and dicts. (23274)torch::List, torch::Dict and torch::Optional, supports dispatch (i.e. registering a different function for CPU and CUDA for the same operator).nn.GRU in script. (23266)pack_padded_sequence and pad_packed_sequence. (23249)torch._C._get_tracing_state in TorchScript. (23248)torch.as_tensor in TorchScript. (23247)Modules. (20708)all builtin. (20521)Final[T] annotated members to __constants__. (21603)save() to scripted Functions. (20386)Constant node. (22007)torch.jit.annotate(). (21390)ModuleList / Sequential. (21306)Module. (19905)Tensor.pin_memory(): only ask for context on current device. (22229)Tensor.view(): suggest using reshape() instead of contiguous() when the input is non-contiguous. (20968)Tensor.numpy(): throw TypeError instead of ValueError if the type isn’t supported. (21608)torch.norm: add support for p="nuc" with dim specified. (21022)torch.qr: support batching of input matrices. (20689)torch.qr: support some parameter akin to NumPy's mode option. (20689)torch.det / torch.logdet / torch.slogdet: added batching support. (22909)torch.cdist: support batching. (20934)torch.symeig: support batching. (21858)torch._dirichlet_grad: support CUDA. (21191)torch.randperm: support torch.float16. (22102)torch.Size is now pickle-able in Python2. (20952)torch.tensor / torch.as_tensor: infer device if input supports Numba’s __cuda_array_interface__. (20584)torch.isinf / torch.isfinite: throw TypeError instead of ValueError when a non-tensor is passed in. (20817)nn.MultiheadedAttention: add functional support. (20415)nn.MultiheadedAttention: added support for key/value to have different number of features. (21288)nn.MultiheadAttention: allow static key/values. (21288)nn.Conv{1,2,3}D: support torch.int64 dtype in forward. (20730, 22594)nn.AvgPool{1,2,3}D: support torch.int64 dtype in forward. (22433)nn.Module: make _save_to_state_dict overrideable. (21933)autograd: Checkpointing of modules inside large fanout networks no longer hits a recursion error. (22397)autograd: Track in-pace changes of Tensors through Module._apply (internal API). (21865)autograd.profiler: Add shape aggregation support. 20035)autograd.profiler: Profile custom c10 ops. (20175)DataLoader: support setting batch_size=0 to disable automatic batching (collation) in DataLoader for easier bulk loading. (19228)DataLoader: add multiprocessing_context parameter. (22990)DataLoader: added error detection for worker_init_fn. (20150)DataLoader: Retry on EINTR. (21723)torch.cuda.set_rng_state / torch.cuda.get_rng_state: accept string as device parameter. (23448)CUDA: add warning when using Turing GPUs and CUDA <= 9000. (21468)CUDA: warn on conditions that can trigger a cuBLAS 9.0 bug. (22034)CPU: Improve CPUAllocator OOM message. (20618)[memory_format]: added support for torch.empty, torch.empty_like, Tensor.contiguous(), Tensor.is_contiguous() to specify / check the order in which dimensions are laid out in memory. (20455, 20558)distributions.MultivariateNormal: fix precision matrix instability. (21366)distributions.transforms.SigmoidTransform: fix numerical instability. (19802)DistributedDataParallel: Support DDP forward/backward calls even if no module parameter is used. (19821)DistributedDataParallel: Only call into reducer if grad is enabled. (19897)DistributedDataParallel: Require finalize DDP backward only when there are indeed gradients computed, this allows application to completely discard DDP outputs and move on to the next iteration. (19901)DistributedDataParallel: Improve DDP backward reduction error messages. (20586)DistributedDataParallel: make DDP failure recoverable. (21591)DistributedDataParallel: Delay reduction of unused parameters until first autograd hook is called. (22219)c10d: support tensors shared across processes. (21449)c10d: ProcessGroupMPI Add device guard around MPI operations. (22446)utils.data.distributed.DistributedSampler: Make shuffling optional. (22479)Tensor.T: added numpy-like support for reversing dimensions. (20598)Tensor.ndim: NumPy equivalent property for the number of dimensions. (20565)Tensor.nonzero: added as_tuple argument (default False) that when True, will return a tuple of Tensors, which matches the behavior of numpy.nonzero. (20293)torch.dtype: support passing in NumPy dtypes as arguments. (21215)torch.normal: add size parameter when called with two floats. (20545)torch.where: add one-argument overload that is an alias for Numpy-like nonzero. (21986)axis instead of dim. (20451)start and step parameters for range in TorchScript. (20795)max_pool2d to symbolic derivatives. (19661)matmul memory usage for certain cases. (23433)__init__ function. (21880)ScriptModule buffer attributes can also cast device/type. (19700)ScriptModule.training an attribute instead of a parameter. (21078)strtod_c compatible with different gcc abi. (21293)nn::PoissonNLLLoss: Added support. (19316)nn::Module: added replace_module API to overwrite submodules in C++ Frontend. (22546)nn:Module::register_module / register_parameter / register_buffer: make public (23196)data::datasets::ChunkDataReader: fix include headers and a vector issue. (19485)data::datasets::ChunkDataset: add new get_batch method. (21797)data::datasets::ChunkDataset: add checkpoint support. (21889)data::datasets::ChunkDataset: add support for cross-chunk shuffling. (22347)data::datasets::ChunkDataset: add sorting policy. (23053)Add support for a number of operators on MKLDNN Tensors including:
Tensor.is_mkldnn: (22386)Tensor.transpose(): (21943)Tensor.zero_(): (20573)torch.empty: (21184)torch.mul: (20575)nn.AdaptiveAvgPool{1,2,3}D: (19818)nn.Sigmoid: (20820)nn.Softmax: (21516)nn.Module: support saving/loading MKLDNN modules. (20799)nn.MaxPool{1,2,3}D: support ceil_mode. (21310)Tensor.index_copy_: fix segfault by properly checking dimension is in range. (21617)Tensor.copy_: Fix a bug where non-blocking was not being respected. (20305)Tensor.clone: Fix an issue with MKLDNN tensors. (20943)torch.cat: Fix segfault with tensors that can't be indexed with 32-bit ints. (21530)torch.range / torch.linspace / torch.logspace: properly respect the current Stream. (21619)torch.lu: return the identity permutation instead of zeros when not using pivoting. (22242)torch.einsum: Fix an issue where the backward pass would potentially be skipped. (22111)torch.cosh: Fix an issue where torch.cos was instead calculated with torch.double dtype and vectorized instructions. (20797)torch.triu / torch.tril: handle strides correctly for in-place versions. (22730).torch.triu / torch.tril: Fix handling of batches > 65535 on CUDA. (21067)torch.inverse / torch.solve / torch.cholesky_solve / torch.triangular_solve: Fix batch sizes > 65535 on CUDA. (21689)torch.histc: return dtype is now the same as the input tensor on CUDA, matching CPU behavior. (20369)torch.histc: properly return 1-dim tensor on CPU with 0-dim input and 1 bin. (21497)torch.randperm: handle non-contiguous out parameter. (23043)torch.unique: Fix empty tensor handling when dim is passed as an argument. (19000)torch.min / torch.max: properly error on empty tensor inputs, as with CPU tensors. (19612).CUDA: fix launch parameters for reductions. (22827).torch.hub: fix an issue with find_module. (20782)autograd: Fix a number of custom autograd Function corner cases by inverting the relationship between PyFunction and THPFunction. (22983)autograd: give “Trying to backward through the graph a second time" error instead of internal assert when the buffers are a list of Tensors (with indexing). (21533)optim.lr_scheduler.CosineAnnealingLR: rename from CosineAnnealingLr. (23242)distributions.Binomial: Fix overflow of log_prob when logits is large. (20679)distributions.SigmoidTransform: Fix numerical issues that could result in inf / -inf return values. (20288)distributions.Categorical.sample: fix a view bug. (23328)CUDA: Give proper error message for bad cuda forks. (23322)pickle: Fix Unpickling error when loading multiple objects from a file. (20270)NCCL: Fix race condition. (23040)nn.Conv{1,2,3}D: fix memory leak on MKLDNN code path. (22392)nn.Conv{1,2,3}D: properly unpickle older pickled versions. (21687)nn.CTCLoss: fix backward on CUDA when 2d target tensor is larger than max_target_length. (20971)nn.CTCLoss: fix some numerical stability issues. (21392)nn.CTCLoss: disable buggy non-deterministic CudNN algorithm. (22977)nn.CTCLoss: fixed empty target handling. (21910, 23298)nn.SyncBatchNorm: fix syncing of running statistics when count size differs between GPUs. (22248)nn.SyncBatchNorm: retain requires_grad value when converting from nn.BatchNorm. (22569)nn.SyncBatchNorm: correctly handle process_group in convert_sync_batchnorm. (19240)nn.MultiheadedAttention: fix for torch.float16 dtype. (21658).nn.EmbeddingBag: fix NaN output when input is empty. (21400)nn.Dropout: fix python crash (with SIGFPE) when called on an empty cuda tensor. (20541)nn.MaxPool: fix output size calculation in some corner cases. (22304)nn.MaxPool: return valid indices if all entries are -inf. (23161)nn.Softmax: respect the current Stream. (22470)nn.LogSoftmax: fix numerical stability issues. (21672)nn.Module.load_state_dict: break ref cycle. (20397)nn.Module: fix loading in 32-bit environments. (20900)nn.utils.rnn.pack_padded_sequence: Fix segfault on empty tensors. (21461)nn.utils.spectral_norm: fix loading state_dict when strict=False. (22545)CudNN: Fix uninitialized PoolWindow on Windows. (22405)nn.parallel.DataParallel: fix error in no_grad mode. (21262)torch.distributed.all_gather: fix errors for views and aliases. (21490)c10d: fix collective communication errors on empty tensors. (20658)deepCopy also copies type information of lists, (23271)dictKeys and dictItems ops on typed dicts return typed lists. (23270)dict key type. (22231)builtin_function_or_method. (22935)__get_state__ to let a user know that ScriptModules can't be deep-copied at the moment.(20885)dropout derivative should respect the train flag. (20760)__constants__ for some nn modules. (21071)ScriptModule.__dir__(). (22426)CompilationUnit::define. (21886)Graph::toString. (21370)NameError with PYTORCH_JIT=0. (20120)pow() bug on overloads. (20824)1 - x in C++ would cause the size of 1 to get hardcoded. (20932)None constants. (23029)_flat_weights bug. (21107)WeakIValueEq. (21891)list() not making a copy. (22093)Module::forward method. (21398)a += b for lists do an in place add. (21896)floor/ceil return ints. (21124)__file__ for torch.ops. (21888)nn::RNN: Fix assertions in bidirectional RNN. (22850).nn::MaxPool / nn::AvgPool: expand incomplete kernel size, as in Python. (22073, 22075)Optim: Fix memory leak when weight_decay is applied to Adam, Adagrad, RMSProp. (23125)Optim::SGD: fix memory leak with weight_decay. (23007)torch::autograd::Scatter / torch::autograd::Gather: Fix nullptr bug. (20286)torch::nn::parallel::data_parallel: fix gradient computation error. (20910)torch.uint8 Tensors is now deprecated in favor of masking via torch.bool Tensors.See the Breaking Changes section for more details about torch.bool Tensors and comparison operators.
torch.masked_select, torch.masked_fill, torch.masked_scatter now expect torch.bool masks rather than torch.uint8.
>>> a = torch.tensor([1, 2, 3])
>>> b = torch.tensor([3, 1, 2])
>>> a.masked_select(tensor([0, 1, 1], dtype=torch.uint8))
UserWarning: masked_select received a mask with dtype torch.uint8,
this behavior is now deprecated, please use a mask with dtype torch.bool instead.
tensor([2, 3])
# instead use torch.bool
>>> a.masked_select(tensor([False, True, True]))
tensor([2, 3])
Comparison operators with out= parameters now expect torch.bool dtype rather than torch.uint8.
>>> a = torch.tensor([1, 2, 3])
>>> b = torch.tensor([3, 1, 2])
>>> res = torch.empty_like(a, dtype=torch.uint8)
>>> torch.gt(a, b, out=res)
UserWarning: torch.gt received 'out' parameter with dtype torch.uint8, this behavior
is now deprecated, please use 'out' parameter with dtype torch.bool instead.
tensor([0, 1, 1], dtype=torch.uint8)
# instead use torch.bool
>>> res = torch.empty_like(a, dtype=torch.bool)
>>> torch.gt(a, b, out=res)
tensor([False, True, True])
autograd.Function (Function without static forward method) is now deprecated>>> class MyLegacyFunction(Function):
>>> def forward(self, x):
>>> return x
>>>
>>> def backward(self, grad_output):
>>> return grad_output
>>>
>>> MyLegacyFunction()(torch.randn((3,), requires_grad=True)
UserWarning: Legacy autograd function with non-static forward method is deprecated
and will be removed in 1.3. Please use new-style autograd function
with static forward method.
# instead use new-style Autograd Function
>>> class MyFunction(Function):
>>> @staticmethod
>>> def forward(ctx, x):
>>> return x
>>>
>>> @staticmethod
>>> def backward(ctx, grad_output):
>>> return grad_output
>>>
>>> MyFunction.apply(torch.randn((3,), requires_grad=True)
See the torch.autograd.Function documentation for more details.
torch.gels: has been renamed to torch.lstsq; torch.gels will work for this release but is now deprecated. (23460)Tensor.copy_: increase broadcasting CUDA copy performance by 25%. (20685)torch.matmul: Optimize the case A.ndim <= 2 && B.ndim >= 3, shows up to 15x speed up. (20448)torch.bmm: Improve performance by up to 3x for small cases on CPU by applying TensorAccessor. (20266)torch.inverse: Move workspace query and allocation outside loop to improve performance by up to 5x. (20904)torch.topk: Optimize CPU perf using parallel and partial sort, up to 6x improvement. (22865)torch.cdist: Improve CPU perf by up to 10x for some cases. (20605)torch.normal: Move normal, normal_means, normal_stddevs, and normal_means_stddevs to ATen, increasing performance by up to 3x. (21287)torch.bernoulli: Speedup bernoulli_scalar_cuda_kernel with grid-stride loop, increasing performance by up to 2x. (21300)torch.coalesce: Use _sparse_coo_tensor_unsafe in coalesce for up to 10x speedup. (21214)torch.sinh / torch.cosh: Parallelize and vectorize on CPU. (21115)torch.lerp: Vectorize on CPU. (22038)torch.eye: Parallelize on CPU. (21077)torch.randperm: Parallelize initialization in randperm on CPU. (21529)nn.Softmax: Add persistent CUDA kernels that increase performance 2-10x on small inputs. (20827)nn.Embedding / nn.EmbeddingBag: Optimize CUDA kernel, increasing performance up to 2.7x. (22016)nn.Linear: optimize BERT model perf by using mkldnn inner product. (21851)nn.Conv{1,2,3}D: improve perf for depthwise convolutions in torch.float16 on Volta and Turing GPUs. (22302)nn.RNN: optimize on CPU by fusing matmul ops. (22512)nn.Upsample: a number of significant perf improvements on CUDA. (21879, 21694).nn.functional.layer_norm: optimize a fast path for layer_norm, increasing perf by up to 4x on CPU. (20345, 20883)mkldnn inner product for nn.Linear() to improve BERT perf. (21851).torch.bool: doc the Boolean tensor type. (21601)torch.as_strided: add docs. (22842)torch.empty_strided: add docs. (23740)torch.lerp: clarify broadcasting requirements. (23268)torch.enable_grad / torch.no_grad / torch.set_grad_enable: clarify interaction between these features. (23310)torch.autograd.grad_mode: Document that no_grad is thread local. (21755)torch.multiprocessing: Explain refcounting of CUDA tensors. (19904)torch.Tensor: Add a warning about memory usage. (20801)torch.utils.data.Dataloader: Document RNG state consumption. (22540)torch.optim.lr_scheduler.CyclicLR: Clarify base_momentum and max_momentum. (20880).tensor.to. (20977)nn.functional / nn.init: Break up NN in docs so they load faster. (21291)nn.functional.conv{1,2,3}d: Remove padding_mode. (20891)nn.functional.upsample / nn.functional.interpolate: add note about overshooting with mode=‘bicubic’. (23321)nn.init.zeros_ / nn.init.ones_: add documentation. (23145)nn.MultiheadAttention: Add documentation for add_bias_kv, add_zero_attn, and attn_mask. (20071)nn.MultiheadAttention: Fix documentation for attention mask shape. (20850)nn.Softmax: Fixed to specify dimension to prevent warning in 1.1.0. (20310)ninja to build instructions. (20079)In PyTorch 1.2, we have added the full support for ONNX Opset 7, 8, 9 and 10 in ONNX exporter, and we have also enhanced the constant folding pass to support Opset 10. The export of ScriptModule has better support. Additionally, users now are able to register their own symbolic to export custom ops, and specify the dynamic dimensions of inputs during export.
Dropout for Opset 10. (20710)Slice and Flip for Opset 10. (20533)Interpolate (Resize) for Opset 10. (21434)torch.arange . (22601)torch.masked_fill. (22521)torch.floor, torch.ceil, torch.log2 and prim::shape. (17895)torch._dim_arange. (20078)torch.randn_like. (20093)torch._standard_gamma. (20126)torch.topk. (21104)__ and__, __or__. (17894)torch.sign. (20470)torch.scatter. (18543)torch.rand. (20559)torch.gather. (21235)torch.cosine_similarity. (21884)torch.sum. (22240)torch.logsumexp. (22306)torch.layer_norm. (22265)torch.min and torch.max with dim. (19689)maxpool with dilations. (18721)RNN with batch_first=True. (19766)Upsample with dynamic input. (20116)torch.full with scalar parameters. (21931)Slice in constant folding optimization. (21811)torch.btrifact: the deprecated info argument has been removed. (14935).
Note: CUDA 8.0 is no longer supported
First-class and native support for visualization and model debugging with TensorBoard, a web application suite for inspecting and understanding training runs, tensors, and graphs. PyTorch now supports TensorBoard logging with a simple from torch.utils.tensorboard import SummaryWriter command. Histograms, embeddings, scalars, images, text, graphs, and more can be visualized across training runs. TensorBoard support is currently experimental. You can browse the docs here.
Attributes can be assigned on a ScriptModule by wrapping them with torch.jit.Attribute and specifying the type. Attributes are similar to parameters or buffers, but can be of any type. They will be serialized along with any paramters/buffers when you call torch.jit.save(), so they are a great way to store arbitrary state in your model. See the docs for more info.
Example:
class Foo(torch.jit.ScriptModule):
def __init__(self, a_dict):
super(Foo, self).__init__(False)
self.words = torch.jit.Attribute([], List[str])
self.some_dict = torch.jit.Attribute(a_dict, Dict[str, int])
@torch.jit.script_method
def forward(self, input: str) -> int:
self.words.append(input)
return self.some_dict[input]
TorchScript now has robust support for list and dictionary types. They behave much like Python lists and dictionaries, supporting most built-in methods, as well as simple comprehensions and for…in constructs.
For more complex stateful operations, TorchScript now supports annotating a class with @torch.jit.script. Classes used this way can be JIT-compiled and loaded in C++ like other TorchScript modules. See the docs for more info.
@torch.jit.script
class Pair:
def __init__(self, first, second)
self.first = first
self.second = second
def sum(self):
return self.first + self.second
nn.parallel.DistributedDataParallel: can now wrap multi-GPU modules, which enables use cases such as model parallel (tutorial) on one server and data parallel (tutorial) across servers.
(19271).
Tensor.set_: the device of a Tensor can no longer be changed via Tensor.set_. This would most commonly happen when setting up a Tensor with the default CUDA device and later swapping in a Storage on a different CUDA device. Instead, set up the Tensor on the correct device from the beginning. (18832).lr_scheduler.step(). (7889).torch.unique: changed the default value of sorted to True. (15379).Type no longer exist; use the functional or Tensor method equivalent. (17991).Backend constructor of TensorOptions no longer exists. (18137).ProcessGroup::getGroupRank has been removed. (19147).torch.tril_indices, torch.triu_indices: added operator with same behavior as NumPy. (14904, 15203).torch.combinations, torch.cartesian_prod: added new itertools-like operators. (9393).torch.repeat_interleave: new operator similar to numpy.repeat. (18395).torch.from_file: new operator similar to Storage.from_file, but returning a tensor. (18688).torch.unique_consecutive: new operator with semantics similar to std::unique in C++. (19060).torch.tril, torch.triu, torch.trtrs: now support batching. (15257, 18025).torch.gather: add support for sparse_grad option. (17182).torch.std, torch.max_values, torch.min_values, torch.logsumexp can now operate over multiple dimensions at once. (14535, 15892, 16475).torch.cdist: added operator equivalent to scipy.spatial.distance.cdist. (16168, 17173).torch.__config__.show(): reports detailed version of all libraries. (18579).nn.MultiheadedAttention: new module implementing MultiheadedAttention from Attention Is All You Need. (18334).nn.functional.interpolate: added support for bicubic. (9849).nn.SyncBatchNorm: support synchronous Batch Normalization. (14267).nn.Conv: added support for Circular Padding via mode='circular'. (17240).nn.EmbeddingBag: now supports trainable `per_sample_weights. (18799).nn.EmbeddingBag: add support for from_pretrained method, as in nn.Embedding. (15273).RNNs: automatically handle unsorted variable-length sequences via enforce_sorted. (15225).nn.Identity: new module for easier model surgery. (19249).torch.bool: added support for torch.bool dtype and Tensors with that dtype (1-byte storage). NumPy conversion is supported, but operations are currently limited. (16810).optim.lr_scheduler.CyclicLR: Support for Cyclical Learning Rate and Momentum. (18001).optim.lr_scheduler.CosineAnnealingWarmRestarts: new scheduler: Stochastic Gradient Descent with Warm Restarts). (17226).torch.distributions: now support multiple inheritance. (16772).quasirandom.SobolEngine: new sampler. (10505).nn.parallel.DistributedDataParallel: now supports modules with unused parameters (e.g. control flow, like adaptive softmax, etc). (18251, 18953).@ignore annotation, which statically tells the TorchScript compiler to ignore the Python function. (#16055)for...in loops on lists. (#16726)...) in Tensor indexing. (#17763)None in Tensor indexing. (#18615)if foo is not None. (#15587)to(), cpu(), and cuda() on ScriptModules. (#15340 , #15904)clear(), pop(), reverse(), copy() , extend(),index(), count(), insert(), remove() ).sort() on lists of specialized type (Tensors, int, float, bool). (#19572)index(), slice(), len())Tensor.to() in TorchScript. ( #15976 )Torch.tensor() in TorchScript. (#14913, #19445)torch.manual_seed() in TorchScript. (#19510)nn.LSTM in TorchScript. (#15744)nn.init in TorchScript. (#19640)hash() builtin. (#18258)min() and max() builtins for numerical types. (#15680)isinstance() builtin, which performs a static type check. (#15076)train() / eval() / is_training() to C++ ScriptModule API. (#16044)std::vector and std::unordered_map as arguments to custom operators. (#17587)nn.Sequential in ModuleList. (#16882)torch.qint8 dtype, torch.quantize_linear conversion function. (18230).MKLDNN tensors via Tensor.to_mkldnn(); operators are currently limited to ResNext101 operators. (17748).torch.min, torch.max, torch.median, torch.mode, torch.kthvalue, torch.symeig, torch.eig, torch.pstrf, torch.qr, torch.geqrf, torch.solve, torch.slogdet, torch.sort, torch.topk, torch.gels, torch.triangular_solve, torch.svd now return namedtuples describing their outputs. (16186, 16950, 17093, 17195, 15429).torch.empty (and other factory functions): now take a pin_memory kwarg; can now pin without going through torch.Storage interface.. (18455).torch.histc: Now supported on CUDA. (15842)torch.unique: Add return_counts. (18391, 18651).torch.logspace: add the ability to specify a base. (19542).torch.set_printoptions: added scientific notation support. (16876).torch.btrifact now handles tensors with greater than 3 dimensions. (14964).torch.kthvalue: now supported on CUDA. (17544).torch.abs: now supported on uint8 and int8 dtypes. (16893).torch.stack, torch.cat: now supported for CPU half tensors. (16389).torch.cross: added support for negative dimensions. (17582).torch.lerp: add support for weight as a Tensor. (17348).torch.transpose: Made consistent with NumPy: 1-d and 0-d arrays are accepted and returned as-is. (17462, 17535).torch.linspace, torch.logspace can now be used with steps=1 and start != end. (14748).torch.cholesky: changed the derivative from a triangular matrix to symmetric matrix. (19116).torch.lerp: Improved numerical stability. (18871).torch.logdet, torch.slogdet: improve numerical precision. (18449).Tensor.__contains__ is now supported. (17733).Tensor.fill_ and torch.zeros now support half on CPU. (17536).Tensor.resize_as_, Tensor.view: now supported on half CPU tensors. (18821).Tensor indexing: allow indexing via NumPy booleans. (14932).nn.EmbeddingBag: enable half precision dense backward. (19293).nn.Embedding: fix dense Embedding to work with double backwards. (9078).nn.MaxPool1d: Allow list and tuples to be passed as output_size. (16489).nn.CTCLoss: support zeroing infinite losses via zero_infinity argument. (16199).nn.Dropout: add support for enabling during eval. (17549).nn.MSELoss: add warning about unexpected broadcasting. (18349).nn.Module.load_state_dict: also return missing_keys and unexpected_keys. (18668).nn.parallel.data_parallel: Enforce devices match device_ids. (17129).torch.device: handle in more places that used to accept only device ordinals. (14929)dtype.int8 tensors can now be converted to NumPy arrays. (14710).nn.functional.gumbel_softmax: allow multidimensional input with dim argument. (13339).nn.functional.cosine_similarity: improved precision. (18250).torch.autograd: Don't keep unnecessary saved_inputs alive, increasing memory efficiency. (16583).torch.autograd.profiler: add Self (non-nested) CPU Time Total, CPU time total (19378).DataLoader: support accepting a custom memory pinning function. (16743).DataLoader: retry libshm on EINTR. (15964).DataLoader: fixed an issue with pin_memory and PackedSequence. (18079)data.utils.collate, data.utils.pin_memory: now preserve namedtuples. (16440)IndexError instead of RuntimeError on many indexing error cases. (17049, 17114).torch.float16 tensor on CPU. (17645).utils.checkpoint.checkpoint: support None as an argument to checkpoint function. (17969).torch.autograd: added more information for one of the variables needed for gradient computation has been modified by an inplace operation exception. (18523).cuda.synchronize: add a device argument. (19573).cuda.reset_max_memory_*: now supported. (15985).distributions.Independent: can now calculate KL Divergence. (17681).torch.distributed.new_group: now supports overriding default backend. (18595).torch.distributed.init_process_group: will now propagate timeout to underlying Store. (16571).nn.Module attributes to __constants__ when they are using in TorchScript. (#18164)torch.save(): Improve error message when you try to save a ScriptModule. (#15321)torch.jit.save(): Improve error message when trying to save a model with Python code. (#16850)__constants__. (#16724)__constants__. (#17167)nn::Module: added Python interop. (13481).autograd::profiler: is now supported. (16580)torch.argsort is now supported in C++. (17099).Tensor.isnan: now supported in C++. (15722).nn::Sequential. (17552).torch::data::transforms::Normalize: now supported in C++. (15891).std::vector<torch::Tensor>. (19677).torch.prod: correct erroneous calculation on large tensors. (15653).torch.mean (and other reductions): fix incorrect calculation on CUDA on large inputs. (16023).nn.Conv: correctly handle non-contiguous inputs on MKLDNN convolution codepath. (16300).Tensor.eq_: Fix erroneous calculation. (15475).torch.mean: Fix fp16 output calculation. (14878).nn.PoissonNLLLoss: Properly handle reduction=None. (17358).Tensor.round is now consistently half to even. (17443).Tensor.resize_: Fix some 0-element cases. (14874).Tensor.numpy: Fix conversion of torch.int8 dtype. (15194).Tensor.grad: correctly handle del. (16525).Tensor.clamp: correctly handle NaN on CUDA. (15479).Tensor.topk: properly set launch bounds on CUDA. (17296).Tensor.kthvalue: treat NaN as bigger than any number. (17824).Tensor.copy_: Properly synchronize on src and dst sreams. (16966).Tensor indexing: Fix incorrect dimension error message. (16495).Tensor.coalesce, Tensor.clone, Tensor.to_dense: fixed for sparse 0-dimensional tensors. (17379).torch.isinf: Don't error out on integral tensors. (15489).torch.argsort, torch.sort: Match NumPy by considering NaNs to be larger than any number. (15886).torch.geqrf, torch.ormqr: when an out parameter is specified, dispatch to the correct function. (16964).torch.cuda.get_device_name / torch.cuda.get_device_capability: Fix handling of optional. (17222).Tensor.tril_ / Tensor.triu_: properly reuse input memory. (17031).torch.arange: fix shape inconsistency between CPU and CUDA. (18462).torch.empty (and other size-based factory functions): properly enforce non-negative sizes. (17077).torch.load: support serializing / deserializing pathlib.Path object. (18562).nn.BatchNorm: correctly handle very large batches. (17047).nn.Softmax / nn.LogSoftmax: fix double backward for torch.half. (17330).nn.Softmax: handle empty inputs in backward. (17259).nn.NLLLoss: Fix crash when ignore_index is out-of-bounds on CPU. (17328).nn.Softmax, nn.LogSoftmax: handle 0-element inputs. (17651).nn.CTCLoss: correct error checking. (16269).nn.Conv: better report convolution size mismatch. (17436).torch.nn.functional.cosine_similarity: fix output sometimes returning result > 1.0. (18168).nn.parallel.data_parallel: Fix handling of buffers that require_grad. (13352).nn.parallel.data_parallel: would previously sometimes frees tensors before all pending operations finish. (18465).torch.distributed.broadcast: fixed repeated calls leading to OOM. (19219).torch.multiprocessing: fix serialization of integer nn.Parameters. (18639).torch.multiprocessing: Fix handling of distributions on CUDA. (16854).torch.nonzero: Fix for 0-dimensional tensors on CUDA. (17406).torch.slogdet: Fix sign requiring grad when input required grad. (16337).torch.cuda.Stream: Properly restore stream on destination device when switching devices. (17439).torch.cuda.Stream: Fixed synchronization issue when used with non-current device. (15689).torch.cuda.Stream: properly change device in stream context manager. (16128).DataLoader: fixed a hang when no data was read and the buffer size is smaller than the chunk size. (17409).DataLoader: _utils.collate.default_collate now converts bool lists to byte Tensors, not integer tensors.
(14669).DataLoader: ensure dataset is indexed by integers. (17649).torch.sparse.mm: Handle transposed dense tensors in backwards. (18737).torch.sparse.sum: Fix parsing of dim. (16517).torch.sparse.mm / torch.sparse.addmm: fix broadcasting and using uninitialized data. (16572).Tensor.to_sparse: Fix for 0-dimensional tensors. (17406).SparseTensor: fix add with non-contiguous values tensors. (18179).compare_exchange_weak in weak_intrusive_ptr. (16302).utils.model_zoo.load_url: Fix race condition. (16578).utils.data.RandomSampler: have len properly take into account num_samples. (15991).torch.distributions: Fix precision issue with expansion that prefers probs over logits. (18614).distributions.dirichlet.Dirichlet: fixed an underflow issue. (17488).distributions.binomial.Binomial.log_prob: fixed numerical stability issue. (15962).Caching Allocator: Free all blocks with outstanding events on OOM-retry. (19222).torch.dtype: fix pickling issue with Python 2. (18045).utils.data.DataLoader: Fix SIGCHLD checking. (19421).optim.Optimizer: Properly copy defaults. (19308).optim.lr_scheduler.CosineAnnealingLR: Fix division-by-zero error. (19180).optim.lr_scheduler.ReduceLROnPlateau: fix bug when the argument to step is reused outside the function.
(16697).cudNN: fix race condition with multiple threads calling into the same device. (15080).cudNN: Properly specify accumulation types. (16825).cuDNN: Fix incorrectly selecting slower algorithms in certain cases. (15881).cuFFT: Properly handle CUDA contexts. (19300)MKLDNN: fix thread safety. (17022).floordiv: Fix integer division and divide-by-zero semantics. (#15813).ord(): Fix handling of utf8 chars. (#19423).requires_grad analysis pass. (#18361).rnn.py. (#18198)._unique_state_dict could contain duplicate Tensors. (#18139).Stream and Event APIs. (15937).extra_cuda_cflags to C++ extensions on Windows. (18638).torch::nn::init::orthogonal_: match Python API. (18915).torch.btrifact: the deprecated info argument has been removed. (14935).torch.potrs has been deprecated, use torch.cholesky_solve instead. Note that upper defaults to False for torch.cholesky_solve, and True for torch.potrs. (15334).torch.pstrf is deprecated; use torch.cholesky instead. Note that upper defaults to False for torch.cholesky, and True for torch.pstrf. (17866).torch.potri is deprecated; use torch.cholesky_inverse instead. Note that upper defaults to False for torch.cholesky_inverse, and True for torch.potri. (19498).torch.btrifact_with_info has been deprecated; use torch.lu with get_infos=True instead.(18435).torch.btrifact has been deprecated; use the new name torch.lu instead. (18435).torch.gesv is deprecated; use the new name `torch.solve instead. (18060).torch.trtrs has been deprecated; use the new name torch.triangular_solve instead. (18213).torch. btriunpack has been deprecated; use the new name torch.lu_unpack instead. (18529).torch.btrisolve has been deprecated; use the new name torch.lu_solve instead. (18726).IntList has been deprecated, use IntArrayRef instead, as it better describes the type and ownership semantics in C++. (16751).Type parameters, e.g. AT_DISPATCH_ALL_TYPES(tensor.type(), ..., are now deprecated; use ScalarType instead, e.g. AT_DISPATCH_ALL_TYPES(tensor.scalar_type(), .... (17527, 17996).variable_tensor_functions have been removed. (15003).nn.BatchNorm CPU inference speed increased up to ~19x.(19152).nn.AdaptiveAvgPool: speed up common-case of size=1 output by ~30x. (17011).nn.EmbeddingBag CPU performance increased by ~4x. (19329).Tensor.copy_: sped up larger tensor copy ~2-3x, small regression in small tensor copy. (18618).torch.nonzero: is now ~2x faster than numpy on CPU. (15190)reduction functions: Speed up some large Tensor cases by 50-80%. (17428).batch_norm fusion for inference. (#15146)layer_norm fusion for inference. (#18266)torch.abs, torch.frac, torch.repiprocal, torch.neg have been vectorized and parallelized (19041).torch.bmm: CPU performance increased by 2x. (19338).torch.sort: CUDA performance increased by ~2x. (19379).torch.cat on CPU is now ~4x faster in the case where inputs are contiguous and dim != 0. (17032).torch.multinomial fixed a 2x performance regression. (17121).torch.empty (and another factory functions): reduce overhead by 20-40%. (17565).torch.linspace has been parallelized on CPU. (15320).torch.logspace has been parallelized on CPU. (15438).torch.range has been parallelized on CPU. (15484).torch.arange has been parallelized on CPU. (15667).torch.load: avoid unnecessary CPU-to-CUDA copy. (17297).reduction functions: improve efficiency on CUDA. (16224, 17040).sparse/dense matrix multiply: improve speed by ~5x. (16905).distributions.MultivariateNormal: sped up. (17294).aten::_convolution now participates in shape analysis. (#16837]randlike. (#14740)adaptive_avg_pool2d. (#15459)erf and erfc. (#15139)layernorm. (#17702)tanh. (#17816)matmul/dropout. (#17523)Tensor.scatter_: add documentation about value parameter. (17467).Tensor.unfold: correctly document dimension parameter, not dim. (19020).Tensor.is_floating_point() is now documented. (15704).torch.cholesky: Fix broken upper example in documentation. (15215).torch.gesv: document out parameter. (15649).torch.mul: better explain elementwise multiplication. (15664).torch.eig, torch.symeig: better explain backwards limitations. (15929).torch.ormqr: fixed output specification. (15694).torch.from_numpy: replaced usage with torch.as_tensor in documentation. (16587).torch.mvlgamma: Fix the constant in the docs. (17045).torch.mode: more precisely describe what is returned. (17069).torch.upsample: documentation now matches torch.interpolate. (17134)torch.arange: correct dtype documentation. (18604)torch.cumprod: document out parameter. (19340).torch.nonzero: document indices being returned lexicographically. (19539).torch.nn.functional.interpolate: better explain aligned_corners parameter. (14806).torch.nn.functional.pad: documentation has been made consistent with other functional ops. (15984).nn.functional.grid_sample: clarify behavior of padding. (19754).nn.TripletMarginLoss: correct type of swap parameter. (18115).nn.CrossEntropyLoss: clarify ignore_index documentation. (18117).nn.CrossEntropyLoss: the input format is more clearly explained. (15990).nn.CTCLoss: Clarify a number of ambiguities. (18415).nn.BCEWithLogitsLoss: add better explanation. (19212).nn.BCEWithLogitsLoss: better explain positive samples. (17258).nn.ModuleList / nn.ParameterList: update documentation. (17731).nn.Module.load_state_dict: correct semantics of strict. (17618)nn.parallel.DataParallel: more accurately specify how different argument types are handled. (15993).nn.parallel.DistributedDataParallel: Clarified batch size requirements. (16010).torch.distributed: Document mixed-precision training. (15440).torch.multiprocessing: Include example multiprocessing code. (16345).torch.autograd: Better explain computing Jacobian-vector product. (15197).torch.cuda.get_rng_state, torch.cuda.set_rng_state: document taking a device object. (14324).torch.device: Fix example of passing device to tensor factory. (16839).DataLoader: update documentation to describe how workers are managed. (18091).reduction arguments to use non-deprecated format. (17300).mark_non_differentiable: document correct semantics. (17891).activations attribute in nn.RNN ONNX export. (19368).There are no breaking changes in this release.
Note: our conda install commands have slightly changed. Version specifiers such as cuda100 in conda install pytorch cuda100 -c pytorch have changed to conda install pytorch cudatoolkit=10.0 -c pytorch
There are no breaking changes in this release.
torch.save and torch.load would initialize the CUDA context on GPU 0 if it hadn't been initialized already, even if the serialized tensors are only on GPU 1.1 ^^ x, where x is a PyTorch scalar (#16687)Some parts of the API may undergo breaking changes during this time.
The JIT is a set of compiler tools for bridging the gap between research in PyTorch
and production. It allows for the creation of models that can run without a dependency on the Python interpreter and which can be optimized more aggressively. Using program annotations existing models can be transformed into Torch Script, a subset of Python that PyTorch can run directly. Model code is still valid Python code and can be debugged with the standard Python toolchain. PyTorch 1.0 provides two ways in which you can make your existing code compatible with the JIT, using torch.jit.trace or torch.jit.script. Once annotated, Torch Script code can be aggressively optimized and it can be serialized for later use in our new C++ API, which doesn't depend on Python at all.
# Write in Python, run anywhere!
@torch.jit.script
def RNN(x, h, W_h, U_h, b_h):
y = []
for t in range(x.size(0)):
h = torch.tanh(x[t] @ W_h + h @ U_h + b_h)
y += [h]
return torch.stack(y), h
As an example, see a tutorial on deploying a seq2seq model, loading an exported model from C++, or browse the docs.
The torch.distributed package and torch.nn.parallel.DistributedDataParallel module are backed by a brand new re-designed distributed library. The main highlights of the new library are:
torch.distributed is performance driven and operates entirely asynchronously for all backends: Gloo, NCCL, and MPI.The C++ frontend is a pure C++ interface to the PyTorch backend that follows the API and architecture of the established Python frontend. It is intended to enable research in high performance, low latency and bare metal C++ applications. It provides equivalents to torch.nn, torch.optim, torch.data and other components of the Python frontend. Here is a minimal side-by-side comparison of the two language frontends:
<p align="center"> <table align="center"> <tr><th>Python</th><th>C++</th></tr> <tr valign="top"> <td><sub><pre lang="python"> import torch <br> model = torch.nn.Linear(5, 1) optimizer = torch.optim.SGD(model.parameters(), lr=0.1) prediction = model.forward(torch.randn(3, 5)) loss = torch.nn.functional.mse_loss(prediction, torch.ones(3, 1)) loss.backward() optimizer.step() </pre></sub></td> <td><sub><pre lang="cpp"> #include <torch/torch.h> <br> torch::nn::Linear model(5, 1); torch::optim::SGD optimizer(model->parameters(), /lr=/0.1); torch::Tensor prediction = model->forward(torch::randn({3, 5})); auto loss = torch::mse_loss(prediction, torch::ones({3, 1})); loss.backward(); optimizer.step(); </pre></sub></td> </tr> </table> </p>
We are releasing the C++ frontend marked as "API Unstable" as part of PyTorch 1.0. This means it is ready to be used for your research application, but still has some open construction sites that will stabilize over the next couple of releases. Some parts of the API may undergo breaking changes during this time.
See https://pytorch.org/cppdocs for detailed documentation on the greater PyTorch C++ API as well as the C++ frontend.
Torch Hub is a pre-trained model repository designed to facilitate research reproducibility.
Torch Hub supports publishing pre-trained models (model definitions and pre-trained weights) to a github repository using a simple hubconf.py file; see hubconf for resnet models in pytorch/vision as an example. Once published, users can load the pre-trained models using the torch.hub.load API.
For more details, see the torch.hub documentation. Expect a more-detailed blog post introducing Torch Hub in the near future!
# Previously: all 0-element tensors are collapsed to shape (0,)
>>> torch.nonzero(torch.zeros(2, 3))
tensor([], dtype=torch.int64)
# Now, proper shape is returned
>>> torch.nonzero(torch.zeros(2, 3))
tensor([], size=(0, 2), dtype=torch.int64)
*) between torch.Tensors and NumPy arrays will now favor dispatching to the torch variant. This may result in different return types. (#9651).numpy conversion no longer implicitly moves a tensor to CPU. Therefore, you may have to explicitly move a CUDA tensor to CPU (tensor.to('cpu')) before an implicit conversion. (#10553).Tensor argument now returns a detached Tensor (i.e. a Tensor where grad_fn is None). This more closely aligns with the intent of the function, which is to return a Tensor with copied data and no history. (#11061,
#11815).(N,) instead of (N, C) to match the behavior of torch.nn.MultiMarginLoss. In addition, it is more numerically stable.
(#9965).TensorOptions as the last argument. For example, replace your call to at::ones(torch::CPU(at::kFloat)), {2, 3}) with torch::ones({2, 3}, at::kCPU). This applies to the following functions:
arange, empty, eye, full, linspace, logspace, ones, rand, randint, randn, randperm, range, zeros.elementwise_mean to mean for loss reduction functions (#13419)>>> torch.empty((0, 2, 4, 0), dtype=torch.float64)
tensor([], size=(0, 2, 4, 0), dtype=torch.float64)
torch.roll operator to match numpy.roll (#13261, #13588, #13874).dtype, similar to numpy.finfo and numpy.iinfo (#12472).Tensor.__cuda_array_interface__ to provide compatibility with numba and other CUDA projects (#11984).Tensor.to_sparse() allows conversion from a dense tensor to a sparse tensor. (#12171)values() and torch.sparse_coo_tensor (with indices and values tensors). E.g., torch.sparse_coo_tensor(i, v).values().sum() is differentiable w.r.t. v. See the updated torch.sparse documentation for details. (#13001).dim argument. (#10423).expand method similar to torch.Tensor.expand. For example: torch.distributions.bernoulli.Bernoulli.expand. (#11341).copy keyword argument. (#12571).dtype accumulation argument. (#11719).compute_uv argument for optionally computing singular vectors (#12517).view) was incorrect with overlapping data locations. (#9538).reduction method. (#10018).__rsub__ now works properly when the CUDA device is not 0. (#12956).replacement=False will not properly throw an error message when there are no more categories to select (#12490).Tensor.__delitem__: fixed a segmentation fault on (#12726).load_from_state_dict now correctly handles 1-dimensional vs 0-dimensional tensors saved from 0.3 versions. (#9781).RuntimeError: storages don't support slicing when loading models saved with PyTorch 0.3. (#11314).reduce parameter. (#12689).eval mode. (#10621).out= parameters correctly, handles expanded tensors correctly, and has corrected argument validity checks on CPU. (#10273).Tensor gave incorrect results on CPU. (#10269).out parameter if it is given. (#9755).replacement=True could select 0 probability events on CUDA. (#9960).NaN.
(#10277).inf / -inf. (#11091).torch.nn.Conv modules with stride and dilation. (#9640).output_size calculation (#12952).Tensor constructors (e.g. torch.FloatTensor(...)) now correctly check their device argument.
(#11669).out parameter is a CPU Tensor for CPU unary ops. (#10358).None. (#12028).dir(torch) has been fixed with Python 3.7. (#10271).replacement=False and the input has fewer nonzero elements than num_samples. (#11933).torch.float16 dtype tensor to .grad. (#11781).can only join a started process error with torch.utils.data.DataLoader. (#11432).unexpected exit in torch.utils.data.DataLoader on KeyboardInterrupt. (#11718).nn.Parameter (#12886)torch.device inputs. (#10189).grad_fns. (#10181).np.int64 to PyTorch scalar. (#9225).eigenvectors=False is passed on CUDA rather than uninitialized data. (#10645).torch.utils.trainer (#12487)torch/torch.h header is deprecated in favor of torch/extension.h, which should be used in all C++ extensions going forward. Including torch/torch.h from a C++ extension will produce a warning. It is safe to batch replace torch/torch.h with torch/extension.h.torch::set_requires_grad. Replacement: at::Tensor now has a set_requires_grad method.torch::requires_grad. Replacement: at::Tensor now has a requires_grad method.torch::getVariableType. Replacement: None.torch.nn.parallel.deprecated.DistributedDataParallel.device, is_cuda, requires_grad, is_leaf and grad.
(#14339)tensors from seq. (#12741)Your coding agent can read these notes before it upgrades. Set up the MCP server →