NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #872 most downloaded on PyPI
Accelerate
Last release 23 days ago
09 Sep 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 54 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
6 years old
83 releases · first in 2020
One column per quarter.
Update deprecated logging warn by @SHi-ON in #881
We are very excited by the newly announced PyTorch 2.0 stack and you can try it using Accelerate on any model by using the dynamo_backend argument of the Accelerator, or when filling your config with accelerate config.
Note that to get the best performance, we recommend:
accelerate config update and accelerate config default. The first will update a config file to have the latest keys added from latter releases of Accelerate, and the second will create a default configuration file automatically mimicking write_default_config() introduced in #851 and #853 by @muellerzraccelerate launch which will show options relevant to the choices shown, such as accelerate launch --multi_gpu will show launch parameters relevant to multi-gpu training.join_uneven_inputs context manager to Accelerator by @Chris-hughes10 in #820default-config command by @muellerzr in #840batch_size by @pacman100 in #861The following contributors have made significant changes to the library over the last release:
join_uneven_inputs context manager to Accelerator (#820)Act on deprecations by @muellerzr in #813
Accelerate now supports Megatron-LM for the three model classes (BERT, GPT-2 and T5). You can learn more in the documentation.
Fixes a bug that returned SIGKILL errors on Windows.
notebook_launcherWith Kaggle now giving instances with two T4 GPUs, Accelerate can leverage this to do multi-gpu training from the notebook
non_blocking kwarg to send_to_device() by @NouamaneTazi in #607infer_auto_device_map by @younesbelkada in #792even_batches keyword to Accelerator by @Chris-hughes10 in #781AcceleratedOptimizer by @pacman100 in #811recurse argument in remove_hook_from_module by @younesbelkada in #812The following contributors have made significant changes to the library over the last release:
even_batches keyword to Accelerator (#781)[Device map] nn.Parameter don't have children in #747 by @patrickvonplaten
Fix num_processes is not defined #746 by @muellerzr
The accelerate command launch did not work well for distributed training using several machines. This is fixed in this version.
The accelerate command launch did not work well for distributed training using several machines. This is fixed in this version.
Instead of prefixing your launch command with CUDA_VISIBLE_DEVICES=xxx you can now specify the GPUs you want to use in your Accelerate config.
The tracebacks are now cleaned up to avoid printing several times the same error, and rich is integrated as an optional dependency.
subprocess from the multi-gpu launcher by @muellerzr in #623notebook_launcher by @pacman100 in #695init_empty_weights to override tensor constructor by @thomasw21 in #699grad_acc_steps from accelerator obj by @pacman100 in #698utils readability fixups by @ryanrussell in #711key_occurrence readability fixup by @ryanrussell in #710hooks readability improvements by @ryanrussell in #712The whole documentation has been revamped, just go look at it here!
The whole documentation has been revamped, just go look at it here!
When doing distributed evaluation, the dataloader loops back at the beginning of the dataset to make batches that have a round multiple of the number of processes. This causes the predictions to be slightly bigger than the length of the dataset, which used to require some truncating. This is all done behind the scenes now if you replace the gather your did in evaluation by gather_for_metrics.
When loading big models for inference, device_map="auto" used to fill the GPUs sequentially, making it hard to use a batch size > 1. It now balances the weights evenly on the GPUs so if you have more GPU space than the model size, you can do predictions with a bigger batch size!
Accelerate now supports M1 GPUs, to learn more about how to setup your environment, see the documentation.
mps device integration by @pacman100 in #596.run in WandBTracker. by @zh-plus in #605set_module_tensor_to_device by @sgugger in #576datasets by @lhoestq in #563accelerate launch by @muellerzr in #5530.6.7 fix by @pacman100 in #544The following contributors have made significant changes to the library over the last release:
Accelerate now handles gradient accumulation if you want, just pass along gradient_accumulation_steps=xxx when instantiating the Accelerator and put a
Accelerate now handles gradient accumulation if you want, just pass along gradient_accumulation_steps=xxx when instantiating the Accelerator and put all your training loop step under a with accelerator.accumulate(model):. Accelerate will then handle the loss re-scaling and gradient accumulation for you (avoiding slowdowns in distributed training when gradients only need to be synced when you want to step). More details in the documentation.
Accelerate now support SageMaker specific brand of data parallelism.
split_batches=True by @sgugger in #509total_batch_size attribute by @pacman100 in #493This release adds two major new features: the DeepSpeed integration has been revamped to match the one in Transformers Trainer, with multiple new opti
This release adds two major new features: the DeepSpeed integration has been revamped to match the one in Transformers Trainer, with multiple new options unlocked, and the TPU integration has been sped up.
This version also officially stops supporting Python 3.6 and requires Python 3.7+
Users can now specify a DeepSpeed config file when they want to use DeepSpeed, which unlocks many new options. More details in the new documentation.
If you're using TPUs we have sped up the dataloaders and models quite a bit, on top of a few bug fixes.
no_sync context wrapper + clean up some more warnings for DDP by @muellerzr in #428This release offers no significant new API, it is just needed to have access to some utils in Transformers.
This release offers no significant new API, it is just needed to have access to some utils in Transformers.
To handle very large models, new functionality has been added in Accelerate:
To handle very large models, new functionality has been added in Accelerate:
load_checkpoint_and_dispatch)See more in the documentation
Add guards for batch size finder 334
v0.7.0: Logging API, FSDP, batch size finder and examples revamp
v0.7.0: Logging API, FSDP, batch size finder and examples revamp
Use any of your favorite logging libraries (TensorBoard, Wandb, CometML...) with just a few lines of code inside your training scripts with Accelerate. All details are in the documentation.
PyTorch recently released a new model wrapper for sharded DDP training called FSDP. This release adds support for it (note that it doesn't work with mixed precision yet). See all caveats in the documentation.
Say goodbye to the CUDA OOM errors with the new find_executable_batch_size decorator. Just decorate your training function and pick a starting batch size, then let Accelerate do the rest.
The Accelerate examples are now split in two: you can find in the base folder a very simple nlp and computer vision examples, as well as complete versions incorporating all features. But you can also browse the examples in the by_feature subfolder, which will show you exactly what code to add for each given feature (checkpointing, tracking, cross-validation etc.)
mixed_precision for launch command by @sgugger in https://github.com/huggingface/accelerate/pull/300lr_scheduler to Accelerator.prepare by @sgugger in https://github.com/huggingface/accelerate/pull/301Full Changelog: https://github.com/huggingface/accelerate/compare/v0.6.0...v0.7.0
The launcher was ignoring the mixed precision attribute of the config since v0.6.0. This patch fixes that.
The launcher was ignoring the mixed precision attribute of the config since v0.6.0. This patch fixes that.
Patches an issue with mixed precision (see #286)
Patches an issue with mixed precision (see #286)
Accelerate now supports bfloat16 mixed precision training. As a result the old --fp16 argument has been deprecated to be replaced by the more generic…
This release adds support for bloat16 mixed precision training (requires PyTorch >= 1.10) and a brand-new checkpoint utility to help with resuming interrupted trainings. We also get a completely revamped documentation frontend.
Save the current state of all your objects (models, optimizers, RNG states) with accelerator.save_state(path_to_checkpoint) and reload everything by calling accelerator.load_state(path_to_checkpoint)
Accelerate now supports bfloat16 mixed precision training. As a result the old --fp16 argument has been deprecated to be replaced by the more generic --mixed-precision.
You can now type accelerate env to have a copy-pastable summary of your environment and default configuration. Very convenient when opening a new issue!
The documentation has been switched to the new Hugging Face frontend, like Transformers and Datasets.
store_true on argparse in nlp example by @monologg in https://github.com/huggingface/accelerate/pull/183set_to_none in Optimizer.zero_grad by @sgugger in https://github.com/huggingface/accelerate/pull/189debug_launcher by @sgugger in https://github.com/huggingface/accelerate/pull/259Full Changelog: https://github.com/huggingface/accelerate/compare/v0.5.1...v0.6.0
convert_to_fp32 returned booleans instead of tensors #173
Fix the two following bugs:
convert_to_fp32 returned booleans instead of tensors #173dispatch_batches=True #175This release introduces support for iterating through a DataLoader only on the main process, that then dispatches the batches to all processes.
This release introduces support for iterating through a DataLoader only on the main process, that then dispatches the batches to all processes.
The motivation behind this come from dataset streaming which introduces two difficulties:
This new feature is activated by default for all IterableDataset.
dispatch_batches #168 (@sgugger)This release adds support for DeepSpeed. While the basics are there to support ZeRO-2, ZeRo-3, as well a CPU and NVME offload, the API might evolve a
This release adds support for DeepSpeed. While the basics are there to support ZeRO-2, ZeRo-3, as well a CPU and NVME offload, the API might evolve a little bit as we polish it in the near future.
It also adds support for multi-node CPU. In both cases, just filling the questionnaire outputted by accelerate config and then launching your script with accelerate launch is enough, there are no changes in the main API.
accelerate test with no config file #79 (@cccntu)optimizer for consistency #81 (@kumapo)unscale_gradients method. #88 (@sgugger)OptimWrapper init #127 (@sgugger)After doing all the data preprocessing in your notebook, you can launch your training loop using the new notebook_launcher functionality. This is espe
After doing all the data preprocessing in your notebook, you can launch your training loop using the new notebook_launcher functionality. This is especially useful for Colab or Kaggle with TPUs! Here is an example on Colab (don't forget to select a TPU runtime).
This launcher also works if you have multiple GPUs on your machine. You just have to pass along num_processes=your_number_of_gpus in the call to notebook_launcher.
Our multi-node training test setup was flawed and the previous releases of 🤗 Accelerate were not working for multi-node distributed training. This is all fixed now and we have ensured to have more robust tests!
set_to_none to AcceleratedOptimizer.zero_grad #43 (@sgugger)Fix a bug preventing the load of a config with accelerate launch
Fix a bug preventing the load of a config with accelerate launch
It's now possible to launch your training script on AWS instances using SageMaker via accelerate launch.
It's now possible to launch your training script on AWS instances using SageMaker via accelerate launch.
To customize how the different objects used for mixed precision or distributed training are instantiated, a new API called KwargsHandler is added. This allows the user to pass along the kwargs that will be passed to those objects if used (and it is ignored if those are not used in the current setup, so the script can still run on any kind of setup).
Trying to gather tensors that are not of the same size across processes resulted in a process hang, a new method Accelerator.pad_across_processes has been added to help with that.
Initial release of 🤗 Accelerate. Checkout the main README or the docs to learn more about it!
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →