NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #830 most downloaded on PyPI
Accelerate
Last release 25 days ago
09 Sep 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 54 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
6 years old
83 releases · first in 2020
One column per quarter.
v1.15.0: FSDP2 activation memory, dtensor improvements
A large batch of FSDP2 work this release: two fixes that cut activation memory at long sequence lengths, tied-embedding support on torch >= 2.13, and a round of checkpointing correctness and scale fixes.
Activation checkpointing was wrapping each child of the matched layer (self_attn, mlp, the norms) instead of the layer itself, so every inter-child activation stayed saved for backward. It now wraps the layer.
There's also a new FSDP2-only activation_checkpointing_offload, which moves the remaining per-layer checkpoint inputs to pinned CPU memory. Gradients are exactly those of plain activation checkpointing:
# fsdp2.yaml
fsdp_config:
fsdp_version: 2
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_activation_checkpointing: true
fsdp_activation_checkpointing_offload: trueaccelerate launch --config_file fsdp2.yaml train.pyactivation_checkpointing_offload: offload checkpointed layer inputs to pinned CPU memory by @qgallouedec in #4175FULL_STATE_DICT dropping every rank's adapter shard except rank 0 by @AmineDiro in #4206Two fixes for DTensor-sharded models, which you hit with FSDP2, tensor parallelism, or any N-D parallelism setup: gradient clipping no longer fails on the foreach op when plain tensors and DTensors are mixed, and prepare_model leaves an already-sharded model where it is:
An entire model can now be dispatched to disk, including tied weights — useful for tools like llm-compressor that compress large models on machines that can't hold them:
Custom trackers can be registered by name and then selected from log_with= like any built-in one:
from accelerate import Accelerator
from accelerate.tracking import register_tracker_class
register_tracker_class(MyTracker) # MyTracker.name == "my_tracker"
accelerator = Accelerator(log_with="my_tracker")Neuron gains a torch dynamo backend (so --torch-compile works with the Transformers Trainer) and MPS is now reported and handled properly by accelerate env and find_executable_batch_size.
neuron backend for torch dynamo by @michaelbenayoun in #4097accelerate env by @xquantize in #4158estimate-memory for timm>=1.0.29 by adding the hf-hub: prefix by @iamsharduld in #4213Full Changelog: v1.14.0...v1.15.0
v1.15.0: FSDP2 activation memory, dtensor improvements Latest
Latest
Compare
This release brings a large batch of FSDP2 fixes and quality-of-life improvements: correct dtype handling on load, sharding of embeddings/norms, QLoRA
This release brings a large batch of FSDP2 fixes and quality-of-life improvements: correct dtype handling on load, sharding of embeddings/norms, QLoRA crash prevention, and a more robust auto-wrap policy.
Accelerate now works end-to-end on AMD ROCm devices. Thanks @Abdennacer-Badaoui!
Further Neuron improvements to reduce recompilation and cover missing device cases.
We improved offloading support for quantized models, including Torchao, int8, and tied-weight handling.
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.13.0...v1.14.0
v1.14.0: AMD ROCm support, FSDP2 hardening
Compare
add MS-AMP deprecation warnings by @neha222222 in https://github.com/huggingface/accelerate/pull/3857
We now have support for AWS Neuron (Trainium/Inferentia) devices. Thanks @michaelbenayoun for adding this.
We've removed IPEX dependency and improved device-agnostic code for XPU.
We've added a bunch of important fixes for FSDP2 users: upcasting only grad-requiring params, better tied embedding errors, DCP optimizer loading, bf16 optimizer step crash fix, and torch < 2.7.0 compatibility.
We've added several fixes to the DeepSpeed + Sequence Parallelism integration introduced in v1.12.0, including evaluation support during SP training and proper process group handling.
We've enhanced FP8 training. Thanks @shimizust for fixing torchao support.
Accelerate now imports faster by deferring heavy dependencies, and torch.compile hooks are disabled lazily.
We now have support for AWS Neuron (Trainium/Inferentia) devices. Thanks @michaelbenayoun for adding this.
We've removed IPEX dependency and improved device-agnostic code for XPU.
We've added a bunch of important fixes for FSDP2 users: upcasting only grad-requiring params, better tied embedding errors, DCP optimizer loading, bf16 optimizer step crash fix, and torch < 2.7.0 compatibility.
We've added several fixes to the DeepSpeed + Sequence Parallelism integration introduced in v1.12.0, including evaluation support during SP training and proper process group handling.
We've enhanced FP8 training. Thanks @shimizust for fixing torchao support.
Accelerate now imports faster by deferring heavy dependencies, and torch.compile hooks are disabled lazily.
v1.13.0: Neuron support, IPEX removal, and distributed training fixes
Compare
Deepspeed Ulysses/ALST is an efficient way of training on long sequences by employing sequence parallelism and attention head parallelism. You can lea
Deepspeed Ulysses/ALST is an efficient way of training on long sequences by employing sequence parallelism and attention head parallelism. You can learn more about this technology in this paper https://arxiv.org/abs/2506.13996 or this deepspeed tutorial https://www.deepspeed.ai/tutorials/ulysses-alst-sequence-parallelism/.
<img width="2368" height="1250" alt="0d8bd9e0" src="https://github.com/user-attachments/assets/b94e90c9-4368-4711-ad57-58de3c714ebc" />
To enable Deepspeed Ulysses, you first need to create ParallelismConfig and setting sp related args:
parallelism_config = ParallelismConfig(
sp_backend="deepspeed",
sp_size=2,
sp_handler=DeepSpeedSequenceParallelConfig(...),
)
Then, you need to make sure to compute the correct loss as described on our docs
...
losses_per_rank = torch.distributed.nn.functional.all_gather(loss, group=sp_group)
good_tokens = (shift_labels != -100).view(-1).sum()
good_tokens_per_rank = torch.distributed.nn.functional.all_gather(good_tokens, group=sp_group)
total_loss = sum(
losses_per_rank[rank] * good_tokens_per_rank[rank]
for rank in range(sp_world_size)
if good_tokens_per_rank[rank] > 0
)
total_good_tokens = sum(good_tokens_per_rank)
loss = total_loss / max(total_good_tokens, 1)
Thanks @S1ro1 for starting this work and for @stas00 for finishing this work. Also thanks @kashif for adding docs and reviewing/testing this PR !
This feature will also be available in HF Trainer thanks for this PR from @stas00: https://github.com/huggingface/transformers/pull/41832
cpu_ram_efficient_loading by @SunMarc in https://github.com/huggingface/accelerate/pull/3816Full Changelog: https://github.com/huggingface/accelerate/compare/v1.11.0...v1.12.0
Remove deprecated FindTiedParametersResult by @cyyever in https://github.com/huggingface/accelerate/pull/3786
We've added support for MXFP8 in our TransformerEngine integration. To use that, you need to set use_mxfp8_block_scaling in fp8_config. See nvidia docs [here]. (https://docs.nvidia.com/deeplearning/transformer-engine/user-guide/examples/fp8_primer.html#MXFP8-and-block-scaling)
BF16 and FP16 support for MPS devices is finally here. You can now pass mixed_precision = "fp16" or "bf16" when training on a mac (fp16 requires torch 2.8 and bf16 requires torch 2.6)
The following PRs add respectively support to ignored_params and no_sync() for FSDPv2:
Mixed precision can now be passed as a dtype string from accelerate cli flag or fsdp_config in accelerate config file:
Some minor updates concerning nd-parallelism.
We've dropped support for python 3.9 as it reached EOL in October.
cpu and offloaded to meta by @Qubitium in https://github.com/huggingface/accelerate/pull/3796with in Accelerator.autocast()instead of __enter__() and __exit__() for more elegant style. by @EquationWalker in https://github.com/huggingface/accelerate/pull/3767SWANLAB_MODE by @SunMarc in https://github.com/huggingface/accelerate/pull/3808Full Changelog: https://github.com/huggingface/accelerate/compare/v1.10.1...v1.11.0
Feat: add to_json by @S1ro1 in #3743
Full Changelog: v1.10.0...v1.10.1
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.10.0...v1.10.1
Training large models across multiple GPUs can be complex, especially when combining different parallelism strategies (e.g TP, CP, DP). To simplify th
Training large models across multiple GPUs can be complex, especially when combining different parallelism strategies (e.g TP, CP, DP). To simplify this process, we've collaborated with Axolotl to introduce an easy-to-use integration that allows you to apply any combination of parallelism strategies directly in your training script. Just pass a ParallelismConfig specifying the size of each parallelism type—it's that simple.
Learn more about how it works in our latest blogpost.
parallelism_config = ParallelismConfig(
dp_shard_size=2,
dp_replicate_size=2,
cp_size=2,
tp_size=2,
)
accelerator = Accelerator(
parallelism_config=parallelism_config,
...
)
model = AutoModelForCausalLM.from_pretrained("your-model-name", device_mesh=accelerator.torch_device_mesh)
model = accelerator.prepare(model)
ParallelismConfig from PartialState by @SunMarc in https://github.com/huggingface/accelerate/pull/3720We've fixed ignored modules attribute. With this, it is now possible to train PEFT model that moe layers that contrains q_proj and v_proj parameters. This is especially important for fine-tuning gpt-oss model.
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.9.0...v1.10.0
Training large models across multiple GPUs can be complex, especially when combining different parallelism strategies (e.g TP, CP, DP). To simplify this process, we've collaborated with Axolotl to introduce an easy-to-use integration that allows you to apply any combination of parallelism strategies directly in your training script. Just pass a ParallelismConfig specifying the size of each parallelism type—it's that simple.
Learn more about how it works in our latest blogpost.
parallelism_config = ParallelismConfig(
dp_shard_size=2,
dp_replicate_size=2,
cp_size=2,
tp_size=2,
)
accelerator = Accelerator(
parallelism_config=parallelism_config,
...
)
model = AutoModelForCausalLM.from_pretrained("your-model-name", device_mesh=accelerator.torch_device_mesh)
model = accelerator.prepare(model)ParallelismConfig from PartialState by @SunMarc in #3720We've fixed ignored modules attribute. With this, it is now possible to train PEFT model that moe layers that contrains q_proj and v_proj parameters. This is especially important for fine-tuning gpt-oss model.
Full Changelog: v1.9.0...v1.10.0
We've added support for a trackio, lightweight, 💯 free experiment tracking Python library built on top of 🤗 Datasets and Spaces.
We've added support for a trackio, lightweight, 💯 free experiment tracking Python library built on top of 🤗 Datasets and Spaces.
Main features are:
space_id.To use it with accelerate, you need to set log_with and initialize the trackers
accelerator = Accelerator(log_with="trackio")
config={"learning_rate": 0.001, "batch_size": 32}
# init_kwargs in order to host the dashboard on spaces
init_kwargs = {"trackio": {"space_id": "hf_username/space_name"}
accelerator.init_trackers("example_project", config=config, init_kwargs=init_kwargs})
Thanks @pcuenca for the integration !
set_module_tensor_to_device Setting tensor while clearing cache is very slow, so we added clear_device option to disable it.
Another small optimization is using non_blocking everywhere and syncing just before returning control to the user. This makes the loading slightly faster.
Accelerator() configuring by @pstjohn in https://github.com/huggingface/accelerate/pull/3677find_executable_batch_size() will no longer halves the batch after every OOM. Instead, we will multiply the batch size by 0.9. This should help user not waste gpu capacity.
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.8.1...v1.9.0
We've added support for a trackio, lightweight, 💯 free experiment tracking Python library built on top of 🤗 Datasets and Spaces.
Main features are:
space_id.To use it with accelerate, you need to set log_with and initialize the trackers
accelerator = Accelerator(log_with="trackio")
config={"learning_rate": 0.001, "batch_size": 32}
# init_kwargs in order to host the dashboard on spaces
init_kwargs = {"trackio": {"space_id": "hf_username/space_name"}
accelerator.init_trackers("example_project", config=config, init_kwargs=init_kwargs})Thanks @pcuenca for the integration !
set_module_tensor_to_device Setting tensor while clearing cache is very slow, so we added clear_device option to disable it.
Another small optimization is using non_blocking everywhere and syncing just before returning control to the user. This makes the loading slightly faster.
Accelerator() configuring by @pstjohn in #3677find_executable_batch_size() will no longer halves the batch after every OOM. Instead, we will multiply the batch size by 0.9. This should help user not waste gpu capacity.
Full Changelog: v1.8.1...v1.9.0
v1.9.0: Trackio support, Model loading speedup, Minor distributed improvements
Compare
Add support for e5e2 and default to hybrid when launcher is used by @IlyasMoutawwakil in https://github.com/huggingface/accelerate/pull/3640
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.8.0...v1.8.1
Full Changelog: v1.8.0...v1.8.1
ipex.optimize is being deprecated. Most optimizations have been upstreamed to PyTorch, and future improvements will land there directly. For users wit…
We've simplified how to prepare FSDPv2 models, as there were too many ways to compose FSDP2 with other features (e.g., FP8, torch.compile, activation checkpointing, etc.). Although the setup is now more restrictive, it leads to fewer errors and a more performant user experience. We’ve also added support for FP8. You can read about the results here. Thanks to @S1ro1 for this contribution!
We updated the CCL_WORKER_COUNT variable and added KMP parameters for Intel CPU users. This significantly improves distributed training performance (e.g., Tensor Parallelism), with up to a 40% speed-up on Intel 4th Gen Xeon when training transformer TP models.
We added support for regional compilation with the DeepSpeed engine. DeepSpeed’s .compile() modifies models in-place using torch.nn.Module.compile(...), rather than the out-of-place torch.compile(...), so we had to account for that. Thanks @IlyasMoutawwakil for this feature!
ipex.optimize is being deprecated. Most optimizations have been upstreamed to PyTorch, and future improvements will land there directly. For users without PyTorch 2.8, we’ll continue to rely on IPEX for now.
We've greatly expanded and stabilized support for Intel XPUs:
We've added support for SwanLab as an experiment tracking backend. Huge thanks to @ShaohonChen for this contribution ! We also deferred all tracker initializations to prevent premature setup of distributed environments.
accelerator().load_state() by @luiz0992 in https://github.com/huggingface/accelerate/pull/3540dtype_byte_size by @SunMarc in https://github.com/huggingface/accelerate/pull/3625Full Changelog: https://github.com/huggingface/accelerate/compare/v1.7.0...v1.8.0
We've simplified how to prepare FSDPv2 models, as there were too many ways to compose FSDP2 with other features (e.g., FP8, torch.compile, activation checkpointing, etc.). Although the setup is now more restrictive, it leads to fewer errors and a more performant user experience. We’ve also added support for FP8. You can read about the results here. Thanks to @S1ro1 for this contribution!
We updated the CCL_WORKER_COUNT variable and added KMP parameters for Intel CPU users. This significantly improves distributed training performance (e.g., Tensor Parallelism), with up to a 40% speed-up on Intel 4th Gen Xeon when training transformer TP models.
We added support for regional compilation with the DeepSpeed engine. DeepSpeed’s .compile() modifies models in-place using torch.nn.Module.compile(...), rather than the out-of-place torch.compile(...), so we had to account for that. Thanks @IlyasMoutawwakil for this feature!
ipex.optimize is being deprecated. Most optimizations have been upstreamed to PyTorch, and future improvements will land there directly. For users without PyTorch 2.8, we’ll continue to rely on IPEX for now.
We've greatly expanded and stabilized support for Intel XPUs:
We've added support for SwanLab as an experiment tracking backend. Huge thanks to @ShaohonChen for this contribution ! We also deferred all tracker initializations to prevent premature setup of distributed environments.
accelerator().load_state() by @luiz0992 in #3540dtype_byte_size by @SunMarc in #3625Full Changelog: v1.7.0...v1.8.0
v1.8.0: FSDPv2 + FP8, Regional Compilation for DeepSpeed, Faster Distributed Training on Intel CPUs, ipex.optimize deprecation
Compare
Remove deprecated PyTorch/XLA APIs by @zpcore in https://github.com/huggingface/accelerate/pull/3484
Instead of compiling the entire model at once, regional compilation targets repeated blocks (such as decoder layers) first. This allows the compiler to cache and reuse optimized code for subsequent blocks, significantly reducing the cold start compilation time typically seen during the first inference. Thanks @IlyasMoutawwakil for the feature ! You can view the full benchmark here, and check out our updated compilation guide for more details!
To enable this feature, set use_regional_compilation=True in the TorchDynamoPlugin configuration.
# Configure the compilation backend
dynamo_plugin = TorchDynamoPlugin(
use_regional_compilation=True,
... # other parameters
)
# Initialize accelerator with the plugin
accelerator = Accelerator(dynamo_plugin=dynamo_plugin)
# This will apply compile_regions to your model
model = accelerator.prepare(model)
We've introduced a new hook that enables per-layer upcasting and downcasting (e.g., for Linear layers) during inference. This allows users to run models with separate storage and compute dtypes, resulting in memory savings. The concept was first implemented in diffusers, where downcasting models to FP8 proved effective without major quality degradation. Contributed by @sayakpaul in https://github.com/huggingface/accelerate/pull/3427
model = ....
storage_dtype = torch.float8_e4m3fn
compute_dtype = torch.bfloat16
attach_layerwise_casting_hooks(
model,
storage_dtype=storage_dtype,
compute_dtype=compute_dtype,
)
This release includes numerous new features and bug fixes. Notably, we’ve added support for FULL_STATE_DICT, a widely used option in FSDP, now enabling .save_pretrained() in transformers to work with FSDP2 wrapped models. QLoRA training is now supported as well but more testing is needed. We have also resolved a backend issue related to parameter offloading to CPU. Additionally, a significant memory spike that occurred when cpu_ram_efficient_loading=True was enabled has been fixed. Several other minor improvements and fixes are also included—see the What’s Changed section for full details.
FULL_STATE_DICT have been enabled by @S1ro1 in https://github.com/huggingface/accelerate/pull/3527cpu_ram_efficient_loading=True by @S1ro1 in https://github.com/huggingface/accelerate/pull/3482We have added a documentation for Intel Gaudi hardware ! The support is already available since v1.5.0 through this PR.
dynamic argumentWe've updated the logic for setting self.dynamic to explicitly preserve None rather than defaulting to False when the USE_DYNAMIC environment variable is unset. This change aligns the behavior with the PyTorch documentation for torch.compile. Thanks to @yafshar for contributing this improvement in #3567.
low_precision_training guide by @sadra-barikbin in https://github.com/huggingface/accelerate/pull/3488torch.distributed.checkpoint.state_dict.set_model_state_dict in load_checkpoint_in_model by @ringohoffman in https://github.com/huggingface/accelerate/pull/3432weights_only=True by @bzhong-solink in https://github.com/huggingface/accelerate/pull/3497cpu_ram_efficient_loading=True by @S1ro1 in https://github.com/huggingface/accelerate/pull/3482accelerator.prepare + IPEX for 2+ nn.Models and/or optim.Optimizers by @mariusarvinte in https://github.com/huggingface/accelerate/pull/3517set_epoch does not take effect. by @hongjx175 in https://github.com/huggingface/accelerate/pull/3556_cast_and_contiguous by @dlvp in https://github.com/huggingface/accelerate/pull/3559cpu_ram_efficient_loading by @SumanthRH in https://github.com/huggingface/accelerate/pull/3307synchronize call for xpu in _gpu_gather by @faaany in https://github.com/huggingface/accelerate/pull/3563Full Changelog: https://github.com/huggingface/accelerate/compare/v1.6.0...v1.7.0
Instead of compiling the entire model at once, regional compilation targets repeated blocks (such as decoder layers) first. This allows the compiler to cache and reuse optimized code for subsequent blocks, significantly reducing the cold start compilation time typically seen during the first inference. Thanks @IlyasMoutawwakil for the feature ! You can view the full benchmark here, and check out our updated compilation guide for more details!
To enable this feature, set use_regional_compilation=True in the TorchDynamoPlugin configuration.
# Configure the compilation backend
dynamo_plugin = TorchDynamoPlugin(
use_regional_compilation=True,
... # other parameters
)
# Initialize accelerator with the plugin
accelerator = Accelerator(dynamo_plugin=dynamo_plugin)
# This will apply compile_regions to your model
model = accelerator.prepare(model)We've introduced a new hook that enables per-layer upcasting and downcasting (e.g., for Linear layers) during inference. This allows users to run models with separate storage and compute dtypes, resulting in memory savings. The concept was first implemented in diffusers, where downcasting models to FP8 proved effective without major quality degradation. Contributed by @sayakpaul in #3427
model = ....
storage_dtype = torch.float8_e4m3fn
compute_dtype = torch.bfloat16
attach_layerwise_casting_hooks(
model,
storage_dtype=storage_dtype,
compute_dtype=compute_dtype,
)This release includes numerous new features and bug fixes. Notably, we’ve added support for FULL_STATE_DICT, a widely used option in FSDP, now enabling .save_pretrained() in transformers to work with FSDP2 wrapped models. QLoRA training is now supported as well but more testing is needed. We have also resolved a backend issue related to parameter offloading to CPU. Additionally, a significant memory spike that occurred when cpu_ram_efficient_loading=True was enabled has been fixed. Several other minor improvements and fixes are also included—see the What’s Changed section for full details.
FULL_STATE_DICT have been enabled by @S1ro1 in #3527cpu_ram_efficient_loading=True by @S1ro1 in #3482We have added a documentation for Intel Gaudi hardware !
The support is already available since v1.5.0 through this PR.
dynamic argumentWe've updated the logic for setting self.dynamic to explicitly preserve None rather than defaulting to False when the USE_DYNAMIC environment variable is unset. This change aligns the behavior with the PyTorch documentation for torch.compile. Thanks to @yafshar for contributing this improvement in #3567.
low_precision_training guide by @sadra-barikbin in #3488torch.distributed.checkpoint.state_dict.set_model_state_dict in load_checkpoint_in_model by @ringohoffman in #3432weights_only=True by @bzhong-solink in #3497cpu_ram_efficient_loading=True by @S1ro1 in #3482accelerator.prepare + IPEX for 2+ nn.Models and/or optim.Optimizers by @mariusarvinte in #3517set_epoch does not take effect. by @hongjx175 in #3556_cast_and_contiguous by @dlvp in #3559cpu_ram_efficient_loading by @SumanthRH in #3307synchronize call for xpu in _gpu_gather by @faaany in #3563Full Changelog: v1.6.0...v1.7.0
v1.7.0 : Regional compilation, Layerwise casting hook, FSDPv2 + QLoRA
Compare
This release introduces the support for FSDPv2 thanks to @S1ro1.
This release introduces the support for FSDPv2 thanks to @S1ro1.
If you are using python code, you need to set fsdp_version=2 in FullyShardedDataParallelPlugin:
from accelerate import FullyShardedDataParallelPlugin, Accelerator
fsdp_plugin = FullyShardedDataParallelPlugin(
fsdp_version=2
# other options...
)
accelerator = Accelerator(fsdp_plugin=fsdp_plugin)
If want to convert a YAML config that contains the FSDPv1 config to FSDPv2 one , use our conversion tool:
accelerate to-fsdp2 --config_file config.yaml --output_file new_config.yaml`
To learn more about the difference between FSDPv1 and FSDPv2, read the following documentation.
We have added initial support for DeepSpeed + TP. Not many changes were required as the DeepSpeed APIs was already compatible. We only needed to make sure that the dataloader was compatible with TP and that we were able to save the TP weights. Thanks @inkcherry for the work ! https://github.com/huggingface/accelerate/pull/3390.
To use TP with deepspeed, you need to update the setting in the deepspeed config file by including tensor_parallel key:
....
"tensor_parallel":{
"autotp_size": ${autotp_size}
},
...
More details in this deepspeed PR.
We've added support for XCCL which is an Intel distributed backend which can be used with XPU devices. More details in this torch PR. Thanks @dvrogozh for the integration !
log_artifact, log_artifacts and log_figure capabilities to the MLflowTracker. by @luiz0992 in https://github.com/huggingface/accelerate/pull/3419Full Changelog: https://github.com/huggingface/accelerate/compare/v1.5.2...v1.6.0
Fixed an issue with torch.get_default_device() requiring a higher version than what we support
Bug Fixes:
torch.get_default_device() requiring a higher version than what we supportpytest import in prodFull Changelog: https://github.com/huggingface/accelerate/compare/v1.5.0...v1.5.2
Nothing published for this version
Adds in HPU accelerator support for 🤗 Accelerate
import torch by @faaany in https://github.com/huggingface/accelerate/pull/3396device=torch.get_default_device() in torch.Generators by @saforem2 in https://github.com/huggingface/accelerate/pull/3420Full Changelog: https://github.com/huggingface/accelerate/compare/v1.4.0...v1.5.0
This release introduces a new FP8 API and brings in a new backend: `torchao`. To use, pass in AORecipeKwargs to the Accelerator while setting mixed_pr
torchao FP8, initial Tensor Parallel support, and memory leak fixestorchao FP8This release introduces a new FP8 API and brings in a new backend: torchao. To use, pass in AORecipeKwargs to the Accelerator while setting mixed_precision="fp8". This is initial support, as it matures we will incorporate more into it (such as accelerate config/yaml) in future releases. See our benchmark examples here
We have intial support for an in-house solution to TP when working with accelerate dataloaders. check out the PR here
memory leak] Replace GradientState -> DataLoader reference with weakrefs by @tomaarsen in https://github.com/huggingface/accelerate/pull/3391tests/test_quantization.py on XPU by @faaany in https://github.com/huggingface/accelerate/pull/3349require_non_xpu test markers by @faaany in https://github.com/huggingface/accelerate/pull/3301memory leak] Replace GradientState -> DataLoader reference with weakrefs by @tomaarsen in https://github.com/huggingface/accelerate/pull/3391get_quantized_model_device_map by @faaany in https://github.com/huggingface/accelerate/pull/3397Full Changelog: https://github.com/huggingface/accelerate/compare/v1.3.0...v1.4.0
As it's been ~2 years since torch 2.0 was first released, we are now requiring this as the minimum version for Accelerate, which similarly was done in
As it's been ~2 years since torch 2.0 was first released, we are now requiring this as the minimum version for Accelerate, which similarly was done in transformers as of its last release.
keep_torch_compile param to unwrap_model and extract_model_from_parallel for distributed compiled model. by @ggoggam in https://github.com/huggingface/accelerate/pull/3282keep_torch_compile param to unwrap_model and extract_model_from_parallel for distributed compiled model. by @ggoggam in https://github.com/huggingface/accelerate/pull/3282Full Changelog: https://github.com/huggingface/accelerate/compare/v1.2.1...v1.3.0
fix: add max_memory to _init_infer_auto_device_map's return statement in https://github.com/huggingface/accelerate/pull/3279 by @Nech-C
Full Changelog: https://github.com/huggingface/accelerate/compare/v1.2.0...v1.2.1
enable find_executable_batch_size on XPU by @faaany in https://github.com/huggingface/accelerate/pull/3236
find_executable_batch_size on XPU by @faaany in https://github.com/huggingface/accelerate/pull/3236numpy._core instead of numpy.core by @qgallouedec in https://github.com/huggingface/accelerate/pull/3247data_loader] Optionally also propagate set_epoch to batch sampler by @tomaarsen in https://github.com/huggingface/accelerate/pull/3246accelerate config prompt text by @faaany in https://github.com/huggingface/accelerate/pull/3268align_module_device, ensure only cpu tensors for get_state_dict_offloaded_model by @kylesayrs in https://github.com/huggingface/accelerate/pull/3217get_state_dict_from_offload by @kylesayrs in https://github.com/huggingface/accelerate/pull/3253preload_module_classes is lost for nested modules by @wejoncy in https://github.com/huggingface/accelerate/pull/3248Update code in tracking documentation by @faaany in https://github.com/huggingface/accelerate/pull/3235
Replaced set/check breakpoint with set/check trigger in the troubleshooting documentation by @relh in https://github.com/huggingface/accelerate/pull/3259
Update set-seed by @faaany in https://github.com/huggingface/accelerate/pull/3228
Fix typo by @faaany in https://github.com/huggingface/accelerate/pull/3221
Use real path for checkpoint by @faaany in https://github.com/huggingface/accelerate/pull/3220
Fixed multiple typos for Tutorials and Guides docs by @henryhmko in https://github.com/huggingface/accelerate/pull/3274
align_module_device, ensure only cpu tensors for get_state_dict_offloaded_model by @kylesayrs in https://github.com/huggingface/accelerate/pull/3217find_executable_batch_size on XPU by @faaany in https://github.com/huggingface/accelerate/pull/3236data_loader] Optionally also propagate set_epoch to batch sampler by @tomaarsen in https://github.com/huggingface/accelerate/pull/3246numpy._core instead of numpy.core by @qgallouedec in https://github.com/huggingface/accelerate/pull/3247accelerate config prompt text by @faaany in https://github.com/huggingface/accelerate/pull/3268get_state_dict_from_offload by @kylesayrs in https://github.com/huggingface/accelerate/pull/3253preload_module_classes is lost for nested modules by @wejoncy in https://github.com/huggingface/accelerate/pull/3248checkpoint by @faaany in https://github.com/huggingface/accelerate/pull/3220Release diff: https://github.com/huggingface/accelerate/compare/v1.1.1...v1.2.0
Nothing published for this version
Allow for a data_seed argument in https://github.com/huggingface/accelerate/pull/3150
data_seed argument in https://github.com/huggingface/accelerate/pull/3150weights_only=True by default for all compatible objects when checkpointing and saving with torch.save in https://github.com/huggingface/accelerate/pull/3036dim input in pad_across_processes in https://github.com/huggingface/accelerate/pull/3114has_offloaded_params utility added in https://github.com/huggingface/accelerate/pull/3188dim input in pad_across_processes by @mariusarvinte in https://github.com/huggingface/accelerate/pull/3114data_seed by @muellerzr in https://github.com/huggingface/accelerate/pull/3150save_model by @muellerzr in https://github.com/huggingface/accelerate/pull/3146weights_only=True by default for all compatible objects by @muellerzr in https://github.com/huggingface/accelerate/pull/3036get_xpu_available_memory by @faaany in https://github.com/huggingface/accelerate/pull/3165has_offloaded_params by @kylesayrs in https://github.com/huggingface/accelerate/pull/3188torch.nn.Module model into account when moving to device by @faaany in https://github.com/huggingface/accelerate/pull/3167torchrun by @faaany in https://github.com/huggingface/accelerate/pull/3166align_module_device by @kylesayrs in https://github.com/huggingface/accelerate/pull/3204Full Changelog: https://github.com/huggingface/accelerate/compare/v1.0.1...v1.1.0
Fixes an issue where the auto values were no longer being parsed when using deepspeed
auto values were no longer being parsed when using deepspeedFull Changelog: https://github.com/huggingface/accelerate/compare/v1.0.0...v1.0.1
With these release notes, we will focus first on the major breaking changes to get your code fixed, followed by what is new specifically between 0.34.…
With accelerate 1.0, we are officially stating that the core parts of the API are now "stable" and ready for the future of what the world of distributed training and PyTorch has to handle. With these release notes, we will focus first on the major breaking changes to get your code fixed, followed by what is new specifically between 0.34.0 and 1.0.
To read more, check out our official blog here
dispatch_batches, split_batches, even_batches, and use_seedable_sampler to the Accelerator() should now be handled by creating an accelerate.utils.DataLoaderConfiguration() and passing this to the Accelerator() instead (Accelerator(dataloader_config=DataLoaderConfiguration(...)))Accelerator().use_fp16 and AcceleratorState().use_fp16 have been removed; this should be replaced by checking accelerator.mixed_precision == "fp16"Accelerator().autocast() no longer accepts a cache_enabled argument. Instead, an AutocastKwargs() instance should be used which handles this flag (among others) passing it to the Accelerator (Accelerator(kwargs_handlers=[AutocastKwargs(cache_enabled=True)]))accelerate.utils.is_tpu_available should be replaced with accelerate.utils.is_torch_xla_availableaccelerate.utils.modeling.shard_checkpoint should be replaced with split_torch_state_dict_into_shards from the huggingface_hub libraryaccelerate.tqdm.tqdm() no longer accepts True/False as the first argument, and instead, main_process_only should be passed in as a named argumentAfter long request, we finally have multiple model DeepSpeed support in Accelerate! (though it is quite early still). Read the full tutorial here, however essentially:
When using multiple models, a DeepSpeed plugin should be created for each model (and as a result, a separate config). a few examples are below:
(Where we train only one model, zero3, and another is used for inference, zero2)
from accelerate import Accelerator
from accelerate.utils import DeepSpeedPlugin
zero2_plugin = DeepSpeedPlugin(hf_ds_config="zero2_config.json")
zero3_plugin = DeepSpeedPlugin(hf_ds_config="zero3_config.json")
deepspeed_plugins = {"student": zero2_plugin, "teacher": zero3_plugin}
accelerator = Accelerator(deepspeed_plugins=deepspeed_plugins)
To then select which plugin to be used at a certain time (aka when calling prepare), we call `accelerator.state.select_deepspeed_plugin("name"), where the first plugin is active by default:
accelerator.state.select_deepspeed_plugin("student")
student_model, optimizer, scheduler = ...
student_model, optimizer, scheduler, train_dataloader = accelerator.prepare(student_model, optimizer, scheduler, train_dataloader)
accelerator.state.select_deepspeed_plugin("teacher") # This will automatically enable zero init
teacher_model = AutoModel.from_pretrained(...)
teacher_model = accelerator.prepare(teacher_model)
For disjoint models, separate accelerators should be used for each model, and their own .backward() should be called later:
for batch in dl:
outputs1 = first_model(**batch)
first_accelerator.backward(outputs1.loss)
first_optimizer.step()
first_scheduler.step()
first_optimizer.zero_grad()
outputs2 = model2(**batch)
second_accelerator.backward(outputs2.loss)
second_optimizer.step()
second_scheduler.step()
second_optimizer.zero_grad()
We've enabled MS-AMP support up to FSDP. At this time we are not going forward with implementing FSDP support with MS-AMP, due to design issues between both libraries that don't make them inter-op easily.
__reduce__ by @byi8220 in https://github.com/huggingface/accelerate/pull/3074_get_named_modules by @faaany in https://github.com/huggingface/accelerate/pull/3052skip_keys usage in forward hooks by @152334H in https://github.com/huggingface/accelerate/pull/3088torch.cuda.amp.GradScaler FutureWarning for pytorch 2.4+ by @Mon-ius in https://github.com/huggingface/accelerate/pull/3132Full Changelog: https://github.com/huggingface/accelerate/compare/v0.34.2...v1.0.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixes an issue where processed DataLoaders could no longer be pickled in #3074 thanks to @byi8220
DataLoaders could no longer be pickled in #3074 thanks to @byi8220default_transformers_cls_names_to_wrap would separate _no_split_modules by characters instead of keeping it as a list of layer names in #3075Full Changelog: https://github.com/huggingface/accelerate/compare/v0.34.0...v0.34.1
Updated Safetensors Requirement: The library now requires safetensors version 0.4.3.
safetensors version 0.4.3.numpy 2.0.0accelerate library will handle this automatically with accelerator.end_training(), or you can do it manually using PartialState().destroy_process_group().transfer_to_npu, ensuring better performance and compatibility.StatefulDataLoader from torchdata, allowing better handling of data loading states. Enable by passing use_stateful_dataloader=True to the DataLoaderConfiguration, and when calling load_state() the DataLoader will automatically be resumed from its last step, no more having to iterate through passed batches.prepare_data_loader() function is now independent of the Accelerator, giving you more flexibility towards which API levels you would like to use.DataLoader states, ensuring smoother training sessions.set_epoch function for MpDeviceLoaderWrapper.TransformerEngine FP8 training, including better defaults for the quantized FP8 weights.TransformerEngine integration works exactly as intended. These scripts run one half using 🤗 Accelerate's integration, the other with raw TransformersEngine, providing users with a nice example of what we do under the hood with accelerate, and a good sanity check to make sure nothing breaks down over time. Find them hereTransformerEngine and accelerate as well. Use docker pull huggingface/accelerate@gpu-fp8-transformerengine to quickly get an environment going.torchpippy no more, long live torch.distributed.pipeliningtorchpippy is now fully integrated into torch core, and as a result we are exclusively supporting the PyTorch implementation from now on[1, n, n] rather than [2, n, n] as before.pipelining no longer supports encoder/decoder models, so the t5 example has been removed.torchpippy potentially if needed.FullyShardedDataParallelPlugin yourself manually with no need for environment patching:from accelerate import FullyShardedDataParallelPlugin
fsdp_plugin = FullyShardedDataParallelPlugin(...)
accelerate launch and need to ensure the env variables are setup properly for model loading:from accelerate.utils import enable_fsdp_ram_efficient_loading, disable_fsdp_ram_efficient_loading
enable_fsdp_ram_efficient_loading()
axolotl library, so very big kudos to their wonderful workstep when loading the state by @muellerzr in https://github.com/huggingface/accelerate/pull/2992find_tied_params for models with shared layers by @qubvel in https://github.com/huggingface/accelerate/pull/2986transformer_engine on import by @oraluben in https://github.com/huggingface/accelerate/pull/3056skip_first_batches support for StatefulDataloader and fix all the tests by @muellerzr in https://github.com/huggingface/accelerate/pull/3068step when loading the state by @muellerzr in https://github.com/huggingface/accelerate/pull/2992find_tied_params for models with shared layers by @qubvel in https://github.com/huggingface/accelerate/pull/2986end_training by @SunMarc in https://github.com/huggingface/accelerate/pull/3012torchdata.stateful_dataloader.StatefulDataLoader within the Accelerator by @byi8220 in https://github.com/huggingface/accelerate/pull/2895prepare_data_loader() from Accelerator by @siddk in https://github.com/huggingface/accelerate/pull/3047transformer_engine on import by @oraluben in https://github.com/huggingface/accelerate/pull/3056skip_first_batches support for StatefulDataloader and fix all the tests by @muellerzr in https://github.com/huggingface/accelerate/pull/3068Small release this month, with key focuses on some added support for backends and bugs:
Small release this month, with key focuses on some added support for backends and bugs:
torch.float8_e4m3fn format dtype_byte_size by @SunMarc in https://github.com/huggingface/accelerate/pull/2945device_map="auto" by @muellerzr in https://github.com/huggingface/accelerate/pull/2914multi_gpu was being set and warning being printed even with num_processes=1 by @HarikrishnanBalagopal in https://github.com/huggingface/accelerate/pull/2921pip caching in CI by @SauravMaheshkar in https://github.com/huggingface/accelerate/pull/2952Full Changelog: https://github.com/huggingface/accelerate/compare/v0.32.1...v0.33.0
Nothing published for this version
Utilize shard saving from the huggingface_hub rather than our own implementation
huggingface_hub rather than our own implementation (https://github.com/huggingface/accelerate/pull/2795)dispatch_model (https://github.com/huggingface/accelerate/pull/2855)Accelerator.step number is now restored when using save_state and load_state (https://github.com/huggingface/accelerate/pull/2765)import accelerate and any other major core import by 68%, now should be only slightly longer than doing import torch (https://github.com/huggingface/accelerate/pull/2845)get_backend and added a clear_device_cache utility (https://github.com/huggingface/accelerate/pull/2857)allreduce. (https://github.com/huggingface/accelerate/pull/2841)log_line_prefix_template optional the notebook_launcher (https://github.com/huggingface/accelerate/pull/2888)accelerate merge-weights, one will be automatically created (https://github.com/huggingface/accelerate/pull/2854).safetensors (https://github.com/huggingface/accelerate/pull/2853)torch>=2.4 (https://github.com/huggingface/accelerate/pull/2825)@require_triton test decorator and enable test_dynamo work on xpu (https://github.com/huggingface/accelerate/pull/2878)load_state_dict not working on xpu and refine xpu safetensors version check (https://github.com/huggingface/accelerate/pull/2879)accelerate launch (https://github.com/huggingface/accelerate/pull/2902)dispatch_model by @panjd123 in https://github.com/huggingface/accelerate/pull/2855test_tracking.ClearMLTest by @faaany in https://github.com/huggingface/accelerate/pull/2863torch_device instead of 0 for device check by @faaany in https://github.com/huggingface/accelerate/pull/2861test_zero3_integration by @faaany in https://github.com/huggingface/accelerate/pull/2864log_line_prefix_template Optional in Elastic Launcher for Backward Compatibility by @yhna940 in https://github.com/huggingface/accelerate/pull/2888require_triton and enable test_dynamo work on xpu by @faaany in https://github.com/huggingface/accelerate/pull/2878load_state_dict for xpu and refine xpu safetensor version check by @faaany in https://github.com/huggingface/accelerate/pull/2879Full Changelog: https://github.com/huggingface/accelerate/compare/v0.31.0...v0.32.0
Set timeout default to PyTorch defaults based on backend by @muellerzr in https://github.com/huggingface/accelerate/pull/2758
timeout default to PyTorch defaults based on backend by @muellerzr in https://github.com/huggingface/accelerate/pull/2758notebook_launcher by @yhna940 in https://github.com/huggingface/accelerate/pull/2788logging to log the actual user call site (instead of the call site inside the logger wrapper) of log functions by @luowyang in https://github.com/huggingface/accelerate/pull/2730notebook_launcher by @yhna940 in https://github.com/huggingface/accelerate/pull/2788get_balanced_memory by @faaany in https://github.com/huggingface/accelerate/pull/2826stage3_prefetch_bucket_size value to an integer by @adk9 in https://github.com/huggingface/accelerate/pull/2814Full Changelog: https://github.com/huggingface/accelerate/compare/v0.30.1...v0.31.0
Fix duplicate environment variable check in multi-cpu condition thanks to @yhna940 in https://github.com/huggingface/accelerate/pull/2752
Full Changelog: https://github.com/huggingface/accelerate/compare/v0.30.0...v0.30.1
Don't use deprecated Repository anymore by @Wauplin in https://github.com/huggingface/accelerate/pull/2658
tqdm wrapper to make it fully passthrough, no need to have tqdm(main_process_only, *args), it is now just tqdm(*args) and you can pass in is_main_process as a kwarg.cann version info to command accelerate env for NPU by @statelesshz in https://github.com/huggingface/accelerate/pull/2689deepspeed-specific Docker image by @muellerzr in https://github.com/huggingface/accelerate/pull/2707. To use, pull the gpu-deepspeed tag docker pull huggingface/accelerate:cuda-deepspeed-nightlyis_train_batch_min type in DeepSpeedPlugin by @yhna940 in https://github.com/huggingface/accelerate/pull/2646free_memory to deal with garbage collection by @muellerzr in https://github.com/huggingface/accelerate/pull/2716execution_device by @faaany in https://github.com/huggingface/accelerate/pull/2612is_train_batch_min type in DeepSpeedPlugin by @yhna940 in https://github.com/huggingface/accelerate/pull/2646Repository anymore by @Wauplin in https://github.com/huggingface/accelerate/pull/2658tqdm: *args should come ahead of main_process_only by @rb-synth in https://github.com/huggingface/accelerate/pull/2654free_memory to deal with garbage collection by @muellerzr in https://github.com/huggingface/accelerate/pull/2716Full Changelog: https://github.com/huggingface/accelerate/compare/v0.29.3...v0.30.0
Nothing published for this version
Fixes issue with backend refactor not working on CPU-based distributed environments by @jiqing-feng: https://github.com/huggingface/accelerate/pull/26
load_checkpoint_and_dispatch needs a strict argumentFull Changelog: https://github.com/huggingface/accelerate/compare/v0.29.2...v0.29.3
Fixes xpu missing parenthesis https://github.com/huggingface/accelerate/pull/2639
Fixed an import which would cause running accelerate CLI to fail if pytest wasn't installed
Fixed an import which would cause running accelerate CLI to fail if pytest wasn't installed
Accelerate can now optimize NUMA affinity, which can help increase throughput on NVIDIA multi-GPU systems. To enable it either follow the prompt durin
accelerate config, set the ACCELERATE_CPU_AFFINITY=1 env variable, or manually using the following:from accelerate.utils import set_numa_affinity
# For GPU 0
set_numa_affinity(0)
Big thanks to @stas00 for the recommendation, request, and feedback during development
set_seed by @muellerzr in https://github.com/huggingface/accelerate/pull/2569BatchSamplerShard by @universuen in https://github.com/huggingface/accelerate/pull/2584notebook_launcher can use multiple GPUs in Google Colab if using a custom instance that supports multiple GPUs by @StefanTodoran in https://github.com/huggingface/accelerate/pull/2561load_checkpoint_in_model behavior when unexpected keys are in the checkpoint by @fxmarty in https://github.com/huggingface/accelerate/pull/2588main_process_ip and master_addr when not using standard as deepspeed launcher by @asdfry in https://github.com/huggingface/accelerate/pull/2495deepspeed, set it with the DS_ENV_FILE environmental variable by @muellerzr in https://github.com/huggingface/accelerate/pull/2566main_process_ip and master_addr when not using standard as deepspeed launcher by @asdfry in https://github.com/huggingface/accelerate/pull/2495load_checkpoint_in_model behavior when unexpected keys are in the checkpoint by @fxmarty in https://github.com/huggingface/accelerate/pull/2588Full Changelog: https://github.com/huggingface/accelerate/compare/v0.28.0...v0.29.0
Introduce a DataLoaderConfiguration and begin deprecation of arguments in the Accelerator `diff +from accelerate import DataLoaderConfiguration +dl_co…
DataLoaderConfiguration and begin deprecation of arguments in the Accelerator+from accelerate import DataLoaderConfiguration
+dl_config = DataLoaderConfiguration(split_batches=True, dispatch_batches=True)
-accelerator = Accelerator(split_batches=True, dispatch_batches=True)
+accelerator = Accelerator(dataloader_config=dl_config)
from accelerate import GradientAccumulationPlugin
plugin = GradientAccumulationPlugin(
+ num_steps=2,
sync_each_batch=sync_each_batch
)
accelerator = Accelerator(gradient_accumulation_plugin=plugin)
launch changesmpirun for multi-cpu training by @dmsuehir in https://github.com/huggingface/accelerate/pull/2493is_torch_tensor over hasattr for torch.compile. by @PhilJd in https://github.com/huggingface/accelerate/pull/2387DataLoaderConfig by @muellerzr in https://github.com/huggingface/accelerate/pull/2441is_namedtuple implementation by @fxmarty in https://github.com/huggingface/accelerate/pull/2475os.path.sep.join path manipulations with a helper by @akx in https://github.com/huggingface/accelerate/pull/2446XLA device type by @will-cromar in https://github.com/huggingface/accelerate/pull/2467Accelerator to detect distributed type from the "LOCAL_RANK" env variable for XPU by @faaany in https://github.com/huggingface/accelerate/pull/2473accelerate launch by @muellerzr in https://github.com/huggingface/accelerate/pull/2498----main_process_port to --main_process_port) by @DerrickWang005 in https://github.com/huggingface/accelerate/pull/2516PYTORCH_NVML_BASED_CUDA_CHECK when calling accelerate.utils.imports.is_cuda_available() by @luiscape in https://github.com/huggingface/accelerate/pull/2524env=os.environ.copy()s by @akx in https://github.com/huggingface/accelerate/pull/2449zero_grad(set_to_none=None) to align with PyTorch by @yongchanghao in https://github.com/huggingface/accelerate/pull/2472Full Changelog: https://github.com/huggingface/accelerate/compare/v0.27.2...v0.28.0
Nothing published for this version
Nothing published for this version
With the latest release of PyTorch 2.2.0, we've guaranteed that there are no breaking changes regarding it
With the latest release of PyTorch 2.2.0, we've guaranteed that there are no breaking changes regarding it
With this release we are excited to announce support for pipeline-parallel inference by integrating PyTorch's PiPPy framework (so no need to use Megatron or DeepSpeed)! This supports automatic model-weight splitting to each device using a similar API to device_map="auto". This is still under heavy development, however the inference side is stable enough that we are ready for a release. Read more about it in our docs and check out the example zoo.
Requires pippy of version 0.2.0 or later (pip install torchpippy -U)
Example usage (combined with accelerate launch or torchrun):
from accelerate import PartialState, prepare_pippy
model = AutoModelForSequenceClassification.from_pretrained("gpt2")
model = prepare_pippy(model, split_points="auto", example_args=(input,))
input = input.to("cuda:0")
with torch.no_grad():
output = model(input)
# The outputs are only on the final process by default
# You can pass in `gather_outputs=True` to prepare_pippy to
# make them available on all processes
if PartialState().is_last_process:
output = torch.stack(tuple(output[0]))
print(output.shape)
This release provides support for utilizing DeepSpeed on XPU devices thanks to @faaany
dispatch_model, and in forward with offloading by @fxmarty in https://github.com/huggingface/accelerate/pull/2330accelerate config by @faaany in https://github.com/huggingface/accelerate/pull/2346block_size picking in megatron_lm_gpt_pretraining example. by @nilq in https://github.com/huggingface/accelerate/pull/2342FP8RecipeKwargs by @sudhakarsingh27 in https://github.com/huggingface/accelerate/pull/2355add_hook_to_module and remove_hook_from_module compatibility with fx.GraphModule by @fxmarty in https://github.com/huggingface/accelerate/pull/2369requires_grad to kwargs when registering empty parameters. by @BlackSamorez in https://github.com/huggingface/accelerate/pull/2376adapter_only option to save_fsdp_model and load_fsdp_model to only save/load PEFT weights by @AjayP13 in https://github.com/huggingface/accelerate/pull/2321split_batches by @izhx in https://github.com/huggingface/accelerate/pull/2344nproc_per_node in the multi gpu test by @faaany in https://github.com/huggingface/accelerate/pull/2422Accelerator to prepare models in eval mode for XPU&CPU by @faaany in https://github.com/huggingface/accelerate/pull/2426Full Changelog: https://github.com/huggingface/accelerate/compare/v0.26.1...v0.27.0
Raise error when using batches of different sizes with dispatch_batches=True by @SunMarc in https://github.com/huggingface/accelerate/pull/2325
dispatch_batches=True by @SunMarc in https://github.com/huggingface/accelerate/pull/2325Full Changelog: https://github.com/huggingface/accelerate/compare/v0.26.0...v0.26.1
This release adds support for the MS-AMP (Microsoft Automatic Mixed Precision Library) into Accelerate as an alternative backend for doing FP8 trainin
This release adds support for the MS-AMP (Microsoft Automatic Mixed Precision Library) into Accelerate as an alternative backend for doing FP8 training on appropriate hardware. It is the default backend of choice. Read more in the docs here. Introduced in https://github.com/huggingface/accelerate/pull/2232 by @muellerzr
In the prior release a new sampler for the DataLoader was introduced that while across seeds does not show statistical differences in the results, repeating the same seed would result in a different end-accuracy that was scary to some users. We have now disabled this behavior by default as it required some additional setup, and brought back the original implementation. To have the new sampling technique (which can provide more accurate repeated results) pass use_seedable_sampler=True to the Accelerator. We will be propagating this up to the Trainer soon.
device_map we've made it possible to not returned grouped key results if desired in https://github.com/huggingface/accelerate/pull/2233device_map="cuda" etc thanks to @younesbelkada in https://github.com/huggingface/accelerate/pull/2254Many improvements to the docs have been made thanks to @stass. Along with this we've made it easier to adjust the config for the sharding strategy and other config values thanks to @pacman100 in https://github.com/huggingface/accelerate/pull/2288
A regression in Accelerate 0.23.0 occurred that showed learning is much slower on multi-GPU setups compared to a single GPU. https://github.com/huggingface/accelerate/pull/2304 has now fixed this thanks to @pacman100
The DeepSpeed integration now also handles auto values better when making a configuration in https://github.com/huggingface/accelerate/pull/2313
Params4bit added to bnb classes in set_module_tensor_to_device() by @poedator in https://github.com/huggingface/accelerate/pull/2315For developers, we've made it much easier to run the tests on different devices with no change to the code thanks to @statelesshz in https://github.com/huggingface/accelerate/pull/2123 and https://github.com/huggingface/accelerate/pull/2235
offload_state_dict=True and dtype is specified by @fxmarty in https://github.com/huggingface/accelerate/pull/2116auto values for comm buffers by @stas00 in https://github.com/huggingface/accelerate/pull/2295offload_state_dict=True and dtype is specified by @fxmarty in https://github.com/huggingface/accelerate/pull/2116Big-Modeling] Harmonize device check to handle corner cases by @younesbelkada in https://github.com/huggingface/accelerate/pull/2254log_images for aim tracker by @Justin900429 in https://github.com/huggingface/accelerate/pull/2257check_tied_parameters_on_same_device by @SunMarc in https://github.com/huggingface/accelerate/pull/2218auto values for comm buffers by @stas00 in https://github.com/huggingface/accelerate/pull/2295prepare_data_loader by @izhx in https://github.com/huggingface/accelerate/pull/2310Params4bit added to bnb classes in set_module_tensor_to_device() by @poedator in https://github.com/huggingface/accelerate/pull/2315Full Changelog: https://github.com/huggingface/accelerate/compare/v0.25.0...v0.26.0
Deprecated runner stuff by @muellerzr in https://github.com/huggingface/accelerate/pull/2152
As of this release, safetensors will be the default format saved when applicable! To read more about safetensors and why it's best to use it for safety (and not pickle/torch.save), check it out here
This release has two new experiment trackers, ClearML and DVCLive!
To use them, just pass clear_ml or dvclive to log_with in the Accelerator init. h/t to @eugen-ajechiloae-clearml and @dberenbaum
FSDP had a huge refactoring so that the interface when using FSDP is the exact same as every other scenario when using accelerate. No more needing to call accelerator.prepare() twice!
We now raise and try to disable P2P communications on consumer GPUs for the 3090 series and beyond. Without this users were seeing timeout issues and the like as NVIDIA dropped P2P support. If using accelerate launch we will automatically disable, and if we sense that it is still enabled on distributed setups using 3090's +, we will raise an error.
When doing .gather(), if tensors are on different devices we explicitly will raise an error (for now only valid on CUDA)
shuffle=True when using multiple GPUs and the new SeedableRandomSampler.save as False by @muellerzr in https://github.com/huggingface/accelerate/pull/2138launch, and pick up in state if a user will face issues. by @muellerzr in https://github.com/huggingface/accelerate/pull/2195Full Changelog: https://github.com/huggingface/accelerate/compare/v0.24.1...v0.25.0
Fixes https://github.com/huggingface/accelerate/issues/2091 by changing how checking for custom samplers is done
One critical issue with Accelerate is training runs were different when using an iterable dataset, no matter what seeds were set. v0.24.0 introduces t
One critical issue with Accelerate is training runs were different when using an iterable dataset, no matter what seeds were set. v0.24.0 introduces the dataloader.set_epoch() function to all Accelerate DataLoaders, where if the underlying dataset (or sampler) has the ability to set the epoch for reproducability it will do so. This is similar to the implementation already existing in transformers. To use:
dataloader = accelerator.prepare(dataloader)
# Say we want to resume at epoch/iteration 2
dataloader.set_epoch(2)
For more information see this PR, we will update the docs on a subsequent release with more information on this API.
save and save_state via the ProjectConfiguration dataclass. See #1953 for more info.bfloat16 mixed precision via torch.autocastall_gather_into_tensor is now used as the main gather operation, reducing memory in the cases of big tensorsdrop_last=True will now properly have the desired affect when performing Accelerator().gather_for_metrics()dispatch_model by @austinapatel in https://github.com/huggingface/accelerate/pull/1971save and save_state via ProjectConfiguration by @muellerzr in https://github.com/huggingface/accelerate/pull/1953torch.autocast for bfloat16 mixed precision by @brcps12 in https://github.com/huggingface/accelerate/pull/2033all_gather_into_tensor by @muellerzr in https://github.com/huggingface/accelerate/pull/1968gather_for_metrics by @muellerzr in https://github.com/huggingface/accelerate/pull/2048Full Changelog: https://github.com/huggingface/accelerate/compare/v0.23.0...v0.24.0
A new model estimation tool to help calculate how much memory is needed for inference has been added. This does not download the pretrained weights, a
A new model estimation tool to help calculate how much memory is needed for inference has been added. This does not download the pretrained weights, and utilizes init_empty_weights to stay memory efficient during the calculation.
Usage directions:
accelerate estimate-memory {model_name} --library {library_name} --dtypes fp16 int8
Or:
from accelerate.commands.estimate import estimate_command_parser, estimate_command, gather_data
parser = estimate_command_parser()
args = parser.parse_args(["bert-base-cased", "--dtypes", "float32"])
output = gather_data(args)
We've made the huggingface_hub library a first-class citizen of the framework! While this is mainly for the model estimation tool, this opens the doors for further integrations should they be wanted
Accelerator Enhancements:gather_for_metrics will now also de-dupe for non-tensor objects. See #1937mixed_precision="bf16" support on NPU devices. See #1949breakpoint API to help when dealing with trying to break from a condition on a single process. See #1940torch.compile support was fixed. See #1919gradient_accumulation_steps to "auto" in your deepspeed config, and Accelerate will use the one passed to Accelerator instead (#1901)accelerate config on npu by @statelesshz in https://github.com/huggingface/accelerate/pull/1895Tests] Finish all todos by @younesbelkada in https://github.com/huggingface/accelerate/pull/1957force_hooks to dispatch_model by @austinapatel in https://github.com/huggingface/accelerate/pull/1969Full Changelog: https://github.com/huggingface/accelerate/compare/v0.22.0...v0.23.0
A new framework has been introduced which can help catch timeout errors caused by distributed operations failing *before* they occur. As this adds a t
A new framework has been introduced which can help catch timeout errors caused by distributed operations failing before they occur. As this adds a tiny bit of overhead, it is an opt-in scenario. Simply run your code with ACCELERATE_DEBUG_MODE="1" to enable this. Read more in the docs, introduced via https://github.com/huggingface/accelerate/pull/1756
Accelerator.load_state can now load the most recent checkpoint automaticallyIf a ProjectConfiguration has been made, using accelerator.load_state() (without any arguments passed) can now automatically find and load the latest checkpoint used, introduced via https://github.com/huggingface/accelerate/pull/1741
In this release multiple new enhancements to distributed gradient accumulation have been added.
accelerator.accumulate() now supports passing in multiple models introduced via https://github.com/huggingface/accelerate/pull/1708.backward() via https://github.com/huggingface/accelerate/pull/1726DataLoaderDispatcher added via https://github.com/huggingface/accelerate/pull/1846no_sync by @NouamaneTazi in https://github.com/huggingface/accelerate/pull/1726get_scale() by patching the step method of optimizer. by @yuxinyuan in https://github.com/huggingface/accelerate/pull/1720set_module_tensor_to_device. by @Narsil in https://github.com/huggingface/accelerate/pull/1731__repr__ of AlignDevicesHook by @KacperWyrwal in https://github.com/huggingface/accelerate/pull/1735KwargsHandler.to_kwargs not working with os.environ initialization in __post_init__ by @CyCle1024 in https://github.com/huggingface/accelerate/pull/1738autocast kwargs and simplify autocast wrapper by @muellerzr in https://github.com/huggingface/accelerate/pull/1740Accelerator.save_state using multi-gpu by @CyCle1024 in https://github.com/huggingface/accelerate/pull/1760max_memory argument is in unexpected order by @ranchlai in https://github.com/huggingface/accelerate/pull/1759is_aim_available() function to not match aim >= 4.0.0 by @alberttorosyan in https://github.com/huggingface/accelerate/pull/1769load_fsdp_optimizer by @awgu in https://github.com/huggingface/accelerate/pull/1755torch.distributed is disabled by @natsukium in https://github.com/huggingface/accelerate/pull/1800get_balanced_memory to avoid OOM by @ranchlai in https://github.com/huggingface/accelerate/pull/1798convert_file_size_to_int by @ranchlai in https://github.com/huggingface/accelerate/pull/1799allow_val_change by @SumanthRH in https://github.com/huggingface/accelerate/pull/1796gather_for_metrics by @dleve123 in https://github.com/huggingface/accelerate/pull/1784load_and_quantize_model arg by @JonathanRayner in https://github.com/huggingface/accelerate/pull/1822init_on_device by @shingjan in https://github.com/huggingface/accelerate/pull/1826unwrap_model and keep_fp32_wrapper=False by @BenjaminBossan in https://github.com/huggingface/accelerate/pull/1838verify_device_map by @Rexhaif in https://github.com/huggingface/accelerate/pull/1842gpu_ids (Rel. Issue #1848) by @devymex in https://github.com/huggingface/accelerate/pull/1850fsdp_with_peak_mem_tracking.py by @pacman100 in https://github.com/huggingface/accelerate/pull/1856init_on_device by @shingjan in https://github.com/huggingface/accelerate/pull/1852DataLoaderDispatcher by @thevasudevgupta in https://github.com/huggingface/accelerate/pull/1846The following contributors have made significant changes to the library over the last release:
Accelerator.accumulate() (#1708)no_sync (#1726)DataLoaderDispatcher (#1846)Full Changelog: https://github.com/huggingface/accelerate/compare/v0.21.0...v0.22.0
You can now quantize any model (no just Transformer models) using Accelerate. This is mainly for models having a lot of linear layers. See the documen
You can now quantize any model (no just Transformer models) using Accelerate. This is mainly for models having a lot of linear layers. See the documentation for more information!
Accelerate now supports Ascend NPUs.
Accelerate now requires Python 3.8+ and PyTorch 1.10+ :
🚨🚨🚨 Spring cleaning: Python 3.8 🚨🚨🚨 by @muellerzr in #1661
🚨🚨🚨 Spring cleaning: PyTorch 1.10 🚨🚨🚨 by @muellerzr in #1662
[doc build] Use secrets by @mishig25 in #1551
Update launch.mdx by @LiamSwayne in #1553
Avoid double wrapping of all accelerate.prepare objects by @muellerzr in #1555
Update README.md by @LiamSwayne in #1556
Fix load_state_dict when there is one device and disk by @sgugger in #1557
Fix tests not being ran on multi-GPU nightly by @muellerzr in #1558
fix the typo when setting the "_accelerator_prepared" attribute by @Yura52 in #1560
[core] Fix possibility to passNoneType objects in prepare by @younesbelkada in #1561
Reset dataloader end_of_datalaoder at each iter by @sgugger in #1562
Update big_modeling.mdx by @LiamSwayne in #1564
[bnb] Fix failing int8 tests by @younesbelkada in #1567
Update gradient sync docs to reflect importance of optimizer.step() by @dleve123 in #1565
Update mixed precision integrations in README by @sgugger in #1569
Raise error instead of warn by @muellerzr in #1568
Introduce listify, fix tensorboard silently failing by @muellerzr in #1570
Check for bak and expand docs on directory structure by @muellerzr in #1571
Perminant solution by @muellerzr in #1577
fix the bug in xpu by @mingxiaoh in #1508
Make sure that we only set is_accelerator_prepared on items accelerate actually prepares by @muellerzr in #1578
Expand prepare() doc by @muellerzr in #1580
Get Torch version using importlib instead of pkg_resources by @catwell in #1585
improve oob performance when use mpirun to start DDP finetune without accelerate launch by @sywangyi in #1575
Update training_tpu.mdx by @LiamSwayne in #1582
Return false if CUDA available by @muellerzr in #1581
fix logger level by @caopulan in #1579
Fix test by @muellerzr in #1586
Update checkpoint.mdx by @LiamSwayne in #1587
FSDP updates by @pacman100 in #1576
Update modeling.py by @ain-soph in #1595
Integration tests by @muellerzr in #1593
Add triggers for CI workflow by @muellerzr in #1597
Remove asking xpu plugin for non xpu devices by @abhilash1910 in #1594
Remove GPU safetensors env variable by @sgugger in #1603
reset end_of_dataloader for dataloader_dispatcher by @megavaz in #1609
fix for arc gpus by @abhilash1910 in #1615
Ignore low_zero option when only device is available by @sgugger in #1617
Fix failing multinode tests by @muellerzr in #1616
Doc to md by @sgugger in #1618
Fix tb issue by @muellerzr in #1623
Fix workflow by @muellerzr in #1625
Fix transformers sync bug with accumulate by @muellerzr in #1624
fixes offload dtype by @SunMarc in #1631
fix: Megatron is not installed. please build it from source. by @yuanwu2017 in #1636
deepspeed z2/z1 state_dict bloating fix by @pacman100 in #1638
Swap disable rich by @muellerzr in #1640
fix autocasting bug by @pacman100 in #1637
fix modeling low zero by @abhilash1910 in #1634
Add skorch to runners by @muellerzr in #1646
add save model by @SunMarc in #1641
Change dispatch_model when we have only one device by @SunMarc in #1648
Doc save model by @SunMarc in #1650
Fix device_map by @SunMarc in #1651
Check for port usage before launch by @muellerzr in #1656
[BigModeling] Add missing check for quantized models by @younesbelkada in #1652
Bump integration by @muellerzr in #1658
TIL by @muellerzr in #1657
docker cpu py version by @muellerzr in #1659
[BigModeling] Final fix for dispatch int8 and fp4 models by @younesbelkada in #1660
remove safetensor dep on shard_checkpoint by @SunMarc in #1664
change the import place to avoid import error by @pacman100 in #1653
Update broken Runhouse link in examples/README.md by @dongreenberg in #1668
Bnb quantization by @SunMarc in #1626
replace save funct in doc by @SunMarc in #1672
Doc big model inference by @SunMarc in #1670
Add docs for saving Transformers models by @deppen8 in #1671
fix bnb tests by @SunMarc in #1679
Fix workflow CI by @muellerzr in #1690
remove duplicate class by @SunMarc in #1691
update readme in examples by @statelesshz in #1678
Fix nightly tests by @muellerzr in #1696
Fixup docs by @muellerzr in #1697
Improve quality errors by @muellerzr in #1698
Move mixed precision wrapping ahead of DDP/FSDP wrapping by @ChenWu98 in #1682
Add offload for 8-bit model by @SunMarc in #1699
Deepcopy on Accelerator to return self by @muellerzr in #1694
Update tracking.md by @stevhliu in #1702
Skip tests when bnb isn't available by @muellerzr in #1706
Fix launcher validation by @abhilash1910 in #1705
Fixes for issue #1683: failed to run accelerate config in colab by @Erickrus in #1692
Fix the bug where DataLoaderDispatcher gets stuck in an infinite wait when the dataset is an IterDataPipe during multi-process training. by @yuxinyuan in #1709
add multi_gpu decorator by @SunMarc in #1712
Modify loading checkpoint behavior by @SunMarc in #1715
fix version by @SunMarc in #1701
Keep old behavior by @muellerzr in #1716
Optimize get_scale to reduce async calls by @muellerzr in #1718
Remove duplicate code by @muellerzr in #1717
New tactic by @muellerzr in #1719
add Comfy-UI by @pacman100 in #1723
add compatibility with peft by @SunMarc in #1725
The following contributors have made significant changes to the library over the last release:
Reset dataloader end_of_datalaoder at each iter in #1562 by @sgugger
fix the typo when setting the "_accelerator_prepared" attribute in #1560 by @Yura52
NoneType objects in prepare in #1561 by @younesbelkadaAvoid double wrapping of all accelerate.prepare objects by @muellerzr in #1555
logging_dir has been fully deprecated, please use project_dir or a Project_configuration
Support has been added to run device_map="auto" on the MPS device. Big model inference also work with models loaded in 4 bits in Transformers.
This version introduces a new Accelerator.split_between_processes utility to help with performing distributed infernece with non-tensorized or non-dataloader workflows. Read more here
LocalSGDlogging_dir has been fully deprecated, please use project_dir or a Project_configuration
core] Introducing CustomDtype enum for custom dtypes by @younesbelkada in #1434state.rank -> process_index by @pcuenca in #1450in_order argument that defaults to False, to log in order. by @JulesGM in #1262register_empty_buffer to match torch args by @NouamaneTazi in #1465split_between_processes by @muellerzr in #1477bnb] Add fp4 support for dispatch by @younesbelkada in #1505The following contributors have made significant changes to the library over the last release:
fix: typing issues, and replace deprecated python typing (Optional, Union) to | by @kiyoon in https://github.com/huggingface/accelerate/pull/1363
Trainer, keep an eye on the repos to see how our progress is coming along!wandb integration now supports logging of images and tables through tracker.log_images and tracker.log_tables respectivelycore] Add Quantization support for dispatch_model by @younesbelkada in https://github.com/huggingface/accelerate/pull/1237has_transfomer_engine_layers by @muellerzr in https://github.com/huggingface/accelerate/pull/1283recursively_apply by @muellerzr in https://github.com/huggingface/accelerate/pull/1286bnb] fix bnb slow test by @younesbelkada in https://github.com/huggingface/accelerate/pull/1292notebook_launcher by @muellerzr in https://github.com/huggingface/accelerate/pull/1293bnb] Fix bnb slow test by @younesbelkada in https://github.com/huggingface/accelerate/pull/1355| by @kiyoon in https://github.com/huggingface/accelerate/pull/1363accelerate env reporting by @muellerzr in https://github.com/huggingface/accelerate/pull/1376Full Changelog: https://github.com/huggingface/accelerate/compare/v0.18.0...v0.19.0
A new GradientAccumulationPlugin has been added to handle more configurations with the GradientState. Specifically you can optionally disable having A
GradientAccumulationPlugin has been added to handle more configurations with the GradientState. Specifically you can optionally disable having Accelerate automatically adjust the length of the scheduler relative to gradient accumulation steps through it. Otherwise Accelerate will now automatically handle ensuring that the schedulers built for non-gradient accumulation will work during gradient accumulationdynamo_backend warning has been silenced.drop_last on linear layers, tied weight loading, and handling of multiple tied parametersfind_tied_parameters now deals with groups of tied parameters (instead of only pairs of them). As a result it now returns a list of list of strings instead of a dictionary.use_orig_params to FullyShardedDataParallelPlugin by @pacman100 in https://github.com/huggingface/accelerate/pull/1184to on modules that wraps accelerate loaded models by @younesbelkada in https://github.com/huggingface/accelerate/pull/1172Full Changelog: https://github.com/huggingface/accelerate/compare/v0.17.1...v0.18.0
Fix CPU error always being raised by @muellerzr in #1175
Flag for deprecation by @muellerzr in #1061
This release fully supports the upcoming PyTorch 2.0 release. You can choose to use torch.compile or not and then customize the options in accelerate.config or via a TorchDynamoPlugin.
This release adds a new PartialState, which contains most of the capabilities of the AcceleratorState however it is designed to be used by the user to assist in any process control mechanisms around it. With this, users also now do not need to have if accelerator.state.is_main_process when utilizing classes such as the Tracking API, as these now will automatically use only the main process for their work by default.
Launching from TPU pods is now supported, please see this issue for more information
accelerate launch by @muellerzr in #1049This release adds experimental support for FP8 mixed precision training, which requires the transformer-engine library as well as a Hopper GPU (or higher).
mps device by default and removing related config by @pacman100 in #1030cpu_offload_with_hook by @sgugger in #1045hidden_size auto value default fixes by @pacman100 in #1060PartialState first class citizen by @muellerzr in #1071additional_args, allowing more flexible configuration and env variable support by @dbpprt in #1113launch for greater extensibility by @Yard1 in #1123torch.distributed module by @mfuntowicz in #1108Accelerator] Fix issue with 8bit models by @younesbelkada in #1155total_limit by @muellerzr in #1165The following contributors have made significant changes to the library over the last release:
launch for greater extensibility (#1123)🚨🚨🚨 Act on deprecations 🚨🚨🚨 by @muellerzr in #917
A new interactive tool has been introduced to the documentation to help users quickly learn how to utilize features of the framework before providing more details on them as shown below:
Not only does it provide a code diff, but it also includes an explanation and links to more resources the user should check out to learn more:
Try it out today in the docs
When resuming training, you can more efficiently skip batches in your dataloader with the new skip_first_batches function (also available as a method on your Accelerator).
A new ZeRO-3 init context manager is added to provide granular control to users in situations involving nested/multiple models. Refactoring of DeepSpeed Config file support to remove ambiguity between it and Accelerate config.
Adding support for auto entries in the DeeSpeed config file to be filled via the accelerate launch command. Try it out today by referring to the section Things to note when using DeepSpeed Config File
deepspeed_config_file by @pacman100 in #941project_dir and limit the number of saved checkpoints by @muellerzr in #916init_on_device by @thomasw21 in #926ACCELERATE_ by @pacman100 in #928load_checkpoint by @sgugger in #920mixed_precision_type property to AcceleratorState by @pacman100 in #935deepspeed_config_file by @pacman100 in #941load_state by @pacman100 in #989Your coding agent can read these notes before it upgrades. Set up the MCP server →