NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4317 most downloaded on PyPI
DeepSpeed library
Last release 17 days ago
16 Sep 2026
Ships on a steady schedule
a new release about every 3 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
6 years old
132 releases · first in 2020
One column per quarter.
Fix bug in bfloat16 optimizer related to checkpointing by @okoge-kaz in https://github.com/microsoft/DeepSpeed/pull/4434
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.11.0...v0.11.1
DeepSpeed-VisualChat: Improve Your Chat Experience with Multi-Round Multi-Image Inputs [English] [中文] [日本語]
set_to_none=true in zero_grad methods by @Jackmin801 in https://github.com/microsoft/DeepSpeed/pull/4438ignore_unused_parameters by @UniverseFly in https://github.com/microsoft/DeepSpeed/pull/4418Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.10.3...v0.11.0
ZeRO-Inference: 20X faster inference through weight quantization and KV cache offloading
non_reentrant_checkpoint fix requires_grad of input must be true for activation checkpoint layer in pipeline train. by @inkcherry in https://github.com/microsoft/DeepSpeed/pull/4224Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.10.2...v0.10.3
MP ZeRO++ by @HeyangQin in https://github.com/microsoft/DeepSpeed/pull/3954
setup.py by @loadams in https://github.com/microsoft/DeepSpeed/pull/4185requires_grad=True by @XuehaiPan in https://github.com/microsoft/DeepSpeed/pull/4138non_reentrant_checkpoint to support inputs require no grad by @hughpu in https://github.com/microsoft/DeepSpeed/pull/4118Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.10.1...v0.10.2
[docs] add zero++ paper link by @jeffra in https://github.com/microsoft/DeepSpeed/pull/3974
DeepSpeedHybridEngine.generate() by @XuehaiPan in https://github.com/microsoft/DeepSpeed/pull/4026load_state_dir non-strict-mode work by @hughpu in https://github.com/microsoft/DeepSpeed/pull/4020# punct in the second sed command by @hughpu in https://github.com/microsoft/DeepSpeed/pull/4061Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.10.0...v0.10.1
[Docs] chrome://tracing is deprecated by @keyboardAnt in https://github.com/microsoft/DeepSpeed/pull/3805
chrome://tracing is deprecated by @keyboardAnt in https://github.com/microsoft/DeepSpeed/pull/3805CHECK_CUDA by @Flamefire in https://github.com/microsoft/DeepSpeed/pull/3854Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.4...v0.10.0
Documentation for DeepSpeed Accelerator Abstraction Interface by @delock in https://github.com/microsoft/DeepSpeed/pull/3184
Flops Profiler to test model.generate() by @CaffreyR in https://github.com/microsoft/DeepSpeed/pull/2515Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.4...v0.9.5
[MiCS] [Fix] saving and loading model checkpoint logic for MiCS sharding by @zarzen in https://github.com/microsoft/DeepSpeed/pull/3440
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.3...v0.9.4
Enable auto TP policy for llama model by @jianan-gu in https://github.com/microsoft/DeepSpeed/pull/3170
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.2...v0.9.3
MiCS implementation by @zarzen in https://github.com/microsoft/DeepSpeed/pull/2964
PipelineEngine.eval_batch result by @nrailgun in https://github.com/microsoft/DeepSpeed/pull/3316Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.1...v0.9.2
Update DS-Chat docs for v0.9.0 by @mrwyattii in https://github.com/microsoft/DeepSpeed/pull/3216
torch.cuda.is_available() check when compiling ops by @jinzhen-lin in https://github.com/microsoft/DeepSpeed/pull/3085Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.9.0...v0.9.1
♻️ replace deprecated functions for communication by @mayank31398 in https://github.com/microsoft/DeepSpeed/pull/2995
pass calls where they aren't needed by @stas00 in https://github.com/microsoft/DeepSpeed/pull/2826nv-transformers-v100 - use the same torch version as transformers CI by @stas00 in https://github.com/microsoft/DeepSpeed/pull/3096NamedTuple when sharding parameters [#3029] by @AlexanderVanEck in https://github.com/microsoft/DeepSpeed/pull/3037Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.8.3...v0.9.0
[deepspeed/autotuner] Bug fix for skipping mbs on gas by @rahilbathwal5 in https://github.com/microsoft/DeepSpeed/pull/2171
logger.warning_once by @stas00 in https://github.com/microsoft/DeepSpeed/pull/3021Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.8.2...v0.8.3
Remove deprecated torch._six imports by @yasyf in https://github.com/microsoft/DeepSpeed/pull/2863
torch._six imports by @yasyf in https://github.com/microsoft/DeepSpeed/pull/2863vendor_id_raw is not provided by @FarzanT in https://github.com/microsoft/DeepSpeed/pull/2836AttributeError in #2853 by @saforem2 in https://github.com/microsoft/DeepSpeed/pull/2854Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.8.1...v0.8.2
CUDA optional deepspeed ops by @tjruwase in https://github.com/microsoft/DeepSpeed/pull/2507
packaging requirement by @carmocca in https://github.com/microsoft/DeepSpeed/pull/2771Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.8.0...v0.8.1
DeepSpeed Data Efficiency: A composable library that makes better use of data, increases training efficiency, and improves model quality
min_loss_scale default by @stas00 in https://github.com/microsoft/DeepSpeed/pull/2660initial_scale_power to 16 by @stas00 in https://github.com/microsoft/DeepSpeed/pull/2663Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.7...v0.8.0
Update the locator for Megatron-LM by @rapsealk in https://github.com/microsoft/DeepSpeed/pull/2564
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.6...v0.7.7
DeepSpeed inference config. (#2459) by @awan-10 in https://github.com/microsoft/DeepSpeed/pull/2472
init_inference() by @aphedges in https://github.com/microsoft/DeepSpeed/pull/2540Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.5...v0.7.6
Fix Bug #2319 by @jomayeri in https://github.com/microsoft/DeepSpeed/pull/2438
scale_attn_by_inverse_layer_idx feature by @hyunwoongko in https://github.com/microsoft/DeepSpeed/pull/2486Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.4...v0.7.5
MOE residual matmult unit test by @samadejacobs in https://github.com/microsoft/DeepSpeed/pull/2323
gptj_residual_add kernels for better readability by @arashb in https://github.com/microsoft/DeepSpeed/pull/2358fused_bias_residual kernels for better readability by @arashb in https://github.com/microsoft/DeepSpeed/pull/2356Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.3...v0.7.4
Add blob storage to CI runners by @mrwyattii in https://github.com/microsoft/DeepSpeed/pull/2260
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.2...v0.7.3
Enable contiguous gradients with Z1+MoE by @siddharth9820 in https://github.com/microsoft/DeepSpeed/pull/2250
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.1...v0.7.2
Fix for distributed tests on pytorch>=1.12 by @mrwyattii in https://github.com/microsoft/DeepSpeed/pull/2141
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.7.0...v0.7.1
DeepSpeed Compression: https://www.microsoft.com/en-us/research/blog/deepspeed-compression-a-composable-library-for-extreme-compression-and-zero-cost-
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.7...v0.7.0
Add Inference support for running the BigScience-BLOOM Architecture by @RezaYazdaniAminabadi in https://github.com/microsoft/DeepSpeed/pull/2083
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.6...v0.6.7
[docs] add 530b paper by @jeffra in https://github.com/microsoft/DeepSpeed/pull/1979
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.5...v0.6.6
GatheredParameters - accept a tuple of params by @stas00 in https://github.com/microsoft/DeepSpeed/pull/1941
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.4...v0.6.5
[fix] Windows installs cannot import fcntl by @mrwyattii in https://github.com/microsoft/DeepSpeed/pull/1921
Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.3...v0.6.4
Fix setup.py crash when torch is not installed. by @PaperclipBadger in https://github.com/microsoft/DeepSpeed/pull/1866
-lcurand to solve undefined symbol: curandCreateGenerator by @stas00 in https://github.com/microsoft/DeepSpeed/pull/1879Full Changelog: https://github.com/microsoft/DeepSpeed/compare/v0.6.1...v0.6.3
Nothing published for this version
Advancing MoE inference and training to power next-generation AI scale
@stas00, @jithunnair-amd, @rraminen, @jeffdaily, @okakarpa, @jfc4050, @raamjad, @aphedges, @SeanNaren, @liamcli, @andriyor, @manuelciosici
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Mixture of Experts (MoE) support
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
[Press release] DeepSpeed: Accelerating large-scale model inference and training via system optimizations and compression
Full precision (fp32) support for ZeRO Stage2 and Stage3
Special thanks to our contributors: @stas00, @SeanNaren, @sdtblck, @wbuchwalter, @ghosthamlet, @zhujiangang,
Deprecated cpu_offload in config JSON, see JSON docs for more details.
ZeRO-Infinity release allowing nvme offload and more!
Deprecated cpu_offload in config JSON, see JSON docs for more details.
Automatic external parameter registration, more details in the ZeRO 3 docs.
Several bug fixes for ZeRO stage 3
Nothing published for this version
Combined release notes since Jan 12th v0.3.10 release
Combined release notes since Jan 12th v0.3.10 release
Nothing published for this version
Nothing published for this version
Deprecate client ability to disable gradient reduction #552
Combined release notes since November 12th v0.3.1 release
deepspeed.init_distributed API, #608, #645, #644@stas00, @gcooper-isi, @g-karthik, @sxjscience, @brettkoonce, @carefree0910, @Justin1904, @harrydrippin
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →