NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1729 most downloaded on PyPI
State-of-the-art diffusion in PyTorch.
Last release 1 months ago
20 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
99 releases · first in 2022
One column per quarter.
🚨 Breaking changes and deprecations
Tip
This release features several new pipelines, including LTX2.5, MiniMax H3, and Wan Animate 2. We're also graduating Modular Diffusers out of the experimental phase and announcing its stable support. Additionally, this release includes minimal support for tensor-parallel. There's a lot more that went down in this release. So, please consult the notes for details.
MiniMax-H3 generates video and its soundtrack together. A single transformer denoises one packed sequence containing the text conditioning, the conditioning media, and the target video and audio latents — there is no separate vocoder and no post-hoc audio pass. Its conditioner is a Qwen3VLForConditionalGeneration whose unnormalized 50th-decoder-layer hidden state is read instead of the last one.
MiniMax-H3 is integrated as Modular Diffusers blocks only — MiniMaxH3Blocks and their MiniMaxH3ModularPipeline are the whole integration. The conversion ships both checkpoint partitions in one repository and exposes three workflows (t2va, fl2va, ref2va) that can be pruned at from_pretrained time so only that task's components are declared and downloaded.
MiniMax Music 3 produces complete songs up to five minutes long from lyrics and a music description, with expressive vocals and long-range structure. It is a hybrid of an autoregressive and a diffusion stage: an 8B Qwen3-based global language model predicts one semantic audio token per frame while a small depth decoder fills in seven residual RVQ codebooks, and their fused hidden states condition a 2.4B flow-matching transformer that produces Flow-VAE latents in overlapping chunks. A DAC-style decoder turns the latents into 44.1 kHz stereo audio.
Stable Audio 3 is a text-to-audio model from Stability AI that generates high-quality stereo audio at 44.1 kHz. It uses a rectified-flow DiT conditioned on a frozen T5Gemma text encoder (via cross-attention) and on duration (a float embedded by StableAudio3DurationEmbedder and used for adaptive layer norm), and decodes with the SAME (Semantically-Aligned Music Encoder) autoencoder, AutoencoderSAME.
Three pipelines ship: StableAudio3Pipeline, StableAudio3AudioToAudioPipeline, and StableAudio3InpaintPipeline.
Thanks to @buffett0323 for the contribution (#14119).
LTX-2.5 reuses the existing LTX2Pipeline / LTX2VideoTransformer3DModel / AutoencoderKLLTX2Video classes — there is no separate pipeline class. The user-visible difference is the text encoder: LTX-2.5 is paired with a Gemma 4 (gemma4_unified) checkpoint, loaded automatically from a converted LTX-2.5 repo.
Lightricks/LTX-2.5-Diffusers ships both the distilled DiT (transformer/) and the full/SFT DiT (transformer_full/), plus everything two-stage generation needs. Alongside the checkpoint support, this release adds:
LTX2VideoDiffusionDecoderModel and LTX2VideoDiffusionDecodePipeline — a second video decoder over the same latent space, so latents are interchangeable between decoders.duration_head that predicts shot length from the text-connector output, so num_frames is auto-predicted by default when the loaded pipeline has one.google/gemma-4-E2B-it checkpoint (enable_prompt_enhancement=True).LTX25AutoBlocks for Modular Diffusers (#14453).Wan-Animate-2 by the Alibaba Wan Team animates a reference character image with the motion of a driving video. The driving video is processed in fixed-length segments: each segment runs a reference-extraction pass that caches the driving segment's K/V in every transformer layer, denoises against that cache, and is decoded inside the loop, because the next segment conditions on the previous segment's decoded tail frames.
Two presets are available — the base checkpoint samples with classifier-free guidance, and the distilled checkpoint samples in few steps without it. Guidance is owned by the pipeline's guider component, so there is no guidance_scale argument.
Thanks to @kelseyee for authoring the integration (#14413).
JoyAI-Image-Edit-Plus extends the JoyAI-Image family (an 8B MLLM paired with a 16B MMDiT) to multi-image instruction-guided editing. It accepts 1–5 reference images plus a text instruction and composes elements from the references into a new image.
Thanks to @tangyanf for the contribution (#14032).
Cosmos 3 landed in 0.39.0 and gets substantially more coverage in this release:
Cosmos3OmniModularPipeline), with Transfer support for precomputed control videos (edge, blur, depth, segmentation, world-scenario maps) generated autoregressively in chunks and stitched automatically.Cosmos3DistilledModularPipeline.Thanks to @yzhautouskay and @atharvajoshi10 for the contributions.
Modular Diffusers is no longer marked experimental (#14525) — the API warning has been dropped.
DiffusionPipeline half.ConditionalPipelineBlocks declare different defaults for the same input, combine_inputs now merges the default to None and records the per-block defaults in a new InputParam.defaults_by_block field. get_block_state falls back to the block's own declared default, so each branch resolves its own default when it actually runs. Docstrings render conflicted defaults as e.g. "defaults to None or 189, depending on the workflow".intermediate_inputs were cleaned up, stale auto-docstrings are now detected in CI, and Mellon custom-block required-input handling was fixed.Important
Please try out Modular Diffusers and let us know about your feedback!
Tensor-parallel inference is now supported for model inference on CUDA and AWS Neuron (Trainium/Inferentia), exposed through the same public API already used for context parallelism:
from diffusers import TensorParallelConfig
pipe.transformer.enable_parallelism(config=TensorParallelConfig(mesh=tp_mesh))Sharding is model-agnostic and driven from a flat _tp_plan, which has been added to the FLUX.1, FLUX.2 and Qwen-Image DiTs. Check out the docs for more details.
kernels-community/aiter-flash-attn-ck Hub kernel, dropping the aiter dependency. Thanks to @Abdennacer-Badaoui._flash_3_varlen_hub backend and a mask-handling fix. Thanks to @zhtmike.DiffusionPipeline.device deduction for split-device pipelinesThe diffusers-cli was reworked for agentic use (#13966) and then cleaned up (#14381): modular_model_index.json is written when saving a custom block so ModularPipeline.from_pretrained can load and run custom blocks as pipelines, auto CPU offload works for Modular pipelines, workflow can be passed to Modular pipelines, and output saving handles multimodal output (e.g. LTX video frames + audio) and batched video.
Skills are now installed through the CLI rather than the Makefile (#14454):
diffusers-cli skills list
diffusers-cli skills add <skill name>Flax* classes and the flax extras are gone.get_peft_kwargs took lora_alpha from the first entry of the rank dict and never revisited it, so every module whose rank differed from the first key's rank got an arbitrary, key-order-dependent scale. Ranks are now mirrored into the alphas when a checkpoint brings no alpha information (the diffusers/PEFT convention: alpha == rank, scale 1.0). Adapters with a declared alpha keep it, and uniform-rank adapters are unaffected. Existing mixed-rank, no-alpha LoRAs will now produce different (correct) results.torch_dtype is deprecated in favour of dtype (#14205, #14313), following transformers. torch_dtype still works but warns, and will be removed in 1.0.0. A torch.dtype alias was added for the docs.dduf_file warns and will be removed in 0.41.0.SECURITY.md was added.different_shapes_for_compilation. Thanks to @jiqing-feng._merged_adapters when it is unfused from all components.visual_cond channels.WanTransformer3DModel and SD3Transformer2DModel hidden states contiguoussnapshot_download with the latest huggingface_hubA large chunk of this release is test modernization: pipeline tests continue migrating to the new mixin structure (Wan, Qwen-Image, FLUX.2, CogVideoX, Stable Diffusion, and the LoRA pipeline tests), tests/others, training tests, and attention-processor tests moved to pytest, model-level and pipeline-level quantization tests were standardized, and an output_shape property was introduced in the pipeline tests. The agent-facing docs and skills under .ai/ were expanded to cover tests, model implementation, and blockset conventions.
_flash_3_varlen_hub mask handling by @zhtmike in #14115_flash_3_varlen_hub backend by @zhtmike in #13809SD3Transformer2DModel hidden states contiguous by @menglcai in #14186kwargs_type input/output by @yiyixuxu in #14157WanTransformer3DModel hidden states contiguous before the block loop by @menglcai in #14236torch_dtype and prefer dtype following transformers. by @sayakpaul in #14205diffusers-cli for agentic use by @DN6 in #13966Note truncated.
Diffusers 0.40.0: New pipelines, tensor-parallel support, improved CLI, and more Latest
Latest
Compare
Cosmos 3 is NVIDIA's unified world foundation model (WFM) for Physical AI — a single omni-model built on a Mixture-of-Transformers (MoT) architecture
Cosmos 3 is NVIDIA's unified world foundation model (WFM) for Physical AI — a single omni-model built on a Mixture-of-Transformers (MoT) architecture that combines world generation, physical reasoning, and action generation, replacing the separate Predict, Reason, and Transfer models from earlier Cosmos releases. A single Cosmos3OmniTransformer runs a Qwen-style language model in parallel with a diffusion generation pathway, joined by a 3D multimodal RoPE. This release also lands video-to-video and action-conditioned generation, and a sound encoder.
Thanks to @atharvajoshi10, @yzhautouskay, and @MaciejBalaNV for the contributions.
Ideogram 4 is a flow-matching text-to-image model that uses a multimodal text encoder and an asymmetric classifier-free guidance scheme: a dedicated unconditional_transformer produces the negative branch with zeroed text features, while the main transformer consumes the full packed text + image sequence. The pipeline ships with structured prompt upsampling and LoRA loading support.
Thanks to @JinLiIdeogram for the contribution.
Krea 2 (K2) is a flow-matching text-to-image model built around a single-stream MMDiT with grouped-query attention. A Qwen3-VL text encoder provides the conditioning — hidden states from twelve decoder layers are tapped per token and fused inside the transformer by a small text-fusion stage — and images are decoded with the Qwen-Image VAE. Both the base (midtrain) and TDM (distilled, few-step) checkpoints are supported, alongside a LoRA DreamBooth trainer.
Thanks to @EleaZhong and @Abhinay1997 for the contribution.
DreamLite is a text-to-image and image-editing model from ByteDance. It pairs a custom 2D U-Net (DreamLiteUNetModel) with the Qwen3-VL multimodal encoder as its prompt / image-instruction encoder, and uses an AutoencoderTiny (TAESD-style) VAE for fast latent encode/decode. A distilled DreamLiteMobilePipeline targets on-device, low-latency generation.
Thanks to @Carlofkl for the contribution.
PRXPixel is a pixel-space text-to-image generation model by Photoroom. A ~7B PRXTransformer2DModel denoises raw RGB images directly — no VAE is needed. The model is conditioned on a Qwen3-VL text encoder and uses flow matching where the transformer predicts the clean image at each step (x-prediction).
Thanks to @DavidBert for the contribution.
Motif-Video is a 2B parameter diffusion transformer for text-to-video and image-to-video generation. It features a three-stage architecture (12 dual-stream + 16 single-stream + 8 DDT decoder layers), Shared Cross-Attention for stable text-video alignment over long sequences, a T5Gemma2 text encoder, and rectified flow matching for velocity prediction.
Thanks to @waitingcheung for the contribution.
AnyFlow from NVIDIA, NUS, and MIT is the first any-step video diffusion framework built on flow maps, enabling a single model (bidirectional or causal) to adapt to arbitrary inference budgets. It ships both bidirectional and FAR causal pipelines built on Wan2.1 backbones, covering text-to-video, image-to-video, and video-to-video.
Thanks to @Enderfga for the contribution.
JoyAI-Image is a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing. It combines an 8B Multimodal LLM with a 16B Multimodal Diffusion Transformer (MMDiT). JoyImageEditPipeline supports general image editing as well as spatial editing capabilities including object move, object rotation, and camera control.
Thanks to @Moran232 for the contribution.
DiffusionGemma is a block-diffusion encoder-decoder language model. A causal encoder reads the clean prompt (and any previously generated blocks) into a KV cache, and a bidirectional decoder denoises a fixed-size "canvas" of tokens by cross-attending to that cache, committing the most confident tokens via the new BlockRefinementScheduler. The released checkpoint is google/diffusiongemma-26B-A4B-it.
Anima is a 2 billion parameter text-to-image model created via a collaboration between CircleStone Labs and Comfy Org. It is focused mainly on anime concepts, characters, and styles, but is also capable of generating a wide variety of other non-photorealistic content.
It reuses the CosmosTransformer3DModel with a Qwen3 text encoder, a T5-token text conditioner, and the AutoencoderKLQwenImage VAE.
Thanks to @rmatif for the contribution.
New LTX2InContextPipeline (in-context LoRA) and LTX2HDRPipeline extend the LTX-2 family with in-context conditioning and HDR video generation.
ErnieImageModularPipeline (#13948) and Ideogram4ModularPipeline (#13980), thanks to @SamuelTallet._dequantize for the TorchAO quantizerAutoPipelineForText2Audiotorch.compile compatibilitysafetensors to 0.8.0torch version is now 2.6txt_seq_lens from qwen transformer. by @sayakpaul in #13674torch.distributed by @hlky in #13673flash_varlen_hub backend by @zhtmike in #13479transformers from main for doc and staging by @sayakpaul in #13723modules_to_not_convert / keep_in_fp32_modules by @dg845 in #13697weighting chunk when using prior preservation in Flux and SD3 LoRA training by @Dev-X25874 in #13743Note truncated.
Diffusers 0.39.0: New image and video pipelines, core library improvements, and more
Compare
[docs] deprecate pipelines by @stevhliu in #13157
LLaDA2 is a family of discrete diffusion language models that generate text through block-wise iterative refinement. Instead of autoregressive token-by-token generation, LLaDA2 starts with a fully masked sequence and progressively unmasks tokens by confidence over multiple refinement steps.
NucleusMoE-Image is a 2B active 17B parameter model trained with efficiency at its core. Our novel architecture highlights the scalability of a sparse MoE architecture for Image generation.
Thanks to @sippycoder for the contribution.
ERNIE-Image is a powerful and highly efficient image generation model with 8B parameters.
Thanks to @HsiaWinter for the contribution.
LongCat-AudioDiT is a text-to-audio diffusion model from Meituan LongCat.
Thanks to @RuixiangMa for the contribution.
ACE-Step 1.5 generates variable-length stereo audio at 48 kHz (10 seconds to 10 minutes) from text prompts and optional lyrics. The full system pairs a Language Model planner with a Diffusion Transformer (DiT) synthesizer; this pipeline wraps the DiT half of that stack, and consists of three components: an AutoencoderOobleck VAE that compresses waveforms into 25 Hz stereo latents, a Qwen3-based text encoder for prompt and lyric conditioning, and an AceStepTransformer1DModel DiT that operates in the VAE latent space using flow matching.
Thanks to @ChuxiJ for the contribution.
Make your Flux.2 decoding faster with this new small decoder model from the Black Forest Labs. You can check it out here. It was contributed by @huemin-art in this PR.
We added modular support for LTX-2 and Hunyuan 1.5.
ring_anything as a new CP backendlru_cache warnings during torch.compile by @jiqing-feng in #13384--with_prior_preservation by @chenyangzhu1 in #133960.8.0-rc.0 by @McPatate in #13470trust_remote_code by @hlky in #13448The following contributors have made significant changes to the library over the last release:
Note truncated.
Diffusers 0.38.0: New image and audio pipelines, Core library improvements, and more
Compare
Fix for loading ModularPipelines with AutoModel type hints in their modular_model_index.json #13271
ModularPipelines with AutoModel type hints in their modular_model_index.json #13271torchvision import in Cosmos Predict 2.5 #13321Your coding agent can read these notes before it upgrades. Set up the MCP server →