NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #274 most downloaded on PyPI
Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Last release 4 days ago
30 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
6 versions withdrawn
withdrawn after publishing
10 years old
243 releases · first in 2016
Deprecate min-max pixels ( #49021 ) by @zucchini-nlp
Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference, handles up to eight speakers, and orders speaker outputs by each speaker's first arrival in the input audio.
The model uses the Arrival-Order Speaker Cache (AOSC) 1 and FIFO queue introduced for Streaming Sortformer 1, 2. A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms. With chunked inference, the maximum audio duration is not limited.
Links: Documentation
NemotronH Omni is a multimodal reasoning model from NVIDIA that pairs the NemotronH hybrid
Mamba-Transformer language model with a RADIO vision encoder and an optional Parakeet-based sound encoder.
Image (and video) patches are projected through a RADIO tower and a pixel-shuffle MLP into the language model's
embedding space at the <image> / <video> context-token positions; audio clips are projected in the same way at
<audio> positions. The result is a single autoregressive model that reasons jointly over text, images, video and
sound.
Links: Documentation
HyperCLOVAX Vision V2 is a multimodal vision-language model developed by NAVER. It combines the HyperClovaX language model backbone with a Qwen2.5-VL vision encoder. The model supports text, image, and video inputs and is capable of chain-of-thought reasoning via built-in thinking tokens (<think>...</think>).
Links: Documentation
GTE was proposed in mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval by Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li and Min Zhang.
GTE is a BERT-style bidirectional encoder that replaces absolute position embeddings with RoPE, uses a gated MLP, and applies layer normalization after each residual connection. The same architecture backs Alibaba's gte-*-v1.5, gte-multilingual-* and gte-en-mlm-* checkpoints as well as Snowflake's snowflake-arctic-embed-m-v2.0.
Links: Documentation
Kernels] Bump version (#48714) by @vasquAutoImageProcessor requiring torchvision when only Pillow is installed (#48616) by @blipbyteMoE] Fix eager EP (#48653) by @vasquminimax failing with output_mismatch (tensor values differ (2)) (#48515) by @sergereview[bot]mistral failing with other (other (2)) (#48429) by @sergereview[bot]MemoryCleanupMixin class for tests (#48681) by @tarekziadeflex_olmo failing with other (other (1)) (#48668) by @sergereview[bot]RTDetrModel/SEWDForCTC loads (wrong base_model_prefix) (#48744) by @peft version requirement (#48716) by @shniuboboedgetam failing with import_or_config (other (12)) (#48322) by @sergereview[bot]AutoModel.from_pretrained not restoring modules_to_save weights (#48595) by @shniuboboreset on the dynamic cache layers (#48809) by @jiqing-fengStaticCache for Mllama and enable torch.compile (#48141) by @jiqing-fengfeat] Allow untying hidden_states[-1] from last_hidden_state via the model config (#48087) by @tomaarsendeepseek_vl failing with other (other (3)) (#48536) by @sergereview[bot]DSA] Only save latents on dsa with indexer as well (#48876) by @vasquNote truncated.
One column per quarter.
Fix some tests by removing the deprecation cycle ( #48503 ) by @Cyrilvallez in [ #48503 ]
Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
token. Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every
token to 8 of them. The context window is 1M tokens.
The architecture combines four features:
kv_lora_rank) that kv_b_proj expands back to one key/value per query head.index_topk keys per query with a lightweight indexer."full"indexer_types run an indexer; "shared" layers reuse the previous full layer's selection.hc_mult parallelThe implementation does not execute the multi-token prediction (MTP) layers. Released checkpoints
keep those weights so that other runtimes can use them for speculative decoding; they are ignored
at load time.
Links: Documentation
VibeVoice is a novel framework for synthesizing high-fidelity, long-form speech with multiple speakers by employing a next-token diffusion approach within a Large Language Model (LLM) structure. It's designed to capture the authentic conversational "vibe" and is particularly suited for generating audio content like podcasts and multi-participant audiobooks.
Links: Documentation
NeoMME is a family of efficient 260M and 800M parameter multimodal-native multilingual foundation encoders from H Company. It processes multilingual text tokens and raw image patches in a single bidirectional Transformer encoder, without a separately pretrained vision tower or causal language model.
NeoMME-Retriever is a model fine-tuned from the NeoMME backbone for visual document retrieval with joint late-interaction and dense objectives. It takes text queries and documents (text or page screenshots) and produces multi-vector embeddings for MeanMaxSim scoring (late-interaction) and mean-pooled embeddings for cosine similarity (dense).
Links: Documentation
Fun-ASR-Nano is an 800M-parameter end-to-end speech recognition model developed by Alibaba DAMO Academy's FunAudioLLM team. It achieves state-of-the-art performance on Chinese, English, and Japanese ASR benchmarks while being significantly smaller than comparable models.
Key features are
Links: Documentation
Kimi Linear is a hybrid linear attention architecture from Moonshot AI, introduced in
Kimi Linear: An Expressive, Efficient Attention Architecture.
At its core is Kimi Delta Attention (KDA), a refinement of Gated DeltaNet
that gives each key channel its own forget gate, so the recurrent state decays per channel instead of per head. KDA is
used in most layers; every fourth layer keeps a full-attention block that reuses DeepSeek-V3's Multi-head Latent
Attention (MLA), and the feed-forward blocks are DeepSeek-V3-style MoE with a shared expert.
Links: Documentation
Canary-1B-v2, a fast, robust multilingual model for Automatic Speech Recognition (ASR) and Speech-to-Text Translation (AST):
Canary reuses the Fast Conformer encoder from Parakeet (loaded through [ParakeetEncoder] / [ParakeetEncoderConfig]) and pairs it with a Transformer decoder that uses fixed sinusoidal positional embeddings, cross-attention to the encoder outputs and tied input/output embeddings. The task is selected through a decoder prompt prefix built by [CanaryProcessor] of the form <|startofcontext|> <|startoftranscript|> <|emo:undefined|> <source_lang> <target_lang> <pnc|nopnc> <|noitn|> <|notimestamp|> <|nodiarize|>, where source_lang == target_lang selects transcription and otherwise selects translation.
Links: Documentation
The NeuCodec model was proposed in Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates.
NeuCodec is a neural audio codec extending on XCodec2. It takes advantage of the following features:
Links: Documentation
Vision rotary embeddings (2D/3D) have been standardized into a unified RoPE frequency computation module, so users with custom vision models relying on attention-layer-level or model-specific RoPE grid interleaving logic must migrate to the new centralized modeling_rope_utils.py implementation.
Generation improvements include a performance optimization that avoids unnecessary accelerator synchronization on every decode step (reducing per-step overhead), and a fix to prevent unconditional downloading of remote hub files during generation. Several correctness fixes were also applied, including enforcing auto-compile cache checks for encoder-decoder models, standardizing past_key_values naming in AfMoE, and resolving flaky export and integration test failures.
Generate] Avoid unconditionally downloading remote hub file (#48620) by @vasqu in [#48620]generation failing with output_mismatch (list output differs (4)) (#48133) by @sergereview[bot] in [#48133]Fixed several cache-related bugs, including a quantized cache issue in VibeVoice, incorrect rejection of non-static cache implementations in VoxtralRealtime, missing auto-compile cache checks for encoder-decoder models, and a silent failure when paged attention is called without a cache. Documentation was also updated to clarify ContinuousBatchingConfig usage and sliding window model limitations.
Kernel support was improved with fixes for nested FLA kernel imports when only fla-core is installed, a warning when hub-kernel functions silently fall back to slower pure-PyTorch reference implementations, and the ability to register standalone functions (e.g., RoPE) in KernelConfig with optional non-inheritance of default mappings. Additional fixes include corrected repository paths for ESMFold2 kernels and updated documentation for KernelConfig customization.
Kernels] Enable functions into kernels registry and allow non inheritance (#48443) by @vasqu in [#48443]Fixed several quantization bugs, including a quant cache issue in VibeVoice, incorrect FP8 embedding handling for Qwen models, missing FP8 tensor parallelism layer overrides, and unnecessary MXFP4 weight dequantization on XPU devices.
shift_labels in decoder-only LLM/VLM losses (#48493) by @qgallouedec in [#48493]generate_flags parsing in transformers chat (#48597) by @SunMarc in [#48597]distogram_head in fp32 as well (#48488) by @kaixuanliu in [#48488]supports_context_parallel to PreTrainedModel (#48442) by @qgallouedec in [#48442]glm4_moe failing with OOM (other (2)) (#48551) by @sergereview[bot] in [#48551]nemotron failing with import_or_config (other (2)) (#48582) by @sergereview[bot] in [#48582]kosmos2 failing with import_or_config (other (2)) (#48552) by @sergereview[bot] in [#48552]Note truncated.
This is a special release as we include GLM! (and a few small fixes)
This is a special release as we include GLM! (and a few small fixes)
GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.
Links: Documentation
Mainly BC behavior for TP and pinning a hf kernel for security reasons 🤗
Full Changelog: v5.16.0...v5.16.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →