NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #225 most downloaded on PyPI
Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Last release 10 days ago
09 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
6 versions withdrawn
withdrawn after publishing
10 years old
242 releases · first in 2016
Biggest breaking change is in transformers chat. This command starts a terminal UI to interact with a chat model. It used to also be able to start a C…
<img width="1800" height="1013" alt="image" src="https://github.com/user-attachments/assets/7b5187d7-6945-4108-a546-6d1d7bfb55e3" />
We have a migration guide that will be continuously updated available on the main branch, please check it out in case you're facing issues: migration guide.
We are excited to announce the initial release of Transformers v5. This is the first major release in five years, and the release is significant: 1200 commits have been pushed to main since the latest minor release. This release removes a lot of long-due deprecations, introduces several refactors that significantly simplify our APIs and internals, and comes with a large number of bug fixes.
We give an overview of our focus for this release in the following blogpost. In these release notes, we'll focus directly on the refactors and new APIs coming with v5.
This release is the full V5 release. It sets in motion something bigger: going forward, starting with v5, we'll now release minor releases every week, rather than every 5 weeks. Expect v5.1 to follow next week, then v5.2 the week that follows, etc.
We're moving forward with this change to ensure you have access to models as soon as they're supported in the library, rather than a few weeks after.
In order to install this release, please do so with the following:
pip install transformers
For us to deliver the best package possible, it is imperative that we have feedback on how the toolkit is currently working for you. Please try it out, and open an issue in case you're facing something inconsistent/a bug.
Transformers version 5 is a community endeavor, and we couldn't have shipped such a massive release without the help of the entire community.
We introduce a new weight loading API in transformers, which significantly improves on the previous API. This
weight loading API is designed to apply operations to the checkpoints loaded by transformers.
Instead of loading the checkpoint exactly as it is serialized within the model, these operations can reshape, merge, and split the layers according to how they're defined in this new API. These operations are often a necessity when working with quantization or parallelism algorithms.
This new API is centered around the new WeightConverter class:
class WeightConverter(WeightTransform):
operations: list[ConversionOps]
source_keys: Union[str, list[str]]
target_keys: Union[str, list[str]]
The weight converter is designed to apply a list of operations on the source keys, resulting in target keys. A common operation done on the attention layers is to fuse the query, key, values layers. Doing so with this API would amount to defining the following conversion:
conversion = WeightConverter(
["self_attn.q_proj", "self_attn.k_proj", "self_attn.v_proj"], # The input layers
"self_attn.qkv_proj", # The single layer as output
operations=[Concatenate(dim=0)],
)
In this situation, we apply the Concatenate operation, which accepts a list of layers as input and returns a single
layer.
This allows us to define a mapping from architecture to a list of weight conversions. Applying those weight conversions
can apply arbitrary transformations to the layers themselves. This significantly simplified the from_pretrained method
and helped us remove a lot of technical debt that we accumulated over the past few years.
This results in several improvements:
Linked PR: https://github.com/huggingface/transformers/pull/41580
Just as we moved towards a single backend library for model definition, we want our tokenizers, and the Tokenizer object to be a lot more intuitive. With v5, tokenizer definition is much simpler; one can now initialize an empty LlamaTokenizer and train it directly on your corpus.
Defining a new tokenizer object should be as simple as this:
from transformers import TokenizersBackend, generate_merges
from tokenizers import pre_tokenizers, Tokenizer
from tokenizers.model import BPE
class Llama5Tokenizer(TokenizersBackend):
def __init__(self, unk_token="<unk>",bos_token="<s>", eos_token="</s>", vocab=None, merges=None ):
if vocab is None:
self._vocab = {
str(unk_token): 0,
str(bos_token): 1,
str(eos_token): 2,
}
else:
self._vocab = vocab
self._merges = merges
self._tokenizer = Tokenizer(
BPE(vocab=self._vocab, merges=self._merges, fuse_unk=True)
)
self._tokenizer.pre_tokenizer = pre_tokenizers.Metaspace(
replacement="▁", prepend_scheme=_get_prepend_scheme(self.add_prefix_space, self), split=False
)
super().__init__(
tokenizer_object=self._tokenizer,
unk_token=unk_token,
bos_token=bos_token,
eos_token=eos_token,
)
Once the tokenizer is defined as above, you can load it with the following: Llama5Tokenizer(). Doing this returns you an empty, trainable tokenizer that follows the definition of the authors of Llama5 (it does not exist yet :wink:).
The above is the main motivation towards refactoring tokenization: we want tokenizers to behave similarly to models: trained or empty, and with exactly what is defined in their class definition.
Up to now, transformers maintained two parallel implementations for many tokenizers:
tokenization_<model>.py) - Python-based implementations, often using SentencePiece as the backend.tokenization_<model>_fast.py) - Rust-based implementations using the 🤗 tokenizers library.In v5, we consolidate to a single tokenizer file per model: tokenization_<model>.py. This file will use the most appropriate backend available:
sentencepiece library. It inherits from PythonBackend.tokenizers. Basically allows adding tokens.MistralCommon's tokenization library. (Previously known as the MistralCommonTokenizer)The AutoTokenizer automatically selects the appropriate backend based on available files and dependencies. This is transparent, you continue to use AutoTokenizer.from_pretrained() as before. This allows transformers to be future-proof and modular to easily support future backends.
We enable users and tokenizer builders to define their own tokenizers from top to bottom. Tokenizers are usually defined using a backend such as tokenizers, sentencepiece or mistral-common, but we offer the possibility to design the tokenizer at a higher-level, without relying on those backends.
To do so, you can import the PythonBackend (which was previously known as PreTrainedTokenizer). This class encapsulates all the logic related to added tokens, encoding, and decoding.
If you want something even higher up the stack, then PreTrainedTokenizerBase is what PythonBackend inherits from. It contains the very basic tokenizer API features:
encodedecodevocab_sizeget_vocabconvert_tokens_to_idsconvert_ids_to_tokensfrom_pretrainedsave_pretrainedStarting with v5, we now enable initializing blank, untrained tokenizers-backed tokenizers:
from transformers import LlamaTokenizer
tokenizer = LlamaTokenizer()
This tokenizer will therefore follow the definition of the LlamaTokenizer as defined in its class definition. It can then be trained on a corpus as can be seen in the tokenizers documentation.
These tokenizers can also be initialized from vocab and merges (if necessary), like the previous "slow" tokenizers:
from transformers import LlamaTokenizer
vocab = {"<unk>": 0, "<s>": 1, "</s>": 2, "hello": 3, "world": 4}
merges = [("h", "e"), ("l", "l"), ("o", " ")]
tokenizer = LlamaTokenizer(vocab=vocab, merges=merges)
This tokenizer will behave as a Llama-like tokenizer, with an updated vocabulary. This allows comparing different tokenizer classes with the same vocab; therefore enabling the comparison of different pre-tokenizers, normalizers, etc.
⚠️ The vocab_file (as in, a path towards a file containing the vocabulary) cannot be used to initialize the LlamaTokenizer as loading from files is reserved to the from_pretrained method.
The batch_decode and decode methods have been unified to reflect behavior of the encode method. Both single and batch decoding now use the same decode method. See an example of the new behavior below:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("t5-small")
inputs = ["hey how are you?", "fine"]
tokenizer.decode(tokenizer.encode(inputs))
Gives:
- 'hey how are you?</s> fine</s>'
+ ['hey how are you?</s>', 'fine</s>']
We expect encode and decode to behave, as two sides of the same coin: encode, process, decode, should work.
[!NOTE] A common use-case would be:
encode,model.generate,decode. However, usinggeneratewould returnlist[list[int]], which would then be incompatible withdecode.
The encode_plus method is deprecated in favor of the single __call__ method.
apply_chat_template returns BatchEncodingPreviously, apply_chat_template returned input_ids for backward compatibility. Starting with v5, it now consistently returns a BatchEncoding dict like other tokenizer methods.
# v5
messages = [
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hi there!"}
]
# Now returns BatchEncoding with input_ids, attention_mask, etc.
outputs = tokenizer.apply_chat_template(messages, return_tensors="pt")
print(outputs.keys()) # dict_keys(['input_ids', 'attention_mask'])
We simplify the serialization of tokenization attributes:
special_tokens_map.json - special tokens are now stored in tokenizer_config.json.added_tokens.json - added tokens are now stored in tokenizer.json.added_tokens_decoder is only stored when there is no tokenizer.json.When loading older tokenizers, these files are still read for backward compatibility, but new saves use the consolidated format. We're gradually moving towards consolidating attributes to fewer files so that other libraries and implementations may depend on them more reliably.
Several models that had identical tokenizers now import from their base implementation:
These modules will eventually be removed altogether.
Removed T5-specific workarounds
The internal _eventually_correct_t5_max_length method has been removed. T5 tokenizers now handle max length consistently with other models.
A few testing changes specific to tokenizers have been applied:
add_tokens, encode, decode) are now centralized and automatically applied across all tokenizers. This reduces test duplication and ensures consistent behaviorFor legacy implementations, the original BERT Python tokenizer code (including WhitespaceTokenizer, BasicTokenizer, etc.) is preserved in bert_legacy.py for reference purposes.
Special Tokens Structure:
SpecialTokensMixin: Merged into PreTrainedTokenizerBase to simplify the tokenizer architecture.special_tokens_map: Now only stores named special token attributes (e.g., bos_token, eos_token). Use extra_special_tokens for additional special tokens (formerly additional_special_tokens). all_special_tokens includes both named and extra tokens.# v4
tokenizer.special_tokens_map # Included 'additional_special_tokens'
# v5
tokenizer.special_tokens_map # Only named tokens
tokenizer.extra_special_tokens # Additional tokens
special_tokens_map_extended and all_special_tokens_extended: Removed. Access AddedToken objects directly from _special_tokens_map or _extra_special_tokens if needed.additional_special_tokens: Still accepted for backward compatibility but is automatically converted to extra_special_tokens.Deprecated Methods:
sanitize_special_tokens(): Already deprecated in v4, removed in v5.prepare_seq2seq_batch(): Deprecated; use __call__() with text_target parameter instead.# v4
model_inputs = tokenizer.prepare_seq2seq_batch(src_texts, tgt_texts, max_length=128)
# v5
model_inputs = tokenizer(src_texts, text_target=tgt_texts, max_length=128, return_tensors="pt")
model_inputs["labels"] = model_inputs.pop("input_ids_target")
BatchEncoding.words(): Deprecated; use word_ids() instead.Removed Methods:
create_token_type_ids_from_sequences(): Removed from base class. Subclasses that need custom token type ID creation should implement this method directly.prepare_for_model(), build_inputs_with_special_tokens(), truncate_sequences(): Moved from tokenization_utils_base.py to tokenization_python.py for PythonBackend tokenizers. TokenizersBackend provides model-ready input via tokenize() and encode(), so these methods are no longer needed in the base class._switch_to_input_mode(), _switch_to_target_mode(), as_target_tokenizer(): Removed from base class. Use __call__() with text_target parameter instead.# v4
with tokenizer.as_target_tokenizer():
labels = tokenizer(tgt_texts, ...)
# v5
labels = tokenizer(text_target=tgt_texts, ...)
parse_response(): Removed from base class.The v5 release significantly improves the performance of the MoE models, as can be seen in the graphs below. We improve and optimize MoE performance through batched and grouped experts implementations, and we optimize them for decoding using batched_mm.
<img width="2048" height="1451" alt="image" src="https://github.com/user-attachments/assets/c3f2e59f-3026-4f56-9a56-36e4eb0fcf73" />
We focus on improving the performance of loading weights on device (which gives speedups up to 6x in tensor parallel situations); this is preliminary work that we'll continue to work on in the coming weeks. Some notable improvements:
dtype updateWe have updated the default dtype for all models loaded with from_pretrained to be auto. This will lead to model instantiations respecting the dtype in which the model was saved, rather than forcing it to load in float 32.
You can, of course, still specify the dtype in which you want to load your model by specifying it as an argument to the from_pretrained method.
The Hugging Face Hub infrastructure has gradually moved to a XET backend. This will significantly simplify uploads and downloads, with higher download and upload speeds, partial uploads, and, most notably, a higher threshold for accepted file sizes on the Hugging Face Hub.
To reflect this, we're increasing the default shard size of models serialized on the Hub to 50GB (up from 5GB).
use_auth_tokenThe use_auth_token argument/parameter is deprecated in favor of token everywhere.
You should be able to search and replace use_auth_token with token and get the same logic.
Linked PR: https://github.com/huggingface/transformers/pull/41666
We decided to remove some features for the upcoming v5 as they are currently only supported in a few old models and no longer integrated in current model additions. It's recommended to stick to v4.x in case you need them. Following features are affected:
We dropped support for two torch APIs:
torchscript in https://github.com/huggingface/transformers/pull/41688torch.fx in https://github.com/huggingface/transformers/pull/41683Those APIs were deprecated by the PyTorch team, and we're instead focusing on the supported APIs dynamo and export.
We clean up the quantization API in transformers, and significantly refactor the weight loading as highlighted above.
We drop support for two quantization arguments that have been deprecated for some time:
load_in_4bitload_in_8bitWe remove them in favor of the quantization_config argument which is much more complete. As an example, here is how
you would load a 4-bit bitsandbytes model using this argument:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(load_in_4bit=True)
model_4bit = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-3B",
device_map="auto",
quantization_config=quantization_config
)
from_xxx_config are deleted. Configs can be init from the __init__ method in the same way. See #41314.mode.rope_parameters, including the rope_theta and rope_type. Model's config.rope_parameters is a simple dictionaty in most cases, and can also be a nested dict in special cases (i.e. Gemma3 and ModernBert) with different rope parameterization for each layer type. Trying to get config.rope_theta will throw an attribute error from now on. See #39847 and #42255config.vocab_size). Users are expected to access keys from their respective sub-configs (config.text_config.vocab_size).model.generate()) will no longer have a generation_config and model.config.generation_config will throw an attribute error.tokenization_<model>.py ) will be removed in favor of using fast tokenizer files tokenization_<model>_fast.py --> will be renamed to tokenization_<model>.py. As fast tokenizers are :hugs:tokenizers - backend, they include a wider range of features that are maintainable and reliable.encode_plus --> __call__batch_decode --> decodeapply_chat_template by default returns naked input_ids rather than a BatchEncoding dict.
This was inconvenient - it should return a BatchEncoding dict like tokenizer.__call__(), but we were stuck with
it for backward compatibility. The method now returns a BatchEncoding.
Linked PRs:
processor_config.json as a nested dict, instead of serializing attributes in their own config files. Loading will be supported for all old format processors (https://github.com/huggingface/transformers/pull/41474)XXXFeatureExtractors classes are completely removed in favor of XXXImageProcessor class for all vision models (https://github.com/huggingface/transformers/pull/41174)XXXFastImageProcessorKwargs is removed in favor of XXXImageProcessorKwargs which will be shared between fast and slow processors (https://github.com/huggingface/transformers/pull/40931)RotaryEmbeddings layers will start returning a dict of tuples, in case the model uses several RoPE configurations (Gemma2, ModernBert). Each value will be a tuple of "cos, sin" per RoPE type.RotaryEmbeddings layer will be unified and accessed via config.rope_parameters. Config attr for rope_theta might not be accessible anymore for some models, and instead will be in config.rope_parameters['rope_theta']. BC will be supported for a while as much as possible, and in the near future we'll gradually move to the new RoPE format (https://github.com/huggingface/transformers/pull/39847)model.language_model. It is recommended to either access the module with model.model.language_model or model.get_decoder(). See #42156kwargs in their forward methodsGreedySearchEncoderDecoderOutput). We now only have 4 output classes built from the following matrix: decoder-only vs encoder-decoder, uses beams vs doesn't use beams (https://github.com/huggingface/transformers/pull/40998)generate doesn't receive any KV Cache argument, the default cache class used is now defined by the model (as opposed to always being DynamicCache) (https://github.com/huggingface/transformers/pull/41505)config.json for any old model, it will be loaded back into model's generation config. Users are expected to access or modify generation parameters only with model.generation_config.do_sample = True.compute_loss_func Handling
compute_loss_func now always takes priority over the model's built-in loss computation, giving users consistent control over custom loss functions.num_items_in_batch in Prediction Step
num_items_in_batch argument is now passed to compute_loss during prediction_step, enabling proper loss scaling during evaluation.report_to now defaults to "none"
TrainingArguments due to low usagemp_parameters -> legacy param that was later on added to the Sagemaker trainer_n_gpu -> not intended for users to set, we will initialize it correctly instead of putting it in the TrainingArgumentsoverwrite_output_dir - > replaced by resume_from_checkpoint, and it was only used in the examples script, no impact on Trainer.logging_dir -> only used for tensorboard, set TENSORBOARD_LOGGING_DIR env var insteadjit_mode_eval -> use use_torch_compile instead, as torchscript is not recommended anymoretpu_num_cores-> It is actually better to remove it, as it is not recommended to set the number of cores. By default, all TPU cores are used . Set TPU_NUM_CORES env var insteadpast_index -> it was only used for a very small number of models that have special architecture like transformersxl + it was not documented at all how to train those modelsray_scope -> only for a minor arg for ray integration. Set RAY_SCOPE var env insteadwarmup_ratio -> use warmup_step instead. We combined both args together by allowing passing float values in warmup_step.TrainingArgumentsfsdp_min_num_params and fsdp_transformer_layer_cls_to_wrap -> use fsdp_configtpu_metrics_debug -> debugpush_to_hub_token -> hub_tokenpush_to_hub_model_id and push_to_hub_organization -> hub_model_idinclude_inputs_for_metrics -> include_for_metricsper_gpu_train_batch_size -> per_device_train_batch_sizeper_gpu_eval_batch_size -> per_device_eval_batch_sizeuse_mps_device -> mps will be used by default if detectedfp16_backend and half_precision_backend -> we will only rely on torch.amp as everything has been upstreamed to torchno_cuda -> use_cpu include_tokens_per_second -> include_num_input_tokens_seenuse_legacy_prediction_loop -> we only use evaluation_loop function from now onTrainertokenizer in initialization -> processing_classmodel_path in train() -> resume_from_checkpointTrainerTraineruse_cache in the model config will be set to False. You can still change the cache value through TrainingArguments usel_cache argument if needed.organization and repo_url from PushToHubMixin. You must pass a repo_id instead.ignore_metadata_errors from PushToMixin. In practice if we ignore errors while loading the model card, we won't be able to push the card back to the Hub so it's better to fail early and not provide the option to fail later.push_to_hub do not accept **kwargs anymore. All accepted parameters are explicitly documented.push_to_hub are now keyword-only to avoid confusion. Only repo_id can be positional since it's the main arg.use_temp_dir argument from push_to_hub. We now use a tmp dir in all cases.Linked PR: https://github.com/huggingface/transformers/pull/42391.
The deprecated transformers-cli ... command was deprecated, transformers ... is now the only CLI entry point.
transformers CLI has been migrated to Typer, making it easier to maintain + adding some nice features out of
the box (improved --help section, autocompletion).
Biggest breaking change is in transformers chat. This command starts a terminal UI to interact with a chat model.
It used to also be able to start a Chat Completion server powered by transformers and chat with it. In this revamped
version, this feature has been removed in favor of transformers serve. The goal of splitting transformers chat
and transformers serve is to define clear boundaries between client and server code. It helps with maintenance
but also makes the commands less bloated. The new signature of transformers chat is:
Usage: transformers chat [OPTIONS] BASE_URL MODEL_ID [GENERATE_FLAGS]...
Chat with a model from the command line.
It works hand in hand with transformers serve, which means that if transformers serve is running on its default endpoint, transformers chat can be launched as follows:
transformers chat HuggingFaceTB/SmolLM3-3B
It can however use any OpenAI API compatible HTTP endpoint:
transformers chat HuggingFaceTB/SmolLM3-3B https://router.huggingface.co/v1
Linked PRs:
run methodThe transformers run (previously transformers-cli run) is an artefact of the past, was not documented nor tested,
and isn't part of any public documentation. We're removing it for now and ask you to please let us know in case
this is a method you are using; in which case we should bring it back with better support.
Linked PR: https://github.com/huggingface/transformers/pull/42447
TRANSFORMERS_CACHE, PYTORCH_TRANSFORMERS_CACHE, and PYTORCH_PRETRAINED_BERT_CACHE have been removed. Please use HF_HOME instead.HUGGINGFACE_CO_EXAMPLES_TELEMETRY, HUGGINGFACE_CO_EXAMPLES_TELEMETRY, HUGGINGFACE_CO_PREFIX, and HUGGINGFACE_CO_RESOLVE_ENDPOINT have been removed. Please use huggingface_hub.constants.ENDPOINT instead.Linked PR: https://github.com/huggingface/transformers/pull/42391.
transformers v5 pins the huggingface_hub version to >=1.0.0. See this migration guide to learn more about this major release. Here are to main aspects to know about:
requests to httpx. This change was made to improve performance and to support both synchronous and asynchronous requests the same way. If you are currently catching requests.HTTPError errors in your codebase, you'll need to switch to httpx.HTTPError.HTTP_PROXY / HTTPS_PROXY environment variableshf_transfer and therefore HF_HUB_ENABLE_HF_TRANSFER have been completed dropped in favor of hf_xet. This should be transparent for most users. Please let us know if you notice any downside!typer-slim has been added as required dependency, used to implement both hf and transformers CLIs.
<img width="809" height="471" alt="image" src="https://github.com/user-attachments/assets/58bb9c70-d481-48ed-ab8f-6553be7c240f" />
The Code World Model (CWM) model was proposed in CWM: An Open-Weights LLM for Research on Code Generation with World Models by Meta FAIR CodeGen Team. CWM is an LLM for code generation and reasoning about code that has, in particular, been trained to better represent and reason about how code and commands affect the state of a program or system. Specifically, we mid-trained CWM on a large number of observation-action trajectories from Python execution traces and agentic interactions in containerized environments. We post-trained with extensive multi-task RL in verifiable coding, math, and multi-turn software engineering environments.
<img width="1505" height="915" alt="image" src="https://github.com/user-attachments/assets/eec48633-f02b-464a-ae5c-c65473387e53" />
SAM3 (Segment Anything Model 3) was introduced in SAM 3: Segment Anything with Concepts.
The SAM3 addition adds four new architectures:
SAM3 performs Promptable Concept Segmentation (PCS) on images. PCS takes text and/or image exemplars as input (e.g., "yellow school bus"), and predicts instance and semantic masks for every single object matching the concept.
Sam3Tracker and Sam3TrackerVideo perform Promptable Visual Segmentation (PVS) on images. PVS takes interactive visual prompts (points, boxes, masks) or text inputs to segment a specific object instance per prompt. This is the task that SAM 1 and SAM 2 focused on, and SAM 3 improves upon it. Sam3Tracker and Sam3TrackerVideo are updated versions of SAM2 Video that maintain the same API while providing improved performance and capabilities.
SAM3 Video performs Promptable Concept Segmentation (PCS) on videos. PCS takes text as input (e.g., "yellow school bus"), and predicts instance and semantic masks for every single object matching the concept, while preserving object identities across video frames. The model combines a detection module (SAM3) with a tracking module (SAM2-style tracker) to enable robust object tracking across video frames using text prompts.
<img width="1080" height="849" alt="image" src="https://github.com/user-attachments/assets/a9fa1b81-114d-4054-9699-5083ac69d830" />
LFM2-MoE is a Mixture-of-Experts (MoE) variant of LFM2. The LFM2 family is optimized for on-device inference by combining short‑range, input‑aware gated convolutions with grouped‑query attention (GQA) in a layout tuned to maximize quality under strict speed and memory constraints.
LFM2‑MoE keeps this fast backbone and introduces sparse MoE feed‑forward networks to add representational capacity without significantly increasing the active compute path. The first LFM2-MoE release is LFM2-8B-A1B, with 8.3B total parameters and 1.5B active parameters. The model excels in quality (comparable to 3-4B dense models) and speed (faster than other 1.5B class models).
<img width="812" height="366" alt="image" src="https://github.com/user-attachments/assets/21c82c6e-cf0a-4d6c-a707-b9e57663ca85" />
The VideoLLaMA3 model is a major update to VideoLLaMA2 from Alibaba DAMO Academy.
<img width="621" height="475" alt="image" src="https://github.com/user-attachments/assets/c9616758-b3aa-41d0-bd58-695966ba146d" />
Audio Flamingo 3 (AF3) is a fully open large audio–language model designed for robust understanding and reasoning over speech, environmental sounds, and music. AF3 pairs a Whisper-style audio encoder with a causal language model and performs replace-in-place audio–text fusion: the processor aligns post-pool audio frames to a dedicated placeholder token and the model replaces those token slots with projected audio embeddings during the forward pass.
The model checkpoint is available at: nvidia/audio-flamingo-3-hf
Highlights:
NanoChat is a compact decoder-only transformer model designed for educational purposes and efficient training. The model features several fundamental architectural innovations which are common in modern transformer models. Therefore, it is a good model to use as a starting point to understand the principles of modern transformer models. NanoChat is a variant of the Llama architecture, with simplified attention mechanism and normalization layers.
<img width="868" height="331" alt="image" src="https://github.com/user-attachments/assets/cd8b82cf-10de-49b0-af2a-28ffbeac6fa7" />
FastVLM is an open-source vision-language model featuring a novel hybrid vision encoder, FastViTHD. Leveraging reparameterizable convolutional layers, scaled input resolution, and a reduced number of visual tokens, FastVLM delivers high accuracy with exceptional efficiency. Its optimized architecture enables deployment even on edge devices, achieving ultra-low TTFT (time to first token) without sacrificing performance.
<img width="3840" height="2160" alt="image" src="https://github.com/user-attachments/assets/712fbe57-a4b6-4bf1-acf6-3ef803f75f0e" />
PaddleOCR-VL is a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-language model (VLM) that integrates a NaViT-style dynamic resolution visual encoder with the ERNIE-4.5-0.3B language model to enable accurate element recognition. This innovative model efficiently supports 109 languages and excels in recognizing complex elements (e.g., text, tables, formulas, and charts), while maintaining minimal resource consumption. Through comprehensive evaluations on widely used public benchmarks and in-house benchmarks, PaddleOCR-VL achieves SOTA performance in both page-level document parsing and element-level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top-tier VLMs, and delivers fast inference speeds. These strengths make it highly suitable for practical deployment in real-world scenarios.
<img width="719" height="541" alt="image" src="https://github.com/user-attachments/assets/2fc5ee26-3c15-451f-bc7f-c779c2a78919" />
PE Audio (Perception Encoder Audio) is a state-of-the-art multimodal model that embeds audio and text into a shared (joint) embedding space. The model enables cross-modal retrieval and understanding between audio and text.
Text input
Audio input
The resulting embeddings can be used for:
<img width="2100" height="1154" alt="image" src="https://github.com/user-attachments/assets/a9343e81-903b-4445-ba0c-61c87830776a" />
Jais2 a next-generation Arabic open-weight LLM trained on the richest Arabic-first dataset to date. Built from the ground up with 8B and 70B parameters, Jais 2 understands Arabic the way it's truly spoken across dialects, cuulutre, and modern expression. It is developed by MBZUAI, Inception and Cerebras Systems and based on the transformer architecture with modifications including:
<img width="5478" height="2102" alt="image" src="https://github.com/user-attachments/assets/abbd93d4-8c4c-4fd9-8fab-58d969d8b296" />
Pixio is a vision foundation model that uses ViT as a feature extractor for multiple downstream tasks like depth estimation, semantic segmentation, feed-forward 3D reconstruction, robotics, and image classification. It is built on the Masked Autoencoder (MAE) pre-training framework, with four minimal yet critical updates: 1) deeper decoder, 2) larger masking granularity, 3) more class tokens, and 4) web-scale curated training data.
<img width="848" height="853" alt="image" src="https://github.com/user-attachments/assets/02817ead-7560-4eea-8f2b-0b959553a3cd" />
The Ernie 4.5 VL MoE model was released in the Ernie 4.5 Model Family release by baidu. This family of models contains multiple different architectures and model sizes. The Vision-Language series in specific is composed of a novel multimodal heterogeneous structure, sharing paremeters across modalities and dedicating parameters to specific modalities. This becomes especially apparent in the Mixture of Expert (MoE) which is composed of
This architecture has the advantage to enhance multimodal understanding without compromising, and even improving, performance on text-related tasks. An more detailed breakdown is given in the Technical Report.
Ernie 4.5] Ernie VL models by @vasqu in https://github.com/huggingface/transformers/pull/39585<img width="1600" height="1029" alt="image" src="https://github.com/user-attachments/assets/d630a900-9ef5-467c-93ab-34f064501b8d" />
GLM-ASR-Nano-2512 is a robust, open-source speech recognition model with 1.5B parameters. Designed for real-world complexity, it outperforms OpenAI Whisper V3 on multiple benchmarks while maintaining a compact size.
Key capabilities include:
Exceptional Dialect Support Beyond standard Mandarin and English, the model is highly optimized for Cantonese (粤语) and other dialects, effectively bridging the gap in dialectal speech recognition.
Low-Volume Speech Robustness Specifically trained for "Whisper/Quiet Speech" scenarios. It captures and accurately transcribes extremely low-volume audio that traditional models often miss.
SOTA Performance Achieves the lowest average error rate (4.10) among comparable open-source models, showing significant advantages in Chinese benchmarks (Wenet Meeting, Aishell-1, etc..).
This model was contributed by Eustache Le Bihan and Yuxuan Zhang. you can check the model card for more details and our github repo.
<img width="5038" height="5860" alt="image" src="https://github.com/user-attachments/assets/af0f2821-3831-4afd-a7dc-26a91f17afd0" />
GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.
<img width="10101" height="7371" alt="image" src="https://github.com/user-attachments/assets/e2de76ab-35a9-4e42-9252-6b3dbf75e991" />
We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the reasoning-centric training framework. We first develop a capable vision foundation model with significant potential through large-scale pre-training, which arguably sets the upper bound for the final performance. We then propose Reinforcement Learning with Curriculum Sampling (RLCS) to unlock the full potential of the model, leading to comprehensive capability enhancement across a diverse range of tasks, including STEM problem solving, video understanding, content recognition, coding, grounding, GUI-based agents, and long document interpretation. In a comprehensive evaluation across 42 public benchmarks, GLM-4.5V achieves state-of-the-art performance on nearly all tasks among open-source models of similar size, and demonstrates competitive or even superior results compared to closed-source models such as Gemini-2.5-Flash on challenging tasks including Coding and GUI Agents. Meanwhile, the smaller GLM-4.1V-9B-Thinking remains highly competitive-achieving superior results to the much larger Qwen2.5-VL-72B on 29 benchmarks. We open-source both GLM-4.1V-9B-Thinking and GLM-4.5V. We further introduce the GLM-4.6V series, open-source multimodal models with native tool use and a 128K context window. A brief overview is available at this https URL. Code, models and more information are released at https://github.com/zai-org/GLM-V
<img width="706" height="242" alt="image" src="https://github.com/user-attachments/assets/2c4770f4-5d48-4d87-a454-b574b648e5ae" />
LW-DETR proposes a light-weight Detection Transformer (DETR) architecture designed to compete with and surpass the dominant YOLO series for real-time object detection. It achieves a new state-of-the-art balance between speed (latency) and accuracy (mAP) by combining recent transformer advances with efficient design choices.
The LW-DETR architecture is characterized by its simple and efficient structure: a plain ViT Encoder, a Projector, and a shallow DETR Decoder. It enhances the DETR architecture for efficiency and speed using the following core modifications:
Efficient ViT Encoder: Uses a plain ViT with interleaved window/global attention and a window-major organization to drastically reduce attention complexity and latency.
Richer Input: Aggregates multi-level features from the encoder and uses a C2f Projector (YOLOv8) to pass two-scale features ( 1 / 8 and 1 / 32 ).
Faster Decoder: Employs a shallow 3-layer DETR decoder with deformable cross-attention for lower latency and faster convergence.
Optimized Queries: Uses a mixed-query scheme combining learnable content queries and generated spatial queries.
<img width="1172" height="661" alt="image" src="https://github.com/user-attachments/assets/cee98e82-b3d0-42a6-b820-3061752ad4a8" />
LightOnOcr combines a Vision Transformer encoder (Pixtral-based) with a lightweight text decoder (Qwen3-based) distilled from high-quality open VLMs. It is optimized for document parsing tasks, producing accurate, layout-aware text extraction from high-resolution pages.
JetMoe Fix jetmoe after #40132 by @ArthurZucker in #41324gemma3 by @Sai-Suraj-27 in #41354PretrainedConfig to PreTrainedConfig by @Cyrilvallez in #41300ModularChecker] QOL for the modular checker by @ArthurZucker in #41361v5] Remove relative position embeddings (for bert like models) by @vasqu in #41170apply_chat_template by @Samoed in #41355test_longcat_generation_cpu by @ydshieh in #41368CB] Refactors the way we access paged by @ArthurZucker in #41370v5] Sync Bert and Bart eager attention by @vasqu in #41248TypeError exception for invalid type by @Sai-Suraj-27 in #41346update_device_map for GPTQ quantizer by @Sai-Suraj-27 in #41328prune_heads by @gante in #41417JetMoe] Fix KV head repetition and padding free by @vasqu in #41423JetMoeIntegrationTest by @ydshieh in #41377past_key_value in BERT-like models by @zucchini-nlp in #41448utils/tf_ops/ by @gante in #41402Attention Masks] Bidirectional masks for encoder and encoder-decoder models by @vasqu in #41265past_index by @SunMarc in #41384report_to default changed to "none" + cleaning deprecated env var by @SunMarc in #41375overwrite_output_dir by @SunMarc in #41323CI] Fix copies on main by @vasqu in #41486jit_mode_eval by @SunMarc in #41376local_rank arg from TrainingArguments by @SunMarc in #41382pickle - BloomTokenizerFast by @ydshieh in #41466glm4v by @Sai-Suraj-27 in #41483truncation to False in Qwen3Omni to avoid default truncation by @BakerBunker in #41473local_rank deletion and some cleaning by @SunMarc in #41504tpu_num_cores by @SunMarc in #41383HunYuanMoEV1IntegrationTest:test_model_generation by @ydshieh in #41373generate delegates default cache initialization to the model by @gante in #41505from_pretrained] Small refactor from_pretrained: move around unrelated stuff by @ArthurZucker in #41445transformers serve by @LysandreJik in #41446logits_to_keep to many older CausalLM models by @philiproeleveld in #41335torch.compile recompiled part of th… by @sywangyi in #41558Docs] Fix changed references by @vasqu in #41614expand_device_map instead of redefining it by @Cyrilvallez in #41608tp_plan in from_pretrained directly by @Cyrilvallez in #41435Executorch] Simplify for encoder models by @vasqu in #41627Ernie 4.5 Moe] Fix Moe and offloading by @vasqu in #41385Masks] Fix mask handling in eager for vision models by @vasqu in #41625utils/check_bad_commit.py by @ydshieh in #41658use_cache default to False by @SunMarc in #41585chat_extras.md to Korean by @Judy-Choi in #39863big_bird.md to Korean by @ssum21 in #40445code_llama.md to Korean by @Judy-Choi in #40558ko-LFM2.md to Korean by @ssum21 in #41502use_auth_token parameter by @Wauplin in #41666Attn] Allow dynamic causality in SDPA via Kwargs by @vasqu in #41692run_name docs in TrainingArguments by @tobiasofsn in #41705utils/check_bad_commit.py by @ydshieh in #41658)videos from image processing classes by @zucchini-nlp in #41607@staticmethod from module-level get_device_and_memory_breakdown by @albertvillanova in #41747Onnx docs] Remove some traces by @vasqu in #41791utils/check_bad_commit.py by @ydshieh in #41658)Clip] Fix masking and enable flash attention on all model types by @vasqu in #41750test_tensor_parallel.py by @3outeille in #41918detectron2 installation in docker files by @ydshieh in #41975autoawq[kernels] installation in quantization docker file by @ydshieh in #41978torchcodec version in quantization docker file by @ydshieh in #41988run slow v2: empty report when there is only one model by @ydshieh in #42002torch+deepspeed docker file by @ydshieh in #41985logging_dir by @SunMarc in #42013deeepspeed in AMD docker file by @ydshieh in #42025huggingface_hub dependency version by @hanouticelina in #42033pr_slow_ci_suggestion.yml after #42023 by @ydshieh in #42049Argument list too long in pr_slow_ci_suggestion.yml by @ydshieh in #42061setattr as well by @zucchini-nlp in #41808Attn Masks] Non-vmap default for attention masks by @vasqu in #41852image_transforms.py by @yaswanth19 in #42044prepare_inputs_for_generation cache slicing condition by @aNote truncated.
One column per quarter.
We are getting closer and closer to the official release! This RC is focused on removing more of the deprecated stuff, fixing some minors issues, doc…
We are getting closer and closer to the official release! This RC is focused on removing more of the deprecated stuff, fixing some minors issues, doc updates.
_get_num_multimodal_tokens by @Abhinavexists in https://github.com/huggingface/transformers/pull/43137BartModelIntegrationTest by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43160auto_doctring in Processors by @yonigozlan in https://github.com/huggingface/transformers/pull/42101BitModelIntegrationTest by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43164Fp8] Fix experts by @vasqu in https://github.com/huggingface/transformers/pull/43154salesforce-ctrl, xlm & gpt-neo model generation tests by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43180Generate] Allow custom config values in generate config by @vasqu in https://github.com/huggingface/transformers/pull/43181Pix2StructIntegrationTest by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43229PhiIntegrationTests by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43214HF_TOKEN directly and remove require_read_token by @ydshieh in https://github.com/huggingface/transformers/pull/43233Owlv2ModelIntegrationTest & OwlViTModelIntegrationTest by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43182add_dates by @yonigozlan in https://github.com/huggingface/transformers/pull/43199Vip-llava model integration test by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43252position_ids in all apply_rotary_pos_emb by @Cyrilvallez in https://github.com/huggingface/transformers/pull/43255_get_test_info in testing_utils.py by @ydshieh in https://github.com/huggingface/transformers/pull/43259Hiera, SwiftFormer & LED Model integration tests by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43225_toctree.yml by @Cyrilvallez in https://github.com/huggingface/transformers/pull/43264PegasusX, Mvp & LED model integration tests by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/43245Full Changelog: https://github.com/huggingface/transformers/compare/v5.0.0rc2...v5.0.0rc3
This release candidate is focused on fixing AutoTokenizer, expanding the dynamic weight loading support, and improving performances with MoEs!
This release candidate is focused on fixing AutoTokenizer, expanding the dynamic weight loading support, and improving performances with MoEs!
<img width="2048" height="1451" alt="image" src="https://github.com/user-attachments/assets/3ed2508e-3eb1-4f13-8717-cd9027d12a39" />
The main issue with the tokenization refactor is that tokenizer_class are now "enforced" when in most cases they are wrong. This took a while to properly isolate and now we try to use TokenizersBackend whenever we can. #42894 has a much more detailed description of the big changes!
TokenizersBackend by @ArthurZucker in https://github.com/huggingface/transformers/pull/42894Tokenizers] Change treatment of special tokens by @vasqu in https://github.com/huggingface/transformers/pull/42903Here we focused on boosting the performances of loading weights on device!
post_init and fix all of them by @Cyrilvallez in https://github.com/huggingface/transformers/pull/42873_init_weights for ALL models by @Cyrilvallez in https://github.com/huggingface/transformers/pull/42309Ernie 4.5] Ernie VL models by @vasqu in https://github.com/huggingface/transformers/pull/39585Mostly around processors!
convert_segmentation_map_to_binary_masks to EoMT by @simonreise in https://github.com/huggingface/transformers/pull/43073Thanks again to everyone !
Full Changelog: https://github.com/huggingface/transformers/compare/v5.0.0rc1...v5.0.0rc2
This release candidate was focused mostly on quantization support with the new dynamic weight loader, and a few notable 🚨 breaking changes🚨:
This release candidate was focused mostly on quantization support with the new dynamic weight loader, and a few notable 🚨 breaking changes🚨:
from_pretrained is now auto!Mostly QOL and fixed + support back CPU offloading.
Mostly added support for fbgemme , quanto,
The dynamic weight loader broke small things, this adds glue for all models but MoEs.
Tokenization needed more refactoring, this time its a lot cleaner!
rope_parameters to empty dict if there is something to put in it by @hmellor in https://github.com/huggingface/transformers/pull/42651We omitted a lot of other commits for clarity, but thanks to everyone and the new contributors!
Full Changelog: https://github.com/huggingface/transformers/compare/v5.0.0rc0...v5.0.0rc1
Biggest breaking change is in transformers chat. This command starts a terminal UI to interact with a chat model. It used to also be able to start a C…
<img width="1800" height="1013" alt="image" src="https://github.com/user-attachments/assets/7b5187d7-6945-4108-a546-6d1d7bfb55e3" />
We are excited to announce the initial release of Transformers v5. This is the first major release in five years, and the release is significant: 800 commits have been pushed to main since the latest minor release. This release removes a lot of long-due deprecations, introduces several refactors that significantly simplify our APIs and internals, and comes with a large number of bug fixes.
We give an overview of our focus for this release in the following blogpost. In these release notes, we'll focus directly on the refactors and new APIs coming with v5.
This release is a release candidate (RC). It is not the final v5 release, and we will push on pypi as a pre-release. This means that the current release is purely opt-in, as installing transformers without specifying this exact release will install the latest version instead (v4.57.3 as of writing).
In order to install this release, please do so with the following:
pip install transformers --pre
For us to deliver the best package possible, it is imperative that we have feedback on how the toolkit is currently working for you. Please try it out, and open an issue in case you're facing something inconsistent/a bug.
Transformers version 5 is a community endeavor, and this is the last mile. Let's ship this together!
[!NOTE] 👀 Nothing is final and things are still actively in movement. We have a section dedicated to what is planned for future release candidates, yet is known not to work in the RC0. Look for "Disclaimers for the RC0".
We'll be eagerly awaiting your feedback in our GitHub issues!
We introduce a new weight loading API in transformers, which significantly improves on the previous API. This
weight loading API is designed to apply operations to the checkpoints loaded by transformers.
Instead of loading the checkpoint exactly as it is serialized within the model, these operations can reshape, merge, and split the layers according to how they're defined in this new API. These operations are often a necessity when working with quantization or parallelism algorithms.
This new API is centered around the new WeightConverter class:
class WeightConverter(WeightTransform):
operations: list[ConversionOps]
source_keys: Union[str, list[str]]
target_keys: Union[str, list[str]]
The weight converter is designed to apply a list of operations on the source keys, resulting in target keys. A common operation done on the attention layers is to fuse the query, key, values layers. Doing so with this API would amount to defining the following conversion:
conversion = WeightConverter(
["self_attn.q_proj", "self_attn.k_proj", "self_attn.v_proj"], # The input layers
"self_attn.qkv_proj", # The single layer as output
operations=[Concatenate(dim=0)],
)
In this situation, we apply the Concatenate operation, which accepts a list of layers as input and returns a single
layer.
This allows us to define a mapping from architecture to a list of weight conversions. Applying those weight conversions
can apply arbitrary transformations to the layers themselves. This significantly simplified the from_pretrained method
and helped us remove a lot of technical debt that we accumulated over the past few years.
This results in several improvements:
While this is being implemented, expect varying levels of support across different release candidates.
Linked PR: https://github.com/huggingface/transformers/pull/41580
Just as we moved towards a single backend library for model definition, we want our tokenizers, and the Tokenizer object to be a lot more intuitive. With v5, tokenizer definition is much simpler; one can now initialize an empty LlamaTokenizer and train it directly on your corpus.
Defining a new tokenizer object should be as simple as this:
from transformers import TokenizersBackend, generate_merges
from tokenizers import pre_tokenizers, Tokenizer
from tokenizers.model import BPE
class Llama5Tokenizer(TokenizersBackend):
def __init__(self, unk_token="<unk>",bos_token="<s>", eos_token="</s>", vocab=None, merges=None ):
if vocab is None:
self._vocab = {
str(unk_token): 0,
str(bos_token): 1,
str(eos_token): 2,
}
else:
self._vocab = vocab
if merges is not None:
self._merges = merges
else:
self._merges = generate_merges(filtered_vocab)
self._tokenizer = Tokenizer(
BPE(vocab=self._vocab, merges=self._merges, fuse_unk=True)
)
self._tokenizer.pre_tokenizer = pre_tokenizers.Metaspace(
replacement="▁", prepend_scheme=_get_prepend_scheme(self.add_prefix_space, self), split=False
)
super().__init__(
tokenizer_object=self._tokenizer,
unk_token=unk_token,
bos_token=bos_token,
eos_token=eos_token,
)
Once the tokenizer is defined as above, you can load it with the following: Llama5Tokenizer(). Doing this returns you an empty, trainable tokenizer that follows the definition of the authors of Llama5 (it does not exist yet :wink:).
The above is the main motivation towards refactoring tokenization: we want tokenizers to behave similarly to models: trained or empty, and with exactly what is defined in their class definition.
Up to now, transformers maintained two parallel implementations for many tokenizers:
tokenization_<model>.py) - Python-based implementations, often using SentencePiece as the backend.tokenization_<model>_fast.py) - Rust-based implementations using the 🤗 tokenizers library.In v5, we consolidate to a single tokenizer file per model: tokenization_<model>.py. This file will use the most appropriate backend available:
sentencepiece library. It inherits from PythonBackend.tokenizers. Basically allows adding tokens.MistralCommon's tokenization library. (Previously known as the MistralCommonTokenizer)The AutoTokenizer automatically selects the appropriate backend based on available files and dependencies. This is transparent, you continue to use AutoTokenizer.from_pretrained() as before. This allows transformers to be future-proof and modular to easily support future backends.
We enable users and tokenizer builders to define their own tokenizers from top to bottom. Tokenizers are usually defined using a backend such as tokenizers, sentencepiece or mistral-common, but we offer the possibility to design the tokenizer at a higher-level, without relying on those backends.
To do so, you can import the PythonBackend (which was previously known as PreTrainedTokenizer). This class encapsulates all the logic related to added tokens, encoding, and decoding.
If you want something even higher up the stack, then PreTrainedTokenizerBase is what PythonBackend inherits from. It contains the very basic tokenizer API features:
encodedecodevocab_sizeget_vocabconvert_tokens_to_idsconvert_ids_to_tokensfrom_pretrainedsave_pretrainedStarting with v5, we now enable initializing blank, untrained tokenizers-backed tokenizers:
from transformers import LlamaTokenizer
tokenizer = LlamaTokenizer()
This tokenizer will therefore follow the definition of the LlamaTokenizer as defined in its class definition. It can then be trained on a corpus as can be seen in the tokenizers documentation.
These tokenizers can also be initialized from vocab and merges (if necessary), like the previous "slow" tokenizers:
from transformers import LlamaTokenizer
vocab = {"<unk>": 0, "<s>": 1, "</s>": 2, "hello": 3, "world": 4}
merges = [("h", "e"), ("l", "l"), ("o", " ")]
tokenizer = LlamaTokenizer(vocab=vocab, merges=merges)
This tokenizer will behave as a Llama-like tokenizer, with an updated vocabulary. This allows comparing different tokenizer classes with the same vocab; therefore enabling the comparison of different pre-tokenizers, normalizers, etc.
⚠️ The vocab_file (as in, a path towards a file containing the vocabulary) cannot be used to initialize the LlamaTokenizer as loading from files is reserved to the from_pretrained method.
The batch_decode and decode methods have been unified to reflect behavior of the encode method. Both single and batch decoding now use the same decode method. See an example of the new behavior below:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("t5-small")
inputs = ["hey how are you?", "fine"]
tokenizer.decode(tokenizer.encode(inputs))
Gives:
- 'hey how are you?</s> fine</s>'
+ ['hey how are you?</s>', 'fine</s>']
We expect encode and decode to behave, as two sides of the same coin: encode, process, decode, should work.
[!NOTE] A common use-case would be:
encode,model.generate,decode. However, usinggeneratewould returnlist[list[int]], which would then be incompatible withdecode.
The encode_plus method is deprecated in favor of the single __call__ method.
apply_chat_template returns BatchEncodingPreviously, apply_chat_template returned input_ids for backward compatibility. Starting with v5, it now consistently returns a BatchEncoding dict like other tokenizer methods.
# v5
messages = [
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hi there!"}
]
# Now returns BatchEncoding with input_ids, attention_mask, etc.
outputs = tokenizer.apply_chat_template(messages, return_tensors="pt")
print(outputs.keys()) # dict_keys(['input_ids', 'attention_mask'])
We simplify the serialization of tokenization attributes:
special_tokens_map.json - special tokens are now stored in tokenizer_config.json.added_tokens.json - added tokens are now stored in tokenizer.json.added_tokens_decoder is only stored when there is no tokenizer.json.When loading older tokenizers, these files are still read for backward compatibility, but new saves use the consolidated format. We're gradually moving towards consolidating attributes to fewer files so that other libraries and implementations may depend on them more reliably.
Several models that had identical tokenizers now import from their base implementation:
These modules will eventually be removed altogether.
Removed T5-specific workarounds
The internal _eventually_correct_t5_max_length method has been removed. T5 tokenizers now handle max length consistently with other models.
A few testing changes specific to tokenizers have been applied:
add_tokens, encode, decode) are now centralized and automatically applied across all tokenizers. This reduces test duplication and ensures consistent behaviorFor legacy implementations, the original BERT Python tokenizer code (including WhitespaceTokenizer, BasicTokenizer, etc.) is preserved in bert_legacy.py for reference purposes.
Special Tokens Structure:
SpecialTokensMixin: Merged into PreTrainedTokenizerBase to simplify the tokenizer architecture.special_tokens_map: Now only stores named special token attributes (e.g., bos_token, eos_token). Use extra_special_tokens for additional special tokens (formerly additional_special_tokens). all_special_tokens includes both named and extra tokens.# v4
tokenizer.special_tokens_map # Included 'additional_special_tokens'
# v5
tokenizer.special_tokens_map # Only named tokens
tokenizer.extra_special_tokens # Additional tokens
special_tokens_map_extended and all_special_tokens_extended: Removed. Access AddedToken objects directly from _special_tokens_map or _extra_special_tokens if needed.additional_special_tokens: Still accepted for backward compatibility but is automatically converted to extra_special_tokens.Deprecated Methods:
sanitize_special_tokens(): Already deprecated in v4, removed in v5.prepare_seq2seq_batch(): Deprecated; use __call__() with text_target parameter instead.# v4
model_inputs = tokenizer.prepare_seq2seq_batch(src_texts, tgt_texts, max_length=128)
# v5
model_inputs = tokenizer(src_texts, text_target=tgt_texts, max_length=128, return_tensors="pt")
model_inputs["labels"] = model_inputs.pop("input_ids_target")
BatchEncoding.words(): Deprecated; use word_ids() instead.Removed Methods:
create_token_type_ids_from_sequences(): Removed from base class. Subclasses that need custom token type ID creation should implement this method directly.clean_up_tokenization(): Removed from base class. Now defined at model class level for models that need it (e.g., PLBart, CLVP, Wav2Vec2).prepare_for_model(), build_inputs_with_special_tokens(), truncate_sequences(): Moved from tokenization_utils_base.py to tokenization_python.py for PythonBackend tokenizers. TokenizersBackend provides model-ready input via tokenize() and encode(), so these methods are no longer needed in the base class._switch_to_input_mode(), _switch_to_target_mode(), as_target_tokenizer(): Removed from base class. Use __call__() with text_target parameter instead.# v4
with tokenizer.as_target_tokenizer():
labels = tokenizer(tgt_texts, ...)
# v5
labels = tokenizer(text_target=tgt_texts, ...)
parse_response(): Removed from base class.Because we are switching from the naive MOE (nn.ModuleList for experts) we currently have an issue with MoEs that have adapters. For more details see https://github.com/huggingface/transformers/issues/42491#issuecomment-3591485649.
We aim for this to be fixed and released in a following release candidate in the week that follows RC0.
We are streamlining the MoE support with vLLM; while this is being implemented, tensor parallelism and expert parallelism aren't working as expected. This is known and actively being worked on.
We aim for this to be fixed and released in a following release candidate in the week that follows RC0.
For anyone inheriting from a transformers PreTrainedModel, the weights are automatically initialized with the common scheme:
@torch.no_grad()
def _init_weights(self, module):
"""
Initialize the weights. This is quite general on purpose, in the spirit of what we usually do. For more complex
initialization scheme, it should be overridden by the derived `PreTrainedModel` class. In case a model adds an explicit
`nn.Parameter`, this method should also be overridden in order to initialize it correctly.
"""
if hasattr(self.config, "initializer_range"):
std = self.config.initializer_range or 0.02
elif hasattr(self.config, "init_std"):
std = self.config.init_std
elif hasattr(self.config, "initializer_factor"):
std = self.config.initializer_factor
else:
# 0.02 is the standard default value across the library
std = getattr(self.config.get_text_config(), "initializer_range", 0.02)
if isinstance(module, (nn.Linear, nn.Conv1d, nn.Conv2d, nn.Conv3d, nn.ConvTranspose1d, nn.ConvTranspose2d)):
if getattr(module, "weight", None) is not None:
init.normal_(module.weight, mean=0.0, std=std)
if getattr(module, "bias", None) is not None:
init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
if getattr(module, "weight", None) is not None:
init.normal_(module.weight, mean=0.0, std=std)
# Here we need the check explicitly, as we slice the weight in the `zeros_` call, so it looses the flag
if module.padding_idx is not None and not getattr(module.weight, "_is_hf_initialized", False):
init.zeros_(module.weight[module.padding_idx])
elif isinstance(module, nn.MultiheadAttention):
# This uses torch's original init
module._reset_parameters()
# We cannot use `isinstance` on the RMSNorms or LayerNorms, as they usually are custom modules which change names
# between modelings (because they are prefixed with the model name)
elif (
isinstance(module, (nn.GroupNorm, nn.BatchNorm1d, nn.BatchNorm2d, nn.BatchNorm3d))
or "LayerNorm" in module.__class__.__name__
or "RMSNorm" in module.__class__.__name__
):
# Norms can exist without weights (in which case they are None from torch primitives)
if hasattr(module, "weight") and module.weight is not None:
init.ones_(module.weight)
if hasattr(module, "bias") and module.bias is not None:
init.zeros_(module.bias)
If you want to avoid that, for now you should just do:
class CustomModel(Qwen3VLForConditionalGeneration):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.action_head = nn.Linear(1024, 7)
self.positional_embedding = nn.Parameter(torch.randn(16, 1152))
self.post_init()
def _init_weights(self, module):
pass
There is a tracker for that here: https://github.com/huggingface/transformers/issues/42418.
use_auth_tokenThe use_auth_token argument/parameter is deprecated in favor of token everywhere.
You should be able to search and replace use_auth_token with token and get the same logic.
Linked PR: https://github.com/huggingface/transformers/pull/41666
We decided to remove some features for the upcoming v5 as they are currently only supported in a few old models and no longer integrated in current model additions. It's recommended to stick to v4.x in case you need them. Following features are affected:
We dropped support for two torch APIs:
torchscript in https://github.com/huggingface/transformers/pull/41688torch.fx in https://github.com/huggingface/transformers/pull/41683Those APIs were deprecated by the PyTorch team, and we're instead focusing on the supported APIs dynamo and export.
We clean up the quantization API in transformers, and significantly refactor the weight loading as highlighted above.
We drop support for two quantization arguments that have been deprecated for some time:
load_in_4bitload_in_8bitWe remove them in favor of the quantization_config argument which is much more complete. As an example, here is how
you would load a 4-bit bitsandbytes model using this argument:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(load_in_4bit=True)
model_4bit = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-3B",
device_map="auto",
quantization_config=quantization_config
)
from_xxx_config are deleted. Configs can be init from the __init__ method in the same way. See #41314.mode.rope_parameters, including the rope_theta and rope_type. Model's config.rope_parameters is a simple dictionaty in most cases, and can also be a nested dict in special cases (i.e. Gemma3 and ModernBert) with different rope parameterization for each layer type. Trying to get config.rope_theta will throw an attribute error from now on. See #39847 and #42255config.vocab_size). Users are expected to access keys from their respective sub-configs (config.text_config.vocab_size).model.generate()) will no longer have a generation_config and model.config.generation_config will throw an attribute error.tokenization_<model>.py ) will be removed in favor of using fast tokenizer files tokenization_<model>_fast.py --> will be renamed to tokenization_<model>.py. As fast tokenizers are :hugs:tokenizers - backend, they include a wider range of features that are maintainable and reliable.encode_plus --> __call__batch_decode --> decodeapply_chat_template by default returns naked input_ids rather than a BatchEncoding dict.
This was inconvenient - it should return a BatchEncoding dict like tokenizer.__call__(), but we were stuck with
it for backward compatibility. The method now returns a BatchEncoding.
Linked PRs:
processor_config.json as a nested dict, instead of serializing attributes in their own config files. Loading will be supported for all old format processors (https://github.com/huggingface/transformers/pull/41474)XXXFeatureExtractors classes are completely removed in favor of XXXImageProcessor class for all vision models (https://github.com/huggingface/transformers/pull/41174)XXXFastImageProcessorKwargs is removed in favor of XXXImageProcessorKwargs which will be shared between fast and slow processors (https://github.com/huggingface/transformers/pull/40931)RotaryEmbeddings layers will start returning a dict of tuples, in case the model uses several RoPE configurations (Gemma2, ModernBert). Each value will be a tuple of "cos, sin" per RoPE type.RotaryEmbeddings layer will be unified and accessed via config.rope_parameters. Config attr for rope_theta might not be accessible anymore for some models, and instead will be in config.rope_parameters['rope_theta']. BC will be supported for a while as much as possible, and in the near future we'll gradually move to the new RoPE format (https://github.com/huggingface/transformers/pull/39847)model.language_model. It is recommended to either access the module with model.model.language_model or model.get_decoder(). See #42156GreedySearchEncoderDecoderOutput). We now only have 4 output classes built from the following matrix: decoder-only vs encoder-decoder, uses beams vs doesn't use beams (https://github.com/huggingface/transformers/pull/40998)generate doesn't receive any KV Cache argument, the default cache class used is now defined by the model (as opposed to always being DynamicCache) (https://github.com/huggingface/transformers/pull/41505)config.json for any old model, it will be loaded back into model's generation config. Users are expected to access or modify generation parameters only with model.generation_config.do_sample = True.compute_loss_func Handling
compute_loss_func now always takes priority over the model's built-in loss computation, giving users consistent control over custom loss functions.num_items_in_batch in Prediction Step
num_items_in_batch argument is now passed to compute_loss during prediction_step, enabling proper loss scaling during evaluation.report_to now defaults to "none"
TrainingArguments due to low usagemp_parameters -> legacy param that was later on added to the Sagemaker trainer_n_gpu -> not intended for users to set, we will initialize it correctly instead of putting it in the TrainingArgumentsoverwrite_output_dir - > replaced by resume_from_checkpoint, and it was only used in the examples script, no impact on Trainer.logging_dir -> only used for tensorboard, set TENSORBOARD_LOGGING_DIR env var insteadjit_mode_eval -> use use_torch_compile instead, as torchscript is not recommended anymoretpu_num_cores-> It is actually better to remove it, as it is not recommended to set the number of cores. By default, all TPU cores are used . Set TPU_NUM_CORES env var insteadpast_index -> it was only used for a very small number of models that have special architecture like transformersxl + it was not documented at all how to train those modelsray_scope -> only for a minor arg for ray integration. Set RAY_SCOPE var env insteadwarmup_ratio -> use warmup_step instead. We combined both args together by allowing passing float values in warmup_step.TrainingArgumentsfsdp_min_num_params and fsdp_transformer_layer_cls_to_wrap -> use fsdp_configtpu_metrics_debug -> debugpush_to_hub_token -> hub_tokenpush_to_hub_model_id and push_to_hub_organization -> hub_model_idinclude_inputs_for_metrics -> include_for_metricsper_gpu_train_batch_size -> per_device_train_batch_sizeper_gpu_eval_batch_size -> per_device_eval_batch_sizeuse_mps_device -> mps will be used by default if detectedfp16_backend and half_precision_backend -> we will only rely on torch.amp as everything has been upstreamed to torchno_cuda -> use_cpu include_tokens_per_second -> include_num_input_tokens_seenuse_legacy_prediction_loop -> we only use evaluation_loop function from now onTrainertokenizer in initialization -> processing_classmodel_path in train() -> resume_from_checkpointTrainerTraineruse_cache in the model config will be set to False. You can still change the cache value through TrainingArguments usel_cache argument if needed.organization and repo_url from PushToHubMixin. You must pass a repo_id instead.ignore_metadata_errors from PushToMixin. In practice if we ignore errors while loading the model card, we won't be able to push the card back to the Hub so it's better to fail early and not provide the option to fail later.push_to_hub do not accept **kwargs anymore. All accepted parameters are explicitly documented.push_to_hub are now keyword-only to avoid confusion. Only repo_id can be positional since it's the main arg.use_temp_dir argument from push_to_hub. We now use a tmp dir in all cases.Linked PR: https://github.com/huggingface/transformers/pull/42391.
The deprecated transformers-cli ... command was deprecated, transformers ... is now the only CLI entry point.
transformers CLI has been migrated to Typer, making it easier to maintain + adding some nice features out of
the box (improved --help section, autocompletion).
Biggest breaking change is in transformers chat. This command starts a terminal UI to interact with a chat model.
It used to also be able to start a Chat Completion server powered by transformers and chat with it. In this revamped
version, this feature has been removed in favor of transformers serve. The goal of splitting transformers chat
and transformers serve is to define clear boundaries between client and server code. It helps with maintenance
but also makes the commands less bloated. The new signature of transformers chat is:
Usage: transformers chat [OPTIONS] BASE_URL MODEL_ID [GENERATE_FLAGS]...
Chat with a model from the command line.
It works hand in hand with transformers serve, which means that if transformers serve is running on its default endpoint, transformers chat can be launched as follows:
transformers chat HuggingFaceTB/SmolLM3-3B
It can however use any OpenAI API compatible HTTP endpoint:
transformers chat HuggingFaceTB/SmolLM3-3B https://router.huggingface.co/v1
Linked PRs:
run methodThe transformers run (previously transformers-cli run) is an artefact of the past, was not documented nor tested,
and isn't part of any public documentation. We're removing it for now and ask you to please let us know in case
this is a method you are using; in which case we should bring it back with better support.
Linked PR: https://github.com/huggingface/transformers/pull/42447
TRANSFORMERS_CACHE, PYTORCH_TRANSFORMERS_CACHE, and PYTORCH_PRETRAINED_BERT_CACHE have been removed. Please use HF_HOME instead.HUGGINGFACE_CO_EXAMPLES_TELEMETRY, HUGGINGFACE_CO_EXAMPLES_TELEMETRY, HUGGINGFACE_CO_PREFIX, and HUGGINGFACE_CO_RESOLVE_ENDPOINT have been removed. Please use huggingface_hub.constants.ENDPOINT instead.Linked PR: https://github.com/huggingface/transformers/pull/42391.
transformers v5 pins the huggingface_hub version to >=1.0.0. See this migration guide to learn more about this major release. Here are to main aspects to know about:
requests to httpx. This change was made to improve performance and to support both synchronous and asynchronous requests the same way. If you are currently catching requests.HTTPError errors in your codebase, you'll need to switch to httpx.HTTPError.HTTP_PROXY / HTTPS_PROXY environment variableshf_transfer and therefore HF_HUB_ENABLE_HF_TRANSFER have been completed dropped in favor of hf_xet. This should be transparent for most users. Please let us know if you notice any downside!typer-slim has been added as required dependency, used to implement both hf and transformers CLIs.
<img width="809" height="471" alt="image" src="https://github.com/user-attachments/assets/58bb9c70-d481-48ed-ab8f-6553be7c240f" />
The Code World Model (CWM) model was proposed in CWM: An Open-Weights LLM for Research on Code Generation with World Models by Meta FAIR CodeGen Team. CWM is an LLM for code generation and reasoning about code that has, in particular, been trained to better represent and reason about how code and commands affect the state of a program or system. Specifically, we mid-trained CWM on a large number of observation-action trajectories from Python execution traces and agentic interactions in containerized environments. We post-trained with extensive multi-task RL in verifiable coding, math, and multi-turn software engineering environments.
<img width="1505" height="915" alt="image" src="https://github.com/user-attachments/assets/eec48633-f02b-464a-ae5c-c65473387e53" />
SAM3 (Segment Anything Model 3) was introduced in SAM 3: Segment Anything with Concepts.
The SAM3 addition adds four new architectures:
SAM3 performs Promptable Concept Segmentation (PCS) on images. PCS takes text and/or image exemplars as input (e.g., "yellow school bus"), and predicts instance and semantic masks for every single object matching the concept.
Sam3Tracker and Sam3TrackerVideo perform Promptable Visual Segmentation (PVS) on images. PVS takes interactive visual prompts (points, boxes, masks) or text inputs to segment a specific object instance per prompt. This is the task that SAM 1 and SAM 2 focused on, and SAM 3 improves upon it. Sam3Tracker and Sam3TrackerVideo are updated versions of SAM2 Video that maintain the same API while providing improved performance and capabilities.
SAM3 Video performs Promptable Concept Segmentation (PCS) on videos. PCS takes text as input (e.g., "yellow school bus"), and predicts instance and semantic masks for every single object matching the concept, while preserving object identities across video frames. The model combines a detection module (SAM3) with a tracking module (SAM2-style tracker) to enable robust object tracking across video frames using text prompts.
<img width="1080" height="849" alt="image" src="https://github.com/user-attachments/assets/a9fa1b81-114d-4054-9699-5083ac69d830" />
LFM2-MoE is a Mixture-of-Experts (MoE) variant of LFM2. The LFM2 family is optimized for on-device inference by combining short‑range, input‑aware gated convolutions with grouped‑query attention (GQA) in a layout tuned to maximize quality under strict speed and memory constraints.
LFM2‑MoE keeps this fast backbone and introduces sparse MoE feed‑forward networks to add representational capacity without significantly increasing the active compute path. The first LFM2-MoE release is LFM2-8B-A1B, with 8.3B total parameters and 1.5B active parameters. The model excels in quality (comparable to 3-4B dense models) and speed (faster than other 1.5B class models).
<img width="812" height="366" alt="image" src="https://github.com/user-attachments/assets/21c82c6e-cf0a-4d6c-a707-b9e57663ca85" />
The VideoLLaMA3 model is a major update to VideoLLaMA2 from Alibaba DAMO Academy.
<img width="621" height="475" alt="image" src="https://github.com/user-attachments/assets/c9616758-b3aa-41d0-bd58-695966ba146d" />
Audio Flamingo 3 (AF3) is a fully open large audio–language model designed for robust understanding and reasoning over speech, environmental sounds, and music. AF3 pairs a Whisper-style audio encoder with a causal language model and performs replace-in-place audio–text fusion: the processor aligns post-pool audio frames to a dedicated placeholder token and the model replaces those token slots with projected audio embeddings during the forward pass.
The model checkpoint is available at: nvidia/audio-flamingo-3-hf
Highlights:
NanoChat is a compact decoder-only transformer model designed for educational purposes and efficient training. The model features several fundamental architectural innovations which are common in modern transformer models. Therefore, it is a good model to use as a starting point to understand the principles of modern transformer models. NanoChat is a variant of the Llama architecture, with simplified attention mechanism and normalization layers.
JetMoe Fix jetmoe after #40132 by @ArthurZucker in #41324gemma3 by @Sai-Suraj-27 in #41354PretrainedConfig to PreTrainedConfig by @Cyrilvallez in #41300ModularChecker] QOL for the modular checker by @ArthurZucker in #41361v5] Remove relative position embeddings (for bert like models) by @vasqu in #41170apply_chat_template by @Samoed in #41355test_longcat_generation_cpu by @ydshieh in #41368CB] Refactors the way we access paged by @ArthurZucker in #41370v5] Sync Bert and Bart eager attention by @vasqu in #41248TypeError exception for invalid type by @Sai-Suraj-27 in #41346update_device_map for GPTQ quantizer by @Sai-Suraj-27 in #41328prune_heads by @gante in #41417JetMoe] Fix KV head repetition and padding free by @vasqu in #41423JetMoeIntegrationTest by @ydshieh in #41377past_key_value in BERT-like models by @zucchini-nlp in #41448utils/tf_ops/ by @gante in #41402Attention Masks] Bidirectional masks for encoder and encoder-decoder models by @vasqu in #41265past_index by @SunMarc in #41384report_to default changed to "none" + cleaning deprecated env var by @SunMarc in #41375overwrite_output_dir by @SunMarc in #41323CI] Fix copies on main by @vasqu in #41486jit_mode_eval by @SunMarc in #41376local_rank arg from TrainingArguments by @SunMarc in #41382pickle - BloomTokenizerFast by @ydshieh in #41466glm4v by @Sai-Suraj-27 in #41483truncation to False in Qwen3Omni to avoid default truncation by @BakerBunker in #41473local_rank deletion and some cleaning by @SunMarc in #41504tpu_num_cores by @SunMarc in #41383HunYuanMoEV1IntegrationTest:test_model_generation by @ydshieh in #41373generate delegates default cache initialization to the model by @gante in #41505from_pretrained] Small refactor from_pretrained: move around unrelated stuff by @ArthurZucker in #41445transformers serve by @LysandreJik in #41446logits_to_keep to many older CausalLM models by @philiproeleveld in #41335torch.compile recompiled part of th… by @sywangyi in #41558Docs] Fix changed references by @vasqu in #41614expand_device_map instead of redefining it by @Cyrilvallez in #41608tp_plan in from_pretrained directly by @Cyrilvallez in #41435Executorch] Simplify for encoder models by @vasqu in #41627Ernie 4.5 Moe] Fix Moe and offloading by @vasqu in #41385Masks] Fix mask handling in eager for vision models by @vasqu in #41625utils/check_bad_commit.py by @ydshieh in #41658use_cache default to False by @SunMarc in #41585chat_extras.md to Korean by @Judy-Choi in #39863big_bird.md to Korean by @ssum21 in #40445code_llama.md to Korean by @Judy-Choi in #40558ko-LFM2.md to Korean by @ssum21 in #41502use_auth_token parameter by @Wauplin in #41666Attn] Allow dynamic causality in SDPA via Kwargs by @vasqu in #41692run_name docs in TrainingArguments by @tobiasofsn in #41705utils/check_bad_commit.py by @ydshieh in #41658)videos from image processing classes by @zucchini-nlp in #41607@staticmethod from module-level get_device_and_memory_breakdown by @albertvillanova in #41747Onnx docs] Remove some traces by @vasqu in #41791utils/check_bad_commit.py by @ydshieh in #41658)Clip] Fix masking and enable flash attention on all model types by @vasqu in #41750test_tensor_parallel.py by @3outeille in #41918detectron2 installation in docker files by @ydshieh in #41975autoawq[kernels] installation in quantization docker file by @ydshieh in #41978torchcodec version in quantization docker file by @ydshieh in #41988run slow v2: empty report when there is only one model by @ydshieh in #42002torch+deepspeed docker file by @ydshieh in #41985logging_dir by @SunMarc in #42013deeepspeed in AMD docker file by @ydshieh in #42025huggingface_hub dependency version by @hanouticelina in #42033pr_slow_ci_suggestion.yml after #42023 by @ydshieh in #42049Argument list too long in pr_slow_ci_suggestion.yml by @ydshieh in #42061setattr as well by @zucchini-nlp in #41808Attn Masks] Non-vmap default for attention masks by @vasqu in #41852image_transforms.py by @yaswanth19 in #42044prepare_inputs_for_generation cache slicing condition by @albertvillanova in #41764T5Gemma] Fix cross attention cache by @vasqu in #41890streaming by @McPatate in #42102pytest<9 for now by @ydshieh in #42162Pop2Piano] Fix cache usage by @vasqu in #42170PEFT] Fix prefix tuning by @vasqu in #41696FqnToConfig by @jcaip in #41894PEFT] Fix the general test for prefix tuning by @vasqu in #42185Pop2Piano] Fix tied weights by @vasqu in #42193BLT] Fix cache usage by @vasqu in #42188test_dynamic_cache_exportability_multiple_run (failing on torch 2.10 nightly) by @ydshieh in #42212AttentionMaskConverter._unmask_unattended for xpu device before by @kaixuanliu in #42230base_model by @zucchini-nlp in #41589batch_size by @ydshieh in #42213batch_size" by @ydshieh in #42258get_decoder() for multimodal and delete redundant code 🔪 by @zucchini-nlp in #42156cwm by @ydshieh in #42261torch.get_autocast_dtype instead of torch.get_autocast_gpu_dtype by @qgallouedec in #42055WhisperFeatureExtractor by @TopCoder2K in #42286CI] Skip EfficientLoFTR test by @vasqu in #42327Attn Masks] Lift bidirectional mask restriction on eager by @vasqu in #42325torch.distributed imports by @Cyrilvallez in #42361Attn Masks] Add skip option for non-packed sequences by @vasqu in #42367Mistral Tokenizers] Fix tokenizer detection by @vasqu in #42389Note truncated.
Another fix for qwen vl models that prevented correctly loading the associated model type - this works together with https://github.com/huggingface/tr
Another fix for qwen vl models that prevented correctly loading the associated model type - this works together with https://github.com/huggingface/transformers/pull/41808 of the previous patch release.
Full Changelog: https://github.com/huggingface/transformers/compare/v4.57.5...v4.57.6
Should not have said last patch :wink: These should be the last remaining fixes that got lost in between patches and the transition to v5.
Should not have said last patch :wink: These should be the last remaining fixes that got lost in between patches and the transition to v5.
Full Changelog: https://github.com/huggingface/transformers/compare/v4.57.4...v4.57.5
Last patch release for v4: We have a few small fixes for remote generation methods (e.g. group beam search), vLLM, and an offline tokenizer fix (if it
Last patch release for v4: We have a few small fixes for remote generation methods (e.g. group beam search), vLLM, and an offline tokenizer fix (if it's already been cached).
Full Changelog: https://github.com/huggingface/transformers/compare/v4.57.3...v4.57.4
There was a hidden bug when loading models with local_files_only=True and a typo related to the recent patch.
There was a hidden bug when loading models with local_files_only=True and a typo related to the recent patch.
The main fix is: https://github.com/huggingface/transformers/commit/b6055550a15a8fab367cf983b743ff68cc58d81a.
We are really sorry that this slipped through, our CIs just did not catch it.
As it affects a lot of users we are gonna yank the previous release
This patch most notably fixes an issue on some Mistral tokenizers. It contains the following commits:
This patch most notably fixes an issue on some Mistral tokenizers. It contains the following commits:
@staticmethod from module-level get_device_and_memory_breakdown (#41747)This patch most notably fixes an issue with an optional dependency (optax), which resulted in parsing errors with poetry. It contains the following fi
This patch most notably fixes an issue with an optional dependency (optax), which resulted in parsing errors with poetry. It contains the following fixes:
[deprecations] Remove generate-related deprecations up to v4.56 by @gante in #40729
<img width="1200" height="511" alt="image" src="https://github.com/user-attachments/assets/3abad6c4-5650-412d-a831-f8a30a5d962e" />
The Qwen3-Next series represents the Qwen team's next-generation foundation models, optimized for extreme context length and large-scale parameter efficiency. The series introduces a suite of architectural innovations designed to maximize performance while minimizing computational cost:
Built on this architecture, they trained and open-sourced Qwen3-Next-80B-A3B — 80B total parameters, only 3B active — achieving extreme sparsity and efficiency.
Despite its ultra-efficiency, it outperforms Qwen3-32B on downstream tasks — while requiring less than 1/10 of the training cost. Moreover, it delivers over 10x higher inference throughput than Qwen3-32B when handling contexts longer than 32K tokens.
For more details, please visit their blog Qwen3-Next (blog post).
<img width="1282" height="392" alt="image" src="https://github.com/user-attachments/assets/9412905b-4083-4994-9000-aa0dbf97eb6f" />
VaultGemma is a text-only decoder model derived from Gemma 2, notably it drops the norms after the Attention and MLP blocks, and uses full attention for all layers instead of alternating between full attention and local sliding attention. VaultGemma is available as a pretrained model with 1B parameters that uses a 1024 token sequence length.
VaultGemma was trained from scratch with sequence-level differential privacy (DP). Its training data includes the same mixture as the Gemma 2 models, consisting of a number of documents of varying lengths. Additionally, it is trained using DP stochastic gradient descent (DP-SGD) and provides a (ε ≤ 2.0, δ ≤ 1.1e-10)-sequence-level DP guarantee, where a sequence consists of 1024 consecutive tokens extracted from heterogeneous data sources. Specifically, the privacy unit of the guarantee is for the sequences after sampling and packing of the mixture.
<img width="3544" height="1886" alt="image" src="https://github.com/user-attachments/assets/5afa70cb-506e-4d56-baa3-30e7522ac653" />
Qwen3-VL is a multimodal vision-language model series, encompassing both dense and MoE variants, as well as Instruct and Thinking versions.
Building upon its predecessors, Qwen3-VL delivers significant improvements in visual understanding while maintaining strong pure text capabilities. Key architectural advancements include: enhanced MRope with interleaved layout for better spatial-temporal modeling, DeepStack integration to effectively leverage multi-level features from the Vision Transformer (ViT), and improved video understanding through text-based time alignment—evolving from T-RoPE to text timestamp alignment for more precise temporal grounding.
These innovations collectively enable Qwen3-VL to achieve superior performance in complex multimodal tasks.
<img width="763" height="468" alt="image" src="https://github.com/user-attachments/assets/289d33e0-6c71-458d-ae07-b7d454ac2adf" />
The LongCatFlash model was proposed in LongCat-Flash Technical Report by the Meituan LongCat Team. LongCat-Flash is a 560B parameter Mixture-of-Experts (MoE) model that activates 18.6B-31.3B parameters dynamically (average ~27B). The model features a shortcut-connected architecture enabling high inference speed (>100 tokens/second) and advanced reasoning capabilities.
The abstract from the paper is the following:
We present LongCat-Flash, a 560 billion parameter Mixture-of-Experts (MoE) language model featuring a dynamic computation mechanism that activates 18.6B-31.3B parameters based on context (average ~27B). The model incorporates a shortcut-connected architecture enabling high inference speed (>100 tokens/second) and demonstrates strong performance across multiple benchmarks including 89.71% accuracy on MMLU and exceptional agentic tool use capabilities.
Tips:
<img width="700" height="414" alt="image" src="https://github.com/user-attachments/assets/7b92ee0f-5f5a-459c-ad4d-e01b5c10202e" />
FlexOlmo is a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on closed datasets, and (2) data-flexible inference, where these parameters along with their associated data can be flexibly included or excluded from model inferences with no further training. FlexOlmo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on closed datasets and later integrated through a new domain-informed routing without any joint training. FlexOlmo is trained on FlexMix, a corpus we curate comprising publicly available datasets alongside seven domain-specific sets, representing realistic approximations of closed sets.
You can find all the original FlexOlmo checkpoints under the FlexOlmo collection.
<img width="2300" height="1400" alt="image" src="https://github.com/user-attachments/assets/ef0605cd-9512-458c-915a-62316e14d90c" />
LFM2-VL first series of vision-language foundation models developed by Liquid AI. These multimodal models are designed for low-latency and device-aware deployment. LFM2-VL extends the LFM2 family of open-weight Liquid Foundation Models (LFMs) into the vision-language space, supporting both text and image inputs with variable resolutions.
LFM2-VL consists of three main components: a language model backbone, a vision encoder, and a multimodal projector. LFM2-VL builds upon the LFM2 backbone, inheriting from either LFM2-1.2B (for LFM2-VL-1.6B) or LFM2-350M (for LFM2-VL-450M). For the vision tower, LFM2-VL uses SigLIP2 NaFlex encoders to convert input images into token sequences. Two variants are implemented:
The encoder processes images at their native resolution up to 512×512 pixels, efficiently handling smaller images without upscaling and supporting non-standard aspect ratios without distortion. Larger images are split into non-overlapping square patches of 512×512 each, preserving detail. In LFM2-VL-1.6B, the model also receives a thumbnail (a small, downscaled version of the original image capturing the overall scene) to enhance global context understanding and alignment. Special tokens mark each patch’s position and indicate the thumbnail’s start. The multimodal connector is a 2-layer MLP connector with pixel unshuffle to reduce image token count.
<img width="1448" height="1062" alt="image" src="https://github.com/user-attachments/assets/af1fbb09-082c-4331-9217-357adb506cbf" />
The BLT model was proposed in Byte Latent Transformer: Patches Scale Better Than Tokens by Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li1, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman†, Srinivasan Iyer. BLT is a byte-level LLM that achieves tokenization-level performance through entropy-based dynamic patching.
The abstract from the paper is the following:
We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency and robustness. BLT encodes bytes into dynamically sized patches, which serve as the primary units of computation. Patches are segmented based on the entropy of the next byte, allocating more compute and model capacity where increased data complexity demands it. We present the first flop controlled scaling study of byte-level models up to 8B parameters and 4T training bytes. Our results demonstrate the feasibility of scaling models trained on raw bytes without a fixed vocabulary. Both training and inference efficiency improve due to dynamically selecting long patches when data is predictable, along with qualitative improvements on reasoning and long tail generalization. Overall, for fixed inference costs, BLT shows significantly better scaling than tokenization-based models, by simultaneously growing both patch and model size.
Dual Model Architecture: BLT consists of two separate trained models:
Dynamic Patching: The model uses entropy-based dynamic patching where:
Local Encoder: Processes byte sequences with cross-attention to patch embeddings
Global Transformer: Processes patch-level representations with full attention across patches
Local Decoder: Generates output with cross-attention back to the original byte sequence
Byte-Level Tokenizer: Unlike traditional tokenizers that use learned vocabularies, BLT's tokenizer simply converts text to UTF-8 bytes and maps each byte to a token ID. There is no need for a vocabulary.
<img width="14084" height="7429" alt="image" src="https://github.com/user-attachments/assets/20d46a43-15f2-42bf-9703-9575f5ca4430" />
The Qwen2.5-Omni model is a unified multiple modalities model proposed in Qwen2.5-Omni Technical Report from Qwen team, Alibaba Group.
Qwen2_5OmniForConditionalGeneration] to generate audio and text output. To generate only one output type, use [Qwen2_5OmniThinkerForConditionalGeneration] for text-only and [Qwen2_5OmniTalkersForConditionalGeneration] for audio-only outputs.Qwen2_5OmniForConditionalGeneration] supports only single batch size at the moment.processor.max_pixels. By default the maximum is set to a very arge value and high resolution visuals will not be resized, unless resolution exceeds processor.max_pixels.~ProcessorMixin.apply_chat_template] method to convert chat messages to model inputs.<img width="1431" height="527" alt="image" src="https://github.com/user-attachments/assets/e831f451-9be3-4b5c-a222-b833a50ceb2a" />
Parakeet models, introduced by NVIDIA NeMo, are models that combine a Fast Conformer encoder with connectionist temporal classification (CTC), recurrent neural network transducer (RNNT) or token and duration transducer (TDT) decoder for automatic speech recognition.
Model Architecture
ParakeetEncoder] for the encoder implementation and details).<img width="949" height="537" alt="image" src="https://github.com/user-attachments/assets/5ca4e73d-5aa9-487d-96e1-92d4f2f4739f" />
The EdgeTAM model was proposed in EdgeTAM: On-Device Track Anything Model Chong Zhou, Chenchen Zhu, Yunyang Xiong, Saksham Suri, Fanyi Xiao, Lemeng Wu, Raghuraman Krishnamoorthi, Bo Dai, Chen Change Loy, Vikas Chandra, Bilge Soran.
EdgeTAM is an efficient adaptation of SAM 2 that introduces a 2D Spatial Perceiver architecture to optimize memory attention mechanisms for real-time video segmentation on mobile devices.
More details to come soon :eyes:
We are introducing Continuous Batching (CB) in this release, we consider it a stable feature. The main use case for CB is batched generation, which makes it very efficient in the context of GRPO training or evaluation. Thanks to CB, researchers or model developers are now free to use transformers in these contexts without having to spin up an additional inference engine.
CB currently supports both full attention and sliding window attention: this means that the vast majority of models are supported, like llama, gemma3, gpt-oss.
CB is also integrated with transformers serve, which means that you can deploy transformers as an OpenAI-compatible HTTP server. Here is a small snippet on how to use it:
import datasets
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers.generation import GenerationConfig
model = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-4B-Instruct-2507", dtype=torch.bfloat16, _attn_implementation="sdpa_paged", device_map="auto"
)
model.generation_config.max_new_tokens = 32
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507", padding_side="left")
dataset = datasets.load_dataset("openai/gsm8k", "socratic", split="test")
tokenized_datasets = dataset.map(lambda x: tokenizer(x["question"]), batched=True)
simple_batch_inputs = [item["input_ids"] for item in tokenized_datasets]
batch_outputs = model.generate_batch(inputs=simple_batch_inputs)
for request in batch_outputs:
print(tokenizer.decode(batch_outputs[request].generated_tokens))
"""
Let's break down the problem step by step:
1. **Total eggs laid per day**:
Janet’s ducks lay **16 eggs per day**
Let's break down the problem step by step:
1. **Blue fiber**: The robe takes **2 bolts** of blue fiber.
2. **White fiber
To determine Josh's profit from flipping the house, let's go step by step.
---
### Step 1: Initial cost of the house
Josh buys the
To find the total distance James runs in a week, we can break down the problem step by step:
1. **Sprints per session**: James runs
To determine how many cups of feed Wendi needs to give her chickens in the final meal of the day, let's go step by step.
"""
check_model_inputs in core VLMs by @zucchini-nlp in #40342input_feature length and attention_mask length in WhisperFeatureExtractor by @BakerBunker in #39221_prepare_generation_config by @manueldeprada in #40715center_crop fast equivalent to slow by @yonigozlan in #40856pytest-rerunfailures<16.0 by @ydshieh in #40561test_all_params_have_gradient=False for DeepseekV2ModelTest by @ydshieh in #40566test_eager_matches_sdpa_inference not run for CLIP by @ydshieh in #40581remi-or to run-slow by @ydshieh in #40590get_*_features methods + update doc snippets by @qubvel in #40555TvpImageProcessingTest::test_slow_fast_equivalence by @ydshieh in #40593siglip flaky test_eager_matches_sdpa_inference by @ydshieh in #40584Tests] Fixup duplicated mrope logic by @vasqu in #40592TokenizerTesterMixin temporarily by @ydshieh in #40611transformers serve by @McPatate in #40479too many request caused by AutoModelTest::test_dynamic_saving_from_local_repo by @ydshieh in #40614JambaModelTest.test_load_balancing_loss by @ydshieh in #40617deepseek_v3.md to Korean by @ssum21 in #39649too many requests in TestMistralCommonTokenizer by @ydshieh in #40623test_prompt_lookup_decoding_matches_greedy_search for voxtral by @ydshieh in #40643LongformerModelTest::test_attention_outputs as flaky by @ydshieh in #40655custom_generate Callables and unify generation args structure by @manueldeprada in #40586check_determinism inside test_determinism by @ydshieh in #40661test_fast_is_faster_than_slow for Owlv2ImageProcessingTest by @ydshieh in #40663test_prompt_lookup_decoding_matches_greedy_search for qwen2_audio by @ydshieh in #40664GitModelTest::test_beam_search_generate by @ydshieh in #40666tolist instead of list comprehension calling .item() by @McPatate in #40646Aimv2ModelTest::test_eager_matches_sdpa_inference_04_fp16_pad_right_sdpa_kernels as flaky by @ydshieh in #40683T5GemmaModelTest::test_eager_matches_sdpa_inference being flaky by @ydshieh in #40702hf_hub_download by @ydshieh in #40710self in post-process methods by @framonmar7 in #40711or for grounding dino mask by @lmarshall12 in #40625Gemma Embedding] Fix SWA by @vasqu in #40700VitMatteImageProcessingTest::test_fast_is_faster_than_slow by @ydshieh in #40713request_id to headers by @McPatate in #40722and/or_mask_function by @Cyrilvallez in #40753--continuous_batching by @McPatate in #40618continue_final_message in apply_chat_template to prevent substring matching issues by @abdokaseb in #40732public.cloud.experiment_url api error by @Zeyi-Lin in #40763PromptLookupCandidateGenerator won't generate forbidden tokens by @gante in #40726test_past_key_values_format and delete overwrites by @gante in #40701generate by @gante in #40375Jetmoe] Fix RoPE by @vasqu in #40819self.loss_function by @qubvel in #40764test_modeling_common.py by @gante in #40854past_key_values by @gante in #40803rsqrt by @thalahors in #40848VaultGemma] Update expectations in integration tests by @vasqu in #40855imageprocessor.md to Korean by @HyunZ118 in #39557Gemma3nAudioFeatureExtractionTest::test_dither by @ydshieh in #40902get_mask_sizes by @Cyrilvallez in #40907Glm4vIntegrationTest by @ydshieh in #40905runner_map by @ydshieh in #40880test_fast_is_faster_than_slow by @ydshieh in #40909Gemma3ForConditionalGeneration compatible with assisted generation by @gante in #40791image_sizes arg and deprecate vision_feature_layer by @yaswanth19 in #40832import torch.utils.checkpoint by @gante in #40934Glm4vMoeIntegrationTest by @ydshieh in #40930Glm4vModelTest::test_eager_matches_fa2_generate by @ydshieh in #40947test_speculative_generation by @ydshieh in #40949The following contributors have made significant changes to the library over the last release:
imageprocessor.md to Korean (#39557)Processor load with multi-processing
This patch most notably fixes an issue with the new dtype argument (replacing torch_dtype) in pipelines!
This patch most notably fixes an issue with the new dtype argument (replacing torch_dtype) in pipelines!
The following commits are breaking changes in workflows that were either buggy or not working as expected.
DINOv3 is a family of versatile vision foundation models that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models.
You can find all the original DINOv3 checkpoints under the DINOv3 collection.
<img width="814" height="658" alt="image" src="https://github.com/user-attachments/assets/740a5c3d-a5a1-45d9-9e4c-d9117837205d" />
he X-Codec model was proposed in Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model by Zhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin, Xu Tan, Zheqi Dai, Qiuqiang Kong, Jianyi Chen, Jiahao Pan, Qifeng Liu, Yike Guo, Wei Xue
The X-Codec model is a neural audio codec that integrates semantic information from self-supervised models (e.g., HuBERT) alongside traditional acoustic information. This enables :
<img width="1958" height="949" alt="image" src="https://github.com/user-attachments/assets/e36552d0-6465-4921-8208-3f7d3c9087f1" />
The Ovis2 is an updated version of the Ovis model developed by the AIDC-AI team at Alibaba International Digital Commerce Group.
Ovis2 is the latest advancement in multi-modal large language models (MLLMs), succeeding Ovis1.6. It retains the architectural design of the Ovis series, which focuses on aligning visual and textual embeddings, and introduces major improvements in data curation and training methods.
<img src="https://cdn-uploads.huggingface.co/production/uploads/637aebed7ce76c3b834cea37/XB-vgzDL6FshrSNGyZvzc.png" width="600">
MetaCLIP 2 is a replication of the original CLIP model trained on 300+ languages. It achieves state-of-the-art (SOTA) results on multilingual benchmarks (e.g., XM3600, CVQA, Babel‑ImageNet), surpassing previous SOTA such as mSigLIP and SigLIP‑2. The authors show that English and non-English worlds can mutually benefit and elevate each other.
<img width="805" height="408" alt="image" src="https://github.com/user-attachments/assets/72eaa441-9362-4a6a-a834-f505d6727a2a" />
Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks. Florence-2 can interpret simple text prompts to perform tasks like captioning, object detection, and segmentation. It leverages the FLD-5B dataset, containing 5.4 billion annotations across 126 million images, to master multi-task learning. The model's sequence-to-sequence architecture enables it to excel in both zero-shot and fine-tuned settings, proving to be a competitive vision foundation model.
<img width="864" height="565" alt="image" src="https://github.com/user-attachments/assets/d09dfe3a-6dda-45a3-8dd3-0254d8503b4e" />
SAM2 (Segment Anything Model 2) was proposed in Segment Anything in Images and Videos by Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, Christoph Feichtenhofer.
The model can be used to predict segmentation masks of any object of interest given an input image or video, and input points or bounding boxes.
<img width="960" height="540" alt="image" src="https://github.com/user-attachments/assets/0ab42e5c-6951-4cbc-9d5d-ff8bf0c2dbf1" />
The Kosmos-2.5 model was proposed in KOSMOS-2.5: A Multimodal Literate Model by Microsoft.
The abstract from the paper is the following:
We present Kosmos-2.5, a multimodal literate model for machine reading of text-intensive images. Pre-trained on large-scale text-intensive images, Kosmos-2.5 excels in two distinct yet cooperative transcription tasks: (1) generating spatially-aware text blocks, where each block of text is assigned its spatial coordinates within the image, and (2) producing structured text output that captures styles and structures into the markdown format. This unified multimodal literate capability is achieved through a shared Transformer architecture, task-specific prompts, and flexible text representations. We evaluate Kosmos-2.5 on end-to-end document-level text recognition and image-to-markdown text generation. Furthermore, the model can be readily adapted for any text-intensive image understanding task with different prompts through supervised fine-tuning, making it a general-purpose tool for real-world applications involving text-rich images. This work also paves the way for the future scaling of multimodal large language models.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/kosmos2_5_ocr.png" alt="drawing" width="600"/>
<img width="674" height="402" alt="image" src="https://github.com/user-attachments/assets/230f83f0-870c-4b31-b8b7-738116761457" />
More information at release 🤗
<img width="858" height="537" alt="image" src="https://github.com/user-attachments/assets/29ccc3c2-9b85-4d89-935a-1e1c28d173fd" />
More information at release 🤗
More information at release 🤗
Beyond a large refactor of the caching system in Transformers, making it much more practical and general, models using sliding window attention/chunk attention do not waste memory anymore when caching past states. It was allowed most notable by:
See the following improvements on memory usage for Mistral (using only sliding layers) and GPT-OSS (1 out of 2 layers is sliding) respectively: <img width="569" height="431" alt="image" src="https://github.com/user-attachments/assets/7f1688f4-b077-4840-a62c-bfa6131fe806" /> <img width="574" height="431" alt="image" src="https://github.com/user-attachments/assets/bb4a284f-961e-413d-b7e1-783bb5d8fb39" />
Beyond memory usage, it will also improve generation/forward speed by a large margin for large contexts, as only necessary states are passed to the attention computation, which is very sensitive to the sequence length.
Since the GPT-OSS release which introduced the MXPF4 quantization type, several improvements have been made to the support, which should now stabilize.
swiglu_limit not passed in for MXFP4 by @danielhanchen in #40197Mxfp4] Add a way to save with a quantization method by @ArthurZucker in #40176Now that we deprecated tensorflow and jax, we felt that torch_dtype was not only misaligned with torch, but was redundant and hard to remember. For this reason, we switched to a much more standard dtype argument!
torch_dtype will still be a valid usage for as long as needed to ensure a smooth transition, but new code should use dtype, and we encourage you to update older code as well!
The following commits are breaking changes in workflows that were either buggy or not working as expected.
On models where the hub checkpoint specifies cache_implementation="hybrid" (static sliding window hybrid cache), UNSETS this value. This will make the model use the dynamic sliding window layers by default.
This default meant that there were widespread super slow 1st generate calls on models with hybrid caches, which should nol onger be the case.
cache_implementation="hybrid" hub defaults by @gante in #40135Cache the computation of sine positional embeddings for MaskFormer; results in a 6% performance improvement.
Adds explicit cache initialization to prepare for the deprecation of the from_legacy_cache utility.
fullgraph=FalseHaving fullgraph set to True during compilation ended up being very restrictive, especially with the arrival of widely-used MoEs.
The DoLa decoding strategy has been moved to the following remote-code repository a few versions ago: https://huggingface.co/transformers-community/dola
The Contrastive Search decoding strategy has been moved to the following remote-code repository a few versions ago: https://huggingface.co/transformers-community/contrastive-search
Both have now been removed from the library as a result.
Flash attention has used sliding window sizes which were off by one. This affected generations that had initially bigger contexts than the sliding window size.
Flash Attention] Fix sliding window size by @vasqu in #40163Torch 2.1 support has been unreliable for some time, so we've now made it official and bumped our minimum version to 2.2.
GptOss fixes for green CI by @gante in #39929utils/check_bad_commit.py failing due to rate limit (requesting api.github.com) by @ydshieh in #39918torch.device('cpu').index being None by @manueldeprada in #39933torchcodec is updated by @ydshieh in #39951triton_kernels dep with kernels instead by @SunMarc in #39926fix_and_overwrite mode of utils/check_docstring.py by @manueldeprada in #39369find_file_type by @yonigozlan in #39897past_key_value to past_key_valueS everywhere by @Cyrilvallez in #39956notification_service.py about time_spent by @ydshieh in #40037notification_service.py about time_spent" by @ydshieh in #40044torchcodec==0.5.0 and use torch 2.8 on daily CI by @ydshieh in #40072time_spent in notification_service.py. by @ydshieh in #40081GPT Big Code] Fix attention scaling by @vasqu in #40041ForConditionalGeneration by @qgallouedec in #39973is_fast to ImageProcessor by @MilkClouds in #39603logger.warning with logger.warning_once in GradientCheckpointingLayer by @qgallouedec in #40091Flash Attention] Fix flash attention integration by @vasqu in #40002custom_generate collections by @gante in #39894tiny_agents.md to Korean by @AhnJoonSung in #39913content inputs for LLMs by @gante in #39829decoding_method argument in generate by @manueldeprada in #40085generation_config by @gante in #40127main_classes/processors.md to Korean by @TaskerJang in #39519jamba.md to Korean by @skwh54 in #39890main_classes/optimizer_schedules.md to Korean by @luckyvickyricky in #39713gpt2.md to Korean by @taemincode in #39808optimizers.md to Korean by @chelsseeey in #40011pipelines.md to Korean by @xhaktm00 in #39577gemma3.md to Korean by @seopp in #39865torch_compile_test and torch_export_test by @ydshieh in #39950self.tokenizer by self.processing_class by @qgallouedec in #40119too long with no output by @ydshieh in #40201model_input_names for PixtralImageProcessor by @rohitrango in #40226chat_template (jinja2) as an extra dependency by @tboerstad in #40128CI] Fix repo consistency by @vasqu in #40249k_proj weight and bias slicing in D-FINE by @notkisk in #40257id=usage to <hfoptions> tag in LayoutLM model card by @Jin-HoMLee in #40273torch.compile tests with fullgraph=True by @ydshieh in #40164FA] Fix dtype in varlen with position ids by @vasqu in #40295fix] Pass adamw optimizer parameters to StableAdamW by @emapco in #40184find_executable_batch_size to match new 0.9 ratio by @MilkClouds in #40206Flash Attention] Fix sliding window size by @vasqu in #40163_tp_plan attribute by @rishub-tamirisa in #39944natten by @ydshieh in #40287GPT OSS] Refactor the tests as it was not properly checking the outputs by @ArthurZucker in #40288get_placeholder_mask in Ovis2 by @thisisiron in #40280/en/model_doc by @gante in #40311/en/model_doc by @gante in #40344test_spm_converter_bytefallback_warning by @ydshieh in #40284FA] Fix some model tests by @vasqu in #40350label_names as an argument to TrainingArguments by @huzaifa-jawad367 in #40353skip_special_tokens in the main text generation pipelines by @gante in #40356dtype instead of torch_dtype everywhere! by @Cyrilvallez in #39782tokenizer_kwargs argument to the text generation pipeline by @Joshua-Chin in #40364transformers TF classes/methods by @gante in #40429models.md to Korean by @Judy-Choi in #39518main by @ydshieh in #40451qwen2_moe tests by @ydshieh in #40494merge to main by @ydshieh in #40503The following contributors have made significant changes to the library over the last release:
get_placeholder_mask in Ovis2 (#40280)There was a mick mack on our side when cherry-picking the commit #40197 which led to a wrong commit in the patch! Sorry everyone 😭
There was a mick mack on our side when cherry-picking the commit #40197 which led to a wrong commit in the patch! Sorry everyone 😭
This patch is just the official fix for #40197!
Focused on stabilizing FlashAttention-2 on Ascend NPU, improving FSDP behavior for generic-task models, fixing MXFP4 integration for GPT-OSS
Focused on stabilizing FlashAttention-2 on Ascend NPU, improving FSDP behavior for generic-task models, fixing MXFP4 integration for GPT-OSS
😢 Well sorry everyone, sometimes shit can happen... 4.55.1 was broken because of 🥁 git merge conflict. I cherry-picked https://github.com/huggingface/
FA2 generations!😢 Well sorry everyone, sometimes shit can happen...
4.55.1 was broken because of 🥁 git merge conflict.
I cherry-picked https://github.com/huggingface/transformers/pull/40002 without having https://github.com/huggingface/transformers/pull/40029 , thus from ..modeling_flash_attention_utils import prepare_fa_kwargs_from_position_ids is missing, and since this is a slow test, nothing caught it.
Will work to remediate and write the post-mortem when yanking the release.
Mostly focused around stabalizing the Mxfp4 for GPTOSS model!
Mostly focused around stabalizing the Mxfp4 for GPTOSS model!
Remove all expired deprecation cycles by @Cyrilvallez in #39725
<img width="2320" height="1160" alt="image" src="https://github.com/user-attachments/assets/4a1cd2f6-dde9-445e-83d9-73f6551e2da2" />
For more detailed information about this model, we recommend reading the following blogpost: https://huggingface.co/blog/welcome-openai-gpt-oss
GPT OSS is a hugely anticipated open-weights release by OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases. It comprises two models: a big one with 117B parameters (gpt-oss-120b), and a smaller one with 21B parameters (gpt-oss-20b). Both are mixture-of-experts (MoEs) and use a 4-bit quantization scheme (MXFP4), enabling fast inference (thanks to fewer active parameters, see details below) while keeping resource usage low. The large model fits on a single H100 GPU, while the small one runs within 16GB of memory and is perfect for consumer hardware and on-device applications.
The following snippet shows simple inference with the 20B model. It runs on 16 GB GPUs when using mxfp4, or ~48 GB in bfloat16.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openai/gpt-oss-20b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto",
)
messages = [
{"role": "user", "content": "How many rs are in the word 'strawberry'?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
generated = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(generated[0][inputs["input_ids"].shape[-1]:]))
The models use attention sinks, a technique the vLLM team made compatible with Flash Attention 3. We have packaged and integrated their optimized kernel in kernels-community/vllm-flash-attn3. At the time of writing, this super-fast kernel has been tested on Hopper cards with PyTorch 2.7 and 2.8. We expect increased coverage in the coming days. If you run the models on Hopper cards (for example, H100 or H200), you need to pip install –upgrade kernels and add the following line to your snippet:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openai/gpt-oss-20b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto",
+ # Flash Attention with Sinks
+ attn_implementation="kernels-community/vllm-flash-attn3",
)
messages = [
{"role": "user", "content": "How many rs are in the word 'strawberry'?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
generated = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(generated[0][inputs["input_ids"].shape[-1]:]))
Even though the 120B model fits on a single H100 GPU (using mxfp4), you can also run it easily on multiple GPUs using accelerate or torchrun. Transformers provides a default parallelization plan, and you can leverage optimized attention kernels as well. The following snippet can be run with torchrun --nproc_per_node=4 generate.py on a system with 4 GPUs:
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers.distributed import DistributedConfig
import torch
model_path = "openai/gpt-oss-120b"
tokenizer = AutoTokenizer.from_pretrained(model_path, padding_side="left")
device_map = {
"tp_plan": "auto", # Enable Tensor Parallelism
}
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype="auto",
attn_implementation="kernels-community/vllm-flash-attn3",
**device_map,
)
messages = [
{"role": "user", "content": "Explain how expert parallelism works in large language models."}
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1000)
# Decode and print
response = tokenizer.decode(outputs[0])
print("Model response:", response.split("<|channel|>final<|message|>")[-1].strip())
If you have a Hopper GPU or better, we recommend you use mxfp4 for the reasons explained above. If you can additionally use Flash Attention 3, then by all means do enable it!
[!TIP] If your GPU is not compatible with mxfp4, then we recommend you use MegaBlocks MoE kernels for a nice speed bump. To do so, you just need to adjust your inference code like this:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openai/gpt-oss-20b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto",
+ # Optimize MoE layers with downloadable MegaBlocksMoeMLP
+ use_kernels=True,
)
messages = [
{"role": "user", "content": "How many rs are in the word 'strawberry'?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
generated = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(generated[0][inputs["input_ids"].shape[-1]:]))
[!TIP] MegaBlocks optimized MoE kernels require the model to run on bfloat16, so memory consumption will be higher than running on mxfp4. We recommend you use mxfp4 if you can, otherwise opt in to MegaBlocks via use_kernels=True.
You can use transformers serve to experiment locally with the models, without any other dependencies. You can launch the server with just: transformers serve
To which you can send requests using the Responses API.
# responses API
curl -X POST http://localhost:8000/v1/responses \
-H "Content-Type: application/json" \
-d '{"input": [{"role": "system", "content": "hello"}], "temperature": 1.0, "stream": true, "model": "openai/gpt-oss-120b"}'
You can also send requests using the standard Completions API:
# completions API
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "system", "content": "hello"}], "temperature": 1.0, "max_tokens": 1000, "stream": true, "model": "openai/gpt-oss-120b"}'
<img width="1920" height="960" alt="image" src="https://github.com/user-attachments/assets/5502cc65-2fc9-49ac-8e15-262aa573b68d" />
Command A Vision is a state-of-the-art multimodal model designed to seamlessly integrate visual and textual information for a wide range of applications. By combining advanced computer vision techniques with natural language processing capabilities, Command A Vision enables users to analyze, understand, and generate insights from both visual and textual data.
The model excels at tasks including image captioning, visual question answering, document understanding, and chart understanding. This makes it a versatile tool for AI practitioners. Its ability to process complex visual and textual inputs makes it useful in settings where text-only representations are imprecise or unavailable, like real-world image understanding and graphics-heavy document processing.
Command A Vision is built upon a robust architecture that leverages the latest advancements in VLMs. It's highly performant and efficient, even when dealing with large-scale datasets. The model's flexibility makes it suitable for a wide range of use cases, from content moderation and image search to medical imaging analysis and robotics.
<img width="838" height="266" alt="image" src="https://github.com/user-attachments/assets/4d1e153c-0586-4650-8e18-c9d08145ce49" />
MM Grounding DINO model was proposed in An Open and Comprehensive Pipeline for Unified Object Grounding and Detection by Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiangtai Li, Xinjiang Wang, Yining Li, Haian Huang>.
MM Grounding DINO improves upon the Grounding DINO by improving the contrastive class head and removing the parameter sharing in the decoder, improving zero-shot detection performance on both COCO (50.6(+2.2) AP) and LVIS (31.9(+11.8) val AP and 41.4(+12.6) minival AP).
You can find all the original MM Grounding DINO checkpoints under the MM Grounding DINO collection. This model also supports LLMDet inference. You can find LLMDet checkpoints under the LLMDet collection.
FastSpeech2Conformer by @bvantuan in #39689classmethod by @zucchini-nlp in #38812CI] Add Eric to comment slow ci by @vasqu in #39601QAPipelineTests::test_large_model_course after #39193 by @ydshieh in #39666Glm4MoeModelTest::test_torch_compile_for_training by @ydshieh in #39670Qwen2AudioForConditionalGeneration.forward() and test_flash_attn_kernels_inference_equivalence by @ebezzam in #39503models/__init__.py for typo checking by @hebangwen in #39745GemmaIntegrationTest::test_model_2b_bf16_dola again by @ydshieh in #39731--gpus all in workflow files by @ydshieh in #39752libcst to extras["testing"] in setup.py by @ydshieh in #39761main_classes/peft.md by @luckyvickyricky in #39515tvp.md to Korean by @Kim-Ju-won in #39578tokenizer.md to Korean by @seopp in #39532pipeline_gradio.md to Korean by @AhnJoonSung in #39520perf_train_gpu_one.md to Korean by @D15M4S in #39552how_to_hack_models.md to Korean by @skwh54 in #39536run_name when none by @qgallouedec in #39695model_results.json by @ydshieh in #39783attn_implementation] remove recursive, allows custom kernels with wrappers by @ArthurZucker in #39823plot_keypoint_matching, make visualize_keypoint_matching as a standard by @sbucaille in #39830TrackioCallback to work when pynvml is not installed by @qgallouedec in #39851is_wandb_available function to verify WandB installation by @qgallouedec in #39875sub_configs by @qubvel in #39855Tokenizer with PreTrainedTokenizerFast in ContinuousBatchProcessor by @qgallouedec in #39858torch.backends.cudnn.allow_tf32 = False for CI by @ydshieh in #39885AutoModelForCausalLM and AutoModelForImageTextToText by @qubvel in #39881ModernBertForMultipleChoice by @netique in #39232Exaone4] Fixes the attn implementation! by @ArthurZucker in #39906The following contributors have made significant changes to the library over the last release:
We had quite a lot of bugs that got through! Release was a bit rushed, sorry everyone! 🤗 Mostly cache fixes, as we now have layered cache, and fixed t
We had quite a lot of bugs that got through! Release was a bit rushed, sorry everyone! 🤗 Mostly cache fixes, as we now have layered cache, and fixed to distributed.
Fix tests due to breaking change in accelerate by @SunMarc in #39451
In order to become the source of truth, we recognize that we need to address two common and long-heard critiques about transformers:
transformers is bloatedtransformers is slowOur team has focused on improving both aspects, and we are now ready to announce this.
The modeling files for the standard Llama models are down to 500 LOC and should be much more readable, keeping just the core of the modeling and hiding the "powerful transformers features."
<img width="1583" height="974" alt="image" src="https://github.com/user-attachments/assets/f1075598-d63e-4184-b3af-c0d4b31cdde5" />
The MoEs are getting some kernel magic, enabling the use of the efficient megablocks kernels, setting a good precedent to allow the community to leverage any of the most powerful kernels developed for quantization as well!
It should also be much more convenient to use with any attention implementation you want. This opens the door to some optimizations such as leveraging flash-attention on Metal (MPS Torch backend).
<img width="2050" height="752" alt="image" src="https://github.com/user-attachments/assets/23ebfb20-7626-46a5-b264-76ffb8b8c811" />
This is but the tip of the iceberg: with the work on kernels that we're heavily pushing forward, expect speed-ups on several backends in the coming months!!
This release also includes the first steps to enabling efficient distributed training natively in transformers. Loading a 100B model takes ~3 seconds on our cluster — we hope this will be the norm for everyone! We are working on distributed checkpointing as well, and want to make sure our API can be easily used for any type of parallelism.
We want the community to benefit from all of the advances, and as always, include all hardware and platforms! We believe the kernels library will give the tools to optimize everything, making a big difference for the industry!
The Ernie 4.5 model was released in the Ernie 4.5 Model Family release by baidu. This family of models contains multiple different architectures and model sizes. This model in specific targets the base text model without mixture of experts (moe) with 0.3B parameters in total. It uses the standard Llama at its core.
Other models from the family can be found at Ernie 4.5 MoE.
<div class="flex justify-center"> <img src="https://ernie.baidu.com/blog/posts/ernie4.5/overview.png"/> </div>
Ernie 4.5] Add ernie text models by @vasqu in #39228Voxtral is an upgrade of Ministral 3B and Mistral Small 3B, extending its language capabilities with audio input support. It is designed to handle tasks such as speech transcription, translation, and audio understanding.
You can read more in Mistral's realease blog post.
The model is available in two checkpoints:
Voxtral builds on Ministral-3B by adding audio processing capabilities:
LFM2 represents a new generation of Liquid Foundation Models developed by Liquid AI, specifically designed for edge AI and on-device deployment.
The models are available in three sizes (350M, 700M, and 1.2B parameters) and are engineered to run efficiently on CPU, GPU, and NPU hardware, making them particularly well-suited for applications requiring low latency, offline operation, and privacy.
The DeepSeek-V2 model was proposed in DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model by DeepSeek-AI Team.
The model uses Multi-head Latent Attention (MLA) and DeepSeekMoE architectures for efficient inference and cost-effective training. It employs an auxiliary-loss-free strategy for load balancing and multi-token prediction training objective. The model can be used for various language tasks after being pre-trained on 14.8 trillion tokens and going through Supervised Fine-Tuning and Reinforcement Learning stages.
ModernBERT Decoder is the same architecture as ModernBERT but trained from scratch with a causal language modeling (CLM) objective. This allows for using the same architecture for comparing encoders and decoders. This is the decoder architecture implementation of ModernBERT, designed for autoregressive text generation tasks.
Like the encoder version, ModernBERT Decoder incorporates modern architectural improvements such as rotary positional embeddings to support sequences of up to 8192 tokens, unpadding to avoid wasting compute on padding tokens, GeGLU layers, and alternating attention patterns. However, it uses causal (unidirectional) attention to enable autoregressive generation.
The Encoder-only Mask Transformer (EoMT) model was introduced in the CVPR 2025 Highlight Paper Your ViT is Secretly an Image Segmentation Model by Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans, Narges Norouzi, Giuseppe Averta, Bastian Leibe, Gijs Dubbelman, and Daan de Geus. EoMT reveals Vision Transformers can perform image segmentation efficiently without task-specific components.
<div style="text-align: center;"> <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/eomt_architecture.png" alt="drawing" width="500"/> </div>
Doge is a series of small language models based on the Doge architecture, aiming to combine the advantages of state-space and self-attention algorithms, calculate dynamic masks from cached value states using the zero-order hold method, and solve the problem of existing mainstream language models getting lost in context. It uses the wsd_scheduler scheduler to pre-train on the smollm-corpus, and can continue training on new datasets or add sparse activation feedforward networks from stable stage checkpoints.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/refs%2Fpr%2F426/transformers/model_doc/doge_architecture.png" alt="drawing" width="600"
The AIMv2 model was proposed in Multimodal Autoregressive Pre-training of Large Vision Encoders by Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor Guilherme Turrisi da Costa, Louis Béthune, Zhe Gan, Alexander T Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby.
The abstract from the paper is the following:
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
The PerceptionLM model was proposed in PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding by Jang Hyun Cho et al. It's a fully open, reproducible model for transparent research in image and video understanding. PLM consists of a vision encoder with a small scale (<8B parameters) LLM decoder.
The EfficientLoFTR model was proposed in Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed by Yifan Wang, Xingyi He, Sida Peng, Dongli Tan and Xiaowei Zhou.
This model consists of matching two images together by finding pixel correspondences. It can be used to estimate the pose between them. This model is useful for tasks such as image matching, homography estimation, etc.
<img width="2772" height="3276" alt="image" src="https://github.com/user-attachments/assets/8fb76e17-9bff-4edc-9ac8-205f4b58a898" />
Evolla is an advanced 80-billion-parameter protein-language generative model designed to decode the molecular language of proteins. It integrates information from protein sequences, structures, and user queries to generate precise and contextually nuanced insights into protein function. Trained on an unprecedented AI-generated dataset of 546 million protein question-answer pairs and 150 billion word tokens, Evolla significantly advances research in proteomics and functional genomics, providing expert-level insights and shedding light on the molecular logic encoded in proteins.
<img width="824" height="1017" alt="image" src="https://github.com/user-attachments/assets/298074a0-c509-4e0b-9adb-c72bac206b18" />
Deepseek-VL was introduced by the DeepSeek AI team. It is a vision-language model (VLM) designed to process both text and images for generating contextually relevant responses. The model leverages LLaMA as its text encoder, while SigLip is used for encoding images.
<img width="692" height="434" alt="image" src="https://github.com/user-attachments/assets/164243ba-1fd8-48b5-b565-3396d640ce0e" />
The xLSTM model was proposed in xLSTM: Extended Long Short-Term Memory by Maximilian Beck*, Korbinian Pöppel*, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter. xLSTM updates the original LSTM architecture to be competitive with Transformer models by introducing exponential gating, matrix memory expansion, and parallelizable training and ingestion.
The 7B model variant was trained by the xLSTM team Maximilian Beck, Korbinian Pöppel, Phillip Lippe, Richard Kurle, Patrick Blies, Sebastian Böck and Sepp Hochreiter at NXAI.
<img width="3750" height="954" alt="image" src="https://github.com/user-attachments/assets/591040d2-d90a-4770-a4a7-0d91e21263fb" />
EXAONE 4.0 model is the language model, which integrates a Non-reasoning mode and Reasoning mode to achieve both the excellent usability of EXAONE 3.5 and the advanced reasoning abilities of EXAONE Deep. To pave the way for the agentic AI era, EXAONE 4.0 incorporates essential features such as agentic tool use, and its multilingual capabilities are extended to support Spanish in addition to English and Korean.
The EXAONE 4.0 model series consists of two sizes: a mid-size 32B model optimized for high performance, and a small-size 1.2B model designed for on-device applications.
We've added Expert Parallel support for Llama4, next release will include it for all model! You can just set a distributed_config with enable_expert_parallel=True. This is enabling efficient training of sparse Mixture-of-Experts (MoE) models across multiple devices. This allows each expert in the MoE layer to run in parallel (instead of previous TP which requires more communication), significantly improving scalability and memory efficiency.
FP-Quant is a quantization method optimized for Blackwell-generation Nvidia GPUs, supporting efficient post-training quantization (PTQ) and quantization-aware training (QAT) of LLMs using MXFP4 and NVFP4 formats.
Currently, only PTQ with MXFP4 is available. You can quantize models on-the-fly using transformers:
from transformers import AutoModelForCausalLM, FPQuantConfig
model = AutoModelForCausalLM.from_pretrained(
"qwen/Qwen3-8B",
quantization_config=FPQuantConfig(),
device_map="cuda",
torch_dtype=torch.bfloat16,
)
FP-Quant requires a Blackwell GPU and runs via the QuTLASS library. No Blackwell GPU? Use FPQuantConfig(pseudoquant=True) to emulate quantization (no QuTLASS needed).
The following results show the inference speedup of QuTLASS MXFP4 over PyTorch BF16 in Transformers. MXFP4 gives consistent speedups across all batch sizes, reaching up to 4× faster at larger scales.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/qwen3-8b-end-to-end-prefill-speedup-mxfp4-vs-bf16-on-rtx5090.svg" alt="drawing" width="600">
The kernels project aims to become the single trusted source for high-performance kernels in the Transformers ecosystem. We're working toward centralizing all kernels on the Hub, so updates, bug fixes, and improvements can happen in one place—no more scattered repos and no compilation headaches!
You can already try it out today by setting use_kernels=True in from_pretrained. Any contributor can build their kernel, register it and use it right away—no extra setup, more on this here
Even better: want to use Flash Attention 3? No need to deal with tricky compilation and missing symbols issues! Just drop in:
model.set_attn_implementation("kernels-community/flash-attn3")
This automatically fetches the right build for your setup (e.g. CUDA and PyTorch versions).
We’re also teaming up with amazing kernel devs from Unsloth, Liger, vLLM, and more to bring their work directly to the Hub—making it easier than ever to access amazing performance with a single line of code.
https://github.com/user-attachments/assets/9928f62b-543c-4b8a-b81b-4a6e262c229e
Over the past few months, we have been putting more and more functionality in the transformers chat utility, which offers a CLI-based app to chat with chat models. We've chosen to push this further by splitting the backend of transformers chat in a new, separate utility called transformers serve.
This is ideal for experimentation purposes, or to run models locally for personal and private use. It does not aim to compete with dedicated inference engines such as vLLM or SGLang.
Models of diverse modalities supported by transformers may be served with the transformers serve CLI. It spawns a local server that offers compatibility with the OpenAI SDK, which is the de-facto standard for LLM conversations and other related tasks. This way, you can use the server from many third party applications, or test it using the transformers chat CLI (docs).
The server supports the following REST APIs:
/v1/chat/completions/v1/responses/v1/audio/transcriptions/v1/modelsRelevant commits:
transformers chat and transformers serve by @LysandreJik in #38443transformers serve by @LysandreJik in #39149generation_config by @gante in #39230transformers serve by @LysandreJik in #39155/v1/audio/transcriptions) by @gante in #39434Significant refactors have been underway in transformers, aiming to reduce the complexity of the code. A metric we follow to see how the refactors impact our code is to follow the number of lines in a given model; we try to reduce it as much as possible, while keeping everything related to the forward pass and model definition in that file.
See the evolution here:
<img width="1200" height="600" alt="image" src="https://github.com/user-attachments/assets/c232bc8d-7d7c-4192-baa8-a60efe5eb2ff" />
Some notable refactors:
KV caches are now defined per layer, enabling new hybrid caches that mix different attention types. CacheProcessors also encapsulate cache quantization and offloading, making them easy to customize.
output_attentions or output_hidden_statesSuch attributes require very specific handling within the forward call, while they're not important to understand how the model works. We remove that code but keep the functionality by providing a better utility to handle it.
We refactor the way to explicitly set the attention implementation so that it has a method dedicated to it.
average_tokens_across_devices by default in TrainingArguments by @Krish0909 in #39395Flex Attn] Fix torch 2.5.1 incompatibilities by @vasqu in #37406test_compare_unprocessed_logit_scores by @ydshieh in #39053t5gemma tests by @ydshieh in #39052layoutlmv3 tests by @ydshieh in #39050Gemma3nProcessorTest by @ydshieh in #39068mistral3 tests by @ydshieh in #38989dots1 tests by @ydshieh in #39088test_is_split_into_words in test_pipelines_token_classification.py by @st81 in #39079test_sdpa_can_dispatch_on_flash by @ydshieh in #39092@lru_cache() to @lru_cache to match styles from #38883. by @rasmi in #39093run-slow by @ydshieh in #39100llama tests by @ydshieh in #39161Dia] Change ckpt path in docs by @vasqu in #39181from_pretrained by @qubvel in #39184fastspeech2_conformer tests by @ydshieh in #39229is not None -> isinstance(..., dict) by @qubvel in #39145segmentation_maps support to MobileNetV2ImageProcessor by @simonreise in #37312tests/generation/test_utils.py by @ydshieh in #39254test_eager_matches sdpa generate and update an integration test for blip-like models by @ydshieh in #39248smollm3 by @gante in #39271PretrainedConfig.__init__ method to make it more explicit by @qubvel in #39158test_generate_compile_model_forward by @ydshieh in #39276datasets 4.0 by @lhoestq in #39156aria tests by @ydshieh in #39277test_torchscript_* for now until the majority of the community ask for it by @ydshieh in #39307stevhliu to the list in self-comment-ci.yml by @ydshieh in #39315src/ for doctest (for now) by @ydshieh in #39316max_length_q and max_length_k types to flash_attn_varlen_func by @HollowMan6 in #37206phi3 tests by @ydshieh in #39312position_ids in masking_utils by @Cyrilvallez in #39310test_sdpa_can_dispatch_on_flash by @ydshieh in #39259timm (for perception_lm) by @ydshieh in #39380/v1/models output payload by @alvarobartt in #39414set_tracer_provider and set_meter_provider calls by @McPatate in #39422JetMoeForCausalLM by @Phoenix-Shen in #37830ContinuousBatchProcessor by @qgallouedec in #39372CI] Fix partially red CI by @vasqu in #39448GemmaIntegrationTest::test_model_2b_bf16_dola by @ydshieh in #39362datasets pin by @gante in #39500args_doc.py to auto_docstring.py by @yonigozlan in #39439_supports_flash_attn_2 in examples and tests by @zucchini-nlp in #39471TypeError instead of ValueError for invalid types by @Sai-Suraj-27 in #38660MambaCache to modeling_mamba.py by @manueldeprada in #38086perf_infer_gpu_multi.md to Korean by @luckyvickyricky in #39441CI] Fix post merge ernie 4.5 by @vasqu in #39561docs/source/ko/_toctree.yml by @jungnerd in #39516supports_static_cache to can_compile_fullgraph by @zucchini-nlp in #39505device_mesh have multiple dim by @S1ro1 in #38949test_export_static_cache by @gante in #39662Ernie 4.5] Post merge adaptations by @vasqu in #39664kyutai tests by @ydshieh in #39416typing.Literal as type of tool parameters or return value by @grf53 in #39633The following contributors have made significant changes to the library over the last release:
segmentation_maps support to MobileNetV2ImageProcessor (#37312)docs/source/ko/_toctree.yml (#39516)A small patch for open telemetry fixes! Sorry for the delay!
A small patch for open telemetry fixes! Sorry for the delay!
** refactor: remove set_tracer_provider and set_meter_provider calls (https://github.com/huggingface/transformers/pull/39422) from @McPatate
[sliding window] revert and deprecate
This patch contains the following bug fixes:
smollm3 (#39271)position_ids in masking_utils (#39310)This patch contains several bug fixes. The following commits are included:
This patch contains several bug fixes. The following commits are included:
Several minimal breaking changes aiming to bring clearer defaults while greatly simplifying the library have been merged.
Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for pre-trained and instruction-tuned variants. These models were trained with data in over 140 spoken languages.
Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the Gemma 3n page.
from transformers import pipeline
import torch
pipe = pipeline(
"image-text-to-text",
torch_dtype=torch.bfloat16,
model="google/gemma-3n-e4b",
device="cuda",
)
output = pipe(
"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg",
text="<image_soft_token> in this image, there is"
)
print(output)
Dia is an opensource text-to-speech (TTS) model (1.6B parameters) developed by Nari Labs. It can generate highly realistic dialogue from transcript including nonverbal communications such as laughter and coughing. Furthermore, emotion and tone control is also possible via audio conditioning (voice cloning).
Model Architecture: Dia is an encoder-decoder transformer based on the original transformer architecture. However, some more modern features such as rotational positional embeddings (RoPE) are also included. For its text portion (encoder), a byte tokenizer is utilized while for the audio portion (decoder), a pretrained codec model DAC is used - DAC encodes speech into discrete codebook tokens and decodes them back into audio.
<img src="https://huggingface.co/datasets/eustlb/documentation-images/resolve/main/kyutai_stt.png"/>
Kyutai STT is a speech-to-text model architecture based on the Mimi codec, which encodes audio into discrete tokens in a streaming fashion, and a Moshi-like autoregressive decoder. Kyutai’s lab has released two model checkpoints:
Read more about the model in the documentation
<div class="flex justify-center"> <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/vjepa.gif" alt="drawing" width="600"/> </div>
V-JEPA 2 is a self-supervised approach to training video encoders developed by FAIR, Meta. Using internet-scale video data, V-JEPA 2 attains state-of-the-art performance on motion understanding and human action anticipation tasks. V-JEPA 2-AC is a latent action-conditioned world model post-trained from V-JEPA 2 (using a small amount of robot trajectory interaction data) that solves robot manipulation tasks without environment-specific data collection or task-specific training or calibration.
Read more about the model in the documentation.
Arcee is a decoder-only transformer model based on the Llama architecture with a key modification: it uses ReLU² (ReLU-squared) activation in the MLP blocks instead of SiLU, following recent research showing improved training efficiency with squared activations. This architecture is designed for efficient training and inference while maintaining the proven stability of the Llama design.
The Arcee model is architecturally similar to Llama but uses x * relu(x) in MLP layers for improved gradient flow and is optimized for efficiency in both training and inference scenarios.
Read more about the model in the documentation.
ColQwen2 is a variant of the ColPali model designed to retrieve documents by analyzing their visual features. Unlike traditional systems that rely heavily on text extraction and OCR, ColQwen2 treats each page as an image. It uses the Qwen2-VL backbone to capture not only text, but also the layout, tables, charts, and other visual elements to create detailed multi-vector embeddings that can be used for retrieval by computing pairwise late interaction similarity scores. This offers a more comprehensive understanding of documents and enables more efficient and accurate retrieval.
Read more about the model in the documentation.
MiniMax is a powerful language model with 456 billion total parameters, of which 45.9 billion are activated per token. To better unlock the long context capabilities of the model, MiniMax adopts a hybrid architecture that combines Lightning Attention, Softmax Attention and Mixture-of-Experts (MoE). Leveraging advanced parallel strategies and innovative compute-communication overlap methods—such as Linear Attention Sequence Parallelism Plus (LASP+), varlen ring attention, Expert Tensor Parallel (ETP), etc., MiniMax's training context length is extended to 1 million tokens, and it can handle a context of up to 4 million tokens during the inference. On various academic benchmarks, MiniMax also demonstrates the performance of a top-tier model.
The architecture of MiniMax is briefly described as follows:
For more details refer to the release blog post.
Read more about the model in the documentation.
T5Gemma (aka encoder-decoder Gemma) was proposed in a research paper by Google. It is a family of encoder-decoder large langauge models, developed by adapting pretrained decoder-only models into encoder-decoder. T5Gemma includes pretrained and instruction-tuned variants. The architecture is based on transformer encoder-decoder design following T5, with improvements from Gemma 2: GQA, RoPE, GeGLU activation, RMSNorm, and interleaved local/global attention.
T5Gemma has two groups of model sizes: 1) Gemma 2 sizes (2B-2B, 9B-2B, and 9B-9B), which are based on the offical Gemma 2 models (2B and 9B); and 2) T5 sizes (Small, Base, Large, and XL), where are pretrained under the Gemma 2 framework following T5 configuration. In addition, we also provide a model at ML size (medium large, ~2B in total), which is in-between T5 Large and T5 XL.
The pretrained varaints are trained with two objectives: prefix language modeling with knowledge distillation (PrefixLM) and UL2, separately. We release both variants for each model size. The instruction-turned varaints was post-trained with supervised fine-tuning and reinforcement learning.
Read more about the model in the documentation.
The GLM-4.1V model architecture is added to transformers; no models have yet been released with that architecture. Stay tuned for the GLM team upcoming releases!
Read more about the model in the documentation.
The FalconH1 model was developed by the TII Pretraining team. A comprehensive research paper covering the architecture, pretraining dynamics, experimental results, and conclusions is forthcoming. You can read more about this series in this website.
Read more about the model in the documentation.
The LightGlue model was proposed in LightGlue: Local Feature Matching at Light Speed by Philipp Lindenberger, Paul-Edouard Sarlin and Marc Pollefeys.
Similar to SuperGlue, this model consists of matching two sets of local features extracted from two images, its goal is to be faster than SuperGlue. Paired with the SuperPoint model, it can be used to match two images and estimate the pose between them. This model is useful for tasks such as image matching, homography estimation, etc.
The abstract from the paper is the following:
We introduce LightGlue, a deep neural network that learns to match local features across images. We revisit multiple design decisions of SuperGlue, the state of the art in sparse matching, and derive simple but effective improvements. Cumulatively, they make LightGlue more efficient - in terms of both memory and computation, more accurate, and much easier to train. One key property is that LightGlue is adaptive to the difficulty of the problem: the inference is much faster on image pairs that are intuitively easy to match, for example because of a larger visual overlap or limited appearance change. This opens up exciting prospects for deploying deep matchers in latency-sensitive applications like 3D reconstruction. The code and trained models are publicly available at this https URL
Read more about the model in the documentation.
The abstract from the report is the following:
Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that activates 14B parameters out of a total of 142B parameters, delivering performance on par with state-of-the-art models while reducing training and inference costs. Leveraging our meticulously crafted and efficient data processing pipeline, dots.llm1 achieves performance comparable to Qwen2.5-72B after pretraining on high-quality corpus and post-training to fully unlock its capabilities. Notably, no synthetic data is used during pretraining. To foster further research, we open-source intermediate training checkpoints spanning the entire training process, providing valuable insights into the learning dynamics of large language models.
Read more about the model in the documentation.
SmolLM3 is a fully open, compact language model designed for efficient deployment while maintaining strong performance. It uses a Transformer decoder architecture with Grouped Query Attention (GQA) to reduce the kv cache, and no RoPE, enabling improved performance on long-context tasks. It is trained using a multi-stage training approach on high-quality public datasets across web, code, and math domains. The model is multilingual and supports very large context lengths. The instruct variant is optimized for reasoning and tool use.
Read more about the model in the documentation.
In previous versions, installing the kernels library would automatically activate the custom kernels added to transformers, because the @use_kernel_forward_from_the_hub decorator directly swapped out the model’s forward method. This implicit behavior caused several issues for users — including problems with torch.compile, non-determinism, and inconsistent outputs.
To address this, we've introduced a new opt-in mechanism called kernelize. You can now enable kernel usage explicitly by passing use_kernels=True to from_pretrained. The use_kernel_forward_from_the_hub decorator now simply stores the kernel name that the user wants to use — and kernelize handles the rest under the hood.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-1B-Instruct",
torch_dtype=torch.bfloat16,
device_map="cuda",
use_kernels=True
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-1B-Instruct")
input = "Hello"
input_ids = tokenizer(input, return_tensors="pt").to(model.device).input_ids
output = model.generate(input_ids, max_new_tokens=100)
print(tokenizer.decode(output[0], skip_special_tokens=True))
More kernels will be added over time — this will be a collaborative, community-driven effort to make transformers lighter and faster 🤗
Support for Flash Attention 3 is added across the most popular models.
Several efforts refactoring the repository are happening in parallel. The direction is to greatly simplify the library, removing unnecessary codepaths. Whilst the efforts are spread across the library, they're particularly visible in each individual models; where non-modeling-specific code will be simplified and eventually removed.
We take the assumption that model-agnostic utilities shouldn't be in the modeling code. Things like the output of attentions, hidden states, router logits, are important for end-users but don't need to be explicitely displayed in the modeling code.
Several minimal breaking changes aiming to bring clearer defaults while greatly simplifying the library have been merged.
dtype for pipelines to auto by @Vaibhavs10 in #38882output_attentions=True and the attn implementation is wrong by @ArthurZucker in #38288Attention] Refactor Attention Interface for Bart-based Models by @vasqu in #38108Attention] Attention refactor for Whisper-based models by @vasqu in #38235compile] re-enable for Qwen-VL models by @zucchini-nlp in #38127forced_decoder_ids by @gante in #38232liger-kernel to docker file by @ydshieh in #38292transformers env output by @yao-matrix in #38274forced_decoder_ids deletion by @gante in #38316beam_indices by @gante in #38259custom_generate and trust_remote_code by @gante in #38304vasqu to self-comment-ci.yml by @ydshieh in #38324FlexAttention] Reenable flex for encoder-decoder and make the test more robust by @vasqu in #38321kernels for AMD docker images by @ydshieh in #38354OPT] Fix attention scaling by @vasqu in #38290get_default_device for torch<2.3 by @Cyrilvallez in #38376utils/notification_service.py by @ydshieh in #38379initialize_weights by @Cyrilvallez in #38382tokenizer -> tokenize by @foldl in #38357generation_config.json as base parameterization by @gante in #38330test_offloaded_cache_implementation) by @gante in #37896pixel_values with inputs_embeds by @dxoigmn in #38334CsmForConditionalGenerationIntegrationTest by @ydshieh in #38424huggingface/transformers by @ydshieh in #38413from_pretrained by @pstjohn in #38155from_args_and_dict ProcessorMixin by @yonigozlan in #38296microsoft/python-type-stubs (post dropping support for Python 3.8) by @Avasam in #38335BatchFeature and BatchEncoding by @lgeiger in #38459Gemma3IntegrationTest by @ydshieh in #38471SinkCache to a custom_generate repo by @gante in #38399Gemma2IntegrationTest by @ydshieh in #38492av by @ydshieh in #38548python3 by @S1ro1 in #38555utils/notification_service.py by @ydshieh in #38556chameleon tests by @ydshieh in #38565utils/notification_service.py for AMD vs Nvidia by @ydshieh in #38563deepseekv3 by @ydshieh in #38562FlexAttn] Fix models with unique characteristics by @vasqu in #38433repository field to benchmarks table by @McPatate in #38582mlm_probability to be set to None when mlm=False in DataCollatorForLanguageModeling by @KameniAlexNea in #38522)isort from dependencies by @Sai-Suraj-27 in #38616return_dict=False giving errors in a few VLM models by @ydshieh in #38519MiniMax (docs and integration tests checkpoint) by @geetu040 in #38575test_initialization by @ydshieh in #38607ColQwen2ModelIntegrationTest by @ydshieh in #38583test_initialization for SwiftFormer by @ydshieh in #38636AriaForConditionalGenerationModelTest on CircleCI by @ydshieh in #38615InternVL integration test by @ydshieh in #38612aya_vision test by @ydshieh in #38674is_bitsandbytes_available() by @ved1beta in #38528llava tests by @ydshieh in #38722None instead of try/except by @zucchini-nlp in #38561average_tokens_across_devices=True and world size = 1 by @qgallouedec in #38785qwen_2_5 omni by @ydshieh in #38658llava_onevision tests by @ydshieh in #38791mllama by @ydshieh in #38704low_cpu_mem_usage by @Cyrilvallez in #38792llava_next tests by @ydshieh in #38813wandb.run.url instead of wandb.run.get_url() (deprecated) by @qgallouedec in #38817align_to_words=True in QuestionAnsweringPipeline can lead to duplicate answers by @yushi2006 in #38761qwen2_5_vl tests by @ydshieh in #38845auxiliary_in_channels default behavior in UperNet by @simonreise in #37540qwen3 tests by @ydshieh in #38862phi4_multimodal tests by @ydshieh in #38816qwen3_moe tests by @ydshieh in #38865raise from e in hub.py utility by @Wauplin in #37241fsmt tests by @ydshieh in #38904FalconMambaIntegrationTests by @ydshieh in #38566ALL_LAYERNORM_LAYERS by @Cyrilvallez in #38922test_initialization by @ydshieh in #38932mistral and mistral3 tests by @ydshieh in #38978is_split_into_words in the TokenClassificationPipeline. by @yushi2006 in #38818rag by @ydshieh in #38585Attention] Small fix on output attentions by @vasqu in #38948require_tf) by @gante in #38944The following contributors have made significant changes to the library over the last release:
liger-kernel to docker file (#38292)vasqu to self-comment-ci.yml (#38324)kernels for AMD docker images (#38354)utils/notification_service.py (#38379)CsmForConditionalGenerationIntegrationTest (#38424)huggingface/transformers (#38413)Gemma3IntegrationTest (#38471)Gemma2IntegrationTest (#38492)av (#38548)utils/notification_service.py (#38556)chameleon tests (#38565)utils/notification_service.py for AMD vs Nvidia (#38563)deepseekv3 (#38562)return_dict=False giving errors in a few VLM models (#38519)test_initialization (#38607)ColQwen2ModelIntegrationTest (#38583)test_initialization for SwiftFormer (#38636)AriaForConditionalGenerationModelTest on CircleCI (#38615)InternVL integration test (#38612)aya_vision test (#38674)llava tests (#38722)qwen_2_5 omni (#38658)llava_onevision tests (#38791)mllama (#38704)llava_next tests (#38813)qwen2_5_vl tests (#38845)qwen3 tests (#38862)phi4_multimodal tests (#38816)qwen3_moe tests (#38865)fsmt tests (#38904)FalconMambaIntegrationTests (#38566)test_initialization (#38932)mistral and mistral3 tests (#38978)rag (#38585)output_attentions=True and the attn implementation is wrong (#38288)transformers env output (#38274)Attention] Refactor Attention Interface for Bart-based Models (#38108)FlexAttention] Reenable flex for encoder-decoder and make the test more robust (#38321)OPT] Fix attention scaling (#38290)Attention] Attention refactor for Whisper-based models (#38235)FlexAttn] Fix models with unique characteristics (#38433)Attention] Small fix on output attentions (#38948)microsoft/python-type-stubs (post dropping support for Python 3.8) (#38335)MiniMax (docs and integration tests checkpoint) (#38575)The following commits are included in that patch release:
The following commits are included in that patch release:
We had to protect the imports again, a series of bad events. Here are the two prs for the patch:
We had to protect the imports again, a series of bad events. Here are the two prs for the patch:
We had to revert #37877 because of a missing flag that was overriding the device map. We re-introduced the changes because they allow native 3D parall
We had to revert #37877 because of a missing flag that was overriding the device map. We re-introduced the changes because they allow native 3D parallel training in Transformers. Sorry everyone for the troubles! 🤗
[chat template] fix security vulnerability by @zucchini-nlp in #37523
<img width="1090" alt="image" src="https://github.com/user-attachments/assets/77f0fe5b-59cd-4fb6-b222-bcc2b35d6406" />
The Qwen2.5-Omni model is a unified multiple modalities model proposed in Qwen2.5-Omni Technical Report from Qwen team, Alibaba Group.
The abstract from the technical report is the following:
We present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. To enable the streaming of multimodal information inputs, both audio and visual encoders utilize a block-wise processing approach. This strategy effectively decouples the handling of long sequences of multimodal data, assigning the perceptual responsibilities to the multimodal encoder and entrusting the modeling of extended sequences to a large language model.
Such a division of labor enhances the fusion of different modalities via the shared attention mechanism. To synchronize the timestamps of video inputs with audio, we organized the audio and video sequentially in an interleaved manner and propose a novel position embedding approach, named TMRoPE (Time-aligned Multimodal RoPE). To concurrently generate text and speech while avoiding interference between the two modalities, we propose Thinker-Talker architecture.
In this framework, Thinker functions as a large language model tasked with text generation, while Talker is a dual-track autoregressive model that directly utilizes the hidden representations from the Thinker to produce audio tokens as output. Both the Thinker and Talker models are designed to be trained and inferred in an end-to-end manner. For decoding audio tokens in a streaming manner, we introduce a sliding-window DiT that restricts the receptive field, aiming to reduce the initial package delay. Qwen2.5-Omni outperforms the similarly sized Qwen2-VL and Qwen2-Audio in both image and audio capabilities. Furthermore, Qwen2.5-Omni achieves state-of-the-art performance on multimodal benchmarks like Omni-Bench.
Notably, Qwen2.5-Omni is the first open-source model to achieve a level of performance in end-to-end speech instruction following that is comparable to its capabilities with text inputs, as evidenced by benchmarks such as MMLU and GSM8K. As for speech generation, Qwen2.5-Omni’s streaming Talker outperform most existing streaming and non-streaming alternatives in robustness and naturalness.
SAM-HQ (High-Quality Segment Anything Model) was proposed in Segment Anything in High Quality by Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu.
The model is an enhancement to the original SAM model that produces significantly higher quality segmentation masks while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability.
SAM-HQ introduces several key improvements over the original SAM model:
The abstract from the paper is the following:
The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced dataset of 44k masks, which takes only 4 hours on 8 GPUs.
Tips:
The GraniteMoeHybrid model builds on top of GraniteMoeSharedModel and Bamba. Its decoding layers consist of state space layers or MoE attention layers with shared experts. By default, the attention layers do not use positional encoding.
<img width="1051" alt="image" src="https://github.com/user-attachments/assets/3274da06-ff44-4bb4-bebf-8bc5f9b72aac" />
The D-FINE model was proposed in D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement by Yansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang, Xiaoyan Sun, Feng Wu
The abstract from the paper is the following:
We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and pretrained models: this https URL.
The Conversational Speech Model (CSM) is the first open-source contextual text-to-speech model released by Sesame. It is designed to generate natural-sounding speech with or without conversational context. This context typically consists of multi-turn dialogue between speakers, represented as sequences of text and corresponding spoken audio.
Model Architecture: CSM is composed of two LLaMA-style auto-regressive transformer decoders: a backbone decoder that predicts the first codebook token and a depth decoder that generates the remaining tokens. It uses the pretrained codec model Mimi, introduced by Kyutai, to encode speech into discrete codebook tokens and decode them back into audio.
The original csm-1b checkpoint is available under the Sesame organization on Hugging Face.
<div class="flex justify-center"> <img src="https://huggingface.co/datasets/eustlb/documentation-images/resolve/main/csm_architecture.png"/> </div>
<img width="697" alt="image" src="https://github.com/user-attachments/assets/022e426e-71bb-40fd-8458-ad3b48432759" />
Trained on a corpus of 4 trillion tokens, this model demonstrates that native 1-bit LLMs can achieve performance comparable to leading open-weight, full-precision models of similar size, while offering substantial advantages in computational efficiency (memory, energy, latency).
Llama Guard 4 is a new multimodal model designed to detect inappropriate content in images and text, whether used as input or generated as output by the model. It’s a dense 12B model pruned from Llama 4 Scout model, and it can run on a single GPU (24 GBs of VRAM). It can evaluate both text-only and image+text inputs, making it suitable for filtering both inputs and outputs of large language models. This enables flexible moderation pipelines where prompts are analyzed before reaching the model, and generated responses are reviewed afterwards for safety. It can also understand multiple languages.
<img width="625" alt="image" src="https://github.com/user-attachments/assets/6d7fd266-f391-4914-bdf9-ebdddb4d3f5f" />
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model proposed in A decoder-only foundation model for time-series forecasting by Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. It is a decoder only model that uses non-overlapping patches of time-series data as input and outputs some output patch length prediction in an autoregressive fashion.
The abstract from the paper is the following:
Motivated by recent advances in large language models for Natural Language Processing (NLP), we design a time-series foundation model for forecasting whose out-of-the-box zero-shot performance on a variety of public datasets comes close to the accuracy of state-of-the-art supervised forecasting models for each individual dataset. Our model is based on pretraining a patched-decoder style attention model on a large time-series corpus, and can work well across different forecasting history lengths, prediction lengths and temporal granularities.
<img width="618" alt="image" src="https://github.com/user-attachments/assets/2c2c1a6c-9c96-4c6c-a3d3-a24b0fc908af" />
The MLCD models were released by the DeepGlint-AI team in unicom, which focuses on building foundational visual models for large multimodal language models using large-scale datasets such as LAION400M and COYO700M, and employs sample-to-cluster contrastive learning to optimize performance. MLCD models are primarily used for multimodal visual large language models, such as LLaVA.
<img width="770" alt="image" src="https://github.com/user-attachments/assets/8cd33a13-7d9c-430b-a822-893d83f09b87" />
The Janus Model was originally proposed in Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation by DeepSeek AI team and later refined in Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. Janus is a vision-language model that can generate both image and text output, it can also take both images and text as input.
[!NOTE] The model doesn't generate both images and text in an interleaved format. The user has to pass a parameter indicating whether to generate text or image.
The abstract from the original paper is the following:
In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.
The abstract from the aforementioned Janus-Pro paper, released afterwards, is the following:
In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strate (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.
The InternVL3 family of Visual Language Models was introduced in InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.
The abstract from the paper is the following:
We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model (LLM) into a multimodal large language model (MLLM) that supports visual inputs, InternVL3 jointly acquires multimodal and linguistic capabilities from both diverse multimodal data and pure-text corpora during a single pre-training stage. This unified training paradigm effectively addresses the complexities and alignment challenges commonly encountered in conventional post-hoc training pipelines for MLLMs. To further improve performance and scalability, InternVL3 incorporates variable visual position encoding (V2PE) to support extended multimodal contexts, employs advanced post-training techniques such as supervised fine-tuning (SFT) and mixed preference optimization (MPO), and adopts test-time scaling strategies alongside an optimized training infrastructure. Extensive empirical evaluations demonstrate that InternVL3 delivers superior performance across a wide range of multi-modal tasks. In particular, InternVL3-78B achieves a score of 72.2 on the MMMU benchmark, setting a new state-of-the-art among open-source MLLMs. Its capabilities remain highly competitive with leading proprietary models, including ChatGPT-4o, Claude 3.5 Sonnet, and Gemini 2.5 Pro, while also maintaining strong pure-language proficiency. In pursuit of open-science principles, we will publicly release both the training data and model weights to foster further research and development in next-generation MLLMs.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/internvl_architecture.png" alt="drawing" width="600"/>
<small> Overview of InternVL3 models architecture, which is the same as InternVL2.5. Taken from the <a href="https://huggingface.co/OpenGVLab/InternVL3-1B">original checkpoint.</a> </small>
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/internvl_overview_performance.png" alt="drawing" width="600"/>
<small> Comparison of InternVL3 performance on OpenCompass against other SOTA VLLMs. Taken from the <a href="https://huggingface.co/OpenGVLab/InternVL3-1B">original checkpoint.</a> </small>
We integrate some kernels in the transformers library via the kernels package: https://github.com/huggingface/kernels
We start with some kernels in the Llama model, and we iterate to identify the best performance optimizations
In the previous release, we've added TP support in order to run distributed inference. However, this is not supported for all quantization methods. We are progressively adding support to it. Right now, only compressed-tensors, fp8 and fp8-fbgemm support it.
From the AutoRound contributors:
AutoRound is an advanced quantization algorithm that delivers strong accuracy, even at 2-bit precision. It leverages sign gradient descent to fine-tune both rounding values and min-max clipping thresholds in just 200 steps ... More details here: https://github.com/intel/auto-round
We have added two new sections to better understand and get started with quantization:
We've added GGUF support to gemma3 family models.
Most Vision Models and VLMs in Transformers can now benefit from fast image processors. By utilizing torch/torchvision functional transforms, these processors offer a substantial speedup when processing images compared to PiL/numpy functions, and support processing on both CPU and CUDA.
The new @auto_docstring decorator makes it easier to add proper documentation when contributing a model without bloating the modeling code:
@auto_docstring: AutoDocstringgenerateWe now support custom generate methods to be loaded from model.generate. The custom generate methods can be stored on the Hub, enabling quick distribution of experiments regarding new caches, decoding methods, heuristics, ...
from transformers import AutoModelForCausalLM, AutoTokenizer
# `generate` with `custom_generate` -> `generate` uses custom code
# note: calling the custom method prints "✨ using a custom generation method ✨"
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", device_map="auto")
inputs = tokenizer(["The quick brown"], return_tensors="pt").to(model.device)
gen_out = model.generate(**inputs, custom_generate="transformers-community/custom_generate_example", trust_remote_code=True)
print(tokenizer.batch_decode(gen_out, skip_special_tokens=True))
You can find the docs here, and all custom generation methods by searching for the custom_generate tag.
The transformers-cli command is updated to be simpler and cleaner, specifically for its chat variant.
The following is now possible and recommended:
transformers chat Qwen/Qwen2.5-3B-Instruct
Additionally, almost any generate flag can now be passed as a positional argument, present and future, as opposed to being limited to a set of hardcoded flags, for example:
transformers chat Qwen/Qwen2.5-0.5B-Instruct do_sample=False max_new_tokens=10
chat] generate parameterization powered by GenerationConfig and UX-related changes by @gante in #38047The agents folder is finally removed from transformers in favour of using smolagents.
We are moving away from torch 2.0 as it has been released more than two years ago.
init empty weights without accelerate by @Cyrilvallez in #37337_init_weights by @Cyrilvallez in #37341GenerationMixin inheritance by default in PreTrainedModel by @gante in #37173_pytree._register_pytree_node and torch.cpu.amp.autocast by @bzhong-solink in #37372kernels to 0.4.3 by @ArthurZucker in #37419rms_norm_eps for the L2Norm for Llama4 by @ArthurZucker in #37418tests/models/ by @ydshieh in #37415fsspec dependency which isn't directly used by transformers by @cyyever in #37318_init_weights() issues - make it work for composite models by @Cyrilvallez in #37070num_logits_to_keep by @Cyrilvallez in #37149from_pretrained by @Cyrilvallez in #37216attn_temperature_tuning by @gmlwns2000 in #37501test_offloaded_cache_implementation on XPU by @yao-matrix in #37514as_tensor) by @ydshieh in #37551test_can_load_with_global_device_set using a subprocess by @ydshieh in #37553xxx_token_id for multimodal tokens by @zucchini-nlp in #37573test_past_key_values_format by @gante in #37614/scripts 🧹 🧹 by @gante in #37676en docs in push CI by @gante in #37677siglip.md to Korean by @devxaitist in #37145Qwen2_5OmniConfig.get_text_config by @shahruk10 in #37690/model_cards 🧹 🧹 by @gante in #37685sacrebleu (and document why) by @gante in #37700qwen2_5_omni] fix flaky tests by @gante in #37721test_nemotron_8b_generation_sdpa by @faaany in #37665embeds_to_talker device in Qwen2.5-Omni by @BakerBunker in #37739AriaForConditionalGenerationIntegrationTest on T4 by @ydshieh in #37746MllamaForConditionalGenerationIntegrationTest by @ydshieh in #37750HybridCache init when device is passed by @gante in #37718GPT2Model StaticCache support by @poedator in #35761torch version by @gante in #37760roberta.md to Korean by @garongkim in #37069keypoint_detection.md to Korean by @rlaalsrl0922 in #36649hub.py by @srai9 in #37796test_generate_continue_from_past_key_values by @gante in #37724electra.md to Korean by @Kim-Ju-won in #36763torch.compile test by @gante in #37894AOPerModuleConfig and include_embedding by @jerryzh168 in #37802load_state_dict by @woct0rdho in #37902gpu_selection.md to Korean by @nsbg in #36757vocab_size access for multimodal models by @kurzdev in #37937max_memory argument when factoring in unused reserved memory by @gante in #37982pad image transform for batched inputs by @sebasv in #37544Optional typing by @qubvel in #38018test_push_to_hub_with_saves_each_epoch for now by @ydshieh in #38022torchscript.md by @Madghostek in #38004test_speculative_decoding_non_distil device-agnostic by @faaany in #38010AutoDocstring] Based on inspect parsing of the signature by @ArthurZucker and @yonigozlan in #33771ready for review by @ydshieh in #37885Trigger CircleCI via GitHub Actions when ready for review` by @ydshieh in #38038Trigger CircleCI via GitHub Actions when "ready for review" by @ydshieh in #37885)kernels from docker images by @ydshieh in #38083require_read_token by @ydshieh in #38093librispeech_asr dataset by @faaany in #38073lr_scheduler_kwargs options to create LR Scheduler when LayerWiseDummyOptimizer is used by @BlackNoodle in #34559past_key_values type hint in model output types by @ChengLyu in #37953check_bad commit.py gives wrong results by @ydshieh in #38107manueldeprada to run_slow whitelist by @manueldeprada in #38126include_embedding flag by @jerryzh168 in #37935SinusoidsPositionEmbedding precision by @BakerBunker in #38151Trigger CircleCI by ready for review by @ydshieh in #38171convert to draft workflow by @ydshieh in #38177fetch_tests CircleCI job by @ydshieh in #38176test_sdpa_equivalence (redundant) by @gante in #37911The following contributors have made significant changes to the library over the last release:
fsspec dependency which isn't directly used by transformers (#37318)test_offloaded_cache_implementation on XPU (#37514)embeds_to_talker device in Qwen2.5-Omni (#37739)SinusoidsPositionEmbedding precision (#38151)siglip.md to Korean (#37145)Nothing published for this version
A mix of bugs were fixed in this patch; very exceptionally, we diverge from semantic versioning to merge GLM-4 in this patch release.
A mix of bugs were fixed in this patch; very exceptionally, we diverge from semantic versioning to merge GLM-4 in this patch release.
This is another round of bug fixes, but they are a lot more minor and outputs were not really affected!
This is another round of bug fixes, but they are a lot more minor and outputs were not really affected!
Since the release of Llama 4, we have fixed a few issues that we are now releasing in patch v4.51.1
Since the release of Llama 4, we have fixed a few issues that we are now releasing in patch v4.51.1
Thanks all for your patience
Deprecate #36741 and map Causal to Conditional by @zucchini-nlp in #36917
Llama 4, developed by Meta, introduces a new auto-regressive Mixture-of-Experts (MoE) architecture.This generation includes two models:
Both models leverage early fusion for native multimodality, enabling them to process text and image inputs. Maverick and Scout are both trained on up to 40 trillion tokens on data encompassing 200 languages (with specific fine-tuning support for 12 languages including Arabic, Spanish, German, and Hindi).
For deployment, Llama 4 Scout is designed for accessibility, fitting on a single server-grade GPU via on-the-fly 4-bit or 8-bit quantization, while Maverick is available in BF16 and FP8 formats. These models are released under the custom Llama 4 Community License Agreement, available on the model repositories
Getting started with Llama 4 using transformers is straightforward. Make sure you have transformers v4.51.0 or later installed:
pip install -U transformers[hf_xet]
Here's a quick example using the instruction-tuned Maverick model responding about two images, using tensor parallel for maximum speed. You need to run this script on an instance with 8 GPUs, using a command like:
torchrun –nproc-per-instance=8 script.py
from transformers import AutoProcessor, Llama4ForConditionalGeneration
import torch
model_id = "meta-llama/Llama-4-Maverick-17B-128E-Instruct"
processor = AutoProcessor.from_pretrained(model_id)
model = Llama4ForConditionalGeneration.from_pretrained(
model_id,
attn_implementation="flex_attention",
device_map="auto",
torch_dtype=torch.bfloat16,
)
url1 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/0052a70beed5bf71b92610a43a52df6d286cd5f3/diffusers/rabbit.jpg"
url2 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/datasets/cat_style_layout.png"
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": url1},
{"type": "image", "url": url2},
{"type": "text", "text": "Can you describe how these two images are similar, and how they differ?"},
]
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
)
response = processor.batch_decode(outputs[:, inputs["input_ids"].shape[-1]:])[0]
print(response)
print(outputs[0])
Make sure to check the model cards on the repos (Llama 4 Maverick (~400B) and Llama 4 Scout (~109B)) for detailed usage instructions, including multimodal examples, specific prompt formats (like system prompts), quantization details, and advanced configuration options!
<img width="898" alt="image" src="https://github.com/user-attachments/assets/847a18d8-0d6a-4767-b45c-3cc9d6ff392e" />
Phi-4-multimodal-instruct is a lightweight open multimodal foundation model that leverages the language, vision, and speech research and datasets used for Phi-3.5 and 4.0 models. The model processes text, image, and audio inputs, generating text outputs, and comes with 128K token context length. The model underwent an enhancement process, incorporating both supervised fine-tuning, direct preference optimization and RLHF (Reinforcement Learning from Human Feedback) to support precise instruction adherence and safety measures. The languages that each modal supports are the following:
DeepSeek-v3 is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
The model is detailed in the following paper.
The DeepSeek-V3 model was proposed in DeepSeek-V3 Technical Report by DeepSeek-AI Team.
The abstract from the paper is the following:
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.
The Qwen3 architecture has been contributed to transformers and is available in v4.51.0. At time of release, the models themselves have not yet been released - stay tuned for a release from the Qwen team!
Model docs are getting a significant overhaul by providing much needed, ready-to-use examples one can copy-paste in their modules/consoles. We will adapt these examples to each model, with the goal of providing relevant examples on a per-model basis.
A very large PR was provided by @nikosanto13 that helped add modular files to all speech models in the library; seeing the difference between each of them is now much simpler, as well as maintenance and eventual refactors.
original_max_position_embeddings to YARN rope_scaling optional keys by @JustinTong0323 in #36877trainer_pt_utils.py docstrings for consistency by @ethanknights in #36912DataCollatorForWholeWordMask by @capemox in #36903uv for installing packages by @Sai-Suraj-27 in #36957networkx==3.2.1 manually in some CircleCI jobs after #36957 by @ydshieh in #37000to_py_obj for python-native numeric lists and scalars by @n0gu-furiosa in #36885qwen2_vl.md to Korean by @MinJu-Ha in #36750AwqConfigTest by @faaany in #37032test_assisted_decoding_in_different_gpu test on XPU by @yao-matrix in #37120_VALID_DICT_FIELDS to class attribute for shared dict parsing in subclasses by @Tavish9 in #36736ModernBERT] Never save 'reference_compile' config; should be set based on end user by @tomaarsen in #36305307 in RequestCounter by @ydshieh in #36953TASK_MAPPING by @saattrupdan in #37107min_new_tokens to prevent flaky length checks by @gante in #37175num_items_in_batch if necessary by @regisss in #36967utils/check_bad_commit.py by @ydshieh in #37272return_tensors in audio chat templates by @zucchini-nlp in #346010.11.2 by @ydshieh in #36962lru_cache for tokenization tests by @ydshieh in #36818return_dict logic to remove complicated if/else paths by @qubvel in #36794The following contributors have made significant changes to the library over the last release:
Thanks to the vllm team we have a few more bugs that slipped in!
Thanks to the vllm team we have a few more bugs that slipped in!
[generate] beam search -- fix output cropping (#37080) by @gante
[blip-2] Fix dtype mismatch when keep in fp32 (#37068) by @zucchini-nlp
Fix PixtralProcessor patch_size when spatial_merge_size is used (#37019)
I completely forgot to put these in the previous patch sorry! Should put the transformers backend in a good spot!
I completely forgot to put these in the previous patch sorry! Should put the transformers backend in a good spot!
[Utils] torch version checks optionally accept dev versions (#36847) by @gante
Fix processor kwargs qwen2 vl (#36890) by @yonigozlan
Fix Pan and Scan on batched images Gemma3 (#36864) by @yonigozlan
Deprecate #36741 and map Causal to Conditional (#36917) by @zucchini-nlp
There were some very minor bugs with the new hub kernels, and with remote code that we had to fix
Deprecate #36741 and map Causal to Conditional (#36917) by @zucchini-nlp
Fix pytorch deform attn path (#36923) by @qubvel
[chameleon] fix num image token check (#36918) by @zucchini-nlp
Fix torch version guard at import (#36907) by @zucchini-nlp
Security fix for benchmark.yml by @ydshieh in #36402
Starting with version v4.49.0, we have been doing model-based releases, additionally to our traditional, software-based monthly releases. These model-based releases provide a tag from which models may be installed.
Contrarily to our software-releases; these are not pushed to pypi and are kept on our GitHub. Each release has a tag attributed to it, such as:
v4.49.0-Gemma-3v4.49.0-AyaVision⚠️ As bugs are identified and fixed on each model, the release tags are updated so that installing from that tag always gives the best experience possible with that model.
Each new model release will always be based on the current state of the main branch at the time of its creation. This ensures that new models start with the latest features and fixes available.
For example, if two models—Gemma-3 and AyaVision—are released from main, and then a fix for gemma3 is merged, it will look something like this:
o---- v4.49.0-Gemma-3 (includes AyaVision, plus main fixes)
/ \
---o--o--o--o--o-- (fix for gemma3) --o--o--o main
\
o---- v4.49.0-AyaVision
We strive to merge model specific fixes on their respective branches as fast as possible!
Gemma 3 is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
The Gemma 3 model was proposed by Google. It is a vision-language model composed by a SigLIP vision encoder and a Gemma 2 language decoder linked by a multimodal linear projection.
It cuts an image into a fixed number of tokens same way as Siglip if the image does not exceed certain aspect ratio. For images that exceed the given aspect ratio, it crops the image into multiple smaller pacthes and concatenates them with the base image embedding.
One particularity is that the model uses bidirectional attention on all the image tokens. Also, the model interleaves sliding window local attention with full causal attention in the language backbone, where each sixth layer is a full causal attention layer.
ShieldGemma 2 is built on Gemma 3, is a 4 billion (4B) parameter model that checks the safety of both synthetic and natural images against key categories to help you build robust datasets and models. With this addition to the Gemma family of models, researchers and developers can now easily minimize the risk of harmful content in their models across key areas of harm as defined below:
We recommend using ShieldGemma 2 as an input filter to vision language models, or as an output filter of image generation systems. To train a robust image safety model, we curated training datasets of natural and synthetic images and instruction-tuned Gemma 3 to demonstrate strong performance.
AyaVision is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
The Aya Vision 8B and 32B models is a state-of-the-art multilingual multimodal models developed by Cohere For AI. They build on the Aya Expanse recipe to handle both visual and textual information without compromising on the strong multilingual textual performance of the original model.
Aya Vision 8B combines the Siglip2-so400-384-14 vision encoder with the Cohere CommandR-7B language model further post-trained with the Aya Expanse recipe, creating a powerful vision-language model capable of understanding images and generating text across 23 languages. Whereas, Aya Vision 32B uses Aya Expanse 32B as the language model.
Key features of Aya Vision include:
Mistral 3.1 is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
Building upon Mistral Small 3 (2501), Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance. With 24 billion parameters, this model achieves top-tier capabilities in both text and vision tasks.
It is ideal for:
SmolVLM-2 is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
SmolVLM2 is an adaptation of the Idefics3 model with two main differences:
SigLIP-2 is heavily referenced in the following model-based release and we recommend reading these if you want all the information relative to that model.
The SigLIP2 model was proposed in SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features by Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner and Xiaohua Zhai.
The model comes in two variants
transformers)PromptDepthAnything is a high-resolution, accurate metric depth estimation model that leverages prompting, inspired by its success in vision-language (VLMs) and large language models (LLMs). Using iPhone LiDAR as a prompt, the model generates precise depth maps at up to 4K resolution, unlocking the potential of depth foundation models.
We add a new tool to transformers to visualize the attention layout of a given model. It only requires a model ID as input, and will load the relevant tokenizer/model and display what the attention mask looks like. Some examples:
from transformers.utils.attention_visualizer import AttentionMaskVisualizer
visualizer = AttentionMaskVisualizer("meta-llama/Llama-3.2-3B-Instruct")
visualizer("A normal attention mask")
visualizer = AttentionMaskVisualizer("mistralai/Mistral-Small-24B-Instruct-2501")
visualizer("A normal attention mask with a long text to see how it is displayed, and if it is displayed correctly")
visualizer = AttentionMaskVisualizer("google/paligemma2-3b-mix-224")
visualizer("<img> You are an assistant.", suffix = "What is on the image?")
visualizer = AttentionMaskVisualizer("google/gemma-2b")
visualizer("You are an assistant. Make sure you print me") # we should have slidiing on non sliding side by side
visualizer = AttentionMaskVisualizer("google/gemma-3-27b-it")
visualizer("<img>You are an assistant. Make sure you print me") # we should have slidiing on non sliding side by side
We are deprecating transformers.agents in favour of the smolagents library. Read more about smolagents here.
We support adding custom quantization method by using the @register_quantization_config and @register_quantizer decorator:
@register_quantization_config("custom")
class CustomConfig(QuantizationConfigMixin):
pass
@register_quantizer("custom")
class CustomQuantizer(HfQuantizer):
pass
quantized_model = AutoModelForCausalLM.from_pretrained(
"facebook/opt-350m", quantization_config=CustomConfig(), torch_dtype="auto"
)
AMD is developing its in-house quantizer named Quark released under MIT license, which supports a broad range of quantization pre-processing, algorithms, dtypes and target hardware. You can now load a model quantized by quark library:
# pip install amd-quark
model_id = "EmbeddedLLM/Llama-3.1-8B-Instruct-w_fp8_per_channel_sym"
model = AutoModelForCausalLM.from_pretrained(model_id)
model = model.to("cuda")
Torchao is augmented with autoquant support, CPU-quantization, as well as new AOBaseConfig object instances for more advanced configuration.
At loading time, the parallelization is now applied module-by-module, so that no memory overhead is required compared to what the final weight distribution will be!
This release includes two speed upgrades to generate:
do_sample=True;from transformers import pipeline
import torch
prompt = "Alice and Bob"
checkpoint = "google/gemma-2-9b"
assistant_checkpoint = "double7/vicuna-68m"
pipe = pipeline(
"text-generation",
model=checkpoint,
assistant_model=assistant_checkpoint,
do_sample=True
)
pipe_output = pipe(prompt, max_new_tokens=50, do_sample=True)
print(pipe_output[0]["generated_text"])
num_beams. The speedup is more visible on smaller models, where model.forward doesn't dominate the total run time.CandidateGenerator by @keyboardAnt, @jmamou, and @gauravjain14 in #35029A significant redesign of our documentation has wrapped-up. The goal was to greatly simplify the transformers documentation, making it much more easy to navigate. Let us know what you think!
The research examples folder that was hosted in transformers is no more. We have moved it out of transformers and in the following repo: github.com/huggingface/transformers-research-projects/
We have updated our flex attention support so as to have it be on-par with our Flash Attention 2 support.
EsmModelIntegrationTest::test_inference_bitsandbytes by @faaany in #36225LlavaForConditionalGenerationModelTest::test_config after #36077 by @ydshieh in #36230/generation by @gante in #36235test_export_to_onnx by @gante in #36241test_fast_is_faster_than_slow by @ydshieh in #36240Speech2TextFeatureExtractor API. by @KarelVesely84 in #34638pt_tf equivalence tests by @gante in #36253test_from_pretrained_low_cpu_mem_usage_equal less flaky by @gante in #36255GenerationTesterMixin inheritance is correct 🐛 🔫 by @gante in #36180main by @ydshieh in #36375is_causal fail with compile by @Cyrilvallez in #36374benchmark.yml by @ydshieh in #36402CandidateGenerator by @keyboardAnt in #35029contents: write by @ydshieh in #36445torch.distributed-compatible DynamicCache by @gante in #36373src/transformers/image_utils.py by @hmellor in #36435hub_retry by @ydshieh in #36449TRUST_REMOTE_CODE for RealmRetriever for security by @ydshieh in #36511input_ids passed to PrefixConstrainedLogitsProcessor is zero by @HiDolen in #36489DataCollatorForLanguageModeling by @capemox in #36457HybridCache] disable automatic compilation by @gante in #36620make fix-copies by @gante in #36664from_pretrained by @Cyrilvallez in #36033meta device by @gante in #36543gc.collect() if only 1 shard is used by @gante in #36721test_eager_matches_sdpa_inference by @gante in #36650generation_config, overwrite default values with the model's base generation_config by @gante in #36684TrainingArguments.torch_empty_cache_steps post_init check by @pkuderov in #36734test_eager_matches_sdpa_inference by @gante in #36740is_decoder usage in PretrainedConfig documentation by @d-kleine in #36724tj-actions/changed-files by @ydshieh in #36795dist": "loadfile" for pytest in CircleCI jobs by @ydshieh in #36811Trainer.collator.tokenizer in when Trainer.processing_class is None by @innerNULL in #36552GenerationMixin by @gante in #36605DataCollatorForLanguageModeling by @capemox in #36497.item in get_batch_samples by @regisss in #36861deformable_detr kernel from the Hub by @danieldk in #36853The following contributors have made significant changes to the library over the last release:
CandidateGenerator (#35029)deformable_detr kernel from the Hub (#36853)Security fix for self-comment-ci.yml by @ydshieh in #35548
Helium-1 preview is a lightweight language model with 2B parameters, targeting edge and mobile devices. It supports the following languages: English, French, German, Italian, Portuguese, Spanish.
<img width="860" alt="image" src="https://github.com/user-attachments/assets/52e91b74-5572-46a6-93e5-058730411675" />
The Qwen2.5-VL model is an update to Qwen2-VL from Qwen team, Alibaba Group.
The abstract from this update is the following:
Qwen2.5-VL marks a major step forward from Qwen2-VL, built upon the latest Qwen2.5 LLM. We’ve accelerated training and testing through the strategic implementation of window attention within the ViT. The ViT architecture itself has been refined with SwiGLU and RMSNorm, aligning it more closely with the LLM’s structure. A key innovation is the expansion of native dynamic resolution to encompass the temporal dimension, in addition to spatial aspects. Furthermore, we’ve upgraded MRoPE, incorporating absolute time alignment on the time axis to allow the model to effectively capture temporal dynamics, regardless of frame rate, leading to superior video understanding.
The SuperGlue model was proposed in SuperGlue: Learning Feature Matching with Graph Neural Networks by Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz and Andrew Rabinovich.
This model consists of matching two sets of interest points detected in an image. Paired with the SuperPoint model, it can be used to match two images and estimate the pose between them. This model is useful for tasks such as image matching, homography estimation, etc.
<img width="424" alt="image" src="https://github.com/user-attachments/assets/1d81983f-f9ce-4d82-adb7-e76098df543a" />
The Granite Vision model is a variant of LLaVA-NeXT, leveraging a Granite language model alongside a SigLIP visual encoder. It utilizes multiple concatenated vision hidden states as its image features, similar to VipLlava. It also uses a larger set of image grid pinpoints than the original LlaVa-NeXT models to support additional aspect ratios.
Zamba2 is a large language model (LLM) trained by Zyphra, and made available under an Apache 2.0 license.
Zamba2-1.2B, Zamba2-2.7B and Zamba2-7B are hybrid models combining state-space models (Specifically Mamba) and transformer, and were trained using next-token prediction. Zamba2 uses shared transformer layers after every 6 mamba blocks. It uses the Mistral v0.1 tokenizer. We came to this architecture after a series of ablations at small scales. Zamba2-1.2B, Zamba2-2.7B and Zamba2-7B were pre-trained on 2T and 3T tokens, respectively.
GOT-OCR2 works on a wide range of tasks, including plain document OCR, scene text OCR, formatted document OCR, and even OCR for tables, charts, mathematical formulas, geometric shapes, molecular formulas and sheet music. While this implementation of the model will only output plain text, the outputs can be further processed to render the desired format, with packages like pdftex, mathpix, matplotlib, tikz, verovio or pyecharts. The model can also be used for interactive OCR, where the user can specify the region to be recognized by providing the coordinates or the color of the region’s bounding box.
DAB-DETR is an enhanced variant of Conditional DETR. It utilizes dynamically updated anchor boxes to provide both a reference query point (x, y) and a reference anchor size (w, h), improving cross-attention computation. This new approach achieves 45.7% AP when trained for 50 epochs with a single ResNet-50 model as the backbone.
DepthPro is a foundation model for zero-shot metric monocular depth estimation, designed to generate high-resolution depth maps with remarkable sharpness and fine-grained details. It employs a multi-scale Vision Transformer (ViT)-based architecture, where images are downsampled, divided into patches, and processed using a shared Dinov2 encoder. The extracted patch-level features are merged, upsampled, and refined using a DPT-like fusion stage, enabling precise depth estimation.
An improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 refines RT-DETR by introducing selective multi-scale feature extraction, a discrete sampling operator for broader deployment compatibility. These improvements yield a 0.3 to 1.4 increase in mAP metrics on the COCO dataset, all while maintaining the same parameter count and frames-per-second (FPS) performance.
Transformers' CLI welcomes a new command: chat. This command starts a conversation with the model of your choosing directly in your terminal.
This feature exists in TRL and has been migrated to transformers for easier usage.
An ongoing work is to standardize the image processors so that their API is equivalent. Additionally, the processors are given a fast variant so that they are never blockers in the image processing pipelines.
In this release, several processors have been standardized and have seen their fast version be contributed.
DPT image processors did not support segmentation_maps, instead only requiring images. This has been fixed.
This adds an argument to the preprocess method, therefore users using arguments as positional arguments with that method may see changed behavior. We recommend using keyword arguments for such methods so as to not be bothered by the addition of new features.
segmentation maps support for DPT image processor by @simonreise in #34345The problem_type in the config.json file was read incorrectly by the pipeline, which mapped single-label to multi-label losses, and vice-versa. This has been fixed.
The description of the pull request is the easiest way to understand the problem, why it exists, and how it is solved; please read the description below:
The ignore_index property of the llava configuration has been removed as it was not serving a purpose.
Quantization has received several improvements and fixes, including the contribution of FP8 quantization and the HIGGS quantization interface.
Additionally, we're replacing the AutoGPTQ implementaiton with GPTQModel from ModelCloud (see repository here)).
GPTQModel originated as major refractor of AutoGPTQ but is now a full-stand-in replacement with cleaner api, up-to-date model support, faster inference, higher quality quants.
max_length by @gante in #36120generate-related objects and methods scheduled for removal in v4.48 by @gante in #35677GenerationConfig(cache_implementation="static") by @gante in #35679SequenceBiasLogitsProcessor by @gante in #35699torch.compile(model.forward) as a fast test by @gante in #34544Pipelines have received several bug fixes and improvements which are detailed below.
test_custom_4d_attention_mask by @ydshieh in #35606EarlyStoppingCallback not require load_best_model_at_end by @muellerzr in #35101test_beam_search_low_memory by @ydshieh in #35611MobileNetV1ModelTest::test_batching_equivalence for now by @ydshieh in #35614Phi] bias should be True by @ArthurZucker in #35650Compile] Only test compiling model forward pass by @ArthurZucker in #35658zero_shot_image_classification documentation guide link in SigLIP by @aretrace in #35671Trainer cannot correctly call torch_jit_model_eval by @Wanguy in #35722pt_to_tf by @gante in #35672check_circleci_user job by @Sai-Suraj-27 in #32866MimiModel with DeepSpeed ZeRO-3 by @anferico in #34735PeftModel by @ambroser53 in #35680MimiModel with DeepSpeed ZeRO-3" by @eustlb in #35755self-comment-ci.yml by @ydshieh in #35548timm import behaviour by @rwightman in #35800test_batching_equivalence's flakiness by @ydshieh in #35729TimmWrapper by @ariG23498 in #35744timm tag to timm-wrapper models. by @pcuenca in #35794get_cached_models by @Wauplin in #35809docs/source/ar/tasks/masked_language_modeling.md into Arabic by @AhmedAlmaghz in #35198benchmark code by @gante in #35730self-comment-ci.yml by @ydshieh in #35816working-directory in self-comment-ci.yml by @ydshieh in #35833head_dim in config extracted from Gemma2 GGUF model by @Isotr0py in #35818tests] remove some flash attention class tests by @ArthurZucker in #35817num_logits_to_keep as Tensor + add flag by @Cyrilvallez in #35757test_pipelines_video_classification that was always failing by @CalOmnie in #35842Rocketknight1 to self-comment-ci.yml by @ydshieh in #35881_supports_static_cache = True for some model classes by @ydshieh in #34975test_generated_length_assisted_generation by @keyboardAnt in #34935unwrap_and_save_reload_schedule to use weights_only=False by @ydshieh in #35952squad_convert_example_to_features to work with numpy v2 by @ydshieh in #35955test_assisted_decoding_matches_greedy_search by @ydshieh in #35951transformers-pytorch-deepspeed-latest-gpu by @ydshieh in #35940Tester object has no attribute '_testMethodName' by @faaany in #35781TimmBackboneModelTest::test_batching_equivalence by @ydshieh in #35971benchmark.yml by @ydshieh in #35974generation / quantization) by @ydshieh in #35341self-comment-ci.yml by @ydshieh in #36030Qwen2VLImageProcessorFast into Qwen2VLProcessor by @yeliudev in #35987past_key_values by @yaswanth19 in #35890test_flash_attn_2_can_dispatch_composite_models by @ydshieh in #36050trainer.md by @faaany in #36066perf_infer_gpu_one.md by @faaany in #36087torch.export and fix some vision models by @qubvel in #35124output_dir Optional in TrainingArguments #27866 by @sambhavnoobcoder in #35735PretrainedConfig and PreTrainedModel by @hmellor in #36091test_initialization for VitPoseBackboneModelTest for now by @ydshieh in #36154get_default_model_revision by @MarcoGorelli in #35982DataCollatorForMultipleChoice from the docs to the package by @bauwenst in #34763check_repository_consistency run faster by MP by @ydshieh in #36175test-save-trainer by @zucchini-nlp in #36191The following contributors have made significant changes to the library over the last release:
docs/source/ar/tasks/masked_language_modeling.md into Arabic (#35198)head_dim in config extracted from Gemma2 GGUF model (#35818)DataCollatorForMultipleChoice from the docs to the package (#34763)This ends the python3.9 issues mostly!
This ends the python3.9 issues mostly!
For some very niche cases, the new rope embedding introduced device failures
num_items_in_batchFinally the fix to Gemma2 is propagated to paligemma2!
Sorry because the fixes for num_items_in_batches are not done yet 😓 To follow along see this PR, a new patch will be available soon!
Sorry because the fixes for num_items_in_batches are not done yet 😓 To follow along see this PR, a new patch will be available soon!
Now, we mostly had BC issue with python version 3.9:
Then we had a small regression for DBRX saving:
Finally we have a fix for gemma and the hybrid attention architectures:
Miscellaneous:
Yet again we are dawned with a gradient accumulation fix! There is also a refactoring of the attention that let a small typo in, we made sure PHI is n
Yet again we are dawned with a gradient accumulation fix! There is also a refactoring of the attention that let a small typo in, we made sure PHI is no longer broken!
Moonshine had a small issue when wrapping generate so we removed that!
🤗
Your coding agent can read these notes before it upgrades. Set up the MCP server →