NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #225 most downloaded on PyPI
Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Last release 11 days ago
09 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
6 versions withdrawn
withdrawn after publishing
10 years old
242 releases · first in 2016
These scripts can be found in the repo using the glob pattern models//convert_*.py. They were a recurring source of vulnerability reports and CVEs bec…
The ModernBert model was proposed in Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference by Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Galalgher, Raja Bisas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Grifin Adams, Jeremy Howard and Iacopo Poli.
It is a refresh of the traditional encoder architecture, as used in previous models such as BERT and RoBERTa.
It builds on BERT and implements many modern architectural improvements which have been developed since its original release, such as:
The Aria model was proposed in Aria: An Open Multimodal Native Mixture-of-Experts Model by Li et al. from the Rhymes.AI team.
Aria is an open multimodal-native model with best-in-class performance across a wide range of multimodal, language, and coding tasks. It has a Mixture-of-Experts architecture, with respectively 3.9B and 3.5B activated parameters per visual token and text token.
We add a TimmWrapper set of classes such that timm models can be loaded in as transformer models into the library.
Here's a general usage example:
import torch
from urllib.request import urlopen
from PIL import Image
from transformers import AutoConfig, AutoModelForImageClassification, AutoImageProcessor
checkpoint = "timm/resnet50.a1_in1k"
img = Image.open(urlopen(
'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png'
))
image_processor = AutoImageProcessor.from_pretrained(checkpoint)
inputs = image_processor(img, return_tensors="pt")
model = AutoModelForImageClassification.from_pretrained(checkpoint)
with torch.no_grad():
logits = model(**inputs).logits
top5_probabilities, top5_class_indices = torch.topk(logits.softmax(dim=1) * 100, k=5)
Thanks to this, timm models now have access to pipelines, as well as Trainer, accelerate device maps, quantization, etc:
import torch
from urllib.request import urlopen
from PIL import Image
from transformers import pipeline
img = Image.open(urlopen(
'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png'
))
pipe = pipeline("image-classification", model="timm/resnet18.a1_in1k")
print(pipe(img))
Pixtral modeling and checkpoint conversion code has been updated to support the new Pixtral-Large model.
The ColPali model was proposed in ColPali: Efficient Document Retrieval with Vision Language Models by Manuel Faysse*, Hugues Sibille*, Tony Wu*, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo (* denotes equal contribution). Work lead by ILLUIN Technology.
In the proposed ColPali approach, the authors leverage VLMs to construct efficient multi-vector embeddings directly from document images (“screenshots”) for document retrieval. They train the model to maximize the similarity between these document embeddings and the corresponding query embeddings, using the late interaction method introduced in ColBERT.
Falcon3 represents a natural evolution from previous releases, emphasizing expanding the models’ science, math, and code capabilities. This iteration includes five base models: Falcon3-1B-Base, Falcon3-3B-Base, Falcon3-Mamba-7B-Base, Falcon3-7B-Base, and Falcon3-10B-Base. In developing these models, the authors incorporated several key innovations aimed at improving the models’ performances while reducing training costs:
One pre-training: They conducted a single large-scale pretraining run on the 7B model, using 2048 H100 GPU chips, leveraging 14 trillion tokens featuring web, code, STEM, and curated high-quality and multilingual data. Depth up-scaling for improved reasoning: Building on recent studies on the effects of model depth, they upscaled the 7B model to a 10B parameters model by duplicating the redundant layers and continuing pre-training with 2TT of high-quality data. This yielded Falcon3-10B-Base which achieves state-of-the-art zero-shot and few-shot performance for models under 13B parameters. Knowledge distillation for better tiny models: To provide compact and efficient alternatives, we developed Falcon3-1B-Base and Falcon3-3B-Base by leveraging pruning and knowledge distillation techniques, using less than 100GT of curated high-quality data, thereby redefining pre-training efficiency.
Bamba-9B is a decoder-only language model based on the Mamba-2 architecture and is designed to handle a wide range of text generation tasks. It is trained from scratch using a two-stage training approach. In the first stage, the model is trained on 2 trillion tokens from the Dolma v1.7 dataset. In the second stage, it undergoes additional training on 200 billion tokens, leveraging a carefully curated blend of high-quality data to further refine its performance and enhance output quality.
Checkout all Bamba-9B model checkpoints here.
ViTPose is a state-of-the-art vision transformer-based model for human pose estimation, introduced by Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao in "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation”.
The model leverages the capabilities of vision transformers to accurately predict 2D human keypoints. Adopting a top-down approach, ViTPose estimates keypoints locations for each detected person, allowing it to be easily used with any object detection model.
The DINOv2 with Registers model was proposed in Vision Transformers Need Registers by Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski.
The Vision Transformer (ViT) is a transformer encoder model (BERT-like) originally introduced to do supervised image classification on ImageNet.
Next, people figured out ways to make ViT work really well on self-supervised image feature extraction (i.e. learning meaningful features, also called embeddings) on images without requiring any labels. Some example papers here include DINOv2 and MAE.
The authors of DINOv2 noticed that ViTs have artifacts in attention maps. It’s due to the model using some image patches as “registers”. The authors propose a fix: just add some new tokens (called “register” tokens), which you only use during pre-training (and throw away afterwards). This results in:
The Emu3 model was proposed in Emu3: Next-Token Prediction is All You Need by Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tiejun Huang, Zhongyuan Wang.
Emu3 sets a new standard in multimodal AI by using next-token prediction to handle images, text, and videos. It simplifies multimodal modeling by tokenizing all data into a unified format and training a single transformer. Visual data is tokenized using vector quantization methods based on VQ-VAE model. Discretized visual tokens are later fused with text token ids for image and text generation.
Emu3 outperforms leading models like SDXL and LLaVA-1.6 in both generation and perception tasks, without relying on diffusion or compositional methods..
A new Cohere update was added through a new "Cohere2" set of classes.
TextNet is a lightweight and efficient architecture designed specifically for text detection, offering superior performance compared to traditional models like MobileNetV3. With variants TextNet-T, TextNet-S, and TextNet-B (6.8M, 8.0M, and 8.9M parameters respectively), it achieves an excellent balance between accuracy and inference speed.
Differential Transformer combines the Llama architecture with Differential Transformer's Attention.
The conversion script needed a few update, while the modeling code was barely changed!
Moonshine is an autoregressive speech recognition encoder-decoder model that improves upon Whisper's architecture. Namely, it replaces absolute position embeddings with Rotary Position Embeddings (RoPE). This allows Moonshine to handle audio inputs of any length, unlike Whisper, which is restricted to fixed 30-second windows. It was introduced by Nat Jeffries, Evan King, Manjunath Kudlur, Guy Nicholson, James Wang, and Pete Warden in Moonshine: Speech Recognition for Live Transcription and Voice Commands .
From the VPTQ contributors:
VPTQ is a novel Post-Training Quantization method that leverages Vector Quantization to high accuracy on LLMs at an extremely low bit-width (<2-bit). VPTQ can compress 70B, even the 405B model, to 1-2 bits without retraining and maintain high accuracy.. More details here: https://github.com/microsoft/vptq
From the contributors:
HIGGS is a new 0-shot quantization algorithm that combines Hadamard preprocessing with MSE-Optimal quantization grids to achieve lower quantization error and SOTA performance. You can find more information in the paper.
Runtime support for HIGGS is implemented through FLUTE, and its library.
This PR adds support for HIGGS+FLUTE into transformers allowing for low-error 0-shot quantization and fast LLM inference.
We merged a cleanup for vision language models, to make sure it all models are standardized.
Many models in Transformers include scripts to convert the original model checkpoints into a Transformers-compatible format. These scripts can be found in the repo using the glob pattern models/**/convert_*.py. They were a recurring source of vulnerability reports and CVEs because many models were originally released using insecure formats like older PyTorch .bin weights or pickle files. The conversion scripts had to open these formats, and this meant that they were vulnerable to maliciously crafted inputs.
In practice, we do not see this as a serious vulnerability. The conversion scripts are never imported or called by the rest of the library; each script is standalone, and so the only way to exploit the vulnerability is to create a malicious checkpoint, induce a user to download it, and then also induce them to manually call a specific conversion script on it.
However, even if there is little practical risk of an exploit, we are aware that open vulnerability reports create a compliance problem for users, and so beginning with this release we will be excluding these conversion scripts from release branches and wheels. They will remain accessible to developers on the main branch.
A regular expression used within the Nougat code has been modified to ensure it does not hang. The method should output the same results but we cannot guarantee it; we recommend upgrading to the latest transformers if you use this model to ensure your code is performance-optimized.
This PR finalizes work that aimes to enable short-form (< 30 secs) and long-form generation using temperature fallback. It is a significant improvement to the whisper codebase, but it does result in the following breaking changes:
➡️ Previously:
• Short-form: Returned a ModelOutput or torch.LongTensor, including decoder input IDs and the EOS token ID.
• Long-form: Returned a Dict or torch.LongTensor, excluding decoder input IDs and the EOS token ID.
➡️ From now on:
Short-form and long-form generation are now treated identically, meaning output differentiation based on these modes is no longer applicable.
Decoder input IDs and EOS token IDs are never returned, except in two specific cases: when return_dict_in_generate=True and (return_timestamps=False or force_unique_generate_call=True).
In this case, the output will be a ModelOutput, which is the result of the underlying call to GenerationMixin’s generate. Indeed, return_timestamps=False ensures no seeking occurs; only a single call to generate is made. Therefore, this output includes both decoder input IDs and the EOS token ID.
In order to have a cleaner, isolated, future-proof code for the attention layers, they have been refactored so as to keep the model attention code within their files; but attention definitions relating to SDPA, Flash Attention, and other types of attention have been moved to a common file.
num_items_in_batch not being an integer by @xspirus in #35115docs/source/ar/community.md into Arabic by @AhmedAlmaghz in #33027AssistedCandidateGenerator for Improved Modularity and Reusability by @keyboardAnt and @jmamou in #35009Thread for SF conversion by @ydshieh in #35236rsfE with pytest by @ydshieh in #35119benchmark job in push-important-models.yml by @ydshieh in #35292benchmarks_entrypoint.py by @McPatate in #34495text by @probicheaux in #35201docs] Add link to ModernBERT Text Classification GLUE finetuning script by @tomaarsen in #35347Mamba2] Fix caching, slow path, and multi-gpu by @vasqu in #35154_make_causal_mask by @jiwoong-choi in #35291weights_only=True with torch.load for transfo_xl by @ydshieh in #35241test_generate_with_static_cache even less flaky by @ydshieh in #34995is_causal is passed explicitly by @Cyrilvallez in #35390PaliGemmaProcessor by @alvarobartt in #35278.github/workflows/self-comment-ci.yml for now by @ydshieh in #35366GPTQ, CompressedTensors] Fix unsafe imports and metada check by @vasqu in #34815ACCELERATE_MIN_VERSION on error by @KSafran in #35189model_accepts_loss_kwargs for timm model by @qubvel in #35257sdpa_kernel by @jla524 in #35410docs/source/ar/tasks/question_answering.md into Arabic by @AhmedAlmaghz in #35196docs/source/ar/tasks/summarization.md into Arabic by @AhmedAlmaghz in #35195sdpa_kernel by @jla524 in #35461The following contributors have made significant changes to the library over the last release:
Thread for SF conversion (#35236)rsfE with pytest (#35119)benchmark job in push-important-models.yml (#35292)weights_only=True with torch.load for transfo_xl (#35241)test_generate_with_static_cache even less flaky (#34995).github/workflows/self-comment-ci.yml for now (#35366)One column per quarter.
We waited a little bit to make sure it was stable, thanks @winglian for double checking and everyone for the fixes!
We waited a little bit to make sure it was stable, thanks @winglian for double checking and everyone for the fixes!
Fix GA loss bugs and add unit test (#35121) Contributed by @techkang and @ArthurZucker.
Fix num_items_in_batch not being an integer (#35115)) Contributed by @xspirus.
Fix FSDP no longer working (#35212) Contributed by @muellerzr.
Don't use no_sync when DeepSpeed doesn't support it for certain ZeRO configurations (#35212) Contributed by @winglian.
Only import torch.distributed if it is available (#35133) Contributed by @GaetanLepage.
[Whisper] Patch float type on MPS (#35295) Contributed by @eustlb. 🔜 we should probably have MPS CIs to avoid repeating this!
…as a raw Jinja file. In general, we'll be gently deprecating multiple templates in future.
PaliGemma 2 and PaliGemma are lightweight open vision-language models (VLM) inspired by PaLI-3, and based on open components like the SigLIP vision model and the Gemma language model. PaliGemma takes both images and text as inputs and can answer questions about images with detail and context, meaning that PaliGemma can perform deeper analysis of images and provide useful insights, such as captioning for images and short videos, object detection, and reading text embedded within images.
PaliGemma 2 is available in 3B, 10B, and 28B parameter sizes, which are based on Gemma 2 2B, 9B, and 27B models, respectively. The original PaliGemma models are available in the 3B size. For more information on Gemma model variants, see the Gemma models list. PaliGemma model variants support different pixel resolutions for image inputs, including 224 x 224, 448 x 448, and 896 x 896 pixels.
<img width="743" alt="image" src="https://github.com/user-attachments/assets/55cda8a6-b463-4a58-b7d3-f7d50ee2fa11">
The I-JEPA model was proposed in Image-based Joint-Embedding Predictive Architecture by Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas. I-JEPA is a self-supervised learning method that predicts the representations of one part of an image based on other parts of the same image. This approach focuses on learning semantic features without relying on pre-defined invariances from hand-crafted data transformations, which can bias specific tasks, or on filling in pixel-level details, which often leads to less meaningful representations.
<img width="413" alt="image" src="https://github.com/user-attachments/assets/561ca9d7-0327-477a-96b8-61d2af0caf34">
<img width="833" alt="image" src="https://github.com/user-attachments/assets/1abdde92-0aae-404a-b83e-77ec8bd13b7f">
The OLMo2 model is the successor of the OLMo model, which was proposed in OLMo: Accelerating the Science of Language Models.
The architectural changes from the original OLMo model to this model are:
Commits:
We add support for Meta's Layer-Skip Llama 3.2 1B model.
The Llama3.2 1B model was continually pretrained with LayerSkip recipe, early exit loss and layer dropout, as presented in Layer Skip: Enabling Early Exit Inference and Self-Speculative Decoding and is capable of performing self-speculative decoding: decode with earlier layers and verify with remaining layers.
<img width="854" alt="image" src="https://github.com/user-attachments/assets/4a9e3596-e44e-419f-804d-9f4d03f8f680">
This PR uses the torch.distributed.tensor.parallel subpackage to implement Tensor Parallel for Llama (as an example).
The motivation is multi-fold:
to make modeling code simple as single-worker case:
all manual TP implementations under if self.config.pretraining_tp > 1 can be removed.
to make tensor parallelism easily accessible by users:
added a model.tensor_parallel(device_mesh) method that allows users to turn a single-proc model into a parallel model. !- Please guide me to a right place to put this function/method if PreTrainedModel is not a preferred place. -!
This is the first PR of many to simplify and enable Tensor Parallel across models.
Python 3.8 reaches end of life, and, as such, we drop it from our CI.
Several improvements have been done to the GGUF support in transformers; notably by adding new architectures to the list of supported architectures.
use_parallel_residual and qkv_bias for StableLM GGUF config extraction by @Isotr0py in #34450We continue the work to improve the speed of fast processors as detailed in this roadmap.
We contribute a fast processor to RT-DETR.
A new pipeline has been added to transformers: image-text-to-text!
the pipeline support the following inputs:
We have had several issues with chat templates because they're stored as single lines in the JSON config files:
processor templates in chat_template.json and tokenizer templates in tokenizer_config.json causing confusionThe solution:
chat_template.jinja file in the repoProcessor classes, so processors should always be able to save their template as a raw Jinja file. In general, we'll be gently deprecating multiple templates in future.chat_template.jinja file is present, it overrides the JSON files. If a tokenizer is loaded with both Jinja and JSON chat templates and resaved, it should save only the Jinja file, and not have any chat_template entry in tokenizer_config.json.For now, we continue saving in the old format by default. I'll probably keep it this way for several versions before making the new format the default, to ensure that most users are able to load the new format before it becomes common. Until then, the new format should mostly be used for testing, to make sure it's ready for deployment when we do the switch.
This PR largely rework the logic we use in the modular converter. It is (hopefully) clearer and maintainable. Instead of going in all directions, adding stuff, then deleting it if not needed, we now do the following:
convert_tokens_to_ids by @winstxnhdw in #34030torch.fx issue related to the new loss_kwargs keyword argument by @michaelbenayoun in #34380test_eager_matches_sdpa_generate by @gante in #34386tensorflow_probability<0.22 in docker files by @ydshieh in #34381"best" for args.save_strategy. by @seanswyi in #31817model_doc/barthez.md to Korean by @Jwaminju in #33980docs/source/ar/fast_tokenizers.md into Arabic by @AhmedAlmaghz in #33034post_process_depth_estimation for GLPN by @alex-bene in #34413head_dim for mixtral model by @wavy-jung in #34281optimizer_cls_and_kwargs to Trainer.__init__ by @apoorvkh in #34358generate tests to the right mixin and delete redundant tests by @gante in #34464gc.collect and cuda.empty_cache by @ydshieh in #34514input_ids-inputs_embeds equivalence check by @gante in #34535test_eager_matches_sdpa_inference less flaky by @ydshieh in #34512docs/source/ar/multilingual.md into Arabic by @AhmedAlmaghz in #33048query_pre_attn_scalar different of num_heads in default gemma2 config by @molbap in #34540isin_mps_friendly can support 0D tensors by @gante in #34538@slow for test_eager_matches_sdpa_inference by @ydshieh in #34558convbert.md to Korean by @ahnjj in #34599timesformer.md to Korean by @mreraser in #33972docs/source/ar/trainer.md into Arabic by @AhmedAlmaghz in #33080Tool.from_space() by @aymeric-roucher in #34561docs/source/ar/torchscript.md into Arabic by @AhmedAlmaghz in #33079continue_final_message=True by @lewtun in #34253patch_size -> num_image_tokens in processing by @zucchini-nlp in #33424empty_cache device-agnostic by @faaany in #34774test_medium_seamless_m4t_pt in subprocess to avoid many failures by @ydshieh in #34812check_training_gradient_checkpointing by @ydshieh in #34806torch.export by @philkuz in #34103max_steps overriding num_train_epochs by @qgallouedec in #34810use_cache by @zucchini-nlp in #34274Deberta/Deberta-v2] Refactor code base to support compile, export, and fix LLM by @ArthurZucker in #22105peft] Given that self.active_adapter is deprecated, avoid using it by @tomaarsen in #34804test_auto_backbone_timm_model_from_pretrained by @ydshieh in #34877docs/source/ar/benchmarks.md into Arabic by @AhmedAlmaghz in #33023FlexAttention] Update gemma2 by @ArthurZucker in #34942get_max_length by @ydshieh in #34971Thread by @ydshieh in #34966release_memory() by @faaany in #34911VisitWebpageTool by @sergiopaniego in #34978save_pretrained for partially offloaded models by @kylesayrs in #34890utils/check_bad_commit.py (for auto ping in CI) by @ydshieh in #34943PixtralImageProcessorFast by @mgoin in #34836.from_pretrained type annotations by @qubvel in #34973FillMaskPipeline.__call__ signature and docstring by @alvarobartt in #35006cu_seqlens when tracing by @xenova in #35016test_eager_matches_sdpa_inference for XPU backend by @dvrogozh in #34889docs/source/ar/notebooks.md into Arabic by @AhmedAlmaghz in #33049BertGeneration by @ydshieh in #35043cuda by @faaany in #35047pad_token_tensor is None in warning by @tshu-w in #34005GPTNeoX] Flex Attention + Refactor by @vasqu in #34896tokenizers] bump to 0.21 by @ArthurZucker in #34972tie_word_embeddings handling for GGUF models by @Isotr0py in #35085trainer] fix the GA model_accepts_loss_kwargs by @ArthurZucker in #34915test_trainer.py) by @ydshieh in #35062The following contributors have made significant changes to the library over the last release:
docs/source/ar/fast_tokenizers.md into Arabic (#33034)docs/source/ar/multilingual.md into Arabic (#33048)docs/source/ar/trainer.md into Arabic (#33080)docs/source/ar/torchscript.md into Arabic (#33079)docs/source/ar/benchmarks.md into Arabic (#33023)PixtralImageProcessorFast (#34836)One small fix for FSDP + gradient accumulation loss issue!
One small fix for FSDP + gradient accumulation loss issue!
Mostly had to finish the gradient accumulation ! Thanks to @techkang and @Ryukijano 🤗
Mostly had to finish the gradient accumulation ! Thanks to @techkang and @Ryukijano 🤗
This is mostly for fx and onnx issues!
This is mostly for fx and onnx issues!
** Fix regression loading dtype #34409 by @SunMarc
** LLaVa: latency issues #34460 by @zucchini-nlp
** Fix pix2struct #34374 by @IlyasMoutawwakil
** Fix onnx non-exposable inplace aten op #34376 by @IlyasMoutawwakil
** Fix torch.fx issue related to the new loss_kwargs keyword argument #34380 by @michaelbenayoun
As part of this, we are cleaning up the input and output signatures for our pipeline classes and deprecating some rarely-used arguments. This is still…
The Moshi model was proposed in Moshi: a speech-text foundation model for real-time dialogue by Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave and Neil Zeghidour.
Moshi is a speech-text foundation model that casts spoken dialogue as speech-to-speech generation. Starting from a text language model backbone, Moshi generates speech as tokens from the residual quantizer of a neural audio codec, while modeling separately its own speech and that of the user into parallel streams. This allows for the removal of explicit speaker turns, and the modeling of arbitrary conversational dynamics. Moshi also predicts time-aligned text tokens as a prefix to audio tokens. This “Inner Monologue” method significantly improves the linguistic quality of generated speech and provides streaming speech recognition and text-to-speech. As a result, Moshi is the first real-time full-duplex spoken large language model, with a theoretical latency of 160ms, 200ms in practice.
Zamba-7B-v1 is a hybrid between state-space models (Specifically Mamba) and transformer, and was trained using next-token prediction. Zamba uses a shared transformer layer after every 6 mamba blocks. It uses the Mistral v0.1 tokenizer. We came to this architecture after a series of ablations at small scales. Zamba-7B-v1 was pre-trained on 1T tokens of text and code data.
<img width="400" alt="zamba" src="https://github.com/user-attachments/assets/a86428b8-4d24-4e5a-bf78-222312693bb2">
The GLM Model was proposed in ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools by GLM Team, THUDM & ZhipuAI.
The abstract from the paper starts with the following:
We introduce ChatGLM, an evolving family of large language models that we have been developing over time. This report primarily focuses on the GLM-4 language series, which includes GLM-4, GLM-4-Air, and GLM-4-9B.
The Idefics3 model was proposed in Building and better understanding vision-language models: insights and future directions by Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Léo Tronchon.
Idefics3 is an adaptation of the Idefics2 model with three main differences:
The PhiMoE model was proposed in Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone by Microsoft.
This model is very similar to Mixtral with the main difference of Phi3LongRoPEScaledRotaryEmbedding, where they are used to extend the context of the rotary embeddings. The query, key and values are fused, and the MLP’s up and gate projection layers are also fused.
This release adds SynthID, a novel state-of-the-art watermarking technique by Google DeepMind. SynthID has a low generation-time computational cost and can be configured to be nearly imperceptible (at the cost of harder watermarking detection). The release also comes with the code to train and run the corresponding detector, which is a machine learning model itself.
from transformers import AutoModelForCausalLM, AutoTokenizer, SynthIDTextWatermarkingConfig
tokenizer = AutoTokenizer.from_pretrained('google/gemma-2-2b', padding_side="left")
model = AutoModelForCausalLM.from_pretrained('google/gemma-2-2b')
# SynthID Text configuration
watermarking_config = SynthIDTextWatermarkingConfig(
keys=[654, 400, 836, 123, 340, 443, 597, 160, 57],
ngram_len=5,
)
# Generation with watermarking
tokenized_prompts = tokenizer(["Once upon a time, "], return_tensors="pt", padding=True)
output_sequences = model.generate(
**tokenized_prompts, watermarking_config=watermarking_config, do_sample=True, max_new_tokens=10
)
watermarked_text = tokenizer.batch_decode(output_sequences, skip_special_tokens=True)
print(watermarked_text)
Docs for applying SynthID watermarking: https://huggingface.co/docs/transformers/internal/generation_utils#transformers.SynthIDTextWatermarkLogitsProcessor Docs for detecting SynthID watermarking: https://huggingface.co/docs/transformers/internal/generation_utils#transformers.SynthIDTextWatermarkDetector
<img width="750" alt="how-synthid-works-high-level" src="https://github.com/user-attachments/assets/c5702b21-e7e6-490d-8fe6-b73783e78e6b">
BitNet is an architecture introduced by Microsoft Research that uses extreme quantization, representing each parameter with only three values: -1, 0, and 1. This results in a model that uses just 1.58 bits per parameter, significantly reducing computational and memory requirements. It replaces traditional Linear layers in Multi-Head Attention and Feed-Forward Networks with specialized layers called BitLinears that use ternary precision (or even binary, in the initial version)
More architectures are now supported in our GGUF loader; GGUF files saved with this architecture can now be loaded directly in transformers to be fine-tuned. We recommend using tooling from llama.cpp to requantize the models after further training has been done.
We are pushing for a unified inference API across multiple libraries. As part of this, we are cleaning up the input and output signatures for our pipeline classes and deprecating some rarely-used arguments. This is still a work-in-progress, but when it's finished, transformers pipelines should exactly match workflows in deployment libraries like transformers.js or TGI, allowing you to seamlessly move from development to production.
Also, pipelines now fully support the Processor class, used by vision-language models. Expect full pipeline support for chatting with VLMs in the very near future!
pipeline able to load processor by @qubvel in #32514ExecuTorch is an end-to-end solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers. It is part of the PyTorch ecosystem and supports the deployment of PyTorch models with a focus on portability, productivity, and performance.
We are collaborating with the executorch team so that 🤗 Transformers models can be exported using torch.export. The goal of this integration is not only to enable export but also to ensure that the exported artifact can be further lowered and optimized to run efficiently in ExecuTorch, particularly for mobile and edge use cases.
<img width="750" alt="how-executorch-works-high-level" src="https://github.com/user-attachments/assets/e353f9c9-b3e8-4172-86e0-c9b0b1bdd17a">
MllamaProcessor] Update errors and API with multiple image by @ArthurZucker in #33715can_generate() recursive check by @gante in #33718image_size in Convnextv2 config by @lucianosrp in #33734clean_up_tokenization_spaces] Pl bart was failing, updating by @ArthurZucker in #33735MllamaImageProcessing] Update doc by @ArthurZucker in #33747load_balancing_loss_func function of modeling_mixtral.py. by @PhilipMay in #33641modular] fixes! by @ArthurZucker in #33820prepare_inputs_for_generation to GenerationMixin by @gante in #33677accelerate dependency error in case of defaulting low_cpu_mem_usage=True by @kylesayrs in #33830compressed_tensors by @kylesayrs in #33828tokenizer kwarg deprecation with decorator by @qubvel in #33887test_static_cache_matches_dynamic as flaky by @gante in #33630SplinterTokenizer unit test by @ariepratama in #32652weights_only flag when loading state_dict by @jerryzh168 in #32481save_pretrained exception to warning by @gante in #33906logits.float() by @ringohoffman in #33902validate_rope by @zucchini-nlp in #33753PR run-slow] by @ArthurZucker in #33939self.position_embeddings->self.position_embedding by @ArthurZucker in #33958char_to_token documentation to note behaviour when trim_offsets is True by @Craigacp in #33919TF] Fix Tensorflow XLA Generation on limited seq_len models by @vasqu in #33903Red CIs] Fix hub failures by @ArthurZucker in #34001pytes collection] Fix flax test collection by @ArthurZucker in #34004gguf.md to Korean by @yijun-lee in #33764swinv2.md to Korean by @mreraser in #33566audio_utils.md to Korean by @yijun-lee in #33802esm.md to Korean by @yijun-lee in #33796time_series_utils.md to Korean by @yijun-lee in #33806pipelines_utils.md to Korean by @yijun-lee in #33809trainer.md to Korean by @yijun-lee in #33797chameleon.md to Korean by @yijun-lee in #33799logging.md to Korean by @chhaewxn in #33543auto.md to Korean by @boyunJang in #33590swin2sr.md to Korean by @mreraser in #33795vit.md to Korean by @mreraser in #33884gemma.md to Korean by @yijun-lee in #33936decoder_config=None by @SunMarc in #34014trainer_seq2seq.py's __init__ type annotations by @benglewis in #34021feature_extractor.md to Korean by @yijun-lee in #33775bertweet.md to Korean by @ahnjj in #33891gpt_neox_japanese.md to Korean by @ahnjj in #33894rag.md to Korean by @chhaewxn in #33989main_classes/quantization.md to Korean by @fabxoe in #33959main_classes/configuration.md to Korean by @fabxoe in #33952model_doc/mamba.md to Korean by @fabxoe in #33626model_doc/autoformer.md to Korean by @fabxoe in #33574model_doc/patchtsmixer.md to Korean by @fabxoe in #33587model_doc/clip.md to Korean by @fabxoe in #33610model_doc/paligemma.md to Korean by @fabxoe in #33612model_doc/llama3.md to Korean by @fabxoe in #33635model_doc/mistral.md to Korean by @fabxoe in #33648model_doc/cohere.md to Korean by @fabxoe in #33885model_doc/dbrx.md to Korean by @fabxoe in #33951model_doc/deberta-v2.md to Korean by @fabxoe in #33968main_classes/onnx.md to Korean by @fabxoe in #33601tokenization_utils.md to Korean by @yijun-lee in #33813swin.md to Korean by @mreraser in #33510file_utils.md to Korean by @yijun-lee in #33803openai-gpt.md to Korean by @yijun-lee in #33801biogpt.md to Korean by @yijun-lee in #33773blip.md to Korean by @cjfghk5697 in #33515image_processing_utils.md to Korean by @yijun-lee in #33804modular_transformers.md to Korean by @yijun-lee in #33772Patch helper] update to not have to checkout main by @ArthurZucker in #34006prepare_inputs_for_generation by @gante in #33870model_doc/bart.md to Korean by @fabxoe in #33893model_doc/deberta.md to Korean by @fabxoe in #33967main_classes/keras_callbacks.md to Korean by @fabxoe in #33955model_doc/mamba2.md to Korean by @fabxoe in #33629main_classes/model.md to Korean by @fabxoe in #33606model_doc/trajectory_transformer.md to Korean by @fabxoe in #33597model_doc/time_series_transformer.md to Korean by @fabxoe in #33596model_doc/informer.md to Korean by @fabxoe in #33585model_doc/graphormer.md to Korean by @fabxoe in #33569modeling_utils.md to Korean by @yijun-lee in #33808main_classes/data_collator.md to Korean by @fabxoe in #33954model_doc/patchtst.md to Korean by @fabxoe in #33589text_generation.md to Korean by @yijun-lee in #33777main_classes/callback.md to Korean by @Jwaminju in #33572generation_utils.md to Korean by @yijun-lee in #33818is_pipeline_test_to_skip method signature by @qubvel in #34067synced_gpus to True when using FullyShardedDataParallel by @ringohoffman in #33483logits to float() by @gante in #34042prepare_inputs_for_generation in encoder-decoder llms by @gante in #34048LlavaNextVideoForConditionalGeneration by @ydshieh in #34070generate calls with synced_gpus by @gante in #34095logits to same device as input_ids by @gante in #34076vivit.md to Korean by @mreraser in #33935gemma2.md to Korean by @yijun-lee in #33937trainer_utils.md to Korean by @yijun-lee in #33817blip-2.md to Korean by @cjfghk5697 in #33516accelerate error caused by 46d09af by @steveepreston in #34197trainer._get_eval_sampler() to support group_by_length arg by @larin92 in #33514prepare_inputs_for_generation by @gante in #34199require_torch_up_to_2_accelerators by @byi8220 in #34201MLFLOW_MAX_LOG_PARAMS to MLflowCallback by @cecheta in #34279executorch.md to Korean by @ahnjj in #33888bert japanese.md to Korean by @ahnjj in #33890model_doc/bartpho.md to Korean by @Jwaminju in #33981The following contributors have made significant changes to the library over the last release:
MllamaProcessor] Update errors and API with multiple image (#33715)clean_up_tokenization_spaces] Pl bart was failing, updating (#33735)MllamaImageProcessing] Update doc (#33747)modular] fixes! (#33820)PR run-slow] (#33939)self.position_embeddings->self.position_embedding (#33958)Red CIs] Fix hub failures (#34001)pytes collection] Fix flax test collection (#34004)Patch helper] update to not have to checkout main (#34006)TF] Fix Tensorflow XLA Generation on limited seq_len models (#33903)LlavaNextVideoForConditionalGeneration (#34070)logits.float() (#33902)synced_gpus to True when using FullyShardedDataParallel (#33483)gguf.md to Korean (#33764)audio_utils.md to Korean (#33802)esm.md to Korean (#33796)time_series_utils.md to Korean (#33806)pipelines_utils.md to Korean (#33809)trainer.md to Korean (#33797)chameleon.md to Korean (#33799)gemma.md to Korean (#33936)feature_extractor.md to Korean (#33775)tokenization_utils.md to Korean (#33813)file_utils.md to Korean (#33803)openai-gpt.md to Korean (#33801)biogpt.md to Korean (#33773)image_processing_utils.md to Korean (#33804)modular_transformers.md to Korean (#33772)modeling_utils.md to Korean (#33808)text_generation.md to Korean (#33777)generation_utils.md to Korean (#33818)gemma2.md to Korean (#33937)trainer_utils.md to Korean (#33817)main_classes/quantization.md to Korean (#33959)main_classes/configuration.md to Korean (#33952)model_doc/mamba.md to Korean (#33626)model_doc/autoformer.md to Korean (#33574)model_doc/patchtsmixer.md to Korean (#33587)model_doc/clip.md to Korean (#33610)model_doc/paligemma.md to Korean (#33612)model_doc/llama3.md to Korean (#33635)model_doc/mistral.md to Korean (#33648)model_doc/cohere.md to Korean (#33885)model_doc/dbrx.md to Korean (#33951)model_doc/deberta-v2.md to Korean (#33968)main_classes/onnx.md to Korean (#33601)model_doc/bart.md to Korean (#33893)model_doc/deberta.md to Korean (#33967)main_classes/keras_callbacks.md to Korean (#33955)model_doc/mamba2.md to Korean (#33629)main_classes/model.md to Korean (#33606)model_doc/trajectory_transformer.md to Korean (#33597)model_doc/time_series_transformer.md to Korean (#33596)model_doc/informer.md to Korean (#33585)model_doc/graphormer.md to Korean (#33569)main_classes/data_collator.md to Korean (#33954)model_doc/patchtst.md to Korean (#33589)Mostly some warnings that were not properly removed ⚠️ :
Mostly some warnings that were not properly removed ⚠️ :
🔴 Had a small regression with dynamic Cache 🔴 *Cache: revert DynamicCache init for BC #33861 by @gante
A small fix for idefic 🐩 :
And a fix for Siglip 🤧 !
[MllamaProcessor] Update errors and API with multiple image (#33715) by @ArthurZucker
Fix typo: depracted -> deprecated by @tomaarsen in #32489
The Llama 3.2-Vision collection of multimodal large language models (LLMs) is a collection of pretrained and instruction-tuned image reasoning generative models in 11B and 90B sizes (text + images in / text out). The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. The models outperform many of the available open source and closed multimodal models on common industry benchmarks.
The Qwen2-VL is a major update from the previous Qwen-VL by the Qwen team.
An extract from the Qwen2-VL blogpost available here is as follows:
Qwen2-VL is the latest version of the vision language models based on Qwen2 in the Qwen model familities. Compared with Qwen-VL, Qwen2-VL has the capabilities of:
The Qwen2-Audio is the new model series of large audio-language models from the Qwen team. Qwen2-Audio is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions.
They introduce two distinct audio interaction modes:
OLMoE is a series of Open Language Models using sparse Mixture-of-Experts designed to enable the science of language models. The team releases all code, checkpoints, logs, and details involved in training these models.
LLaVA-Onevision is a Vision-Language Model that can generate text conditioned on one or several images/videos. The model consists of SigLIP vision encoder and a Qwen2 language backbone. The images are processed with anyres-9 technique where the image is split into 9 patches to better process high resolution images and capture as much details as possible. However, videos are pooled to a total sequence length of 196 tokens each frame for more memory efficient computation. LLaVA-Onevision is available in three sizes: 0.5B, 7B and 72B and achieves remarkable performance on benchmark evaluations.
The FalconMamba model was proposed by TII UAE (Technology Innovation Institute) in their release.
The model has been trained on approximtely 6T tokens consisting a mixture of many data sources such as RefineWeb, Cosmopedia and Math data.
The team releases an accompanying blog post.
he Granite model was proposed in Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler by Yikang Shen, Matthew Stallone, Mayank Mishra, Gaoyuan Zhang, Shawn Tan, Aditya Prasad, Adriana Meza Soria, David D. Cox and Rameswar Panda.
PowerLM-3B is a 3B state-of-the-art small language model trained with the Power learning rate scheduler. It is trained on a wide range of open-source and synthetic datasets with permissive licenses. PowerLM-3B has shown promising results compared to other models in the size categories across various benchmarks, including natural language multi-choices, code generation, and math reasoning.
The GraniteMoe model was proposed in Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler by Yikang Shen, Matthew Stallone, Mayank Mishra, Gaoyuan Zhang, Shawn Tan, Aditya Prasad, Adriana Meza Soria, David D. Cox and Rameswar Panda.
PowerMoE-3B is a 3B sparse Mixture-of-Experts (sMoE) language model trained with the Power learning rate scheduler. It sparsely activates 800M parameters for each token. It is trained on a mix of open-source and proprietary datasets. PowerMoE-3B has shown promising results compared to other dense models with 2x activate parameters across various benchmarks, including natural language multi-choices, code generation, and math reasoning.
The Descript Audio Codec (DAC) model is a powerful tool for compressing audio data, making it highly efficient for storage and transmission. By compressing 44.1 KHz audio into tokens at just 8kbps bandwidth, the DAC model enables high-quality audio processing while significantly reducing the data footprint. This is particularly useful in scenarios where bandwidth is limited or storage space is at a premium, such as in streaming applications, remote conferencing, and archiving large audio datasets.
The Pixtral model was released by the Mistral AI team. Pixtral is a multimodal model, taking images and text as input, and producing text as output. This model follows the Llava family, meaning image embeddings are placed instead of the [IMG] token placeholders.
The model uses PixtralVisionModel for its vision encoder, and MistralForCausalLM for its language decoder. The main contribution is the 2d ROPE (rotary postiion embeddings) on the images, and support for arbitrary image sizes (the images are not padded together nor are they resized).
The Mimi model was proposed in Moshi: a speech-text foundation model for real-time dialogue by Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave and Neil Zeghidour. Mimi is a high-fidelity audio codec model developed by the Kyutai team, that combines semantic and acoustic information into audio tokens running at 12Hz and a bitrate of 1.1kbps. In other words, it can be used to map audio waveforms into “audio tokens”, known as “codebooks”.
The OmDet-Turbo model was proposed in Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head by Tiancheng Zhao, Peng Liu, Xuan He, Lu Zhang, Kyusong Lee. OmDet-Turbo incorporates components from RT-DETR and introduces a swift multimodal fusion module to achieve real-time open-vocabulary object detection capabilities while maintaining high accuracy. The base model achieves performance of up to 100.2 FPS and 53.4 AP on COCO zero-shot.
GGUF support continues to be enhanced in the library by offering a way to load GGUF models within transformers by unquantizing them, before re-quantizing them for re-use within the GGUF/GGML ecosystem.
An ongoing effort is to add the ability to use torchao as a quantization backend. Future PRs will enable saving and fine-tuning with peft.
The Liger kernel is now supported in the Trainer class.
This PR introduces Modularity for transformers, which has always been prohibited when working with transformers (see blog post for the accompanying design philosophy).
The core idea behind this PR is to facilitate model addition by enabling Pythonic inheritance while keeping true to our single-file policy in which models/processors must be contained within a single file, enabling working around the object without going through 10 layers of abstractions.
It is heavily recommended to read the PR description in order to understand the depth of the change: https://github.com/huggingface/transformers/pull/33248
transformers: modularity and inheritance for new model additions by @ArthurZucker in #33248Agents continue being improved at each release; this time making it much simpler to leverage a local engine through a local Transformers Engine.
This PR adds to all decoder-only models (except for XLNet) support for dynamic cache.
The documentation for the Dynamic cache can be found here, and documentation related to the KV cache in transformers in general can be found here.
We've made several updates to our handling of chat models and chat templates. The most noticeable change is that assistant prefill is now supported. This means you can end a chat with an assistant message, and the model will continue that message instead of starting a new one, allowing you to guide the model's response:
pipe = pipeline("text-generation", model_checkpoint)
chat = [
{"role": "user", "content": "Can you format the answer in JSON?"},
{"role": "assistant", "content": '{"name": "'}
]
output = pipe(chat) # The model will continue outputting JSON!
We've also enabled several new functionalities in Jinja that will allow more powerful templates in future, including Loop Controls and a strftime_now function that can get the current date and time, which is commonly used in system messages. For more details, see the updated chat template docs.
mask_generation.md to Korean by @jeongiin in #32257idefics.md to Korean by @boyunJang in #32258image_to_image.md to Korean by @shinhyunji36 in #32327gptq.md to Korean by @1kmmk1 in #32293prompting.md to Korean by @chhaewxn in #32294quantization/quanto.md to Korean by @fabxoe in #32281image_feature_extraction.md to Korean by @mreraser in #32239chat_templating.md to Korean by @enchantee00 in #32362_supports_sdpa to True by @pocca2048 in #32457ko-llm_tutorial_optimization.md to Korean by @010kim in #32372trainer.md to Korean by @cjfghk5697 in #32260eetq.md to Korean by @jun048098 in #32352fsdp.md to Korean by @win2dvp21 in #32261bitsandbytes.md to Korean by @SeungAhSon in #32408inputs_embeds as input by @molbap in #32493test_static_cache_exportability with torch 2.4.0 by @guangy10 in #32516agent.md to Korean by @Jwaminju in #32351encodec model names by @Sai-Suraj-27 in #32581.push_to_hub(..., create_pr=True, revision="my-branch") when creating PR on not-owned repo by @Wauplin in #32094deepspeed.md to Korean by @4N3MONE in #32431awq.mdto Korean by @ahnjj in #32324test_find_base_model_checkpoint by @Sai-Suraj-27 in #32638is_torch_mps_available() function to include min_version argument by @Sai-Suraj-27 in #32545transformers tag to the modelcard by @LysandreJik in #32623WhisperGenerationMixin by @faaany in #32316test_tokenization_utils.py by @Sai-Suraj-27 in #32601tests/utils/test_add_new_model_like.py by @Sai-Suraj-27 in #32678JetMoeIntegrationTest by @ydshieh in #32332doctest_glob by @Sai-Suraj-27 in #32475 falcon-mamba-7b model checkpoint name by @Sai-Suraj-27 in #32837LogitsWarper and LogitsProcessor by @gante in #32626batch_size instead of max_batch_size by @gante in #32657to in DoLa body, causing exceptions in multi-gpu generation by @gante in #32856test_sdpa_can_compile_dynamic device-agnostic by @faaany in #32519whisper-large-v2 model link in docs by @Sai-Suraj-27 in #32871norm_before_gate usage by @vasqu in #32686tensor.norm() with decomposed version for CLIP executorch export by @qubvel in #32887return_timestamps when return_timestamps is not passed to generate function by @hrl in #31296huggingface_hub installation to workflows by @Sai-Suraj-27 in #32891exceptions.ConnectionError by @younesbelkada in #31469AttributeError raised when using Trainer with eval_on_start=True in Jupyter Notebook. by @fshp971 in #32849Processor.save_pretrained caused by #31691 by @leloykun in #32921use_cache=False by @gante in #32863PretrainedConfig from saving generate parameters; Update deprecations in generate-related code 🧹 by @gante in #32659atol in test_forward_with_num_logits_to_keep by @gante in #33093isin_mps_friendly, a wrapper function for torch.isin by @gante in #33099pydantic required version in dockerfiles to make it compatible with DeepSpeed by @Sai-Suraj-27 in #33105efficientnet pipeline timeout and prevent future similar issues due to large image size by @gante in #33123conversations.md to Korean by @newfull5 in #32468llm_optims.md to Korean by @yijun-lee in #32325return_dict_in_generate is False but should be True by @gante in #33146bitsandbytes) in docstrings by @rapsealk in #33230torch.from_numpy() to create tensors for np.ndarrays by @shinyano in #33201num_logits_to_keep in composite models by @zucchini-nlp in #33168FalconMamba training issues due to incompatible kernels by @younesbelkada in #33195torch.jit.trace for interpolate_pos_encoding in all vision models by @xenova in #33226inputs_embeds by @zucchini-nlp in #32932transformers[en] Documentation by @nnilayy in #33350FalconMambaForCausalLM by @younesbelkada in #33381FbgemmFp8Linear not preserving tensor shape by @vgel in #33239Zero-shot object detection documentation by @sergiopaniego in #33430SSH into runner info. to DM by @ydshieh in #33346train with a script by @faaany in #33423padding_side as call time kwargs by @zucchini-nlp in #33385Agents and tools documentation links typos by @sergiopaniego in #33471Agents, supercharged - Multi-agents, External tools, and more docs typo fixed by @sergiopaniego in #33478docs/source/ar/_toctree.yml by @AhmedAlmaghz in #32696accelerator.use_fp16 in examples by @hlky in #33513sequences_scores in the Whisper beam search output by @Nik-Kras in #32970model.config and model.generation_config 🔫 by @gante in #33480past_key_values is None by @gante in #33541attention_mask is 2D by @gante in #33575Mamba2] Move dt calculations to kernel by @vasqu in #33520gemma2 when instantiating a new cache by @gante in #33595test_generate_from_inputs_embeds_decoder_only by @gante in #33602torch_job by @ydshieh in #33593PreTrainedModel inheriting from GenerationMixin by @gante in #33203cache_implementation) by @gante in #33684The following contributors have made significant changes to the library over the last release:
chat_templating.md to Korean (#32362)ko-llm_tutorial_optimization.md to Korean (#32372)trainer.md to Korean (#32260)exceptions.ConnectionError (#31469)FalconMamba training issues due to incompatible kernels (#33195)FalconMambaForCausalLM (#33381)deepspeed.md to Korean (#32431)docs/source/ar/_toctree.yml (#32696)Patch release v4.44.2, mostly 2 regressions that were not caught for Jamba and for processors!
Patch release v4.44.2, mostly 2 regressions that were not caught for Jamba and for processors!
is_torchdynamo_compiling -- cast a wide exception net (#32476) by @gante
Full Changelog: https://github.com/huggingface/transformers/compare/v4.44.0...v4.44.1
fix: Replaced deprecated unittest method with the correct one by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32198
This release comes a bit early in our cycle because we wanted to ship important and requested models along with improved performances for everyone!
All of these are included with examples in the awesome https://github.com/huggingface/local-gemma repository! 🎈 We tried to share examples of what is now possible with all the shipped features! Kudos to @gante, @sanchit-gandhi and @xenova
Generate: end-to-end compilation #30788 by @gante: model.generate now supports compiling! There are a few limitations, but here is a small snippet:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
import copy
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Meta-Llama-3.1-8B", torch_dtype=torch.bfloat16, device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3.1-8B")
# compile generate
compiled_generate = torch.compile(model.generate, fullgraph=True, mode="reduce-overhead")
# compiled generate does NOT accept parameterization except a) model inputs b) a generation config
generation_config = copy.deepcopy(model.generation_config)
generation_config.pad_token_id = model.config.eos_token_id
model_inputs = tokenizer(["Write a poem about the market crashing in summer"], return_tensors="pt")
model_inputs = model_inputs.to(model.device)
output_compiled = compiled_generate(**model_inputs, generation_config=generation_config)
print(output_compiled)
model.forward = torch.compile(model.forward, mode="reduce-overhead", fullgraph=True)cache_implementation="offloaded" when calling from_pretrained or using this:from transformers import GenerationConfig
gen_config = GenerationConfig(cache_implementation="offloaded", # other generation options such as num_beams=4,num_beam_groups=2,num_return_sequences=4,diversity_penalty=1.0,max_new_tokens=50,early_stopping=True)
outputs = model.generate(inputs["input_ids"],generation_config=gen_config)
pytorch team gave us a great gift: you can now use torch.export directly compatible with Executorch! Find examples here.
This also unlocks support for prompt reuse:
import os, torch, copy
from transformers import AutoModelForCausalLM, AutoTokenizer, DynamicCache
device = "cuda"
ckpt = "meta-llama/Meta-Llama-3.1-8B-Instruct"
INITIAL_PROMPT = "From now on, you are going to answer all my questions with historical details. Make sure to always add a bit of french here and there, for style."
model = AutoModelForCausalLM.from_pretrained(ckpt, torch_dtype=torch.float16)
model.to(device)
tokenizer = AutoTokenizer.from_pretrained(ckpt)
prompt_cache = DynamicCache()
inputs = tokenizer(INITIAL_PROMPT, return_tensors="pt").to("cuda")
prompt_cache = model(**inputs, past_key_values = prompt_cache).past_key_values
prompt = "Why are french people obsessed with french?"
new_inputs = tokenizer(INITIAL_PROMPT + prompt, return_tensors="pt").to("cuda")
past_key_values = copy.deepcopy(prompt_cache)
outputs = model.generate(**new_inputs, past_key_values=past_key_values,max_new_tokens=20)
response = tokenizer.batch_decode(outputs)[0]
print(response)
prompt = "What is the best city to swim in?"
new_inputs = tokenizer(INITIAL_PROMPT + prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**new_inputs, past_key_values=copy.deepcopy(prompt_cache),max_new_tokens=20)
response = tokenizer.batch_decode(outputs)[0]
Gemma 2: support assisted generation #32357 by @gante
We now have a 2B Gemma 2 model -- a perfect sidekick for the 27B with assisted generation. We've enabled assisted generation in gemma 2, with a caveat: assisted generation currently requires the use of a windowless cache (as opposed to the default cache for gemma 2), so you might observe some output mismatch on long sequences. Read more about it here.
# transformers assisted generation reference:
# https://huggingface.co/docs/transformers/main/en/llm_optims#speculative-decoding
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# we DON’T recommend using the 9b model with the 2b model as its assistant
assistant_model_name = 'google/gemma-2-2b-it'
reference_model_name = 'google/gemma-2-27b-it'
tokenizer = AutoTokenizer.from_pretrained(reference_model_name)
model = AutoModelForCausalLM.from_pretrained(
reference_model_name, device_map='auto', torch_dtype=torch.bfloat16
)
assistant_model = AutoModelForCausalLM.from_pretrained(
assistant_model_name, device_map='auto', torch_dtype=torch.bfloat16
)
model_inputs = tokenizer("Einstein's theory of relativity states", return_tensors="pt").to(model.device)
generation_options = {
"assistant_model": assistant_model,
"do_sample": True,
"temperature": 0.7,
"max_new_tokens": 64,
}
outputs = model.generate(**model_inputs, **generation_options)
tokenizer.batch_decode(outputs, skip_special_tokens=True)
Nemotron-4-340B-Instruct is a large language model (LLM) that can be used as part of a synthetic data generation pipeline to create training data that helps researchers and developers build their own LLMs. It is a fine-tuned version of the Nemotron-4-340B-Base model, optimized for English-based single and multi-turn chat use-cases. It supports a context length of 4,096 tokens.
The conversion script should be able to cover Minitron and Nemotron, thanks and kudos to @suiyoubi. See:
Codestral is trained on a diverse dataset of 80+ programming languages, including the most popular ones, such as Python, Java, C, C++, JavaScript, and Bash. It also performs well on more specific ones like Swift and Fortran. This broad language base ensures Codestral can assist developers in various coding environments and projects.
Codestral saves developers time and effort: it can complete coding functions, write tests, and complete any partial code using a fill-in-the-middle mechanism. Interacting with Codestral will help level up the developer’s coding game and reduce the risk of errors and bugs.
It's mamba2 architecture, was a bit of a pain to remove all einops but hope we made it better for everyone!
We removed the chat template in the code, they should all be on the hub!
Our great @sanchit-gandhi worked on porting the recent compile upgrades to long form decoding in
ruff to the latest version by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/31926unittest method with the correct one by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32198eos for assisted decoding by @zucchini-nlp in https://github.com/huggingface/transformers/pull/31301object base class by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32230target_sizes is None in post_process_image_guided_detection for owlv2 by @catalys1 in https://github.com/huggingface/transformers/pull/31934static cache implementation is not compatible with attn_implementation==flash_attention_2 by @faaany in https://github.com/huggingface/transformers/pull/32039convert_blip_checkpoint function call by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32262check_docstrings by @gante in https://github.com/huggingface/transformers/pull/32259p_mask a numpy array before passing to select_starts_ends by @faaany in https://github.com/huggingface/transformers/pull/32076gguf==0.9.1 by @Isotr0py in https://github.com/huggingface/transformers/pull/32298fetch-depth: 0 in trufflehog checkout step by @McPatate in https://github.com/huggingface/transformers/pull/31663inv_freq assignment by @gante in https://github.com/huggingface/transformers/pull/323303-5x faster torch.compile forward compilation for autoregressive decoder models by @fxmarty in https://github.com/huggingface/transformers/pull/32227
staticmethods with self as first argument by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32361speech dep to the consistency docker image by @gante in https://github.com/huggingface/transformers/pull/32374transformers/examples/flax/language-modeling/t5_tokenizer_model.py. by @fshp971 in https://github.com/huggingface/transformers/pull/32157test_embeded_special_tokens for luke and mluke models by @Sai-Suraj-27 in https://github.com/huggingface/transformers/pull/32413preprocess with decorator by @qubvel in https://github.com/huggingface/transformers/pull/32024Full Changelog: https://github.com/huggingface/transformers/compare/v4.43.4...v4.44.0
There was a mick mack, now deepseep issues are properly pushed with:
There was a mick mack, now deepseep issues are properly pushed with:
🤗 Enjoy holidays
Patch release v4.43.3: We still saw some bugs so @zucchini-nlp added: ~- Resize embeds with DeepSpeed #32214~
Patch release v4.43.3: We still saw some bugs so @zucchini-nlp added: ~- Resize embeds with DeepSpeed #32214~
Other fixes:
Fix float8_e4m3fn in modeling_utils
Nothing published for this version
…their performance. In practice, this is a breaking change.
The Llama 3.1 models are released by Meta and come in three flavours: 8B, 70B, and 405B.
To get an overview of Llama 3.1, please visit the Hugging Face announcement blog post.
We release a repository of llama recipes to showcase usage for inference, total and partial fine-tuning of the different variants.
The Chameleon model was proposed in Chameleon: Mixed-Modal Early-Fusion Foundation Models by META AI Chameleon Team. Chameleon is a Vision-Language Model that use vector quantization to tokenize images which enables the model to generate multimodal output. The model takes images and texts as input, including an interleaved format, and generates textual response.
The ZoeDepth model was proposed in ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth by Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, Matthias Müller. ZoeDepth extends the DPT framework for metric (also called absolute) depth estimation. ZoeDepth is pre-trained on 12 datasets using relative depth and fine-tuned on two domains (NYU and KITTI) using metric depth. A lightweight head is used with a novel bin adjustment design called metric bins module for each domain. During inference, each input image is automatically routed to the appropriate head using a latent classifier.
Hiera was proposed in Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles by Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei, Haoqi Fan, Po-Yao Huang, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, Jitendra Malik, Yanghao Li, Christoph Feichtenhofer
The paper introduces “Hiera,” a hierarchical Vision Transformer that simplifies the architecture of modern hierarchical vision transformers by removing unnecessary components without compromising on accuracy or efficiency. Unlike traditional transformers that add complex vision-specific components to improve supervised classification performance, Hiera demonstrates that such additions, often termed “bells-and-whistles,” are not essential for high accuracy. By leveraging a strong visual pretext task (MAE) for pretraining, Hiera retains simplicity and achieves superior accuracy and speed both in inference and training across various image and video recognition tasks. The approach suggests that spatial biases required for vision tasks can be effectively learned through proper pretraining, eliminating the need for added architectural complexity.
Our ReactAgent has a specific way to return its final output: it calls the tool final_answer, added to the user-defined toolbox upon agent initialization, with the answer as the tool argument. We found that even for a one-shot agent like CodeAgent, using a specific final_answer tools helps the llm_engine find what to return: so we generalized the final_answer tool for all agents.
Now if your code-based agent (like ReactCodeAgent) defines a function at step 1, it will remember the function definition indefinitely. This means your agent can create its own tools for later re-use!
This is a transformative PR: it allows the agent to regularly run a specific step for planning its actions in advance. This gets activated if you set an int for planning_interval upon agent initialization. At step 0, a first plan will be done. At later steps (like steps 3, 6, 9 if you set planning_interval=3 ), this plan will be updated by the agent depending on the history of previous steps. More detail soon!
A significant RoPE refactor was done to make it model agnostic and more easily adaptable to any architecture. It is only applied to Llama for now but will be applied to all models using RoPE over the coming days.
🚨🚨 This PR changes the code to rely on the tokenizer's defaults when these flags are unset. This means some models using TextGenerationPipeline previously did not add a <bos> by default, which (negatively) impacted their performance. In practice, this is a breaking change.
Example of a script changed as a result of this PR:
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2-9b-it")
model = AutoModelForCausalLM.from_pretrained("google/gemma-2-9b-it", torch_dtype=torch.bfloat16, device_map="auto")
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
print(pipe("Foo bar"))
get_seq_length method by @sanchit-gandhi in #31661keras-nlp<0.14 pin by @gante in #31684tets/test_xxx_utils.py) to tests/utils by @ydshieh in #31730pytest_num_workers=4 for some CircleCI jobs by @ydshieh in #31764sdpa support for SigLIP by @qubvel in #31499TFBlipModelTest::test_pipeline_image_to_text by @ydshieh in #31827TrainingArguments by @andstor in #31812vocab_size in other two VLMs by @zucchini-nlp in #31681.generate() by @voidism in #29619_init_weights for ResNetPreTrainedModel by @ydshieh in #31851_init_weights for ResNetPreTrainedModel" by @ydshieh in #31868duplicate field definitions in some classes by @Sai-Suraj-27 in #31888push_to_hub=True in TrainingArguments by @SunMarc in #31808warnings in a with block to avoid flaky tests by @ydshieh in #31893ConvertSlow] make sure the order is preserved for addedtokens by @ArthurZucker in #31902Gemma2] Support FA2 softcapping by @ArthurZucker in #318871st argument name in classmethods by @Sai-Suraj-27 in #31907SlidingWindowCache.reset() by @gante in #31917Trainer.get_optimizer_cls_and_kwargs to be overridden by @apoorvkh in #31875GenerationMixin.generate compatibility with pytorch profiler by @fxmarty in #31935Cache and cache_position being default by @gante in #31898sigmoid_focal_loss() function call by @Sai-Suraj-27 in #31951logits_warper update in models with custom generate fn by @gante in #31957create_repo() function call by @Sai-Suraj-27 in #31947test_stage3_nvme_offload by @faaany in #31881src/transformers/__init__.py by @Sai-Suraj-27 in #31993log messages that are resulting in TypeError due to too many arguments by @Sai-Suraj-27 in #32017SeamlessM4Tv2ConformerEncoderLayer.forward() when gradient checkpointing is enabled by @anferico in #31945sdpa and FA2 for CLIP by @qubvel in #31940numpy<2.0 by @ydshieh in #32018head_dim through config (and do not require head_dim * num_heads == hidden_size) by @xenova in #32050duplicate entries in a dictionary by @Sai-Suraj-27 in #32041huggingface_hub 0.24 by @Wauplin in #32054mktemp() function by @Sai-Suraj-27 in #32123ko/_toctree.yml and remove custom_tools.md to reflect latest changes by @jungnerd in #31969TypeError instead of ValueError for invalid type by @Sai-Suraj-27 in #32111trust_remote_code when loading Libri Dummy by @sanchit-gandhi in #31748GPTNeoX and GPT2 by @vasqu in #31944The following contributors have made significant changes to the library over the last release:
.generate() (#29619)but also fix the sliding window for long context and other typos.
but also fix the sliding window for long context and other typos.
Was off last week could not get this out, thanks all for your patience 🥳
After experimenting, we noticed that for the 27b model mostly, softcapping is a must. So adding it back (it should have been there, but an error on my
After experimenting, we noticed that for the 27b model mostly, softcapping is a must. So adding it back (it should have been there, but an error on my side made it disappear) sorry all! 😭
Thanks to our 2 contributors for their prompt fixing mostly applies for training and FA2!
Thanks to our 2 contributors for their prompt fixing mostly applies for training and FA2!
[HybridCache] Fix get_seq_length method
Patch release for commit:
This has the possibility of being a slight breaking change if users are creating models and relying on out_indices on being a tuple. As this property…
The Gemma2 model was proposed in Gemma2: Open Models Based on Gemini Technology and Research by Gemma2 Team, Google. Gemma2 models are trained on 6T tokens, and released with 2 versions, 2b and 7b.
The abstract from the paper is the following:
This work introduces Gemma2, a new family of open language models demonstrating strong performance across academic benchmarks for language understanding, reasoning, and safety. We release two sizes of models (2 billion and 7 billion parameters), and provide both pretrained and fine-tuned checkpoints. Gemma2 outperforms similarly sized open models on 11 out of 18 text-based tasks, and we present comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of our model development. We believe the responsible release of LLMs is critical for improving the safety of frontier models, and for enabling the next wave of LLM innovations
The RT-DETR model was proposed in DETRs Beat YOLOs on Real-time Object Detection by Wenyu Lv, Yian Zhao, Shangliang Xu, Jinman Wei, Guanzhong Wang, Cheng Cui, Yuning Du, Qingqing Dang, Yi Liu.
RT-DETR is an object detection model that stands for “Real-Time DEtection Transformer.” This model is designed to perform object detection tasks with a focus on achieving real-time performance while maintaining high accuracy. Leveraging the transformer architecture, which has gained significant popularity in various fields of deep learning, RT-DETR processes images to identify and locate multiple objects within them.
The InstructBLIP model was proposed in InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi. InstructBLIP leverages the BLIP-2 architecture for visual instruction tuning.
InstructBLIP uses the same architecture as BLIP-2 with a tiny but important difference: it also feeds the text prompt (instruction) to the Q-Former.
The LLaVa-NeXT-Video model was proposed in LLaVA-NeXT: A Strong Zero-shot Video Understanding Model by Yuanhan Zhang, Bo Li, Haotian Liu, Yong Jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, Chunyuan Li. LLaVa-NeXT-Video improves upon LLaVa-NeXT by fine-tuning on a mix if video and image dataset thus increasing the model’s performance on videos.
LLaVA-NeXT surprisingly has strong performance in understanding video content in zero-shot fashion with the AnyRes technique that it uses. The AnyRes technique naturally represents a high-resolution image into multiple images. This technique is naturally generalizable to represent videos because videos can be considered as a set of frames (similar to a set of images in LLaVa-NeXT). The current version of LLaVA-NeXT makes use of AnyRes and trains with supervised fine-tuning (SFT) on top of LLaVA-Next on video data to achieves better video understanding capabilities.The model is a current SOTA among open-source models on VideoMME bench.
A very significant change makes its way within the transformers codebase, introducing a new way to add models to transformers. We recommend reading the description of the PR below, but here is the gist of it:
The diff_converter tool is here to replace our old Copied from statements, while keeping our core transformers philosophy:
- single model single file
- explicit code
- standardization of modeling code
- readable and educative code
- simple code
- least amount of modularity
This additionally unlocks the ability to very quickly see the differences between new architectures that get developed. While many architectures are similar, the "single model, single file" policy can obfuscate the changes. With this diff converter, we want to make the changes between architectures very explicit.
We've made major updates to our support for tool-use and RAG models. We can now automatically generate JSON schema descriptions for Python functions which are suitable for passing to tool models, and we've defined a standard API for tool models which should allow the same tool inputs to be used with many different models. Models will need updates to their chat templates to support the new API, and we're targeting the Nous-Hermes, Command-R and Mistral/Mixtral model families for support in the very near future. Please see the updated chat template docs for more information.
If you are the owner of a model that supports tool use, but you're not sure how to update its chat template to support the new API, feel free to reach out to us for assistance with the update, for example on the Hugging Face Discord server. Ping Matt and yell key phrases like "chat templates" and "Jinja" and your issue will probably get resolved.
We further the support of GGUF files to offer fine-tuning within the python/HF ecosystem, before converting them back to the GGUF/GGML/llama.cpp libraries.
A new optimizer is added in the Trainer.
Several improvements are done related to quantization: a new cache (the quantized KV cache) is added, offering the ability to convert the cache of generative models, further reducing the memory requirements.
Additionally, the documentation related to quantization is entirely redone with the aim of helping users choose which is the best quantization method.
New instance segmentation examples are added by @qubvel
As a notable improvement to the HF vision models that leverage backbones, we enable leveraging HF pretrained model weights as backbones, with the following API:
from transformers import MaskFormerConfig, MaskFormerForInstanceSegmentation
config = MaskFormerConfig(backbone="microsoft/resnet-50", use_pretrained_backbone=True)
model = MaskFormerForInstanceSegmentation(config)
Additionally, we thank @Cyrilvallez for diving into our generate method and greatly reducing the memory requirements.
generate() 🔥🔥🔥 by @Cyrilvallez in #30536Both the ConversationalPipeline and the Conversation object have been deprecated for a while, and are due for removal in 4.42, which is the upcoming version.
The TextGenerationPipeline is recommended for this use-case, and now accepts inputs in the form of the OpenAI API.
Removes duplicate softmax application in FLAVA attention. Likely to have a small change on the outputs but flagging with 🚨 as it will change a bit.
ignore_index attribute of the loss is updated to -100timm being updatedRecent updates to timm changed the type of the attribute model.feature_info.out_indices. Previously, out_indices would reflect the input type of out_indices on the create_model call i.e. either tuple or list. Now, this value is always a tuple.
As list are more useful and consistent for us -- we cannot save tuples in configs, they must be converted to lists first -- we instead choose to cast out_indices to always be a list.
This has the possibility of being a slight breaking change if users are creating models and relying on out_indices on being a tuple. As this property only happens when a new model is created, and not if it's saved and reloaded (because of the config), then I think this has a low chance of having much of an impact.
mamba slow forward by @vasqu in #30691tokenizer_class = "AutoTokenizer" Llava Family by @ArthurZucker in #30912optimum-benchmark by @ydshieh in #30615torch.use_deterministic_algorithms for XPU by @faaany in #30774MptIntegrationTests expected outputs by @ydshieh in #30989uv==0.1.45 by @ydshieh in #31006test_model_parallelism device-agnostic by @faaany in #30844test_model_parallelism for 2 model test classes by @ydshieh in #31067@main by @ydshieh in #31065ninja from docker image build by @ydshieh in #31080accelerate as a hard requirement by @younesbelkada in #31090OPTForQuestionAnswering by @younesbelkada in #31092test_multi_gpu_data_parallel_forward for vit and deit by @ydshieh in #31086HF_HUB_OFFLINE + fix has_file in offline mode by @Wauplin in #31016transformers-cli env reporting by @statelesshz in #31003load_in_8bit with bnb config by @younesbelkada in #31136IS_GITHUB_CI by @younesbelkada in #31147GemmaModel] fix small typo by @ArthurZucker in #31202test_compile_static_cache by @ydshieh in #30991mistral.py::Mask4DTestHard by @ydshieh in #31212MistralIntegrationTest by @ydshieh in #31231BlipModel by @younesbelkada in #31235name 'torch' is not defined in bitsandbytes integration by @jamesbraza in #31243benchmark job in push-important-models.yml by @ydshieh in #31259SwitchTransformer] Significant performance improvement on MoE blocks by @ranggihwang in #31173cached_download to hf_hub_download in remaining occurrences by @Wauplin in #31284str should be used not int when setting env variables by @statelesshz in #31272decoder_attention_mask shape by @ylacombe in #28071inputs_embeds padding logger.warning to logger.warning_once by @naimenz in #31411tokenizer being popped twice by @gante in #31427TestDeepSpeedModelZoo device-agnostic by @faaany in #31402dataloader_persistent_workers=True by @bastienlc in #30627Qwen2ForTokenClassification by @kevinhu in #31440generate call from local path by @gante in #31470PreTrainedTokenizerFast loading time when there are many added tokens by @ydshieh in #31404metric_for_best_model errors by @tomaarsen in #31450GPT2] Add SDPA support by @vasqu in #31172test_config_object to test_ds_config_object by @faaany in #31403torch.compile support for AQLM by @younesbelkada in #31473wandb integration with SetFit model by @timothepearce in #30021tokenization_utils_base.py's docstring by @sadra-barikbin in #31510spectrogram_batch by @ravenouse in #27159TrainingArguments by @qgallouedec in #31503_no_split_module by @zucchini-nlp in #31566i18n by @SauravMaheshkar in #31584self.projection call in VivitTubeletEmbeddings by @v-iashin in #31632GPT-NeoX] Add SDPA support by @vasqu in #31031past_key_values passed as kwargs by @gante in #31644The following contributors have made significant changes to the library over the last release:
mamba slow forward (#30691)GPT2] Add SDPA support (#31172)GPT-NeoX] Add SDPA support (#31031)generate() 🔥🔥🔥 (#30536)spectrogram_batch (#27159)Mostly fixing some stuff related to trust_remote_code=True and from_pretrained
Mostly fixing some stuff related to trust_remote_code=True and from_pretrained
The local_file_only was having a hard time when a .safetensors file did not exist. This is not expected and instead of trying to convert, we should just fallback to loading the .bin files.
The causal mask and label creation was causing label leaks when training. Kudos to @probicheaux for finding and reporting!
The causal mask and label creation was causing label leaks when training. Kudos to @probicheaux for finding and reporting!
Other fixes:
Reverted https://github.com/huggingface/transformers/commit/4ab7a28216211571fdddba414d4edd8426ab6489
🚨🚨🚨Deprecate evaluation_strategy to eval_strategy🚨🚨🚨 by @muellerzr in https://github.com/huggingface/transformers/pull/30190
The Phi-3 model was proposed in Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone by Microsoft.
TLDR; Phi-3 introduces new ROPE scaling methods, which seems to scale fairly well! A 3b and a Phi-3-mini is available in two context-length variants—4K and 128K tokens. It is the first model in its class to support a context window of up to 128K tokens, with little impact on quality.
<img width="1599" alt="image" src="https://github.com/huggingface/transformers/assets/48595927/0f37c6b0-b118-453c-ac64-6e45aa291d0a">
JetMoe-8B is an 8B Mixture-of-Experts (MoE) language model developed by Yikang Shen and MyShell. JetMoe project aims to provide a LLaMA2-level performance and efficient language model with a limited budget. To achieve this goal, JetMoe uses a sparsely activated architecture inspired by the ModuleFormer. Each JetMoe block consists of two MoE layers: Mixture of Attention Heads and Mixture of MLP Experts. Given the input tokens, it activates a subset of its experts to process them. This sparse activation schema enables JetMoe to achieve much better training throughput than similar size dense models. The training throughput of JetMoe-8B is around 100B tokens per day on a cluster of 96 H100 GPUs with a straightforward 3-way pipeline parallelism strategy.
<img width="1559" alt="image" src="https://github.com/huggingface/transformers/assets/48595927/cc83ce99-7a61-4d04-a234-3f68e6c0fafd">
PaliGemma is a lightweight open vision-language model (VLM) inspired by PaLI-3, and based on open components like the SigLIP vision model and the Gemma language model. PaliGemma takes both images and text as inputs and can answer questions about images with detail and context, meaning that PaliGemma can perform deeper analysis of images and provide useful insights, such as captioning for images and short videos, object detection, and reading text embedded within images.
More than 120 checkpoints are released see the collection here !
<img width="1064" alt="image" src="https://github.com/huggingface/transformers/assets/48595927/23584b9a-6c36-46f5-8700-32f402c0f674">
Video-LLaVA exhibits remarkable interactive capabilities between images and videos, despite the absence of image-video pairs in the dataset.
💡 Simple baseline, learning united visual representation by alignment before projection With the binding of unified visual representations to the language feature space, we enable an LLM to perform visual reasoning capabilities on both images and videos simultaneously. 🔥 High performance, complementary learning with video and image Extensive experiments demonstrate the complementarity of modalities, showcasing significant superiority when compared to models specifically designed for either images or videos.
<img width="532" alt="image" src="https://cdn-uploads.huggingface.co/production/uploads/62441d1d9fdefb55a0b7d12c/cLniWc__KECBBesliHKhd.png">
<img width="1024" alt="image" src="https://falconllm.tii.ae/assets/images/table-1___.jpeg">
Two new models from TII-UAE! They published a blog-post with more details! Falcon2 introduces parallel mlp, and falcon VLM uses the Llava framework
from_pretrained support<img width="1064" alt="image" src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/gguf-spec.png">
You can now load most of the GGUF quants directly with transformers' from_pretrained to convert it to a classic pytorch model. The API is simple:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF"
filename = "tinyllama-1.1b-chat-v1.0.Q6_K.gguf"
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
We plan more closer integrations with llama.cpp / GGML ecosystem in the future, see: https://github.com/huggingface/transformers/issues/27712 for more details
v4.41.0 introduces a significant refactor of the Agents framework.
With this release, we allow you to build state-of-the-art agent systems, including the React Code Agent that writes its actions as code in ReAct iterations, following the insights from Wang et al., 2024
Just install with pip install "transformers[agents]". Then you're good to go!
from transformers import ReactCodeAgent
agent = ReactCodeAgent(tools=[])
code = """
list=[0, 1, 2]
for i in range(4):
print(list(i))
"""
corrected_code = agent.run(
"I have some code that creates a bug: please debug it and return the final code",
code=code,
)
In this release we support new quantization methods: HQQ & EETQ contributed by the community. Read more about how to quantize any transformers model using HQQ & EETQ in the dedicated documentation section
dequantize API for bitsandbytes modelsIn case you want to dequantize models that have been loaded with bitsandbytes, this is now possible through the dequantize API (e.g. to merge adapter weights)
dequantize API for bitsandbytes quantized models by @younesbelkada in https://github.com/huggingface/transformers/pull/30806API-wise, you can achieve that with the following:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer
model_id = "facebook/opt-125m"
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=BitsAndBytesConfig(load_in_4bit=True))
tokenizer = AutoTokenizer.from_pretrained(model_id)
model.dequantize()
text = tokenizer("Hello my name is", return_tensors="pt").to(0)
out = model.generate(**text)
print(tokenizer.decode(out[0]))
min_p sampling by @gante in https://github.com/huggingface/transformers/pull/30639Gemma work with torch.compile by @ydshieh in https://github.com/huggingface/transformers/pull/30775BERT] Add support for sdpa by @hackyon in https://github.com/huggingface/transformers/pull/28802Addition of fine-tuning script for object detection models
Add interpolation of embeddings. This enables predictions from pretrained models on input images of sizes different than those the model was originally trained on. Simply pass interpolate_pos_embedding=True when calling the model.
Added for: BLIP, BLIP 2, InstructBLIP, SigLIP, ViViT
import requests
from PIL import Image
from transformers import Blip2Processor, Blip2ForConditionalGeneration
image = Image.open(requests.get("https://huggingface.co/hf-internal-testing/blip-test-image/resolve/main/demo.jpg", stream=True).raw)
processor = Blip2Processor.from_pretrained("Salesforce/blip2-opt-2.7b")
model = Blip2ForConditionalGeneration.from_pretrained(
"Salesforce/blip2-opt-2.7b",
torch_dtype=torch.float16
).to("cuda")
inputs = processor(images=image, size={"height": 500, "width": 500}, return_tensors="pt").to("cuda")
predictions = model(**inputs, interpolate_pos_encoding=True)
# Generated text: "a woman and dog on the beach"
generated_text = processor.batch_decode(predictions, skip_special_tokens=True)[0].strip()
evaluation_strategy to eval_strategy🚨🚨🚨 by @muellerzr in https://github.com/huggingface/transformers/pull/30190LlamaTokenizerFast] Refactor default llama by @ArthurZucker in https://github.com/huggingface/transformers/pull/28881prev_ci_results by @ydshieh in https://github.com/huggingface/transformers/pull/30313pad token id in pipeline forward arguments by @zucchini-nlp in https://github.com/huggingface/transformers/pull/30285jnp import in utils/generic.py by @ydshieh in https://github.com/huggingface/transformers/pull/30322AssertionError in clip conversion script by @ydshieh in https://github.com/huggingface/transformers/pull/30321pad_token_id again by @zucchini-nlp in https://github.com/huggingface/transformers/pull/30338Llama family, fix use_cache=False generation by @ArthurZucker in https://github.com/huggingface/transformers/pull/30380-rs to show skip reasons by @ArthurZucker in https://github.com/huggingface/transformers/pull/30318require_torch_sdpa for test that needs sdpa support by @faaany in https://github.com/huggingface/transformers/pull/30408LlamaTokenizerFast] Refactor default llama by @ArthurZucker in https://github.com/huggingface/transformers/pull/28881Llava] + CIs fix red cis and llava integration tests by @ArthurZucker in https://github.com/huggingface/transformers/pull/30440paths filter to avoid the chance of being triggered by @ydshieh in https://github.com/huggingface/transformers/pull/30453utils/check_if_new_model_added.py by @ydshieh in https://github.com/huggingface/transformers/pull/30456research_project] Most of the security issues come from this requirement.txt by @ArthurZucker in https://github.com/huggingface/transformers/pull/29977WandbCallback with third parties by @tomaarsen in https://github.com/huggingface/transformers/pull/30477SourceFileLoader.load_module() in dynamic module loading by @XuehaiPan in https://github.com/huggingface/transformers/pull/30370HfQuantizer quant method update by @younesbelkada in https://github.com/huggingface/transformers/pull/30484bitsandbytes error formatting ("Some modules are dispatched on ...") by @kyo-takano in https://github.com/huggingface/transformers/pull/30494dtype_byte_size to handle torch.float8_e4m3fn/float8_e5m2 types by @mgoin in https://github.com/huggingface/transformers/pull/30488DETR] Remove timm hardcoded logic in modeling files by @amyeroberts in https://github.com/huggingface/transformers/pull/29038_load_best_model by @muellerzr in https://github.com/huggingface/transformers/pull/30553use_cache in kwargs for GPTNeoX by @zucchini-nlp in https://github.com/huggingface/transformers/pull/30538use_square_size after loading by @ydshieh in https://github.com/huggingface/transformers/pull/30567output_router_logits in SwitchTransformers by @lausannel in https://github.com/huggingface/transformers/pull/30573contiguous() in clip checkpoint conversion script by @ydshieh in https://github.com/huggingface/transformers/pull/30613generate-related rendering issues by @gante in https://github.com/huggingface/transformers/pull/30600StoppingCriteria autodocs by @gante in https://github.com/huggingface/transformers/pull/30617SinkCache on Llama models by @gante in https://github.com/huggingface/transformers/pull/30581None as attention when layer is skipped by @jonghwanhyeon in https://github.com/huggingface/transformers/pull/30597TextGenerationPipeline._sanitize_parameters from overriding previously provided parameters by @yting27 in https://github.com/huggingface/transformers/pull/30362CI update] Try to use dockers and no cache by @ArthurZucker in https://github.com/huggingface/transformers/pull/29202resume_download deprecation by @Wauplin in https://github.com/huggingface/transformers/pull/30620cache_position initialisation for generation with use_cache=False by @nurlanov-zh in https://github.com/huggingface/transformers/pull/30485forward in Idefics2ForConditionalGeneration with correct ignore_index value by @zafstojano in https://github.com/huggingface/transformers/pull/30678workflow_id in utils/get_previous_daily_ci.py by @ydshieh in https://github.com/huggingface/transformers/pull/30695prev_ci_results to ci_results by @ydshieh in https://github.com/huggingface/transformers/pull/30697model.active_adapters() instead of deprecated model.active_adapter whenever possible by @younesbelkada in https://github.com/huggingface/transformers/pull/30738actions/post-slack with centrally defined workflow by @younesbelkada in https://github.com/huggingface/transformers/pull/30737model_parallel = False to T5ForTokenClassification and MT5ForTokenClassification by @retarfi in https://github.com/huggingface/transformers/pull/30763WhisperGenerationMixin by @cifkao in https://github.com/huggingface/transformers/pull/29688Optional in typing. by @xkszltl in https://github.com/huggingface/transformers/pull/30821torch 2.3 for CI by @ydshieh in https://github.com/huggingface/transformers/pull/30837Cache but not static cache by @gante in https://github.com/huggingface/transformers/pull/30800Full Changelog: https://github.com/huggingface/transformers/compare/v4.40.2...v4.41.0
Fix copies for DBRX - neuron fix
Thanks @michaelbenayoun !
Kudos to @pcuenca for the prompt fix in:
Kudos to @pcuenca for the prompt fix in:
To support EosTokenCriteria on MPS while pytorch adds this functionality.
[generate] fix breaking change for patch by @ArthurZucker in #29976
Llama 3 is supported in this release through the Llama 2 architecture and some fixes in the tokenizers library.
<img src="https://huggingface.co/HuggingFaceM4/idefics-80b/resolve/main/assets/IDEFICS.png" alt="drawing" width="300"/>
The Idefics2 model was created by the Hugging Face M4 team and authored by Léo Tronchon, Hugo Laurencon, Victor Sanh. The accompanying blog post can be found here.
Idefics2 is an open multimodal model that accepts arbitrary sequences of image and text inputs and produces text outputs. The model can answer questions about images, describe visual content, create stories grounded on multiple images, or simply behave as a pure language model without visual inputs. It improves upon IDEFICS-1, notably on document understanding, OCR, or visual reasoning. Idefics2 is lightweight (8 billion parameters) and treats images in their native aspect ratio and resolution, which allows for varying inference efficiency.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/recurrent-gemma.png" alt="drawing" width="600"/>
<small> Recurrent Gemma architecture. Taken from the <a href="https://arxiv.org/pdf/2402.19427.pdf">original paper.</a> </small>
The Recurrent Gemma model was proposed in RecurrentGemma: Moving Past Transformers for Efficient Open Language Models by the Griffin, RLHF and Gemma Teams of Google.
The abstract from the paper is the following:
We introduce RecurrentGemma, an open language model which uses Google’s novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve excellent performance on language. It has a fixed-sized state, which reduces memory use and enables efficient inference on long sequences. We provide a pre-trained model with 2B non-embedding parameters, and an instruction tuned variant. Both models achieve comparable performance to Gemma-2B despite being trained on fewer tokens.
Jamba is a pretrained, mixture-of-experts (MoE) generative text model, with 12B active parameters and an overall of 52B parameters across all experts. It supports a 256K context length, and can fit up to 140K tokens on a single 80GB GPU.
As depicted in the diagram below, Jamba’s architecture features a blocks-and-layers approach that allows Jamba to successfully integrate Transformer and Mamba architectures altogether. Each Jamba block contains either an attention or a Mamba layer, followed by a multi-layer perceptron (MLP), producing an overall ratio of one Transformer layer out of every eight total layers.
Jamba introduces the first HybridCache object that allows it to natively support assisted generation, contrastive search, speculative decoding, beam search and all of the awesome features from the generate API!
DBRX is a transformer-based decoder-only large language model (LLM) that was trained using next-token prediction. It uses a fine-grained mixture-of-experts (MoE) architecture with 132B total parameters of which 36B parameters are active on any input.
It was pre-trained on 12T tokens of text and code data. Compared to other open MoE models like Mixtral-8x7B and Grok-1, DBRX is fine-grained, meaning it uses a larger number of smaller experts. DBRX has 16 experts and chooses 4, while Mixtral-8x7B and Grok-1 have 8 experts and choose 2.
This provides 65x more possible combinations of experts and the authors found that this improves model quality. DBRX uses rotary position encodings (RoPE), gated linear units (GLU), and grouped query attention (GQA).
The OLMo model was proposed in OLMo: Accelerating the Science of Language Models by Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi.
OLMo is a series of Open Language Models designed to enable the science of language models. The OLMo models are trained on the Dolma dataset. We release all code, checkpoints, logs (coming soon), and details involved in training these models.
Qwen2MoE is the new model series of large language models from the Qwen team. Previously, we released the Qwen series, including Qwen-72B, Qwen-1.8B, Qwen-VL, Qwen-Audio, etc.
Model Details Qwen2MoE is a language model series including decoder language models of different model sizes. For each size, we release the base language model and the aligned chat model. Qwen2MoE has the following architectural choices:
Qwen2MoE is based on the Transformer architecture with SwiGLU activation, attention QKV bias, group query attention, mixture of sliding window attention and full attention, etc. Additionally, we have an improved tokenizer adaptive to multiple natural languages and codes. Qwen2MoE employs Mixture of Experts (MoE) architecture, where the models are upcycled from dense language models. For instance, Qwen1.5-MoE-A2.7B is upcycled from Qwen-1.8B. It has 14.3B parameters in total and 2.7B activated parameters during runtime, while it achieves comparable performance with Qwen1.5-7B, with only 25% of the training resources.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/grouding_dino_architecture.png" alt="drawing" width="600"/>
<small> Taken from the <a href="https://arxiv.org/pdf/2303.05499.pdf">original paper.</a> </small>
The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang. Grounding DINO extends a closed-set object detection model with a text encoder, enabling open-set object detection. The model achieves remarkable results, such as 52.5 AP on COCO zero-shot.
Static pretrained maps have been removed from the library's internals and are currently deprecated. These used to reflect all the available checkpoints for a given architecture on the Hugging Face Hub, but their presence does not make sense in light of the huge growth of checkpoint shared by the community.
With the objective of lowering the bar of model contributions and reviewing, we first start by removing legacy objects such as this one which do not serve a purpose.
Processors are ungoing changes in order to uniformize them and make them clearer to use.
Pipelines can now be pushed to Hub using a convenient push_to_hub method.
push_to_hub to pipeline by @not-lain in #29172Thanks to the community contribution, Flash Attention 2 has been integrated for more architectures
- and the from custom_tools.md by @windsonsea in #29767-OO mode for docstring_decorator by @matthid in #29689Latest PyTorch + TensorFlow [dev] by @ydshieh in #29764LlavaNext] Fix llava next unsafe imports by @ArthurZucker in #29773set_seed by @muellerzr in #29778torch_dtype in the run_mlm example by @jla524 in #29776bos token to Blip generations by @zucchini-nlp in #29642quality] update quality check to make sure we check imports 😈 by @ArthurZucker in #29771vocab_size by @fxmarty in #29389AssistedCandidateGenerator by @gante in #29787cleanup] vestiges of causal mask by @ArthurZucker in #29806SuperPoint] Fix doc example by @amyeroberts in #29816bos_token_id is None during the generation with inputs_embeds by @LZHgrla in #29772cosine_with_min_lr scheduler in Trainer by @liuyanyi in #29341num_attention_heads != num_key_value_heads in Flax Llama Implementation by @bminixhofer in #29557slow_forward gradient fix by @vasqu in #29563eos_token_id to stopping criteria by @zucchini-nlp in #29459make fix-copies] update and help by @ArthurZucker in #29924GptNeox] don't gather on pkv when using the trainer by @ArthurZucker in #29892pipeline]. Zero shot add doc warning by @ArthurZucker in #29845xpu to the testing documentation by @faaany in #29894torch.testing.assert_allclose by torch.testing.assert_close by @gante in #29915Mamba] from pretrained issue with self.embeddings by @ArthurZucker in #29851TokenizationLlama] fix the way we convert tokens to strings to keep leading spaces 🚨 breaking fix by @ArthurZucker in #29453BC] Fix BC for other libraries by @ArthurZucker in #29934LlamaSlowConverter] Slow to Fast better support by @ArthurZucker in #29797StableLm] Add QK normalization and Parallel Residual Support by @jon-tow in #29745test_eager_matches_sdpa_generate flaky for some models by @ydshieh in #29479run_qa.py by @jla524 in #29867BC] Fix BC for AWQ quant by @TechxGenus in #29965ImageToTextPipelineTests.test_conditional_generation_llava by @faaany in #29975generate] fix breaking change for patch by @ArthurZucker in #29976_replace_with_bnb_linear by @SunMarc in #29958skip_special_tokens for Wav2Vec2CTCTokenizer._decode by @msublee in #29311remove_columns in text-classification example by @mariosasko in #29351tests/utils/tiny_model_summary.json by @ydshieh in #29941kwargs handling in generate_with_fallback by @cifkao in #29225WhisperNoSpeechDetection when recomputing scores by @cifkao in #29248Main CIs] Fix the red cis by @ArthurZucker in #30022ProcessingIdefics] Attention mask bug with padding by @byi8220 in #29449whisper to IMPORTANT_MODELS by @ydshieh in #30046test_encode_decode_fast_slow_all_tokens for now by @ydshieh in #30044llm_int8_enable_fp32_cpu_offload=True...." instead of "load_in_8bit_fp32_cpu_offload=True". by @miRx923 in #30013torch.fx symbolic tracing for LLama by @michaelbenayoun in #30047require_bitsandbytes marker by @faaany in #30116mps as device for Pipeline class by @fnhirwa in #30080itemize by @younesbelkada in #30162ruff configuration to avoid deprecated configuration warning by @Sai-Suraj-27 in #30179CI] Add new workflow to run slow tests of important models on push main if they are modified by @younesbelkada in #29235logger.warn with logger.warning by @Sai-Suraj-27 in #30197RecurrentGemmaIntegrationTest.test_2b_sample by @ydshieh in #30222assertEquals with assertEqual by @Sai-Suraj-27 in #30241typing.Text with str by @Sai-Suraj-27 in #30230type annotation for compatability with python 3.8 by @Sai-Suraj-27 in #30243docs/source/en) by @ydshieh in #30247require_torch_multi_gpu flag by @faaany in #30250ko/_toctree.yml by @jungnerd in #30062raise statement by @Sai-Suraj-27 in #30275Idefics2's doc example by @ydshieh in #30274ExamplesTests::test_run_translation by @ydshieh in #30281Fatal Python error: Bus error in ZeroShotAudioClassificationPipelineTests by @ydshieh in #30283The following contributors have made significant changes to the library over the last release:
generate fix breaking change for patch #29976
The AWQ issue persisted, and there was a regression reported with beam search and input embeddings.
Series of fixes for backwards compatibility (AutoAWQ and other quantization libraries, imports from trainer_pt_utils) and functionality (LLaMA tokeniz
Series of fixes for backwards compatibility (AutoAWQ and other quantization libraries, imports from trainer_pt_utils) and functionality (LLaMA tokenizer conversion)
BC] Fix BC for other libraries #29934LlamaSlowConverter] Slow to Fast better support #29797Patch release to fix some breaking changes to LLaVA model, fixes/cleanup for Cohere & Gemma and broken doctest
Patch release to fix some breaking changes to LLaVA model, fixes/cleanup for Cohere & Gemma and broken doctest
vocab_size #29389cleanup] vestiges of causal mask #29806SuperPoint] Fix doc example (https://github.com/huggingface/transformers/pull/29816)The PRs below introduced slightly breaking changes that we believed was necessary for the repository; if these seem to impact your usage of transforme…
The Llama, Cohere and the Gemma model both no longer cache the triangular causal mask unless static cache is used. This was reverted by #29753, which fixes the BC issues w.r.t speed , and memory consumption, while still supporting compile and static cache. Small note, fx is not supported for both models, a patch will be brought very soon!
Command-R is a generative model optimized for long context tasks such as retrieval augmented generation (RAG) and using external APIs and tools. It is designed to work in concert with Cohere's industry-leading Embed and Rerank models to provide best-in-class integration for RAG applications and excel at enterprise use cases. As a model built for companies to implement at scale, Command-R boasts:
Llava next is the next version of Llava, which includes better support for non padded images, improved reasoning, OCR, and world knowledge. LLaVA-NeXT even exceeds Gemini Pro on several benchmarks.
Compared with LLaVA-1.5, LLaVA-NeXT has several improvements:
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/llava_next_overview.png" alt="drawing" width="600"/>
<small> LLaVa-NeXT incorporates a higher input resolution by encoding various patches of the input image. Taken from the <a href="https://arxiv.org/abs/2310.03744">original paper.</a> </small>
The MusicGen Melody model was proposed in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi and Alexandre Défossez.
MusicGen Melody is a single stage auto-regressive Transformer model capable of generating high-quality music samples conditioned on text descriptions or audio prompts. The text descriptions are passed through a frozen text encoder model to obtain a sequence of hidden-state representations. MusicGen is then trained to predict discrete audio tokens, or audio codes, conditioned on these hidden-states. These audio tokens are then decoded using an audio compression model, such as EnCodec, to recover the audio waveform.
Through an efficient token interleaving pattern, MusicGen does not require a self-supervised semantic representation of the text/audio prompts, thus eliminating the need to cascade multiple models to predict a set of codebooks (e.g. hierarchically or upsampling). Instead, it is able to generate all the codebooks in a single forward pass.
The PVTv2 model was proposed in PVT v2: Improved Baselines with Pyramid Vision Transformer by Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. As an improved variant of PVT, it eschews position embeddings, relying instead on positional information encoded through zero-padding and overlapping patch embeddings. This lack of reliance on position embeddings simplifies the architecture, and enables running inference at any resolution without needing to interpolate them.
The UDOP model was proposed in Unifying Vision, Text, and Layout for Universal Document Processing by Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, Mohit Bansal. UDOP adopts an encoder-decoder Transformer architecture based on T5 for document AI tasks like document image classification, document parsing and document visual question answering.
<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/udop_architecture.jpg" alt="drawing" width="600"/>
<small> UDOP architecture. Taken from the <a href="https://arxiv.org/abs/2212.02623">original paper.</a> </small>
This model is a new paradigm architecture based on state-space-models, rather than attention like transformer models. The checkpoints are compatible with the original ones
Add Mamba] Adds support for the Mamba models by @ArthurZucker in #28094StarCoder2 is a family of open LLMs for code and comes in 3 different sizes with 3B, 7B and 15B parameters. The flagship StarCoder2-15B model is trained on over 4 trillion tokens and 600+ programming languages from The Stack v2. All models use Grouped Query Attention, a context window of 16,384 tokens with a sliding window attention of 4,096 tokens, and were trained using the Fill-in-the-Middle objective.
The SegGPT model was proposed in SegGPT: Segmenting Everything In Context by Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, Tiejun Huang. SegGPT employs a decoder-only Transformer that can generate a segmentation mask given an input image, a prompt image and its corresponding prompt mask. The model achieves remarkable one-shot results with 56.1 mIoU on COCO-20 and 85.6 mIoU on FSS-1000.
With Galore, you can pre-train large models on consumer-type hardwares, making LLM pre-training much more accessible to anyone from the community.
Our approach reduces memory usage by up to 65.5% in optimizer states while maintaining both efficiency and performance for pre-training on LLaMA 1B and 7B architectures with C4 dataset with up to 19.7B tokens, and on fine-tuning RoBERTa on GLUE tasks. Our 8-bit GaLore further reduces optimizer memory by up to 82.5% and total training memory by 63.3%, compared to a BF16 baseline. Notably, we demonstrate, for the first time, the feasibility of pre-training a 7B model on consumer GPUs with 24GB memory (e.g., NVIDIA RTX 4090) without model parallel, checkpointing, or offloading strategies.
Galore is based on low rank approximation of the gradients and can be used out of the box for any model.
Below is a simple snippet that demonstrates how to pre-train mistralai/Mistral-7B-v0.1 on imdb:
import torch
import datasets
from transformers import TrainingArguments, AutoConfig, AutoTokenizer, AutoModelForCausalLM
import trl
train_dataset = datasets.load_dataset('imdb', split='train')
args = TrainingArguments(
output_dir="./test-galore",
max_steps=100,
per_device_train_batch_size=2,
optim="galore_adamw",
optim_target_modules=["attn", "mlp"]
)
model_id = "mistralai/Mistral-7B-v0.1"
config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_config(config).to(0)
trainer = trl.SFTTrainer(
model=model,
args=args,
train_dataset=train_dataset,
dataset_text_field='text',
max_seq_length=512,
)
trainer.train()
Quanto has been integrated with transformers ! You can apply simple quantization algorithms with few lines of code with tiny changes. Quanto is also compatible with torch.compile
Check out the announcement blogpost for more details
Exllama and AWQ combined together for faster AWQ inference - check out the relevant documentation section for more details on how to use Exllama + AWQ.
Allow models saved or fine-tuned with Apple’s MLX framework to be loaded in transformers (as long as the model parameters use the same names), and improve tensor interoperability. This leverages MLX's adoption of safetensors as their checkpoint format.
Notable memory reduction in Gemma/LLaMa by changing the causal mask buffer type from int64 to boolean.
torch.bool instead of torch.int64 for non-persistant causal mask buffer by @fxmarty in #29241The PRs below introduced slightly breaking changes that we believed was necessary for the repository; if these seem to impact your usage of transformers, we recommend checking out the PR description to get more insights in how to leverage the new behavior.
Gemma] Fix bad rebase with transformers main by @younesbelkada in #29170torch.compile with fullgraph=True when attention_mask input is used by @fxmarty in #29211Doc] update model doc qwen2 by @ArthurZucker in #29238is_vision_available result by @bmuskalla in #29280DS_DISABLE_NINJA=1 by @ydshieh in #29290non_device_test pytest mark to filter out non-device tests by @fxmarty in #29213dtype and device extraction for CUDA graph generation for quantizers compatibility by @BlackSamorez in #29079attn_implementation documentation by @fxmarty in #29295GenerationMixin's docstring by @sadra-barikbin in #29277Gemma / CI] Make sure our runners have access to the model by @younesbelkada in #29242require_read_token] fix typo by @ArthurZucker in #29345T5 and Llama Tokenizer] remove warning by @ArthurZucker in #29346Llama ROPE] Fix torch export but also slow downs in forward by @ArthurZucker in #29198output_router_logits during inference by @LeonardoEmili in #29249CI / starcoder2] Change starcoder2 path to correct one for slow tests by @younesbelkada in #29359CI]: Fix failing tests for peft integration by @younesbelkada in #29330CI] require_read_token in the llama FA2 test by @younesbelkada in #29361get_values(MODEL_MAPPING) by @ydshieh in #29362offload_buffers parameter of accelerate to PreTrainedModel.from_pretrained method by @notsyncing in #28755quantization / ESM] Fix ESM 8bit / 4bit with bitsandbytes by @younesbelkada in #29329Llama + AWQ] fix prepare_inputs_for_generation 🫠 by @ArthurZucker in #29381YOLOS] Fix - return padded annotations by @amyeroberts in #29300AutoProcessor by @JingyaHuang in #29169post_process_instance_segmentation for panoptic tasks by @nickthegroot in #29304Generation] Fix some issues when running the MaxLength criteria on CPU by @younesbelkada in #29317UdopTokenizer] Fix post merge imports by @ArthurZucker in #29451Udop imports] Processor tests were not run. by @ArthurZucker in #29456import_path location by @loadams in #29154offload_weight() takes from 3 to 4 positional arguments but 5 were given by @faaany in #29457Docs / Awq] Add docs on exllamav2 + AWQ by @younesbelkada in #29474docs] Add starcoder2 docs by @younesbelkada in #29454pad_to_multiple_of by @gante in #29462TextGenerationPipeline.__call__ docstring by @alvarobartt in #29491inputs as kwarg in TextClassificationPipeline by @alvarobartt in #29495VisionEncoderDecoder Positional Arg by @nickthegroot in #29497require_sacremoses decorator by @faaany in #29504torch_device instead of auto for model testing by @faaany in #29531TrainingArguments by @yundai424 in #29189n_gpu in TrainerIntegrationTest::test_train_and_eval_dataloaders for XPU by @faaany in #29307warning_advice for tensorflow warning by @winstxnhdw in #29540Mamba doc] Post merge updates by @ArthurZucker in #29472Docs] fixed minor typo by @j-gc in #29555main branch by @ydshieh in #28816max_position_embeddings in the translation example by @gante in #29600Gemma] Supports converting directly in half-precision by @younesbelkada in #29529MaskFormer, Mask2Former] Use einsum where possible by @amyeroberts in #29544Mask2Former] Move normalization for numerical stability by @amyeroberts in #29542test_trainer_log_level_replica to run on accelerators with more than 2 devices by @faaany in #29609multi_gpu_data_parallel_forward for MusicgenTest by @ydshieh in #29632PEFT] Fix save_pretrained to make sure adapters weights are also saved on TPU by @shub-kris in #29388dataset_revision argument to RagConfig by @ydshieh in #29610cache_position update in generate by @gante in #29467generation_config by @gante in #29675filter_models by @ydshieh in #29673glue to nyu-mll/glue by @lhoestq in #29679filter_models" by @ydshieh in #29682filter_models by @ydshieh in #29710bnb] Make unexpected_keys optional by @younesbelkada in #29420gradio.Interface.from_pipeline by @abidlabs in #29684The following contributors have made significant changes to the library over the last release:
We mostly made sure that performances are not affected by the new change of paradigm with ROPE. Fixed the ROPE computation (should always be in float3
We mostly made sure that performances are not affected by the new change of paradigm with ROPE. Fixed the ROPE computation (should always be in float32) and the causal_mask dtype was set to bool to take less RAM.
YOLOS had a regression, and Llama / T5Tokenizer had a warning popping for random reasons
[Gemma] Fix eager attention #29187 by @sanchit-gandhi
TLDR:
- attn_output = attn_output.reshape(bsz, q_len, self.hidden_size)
+ attn_output = attn_output.view(bsz, q_len, -1)
Remove deprecated eager_serving fn by @Rocketknight1 in #28665
Gemma is a new opensource Language Model series from Google AI that comes with a 2B and 7B variant. The release comes with the pre-trained and instruction fine-tuned versions and you can use them via AutoModelForCausalLM, GemmaForCausalLM or pipeline interface!
Read more about it in the Gemma release blogpost: https://hf.co/blog/gemma
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2b")
model = AutoModelForCausalLM.from_pretrained("google/gemma-2b", device_map="auto", torch_dtype=torch.float16)
input_text = "Write me a poem about Machine Learning."
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
You can use the model with Flash Attention, SDPA, Static cache and quantization API for further optimizations !
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2b")
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2b", device_map="auto", torch_dtype=torch.float16, attn_implementation="flash_attention_2"
)
input_text = "Write me a poem about Machine Learning."
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2b")
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2b", device_map="auto", load_in_4bit=True
)
input_text = "Write me a poem about Machine Learning."
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2b")
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2b", device_map="auto"
)
model.generation_config.cache_implementation = "static"
input_text = "Write me a poem about Machine Learning."
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids)
The Depth Anything model was proposed in Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data by Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao. Depth Anything is based on the DPT architecture, trained on ~62 million images, obtaining state-of-the-art results for both relative and absolute depth estimation.
StableLM 3B 4E1T was proposed in StableLM 3B 4E1T: Technical Report by Stability AI and is the first model in a series of multi-epoch pre-trained language models.
StableLM 3B 4E1T is a decoder-only base language model pre-trained on 1 trillion tokens of diverse English and code datasets for four epochs. The model architecture is transformer-based with partial Rotary Position Embeddings, SwiGLU activation, LayerNorm, etc.
The team also provides StableLM Zephyr 3B, an instruction fine-tuned version of the model that can be used for chat-based applications.
StableLM by @jon-tow in #28810Static past key value cache allows LlamaForCausalLM' s forward pass to be compiled using torch.compile !
This means that (cuda) graphs can be used for inference, which speeds up the decoding step by 4x!
A forward pass of Llama2 7B takes around 10.5 ms to run with this on an A100! Equivalent to TGI performances! ⚡️
Core generation] Adds support for static KV cache by @ArthurZucker in #27931CLeanup] Revert SDPA attention changes that got in the static kv cache PR by @ArthurZucker in #29027⚠️ Support for generate is not included yet. This feature is experimental and subject to changes in subsequent releases.
from transformers import AutoTokenizer, AutoModelForCausalLM, StaticCache
import torch
import os
# compilation triggers multiprocessing
os.environ["TOKENIZERS_PARALLELISM"] = "true"
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-hf")
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf",
device_map="auto",
torch_dtype=torch.float16
)
# set up the static cache in advance of using the model
model._setup_cache(StaticCache, max_batch_size=1, max_cache_len=128)
# trigger compilation!
compiled_model = torch.compile(model, mode="reduce-overhead", fullgraph=True)
# run the model as usual
input_text = "A few facts about the universe: "
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda").input_ids
model_outputs = compiled_model(input_ids)
HfQuantizer makes it easy for quantization method researchers and developers to add inference and / or quantization support in 🤗 transformers. If you are interested in adding the support for new methods, please refer to this documentation page: https://huggingface.co/docs/transformers/main/en/hf_quantizer
HfQuantizer class for quantization-related stuff in modeling_utils.py by @poedator in #26610HfQuantizer] Move it to "Developper guides" by @younesbelkada in #28768HFQuantizer] Remove check_packages_compatibility logic by @younesbelkada in #28789AQLM is a new quantization method that enables no-performance degradation in 2-bit precision. Check out this demo about how to run Mixtral in 2-bit on a free-tier Google Colab instance: https://huggingface.co/posts/ybelkada/434200761252287
The canonical repositories on the hugging face hub (models that did not have an organization, like bert-base-cased), have been moved under organizations.
You can find the entire list of models moved here: https://huggingface.co/collections/julien-c/canonical-models-65ae66e29d5b422218567567
Redirection has been set up so that your code continues working even if you continue calling the previous paths. We, however, still encourage you to update your code to use the new links so that it is entirely future proof.
The Mistral model was added to the library in Flax.
With Keras 3 becoming the standard version of Keras in TensorFlow 2.16, we've made some internal changes to maintain compatibility. We now have full compatibility with TF 2.16 as long as the tf-keras compatibility package is installed. We've also taken the opportunity to do some cleanup - in particular, the objects like BatchEncoding that are returned by our tokenizers and processors can now be directly passed to Keras methods like model.fit(), which should simplify a lot of code and eliminate a long-standing source of annoyances.
Enable loading in pretrained backbones in a new model, where all other weights are randomly initialized. Note: validation checks are still in place when creating a config. Passing in use_pretrained_backbone will raise an error. You can override by setting
config.use_pretrained_backbone = True after creating a config. However, it is not yet guaranteed to be fully backwards compatible.
from transformers import MaskFormerConfig, MaskFormerModel
config = MaskFormerConfig(
use_pretrained_backbone=False,
backbone="microsoft/resnet-18"
)
config.use_pretrained_backbone = True
# Both models have resnet-18 backbone weights and all other weights randomly
# initialized
model_1 = MaskFormerModel(config)
model_2 = MaskFormerModel(config)
Introduce a helper function load_backbone to load a backbone from a backbone's model config e.g. ResNetConfig, or from a model config which contains backbone information. This enables cleaner modeling files and crossloading between timm and transformers backbones.
from transformers import ResNetConfig, MaskFormerConfig
from transformers.utils.backbone_utils import load_backbone
# Resnet defines the backbone model to load
config = ResNetConfig()
backbone = load_backbone(config)
# Maskformer config defines a model which uses a resnet backbone
config = MaskFormerConfig(use_timm_backbone=True, backbone="resnet18")
backbone = load_backbone(config)
config = MaskFormerConfig(backbone_config=ResNetConfig())
backbone = load_backbone(config)
Backbone] Use load_backbone instead of AutoBackbone.from_config by @amyeroberts in #28661Add in API references, list supported backbones, updated examples, clarification and moving information to better reflect usage and docs
Llava] Update convert_llava_weights_to_hf.py script by @isaac-vidas in #28617GPTNeoX] Fix GPTNeoX + Flash Attention 2 issue by @younesbelkada in #28645SigLIP] Only import tokenizer if sentencepiece available by @amyeroberts in #28636PartialState().default_device as it has been officially released by @statelesshz in #27256tensor_size - fix copy/paste error msg typo by @scruel in #28660CodeGenTokenizer by @cmathw in #28628GenerationConfig, now the generation_config.json can be loaded successfully by @ParadoxZW in #28604chore] Add missing space in warning by @tomaarsen in #28695Vilt] align input and model dtype in the ViltPatchEmbeddings forward pass by @faaany in #28633docs] Improve visualization for vertical parallelism by @petergtz in #28583LocalEntryNotFoundError during processor_config.json loading by @ydshieh in #28709docs] Update preprocessing.md by @velaia in #28719weights_only by @ydshieh in #28725GatedRepoError to use cache file (fix #28558). by @scruel in #28566Siglip] protect from imports if sentencepiece not installed by @amyeroberts in #28737DepthEstimationPipeline's docstring by @ydshieh in #28733Block. by @xkszltl in #28727load_in_8bit and load_in_4bit at the same time by @osanseviero in #28266bnb] Fix bnb slow tests by @younesbelkada in #28788torch.arange dtype on float usage to avoid incorrect initialization by @gante in #28760is_torch_bf16_available_on_device more strict by @ydshieh in #28796-v for pytest on CircleCI by @ydshieh in #28840test_encoder_decoder_model_generate for vision_encoder_deocder as flaky by @amyeroberts in #28842Doc] update contribution guidelines by @ArthurZucker in #28858save_only_model with load_best_model_at_end for DeepSpeed/FSDP by @pacman100 in #28866FastSpeech2ConformerModelTest and skip it on CPU by @ydshieh in #28888torchaudio get the correct version in torch_and_flax_job by @ydshieh in #28899logging_first_step by removing "evaluate" by @Sai-Suraj-27 in #28884Exception when trying to generate 0 tokens ⚠️ by @danielkorat in #28621torch_dtype as str to actual torch data type (i.e. "float16" …to torch.float16) by @KossaiSbai in #28208pipelines] updated docstring with vqa alias by @cmahmut in #28951test_save_load_fast_init_from_base as flaky by @gante in #28930NllbTokenizer] refactor with added tokens decoder by @ArthurZucker in #27717DETR] Update the processing to adapt masks & bboxes to reflect padding by @amyeroberts in #28363quantization_config is in config but not passed as an arg by @younesbelkada in #28988AutoQuantizer]: enhance trainer + not supported quant methods by @younesbelkada in #28991Doc] Fix docbuilder - make BackboneMixin and BackboneConfigMixin importable from utils. by @amyeroberts in #29002test_trainer to float32 by @statelesshz in #28920Trainer / tags]: Fix trainer + tags when users do not pass "tags" to trainer.push_to_hub() by @younesbelkada in #29009logger.warning + inline with recent refactor by @younesbelkada in #29039test_save_load_low_cpu_mem_usage tests by @amyeroberts in #29043generation/utils.py::GenerateEncoderDecoderOutput's docstring by @sadra-barikbin in #29044auto_find_batch_size isn't yet supported with DeepSpeed/FSDP. Raise error accrodingly. by @pacman100 in #29058Awq] Add peft support for AWQ by @younesbelkada in #28987bnb / tests]: Fix currently failing bnb tests by @younesbelkada in #29092bert-base-cased tokenizer configuration test by @LysandreJik in #29105examples/pytorch/text-classification/run_classification.py by @Ja1Zhou in #29072pipelines/base.py::Pipeline::_sanitize_parameters()'s docstring by @sadra-barikbin in #29102gradient_checkpointing] default to use it for torch 2.3 by @ArthurZucker in #28538Trainer / bnb]: Add RMSProp from bitsandbytes to HF Trainer by @younesbelkada in #29082bnb / tests] Propagate the changes from #29092 to 4-bit tests by @younesbelkada in #29122cuda kernels] only compile them when initializing by @ArthurZucker in #29133PEFT / Trainer ] Handle better peft + quantized compiled models by @younesbelkada in #29055Core tokenization] add_dummy_prefix_space option to help with latest issues by @ArthurZucker in #28010pipeline] Add pool option to image feature extraction pipeline by @amyeroberts in #28985The following contributors have made significant changes to the library over the last release:
HfQuantizer class for quantization-related stuff in modeling_utils.py (#26610)StableLM (#28810)Protecting the imports for SigLIP's tokenizer if sentencepiece isn't installed
Selection of fixes
torch.loadCommits
A patch release to resolve import errors from removed custom types in generation utils
A patch release to resolve import errors from removed custom types in generation utils
[Flax BERT] Update deprecated 'split' method by @sanchit-gandhi in #28012
Qwen2 is the new model series of large language models from the Qwen team. Previously, the Qwen series was released, including Qwen-72B, Qwen-1.8B, Qwen-VL, Qwen-Audio, etc.
Qwen2 is a language model series including decoder language models of different model sizes. For each size, we release the base language model and the aligned chat model. It is based on the Transformer architecture with SwiGLU activation, attention QKV bias, group query attention, mixture of sliding window attention and full attention, etc. Additionally, we have an improved tokenizer adaptive to multiple natural languages and codes.
Phi-2 is a transformer language model trained by Microsoft with exceptionally strong performance for its small size of 2.7 billion parameters. It was previously available as a custom code model, but has now been fully integrated into transformers.
phi-2 example by @susnato in #28392softmax_scale in PhiFlashAttention2. by @gugarosa in #28537The SigLIP model was proposed in Sigmoid Loss for Language Image Pre-Training by Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer. SigLIP proposes to replace the loss function used in CLIP by a simple pairwise sigmoid loss. This results in better performance in terms of zero-shot classification accuracy on ImageNet.
The VipLlava model was proposed in Making Large Multimodal Models Understand Arbitrary Visual Prompts by Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee.
VipLlava enhances the training protocol of Llava by marking images and interact with the model using natural cues like a “red bounding box” or “pointed arrow” during training.
The FastSpeech2Conformer model was proposed with the paper Recent Developments On Espnet Toolkit Boosted By Conformer by Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang.
FastSpeech 2 is a non-autoregressive model for text-to-speech (TTS) synthesis, which develops upon FastSpeech, showing improvements in training speed, inference speed and voice quality. It consists of a variance adapter; duration, energy and pitch predictor and waveform and mel-spectrogram decoder.
The Wav2Vec2-BERT model was proposed in Seamless: Multilingual Expressive and Streaming Speech Translation by the Seamless Communication team from Meta AI.
This model was pre-trained on 4.5M hours of unlabeled audio data covering more than 143 languages. It requires finetuning to be used for downstream tasks such as Automatic Speech Recognition (ASR), or Audio Classification.
Enables saving and loading transformers models in 4bit formats - you can now push bitsandbytes 4-bit weights on Hugging Face Hub. To save 4-bit models and push them on the hub, simply install the latest bitsandbytes package from pypi pip install -U bitsandbytes, load your model in 4-bit precision and call save_pretrained / push_to_hub. An example repo here
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "facebook/opt-125m"
model = AutoModelForCausalLM.from_pretrained(model_id, load_in_4bit=True)
model.push_to_hub("ybelkada/opt-125m-bnb-4bit")
Docs] Add 4-bit serialization docs by @younesbelkada in #28182Enable passing in 4D attention masks to models that support it. This is useful for reducing memory footprint of certain generation tasks.
attention_mask support by @poedator in #27539Ability to customise which modules are quantized and which are not.
Awq] Enable the possibility to skip quantization for some target modules by @younesbelkada in #27950modules_in_block_to_quantize arg in GPTQconfig by @SunMarc in #27956Added fused modules support
Awq] Add llava fused modules support by @younesbelkada in #28239Mixtral / Awq] Add mixtral fused modules for Awq by @younesbelkada in #28240Llava / Vip-Llava] Add SDPA into llava by @younesbelkada in #28107Mixtral & Mistral] Add support for sdpa by @ArthurZucker in #28133All decoding strategies (temperature fallback, compression/log-prob/no-speech threshold, ...) of OpenAI's long-form transcription (see: https://github.com/openai/whisper or section 4.5 in paper) have been added. Contrary to https://github.com/openai/whisper, Transformers long-form transcription is fully compatible with pure FP16 and Batching!
For more information see: https://github.com/huggingface/transformers/pull/27658.
Assisted generation was reworked to accept arbitrary sources of candidate sequences. This enabled us to smoothly integrate ngram speculation, and opens the door for new candidate generation methods. Additionally, we've added the speculative decoding strategy on top of assisted generation: when you call assisted generation with an assistant model and do_sample=True, you'll benefit from the faster speculative decoding sampling 🏎️💨
assisted_decoding now accepts arbitrary candidate generators by @gante in #27751generate for the assistant by @gante in #28031Adding pickle protection via weights_only=True in the torch.load calls.
Unlike PyTorch, TensorFlow models build their weights "lazily" after model initialization, using the shape of their inputs to figure out what their weight shapes should be. We previously needed a full forward pass through TF models to ensure that all layers received an input they could use to build their weights, but with this change we now have proper build() methods that can correctly infer shapes and build model weights. This avoids a whole range of potential issues, as well as significantly accelerating model load times.
The last version to support PyTorch 1.10 was 4.36.x. As it has been more than 2 years, and we're looking forward to using features available in PyTorch 1.11 and up, we do not support PyTorch 1.10 for v4.37 (i.e. we don't run the tests against torch 1.10).
You can now add custom tags into your model before pushing it on the Hub! This enables you to filter models that contain that tag on the Hub with a simple URL filter. For example if you want to filter models that have trl tag you can search: https://huggingface.co/models?other=trl&sort=created
core/ FEAT] Add the possibility to push custom tags using PreTrainedModel itself by @younesbelkada in #28405 - e.g.from transformers import AutoModelForCausalLM
model_name = "HuggingFaceM4/tiny-random-LlamaForCausalLM"
model = AutoModelForCausalLM.from_pretrained(model_name)
model.add_model_tags(["tag-test"])
model.push_to_hub("llama-tagged")
Mixtral] Change mistral op order by @younesbelkada in #27955Tokenizer Serialization] Fix the broken serialisation by @ArthurZucker in #27099Whisper] raise better errors by @ArthurZucker in #27971CI slow] Fix expected values by @ArthurZucker in #27999SeamlessM4TTokenizer] Safe import by @ArthurZucker in #28026core / modeling] Fix training bug with PEFT + GC by @younesbelkada in #28031test_retain_grad_hidden_states_attentions is flaky by @gante in #28035FA-2] Fix fa-2 issue when passing config to from_pretrained by @younesbelkada in #28043Modeling / Mixtral] Fix GC + PEFT issues with Mixtral by @younesbelkada in #28061Mixtral] update conversion script to reflect new changes by @younesbelkada in #28068test_retain_grad_hidden_states_attentions by @ylacombe in #28060low_cpu_mem_usage Flag Conflict with DeepSpeed Zero 3 in from_pretrained for Models with keep_in_fp32_modules" by @kotarotanahashi in #27762DISABLE_TELEMETRY is used by @Wauplin in #28113Mixtral] Fix loss + nits by @ArthurZucker in #28115CLIPConfig by @ydshieh in #28108input_embeds docstring in encoder-decoder architectures by @gante in #28168docs/source/en/perf_infer_gpu_one.md by @ydshieh in #28198training_args.py fix missing import with accelerate with version accelerate==0.20.1 by @michaelfeil in #28171feature_extractor_type when loading an image processor file by @ydshieh in #28195Llava] Fix llava index errors by @younesbelkada in #28032from_pretrained under ZeRO-3 by @XuehaiPan in #28245_merge_input_ids_with_image_features for llava model by @VictorSanh in #28333DeepSpeed when using auto find batch size by @muellerzr in #28088cache_dir for evaluate.load() in example scripts by @aphedges in #28422TFTrainer by @gante in #28483chore] Update warning text, a word was missing by @tomaarsen in #28017finetuned_from if it is a local path by @ydshieh in #28482task arg in load_dataset in image-classification example by @regisss in #28408TokenizationUtils] Fix add_special_tokens when the token is already there by @ArthurZucker in #28520TokenizationRoformerFast] Fix the save and loading by @ArthurZucker in #28527SpeechT5Tokenization] Add copied from and fix the convert_tokens_to_string to match the fast decoding scheme by @ArthurZucker in #28522Processor by @ydshieh in #27761weights_only only if torch >= 1.13 by @ydshieh in #28506Core Tokenization] Support a fix for spm fast models by @ArthurZucker in #26678LoggingLevel context manager in 3 tests by @ydshieh in #28575processor_config.json if a processor has no extra attribute by @ydshieh in #28584The following contributors have made significant changes to the library over the last release:
attention_mask support (#27539)Patch release to resolve some critical issues relating to the recent cache refactor, flash attention refactor and training in the multi-gpu and multi-
Patch release to resolve some critical issues relating to the recent cache refactor, flash attention refactor and training in the multi-gpu and multi-node settings:
config to from_pretrained with FA #28043A patch release for critical torch issues mostly:
A patch release for critical torch issues mostly:
🔥
Update deprecated torch.range in test_modeling_ibert.py by @kit1980 in #27355
Mixtral is the new open-source model from Mistral AI announced by the blogpost Mixtral of Experts. The model has been proven to have comparable capabilities to Chat-GPT according to the benchmark results shared on the release blogpost.
<img src="https://github.com/huggingface/transformers/assets/49240599/f6ee43a9-0f74-4473-a02d-f8b3abc6d614" width="500">
The architecture is a sparse Mixture of Experts with Top-2 routing strategy, similar as NllbMoe architecture in transformers. You can use it through AutoModelForCausalLM interface:
>>> import torch
>>> from transformers import AutoModelForCausalLM, AutoTokenizer
>>> model = AutoModelForCausalLM.from_pretrained("mistralai/Mixtral-8x7B", torch_dtype=torch.float16, device_map="auto")
>>> tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-8x7B")
>>> prompt = "My favourite condiment is"
>>> model_inputs = tokenizer([prompt], return_tensors="pt").to(device)
>>> model.to(device)
>>> generated_ids = model.generate(**model_inputs, max_new_tokens=100, do_sample=True)
>>> tokenizer.batch_decode(generated_ids)[0]
The model is compatible with existing optimisation tools such Flash Attention 2, bitsandbytes and PEFT library. The checkpoints are release under mistralai organisation on the Hugging Face Hub.
Llava is an open-source chatbot trained by fine-tuning LlamA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model, based on the transformer architecture. In other words, it is an multi-modal version of LLMs fine-tuned for chat / instructions.
<img src="https://cdn-uploads.huggingface.co/production/uploads/62441d1d9fdefb55a0b7d12c/FPshq08TKYD0e-qwPLDVO.png" width="800">
The Llava model was proposed in Improved Baselines with Visual Instruction Tuning by Haotian Liu, Chunyuan Li, Yuheng Li and Yong Jae Lee.
Llava] Add Llava to transformers by @younesbelkada in #27662The integration also includes BakLlava which is a Llava model trained with Mistral backbone.
The mode is compatible with "image-to-text" pipeline:
from transformers import pipeline
from PIL import Image
import requests
model_id = "llava-hf/llava-1.5-7b-hf"
pipe = pipeline("image-to-text", model=model_id)
url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/ai2d-demo.jpg"
image = Image.open(requests.get(url, stream=True).raw)
prompt = "USER: <image>\nWhat does the label 15 represent? (1) lava (2) core (3) tunnel (4) ash cloud\nASSISTANT:"
outputs = pipe(image, prompt=prompt, generate_kwargs={"max_new_tokens": 200})
print(outputs)
And you can find all Llava weights under llava-hf organisation on the Hub.
SeamlessM4T-v2 is a collection of models designed to provide high quality translation, allowing people from different linguistic communities to communicate effortlessly through speech and text. It is an improvement on the previous version and was proposed in Seamless: Multilingual Expressive and Streaming Speech Translation by the Seamless Communication team from Meta AI.
For more details on the differences between v1 and v2, refer to section Difference with SeamlessM4T-v1.
SeamlessM4T enables multiple tasks without relying on separate models:
The PatchTST model was proposed in A Time Series is Worth 64 Words: Long-term Forecasting with Transformers by Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong and Jayant Kalagnanam.
At a high level, the model vectorizes time series into patches of a given size and encodes the resulting sequence of vectors via a Transformer that then outputs the prediction length forecast via an appropriate head. The model is illustrated in the following figure:
The PatchTSMixer model was proposed in TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting by Vijay Ekambaram, Arindam Jati, Nam Nguyen, Phanwadee Sinthong and Jayant Kalagnanam.
PatchTSMixer is a lightweight time-series modeling approach based on the MLP-Mixer architecture. In this HuggingFace implementation, we provide PatchTSMixer’s capabilities to effortlessly facilitate lightweight mixing across patches, channels, and hidden features for effective multivariate time-series modeling. It also supports various attention mechanisms starting from simple gated attention to more complex self-attention blocks that can be customized accordingly. The model can be pretrained and subsequently used for various downstream tasks such as forecasting, classification and regression.
The CLVP (Contrastive Language-Voice Pretrained Transformer) model was proposed in Better speech synthesis through scaling by James Betker.
The Phi-1 model was proposed in Textbooks Are All You Need by Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee and Yuanzhi Li.
The Phi-1.5 model was proposed in Textbooks Are All You Need II: phi-1.5 technical report by Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar and Yin Tat Lee.
The text-visual prompting (TVP) framework was proposed in the paper Text-Visual Prompting for Efficient 2D Temporal Video Grounding by Yimeng Zhang, Xin Chen, Jinghan Jia, Sijia Liu, Ke Ding.
This research addresses temporal video grounding (TVG), which is the process of pinpointing the start and end times of specific events in a long video, as described by a text sentence. Text-visual prompting (TVP), is proposed to enhance TVG. TVP involves integrating specially designed patterns, known as ‘prompts’, into both the visual (image-based) and textual (word-based) input components of a TVG model. These prompts provide additional spatial-temporal context, improving the model’s ability to accurately determine event timings in the video. The approach employs 2D visual inputs in place of 3D ones. Although 3D inputs offer more spatial-temporal detail, they are also more time-consuming to process. The use of 2D inputs with the prompting method aims to provide similar levels of context and accuracy more efficiently.
Depth estimation is added to the DINO v2 implementation.
AMD's ROCm GPU architecture is now supported across the board and fully tested in our CI with MI210/MI250 GPUs. We further enable specific hardware acceleration for ROCm in Transformers, such as Flash Attention 2, GPTQ quantization and DeepSpeed.
scaled_dot_product_attention native supportPyTorch's torch.nn.functional.scaled_dot_product_attention operator is now supported in the most-used Transformers models and used by default when using torch>=2.1.1, allowing to dispatch on memory-efficient attention and Flash Attention backend implementations with no other package than torch required. This should significantly speed up attention computation on hardware that that supports these fastpath.
While Transformers automatically handles the dispatch to use SDPA when available, it is possible to force the usage of a given attention implementation ("eager" being the manual implementation, where each operation is implemented step by step):
# or `attn_implementation="sdpa", or `attn_implementation="flash_attention_2"`
model = AutoModelForSpeechSeq2Seq.from_pretrained("openai/whisper-tiny", attn_implementation="eager")
Training benchmark, run on A100-SXM4-80GB.
| Model | Batch size | Sequence length | Time per batch ("eager", s) |
Time per batch ("sdpa", s) |
Speedup | Peak memory ("eager", MB) |
Peak memory ("sdpa", MB) |
Memory savings |
|---|---|---|---|---|---|---|---|---|
| llama2 7b | 4 | 1024 | 1.065 | 0.90 | 19.4% | 73878.28 | 45977.81 | 60.7% |
| llama2 7b | 4 | 2048 | OOM | 1.87 | / | OOM | 78394.58 | SDPA does not OOM |
| llama2 7b | 1 | 2048 | 0.64 | 0.48 | 32.0% | 55557.01 | 29795.63 | 86.4% |
| llama2 7b | 1 | 3072 | OOM | 0.75 | / | OOM | 37916.08 | SDPA does not OOM |
| llama2 7b | 1 | 4096 | OOM | 1.03 | / | OOM | 46028.14 | SDPA does not OOM |
| llama2 7b | 2 | 4096 | OOM | 2.05 | / | OOM | 78428.14 | SDPA does not OOM |
Inference benchmark, run on A100-SXM4-80GB.
| Model | Batch size | Prompt length | Num new tokens | Per token latency "eager" (ms) |
Per token latency "sdpa" (ms) |
Speedup |
|---|---|---|---|---|---|---|
| llama2 13b | 1 | 1024 | 1 (prefill) | 178.66 | 159.36 | 12.11% |
| llama2 13b | 1 | 100 | 100 | 40.35 | 37.62 | 7.28% |
| llama2 13b | 8 | 100 | 100 | 40.55 | 38.06 | 6.53% |
| Whisper v3 large | 1 | / | 62 | 20.05 | 18.90 | 6.10% |
| Whisper v3 large | 8 | / | 77 | 25.42 | 24.77 | 2.59% |
| Whisper v3 large | 16 | / | 77 | 28.51 | 26.32 | 8.34% |
We are rolling out a new abstraction for the past_key_values cache, which enables the use of different types of caches. For now, only llama and llama-inspired architectures (mistral, persimmon, phi) support it, with other architectures scheduled to have support in the next release. By default, a growing cache (DynamicCache) is used, which preserves the existing behavior.
This release also includes a new SinkCache cache, which implements the Attention Sinks paper. With SinkCache, the model is able to continue generating high-quality text well beyond its training sequence length! Note that it does not expand the context window, so it can’t digest very long inputs — it is suited for streaming applications such as multi-round dialogues. Check this colab for an example.
Cache abstraction and Attention Sinks support by @tomaarsen in #26681We continue toggling features enabling safetensors as a default across the board, in PyTorch, Flax, and TensorFlow.
When using PyTorch model and forcing the load of safetensors file with use_safetensors=True, if the repository does not contain the safetensors files, they will now be converted on-the-fly server-side.
from_pt flag when loading with safetensors by @LysandreJik in #27394We now disallow the use of pickle.load internally for security purposes. To circumvent this, you can use the TRUST_REMOTE_CODE=True command to indicate that you would still like to load it.
pickle.load unless TRUST_REMOTE_CODE=True by @ydshieh in #27776In the previous implementation of beam search, when length_penalty is active, the beam score for decoder-only models was penalized by the total length of both prompt and generated sequence. However, the length of prompt should not be included in the penalization step -- this release fixes it.
AttentionMaskConverter compatible with torch.compile(..., fullgraph=True) by @fxmarty in #27868PEFT / Tests ] Fix peft integration failing tests by @younesbelkada in #27258Docs / SAM ] Reflect correct changes to run inference without OOM by @younesbelkada in #27268FA2] Add flash attention for for DistilBert by @susnato in #26489PretrainedTokenizer] add some of the most important functions to the doc by @ArthurZucker in #27313Kosmos2Processor batch mode by @ydshieh in #27323FA2] Add flash attention for GPT-Neo by @susnato in #26486Whisper] Add conversion script for the tokenizer by @ArthurZucker in #27338gpt_bigcode. by @susnato in #27348Whisper] Nit converting the tokenizer by @ArthurZucker in #27349Kosmos-2 device issue by @ydshieh in #27346from_pt=True by @ydshieh in #27372torch.range in test_modeling_ibert.py by @kit1980 in #27355CodeLlamaTokenizer] Nit, update init to make sure the AddedTokens are not normalized because they are special by @ArthurZucker in #27359pyproject.toml by @ydshieh in #27366pytest.mark directly by @ydshieh in #27390FuyuConfig by @ydshieh in #27399Owlv2 checkpoint name and a default value in Owlv2VisionConfig by @ydshieh in #27402circleci/create_circleci_config.py is modified by @ydshieh in #27413Quantization] Add str to enum conversion for AWQ by @younesbelkada in #27320AttentionMaskConverter] ]Fix-mask-inf by @ArthurZucker in #27114examples_torch_job faster by @ydshieh in #27437utils/not_doctested.txt by @ydshieh in #27459Llama + Mistral] Add attention dropout by @ArthurZucker in #27315gradient_checkpointing_kwargs by @tomaszcichy98 in #27470python-Levenshtein for nougat in CI image by @ydshieh in #27465AWQ ] Addresses TODO for awq tests by @younesbelkada in #27467Peft] modules_to_save support for peft integration by @younesbelkada in #27466CI-test_torch] skip test_tf_from_pt_safetensors for 4 models by @ArthurZucker in #27481ExponentialDecayLengthPenalty doctest by @gante in #27485GenerationConfig.from_pretrained can return unused kwargs by @gante in #27488CI-test_torch] skip test_tf_from_pt_safetensors and test_assisted_decoding_sample by @ArthurZucker in #27508CircleCI] skip test_assisted_decoding_sample for everyone by @ArthurZucker in #27511tokenizers] update tokenizers version pin by @ArthurZucker in #27494PretrainedConfig] Improve messaging by @ArthurZucker in #27438en/model_doc docs to Japanese. by @Yuki-Imajuku in #27401pytest] Avoid flash attn test marker warning by @ArthurZucker in #27509usedforsecurity=False in hashlib methods (FIPS compliance) by @Wauplin in #27483latest-pytorch-amd for now by @ydshieh in #27541Styling] stylify using ruff by @ArthurZucker in #27144convert_hf_to_openai.py script to Whisper documentation resources by @zuazo in #27590FA-2] Add fa2 support for from_config by @younesbelkada in #26914large-v3 version support by @flyingleafe in #27336core / gradient_checkpointing] add support for old GC method by @younesbelkada in #27610past_key_values in generate by @gante in #27612init_git_repo by @statelesshz in #27617use_cache=True in Flash Attention tests by @fxmarty in #27635resize_token_embeddings by @czy-orange in #26861)dependency] update pillow pins by @ArthurZucker in #27409max_steps documentation regarding the end-of-training condition by @qgallouedec in #27624FA2] Add flash attention for opt by @susnato in #26414save_only_model arg and simplifying FSDP integration by @pacman100 in #27652TransfoXL by @ydshieh in #27607DocString] Support a revision in the docstring add_code_sample_docstrings to facilitate integrations by @ArthurZucker in #27645TVPModelTest by @ydshieh in #27695en/model_doc to JP by @rajveer43 in #27264TransfoXLTokenizer.__init__ by @ydshieh in #27721tests/utils/tiny_model_summary.json is modified by @ydshieh in #27693~transformer. -> ~transformers. by @tomaarsen in #27740check_runner_status.yml by @ydshieh in #27767GenerationConfig throws an exception when generate args are passed by @gante in #27757persistent_workers parameter to TrainingArguments by @Sorrow321 in #27189ModelOnTheFlyConversionTester] Mark as slow for now by @ArthurZucker in #27823TvpModelIntegrationTests by @ydshieh in #27792Owlv2ModelIntegrationTest::test_inference_object_detection by @ydshieh in #27793en/tasks folder docs to Japanese 🇯🇵 by @rajveer43 in #27098ruff==0.1.5 by @ydshieh in #27849ClipVision] accelerate support for clip-vision by @younesbelkada in #27851VitDetModelTester.get_config to use pretrain_image_size by @ydshieh in #27831Docs] Update broken image on fused modules by @younesbelkada in #27856_keep_in_fp32_modules being modified by @ydshieh in #27867Flash Attention 2] Add flash attention 2 for GPT-Neo-X by @younesbelkada in #26463blip to clap) 🇯🇵 by @rajveer43 in #27673FA-2] Add Flash Attention to Phi by @susnato in #27661# Ignore copy by @ydshieh in #27328create_model_card to properly save peft details when using Trainer with PEFT by @pacman100 in #27754get_default_device to v4.38 by @statelesshz in #27848model_doc files from clip to cpm to JP by @rajveer43 in #27774prefix_allowed_tokens_fn return empty set of tokens by @Saibo-creator in #27797test_initialization as flaky in 2 model tests by @ydshieh in #27906notification_service.py by @ydshieh in #27903FillMaskPipelineTests by @ydshieh in #27889resume_from_checkpoint to handle auto_find_batch_size by @muellerzr in #27568The following contributors have made significant changes to the library over the last release:
FA2] Add flash attention for for DistilBert (#26489)FA2] Add flash attention for GPT-Neo (#26486)gpt_bigcode. (#27348)FA2] Add flash attention for opt (#26414)FA-2] Add Flash Attention to Phi (#27661)en/model_doc docs to Japanese. (#27401)en/model_doc to JP (#27264)en/tasks folder docs to Japanese 🇯🇵 (#27098)blip to clap) 🇯🇵 (#27673)model_doc files from clip to cpm to JP (#27774)A patch release was made for the following commit:
A patch release was made for the following commit:
tokenizers] update tokenizers version pin #27494to fix all the issues with versioning regarding tokenizers and huggingface_hub
Fix FA2 import + deprecation cycle
A patch release was made for the following three commits:
remove SharedDDP as it is deprecated by @statelesshz in #25702
Distil-Whisper is a distilled version of Whisper that is 6 times faster, 49% smaller, and performs within 1% word error rate (WER) on out-of-distribution data. It was proposed in the paper Robust Knowledge Distillation via Large-Scale Pseudo Labelling.
Distil-Whisper copies the entire encoder from Whisper, meaning it retains Whisper's robustness to different audio conditions. It only copies 2 decoder layers, which significantly reduces the time taken to auto-regressively generate text tokens:
<img src="https://huggingface.co/datasets/distil-whisper/figures/resolve/main/architecture.png" width="800">
Distil-Whisper is MIT licensed and directly available in the Transformers library with chunked long-form inference, Flash Attention 2 support, and Speculative Decoding. For details on using the model, refer to the following instructions.
Joint work from @sanchit-gandhi, @patrickvonplaten and @srush.
The Fuyu model was created by ADEPT, and authored by Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, Sağnak Taşırlar.
The authors introduced Fuyu-8B, a decoder-only multimodal model based on the classic transformers architecture, with query and key normalization. A linear encoder is added to create multimodal embeddings from image inputs.
By treating image tokens like text tokens and using a special image-newline character, the model knows when an image line ends. Image positional embeddings are removed. This avoids the need for different training phases for various image resolutions. With 8 billion parameters and licensed under CC-BY-NC, Fuyu-8B is notable for its ability to handle both text and images, its impressive context size of 16K, and its overall performance.
Joint work from @molbap, @pcuenca, @amyeroberts, @ArthurZucker
The SeamlessM4T model was proposed in SeamlessM4T — Massively Multilingual & Multimodal Machine Translation by the Seamless Communication team from Meta AI.
SeamlessM4T is a collection of models designed to provide high quality translation, allowing people from different linguistic communities to communicate effortlessly through speech and text.
SeamlessM4T enables multiple tasks without relying on separate models:
SeamlessM4TModel can perform all the above tasks, but each task also has its own dedicated sub-model.
The KOSMOS-2 model was proposed in Kosmos-2: Grounding Multimodal Large Language Models to the World by Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Furu Wei.
<div style="text-align: center"> <img src="https://huggingface.co/microsoft/kosmos-2-patch14-224/resolve/main/annotated_snowman.jpg" width="500" height="400" > </div>
KOSMOS-2 is a Transformer-based causal language model and is trained using the next-word prediction task on a web-scale dataset of grounded image-text pairs GRIT. The spatial coordinates of the bounding boxes in the dataset are converted to a sequence of location tokens, which are appended to their respective entity text spans (for example, a snowman followed by <patch_index_0044><patch_index_0863>). The data format is similar to “hyperlinks” that connect the object regions in an image to their text span in the corresponding caption.
Kosmos-2 model by @ydshieh in #24709OWLv2 was proposed in Scaling Open-Vocabulary Object Detection by Matthias Minderer, Alexey Gritsenko, Neil Houlsby. OWLv2 scales up OWL-ViT using self-training, which uses an existing detector to generate pseudo-box annotations on image-text pairs. This results in large gains over the previous state-of-the-art for zero-shot object detection.
torch serialization 🚨🚨🚨Version v4.35.0 now puts safetensors serialization by default. This is a significant change targeted at making users of the Hugging Face Hub, transformers, and any downstream library leveraging it safer.
The safetensors library is a safe serialization framework for machine learning tensors. It has been audited and will become the default serialization framework for several organizations (Hugging Face, EleutherAI, Stability AI).
It was already the default loading mechanism since v4.30.0 and would therefore already default to loading model.safetensors files instead of pytorch_model.bin if these were present in the repository.
With v4.35.0, any call to save_pretrained for torch models will now save a safetensors file. This safetensors file is in the PyTorch format, but can be loaded in TensorFlow and Flax models alike.
⚠️ If you run into any issues with this, please let us know ASAP in the issues so that we may help you. Namely, the following errors may indicate something is up:
safetensors file and having a warning mentioning missing weights unexpectedlysafetensorsIf you wish to continue saving files in the .bin format, you can do so by specifying safe_serialization=False in all your save_pretrained calls.
Chat templates have been expanded with the addition of the add_generation_prompt argument to apply_chat_template(). This has also enabled us to rework the ConversationalPipeline class to use chat templates. Any model with a chat template is now automatically usable through ConversationalPipeline.
Two new guides on LLMs were added the library:
Exllama-v2 provides better GPTQ kernel for higher throughput and lower latency for GPTQ models. The original code can be found here.
You will need the latest versions of optimum and auto-gptq. Read more about the integration here.
AWQ is a new and popular quantization scheme, already used in various libraries such as TGI, vllm, etc. and known to be faster than GPTQ models according to some benchmarks. The original code can be found here and here you can read more about the original paper.
We support AWQ inference with original kernels as well as kernels provided through autoawq package that you can simply install with pip install autoawq.
core / Quantization ] AWQ integration by @younesbelkada in #27045We also provide an example script on how to push quantized weights on the hub on the original repository.
Read more about the benchmarks and the integration here
You can now run GPTQ models on CPU using the latest version of auto-gptq thanks to @vivekkhandelwal1 !
We refactored the attention mask logic for major models in transformers. For instance, we removed padding_mask argument which was ambiguous for some users
padding_mask and instead use a 2D->4D Attn Mask Mapper by @patrickvonplaten in #26792Gpt-bigcode (starcoder), whisper, Bart and MBart now supports FA-2 ! Use it by simply passing use_flash_attention_2=True to from_pretrained. Some bugfixes with respect to mixed precision training with FA2 have been also addressed.
gpt_bigcode by @susnato in #26479FA2] Fix flash attention 2 fine-tuning with Falcon by @younesbelkada in #26852A bugfix with respect to fine-tuning with FA-2 in bfloat16 was addressed. You should now smoothly fine-tune FA-2 models in bfloat16 using quantized base models.
Quantization] Store the original dtype in the config as a private attribute 🚨🚨🚨 by @younesbelkada in #26761FA-2] Final fix for FA2 dtype by @younesbelkada in #26846NEFTune is a new technique to boost Supervised Fine-tuning performance by adding random noise on the embedding vector. Read more about it on the original paper here
We propose a very simple API for users to benefit from this technique, simply pass a valid neftune_noise_alpha parameter to TrainingArguments
Read more about the API here
We have refactored the gradient checkpointing API so that users can pass keyword arguments supported by torch.utils.checkpoint.checkpoint directly through gradient_checkpointing_kwargs when calling gradient_checkpointing_enable(), e.g.
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("facebook/opt-125m")
model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
gradient_checkpointing_kwargs is also supported with Trainer through TrainingArguments.
Trainer / GC] Add gradient_checkpointing_kwargs in trainer and training arguments by @younesbelkada in #27068core] Refactor of gradient_checkpointing by @younesbelkada in #27020core/ GC / tests] Stronger GC tests by @younesbelkada in #27124The refactor should be totally backward compatible with previous behaviour. For superusers, you can still use the attribute gradient_checkpointing on model's submodules to control the activation / deactivation of gradient_checkpointing.
Quantization] Store the original dtype in the config as a private attribute 🚨🚨🚨 by @younesbelkada in #26761Nougat] from transformers import * by @ArthurZucker in #26562semantic_segmentation.md to Korean by @jungnerd in #26515NougatProcessor] Fix the default channel by @ArthurZucker in #26608GPTNeoX] Faster rotary embedding for GPTNeoX (based on llama changes) by @ArthurZucker in #25830use_cache=False before creating presents which relies on use_cache by @yundai424 in #26328main due to torch 2.1 by @ydshieh in #26607ModelOutput serializable by @cbensimon in #26493core] fix silent bug keep_in_fp32 modules by @younesbelkada in #26589transformers-pytorch-gpu docker build by @ydshieh in #26615pytorch-quantization in Doc Builder docker file by @ydshieh in #26622views of position_ids by @ramiro050 in #26059MusicgenTest .test_pipeline_text_to_audio by @ydshieh in #26586LlamaTokenizerFast] Adds edge cases for the template processor by @ArthurZucker in #26606AlbertConfig by @ydshieh in #26636CLIPImageProcessor by @isaac-chung in #26676CLIP by @isaac-chung in #26691LlamaConfig by @pavaris-pm in #26685jnp.array in types with jnp.ndarray. by @hvaara in #26703Copied from for test files by @ydshieh in #26713SwinModel docstring fix by @shivanandmn in #26679use_cuda_amp is no more available by @pacman100 in #26731no_trainer scripts by @muellerzr in #26733torch==2.1.0 by @ydshieh in #26735LlamaTokenizer and LlamaTokenizerFast by @minhoryang in #26669CodeLlamaTokenizer by @Bojun-Feng in #26709Blip2ForConditionalGeneration by @ydshieh in #26737PersimmonIntegrationTest OOM by @ydshieh in #26750MistralIntegrationTest OOM by @ydshieh in #26754UniSpeech, UniSpeechSat, Wav2Vec2ForCTC by @gizemt in #26664GPT2 and Whisper by @McDonnellJoseph in #26642PerceiverModelIntegrationTest::test_inference_masked_lm by @ydshieh in #26760core] Fix fa-2 import by @younesbelkada in #26785TrainerIntegrationFSDP::test_basic_run_with_cpu_offload if torch < 2.1 by @ydshieh in #26764big_models.md to Korean by @wonhyeongseo in #26245IdeficsProcessorTest.test_tokenizer_padding by @ydshieh in #26779RwkvConfig by @Bojun-Feng in #26782DPRConfig by @AVAniketh0905 in #26674Flava] Fix flava doc by @younesbelkada in #26789CanineConfig by @Sparty in #26771CodeLlamaTokenizerFast by @Bojun-Feng in #26666en/internal folder docs to Japanese 🇯🇵 by @rajveer43 in #26747Tokenizer] Fix slow and fast serialization by @ArthurZucker in #26570FA-2] Revert suggestion that broke FA2 fine-tuning with quantized models by @younesbelkada in #26916ChineseCLIP by @Sparty in #26880CodeGen by @daniilgaltsev in #26821FA-2 / Mistral] Supprot fa-2 + right padding + forward by @younesbelkada in #26912max_shard_size to smaller value by @younesbelkada in #26942NLLB-MoE] Fix NLLB MoE 4bit inference by @younesbelkada in #27012SeamlessM4T] fix copies with NLLB MoE int8 by @ArthurZucker in #27018pipeline_tutorial.md to chinese by @jiaqiw09 in #26954preprocessing.md to Chinese by @jiaqiw09 in #26955default_to_square_for_size to CLIPImageProcessor by @ydshieh in #26965TFxxxxForSequenceClassifciation] Fix the eager mode after #25085 by @ArthurZucker in #25751docs] Add MaskGenerationPipeline in docs by @younesbelkada in #27063set_epoch for Accelerate-based dataloaders by @muellerzr in #26850flash_attn version to 2.1 by @younesbelkada in #27079T5Tokenizer] Fix fast and extra tokens by @ArthurZucker in #27085core/ gradient_checkpointing] Refactor GC - part 2 by @younesbelkada in #27073FA2/ Mistral] Revert previous behavior with right padding + forward by @younesbelkada in #27125"common_voice" by @ydshieh in #27147tests / Quantization] Fix bnb test by @younesbelkada in #27145copied from by @ydshieh in #27149en/main_classes folder docs to Japanese 🇯🇵 by @rajveer43 in #26894get_default_device in tools/base.py by @statelesshz in #26774tiny_model_summary.json is modified by @ydshieh in #27175Quantization / tests ] Fix bnb MPT test by @younesbelkada in #27178StarCoder by @susnato in #27182core / Quantization] Fix for 8bit serialization tests by @younesbelkada in #27234The following contributors have made significant changes to the library over the last release:
semantic_segmentation.md to Korean (#26515)get_default_device in tools/base.py (#26774)en/internal folder docs to Japanese 🇯🇵 (#26747)en/main_classes folder docs to Japanese 🇯🇵 (#26894)pipeline_tutorial.md to chinese (#26954)preprocessing.md to Chinese (#26955)A patch release was made for the following three commits:
A patch release was made for the following three commits:
These are not breaking changes per se but rather bugfixes. However, we understand that this may result in some workflow changes so we highlight them b…
Mistral-7B-v0.1 is a decoder-based LM with the following architectural choices:
The authors introduced Persimmon-8B, a decoder model based on the classic transformers architecture, with query and key normalization. Persimmon-8B is a fully permissively licensed model with approximately 8 billion parameters, released under the Apache license. Some of the key attributes of Persimmon-8B are long context size (16K), performance, and capabilities for multimodal extensions.
Persimmon] Add support for persimmon by @ArthurZucker in #26042BROS stands for BERT Relying On Spatiality. It is an encoder-only Transformer model that takes a sequence of tokens and their bounding boxes as inputs and outputs a sequence of hidden states. BROS encode relative spatial information instead of using absolute spatial information.
ViTMatte leverages plain Vision Transformers for the task of image matting, which is the process of accurately estimating the foreground object in images and videos.
Nougat uses the same architecture as Donut, meaning an image Transformer encoder and an autoregressive text Transformer decoder to translate scientific PDFs to markdown, enabling easier access to them.
We've added a new template feature for chat models. This allows the formatting that a chat model was trained with to be saved with the model, ensuring that users can exactly reproduce that formatting when they want to fine-tune the model or use it for inference. For more information, see our template documentation.
Tokenizer] attemp to fix add_token issues by @ArthurZucker in #23909🚨Workflow Changes 🚨:
These are not breaking changes per se but rather bugfixes. However, we understand that this may result in some workflow changes so we highlight them below.
➕ Most visible features:
tokenizer.added_tokens_decoder for both fast and slow tokenizers. Moreover, additional tokens that were already part of the initial vocab are also found there.from_pretrained, faster add_tokens because special and non special can be mixed together and the trie is not always rebuilt.added_tokens_decoder/encoder.tokenizer_config.jsonFor any issues relating to this, make sure to open a new issue and ping @ArthurZucker.
FA2 support added to transformers for most popular architectures (llama, mistral, falcon) architectures actively being contributed in this issue (https://github.com/huggingface/transformers/issues/26350). Simply pass use_flash_attention_2=True when calling from_pretrained
In the future, PyTorch will support Flash Attention 2 through torch.scaled_dot_product_attention, users would be able to benefit from both (transformers core & transformers + SDPA) implementations of Flash Attention-2 with simple changes (model.to_bettertransformer() and force-dispatch the SDPA kernel to FA-2 in the case of SDPA)
core ] Integrate Flash attention 2 in most used models by @younesbelkada in #25598For our future plans regarding integrating F.sdpa from PyTorch in core transformers, see here: https://github.com/huggingface/transformers/issues/26557
Support for lazy loading integration libraries has been added. This will drastically speed up importing transformers and related object from the library.
Example before this change:
2023-09-11 11:07:52.010179: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT
python3 -c "from transformers import CLIPTextModel" 3.31s user 3.06s system 220% cpu 2.893 total
After this change:
python3 -c "from transformers import CLIPTextModel" 1.70s user 1.49s system 220% cpu 1.447 total
test_load_img_url_timeout by @ydshieh in #25976Pop2Piano space demo. by @susnato in #25975generation_config by @gante in #25987CI] Fix red CI and ERROR failed should show by @ArthurZucker in #25995VITS] tokenizer integration test: fix revision did not exist by @ArthurZucker in #25996llm_tutorial.md to Korean by @harheem in #25791tgs speed metrics by @CokeDong in #25858activation_dropout and fix DropOut docs for SEW-D by @gau-nernst in #26031llama.md to Korean by @harheem in #26044CodeLlamaTokenizerFast] Fix fix set_infilling_processor to properly reset by @ArthurZucker in #26041CITests] skip failing tests until #26054 is merged by @ArthurZucker in #26063core] Import tensorflow inside relevant methods in trainer_utils by @younesbelkada in #26106generation_config is untouched by @gante in #25962llama2.md to Korean by @mjk0618 in #26047contributing.md to Korean by @mjk0618 in #25877MarianTokenizer to remove metaspace character in decode by @tanaymeh in #26091core] fix 4bit num_parameters by @younesbelkada in #26132RWKV] Final fix RWMV 4bit by @younesbelkada in #26134generation_config.max_length is set to None by @gante in #26147test_finetune_bert2bert by @ydshieh in #25984beam_scores shape when token scores shape changes after logits_processor by @BakerBunker in #25980accelerate > 0.20.3 by @sam-scale in #26060PEFT] Fix PEFT + gradient checkpointing by @younesbelkada in #25846convert_bros_to_pytorch.py by @ydshieh in #26212utils/documentation_tests.txt by @ydshieh in #26213ctrl to Salesforce/ctrl by @julien-c in #26183whisper.md to Korean by @nuatmochoi in #26002Error not captured in PR doctesting by @ydshieh in #26215Trainer] Refactor trainer + bnb logic by @younesbelkada in #26248ALL_LAYERNORM_LAYERS by @shijie-wu in #26227model._keep_in_fp32_modules is set even when accelerate is not installed by @fxmarty in #26225store_test_results by @ydshieh in #26223audio_classification.mdx to Korean by @gabrielwithappy in #26200RMSProp optimizer by @natolambert in #26425transformers is installed without tokenizers by @urialon in #26236FA / tests] Add use_cache tests for FA models by @younesbelkada in #26415PEFT] Fix PEFT multi adapters support by @younesbelkada in #26407runs-on in workflow files by @ydshieh in #26435debugging.md to Korean by @wonhyeongseo in #26246perf_train_gpu_many.md to Korean by @wonhyeongseo in #26244cos_sin device issue in Falcon model by @ydshieh in #26448PEFT] introducing adapter_kwargs for loading adapters from different Hub location (subfolder, revision) than the base model by @younesbelkada in #26270PEFT] Pass token when calling find_adapter_config by @younesbelkada in #26488core/ auto ] Fix bnb test with code revision + bug with code revision by @younesbelkada in #26431PEFT] Protect adapter_kwargs check by @younesbelkada in #26537tokenizer_summary.md to Korean by @wonhyeongseo in #26243configuration_encoder_decoder.py by @SrijanSahaySrivastava in #26519The following contributors have made significant changes to the library over the last release:
debugging.md to Korean (#26246)perf_train_gpu_many.md to Korean (#26244)tokenizer_summary.md to Korean (#26243)A patch release was made for the following three commits:
A patch release was made for the following three commits:
A patch release was done for these two commits:
A patch release was done for these two commits:
[VITS] Handle deprecated weight norm by @sanchit-gandhi in #25946
Falcon is a class of causal decoder-only models built by TII. The largest Falcon checkpoints have been trained on >=1T tokens of text, with a particular emphasis on the RefinedWeb corpus. They are made available under the Apache 2.0 license.
Falcon’s architecture is modern and optimized for inference, with multi-query attention and support for efficient attention variants like FlashAttention. Both ‘base’ models trained only as causal language models as well as ‘instruct’ models that have received further fine-tuning are available.
Falcon] Remove SDPA for falcon to support earlier versions of PyTorch (< 2.0) by @younesbelkada in #25947Code Llama, is a family of large language models for code based on Llama 2, providing state-of-the-art performance among open models, infilling capabilities, support for large input contexts, and zero-shot instruction following ability for programming tasks.
CodeLlama] Add support for CodeLlama by @ArthurZucker in #25740CodeLlama] Fix CI by @ArthurZucker in #25890ViTDet reuses the ViT model architecture, adapted to object detection.
DINO v2 is the next iteration of the DINO model. It is added as a backbone class, allowing it to be re-used in downstream models.
VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) is an end-to-end speech synthesis model that predicts a speech waveform conditional on an input text sequence. It is a conditional variational autoencoder (VAE) comprised of a posterior encoder, decoder, and conditional prior.
Refactor] Move third-party related utility files into integrations/ folder 🚨🚨🚨 by @younesbelkada in #25599Moves all third party libs (outside HF ecosystem) related utility files inside integrations/ instead of having them in transformers directly.
In order to get the previous usage you should be changing your call to the following:
- from transformers.deepspeed import HfDeepSpeedConfig
+ from transformers.integrations import HfDeepSpeedConfig
TRANSFORMERS_TEST_BACKEND by @vvvm23 in #25655SPM] Patch spm Llama and T5 by @ArthurZucker in #25656GPTNeo] Add input_embeds functionality to gpt_neo Causal LM by @ArthurZucker in #25664utils/documentation_tests.txt by @ydshieh in #25680pad_token check condition by @ydshieh in #25685inputs_embeds by @gante in #25687configuration_gpt2.py by @susnato in #25676LlamaTokenizer] make unk_token_length a property by @ArthurZucker in #25689test_batch_generation for bloom by @ydshieh in #25718PEFT] Fix peft version by @younesbelkada in #25710AutoGPTQ] Add correct installation of GPTQ library + fix slow tests by @younesbelkada in #25713do_sample=False when temperature=0.0 by @gante in #25722from_pretrained] Simpler code for peft by @ArthurZucker in #25726from_pretrained] Fix failing PEFT tests by @younesbelkada in #25733visual_question_answering.md to Korean by @wonhyeongseo in #25679PEFT] Fix PeftConfig save pretrained when calling add_adapter by @younesbelkada in #25738Sentencepiece] make sure legacy do not require protobuf by @ArthurZucker in #25684HammingDiversityLogitsProcessor by @gante in #25756LlamaFamiliy] add a tip about dtype by @ArthurZucker in #25794hidden_act by @stas00 in #25787Docs] More clarifications on BT + FA by @younesbelkada in #25823LlamaTokenizer] tokenize nits. by @ArthurZucker in #25793model_memory_anatomy.md to Korean by @mjk0618 in #25755add_new_pipeline.md to Korean by @heuristicwave in #25498community.md to Korean by @sim-so in #25674Pop2Piano checkpoints by @susnato in #25827generate() return True in can_generate() by @gante in #25838generation_strategies.md by @gante in #25874stage3_gather_16bit_weights_on_model_save=False by @pacman100 in #25817TokenizerFast] can_save_slow_tokenizer as a property for when vocab_file's folder was removed by @ArthurZucker in #25626InstructBlip] FINAL Fix instructblip test by @younesbelkada in #25887setup.py by @ydshieh in #25893is_tensor by @sgugger in #25871ViTDet by @ydshieh in #25913The following contributors have made significant changes to the library over the last release:
Nothing published for this version
Patch release including several patches from v4.31.0, listed below:
Patch release including several patches from v4.31.0, listed below:
We continue the deprecation of models that was introduced in https://github.com/huggingface/transformers/pull/24787.
The IDEFICS model was proposed in OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents by Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh
IDEFICS is the first open state-of-the-art visual language model at the 80B scale!
The model accepts arbitrary sequences of image and text and produces text, similarly to a multimodal ChatGPT.
Blogpost: hf.co/blog/idefics Playground: HuggingFaceM4/idefics_playground
MPT has been added and is now officially supported within Transformers. The repositories from MosaicML have been updated to work best with the model integration within Transformers.
MPT] Add MosaicML's MPT model to transformers by @ArthurZucker & @younesbelkada in #24629GPTQ quantization is now supported in Transformers, through the optimum library. The backend relies on the auto_gptq library, from which we use the GPTQ and QuantLinear classes.
See below for an example of the API, quantizing a model using the new GPTQConfig configuration utility.
from transformers import AutoModelForCausalLM, AutoTokenizer, GPTQConfig
model_name = "facebook/opt-125m"
tokenizer = AutoTokenizer.from_pretrained(model_name)
config = GPTQConfig(bits=4, dataset = "c4", tokenizer=tokenizer, group_size=128, desc_act=False)
# works also with device_map (cpu offload works but not disk offload)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, quantization_config=config)
Most models under TheBloke namespace with the suffix GPTQ should be supported, for example, to load a GPTQ quantized model on TheBloke/Llama-2-13B-chat-GPTQ simply run (after installing latest optimum and auto-gptq libraries):
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "TheBloke/Llama-2-13B-chat-GPTQ"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
For more information about this feature, we recommend taking a look at the following announcement blogpost: https://huggingface.co/blog/gptq-integration
A new pipeline, dedicated to text-to-audio and text-to-speech models, has been added to Transformers. It currently supports the 3 text-to-audio models integrated into transformers: SpeechT5ForTextToSpeech, MusicGen and Bark.
See below for an example:
from transformers import pipeline
classifier = pipeline(model="suno/bark")
output = pipeline("Hey it's HuggingFace on the phone!")
audio = output["audio"]
sampling_rate = output["sampling_rate"]
Classifier-Free Guidance decoding is a text generation technique developed by EleutherAI, announced in this paper. With this technique, you can increase prompt adherence in generation. You can also set it up with negative prompts, ensuring your generation doesn't go in specific directions. See its docs for usage instructions.
A new task guide going into Visual Question Answering has been added to Transformers.
We continue the deprecation of models that was introduced in https://github.com/huggingface/transformers/pull/24787.
By deprecating, we indicate that we will stop maintaining such models, but there is no intention of actually removing those models and breaking support for them (they might one day move into a separate repo/on the Hub, but we would still add the necessary imports to make sure backward compatibility stays). The main point is that we stop testing those models. The usage of the models drives this choice and aims to ease the burden on our CI so that it may be used to focus on more critical aspects of the library.
There are ongoing efforts to translate the transformers' documentation in other languages. These efforts are driven by groups independent to Hugging Face, and their work is greatly appreciated further to lower the barrier of entry to ML and Transformers.
If you'd like to kickstart such an effort or help out on an existing one, please feel free to reach out by opening an issue.
tasks/document_question_answering.md to Korean by @jungnerd in #24588quicktour.md by @wonhyeongseo in #24664serialization.md by @wonhyeongseo in #24686testing.md to Korean by @Sunmin0520 in #24900perf_train_cpu.md to Korean by @seank021 in #24911<tf_xla>.md to Korean by @54data in #24904perf_hardware.md to Korean by @augustinLib in #24966hpo_train.md to Korean by @harheem in #24968perf_infer_cpu.md to Korean by @junejae in #24920transformers_agents.md to Korean by @sim-so in #24881perf_infer_gpu_many.md to Korean by @heuristicwave in #24943perf_infer_gpu_one.md to Korean by @eenzeenee in #24978add_tensorflow_model.md to Korean by @keonju2 in #25017perf_train_cpu_many.md to Korean by @nuatmochoi in #24923add_new_model.md to Korean by @mjk0618 in #24957model_summary.md to Korean by @0525hhgus in #24625philosophy.md to Korean by @TaeYupNoh in #25010perf_train_tpu_tf.md to Korean by @0525hhgus in #25433Addition of input_data_format argument to image transforms and ImageProcessor methods, allowing the user to explicitly set the data format of the images being processed. This enables processing of images with non-standard number of channels e.g. 4 or removes error which occur when the data format was inferred but the channel dimension was ambiguous.
import numpy as np
from transformers import ViTImageProcessor
img = np.random.randint(0, 256, (4, 6, 3))
image_processor = ViTImageProcessor()
inputs = image_processor(img, image_mean=0, image_std=1, input_data_format="channels_first")
torch.scaled_dot_product_attention & Flash AttentionUsers are not aware that it is possible to force dispatch torch.scaled_dot_product_attention method from torch to use Flash Attention kernels. This leads to considerable speedup and memory saving, and is also compatible with quantized models. We decided to make this explicit to users in the documentation.
In a nutshell, one can just run:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("facebook/opt-350m")
model = AutoModelForCausalLM.from_pretrained("facebook/opt-350m").to("cuda")
# convert the model to BetterTransformer
model.to_bettertransformer()
input_text = "Hello my dog is cute and"
inputs = tokenizer(input_text, return_tensors="pt").to("cuda")
+ with torch.backends.cuda.sdp_kernel(enable_flash=True, enable_math=False, enable_mem_efficient=False):
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
to enable Flash-attenion in their model. However, this feature does not support padding yet.
Users will no longer encounter CPU RAM OOM when using FSDP to train very large models in multi-gpu or multi-node multi-gpu setting.
Users no longer have to pass fsdp_transformer_layer_cls_to_wrap as the code now use _no_split_modules by default which is available for most of the popular models. DeepSpeed Z3 init now works properly with Accelerate Launcher + Trainer.
Trainer classThe default optimizer in the Trainer class has been updated to be adam_torch rather than our own adam_hf, as the official Torch optimizer is more robust and fixes some issues.
In order to keep the old behavior, ensure that you pass "adamw_hf" as the optim value in your TrainingArguments.
adamw_hf to adamw_torch 🚨🚨🚨 by @muellerzr in #25109There was an issue with the definition of the rescale of values with ViVit and EfficientNet. These have been fixed, but will result in different model outputs for both of these models. To understand the change and see what needs to be done to obtain previous results, please take a look at the following PR.
The EfficientNetForImageClassification model class did not follow conventions and added a softmax to the model logits. This was removed so that it respects the convention set by other models.
In order to obtain previous results, pass the model logits through a softmax.
Some SPM models had issues with their management of added tokens. Namely the Llama and T5, among others, were behaving incorrectly. These have been updated in https://github.com/huggingface/transformers/pull/25224.
An option to obtain the previous behavior was added through the legacy flag, as explained in the PR linked above.
SPM] Finish fix spm models 🚨🚨🚨 by @ArthurZucker in #25224use_cache=True by @ydshieh in #24893test_model_parallelism for FalconModel by @ydshieh in #24914Llama2] replace self.pretraining_tp with self.config.pretraining_tp by @younesbelkada in #24906image_processing_vilt.py wrong default documented by @stas00 in #24931main_input_name in src/transformers/keras_callbacks.py by @ydshieh in #24916LogitsProcessor class by @shauray8 in #24848RWKV] Add Gradient Checkpointing support for RWKV by @younesbelkada in #24955Parameter.ds_numel by @apoorvkh in #24942LlamaConfig] Nit: pad token should be None by default by @ArthurZucker in #24958llama tokenization doctest by @ydshieh in #24990bnb] Add simple check for bnb import by @younesbelkada in #24995Llama] remove persistent inv_freq tensor by @ArthurZucker in #24998logging.py] set default stderr path if None by @ArthurZucker in #25033TrainingArgs to wandb.config without sanitization. by @parambharat in #250358bit] Fix 8bit corner case with Blip2 8bit by @younesbelkada in #25047RWKV] Add note in doc on RwkvStoppingCriteria by @ArthurZucker in #25055TF32 flag for PyTorch cuDNN backend by @XuehaiPan in #25075per_gpu_eval_batch_size with per_device_eval_batch_size in readme of multiple-choice task by @statelesshz in #25078generate] Only warn users if the generation_config's max_length is set to the default value by @ArthurZucker in #25030ForSequenceClassification] Support left padding by @ArthurZucker in #24979TF] Also apply patch to support left padding by @ArthurZucker in #25085test_model_is_small by @connor-henderson in #25087PreTrainedTokenizerFast] Keep properties from fast tokenizer by @ArthurZucker in #25053MusicgenForConditionalGeneration tests by @ydshieh in #25091T5, MT5, UMT5] Add [T5, MT5, UMT5]ForSequenceClassification by @sjrl in #24726PvtModelIntegrationTest::test_inference_fp16 by @ydshieh in #25106use_auth_token -> token by @ydshieh in #25083T5/LlamaTokenizer] default legacy to None to not always warn by @ArthurZucker in #25131MptConfig] support from pretrained args by @ArthurZucker in #25116token things by @ydshieh in #25146.push_to_hub and cleanup get_full_repo_name usage by @Wauplin in #25120use_auth_token -> token in example scripts by @ydshieh in #25167Mpt] Fix mpt slow test by @younesbelkada in #25170InstructBlip] Fix instructblip slow test by @younesbelkada in #25171_prepare_output_docstrings by @ydshieh in #25202PreTrainedModel] Wrap cuda and to method correctly by @younesbelkada in #25206all_model_classes in FlaxBloomGenerationTest by @ydshieh in #25211pipeline] revisit device check for pipeline by @younesbelkada in #25207Pix2Struct] Fix pix2struct cross attention by @younesbelkada in #25200Docs/quantization] Clearer explanation on how things works under the hood. + remove outdated info by @younesbelkada in #25216MPT] Add require_bitsandbytes on MPT integration tests by @younesbelkada in #25201Detr] Fix detr BatchNorm replacement issue by @younesbelkada in #25230token arugment in example scripts by @ydshieh in #25172pytest_options={"rA": None} in CI by @ydshieh in #25263num_hidden_layers=2 🚀🚀🚀 by @ydshieh in #25266pytest_num_workers=8 for torch/tf jobs by @ydshieh in #25274report_to logging integrations in docstring by @tomaarsen in #25281bark could have tiny model by @ydshieh in #25290trust_remote_code in example scripts by @Jackmin801 in #25248Repository to upload_folder by @sgugger in #25095NoRepeatNGramLogitsProcessor Example for LogitsProcessor class by @Rishab26 in #25186torch.compile() for vision models by @merveenoyan in #24748test_model_parallelism by @ydshieh in #25359token in example template by @ydshieh in #25351torch_job worker(s) crashing by @ydshieh in #25374token by @ydshieh in #25382OneFormerModelTest.test_model_with_labels by @ydshieh in #25383TopPLogitsWarper by @chiral-carbon in #25361device_map is passed by @gante in #25413torch.compile() docs by @merveenoyan in #25432examples to tests to run when setup.py is modified by @ydshieh in #25437main on PRs/branches if setup.py is not modified by @ydshieh in #25445main on PRs/branches" by @ydshieh in #25466auxiliary_head is None in UperNetPreTrainedModel by @mmurray in #25514MaskFormerModelIntegrationTest OOM by @ydshieh in #25544torch.fx tests on nightly CI by @ydshieh in #25549test_onnx_runtime_optimize for now by @ydshieh in #25560Docs] Fix un-rendered images by @younesbelkada in #25561TRANSFORMERS_TEST_DEVICE by @vvvm23 in #25506test_beam_search_xla_generate_simple for T5 by @ydshieh in #25566resize_embedding] Introduce pad_to_multiple_of and guidance by @ArthurZucker in #25088SwitchTransformers] Remove unused module by @ArthurZucker in #25427NllbMoe] Update code to properly support loss computation by @ArthurZucker in #25429Tests] Fix failing 8bit test by @younesbelkada in #25564test_contrastive_generate for TFXLNet by @ydshieh in #25574Docs / BetterTransformer ] Added more details about flash attention + SDPA by @younesbelkada in #25265.cuda with .to(torch_device) in tests by @vvvm23 in #25571split_special_tokens] Add support for split_special_tokens argument to encode by @ArthurZucker in #25081Llama] remove prompt and fix prefix finetuning by @ArthurZucker in #25565TokenizerFast] Fix setting prefix space in init by @ArthurZucker in #25563resize_token_embeddings by @SunMarc in #25596The following contributors have made significant changes to the library over the last release:
quicktour.md (#24664)serialization.md (#24686)testing.md to Korean (#24900)T5, MT5, UMT5] Add [T5, MT5, UMT5]ForSequenceClassification (#24726)trust_remote_code in example scripts (#25248)add_new_model.md to Korean (#24957)Transformers is growing a lot and to ease a bit the burden of maintenance on our side, we have taken the decision to deprecate models that are not use…
Llama 2 was proposed in LLaMA: Open Foundation and Fine-Tuned Chat Models by Hugo Touvron et al. It builds upon the Llama architecture adding Grouped Query Attention for efficient inference.
The MusicGen model was proposed in the paper Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi and Alexandre Défossez.
MusicGen is a single stage auto-regressive Transformer model capable of generating high-quality music samples conditioned on text descriptions or audio prompts. The text descriptions are passed through a frozen text encoder model to obtain a sequence of hidden-state representations. MusicGen is then trained to predict discrete audio tokens, or audio codes, conditioned on these hidden-states. These audio tokens are then decoded using an audio compression model, such as EnCodec, to recover the audio waveform.
Through an efficient token interleaving pattern, MusicGen does not require a self-supervised semantic representation of the text/audio prompts, thus eliminating the need to cascade multiple models to predict a set of codebooks (e.g. hierarchically or upsampling). Instead, it is able to generate all the codebooks in a single forward pass.
Bark is a transformer-based text-to-speech model proposed by Suno AI in suno-ai/bark.
The MMS model was proposed in Scaling Speech Technology to 1,000+ Languages by Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang, Wei-Ning Hsu, Alexis Conneau, Michael Auli
The EnCodec neural codec model was proposed in High Fidelity Neural Audio Compression by Alexandre Défossez, Jade Copet, Gabriel Synnaeve, Yossi Adi.
The InstructBLIP model was proposed in InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi. InstructBLIP leverages the BLIP-2 architecture for visual instruction tuning.
The UMT5 model was proposed in UniMax: Fairer and More Effective Language Sampling for Large-Scale Multilingual Pretraining by Hyung Won Chung, Xavier Garcia, Adam Roberts, Yi Tay, Orhan Firat, Sharan Narang, Noah Constant.
Umt5] Add google's umt5 to transformers by @ArthurZucker in #24477The MRA model was proposed in Multi Resolution Analysis (MRA) for Approximate Self-Attention by Zhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn M Fung, and Vikas Singh.
The Vivit model was proposed in ViViT: A Video Vision Transformer by Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, Cordelia Schmid. The paper proposes one of the first successful pure-transformer based set of models for video understanding.
The last version to support Python 3.7 was 4.30.x, as it reached end-of-life on June 27, 2023 and is no longer supported by the Python Software Foundation.
The last version to support PyTorch 1.9 was 4.30.x. As it has been more than 2 years, and we're looking forward to using features available in PyTorch 1.10 and up, we do not support PyTorch 1.9 for v4.31 and up.
This PR adds RoPE scaling to the LLaMa and GPTNeoX families of models. It allows us to extrapolate and go beyond the original maximum sequence length (e.g. 2048 tokens on LLaMA), without fine-tuning. It offers two strategies:
Tools now return a type that is specific to agents. This type can return a serialized version of itself (a string), that either points to a file on-disk or to the object's content. This should make interaction with text-based systems much simpler.
Models with potentially tied weights dropped off some keys from the state dict even when the weights were not tied. This has now been fixed and more generally, the whole experience of loading a model with state dict that don't match exactly should be improved in this release.
This PR adds a method of predicting timestamps at the word (or even token) level, by analyzing the cross-attentions and applying dynamic time warping.
A new auto model is added, AutoModelForTextEncoding. It is to be used when you want to extract the text encoder from an encoder-decoder architecture.
Transformers is growing a lot and to ease a bit the burden of maintenance on our side, we have taken the decision to deprecate models that are not used a lot. Those models will never actually disappear from the library, but we will stop testing them or accepting PRs modifying them. (enfin ça The criteria to identify models to deprecate was less than 1,000 unique downloads in the last 30 days for models that are at least one year old. The list of deprecated models is:
Fixes an issue with stripped spaces for the T5 family tokenizers. If this impacts negatively inference/training with your models, please let us know by opening an issue.
T5Tokenize] Fix T5 family tokenizers⚠️⚠️ by @ArthurZucker in #24565add trust_remote_code option to CLI download cmd by @radames in #24097
Fix typo in Llama docstrings by @Kh4L in #24020
Avoid GPT-2 daily CI job OOM (in TF tests) by @ydshieh in #24106
[Lllama] Update tokenization code to ensure parsing of the special tokens [core] by @ArthurZucker in #24042
PLAM => PaLM by @xingener in #24129
[bnb] Fix bnb config json serialization by @younesbelkada in #24137
Correctly build models and import call_context for older TF versions by @Rocketknight1 in #24138
Generate: PT's top_p enforces min_tokens_to_keep when it is 1 by @gante in #24111
fix bugs with trainer by @pacman100 in #24134
Fix TF Rag OOM issue by @ydshieh in #24122
Fix SAM OOM issue on CI by @ydshieh in #24125
Fix XGLM OOM on CI by @ydshieh in #24123
[SAM] Fix sam slow test by @younesbelkada in #24140
[lamaTokenizerFast] Update documentation by @ArthurZucker in #24132
[BlenderBotSmall] Update doc example by @ArthurZucker in #24092
Fix Pipeline CI OOM issue by @ydshieh in #24124
[documentation] grammatical fixes in image_classification.mdx by @LiamSwayne in #24141
Fix typo in streamers.py by @freddiev4 in #24144
[tests] fix bitsandbytes import issue by @stas00 in #24151
Avoid OOM in doctest CI by @ydshieh in #24139
Fix Wav2Vec2 CI OOM by @ydshieh in #24190
Fix push to hub by @NielsRogge in #24187
Change ProgressCallback to use dynamic_ncols=True by @gmlwns2000 in #24101
[i18n]Translated "attention.mdx" to korean by @kihoon71 in #23878
Generate: force caching on the main model, in assisted generation by @gante in #24177
Fix device issue in OpenLlamaModelTest::test_model_parallelism by @ydshieh in #24195
Update GPTNeoXLanguageGenerationTest by @ydshieh in #24193
typo: fix typos in CONTRIBUTING.md and deepspeed.mdx by @zsj9509 in #24184
Generate: detect special architectures when loaded from PEFT by @gante in #24198
🌐 [i18n-KO] Translated tasks_summary.mdx to Korean by @kihoon71 in #23977
🚨🚨🚨 Replace DataLoader logic for Accelerate in Trainer, remove unneeded tests 🚨🚨🚨 by @muellerzr in #24028
Fix _load_pretrained_model by @SunMarc in #24200
Fix steps bugs in no trainer examples by @Ethan-yt in #24197
Skip RWKV test in past CI by @ydshieh in #24204
Remove unnecessary aten::to overhead in llama by @fxmarty in #24203
Update WhisperForAudioClassification doc example by @ydshieh in #24188
Finish dataloader integration by @muellerzr in #24201
Add the number of model test failures to slack CI report by @ydshieh in #24207
fix: TextIteratorStreamer cannot work with pipeline by @yuanwu2017 in #23641
Update (TF)SamModelIntegrationTest by @ydshieh in #24199
Improving error message when using use_safetensors=True. by @Narsil in #24232
Safely import pytest in testing_utils.py by @amyeroberts in #24241
fix overflow when training mDeberta in fp16 by @sjrl in #24116
deprecate use_mps_device by @pacman100 in #24239
Tied params cleanup by @sgugger in #24211
[Time Series] use mean scaler when scaling is a boolean True by @kashif in #24237
TF: standardize test_model_common_attributes for language models by @gante in #23457
Generate: GenerationConfig can overwrite attributes at from_pretrained time by @gante in #24238
Add torch >=1.12 requirement for Tapas by @ydshieh in #24251
Update urls in warnings for rich rendering by @IvanReznikov in #24136
Fix how we detect the TF package by @Rocketknight1 in #24255
Stop storing references to bound methods via tf.function by @Rocketknight1 in #24146
Skip GPT-J fx tests for torch < 1.12 by @ydshieh in #24256
docs wrt using accelerate launcher with trainer by @pacman100 in #24250
update FSDP save and load logic by @pacman100 in #24249
Fix URL in comment for contrastive loss function by @taepd in #24271
QA doc: import torch before it is used by @ByronHsu in #24228
Skip some TQAPipelineTests tests in past CI by @ydshieh in #24267
TF: CTRL with native embedding layers by @gante in #23456
Adapt Wav2Vec2 conversion for MMS lang identification by @patrickvonplaten in #24234
Update check of core deps by @sgugger in #24277
Pix2StructImageProcessor requires torch>=1.11.0 by @ydshieh in #24270
Fix Debertav2 embed_proj by @WissamAntoun in #24205
Clean up old Accelerate checks by @sgugger in #24279
Fix bug in slow tokenizer conversion, make it a lot faster by @stephantul in #24266
Fix check_config_attributes: check all configuration classes by @ydshieh in #24231
Fix LLaMa beam search when using parallelize by @FeiWang96 in #24224
remove unused is_decoder parameter in DetrAttention by @JayL0321 in #24226
Split common test from core tests by @sgugger in #24284
[fix] bug in BatchEncoding.getitem by @flybird1111 in #24293
Fix image segmentation tool bug by @amyeroberts in #23897
[Docs] Improve docs for MMS loading of other languages by @patrickvonplaten in #24292
Update README_zh-hans.md by @CooperFu in #24181
deepspeed init during eval fix by @pacman100 in #24298
[EnCodec] Changes for 32kHz ckpt by @sanchit-gandhi in #24296
[Docs] Fix the paper URL for MMS model by @hitchhicker in #24302
Update tokenizer_summary.mdx (grammar) by @belladoreai in #24286
Beam search type by @jprivera44 in #24288
Make can_generate as class method by @ydshieh in #24299
Update test versions on README.md by @sqali in #24307
[SwitchTransformers] Fix return values by @ArthurZucker in #24300
Fix functional TF Whisper and modernize tests by @Rocketknight1 in #24301
Big TF test cleanup by @Rocketknight1 in #24282
Fix ner average grouping with no groups by @Narsil in #24319
Fix ImageGPT doc example by @amyeroberts in #24317
Add test for proper TF input signatures by @Rocketknight1 in #24320
Adding ddp_broadcast_buffers argument to Trainer by @TevenLeScao in #24326
error bug on saving distributed optim state when using data parallel by @xshaun in #24108
🌐 [i18n-KO] Fixed tutorial/preprocessing.mdx by @sim-so in #24156
pin apex to a speicifc commit (for DeepSpeed CI docker image) by @ydshieh in #24351
byebye Hub connection timeout by @ydshieh in #24350
Clean up disk sapce during docker image build for transformers-pytorch-gpu by @ydshieh in #24346
Fix KerasMetricCallback: pass generate_kwargs even if use_xla_generation is False by @Kripner in #24333
Fix device issue in SwitchTransformers by @ydshieh in #24352
Update MMS integration docs by @vineelpratap in #24311
Make AutoFormer work with previous torch version by @ydshieh in #24357
Fix ImageGPT doctest by @amyeroberts in #24353
Fix link to documentation in Install from Source by @SoyGema in #24336
docs: add BentoML to awesome-transformers by @aarnphm in #24344
[Doc Fix] Fix model name path in the transformers doc for AutoClasses by @riteshghorse in #24329
Fix the order in GPTNeo's docstring by @qgallouedec in #24358
Respect explicitly set framework parameter in pipeline by @denis-ismailaj in #24322
Allow passing kwargs through to TFBertTokenizer by @Rocketknight1 in #24324
Fix resuming PeftModel checkpoints in Trainer by @llohann-speranca in #24274
TensorFlow CI fixes by @Rocketknight1 in #24360
Update tiny models for pipeline testing. by @ydshieh in #24364
[modelcard] add audio classification to task list by @sanchit-gandhi in #24363
[Whisper] Make tests faster by @sanchit-gandhi in #24105
Rename test to be more accurate by @sgugger in #24374
Add a check in ImageToTextPipeline._forward by @ydshieh in #24373
[Tokenizer doc] Clarification about add_prefix_space by @ArthurZucker in #24368
style: add BitsAndBytesConfig repr function by @aarnphm in #24331
Better test name and enable pipeline test for pix2struct by @ydshieh in #24377
Skip a tapas (tokenization) test in past CI by @ydshieh in #24378
[Whisper Docs] Nits by @ArthurZucker in #24367
[GPTNeoX] Nit in config by @ArthurZucker in #24349
[Wav2Vec2 - MMS] Correct directly loading adapters weights by @patrickvonplaten in #24335
Migrate doc files to Markdown. by @sgugger in #24376
Update deprecated torch.ger by @kit1980 in #24387
[docs] Fix NLLB-MoE links by @stevhliu in #24388
Add ffmpeg for doc_test_job on CircleCI by @ydshieh in #24397
byebye Hub connection timeout - Recast by @ydshieh in #24399
fix type annotation for debug arg by @Bearnardd in #24033
[Trainer] Fix optimizer step on PyTorch TPU by @cowanmeg in #24389
Fix gradient checkpointing + fp16 autocast for most models by @younesbelkada in #24247
Clean up dist import by @muellerzr in #24402
Check auto mappings could be imported via from transformers by @ydshieh in #24400
Remove redundant code from TrainingArgs by @muellerzr in #24401
Explicit arguments in from_pretrained by @ydshieh in #24306
[ASR pipeline] Check for torchaudio by @sanchit-gandhi in #23953
TF safetensors reduced mem usage by @Rocketknight1 in #24404
Skip test_conditional_generation_pt_pix2struct in Past CI (torch < 1.11) by @ydshieh in #24417
[bnb] Fix bnb serialization issue with new release by @younesbelkada in #24416
Revert "Fix gradient checkpointing + fp16 autocast for most models" by @younesbelkada in #24420
Fix save_cache version in config.yml by @ydshieh in #24419
Update RayTune doc link for Hyperparameter tuning by @JoshuaEPSamuel in #24422
TF CI fix for Segformer by @Rocketknight1 in #24426
Refactor hyperparameter search backends by @alexmojaki in #24384
Clarify batch size displayed when using DataParallel by @sgugger in #24430
Save site-packages as cache in CircleCI job by @ydshieh in #24424
[llama] Fix comments in weights converter by @weimingzha0 in #24436
[Trainer] Fix .to call on 4bit models by @younesbelkada in #24444
fix the grad_acc issue at epoch boundaries by @pacman100 in #24415
Replace python random with torch.rand to enable dynamo.export by @BowenBao in #24434
Fix typo by @siryuon in #24440
Fix some TFWhisperModelIntegrationTests by @ydshieh in #24428
fixes issue when saving fsdp via accelerate's FSDP plugin by @pacman100 in #24446
Allow dict input for audio classification pipeline by @sanchit-gandhi in #23445
Update JukeboxConfig.from_pretrained by @ydshieh in #24443
Improved keras imports by @Rocketknight1 in #24448
add missing alignment_heads to Whisper integration test by @hollance in #24487
Fix tpu_metrics_debug by @cowanmeg in #24452
Update AlbertModel type annotation by @amyeroberts in #24450
[pipeline] Fix str device issue by @younesbelkada in #24396
when resume from peft checkpoint, the model should be trainable by @sywangyi in #24463
deepspeed z1/z2 state dict fix by @pacman100 in #24489
Update InstructBlipModelIntegrationTest by @ydshieh in #24490
Update token_classification.md by @condor-cp in #24484
Add support for for loops in python interpreter by @sgugger in #24429
[InstructBlip] Add accelerate support for instructblip by @younesbelkada in #24488
Compute dropout_probability only in training mode by @ydshieh in #24486
Fix 'local_rank' AttiributeError in Trainer class by @mocobeta in #24297
Compute dropout_probability only in training mode (SpeechT5) by @ydshieh in #24498
Fix link in utils by @SoyGema in #24501
🚨🚨 Fix group beam search by @hukuda222 in #24407
Generate: group_beam_search requires diversity_penalty>0.0 by @gante in #24456
Generate: min_tokens_to_keep has to be >= 1 by @gante in #24453
Fix TypeError: Object of type int64 is not JSON serializable by @xiaoli in #24340
Fix poor past ci by @ydshieh in #24485
🌐 [i18n-KO] Translated tflite.mdx to Korean by @0525hhgus in #24435
use accelerate autocast in jit eval path, since mix precision logic is… by @sywangyi in #24460
Update hyperparameter_search.py by @pacman100 in #24515
[T5] Add T5ForQuestionAnswering and MT5ForQuestionAnswering by @sjrl in #24481
set model to training mode before accelerate.prepare by @sywangyi in #24520
Update huggingface_hub commit sha by @ydshieh in #24527
Find module name in an OS-agnostic fashion by @sgugger in #24526
Fix LR scheduler based on bs from auto bs finder by @muellerzr in #24521
[Mask2Former] Remove SwinConfig by @NielsRogge in #24259
Allow backbones not in backbones_supported - Maskformer Mask2Former by @amyeroberts in #24532
Fix Typo by @tony9402 in #24530
Finishing tidying keys to ignore on load by @sgugger in #24535
Add bitsandbytes support for gpt2 models by @DarioSucic in #24504
⚠️ Time to say goodbye to py37 by @ydshieh in #24091
Unpin DeepSpeed and require DS >= 0.9.3 by @ydshieh in #24541
Allow for warn_only selection in enable_full_determinism by @Frank995 in #24496
Fix typing annotations for FSDP and DeepSpeed in TrainingArguments by @mryab in #24549
Update PT/TF weight conversion after #24030 by @ydshieh in #24547
Update EncodecIntegrationTest by @ydshieh in #24553
[gpt2-int8] Add gpt2-xl int8 test by @younesbelkada in #24543
Fix processor init bug if image processor undefined by @amyeroberts in #24554
[InstructBlip] Add instruct blip int8 test by @younesbelkada in #24555
Update PT/Flax weight conversion after #24030 by @ydshieh in #24556
Make PT/Flax tests could be run on GPU by @ydshieh in #24557
Update masked_language_modeling.md by @condor-cp in #24560
Fixed OwlViTModel inplace operations by @pasqualedem in #24529
Update old existing feature extractor references by @amyeroberts in #24552
Fix Typo by @tony9402 in #24559
Fix annotations by @tony9402 in #24571
Docs: 4 bit doc corrections by @gante in #24572
Revert "Fix typing annotations for FSDP and DeepSpeed in TrainingArguments" by @sgugger in #24574
Update some torchscript tests after #24505 by @ydshieh in #24566
Removal of deprecated vision methods and specify deprecation versions by @amyeroberts in #24570
Fix ESM models buffers by @sgugger in #24576
Check all objects are equally in the main __init__ file by @ydshieh in #24573
Fix annotations by @tony9402 in #24582
fix peft ckpts not being pushed to hub by @pacman100 in #24578
Udate link to RunHouse hardware setup documentation. by @BioGeek in #24590
Show a warning for missing attention masks when pad_token_id is not None by @hackyon in #24510
Make (TF) CI faster (test only a subset of model classes) by @ydshieh in #24592
Speed up TF tests by reducing hidden layer counts by @Rocketknight1 in #24595
[several models] improve readability by @stas00 in #24585
Use protobuf 4 by @ydshieh in #24599
Limit Pydantic to V1 in dependencies by @lig in #24596
🌐 [i18n-KO] Translated perplexity.mdx to Korean by @HanNayeoniee in #23850
[Time-Series] Added blog-post to tips by @elisim in #24482
Pin Pillow for now by @ydshieh in #24633
Fix loading dataset docs link in run_translation.py example by @SoyGema in #24594
Generate: multi-device support for contrastive search by @gante in #24635
Generate: force cache with inputs_embeds forwarding by @gante in #24639
precompiled_charsmap checking before adding to the normalizers' list for XLNetTokenizerFast conversion. by @shahad-mahmud in #24618
Fix audio feature extractor deps by @sanchit-gandhi in #24636
llama fp16 torch.max bug fix by @prathikr in #24561
documentation_tests.txt - sort filenames alphabetically by @amyeroberts in #24647
Update warning messages reffering to post_process_object_detection by @rafaelpadilla in #24649
Add finetuned_from property in the autogenerated model card by @sgugger in #24528
Make warning disappear for remote code in pipelines by @sgugger in #24603
Fix EncodecModelTest::test_multi_gpu_data_parallel_forward by @ydshieh in #24663
Fix VisionTextDualEncoderIntegrationTest by @ydshieh in #24661
Add is_torch_mps_available function to utils by @NripeshN in #24660
Unpin huggingface_hub by @ydshieh in #24667
Fix model referenced and results in documentation. Model mentioned was inaccessible by @rafaelpadilla in #24609
Add Nucleotide Transformer notebooks and restructure notebook list by @Rocketknight1 in #24669
LlamaTokenizer should be picklable by @icyblade in #24681
Add dropouts to GPT-NeoX by @ZHAOTING in #24680
DeepSpeed/FSDP ckpt saving utils fixes and FSDP training args fixes by @pacman100 in #24591
Avoid import sentencepiece_model_pb2 in utils.__init__.py by @ydshieh in #24689
Fix integration with Accelerate and failing test by @muellerzr in #24691
[MT5] Fix CONFIG_MAPPING issue leading it to load umt5 class by @ArthurZucker in #24678
Fix flaky test_for_warning_if_padding_and_no_attention_mask by @ydshieh in #24706
Whisper: fix prompted max length by @gante in #24666
Enable conversational pipeline for GPTSw3Tokenizer by @saattrupdan in #24648
[T5] Adding model_parallel = False to T5ForQuestionAnswering and MT5ForQuestionAnswering by @sjrl in #24684
Docs: change some input_ids doc reference from BertTokenizer to AutoTokenizer by @gante in #24730
add link to accelerate doc by @SunMarc in #24601
[Patch-t5-tokenizer] Patches the changes on T5 to make sure previous behaviour is still valide for beginning of words by @ArthurZucker in #24622
Fix typo in LocalAgent by @jamartin9 in #24736
fix: Text splitting in the BasicTokenizer by @connor-henderson in #22280
Docs: add kwargs type to fix formatting by @gante in #24733
add gradient checkpointing for distilbert by @jordane95 in #24719
Skip keys not in the state dict when finding mismatched weights by @sgugger in #24749
Fix non-deterministic Megatron-LM checkpoint name by @janEbert in #24674
[InstructBLIP] Fix bos token of LLaMa checkpoints by @NielsRogge in #24492
Skip some slow tests for doctesting in PRs (Circle)CI by @ydshieh in #24753
Fix lr scheduler not being reset on reruns by @muellerzr in #24758
:bug: Handle empty gen_kwargs for seq2seq trainer prediction_step function by @gkumbhat in #24759
Allow existing configs to be registered by @sgugger in #24760
Unpin protobuf in docker file (for daily CI) by @ydshieh in #24761
Fix eval_accumulation_steps leading to incorrect metrics by @muellerzr in #24756
Add MobileVitV2 to doctests by @amyeroberts in #24771
Docs: Update logit processors call docs by @gante in #24729
Replacement of 20 asserts with exceptions by @Baukebrenninkmeijer in #24757
Update default values of bos/eos token ids in CLIPTextConfig by @ydshieh in #24773
Fix pad across processes dim in trainer and not being able to set the timeout by @muellerzr in #24775
gpt-bigcode: avoid zero_ to support Core ML by @pcuenca in #24755
Remove WWT from README by @LysandreJik in #24672
Rm duplicate pad_across_processes by @muellerzr in #24780
Revert "Unpin protobuf in docker file (for daily CI)" by @ydshieh in #24800
Removing unnecessary device=device in modeling_llama.py by @Liyang90 in #24696
[fix] Change the condition of ValueError in "convert_checkpoint_from_transformers_to_megatron" by @SeongBeomLEE in #24769
[DOC] Clarify relationshi load_best_model_at_end and save_total_limit by @BramVanroy in #24614
Upgrade jax/jaxlib/flax pin versions by @ydshieh in #24791
Fix MobileVitV2 doctest checkpoint by @amyeroberts in #24805
Skip torchscript tests for MusicgenForConditionalGeneration by @ydshieh in #24782
Generate: add SequenceBiasLogitsProcessor by @gante in #24334
Add accelerate version in transformers-cli env by @amyeroberts in #24806
Fix typo 'submosules' by @dymil in #24809
Remove Falcon docs for the release until TGI is ready by @Rocketknight1 in #24808
Update setup.py to be compatible with pipenv by @georgiemathews in #24789
Use _BaseAutoModelClass's register method by @fadynakhla in #24810
Run hub tests by @sgugger in #24807
Copy code when using local trust remote code by @sgugger in #24785
Fixing double use_auth_token.pop (preventing private models from being visible). by @Narsil in #24812
set correct model input names for gptsw3tokenizer by @DarioSucic in #24788
Check models used for common tests are small by @sgugger in #24824
[🔗 Docs] Fixed Incorrect Migration Link by @kadirnar in #24793
deprecate sharded_ddp training argument by @statelesshz in #24825
🌐 [i18n-KO] Translated custom_tools.mdx to Korean by @sim-so in #24580
Remove unused code in GPT-Neo by @namespace-Pt in #24826
Add Multimodal heading and Document question answering in task_summary.mdx by @y3sar in #23318
Fix is_vision_available by @ydshieh in #24853
Fix comments for _merge_heads by @bofenghuang in #24855
fix broken links in READMEs by @younesbelkada in #24861
Add TAPEX to the list of deprecated models by @sgugger in #24859
Fix token pass by @sgugger in #24862
The following contributors have made significant changes to the library over the last release:
tutorial/preprocessing.mdx (#24156)custom_tools.mdx to Korean (#24580)Fix push to hubby @NielsRogge in #24187
Fix bnb config json serialization in #24137 by @younesbelkada
Act on deprecations in Accelerate no_trainer examples by @muellerzr in #24053
Transformers has just reached 100k stars on GitHub, and to celebrate we wanted to highlight 100 projects in the vicinity of transformers and we have decided to create an awesome-transformers page to do just that.
We accept PRs to add projects to the list!
By leveraging the bitsandbytes library by @TimDettmers, we add 4-bit support to transformers models!
The Agents framework has been improved and continues to be stabilized. Among bug fixes, here are the important new features that were added:
transformers instead of relying on APIs.AzureOpenAiAgent class to support Azure OpenAI agents.The safetensors library is a safe serialization framework for machine learning tensors. It has been audited and will become the default serialization framework for several organizations (Hugging Face, EleutherAI, Stability AI).
It has now become a core dependency of transformers.
safetensors a core dependency. by @Narsil in #23254The SwiftFormer paper introduces a novel efficient additive attention mechanism that effectively replaces the quadratic matrix multiplication operations in the self-attention computation with linear element-wise multiplications. A series of models called ‘SwiftFormer’ is built based on this, which achieves state-of-the-art performance in terms of both accuracy and mobile inference speed. Even their small variant achieves 78.5% top-1 ImageNet1K accuracy with only 0.8 ms latency on iPhone 14, which is more accurate and 2× faster compared to MobileViT-v2.
This model augments the Transformer as a deep decomposition architecture, which can progressively decompose the trend and seasonal components during the forecasting process.
MobileViTV2 is the second version of MobileViT, constructed by replacing the multi-headed self-attention in MobileViT with separable self-attention.
PerSAM proposes a minimal modification to SAM to allow dreambooth-like personalization, enabling to segment concepts in new images using just one example.
We add support for loading timm weights within the AutoBackbone API in transformers. timm models can be instantiated through the TimmBackbone class, and then used with any vision model that needs a backbone.
We add conditional text generation to the image to text pipeline; allowing the model to continue generating an initial text prompt according to an image.
A major rework of the internals of the Trainer is underway, leveraging accelerate instead of redefining them in transformers. This should unify both framework and lead to increased interoperability and more efficient development.
accelerator.prepare by @pacman100 in #23914chore: allow protobuf 3.20.3 requirement by @jose-turintech in #22759
Fix link displayed for custom tools by @sgugger in #23274
Remove missplaced test file by @sgugger in #23275
Bring back the PR Refactor doctests + add CI to main by @ydshieh in #23271
[gpt] Gpt2 fix half precision causal mask by @younesbelkada in #23256
Temporary tolerance fix for flaky whipser PT-TF equiv. test by @amyeroberts in #23257
Add top_k argument to post-process of conditional/deformable-DETR by @CreatlV in #22787
transformers-cli -> huggingface-cli by @AlpinDale in #23276
Temporarily increase tol for PT-FLAX whisper tests by @amyeroberts in #23288
Added missing " in CHAT_PROMPT_TEMPLATE by @galatolofederico in #23287
Update custom_tools.mdx: fix link by @mishig25 in #23292
Update transformers_agents.mdx by @mishig25 in #23289
Convert numpy arrays to lists before saving the evaluation metrics as json by @harisankar95 in #23268
Fix doctest files fetch issue by @ydshieh in #23277
skip test_run_squad_no_trainer for now by @ydshieh in #23302
Better check for packages availability by @apbard in #23163
Add gradient_checkpointing parameter to FlaxWhisperEncoder by @raghavanone in #23300
Agents extras by @LysandreJik in #23301
Fix broken links in the agent docs by @sgugger in #23297
Fix typo in gradio-tools docs by @freddyaboulton in #23305
Fix image segmentation tool test by @sgugger in #23306
unpin tf prob by @ydshieh in #23293
Revert "search buffers for dtype" by @sgugger in #23308
Remove LanguageIdentificationTool in __init__.py as we don't have it yet by @ydshieh in #23326
Fix docker image (caused by tensorflow_text) by @ydshieh in #23321
Compute the mask in-place, with less memory reads, and on CUDA on XLNetLMHeadModel by @lezcano in #23332
Only add files with modification outside doc blocks by @ydshieh in #23327
[docs] Fix Agents and Tools docstring by @stevhliu in #23313
OR am I crazy? by @hwuebben in #23295
Handle padding warning in generation when using inputs_embeds by @zrthxn in #23131
replaced assert with raise ValueError for t5, switch_transformers, pix2struct, mt5, longt5, gptsan_japanese. by @susnato in #23273
Use cu118 with cudnn >= 8.6 in docker file by @ydshieh in #23339
Removing one of the twice defined position_embeddings in LongFormer by @GregorySenay in #23343
Fix issue introduced in PR #23163 by @ydshieh in #23363
Typo suggestion by @richardachen in #23360
Fix some is_xxx_available by @ydshieh in #23365
Fix BigBirdForMaskedLM doctest by @ydshieh in #23369
Fix OwlViTForObjectDetection.image_guided_detection doc example by @ydshieh in #23370
Revert "Only add files with modification outside doc blocks" by @ydshieh in #23371
[Bugfix] OPTDecoderLayer does not return attentions when gradient_checkpointing and training is enabled. by @gmlwns2000 in #23367
Skip failing AlignModelTest::test_multi_gpu_data_parallel_forward by @ydshieh in #23374
Fix test typos - audio feature extractors by @LWprogramming in #23310
Added type hints for Graphormer pytorch version by @dewasahu2003 in #23073
Replace NumPy Operations with JAX NumPy Equivalents for JIT Compilation Compatibility by @gojiteji in #23356
Use mkstemp to replace deprecated mktemp by @ready-research in #23372
Fix RwkvModel by @ydshieh in #23392
Update test_batched_inference_image_captioning_conditioned by @ydshieh in #23391
OPT/BioGPT: Improved attention mask shape exception by @gante in #23270
Fix chat prompt in HFAgent by @IvanSedykh in #23335
🌐 [i18n-KO] Translated asr.mdx to Korean by @sim-so in #23106
Minor fixes in transformers-tools by @Wauplin in #23364
[Pix2Struct] Add conditional generation on docstring example by @younesbelkada in #23399
Generate: faster can_generate check on TF and Flax by @gante in #23398
[AutoModel] fix torch_dtype=auto in from_pretrained by @stas00 in #23379
Docs: add link to assisted generation blog post by @gante in #23397
Build with non Python files by @sgugger in #23405
Generate: add test to check KV format by @gante in #23403
Replace appends with list comprehension. by @ttsugriy in #23359
Fix smdistributed check by @sgugger in #23414
Why crash the whole run when HFHub gives a 50x error? by @ropoctl in #23320
Run doctest (in PRs) only when some doc example(s) are modified by @ydshieh in #23387
Update ConvNextV2ModelIntegrationTest::test_inference_image_classification_head by @ydshieh in #23402
Fix a typo in HfAgent docstring. by @ttsugriy in #23420
Use dict.items to avoid unnecessary lookups. by @ttsugriy in #23415
Update 3 docker files to use cu118 by @ydshieh in #23406
[SAM] fix sam slow test by @younesbelkada in #23376
Return early once stop token is found. by @ttsugriy in #23421
[Reland] search model buffers for dtype as the last resort by @cyyever in #23319
Add Missing tokenization test [electra] by @IMvision12 in #22997
Small fixes and link in the README by @LysandreJik in #23428
TF: embeddings out of bounds check factored into function by @gante in #23427
Update Bigbird Pegasus tests by @ydshieh in #23431
Encoder-Decoder: add informative exception when the decoder is not compatible by @gante in #23426
Remove hardcoded prints in Trainer by @hugoabonizio in #23432
Fix device issue in SwiftFormerModelIntegrationTest::test_inference_image_classification_head by @ydshieh in #23435
Generate: skip left-padding tests on old models by @gante in #23437
remove unnecessary print in gpt neox sequence classifier by @cfhammill in #23433
🌐 [i18n-KO] Translated tasks/zero_shot_object_detection.mdx to Korean by @HanNayeoniee in #23430
Fix (skip) a pipeline test for RwkvModel by @ydshieh in #23444
Fix DecisionTransformerConfig doctring by @joaoareis in #23450
TF: GPT2 with native embedding layers by @gante in #23436
Make RwkvModel accept attention_mask but discard it internally by @ydshieh in #23442
Less flaky test_assisted_decoding_matches_greedy_search by @ydshieh in #23451
Update tiny models and pipeline tests by @ydshieh in #23446
Properly guard PyTorch stuff by @sgugger in #23452
Add an option to log result from the Agent by @sgugger in #23454
Clean up CUDA kernels by @sgugger in #23455
fix bug in group_texts function, that was inserting short batches by @BodaSadalla98 in #23429
feat: Whisper prompting by @connor-henderson in #22496
README: Fix affiliation for MEGA by @julien-c in #23394
Remove .data usages in optimizations.py by @alanwaketan in #23417
TF port of the Segment Anything Model (SAM) by @Rocketknight1 in #22970
[RWKV] Rwkv fix for 8bit inference by @younesbelkada in #23468
Use config to set name and description if not present by @sgugger in #23473
Fix transformers' DeepSpeed CI job by @ydshieh in #23463
Fix PretrainedConfig min_length docstring by @joaoareis in #23471
Fix: Change tensors to integers for torch.dynamo and torch.compile compatibility by @loevlie in #23475
[Blip] Remove redundant shift right by @younesbelkada in #23153
Fix DeepSpeed stuff in the nightly CI by @ydshieh in #23478
Fix confusing transformers installation in CI by @ydshieh in #23465
Fix tests/repo_utils/test_get_test_info.py by @ydshieh in #23485
Debug example code for MegaForCausalLM by @Tylersuard in #23382
Remove erroneous img closing tag by @xenova in #23646
Fix tensor device while attention_mask is not None by @zspo in #23538
Fix accelerate logger bug by @younesbelkada in #23650
Bugfix: LLaMA layer norm incorrectly changes input type and consumers lots of memory by @TimDettmers in #23535
Fix wav2vec2 is_batched check to include 2-D numpy arrays by @LWprogramming in #23223
changing the requirements to a cpu torch version that works by @sshahrokhi in #23483
Fix SAM tests and use smaller checkpoints by @Rocketknight1 in #23656
Update workflow files by @ydshieh in #23658
small fix to remove unused eos in processor when it's not used. by @Narsil in #23408
Fix typo in a parameter name for open llama model by @aaalexlit in #23637
Fix PyTorch SAM tests by @ydshieh in #23682
🌐 [i18n-KO] Translated tasks/monocular_depth_estimation.mdx to Korean by @HanNayeoniee in #23621
Fix a BridgeTower test by @ydshieh in #23694
[SAM] Fixes pipeline and adds a dummy pipeline test by @younesbelkada in #23684
TF version compatibility fixes by @Rocketknight1 in #23663
[Blip] Fix blip doctest by @younesbelkada in #23698
is_batched fix for remaining 2-D numpy arrays by @LWprogramming in #23309
Skip TFCvtModelTest::test_keras_fit_mixed_precision for now by @ydshieh in #23699
fix: load_best_model_at_end error when load_in_8bit is True by @dkqkxx in #23443
Fix some docs what layerdrop does by @zspo in #23691
add GPTJ/bloom/llama/opt into model list and enhance the jit support by @sywangyi in #23291
Paged Optimizer + Lion Optimizer for Trainer by @TimDettmers in #23217
Export to ONNX doc refocused on using optimum, added tflite by @MKhalusova in #23434
fix: use bool instead of uint8/byte in Deberta/DebertaV2/SEW-D to make it compatible with TensorRT by @uchuhimo in #23683
fix gptj could not jit.trace in GPU by @sywangyi in #23317
Better TF docstring types by @Rocketknight1 in #23477
Minor awesome-transformers.md fixes by @pagarsky in #23453
TF SAM memory reduction by @Rocketknight1 in #23732
fix: delete duplicate sentences in document_question_answering.mdx by @jungnerd in #23735
fix: Whisper generate, move text_prompt_ids trim up for max_new_tokens calculation by @connor-henderson in #23724
Overhaul TF serving signatures + dummy inputs by @Rocketknight1 in #23234
[Whisper] Reduce batch size in tests by @sanchit-gandhi in #23736
Fix the regex in get_imports to support multiline try blocks and excepts with specific exception types by @dakinggg in #23725
Remove the last few TF serving sigs by @Rocketknight1 in #23738
Fix pip install --upgrade accelerate command in modeling_utils.py by @tloen in #23747
Fix psuh_to_hub in Trainer when nothing needs pushing by @sgugger in #23751
Revamp test selection for the example tests by @sgugger in #23737
[LongFormer] code nits, removed unused parameters by @ArthurZucker in #23749
Fix is_ninja_available() by @niltok in #23752
[Nllb-Moe] Fix nllb moe accelerate issue by @younesbelkada in #23758
[OPT] Doc nit, using fast is fine by @ArthurZucker in #23789
Fix RWKV backward on GPU by @sgugger in #23774
Update trainer.mdx class_weights example by @amitportnoy in #23787
no_cuda does not take effect in non distributed environment by @sywangyi in #23795
Fix no such file or directory error by @RissyRan in #23783
Enable code-specific revision for code on the Hub by @sgugger in #23799
add type hint in pipeline model argument by @y3sar in #23740
TF SAM shape flexibility fixes by @Rocketknight1 in #23842
fix Whisper tests on GPU by @hollance in #23753
🌐 [i18n-KO] Translated fast_tokenizers.mdx to Korean by @KIHOON71 in #22956
[i18n-KO] Translated video_classification.mdx to Korean by @KIHOON71 in #23026
🌐 [i18n-KO] Translated troubleshooting.mdx to Korean by @0525hhgus in #23166
Adds a FlyteCallback by @peridotml in #23759
Update collating_graphormer.py by @clefourrier in #23862
[LlamaTokenizerFast] nit update post_processor on the fly by @ArthurZucker in #23855
#23388 Issue: Update RoBERTa configuration by @vijethmoudgalya in #23863
[from_pretrained] imporve the error message when _no_split_modules is not defined by @ArthurZucker in #23861
Editing issue with pickle def with lambda function by @Natyren in #23869
Adds AutoProcessor.from_pretrained support for MCTCTProcessor by @Ubadub in #23856
🌐 [i18n-KO] Translated pad_truncation.mdx to Korean by @sim-so in #23823
Fix bug leading to missing token in GPTSanJapaneseTokenizer by @passaglia in #23883
Fix last instances of kbit -> quantized by @sgugger in #23797
fix(configuration_llama): add keys_to_ignore_at_inference to LlamaConfig by @calico-1226 in #23891
Fix Trainer when model is loaded on a different GPU by @sgugger in #23792
Support shared tensors by @thomasw21 in #23871
ensure banned_mask and indices in same device by @cauyxy in #23901
Unpin numba by @sanchit-gandhi in #23162
[bnb] add warning when no linear by @younesbelkada in #23894
fix: Replace add_prefix_space in get_prompt_ids with manual space for FastTokenizer compatibility by @connor-henderson in #23796
[RWKV] Fix RWKV 4bit by @younesbelkada in #23910
add conditional statement for auxiliary loss calculation by @harisankar95 in #23899
Raise error if loss can't be calculated - ViT MIM by @amyeroberts in #23872
Empty circleci config by @sgugger in #23913
Bug fix - flip_channel_order for channels first images by @amyeroberts in #23701
Re-enable squad test by @sgugger in #23912
Update the update metadata job to use upload_folder by @sgugger in #23917
[PushToHub] Make it possible to upload folders by @NielsRogge in #23920
Skip device placement for past key values in decoder models by @sgugger in #23919
[Flax Whisper] Update decode docstring by @sanchit-gandhi in #23908
Effectively allow encoder_outputs input to be a tuple in pix2struct by @fxmarty in #23932
Fix doc string nits by @sheonhan in #23929
Pin rhoknp by @sgugger in #23937
rename DocumentQuestionAnsweringTool parameter input to match docstring by @Adam-D-Lewis in #23939
Update stale.yml to use HuggingFaceBot by @LysandreJik in #23941
Make TF ESM inv_freq non-trainable like PyTorch by @Rocketknight1 in #23940
Revert "Update stale.yml to use HuggingFaceBot" by @LysandreJik in #23943
#23675 Registering Malay language by @soongbren in #23689
Modify device_map behavior when loading a model using from_pretrained by @SunMarc in #23922
use _make_causal_mask in clip/vit models by @kashif in #23942
Fix ReduceLROnPlateau object has no attribute 'get_last_lr' by @wasupandceacar in #23944
[MMS] Scaling Speech Technology to 1,000+ Languages | Add attention adapter to Wav2Vec2 by @patrickvonplaten in #23813
add new mms functions to doc by @patrickvonplaten in #23954
🌐 [i18n-KO] Translated object_detection.mdx to Korean by @KIHOON71 in #23164
Trainer: fixed evaluate raising KeyError for ReduceLROnPlateau by @claudius-kienle in #23952
[Whisper Tokenizer] Skip special tokens when decoding with timestamps by @sanchit-gandhi in #23945
Add an option to reduce compile() console spam by @Rocketknight1 in #23938
Added time-series blogs to the models by @elisim in #23857
Fix typo in doc comment of BitsAndBytesConfig by @ledyba in #23978
Skip test_multi_gpu_data_parallel_forward for MobileViTV2ModelTest by @ydshieh in #24017
Update README.md by @ydshieh in #24022
Auto tokenizer registration by @Bearnardd in #23965
expose safe_serialization argument in the pipeline API by @yessenzhar in #23775
Pix2Struct: fix wrong broadcast axis of attention mask in visual encoder by @affjljoo3581 in #23976
TensorBoard callback no longer adds hparams by @bri25yu in #23999
🌐 [i18n-KO] Translated tasks_explained.mdx to Korean by @0525hhgus in #23844
Fix MobileViTV2 checkpoint name by @ydshieh in #24018
Pin deepspeed to 0.9.2 for now by @ydshieh in #24024
🌐 [i18n-KO] Translated language-modeling.mdx by @wonhyeongseo in #23969
🌐 [i18n-KO] Translated bertology.mdx to Korean by @wonhyeongseo in #23968
Add check for tied parameters by @SunMarc in #24029
Fixing single candidate_label return. by @Narsil in #24023
Use TruncatedNormal from Keras initializers by @hvaara in #24036
Prevent ZeroDivisionError on trainer.evaluate if model and dataset are tiny by @tomaarsen in #24049
Modification of one text example file should trigger said test by @sgugger in #24051
Tiny fix for check_self_hosted_runner.py by @ydshieh in #24052
Reduce memory usage in TF building by @Rocketknight1 in #24046
Move TF building to an actual build() method by @Rocketknight1 in #23760
Use new parametrization based weight norm if available by @ezyang in #24030
bring back filtered_test_list_cross_tests.txt by @ydshieh in #24055
Fix device placement for model-parallelism in generate for encoder/de… by @sgugger in #24025
Remote code improvements by @sgugger in #23959
Generate: increase left-padding test atol by @gante in #23448
[Wav2Vec2] Fix torch srcipt by @patrickvonplaten in #24062
Add support for non-rust implemented tokenization for __getitem__ method. by @jacklanda in #24039
Support PEFT models when saving the model using trainer by @younesbelkada in #24073
[Hub] Add safe_serialization in push_to_hub by @younesbelkada in #24074
Fix is_optimum_neuron_available by @michaelbenayoun in #23961
[bnb] Fix bnb skip modules by @younesbelkada in #24043
Be nice to TF by @ydshieh in #24076
Make the TF dummies even smaller by @Rocketknight1 in #24071
[doc build] Use secrets by @mishig25 in #24079
Fix expected value in tests of the test fetcher by @sgugger in #24077
Update delete_doc_comment_trigger.yml by @mishig25 in #24084
Do not prepare lr scheduler as it as the right number of steps by @sgugger in #24088
Fix a tiny typo in WhisperForConditionalGeneration::generate docstring by @sadra-barikbin in #24045
[Trainer] Correct behavior of _load_best_model for PEFT models by @younesbelkada in #24103
The following contributors have made significant changes to the library over the last release:
fast_tokenizers.mdx to Korean (#22956)Fixes the package so non-Python files (like CUDA kernels) are properly included.
Fixes the package so non-Python files (like CUDA kernels) are properly included.
Reverts a regression in the FSDP integration. Add pip install transformers["agent"] to have all dependencies agents rely on. Fixes the documentation a
Reverts a regression in the FSDP integration.
Add pip install transformers["agent"] to have all dependencies agents rely on.
Fixes the documentation about agents.
This releases has three breaking changes compared to version v4.28.0.
Transformers Agent is a new API that lets you use the library and Diffusers by prompting an agent (which is a large language model) in natural language. That agent will then output code using a set of predefined tools, leveraging the appropriate (and state-of-the-art) models for the task the user wants to perform. It is fully multimodal and extensible by the community. Learn more in the docs
SAM (Segment Anything Model) was proposed in Segment Anything by Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alex Berg, Wan-Yen Lo, Piotr Dollar, Ross Girshick.
The model can be used to predict segmentation masks of any object of interest given an input image.
SAM] Correct arxiv link by @younesbelkada in #22886SAM] Change to facebook/sam-vit-base by @younesbelkada in #22891SAM] Add sam doc by @younesbelkada in #22984DocumentQuestionAnsweringPipeline only for fast ⚡ tokenizers by @ydshieh in #22745automatic-mask-generation pipeline for Segment Anything Model (SAM) by @ArthurZucker in #22840RWKV suggests a tweak in the traditional Transformer attention to make it linear. This way, the model can be used as recurrent network: passing inputs for timestamp 0 and timestamp 1 together is the same as passing inputs at timestamp 0, then inputs at timestamp 1 along with the state of timestamp 0 (see example below).
This can be more efficient than a regular Transformer and can deal with sentence of any length (even if the model uses a fixed context length for training).
The FocalNet model was proposed in Focal Modulation Networks by Jianwei Yang, Chunyuan Li, Xiyang Dai, Lu Yuan, Jianfeng Gao. FocalNets completely replace self-attention (used in models like ViT and Swin) by a focal modulation mechanism for modeling token interactions in vision. The authors claim that FocalNets outperform self-attention based models with similar computational costs on the tasks of image classification, object detection, and segmentation.
The Open-Llama model was proposed in Open-Llama project by community developer s-JoL.
The model is mainly based on LLaMA with some modifications, incorporating memory-efficient attention from Xformers, stable embedding from Bloom, and shared input-output embedding from PLAM. And the model is pre-trained on both Chinese and English, which gives it better performance on Chinese language tasks.
Assisted generation is a new technique that lets you speed up generation with large language models by using a smaller model as assistant. The assistant model will be the ones doing multiple forward pass while the LLM will merely validate the tokens proposed by the assistant. This can lead to speed-ups up to 10x!
To avoid duplicating the model code in multiple repos when using the code on the Hub feature, loading such models will now save in their config the repo in which the code is. This way there is one source of ground truth for code on the Hub models.
This releases has three breaking changes compared to version v4.28.0.
The first one focuses on fixing training issues for Pix2Struct. This slightly affects the results, but should result in the model training much better.
Pix2Struct] Attempts to fix training issues 🚨🚨🚨 by @younesbelkada in #23004The second one is aligning the ignore index in the LUKE model to other models in the library. This breaks the convention that models should stick to their original implementation, but it was necessary in order to align with other transformers in the library
Finally, the third breaking change aims to harmonize the training procedure for most of recent additions in transformers. It should be users' responsibility to fill_mask the padding tokens of the labels with the correct value. This PR addresses the issue that was raised by other architectures such as Luke or Pix2Struct
Blip] remove labels masking by @younesbelkada in #23024torch_dtype to str when saved_model=True in save_pretrained for TF models by @ydshieh in #22740training.mdx to Korean by @gabrielwithappy in #22670DS_BUILD_AIO=1 by @ydshieh in #22741Deta in #22437 by @ydshieh in #22750serving_output for TF composite models (encoder-decoder like models) by @ydshieh in #22743sequence_classification.mdx to Korean by @0525hhgus in #22655CpmAnt model by @ydshieh in #22766tutorial/proprecssing.mdx to Korean by @sim-so in #22578test_word_time_stamp_integration for Wav2Vec2ProcessorWithLMTest by @ydshieh in #22800custom_models.mdx to Korean by @HanNayeoniee in #22534_toctree.yml by @jungnerd in #22549tasks/translation.mdx to Korean by @wonhyeongseo in #22805LayoutLMv2 and LayoutLMv3 in some pipeline tests by @ydshieh in #22774PartialState as the device handler in the Trainer by @muellerzr in #22752auto_tutorial, training by @gabrielwithappy in #22796main by @ydshieh in #22823test_eos_token_id_int_and_list_top_k_top_sampling by @ydshieh in #22826accelerate@main in CI by @ydshieh in #22859is_symbolic_tensor predicate by @hvaara in #22878FillMaskPipelineTests by @ydshieh in #22894accelerate.mdx to Korean by @0525hhgus in #22830tasks/masked_language_modeling.mdx to Korean by @HanNayeoniee in #22838tasks/summarization.mdx to Korean by @sim-so in #22783test_codegen_sample_max_time as flaky by @ydshieh in #22953stride is too high in TokenClassificationPipeline by @boyleconnor in #22942_load_pretrained_model by @hanrui1sensetime in #22947run_scripts.mdx to Korean by @HanNayeoniee in #22793create_a_model doc to Korean by @gabrielwithappy in #22754accelerete@main in PyTorch Past CI jobs by @ydshieh in #22963DeepSpeed CI job link in Past CI by @ydshieh in #22967tasks/masked_language_modeling.mdx by @HanNayeoniee in #22965DocTest] Fix correct checkpoint by @younesbelkada in #22988serialization.mdx to Korean by @wonhyeongseo in #22806tasks/image_captioning.mdx to Korean by @sim-so in #22943token_classification.mdx to Korean by @0525hhgus in #22945PEFT] Add HFTracer support for PEFT by @younesbelkada in #23006Pix2Struct] Fix pix2struct doctest by @younesbelkada in #23023multilingual.mdx to Korean by @HanNayeoniee in #23008test_offline_mode_pipeline_exception by @ydshieh in #23022BridgeTowerModelTester by @ydshieh in #23029_test_xla_generate less flaky by @ydshieh in #22996model_sharing.mdx to Korean by @0525hhgus in #22991bigbird test file by @ydshieh in #23040BridgeTower by @ydshieh in #23039BioGPTForSequenceClassification by @awinml in #22253convnext init by @IMvision12 in #23078tasks/image_classification.mdx to Korean by @0525hhgus in #23048tasks/question_answering.mdx to Korean by @jungnerd in #23012tasks/zero_shot_image_classification.mdx to Korean by @HanNayeoniee in #23065torchscript.mdx to Korean by @sim-so in #23060Flava] Fix flava torch.distributed.nn.functional import all_gather issue by @younesbelkada in #23108Pix2Struct model to set Pix2StructTextModel to is_decoder=True by @gbarello-uipath in #23051Doctest] Fix pix2struct doctest by @younesbelkada in #23121X | Y syntax for HfArgumentParser for Python 3.10+ by @XuehaiPan in #23126_toctree.yml by @HanNayeoniee in #23112symbolic_trace by @regisss in #23105GPT-J] Fix causal mask dtype by @younesbelkada in #23147max_length warning when it is not set by @gante in #23139multiple_choice.mdx by @gabrielwithappy in #23064no_trainer scripts to pre-train Vision Transformers by @awinml in #23156logging_steps, eval_steps, and save_steps by @konstantinjdobler in #23235from_config by @DyeKuu in #23246tensorflow-probability in docker files by @ydshieh in #23260The following contributors have made significant changes to the library over the last release:
training.mdx to Korean (#22670)auto_tutorial, training (#22796)create_a_model doc to Korean (#22754)multiple_choice.mdx (#23064)sequence_classification.mdx to Korean (#22655)accelerate.mdx to Korean (#22830)token_classification.mdx to Korean (#22945)model_sharing.mdx to Korean (#22991)tasks/image_classification.mdx to Korean (#23048)tutorial/proprecssing.mdx to Korean (#22578)tasks/summarization.mdx to Korean (#22783)tasks/image_captioning.mdx to Korean (#22943)torchscript.mdx to Korean (#23060)custom_models.mdx to Korean (#22534)tasks/masked_language_modeling.mdx to Korean (#22838)run_scripts.mdx to Korean (#22793)tasks/masked_language_modeling.mdx (#22965)multilingual.mdx to Korean (#23008)tasks/zero_shot_image_classification.mdx to Korean (#23065)_toctree.yml (#23112)tasks/translation.mdx to Korean (#22805)serialization.mdx to Korean (#22806)BioGPTForSequenceClassification (#22253)no_trainer scripts to pre-train Vision Transformers (#23156)Your coding agent can read these notes before it upgrades. Set up the MCP server →