NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #225 most downloaded on PyPI
Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Last release 12 days ago
09 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
6 versions withdrawn
withdrawn after publishing
10 years old
242 releases · first in 2016
Fixes a regression for DETA models
Fixes a regression for DETA models
Adapt find_tied_parameters to handle breaking change in Accelerate by @sgugger in #22360
The LLaMA model was proposed in LLaMA: Open and Efficient Foundation Language Models. It is a collection of foundation language models ranging from 7B to 65B parameters. You can request access to the weights here then use the conversion script to generate a checkpoint compatible with Hugging Face
Pix2Struct is a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct has been fine-tuned on various tasks and datasets, ranging from image captioning and visual question answering (VQA) over different inputs (books, charts, science diagrams) to captioning UI components, and others.
transformers by @younesbelkada in #22528MEGA proposes a new approach to self-attention with each encoder layer having a multi-headed exponential moving average in addition to a single head of standard dot-product attention, giving the attention mechanism stronger positional biases. This allows MEGA to perform competitively to Transformers on standard benchmarks including LRA while also having significantly fewer parameters. MEGA’s compute efficiency allows it to scale to very long sequences, making it an attractive option for long-document NLP tasks.
The model is a an optimized GPT2 model with support for Multi-Query Attention.
The mixture of experts version of the NLLB release has been added to the library.
NLLB-MoE Adds the moe model by @ArthurZucker in #22024bnb] Let's make serialization of int8 models possible by @younesbelkada in #22177You can now push 8bit models and/or load 8bit models directly from the Hub, save memory and load your 8bit models faster! An example repo here
Notes from the PR:
The BLIP image processor incorrectly passed in the dimensions to resize in the order (width, height). This is reordered to be correct.
In most cases, this won't have an effect as the default height and width are the same. However, this is not backwards compatible for custom configurations with different height, width settings and direct calls to the resize method with different height, width values.
The big problem was the prefix and suffix tokens of the NLLB tokenizer.
Previous behaviour:
>>> from transformers import NllbTokenizer
>>> tokenizer = NllbTokenizer.from_pretrained("facebook/nllb-200-distilled-600M")
>>> tokenizer("How was your day?").input_ids
[13374, 1398, 4260, 4039, 248130, 2, 256047]
>>> # 2: '</s>'
>>> # 256047 : 'eng_Latn'
New behaviour
>>> from transformers import NllbTokenizer
>>> tokenizer = NllbTokenizer.from_pretrained("facebook/nllb-200-distilled-600M")
>>> tokenizer("How was your day?").input_ids
[256047, 13374, 1398, 4260, 4039, 248130, 2]
In case you have pipelines that were relying on the old behavior, here is how you would enable it once again:
>>> from transformers import NllbTokenizer
>>> tokenizer = NllbTokenizer.from_pretrained("facebook/nllb-200-distilled-600M", legacy_behaviour = True)
[NLLB Tokenizer] Fix the prefix tokens 🚨🚨🚨 by @ArthurZucker in #22313The BLIP model is now available in TensorFlow.
As the title says, this PR adds the possibility to export TF generate with a TF-native tokenizer -- the full thing in a single TF graph.
A new task guide has been added, focusing on depth-estimation.
MgpstrModelIntegrationTest by @ydshieh in #22195XGLM] Add accelerate support for XGLM by @younesbelkada in #22207dash==2.8.1 for now for daily CI by @ydshieh in #22227dash==2.8.1 for now for daily CI" by @ydshieh in #22233TFCvtModel by @gcuder in #22267max_memory for device_map strategies by @sgugger in #22311generate(synced_gpus=True, ...) by @stas00 in #22242MBart] Add accelerate support for MBart by @younesbelkada in #22309torch<1.10 by @stas00 in #22370cmake dependencies in CI by @gante in #22383bnb] Force requires_grad to be False by @younesbelkada in #22396causal_mask is created directly on device by @jeffra in #22378bnb] fix bnb failing test by @younesbelkada in #22439Generate] Add conditional generation for multimodal models by @younesbelkada in #22424Pix2Struct] Fix slow test by @younesbelkada in #22448model_type update for auto mapping by @ArthurZucker in #22470max_position_embeddings by @gante in #22471eos_token_id < 0 checks in generate() from ValueError to warning by @lewtun in #22472Wav2Vec2ProcessorWithLM doc example by @ydshieh in #22474TextIteratorStreamer (streamer for gradio) by @gante in #22501Trainer] Force is_model_parallel when model is loaded in multiple GPUs using accelerate by @younesbelkada in #22532T5] Enable naive Pipeline Parallelism training for T5 by @younesbelkada in #22535distutils usage by @XuehaiPan in #22531pyproject.toml by @XuehaiPan in #22539bnb] Fix typo by @younesbelkada in #22556_no_split_modules for Whisper model by @pacman100 in #22486TextIteratorStreamer timeout by @gante in #22576accelerate_tests mark warnings by @gante in #22585_toctree.yml by @wonhyeongseo in #22581pipeline_model_mapping systematically by @ydshieh in #22180bnb] 8bit models should not be converted to DDP by @younesbelkada in #22628Blip] Fix slow tests and doctests with correct values by @younesbelkada in #22632autoclass_tutorial to Korean and Fix the typo of quicktour by @gabrielwithappy in #22533MegaModel CI by @ydshieh in #22652pipeline_tutorial.mdx to Korean by @wonhyeongseo in #22508MarkupLM tests' expected values by @ydshieh in #22667torch.distributed group initialization for torch_neuron disabled when optimum-neuron is installed by @michaelbenayoun in #22728The following contributors have made significant changes to the library over the last release:
_toctree.yml (#22581)pipeline_tutorial.mdx to Korean (#22508)One column per quarter.
This patch fixes a regression with FlauBERT and XLM models.
This patch fixes a regression with FlauBERT and XLM models.
Enforce max_memory for device_map strategies by @sgugger in #22311
Enforce max_memory for device_map strategies by @sgugger in #22311
Fix balanced and auto device_map by @sgugger in #22271
Patches two unwanted breaking changes that were released in v4.27.0:
Patches two unwanted breaking changes that were released in v4.27.0:
This PR deprecated the parallelize API which has been replaced by accelerate months ago. We recommend loading the model using the device_map attribute…
The goal of this model is to build a bridge between each uni-modal encoder and the cross-modal encoder to enable comprehensive and detailed interaction at each layer of the cross-modal encoder thus achieving remarkable performance on various downstream tasks with almost negligible additional performance and computational costs.
The Whisper model was integrated a few releases ago. This release offers significant performance optimizations when generating with timestamps. This was made possible by rewriting the generate() function of Whisper, which now uses the generation_config and implementing a batched timestamp prediction. The language and task can now also be setup when calling generate(). For more details about this refactoring checkout this colab.
Notably, whisper is also now supported in Flax 🚀 thanks to @andyehrenberg ! More whisper related commits:
WhisperModelTest by @ydshieh in #21883do_normalize by @ArthurZucker in #21263model_split_percents for WhisperModelTest by @ydshieh in #21922WhisperFeatureExtractor by @bofenghuang in #21938WhisperEncoderModelTest by @ydshieh in #22060Whiper] add get_input_embeddings to WhisperForAudioClassification by @younesbelkada in #22133DETA (short for Detection Transformers with Assignment) improves Deformable DETR by replacing the one-to-one bipartite Hungarian matching loss with one-to-many label assignments used in traditional detectors with non-maximum suppression (NMS). This leads to significant gains of up to 2.5 mAP.
The SpeechT5 framework consists of a shared encoder-decoder network and six modal-specific (speech/text) pre/post-nets. After preprocessing the input speech/text through the pre-nets, the shared encoder-decoder network models the sequence-to-sequence transformation, and then the post-nets generate the output in the speech/text modality based on the output of the decoder.
XLM-V is multilingual language model with a one million token vocabulary trained on 2.5TB of data from Common Crawl (same as XLM-R).
BLIP-2 leverages frozen pre-trained image encoders and large language models (LLMs) by training a lightweight, 12-layer Transformer encoder in between them, achieving state-of-the-art performance on various vision-language tasks. Most notably, BLIP-2 improves upon Flamingo, an 80 billion parameter model, by 8.7% on zero-shot VQAv2 with 54x fewer trainable parameters.
X-MOD extends multilingual masked language models like XLM-R to include language-specific modular components (language adapters) during pre-training. For fine-tuning, the language adapters in each transformer layer are frozen.
ERNIE-M is a new training method that encourages the model to align the representation of multiple languages with monolingual corpora, to overcome the constraint that the parallel corpus size places on the model performance.
The Textless Vision-Language Transformer (TVLT) is a model that uses raw visual and audio inputs for vision-and-language representation learning, without using text-specific modules such as tokenization or automatic speech recognition (ASR). It can perform various audiovisual and vision-language tasks like retrieval, question answering, etc.
CLAP (Contrastive Language-Audio Pretraining) is a neural network trained on a variety of (audio, text) pairs. It can be instructed in to predict the most relevant text snippet, given an audio, without directly optimizing for the task. The CLAP model uses a SWINTransformer to get audio features from a log-Mel spectrogram input, and a RoBERTa model to get text features. Both the text and audio features are then projected to a latent space with identical dimension. The dot product between the projected audio and text features is then used as a similar score.
CLAP] Fix few broken things by @younesbelkada in #21670GPTSAN is a Japanese language model using Switch Transformer. It has the same structure as the model introduced as Prefix LM in the T5 paper, and support both Text Generation and Masked Language Modeling tasks. These basic tasks similarly can fine-tune for translation or summarization.
EfficientNets are a family of image classification models, which achieve state-of-the-art accuracy, yet being an order-of-magnitude smaller and faster than previous models.
ALIGN is a multi-modal vision and language model. It can be used for image-text similarity and for zero-shot image classification. ALIGN features a dual-encoder architecture with EfficientNet as its vision encoder and BERT as its text encoder, and learns to align visual and text representations with contrastive learning. Unlike previous work, ALIGN leverages a massive noisy dataset and shows that the scale of the corpus can be used to achieve SOTA representations with a simple recipe.
Informer is a method to be applied to long-sequence time-series forecasting. This method introduces a Probabilistic Attention mechanism to select the “active” queries rather than the “lazy” queries and provides a sparse Transformer thus mitigating the quadratic compute and memory requirements of vanilla attention.
safetensors is a safe format of serialization of tensors, which has been supported in transformers as a first-class citizen for the past few versions.
This change enables explicitly forcing the from_pretrained method to use or not to use safetensors. This unlocks a few use-cases, notably the possibility to enforce loading only from this format, limiting security risks.
Example of usage:
from transformers import AutoModel
# As of version v4.27.0, this loads the `pytorch_model.bin` by default if `safetensors` is not installed.
# It loads the `model.safetensors` file if `safetensors` is installed.
model = AutoModel.from_pretrained('bert-base-cased')
# This forces the load from the `model.safetensors` file.
model = AutoModel.from_pretrained('bert-base-cased', use_safetensors=True)
# This forces the load from the `pytorch_model.bin` file.
model = AutoModel.from_pretrained('bert-base-cased', use_safetensors=False)
This PR adds a "variant" keyword argument to PyTorch's from_pretrained and save_pretrained so that multiple weight variants can be saved in the model repo.
Example of usage with the model hosted in this folder on the Hub:
from transformers import CLIPTextModel
path = "huggingface/the-no-branch-repo" # or ./text_encoder if local
# Loads the `no_ema` variant. This loads the `pytorch_model.fp16.bin` file from this folder.
model = CLIPTextModel.from_pretrained(path, subfolder="text_encoder", variant="fp16")
# This loads the no-variant checkpoint, loading the `pytorch_model.bin` file from this folder.
model = CLIPTextModel.from_pretrained(path, subfolder="text_encoder")
The bitsandbytes integration is overhauled, now offering a new configuration: the BytsandbytesConfig.
Read more about it in the documentation.
bnb] Introducing BitsAndBytesConfig by @younesbelkada in #21579bnb] fix bnb decoders bug by @younesbelkada in #21688This PR enables the user to make use of the PyTorch/XLA implementation of FSDP, including the newly added auto-wrap feature. Four arguments have been added to training_args.py to facilitate this functionality:
xla_fsdp: this flag is a string containing the location of a .json file which specifies the FSDP arguments the user wants to use when wrapping their model.xla_fsdp_min_num_params: this flag is an int which will set a size-based automatic wrapping policy which automatically FSDP wraps any module with at least xla_fsdp_min_num_params many parameters.xla_fsdp_transformer_layer_cls_to_wrap: this flag is a list of (case-sensitive) strings which will set a layer-class-based automatic wrapping policy which automatically FSDP wraps any module whose name matches one of the listed strings.xla_fsdp_grad_ckpt: this flag is a bool which determines whether gradient checkpointing is enabled for the automatically wrapped layers.Generate
This PR standardizes beam search behavior across all three frameworks through early_stopping. PyTorch is unchanged, but TensorFlow and Flax users will see a significant speedup if they keep the default generation parameters.
There are, however, minor differences in outputs of the .generate method with beam search on TensorFlow and Flax. It should be very small and will come with significant speedups, but in case it breaks your workflow, we recommend you downgrade to a previous version and let us know in a GitHub issue so that we may investigate what is going on.
Single model initialization
Model initialization has problems which led to the initialization being incoherent across models and across initialization techniques. This is technically a bugfix, but as it may result in your models being initialized with different values, we think it best to highlight it here.
This PR deprecated the parallelize API which has been replaced by accelerate months ago. We recommend loading the model using the device_map attribute and setting it to balanced to obtain the previous behavior.
Setting your own device_map is still permitted, but it needs to be a dictionary from module name to device, for example:
device_map = {'h.0': 0, 'h.1': 1, ...}
A new pipeline focused on zero-shot audio classification is added to the repository.
The task and model summaries have been refactored to take into account the larger number of tasks and models we now have.
t5] Fix T5 inference in float16 + bnb error by @younesbelkada in #21281TrainingArguments.label_names docs to reflect the correct default value behaviour by @fredtcaroli in #21288ImageProcessor in place of FeatureExtractor for pipelines by @Narsil in #20851oneformer. by @Narsil in #21292EfficientFormer by @ydshieh in #21294OneFormerModelIntegrationTest expected values by @ydshieh in #21295Blenderbot doctest by @younesbelkada in #21297past in prepare inputs for generation by @ArthurZucker in #21296model_class.__name__ and compare against XXX_MAPPING_NAMES by @ydshieh in #21304utils/documentation_tests.txt by @ydshieh in #21315TFEncoderDecoder tests by @ydshieh in #21301compute_transition_scores examples by @gante in #21323Perceiver doctest by @younesbelkada in #21318RobertaPreLayerNorm doctest by @ydshieh in #21337GitModelIntegrationTest.test_batched_generation device issue by @ydshieh in #21362max_length and max_new_tokens coexistence by @gante in #21347run_(clm|mlm).py examples] add streaming dataset support by @stas00 in #21343layer_norm_eps in some models by @ydshieh in #21336max_position_embeddings or max_target_positions by @gante in #21389Graphormer and fix its torchscript test failures by @ydshieh in #21380inputs_embeds by @gante in #214051.13.1 in push/schedule CI by @ydshieh in #21421bnb] Fine-tuning HF 8-bit models by @younesbelkada in #21290is_flaky by @ydshieh in #21426inputs_embeds support for .generate() with BLOOM models by @akreal in #21430ConvBertModelTest test by @ydshieh in #21438SpeechT5ForSpeechToSpeechIntegrationTests device issue by @ydshieh in #21460PushToHubCallback import in Share a model docs by @ireneisdoomed in #21457more_itertools dependency. by @Narsil in #21473prepare_inputs_for_generation by @gante in #21477past in favor of pat_key_values by @ArthurZucker in #21443Doc] Fix int8 docs by @younesbelkada in #21487GPT2TokenizerFast to the list of tokenizer to use for OPT. by @ArthurZucker in #20823compute_transition_scores by @gante in #21341report_to none by @stas00 in #21505image_processor in pipeline. by @Narsil in #21513eos_token_ids in model.generate(...) by @tokestermw in #21461__len__ method to _LazyAutoMapping by @ydshieh in #21522.generate() signature == PT .generate() signature by @gante in #21525.generate() can now be exported with dynamic length by @gante in #21474pipeline] A simple fix for half-precision & 8bit models by @younesbelkada in #21479torch_dtype="auto" to look up config.torch_dtype first, expand docs by @stas00 in #21524config.hidden_size by @stas00 in #21504Blip2] Add int8 support for blip2-flan-t5-xxl by @younesbelkada in #21574inputs_embeds support when generating with GPT-J by @dimitry12 in #21575test_constrained_beam_search_generate_dict_output by @gante in #21561bnb] Let's make the daily CI green 🍏 by @younesbelkada in #21597requires_grad on input embedding to train on top of frozen layers by @younesbelkada in #21598max_length is reached." from InfNaNLogitsProcessor documentation by @mmcdermott in #21634ImageProcessor] Refactor default mean & std to OPENAI_CLIP_MEAN & OPENAI_CLIP_STD by @younesbelkada in #21425BLIP] update blip path on slow tests by @younesbelkada in #21476PROCESSOR_MAPPING_NAMES and add tests by @ydshieh in #21703get_class_in_module by @ydshieh in #21709MBart] Fix cross attention mask check by @younesbelkada in #21730gptsan_japanese from doctest list to avoid GPU OOM by @ydshieh in #21722BigBirdForQuestionAnswering by @ydshieh in #21723ErnieMEmbeddings device issue by @ydshieh in #21726GPTSanJapaneseModel by @ydshieh in #21731GPTNeo] Fix gradient checkpointing bug by @younesbelkada in #21733max_length and num_beams by @bofenghuang in #21740concrete_args from outside available by @lygztq in #21775tests] add accelerate marker by @younesbelkada in #21743PerceiverFourierPositionEncoding with fp16 by @fxmarty in #21787ruff==0.0.253 by @ydshieh in #21828logger.warning_once and use it for grad checkpointing code by @stas00 in #21804MobileViTModelTest to TFMobileViTModelTest by @ydshieh in #21825T5] Fix torchquant issue by @younesbelkada in #21843Blip2] Add Blip2Model by @younesbelkada in #21817Blip2] Fix Blip-2 multi gpu by @younesbelkada in #21707PipelineTestCaseMeta 🚀 by @ydshieh in #21516Blip] Fix blip doctest by @younesbelkada in #21868test_load_default_pipelines_pt for ClapModel by @ydshieh in #21886inputs_embeds functionality when generating with BioGPT by @sidkiblawi in #21889d_kv by @ArthurZucker in #21896BridgeTowerModelTest by @ydshieh in #21908repo_utils_job by @ydshieh in #21928AlignModelTest tests by @ydshieh in #21923check_repo.py due to missing backends by @ydshieh in #21930XLMProphetNetModelIntegrationTest by @ydshieh in #21957torch.allclose for some tests by @ydshieh in #21966test_xglm_sample by @ydshieh in #21975Jukebox tests by @ydshieh in #21984notification_service.py by @ydshieh in #21992test_multi_gpu_data_parallel_forward for some model tests by @ydshieh in #21991AudioClassificationPipelineTests::test_small_model_pt for PT 2.0.0 by @ydshieh in #22023bnb] Fix bnb error message by @younesbelkada in #22026text_config_dict and vision_config_dict being saved for CLIP-like models by @ydshieh in #22035BridgeTower tests slow for now by @ydshieh in #22039huggingface_hub warnings in CI report by @ydshieh in #22054image_processing_donut to match code by @vermouthmjl in #22033Blip2] skip accelerate test by @younesbelkada in #22124is_pipeline_test_to_skip to specific model test classes by @ydshieh in #21999--optim adamw_torch_fused for pt-2.0+ by @stas00 in #22144The following contributors have made significant changes to the library over the last release:
ESM openfold_utils type hints by @ringohoffman in #20544
In the vision integration, all feature extractor classes have been deprecated to be renamed to ImageProcessor. The old feature extractors will be full…
GenerationConfigThe generate method has multiple arguments whose defaults were lying in the model config. We have now decoupled these in a separate generation config, which makes it easier to store different sets of parameters for a given model, with different generation strategies. While we will keep supporting generate arguments in the model configuration for the foreseeable future, it is now recommended to use a generation config. You can learn more about its uses here and its documentation here.
GenerationConfig as the basis for .generate() parametrization by @gante in #20388GenerationConfig as the basis for .generate() parametrization by @gante in #20994GenerationConfig as the basis for .generate() parametrization by @gante in #21007ImageProcessorIn the vision integration, all feature extractor classes have been deprecated to be renamed to ImageProcessor. The old feature extractors will be fully removed in version 5 of Transformers and new vision models will only implement the ImageProcessor class, so be sure to switch your code to this new name sooner rather than later!
AltCLIP is a variant of CLIP obtained by switching the text encoder with a pretrained multilingual text encoder (XLM-Roberta). It has very close performances with CLIP on almost all tasks, and extends the original CLIP’s capabilities to multilingual understanding.
BLIP is a model that is able to perform various multi-modal tasks including visual question answering, image-text retrieval (image-text matching) and image captioning.
BioGPT is a domain-specific generative pre-trained Transformer language model for biomedical text generation and mining. BioGPT follows the Transformer language model backbone, and is pre-trained on 15M PubMed abstracts from scratch.
BiT is a simple recipe for scaling up pre-training of ResNet-like architectures (specifically, ResNetv2). The method results in significant improvements for transfer learning.
EfficientFormer proposes a dimension-consistent pure transformer that can be run on mobile devices for dense prediction tasks like image classification, object detection and semantic segmentation.
GIT is a decoder-only Transformer that leverages CLIP’s vision encoder to condition the model on vision inputs besides text. The model obtains state-of-the-art results on image captioning and visual question answering benchmarks.
GPT-Sw3 is a collection of large decoder-only pretrained transformer language models that were developed by AI Sweden in collaboration with RISE and the WASP WARA for Media and Language. GPT-Sw3 has been trained on a dataset containing 320B tokens in Swedish, Norwegian, Danish, Icelandic, English, and programming code. The model was pretrained using a causal language modeling (CLM) objective utilizing the NeMo Megatron GPT implementation.
Graphormer is a Graph Transformer model, modified to allow computations on graphs instead of text sequences by generating embeddings and features of interest during preprocessign and collation, then using a modified attention.
Mask2Former is a unified framework for panoptic, instance and semantic segmentation and features significant performance and efficiency improvements over MaskFormer.
OneFormer is a universal image segmentation framework that can be trained on a single panoptic dataset to perform semantic, instance, and panoptic segmentation tasks. OneFormer uses a task token to condition the model on the task in focus, making the architecture task-guided for training, and task-dynamic for inference.
The RoBERTa-PreLayerNorm model is identical to RoBERTa but uses the --encoder-normalize-before flag in fairseq.
Swin2R improves the SwinIR model by incorporating Swin Transformer v2 layers which mitigates issues such as training instability, resolution gaps between pre-training and fine-tuning, and hunger on data.
TimeSformer is the first video transformer. It inspired many transformer based video understanding and classification papers.
UPerNet is a general framework to effectively segment a wide range of concepts from images, leveraging any vision backbone like ConvNeXt or Swin.
ViT hybrid is a slight variant of the plain Vision Transformer, by leveraging a convolutional backbone (specifically, BiT) whose features are used as initial “tokens” for the Transformer. It’s the first architecture that attains similar results to familiar convolutional architectures.
Breaking a bit the one model per file policy, we introduce backbones (mainly for vision models) which can then be re-used in more complex models like DETR, MaskFormer, Mask2Former etc.
BeitDropPath layers by @younesbelkada in #20587natten with CUDA version by @ydshieh in #20546FEATURE_EXTRACTOR_MAPPING_NAMES by @ydshieh in #20551require_torch to 2 pipeline tests by @ydshieh in #20585tensorflow_probability for TF pipeline CI by @ydshieh in #20586classifier_dropout in config by @ydshieh in #20596set-output by $GITHUB_OUTPUT by @ydshieh in #20547.to function for ImageProcessors by @younesbelkada in #20536AutomaticSpeechRecognitionPipelineTests.run_pipeline_test by @ydshieh in #20597natten installation in docker file by @ydshieh in #206328bitmodels by @younesbelkada in #20651ViTHybrid] + [BiT] cleaner __init__ by @younesbelkada in #20649run_pipeline_test by @ydshieh in #20623dpt-hybrid support by @younesbelkada in #20645BiT] Small patch fix by @younesbelkada in #20657BackboneMixin by @ydshieh in #20660ViTHybrid] Fix accelerate slow tests by @younesbelkada in #20679test_tokenization_led by @IMvision12 in #20568test_multi_gpu_data_parallel_forward for MaskFormerSwinModelTest by @ydshieh in #20688ViTHybrid] fix last accelerate slow test by @younesbelkada in #20705accelerate support for LongT5 models by @pszemraj in #20341AutoModelTest.test_model_from_pretrained by @ydshieh in #20730layoutlm_job to exotic_models_job by @ydshieh in #20736keep_in_fp32_modules support by @younesbelkada in #20683torch_tensorrt in DeepSpeed CI image for now by @ydshieh in #20758() in some usage of is_flaky by @ydshieh in #20749torch-tensorrt 1.3.0 for DeepSpeed CI by @ydshieh in #20764pipeline test by @younesbelkada in #20778layoutlm by @Narsil in #20776apex in DeepSpeed CI image by @ydshieh in #20788IMAGE_PROCESSOR_MAPPING by @younesbelkada in #20790sentencepiece in DeepSpeed CI image by @ydshieh in #20795Vision] [Refactor] Initialize weights on the correct place by @younesbelkada in #20803max_position_embeddings in config classes by @ydshieh in #20836use_cache in config classes by @ydshieh in #20844use_fast parameter in docstring by @stevhliu in #20840config.num_channels in CLIP-like modeling files by @ydshieh in #20857evaluate to the list of libraries required in generated notebooks by @MKhalusova in #20850LevitModelTest.test_problem_types by @ydshieh in #20859HubertModelIntegrationTest.test_inference_keyword_spotting by @ydshieh in #20863FSMT] Make it compatible with xxxForConditionalGeneration models by @younesbelkada in #20825MobileNet-v2] Fix ONNX typo by @younesbelkada in #20860fp16 for asr pipeline. by @Narsil in #20864T5] fix fp16 loading issue by @younesbelkada in #20878WhisperFeatureExtractor by @bofenghuang in #20936distributed_concat] ensure all_gather's inputs are contiguous by @stas00 in #20951AutomaticSpeechRecognitionPipeline by @bofenghuang in #20952_reorder_cache by @gante in #20964MinNewTokensLengthLogitsProcessor for .generate method #20814 by @kotikkonstantin in #20892decoder_attention_mask in generate function by @samuelpullely in #20726BLIP] Fix daily CI failing test by @younesbelkada in #20877past with past_key_values by @ArthurZucker in #20944documentation_tests.txt by @ydshieh in #21036torchscript tests for AltCLIP by @ydshieh in #21102min_new_tokens argument in generate() (implementation based on MinNewTokensLengthLogitsProcessor) by @silverriver in #21044RealmModelIntegrationTest.test_inference_open_qa by @ydshieh in #21136TFTapasEmbeddings by @ydshieh in #21107use_cache from model_kwargs by @gante in #21149installation.mdx to Korean by @wonhyeongseo in #20948test_save_pretrained_signatures slow test by @ydshieh in #21105blip support for training by @younesbelkada in #21021Mask2FormerForUniversalSegmentation by @ydshieh in #21175UperNetModelIntegrationTest by @ydshieh in #21192CVT] Fix module initialization issue by @younesbelkada in #21193automatic-speech-recognition asr for Whisper. by @Narsil in #21196huggingface_hub version by @ydshieh in #21212CONFIG_ARCHIVE_MAP_MAPPING_NAMES by @ydshieh in #21207GPTJ doctest by @ydshieh in #21213parallelism for CircleCI jobs work - but keep it 1 for now by @ydshieh in #21157BLIP] fix docstring for BlipTextxxx by @younesbelkada in #21224BLIP] fix doctest by @younesbelkada in #21217.save_pretrained() by @gante in #21264The following contributors have made significant changes to the library over the last release:
[Proposal] Breaking change zero-shot-object-detection for improved consistency. by @Narsil in #20280
We are very excited by the newly announced PyTorch 2.0 stack. You can enable torch.compile on any of our models, and get support with the Trainer (and in all our PyTorch examples) by using the torchdynamo training argument. For instance, just add --torchdynamo inductor when launching those examples from the command line.
This API is still experimental and may be subject to changes as the PyTorch 2.0 stack matures.
Note that to get the best performance, we recommend:
--pad_to_max_length in our examples)The Audio Spectrogram Transformer model was proposed in AST: Audio Spectrogram Transformer by Yuan Gong, Yu-An Chung, James Glass. The Audio Spectrogram Transformer applies a Vision Transformer to audio, by turning audio into an image (spectrogram). The model obtains state-of-the-art results for audio classification.
The Jukebox model was proposed in Jukebox: A generative model for music by Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, Ilya Sutskever. It introduces a generative music model which can produce minute long samples that can be conditionned on an artist, genres and lyrics.
The SwitchTransformers model was proposed in Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity by William Fedus, Barret Zoph, Noam Shazeer.
It is the first MoE model supported in transformers, with the largest checkpoint currently available currently containing 1T parameters.
The RoCBert model was proposed in RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining by HuiSu, WeiweiShi, XiaoyuShen, XiaoZhou, TuoJi, JiaruiFang, JieZhou. It’s a pretrained Chinese language model that is robust under various forms of adversarial attacks.
The CLIPSeg model was proposed in Image Segmentation Using Text and Image Prompts by Timo Lüddecke and Alexander Ecker. CLIPSeg adds a minimal decoder on top of a frozen CLIP model for zero- and one-shot image segmentation.
NAT was proposed in Neighborhood Attention Transformer by Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi.
It is a hierarchical vision transformer based on Neighborhood Attention, a sliding-window self attention pattern.
DiNAT was proposed in Dilated Neighborhood Attention Transformer by Ali Hassani and Humphrey Shi.
It extends NAT by adding a Dilated Neighborhood Attention pattern to capture global context, and shows significant performance improvements over it.
The MobileNet model was proposed in MobileNetV2: Inverted Residuals and Linear Bottlenecks by Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, Liang-Chieh Chen.
The MobileNet model was proposed in MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications by Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, Hartwig Adam.
Image processors replace feature extractors as the processing class for computer vision models.
Important changes:
size parameter is now a dictionary of {"height": h, "width": w}, {"shortest_edge": s}, {"shortest_egde": s, "longest_edge": l} instead of int or tuple.data_format flag. You can now specify if you want your images to be returned in "channels_first" - NCHW - or "channels_last" - NHWC - format.do_resize can be passed directly to the preprocess method instead of modifying the class attribute: image_processor([image_1, image_2], do_resize=False, return_tensors="pt", data_format="channels_last")return_tensors unset will return a list of numpy arrays.The classes are backwards compatible and can be created using existing feature extractor configurations - with the size parameter converted.
We're adding support for a general AutoBackbone class, which turns any vision model (like ConvNeXt, Swin Transformer) into a backbone to be used with frameworks like DETR and Mask R-CNN. The design is in early stages and we welcome feedback.
safetensors offloadingIf the model you are using has a safetensors checkpoint and you have the library installed, offload to disk will take advantage of this to be more memory efficient and roughly 33% faster.
generate methodconvert_tokens_to_string by @beneyal in #15775tokenizer_type to avoid error when loading checkpoint back by @pacman100 in #20062VisionTextDualEncoderProcessorTest by @ydshieh in #20098generate_dummy_inputs for ImageGPTOnnxConfig by @ydshieh in #20103CLIPSegModelTester by @ydshieh in #20134_keys_to_ignore. by @Narsil in #20042RoCBertTokenizer to TOKENIZER_MAPPING_NAMES by @ydshieh in #20141object-detection. by @Narsil in #20143OnnxConfig.generate_dummy_inputs to check ImageProcessingMixin by @ydshieh in #20157inspect in ASR pipeline and make WhisperEncoder just nice to use. by @Narsil in #19571test_save_load_fast_init_from_base as is_flaky by @ydshieh in #20200ImageSegmentationPipelineTests less flaky by @ydshieh in #20147accelerate support for ViT family by @younesbelkada in #20174authorized_missing_keysin favor of _keys_to_ignore_on_load_missing by @ArthurZucker in #20228run_clip.py by @ydshieh in #20234audio-classification example in the doc. by @Narsil in #20235fill-mask pipeline. by @Narsil in #20241feature-extraction. by @Narsil in #20240depth-estimation pipeline. by @Narsil in #20237table-question-answering pipeline. by @Narsil in #20260image-segmentation pipeline. by @Narsil in #20256text2text-generation pipeline. by @Narsil in #20261text-generation pipeline. by @Narsil in #20264image-classification pipeline. by @Narsil in #20254zero-shot-image-classification pipeline. by @Narsil in #20272zero-shot-classification pipeline. by @Narsil in #20268visual-question-answering pipeline. by @Narsil in #20266text-classification pipeline. by @Narsil in #20262question-answering pipeline. by @Narsil in #20259image-to-text pipeline. by @Narsil in #20257token-classification pipeline. by @Narsil in #20265zero-shot-object-detection pipeline doctest. by @Narsil in #20274object-detection pipeline. by @Narsil in #20258PushToHubCallback by @gante in #20231ImageProcessor by @ydshieh in #20298zero-shot-object-detection for improved consistency. by @Narsil in #20280model_kwargs can also be an input to prepare_inputs_for_generation by @gante in #20353keys_to_ignore for M2M100 by @younesbelkada in #20381accelerate support for ESM by @younesbelkada in #20379accelerate tests for esmfold by @younesbelkada in #20387bnb bug by @younesbelkada in #20408ValueError when trying to cast or assign by @younesbelkada in #20409model_max_length when saving tokenizers by @ydshieh in #20401accelerate support for OwlViT by @younesbelkada in #20411word_to_tokens docstring format by @SaulLu in #20450CLIPSegModelIntegrationTest by @ydshieh in #20467contrastive_loss by @ydshieh in #20455attention_mask truncation in whisper by @ydshieh in #20488add_special_tokens more clear by @ydshieh in #20424galactica models by @younesbelkada in #20390AutomaticSpeechRecognitionPipeline doc example by @ydshieh in #20512natten for CI by @ydshieh in #20511PLBart doctest by @ydshieh in #20527ConditionalDetrForSegmentation doc example by @ydshieh in #20531ZeroShotObjectDetectionPipeline doc example by @ydshieh in #20528The following contributors have made significant changes to the library over the last release:
Nothing published for this version
The following changes are bugfixes that we have chosen to fix even if it changes the resulting behavior. We mark them as breaking changes, so if you a…
ESM-2 and ESMFold are new state-of-the-art Transformer protein language and folding models from Meta AI's Fundamental AI Research Team (FAIR). ESM-2 is trained with a masked language modeling objective, and it can be easily transferred to sequence and token classification tasks for proteins. Checkpoints exist in various sizes, from 8 million parameters up to a huge 15 billion parameter model.
ESMFold is a state-of-the-art single sequence protein folding model which produces high accuracy predictions significantly faster. Unlike previous protein folding tools like AlphaFold2 and openfold, ESMFold uses a pretrained protein language model to generate token embeddings that are used as input to the folding model, and so does not require a multiple sequence alignment (MSA) of related proteins as input. As a result, proteins can be folded in a single forward pass of the model without requiring any external databases or search/alignment tools to be present at inference time. This hugely reduces the time and compute requirements for folding.
Transformer protein language models were introduced in the paper Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences by Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus.
ESMFold was introduced in the paper Language models of protein sequences at the scale of evolution enable accurate structure prediction by Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, and Alexander Rives.
LiLT allows to combine any pre-trained RoBERTa text encoder with a lightweight Layout Transformer, to enable LayoutLM-like document understanding for many languages.
It was proposed in LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding by Jiapeng Wang, Lianwen Jin, Kai Ding.
FLAN-T5 is an enhanced version of T5 that has been finetuned on a mixture of tasks.
It was released in the paper Scaling Instruction-Finetuned Language Models by Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei.
flan-t5 documentation page by @younesbelkada in #19892Table Transformer is a model that can perform table extraction and table structure recognition from unstructured documents based on the DETR architecture.
It was proposed in PubTables-1M: Towards comprehensive table extraction from unstructured documents by Brandon Smock, Rohith Pesala, Robin Abraham.
Contrastive search decoding is a new state-of-the-art generation method which aims at reducing the repetitive patterns in which generation models often fall.
It was introduced in A Contrastive Framework for Neural Text Generation by Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, Nigel Collier.
We continue to explore the new serialization format not using Pickle via the safetensors library, this time by adding support for TensorFlow models. More checkpoints have been converted to this format. Support is still experimental.
The following changes are bugfixes that we have chosen to fix even if it changes the resulting behavior. We mark them as breaking changes, so if you are using this part of the codebase, we recommend you take a look at the PRs to understand what changes were done exactly.
TFWrappedEmbeddings (breaking: TF embedding initialization updated for encoder-decoder models) by @gante in #19263pipeline by @ArthurZucker in #19482nested_XXX functions to mappings/dicts. by @Guillem96 in #19455TFGroupViT CI by @ydshieh in #19461DeiT and TFGroupViT by @ydshieh in #19466WhisperModelIntegrationTests.test_large_batched_generation by @ydshieh in #19472XLMRoberta model and config independent from Roberta by @asofiaoliveira in #19359get_embedding dtype at init. time by @ydshieh in #19473XLMProphet model from Prophet by @srhrshr in #19406generate & device_map=auto & half precision models by @younesbelkada in #19468python3 instead of python in push CI setup job by @ydshieh in #19492OPTForQuestionAnswering doctest by @ydshieh in #19479configuration_bert.py to doctest by @ydshieh in #19485MobileBert tokenizers independent from Bert by @501Good in #19531MarkupLMForMaskedLM from MODEL_WITH_LM_HEAD_MAPPING_NAMES by @ydshieh in #19534configuration_yolos.py by @daspartho in #19539Add configuration_whisper.py by @daspartho in #19540getattribute_from_module can't find anything by @ydshieh in #19535MarkupLMConfig by @ydshieh in #19547configuration_vit.py by @daspartho in #19561ViTMAE and YOSO by @grgkaran03 in #19567DebertaV2ForMultipleChoice Pytorch by @IMvision12 in #19536configuration_blenderbot.py by @grgkaran03 in #19577configuration_blenderbot_small.py by @grgkaran03 in #19589configuration_bigbird_pegasus.py and configuration_big_bird.py by @Xabilahu in #19606test_tf_encode_plus_sent_to_model for TAPAS by @ydshieh in #19559ImageToTextPipelineTests.test_small_model_tf by @ydshieh in #19565FlaubertTokenizer by @ydshieh in #19552configuration_resnet.py by @daspartho in #19620. in layer name by @ArthurZucker in #19124configuration_data2vec_text.py by @daspartho in #19636configuration_trocr.py by @thliang01 in #19658DocumentQuestionAnsweringPipeline by @ankrgyl in #19584facebook/ by @Rocketknight1 in #19675configuration_data2vec_vision.py by @daspartho in #19637VisualBertConfig doc example by @ydshieh in #19692image-segmentation pipeline tests. by @Narsil in #19710configuration_pegasus_x.py by @mukesh663 in #19725accelerate support for Whisper by @younesbelkada in #19697FlavaConfig and FNetConfig by @ndrohith09 in #19724configuration_clip.py by @daspartho in #19647configuration_wavlm.py by @juancopi81 in #19749configuration_decision_transformer.py by @Xabilahu in #19751configuration_detr.py by @Xabilahu in #19752image-segmentation pipeline: re-enable small_model_pt test. by @Narsil in #19716ImageToTextPipelineTests.test_small_model_tf by @ydshieh in #19785test_torchscrip_xxx CI by updating _create_and_check_torchscript by @ydshieh in #19786MaskFormerConfig doctest by @sha016 in #19817configuration_plbart.py by @ayaka14732 in #19809configuration_poolformer.py by @ayaka14732 in #19808configuration_electra.py by @ayaka14732 in #19807configuration_nezha.py by @ayaka14732 in #19810LEDModelIntegrationTests expected values by @ydshieh in #19841MarkupLM by @ydshieh in #19845GenerationMixin.contrastive_search by @ydshieh in #19863image-segmentation pipeline. by @Narsil in #19727MCTCT, MBart and LayoutLM by @Revanth2002 in #19889max_diff in test_save_load_fast_init_to_base by @ydshieh in #19849accelerate support for RoBERTa family by @younesbelkada in #19906accelerate support for M2M100 by @younesbelkada in #19912accelerate support for BART-like models by @younesbelkada in #19927The following contributors have made significant changes to the library over the last release:
XLMRoberta model and config independent from Roberta (#19359)XLMProphet model from Prophet (#19406)DebertaV2ForMultipleChoice Pytorch (#19536)MobileBert tokenizers independent from Bert (#19531)configuration_pegasus_x.py (#19725)Fix a revert introduced by mistake making the "automatic-speech-recognition" for Whisper.
Fix a revert introduced by mistake making the "automatic-speech-recognition" for Whisper.
Update Protobuf dependency version to fix known vulnerability by @qthequartermasterman in #19247
The Whisper model was proposed in Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever.
Whisper is an encoder-decoder Transformer trained on 680,000 hours of labeled (transcribed) audio. The model shows impressive performance and robustness in a zero-shot setting, in multiple languages.
The Deformable DETR model was proposed in Deformable DETR: Deformable Transformers for End-to-End Object Detection by Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, Jifeng Dai.
Deformable DETR mitigates the slow convergence issues and limited feature spatial resolution of the original DETR by leveraging a new deformable attention module which only attends to a small set of key sampling points around a reference.
The Conditional DETR model was proposed in Conditional DETR for Fast Training Convergence by Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, Jingdong Wang.
Conditional DETR presents a conditional cross-attention mechanism for fast DETR training. Conditional DETR converges 6.7× to 10× faster than DETR.
The Time Series Transformer model is a vanilla encoder-decoder Transformer for time series forecasting.
The model is trained in a similar way to how one would train an encoder-decoder Transformer (like T5 or BART) for machine translation; i.e. teacher forcing is used. At inference time, one can autoregressively generate samples, one time step at a time.
:warning: This is a recently introduced model and modality, so the API hasn't been tested extensively. There may be some bugs or slight breaking changes to fix it in the future. If you see something strange, file a Github Issue.
The ViTMSN model was proposed in Masked Siamese Networks for Label-Efficient Learning by Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, Nicolas Ballas.
MSN (masked siamese networks) consists of a joint-embedding architecture to match the prototypes of masked patches with that of the unmasked patches. With this setup, the method yields excellent performance in the low-shot and extreme low-shot regimes for image classification, outperforming other self-supervised methods such as DINO. For instance, with 1% of ImageNet-1K labels, the method achieves 75.7% top-1 accuracy.
The MarkupLM model was proposed in MarkupLM: Pre-training of Text and Markup Language for Visually-rich Document Understanding by Junlong Li, Yiheng Xu, Lei Cui, Furu Wei.
MarkupLM is BERT, but applied to HTML pages instead of raw text documents. The model incorporates additional embedding layers to improve performance, similar to LayoutLM.
The model can be used for tasks like question answering on web pages or information extraction from web pages. It obtains state-of-the-art results on 2 important benchmarks: WebSRC and SWDE.
We explore a new serialization format not using Pickle that we can then leverage in the three frameworks we support: PyTorch, TensorFlow, and JAX. We leverage the safetensors library for that.
Support is for PyTorch models only at this stage, and still experimental.
The processors for computer vision have been overhauled to ensure they have consistent naming, input arguments and outputs.
:warning: The existing methods that are superseded by the introduced methods post_process_object_detection, post_process_semantic_segmentation, post_process_instance_segmentation, post_process_panoptic_segmentation are now deprecated.
The following changes are bugfixes that we have chosen to fix even if it changes the resulting behavior. We mark them as breaking changes, so if you are using this part of the codebase, we recommend you take a look at the PRs to understand what changes were done exactly..
Breaking change for ViT parameter initialization
Breaking change for the top_p argument of the TopPLogitsWarper of the generate method.
OPT and BLOOM now have question answering heads available.
OPTForQuestionAnswering by @clementapa in #19402BloomForQuestionAnswering by @younesbelkada in #19310There is now a zero-shot object detection pipeline.
The GroupViT model is now available in TensorFlow.
test_save_load for TFViTMAEModelTest by @ydshieh in #19040torchdynamo tests by @ydshieh in #19056use_cache by @younesbelkada in #19060LeViT checkpoint by @ydshieh in #19069test_export_to_onnx for LongT5 if torch < 1.11 by @ydshieh in #19122max_eval_samples by @lvwerra in #18722tokenizers release ! by @Narsil in #19139accelerate support for ViLT by @younesbelkada in #18683assertAlmostEqual in BloomEmbeddingTest.test_logits by @ydshieh in #19200cur_len in generation_utils.py by @ekagra-ranjan in #18874math.pi instead of torch.pi in MaskFormer by @ydshieh in #19201TFDeiTForImageClassification by @ydshieh in #19173m2m_100.mdx doc example missing labels by @Mustapha-AJEGHRIR in #19149hf_raise_for_status instead of deprecated _raise_for_status by @Wauplin in #19244HFValidationError in TrainingSummary by @ydshieh in #19252ViTMSNForImageClassification by @sayakpaul in #19183beautifulsoup4 to the dependency list by @ydshieh in #19253BloomConfig docstring by @younesbelkada in #19336ConvBert Tokenizer independent from bert Tokenizer by @IMvision12 in #19347Bert interdependency in tokenization_electra.py by @OtherHorizon in #19356Camembert TF version independent from Roberta by @Mustapha-AJEGHRIR in #19364ViTMSNForImageClassification doctest by @ydshieh in #19275BloomEmbeddingTest.test_embeddings for PyTorch < 1.10 by @ydshieh in #19261add_new_model.mdx by @Steboss89 in #18713The following contributors have made significant changes to the library over the last release:
ViTMSNForImageClassification (#19183)ConvBert Tokenizer independent from bert Tokenizer (#19347)m2m_100.mdx doc example missing labels (#19149)Camembert TF version independent from Roberta (#19364)add_new_model.mdx (#18713)Fixes a bug where a cached tokenizer/model was not accessible anymore offline (either forcing offline mode or because of an internet issue).
Fixes a bug where a cached tokenizer/model was not accessible anymore offline (either forcing offline mode or because of an internet issue).
Patch release for the following PRs:
Fix breaking change in onnxruntime for ONNX quantization by @severinsimmler in #18336
The Swin Transformer V2 model was proposed in Swin Transformer V2: Scaling Up Capacity and Resolution by Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, Baining Guo.
Swin Transformer v2 improves the original Swin Transformer using 3 main techniques: 1) a residual-post-norm method combined with cosine attention to improve training stability; 2) a log-spaced continuous position bias method to effectively transfer models pre-trained using low-resolution images to downstream tasks with high-resolution inputs; 3) A self-supervised pre-training method, SimMIM, to reduce the needs of vast labeled images.
The VideoMAE model was proposed in VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training by Zhan Tong, Yibing Song, Jue Wang, Limin Wang. VideoMAE extends masked auto encoders (MAE) to video, claiming state-of-the-art performance on several video classification benchmarks.
VideoMAE is an extension of ViTMAE for video.
The Donut model was proposed in OCR-free Document Understanding Transformer by Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, Seunghyun Park. Donut consists of an image Transformer encoder and an autoregressive text Transformer decoder to perform document understanding tasks such as document image classification, form understanding and visual question answering.
The PEGASUS-X model was proposed in Investigating Efficiently Extending Transformers for Long Input Summarization by Jason Phang, Yao Zhao and Peter J. Liu.
PEGASUS-X (PEGASUS eXtended) extends the PEGASUS models for long input summarization through additional long input pretraining and using staggered block-local attention with global tokens in the encoder.
The X-CLIP model was proposed in Expanding Language-Image Pretrained Models for General Video Recognition by Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, Haibin Ling. X-CLIP is a minimal extension of CLIP for video. The model consists of a text encoder, a cross-frame vision encoder, a multi-frame integration Transformer, and a video-specific prompt generator.
X-CLIP is a minimal extension of CLIP for video-language understanding.
ERNIE is a series of powerful models proposed by baidu, especially in Chinese tasks, including ERNIE1.0, ERNIE2.0, ERNIE3.0, ERNIE-Gram, ERNIE-health, etc. These models are contributed by nghuyong and the official code can be found in PaddleNLP (in PaddlePaddle).
MobileViT and LayoutLMv3 are now available in TensorFlow.
A new question answering head was added for the LayoutLM model.
Two new pipelines are available in transformers: a document question answering pipeline, as well as an image to text generation pipeline.
There is now Mac M1 support in PyTorch in transformers in pipelines and the Trainer.
pipeline support for device="mps" (or any other string) by @julien-c in #18494mps integration by @pacman100 in #18598Starting from version v4.22.0, we'll now officially support PyTorch and TensorFlow versions that were released up to two years ago. Versions older than two years-old will not be supported going forward.
We're making this change as we begin actively testing transformers compatibility on older versions. This project can be followed here.
The generate method now starts enforcing stronger validation in order to ensure proper usage.
model_kwargs (and catch typos in generate arguments) by @gante in #18261model_kwargs on TF (and catch typos in generate arguments) by @gante in #18651The as_target_tokenizer and as_target_processor context managers have been deprecated. The new API is to use the call method of the tokenizer/processor with keyword arguments. For instance:
with tokenizer.as_target_tokenizer():
encoded_labels = tokenizer(labels, padding=True)
becomes
encoded_labels = tokenizer(text_target=labels, padding=True)
as_target context managers by direct calls by @sgugger in #18325Bits and bytes is now integrated within transformers. This feature can reduce the size of large models by up to 2, with low loss in precision.
bitsandbytes integration by @younesbelkada in #18579bitsandbytes - Linear8bitLt integration into transformers models by @younesbelkada in #17901Models that have sharded checkpoints in PyTorch can be loaded in Flax.
The TensorFlow examples have been rewritten to support all recent features developped in the past months.
DeBERTa-v2 is now trainable with XLA.
position_ids by @thomasw21 in #18342test_load_default_pipelines_tf test error by @ydshieh in #18422generate docstring by @JoaoLages in #18198trust_remote_code and ignore it in PreTrainedModel.from_pretrained by @ydshieh in #18428TF_MODEL_FOR_SEMANTIC_SEGMENTATION_MAPPING by @ydshieh in #18469TFSwinLayer to increase serving compatibility by @harrydrippin in #18352test_dbmdz_english by updating expected values by @ydshieh in #18482quicktour.mdx for resampy 0.3.0 by @ydshieh in #18484transformers-cli login => huggingface-cli login by @julien-c in #18490run_call_with_unpacked_inputs by @ydshieh in #18541align_to_words param to qa pipeline. by @Narsil in #18010load_state_dict by @pacman100 in #18596TFAutoModelForSemanticSegmentation to the main __init__.py by @ydshieh in #18600detectron2 for CircleCI tests by @ydshieh in #18680onnxruntime for ONNX quantization by @severinsimmler in #18336model.tie_weights() should be applied after accelerator.prepare() by @Gladiator07 in #18676**model_kwargs in sample tests by @gante in #18696microsoft/tapex-base-finetuned-wtq by @Narsil in #18711add_tokens docstring by @SaulLu in #18687XGLMModel by @stancld in #16543__call__ method is faster than encode + pad for a fast tokenizer by @SaulLu in #18693torch_fx tests by @ydshieh in #18547test_cached_files_are_used_when_internet_is_down by @Wauplin in #18804HfArgumentParser.parse_{dict,json_file} to raise an Exception when there extra keys by @FelixSchneiderZoom in #18692state_dict to release memory as early as possible by @ydshieh in #18832Seq2SeqTrainer by @kumapo in #18786test_tf_encode_plus_sent_to_model for LayoutLMv3 by @ydshieh in #18898Perplexity of fixed-length models by @ekagra-ranjan in #18906_create_and_check_torch_fx_tracing in specific test files by @ydshieh in #18667decoder_position_ids from check_decoder_model_past_large_inputs by @ydshieh in #18980require_tf for TFOPTGenerationTest by @ydshieh in #19010The following contributors have made significant changes to the library over the last release:
XGLMModel (#16543)Patch release to add a disclaimer about torch models hosted on the Hub. See autoclass tutorial.
Patch release to add a disclaimer about torch models hosted on the Hub. See autoclass tutorial.
Fix a regression in the TableQA pipeline: Fix a regression in Trainer checkpoint loading: #18428
Fix a regression in the TableQA pipeline: Fix a regression in Trainer checkpoint loading: #18428
Fix a regression in Trainer checkpoint loading: #18470
Fix a regression in Trainer checkpoint loading: #18470
Generate: deprecate default max_length by @gante in #18018
The TensorFlow text generation method can now be wrapped with tf.function and compiled to XLA. You should be able to achieve up to 100x speedup this way. See our blog post and our benchmarks. You can also see XLA generation in action in our example notebooks, particularly for summarization and translation.
import tensorflow as tf
from transformers import AutoTokenizer, TFAutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = TFAutoModelForSeq2SeqLM.from_pretrained("t5-small")
# Main changes with respect to the original generate workflow: `tf.function` and `pad_to_multiple_of`
xla_generate = tf.function(model.generate, jit_compile=True)
tokenization_kwargs = {"pad_to_multiple_of": 32, "padding": True, "return_tensors": "tf"}
# The first prompt will be slow (compiling), the others will be very fast!
input_prompts = [
f"translate English to {language}: I have four cats and three dogs."
for language in ["German", "French", "Romanian"]
]
for input_prompt in input_prompts:
tokenized_inputs = tokenizer([input_prompt], **tokenization_kwargs)
generated_text = xla_generate(**tokenized_inputs, max_new_tokens=32)
print(tokenizer.decode(generated_text[0], skip_special_tokens=True))
max_length by @gante in #18018tf.TensorArray by @gante in #17801The OWL-ViT model (short for Vision Transformer for Open-World Localization) was proposed in Simple Open-Vocabulary Object Detection with Vision Transformers by Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, and Neil Houlsby. OWL-ViT is an open-vocabulary object detection network trained on a variety of (image, text) pairs. It can be used to query an image with one or multiple text queries to search for and detect target objects described in text.
The NLLB model was presented in No Language Left Behind: Scaling Human-Centered Machine Translation by Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. No Language Left Behind (NLLB) is a model capable of delivering high-quality translations directly between any pair of 200+ languages — including low-resource languages like Asturian, Luganda, Urdu and more.
The MobileViT model was proposed in MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer by Sachin Mehta and Mohammad Rastegari. MobileViT introduces a new layer that replaces local processing in convolutions with global processing using transformers.
The Nezha model was proposed in NEZHA: Neural Contextualized Representation for Chinese Language Understanding by Junqiu Wei et al. NEZHA is a language model based on BERT with a collection of proven improvements, which include Functional Relative Positional Encoding as an effective positional encoding scheme, Whole Word Masking strategy, Mixed Precision Training and the LAMB Optimizer in training the models.
The GroupViT model was proposed in GroupViT: Semantic Segmentation Emerges from Text Supervision by Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, Xiaolong Wang. Inspired by CLIP, GroupViT is a vision-language model that can perform zero-shot semantic segmentation on any given vocabulary categories, inspired by CLIP.
The MVP model was proposed in MVP: Multi-task Supervised Pre-training for Natural Language Generation by Tianyi Tang, Junyi Li, Wayne Xin Zhao and Ji-Rong Wen. MVP is a generative language model, pre-trained on a labeled pre-training corpus from 45 datasets over seven generation tasks. For each task, the model is further pre-trained using specific soft prompts to stimulate the model capacity in performing a specific task.
The CodeGen model was proposed in A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. CodeGen is an autoregressive language model for program synthesis trained sequentially on The Pile, BigQuery, and BigPython.
The UL2 model was presented in Unifying Language Learning Paradigms by Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, Donald Metzler. UL2 is a unified framework for pretraining models that are universally effective across datasets and setups. UL2 uses Mixture-of-Denoisers (MoD), a pre-training objective that combines diverse pre-training paradigms together. UL2 introduces a notion of mode switching, wherein downstream fine-tuning is associated with specific pre-training schemes.
This adds the ability to support custom pipelines on the Hub and share it with everyone else. Like the code in the Hub feature for models, tokenizers etc., the user has to add trust_remote_code=True when they want to use it. Apart from this, the best way to get familiar with the feature is to look at the added documentation.
This adds a CLI to convert PT weights into TF weights, validate them, and (optionally) open a PR.
The following models have been ported to be used in TensorFlow: SegFormer, DeiT, ResNet and RegNet.
Additionally, our TF models now support loading sharded checkpoints:
The following models have been ported to be used in JAX:
Additionally, our JAX models now support loading sharded checkpoints:
The following models now have a brand new head for new tasks:
A continued community effort provides ONNX converters for an increasing number of models.
A community effort aiming to translate the documentation in several languages has been continued.
KerasMetricCallback to use XLA generation by @Rocketknight1 in #18265--make-reports by @ydshieh in #18250saved_model=True by @amyeroberts in #18153take_along_axis is computed in DeBERTa to stop confusing XLA by @Rocketknight1 in #18256(...)EncoderDecoder models by @gante in #18097no_trainer CI by @muellerzr in #18242LayoutXLM docstrings by @qqaatw in #17038device_map directly in pipeline(..) function. by @Narsil in #17902text-generation pipeline. by @Narsil in #18131pt_to_tf test by @gante in #18108TFGenerationMixin.seed_generator so it's not created at import by @Rocketknight1 in #18044bias keyword argument in TFDebertaEmbeddings by @WissamAntoun in #17940hf-internal-testing) by @mishig25 in #17939return_all_scores introduced in #17606 by @Narsil in #17906group_texts function, drop last block if smaller than block_size by @billray0259 in #17908test_number_of_steps_in_training_with_ipex by @ydshieh in #17889test_inference_instance_segmentation_head by @ydshieh in #17872test_multi_gpu_data_parallel_forward for MaskFormer by @ydshieh in #17864create_commit by @gante in #17755generate, to Seq2SeqTrainer methods evaluate and predict by @eranhirs in #17805top_k_top_p_filtering having unexpected behavior by @unifyh in #17744The following contributors have made significant changes to the library over the last release:
This patch releases fixes a bug in the OPT models and makes Transformers compatible with huggingface_hub version 0.8.1.
This patch releases fixes a bug in the OPT models and makes Transformers compatible with huggingface_hub version 0.8.1.
Fix tests of mixed precision now that experimental is deprecated by @Rocketknight1 in #17300
You can now use the big model inference of Accelerate directly in any call to from_pretrained by specifying device_map="auto" (or your own device_map). It will automatically load the model taking advantage of your GPU(s) then offloading what doesn't fit in RAM, or even on the hard drive if you don't have RAM. Your model can then be used normally for inference without anything else to do.
from transformers import AutoModelForSeq2SeqLM
model = AutoModelForSeq2SeqLM.from_pretrained(
"bigscience/T0pp", revision="sharded", device_map="auto"
)
from_pretrained for big model inference by @sgugger in #17341The BLOOM model has been proposed with its various versions through the BigScience Workshop. The architecture of BLOOM is essentially similar to GPT3 (auto-regressive model for next token prediction), but has been trained on different 46 languages including code.
The Convolutional vision Transformer (CvT) improves the Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both designs.
GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile, whose weights are made freely and openly available to the public through a permissive license. GPT-NeoX-20B is a particularly powerful few-shot reasoner and gains far more in performance when evaluated five-shot than similarly sized GPT-3 and FairSeq models.
LayoutLMv3 simplifies LayoutLMv2 by using patch embeddings (as in ViT) instead of leveraging a CNN backbone, and pre-trains the model on 3 objectives: masked language modeling (MLM), masked image modeling (MIM) and word-patch alignment (WPA).
LeViT improves the Vision Transformer (ViT) in performance and efficiency by a few architectural differences such as activation maps with decreasing resolutions in Transformers and the introduction of an attention bias to integrate positional information.
LongT5 model is an extension of T5 model, and it enables using one of the two different efficient attention mechanisms - (1) Local attention, or (2) Transient-Global attention. It is capable of handling input sequences of a length up to 16,384 tokens.
LongT5 model by @stancld in #16792The M-CTC-T model is a 1B-param transformer encoder, with a CTC head over 8065 character labels and a language identification head over 60 language ID labels. It is trained on Common Voice (version 6.1, December 2020 release) and VoxPopuli. After training on Common Voice and VoxPopuli, the model is trained on Common Voice only. The labels are unnormalized character-level transcripts (punctuation and capitalization are not removed). The model takes as input Mel filterbank features from a 16Khz audio signal.
This Transformer is used for deep reinforcement learning. To use it, you need to create sequences from actions, states and rewards from all previous timesteps. This model will treat all these elements together as one big sequence (a trajectory).
The Wav2Vec2-Conformer is an updated version of fairseq S2T: Fast Speech-to-Text. It requires more parameters than Wav2Vec2, but also yields an improved word error rate.
Data2VecVision for semantic segmentation, OPT and Swin are now available in TensorFlow.
OPT is now available in Flax.
A community effort has been started to translate the documentation in two new languages: Italian and Portuguese.
BloomForSequenceClassification and BloomForTokenClassification classes by @haileyschoelkopf in #17639RepoNotFoundError when not authenticated by @SBrandeis in #17651float16. by @Narsil in #17637top_k argument to text-classification pipeline. by @Narsil in #17606train_new_from_iterator in the case of byte-level tokenizers by @SaulLu in #17549pt-to-tf by @gante in #17588tokenizer type annotation in pipeline(...) by @willfrey in #17500PreTrainedTokenizerBase.add_tokens() by @Witiko in #17119test_inference_no_head by @ydshieh in #17395imageGPT auto feature extractor. by @Narsil in #16871device_map="auto" to OPT by @sgugger in #17382batch_size test to QA pipeline. by @Narsil in #17330max_seq_len in QA pipeline by @Narsil in #17316test_torch_encode_plus_sent_to_model by @SaulLu in #17231The following contributors have made significant changes to the library over the last release:
LongT5 model (#16792)Fixes the errors message when trying to access a repo that does not exist (started to break due to changes in Hub API).
Fixes the errors message when trying to access a repo that does not exist (started to break due to changes in Hub API).
[🐛]Properly raise RepoNotFoundError when not authenticated #17651[
This patch release fixes the install of protobuf when a user wants to do pip install transformers[sentencepiece].
This patch release fixes the install of protobuf when a user wants to do pip install transformers[sentencepiece].
Patch release for the following PRs/commits:
Patch release for the following PRs/commits:
Fix Trainer for Datasets that don't have dict items #17239
Fix Trainer for Datasets that don't have dict items #17239
bert: properly mention deprecation of TF2 conversion script by @stefan-it in #16171
Disclaimer: this release is the first release with no Python 3.6 support.
The OPT model was proposed in Open Pre-trained Transformer Language Models by Meta AI. OPT is a series of open-sourced large causal language models which perform similar in performance to GPT3.
The FLAVA model was proposed in FLAVA: A Foundational Language And Vision Alignment Model by Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela and is accepted at CVPR 2022.
The paper aims at creating a single unified foundation model which can work across vision, language as well as vision-and-language multimodal tasks.
The YOLOS model was proposed in You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection by Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, Wenyu Liu. YOLOS proposes to just leverage the plain Vision Transformer (ViT) for object detection, inspired by DETR. It turns out that a base-sized encoder-only Transformer can also achieve 42 AP on COCO, similar to DETR and much more complex frameworks such as Faster R-CNN.
The RegNet model was proposed in Designing Network Design Spaces by Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, Piotr Dollár.
The authors design search spaces to perform Neural Architecture Search (NAS). They first start from a high dimensional search space and iteratively reduce the search space by empirically applying constraints based on the best-performing models sampled by the current search space.
The TAPEX model was proposed in TAPEX: Table Pre-training via Learning a Neural SQL Executor by Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou. TAPEX pre-trains a BART model to solve synthetic SQL queries, after which it can be fine-tuned to answer natural language questions related to tabular data, as well as performing table fact checking.
The Data2Vec model was proposed in data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language by Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu and Michael Auli. Data2Vec proposes a unified framework for self-supervised learning across different data modalities - text, audio and images. Importantly, predicted targets for pre-training are contextualized latent representations of the inputs, rather than modality-specific, context-independent targets.
The vision model is added in v4.19.0.
PyTorch recently upstreamed the Fairscale FSDP into PyTorch Distributed with additional optimizations. This PR is aimed at integrating it into Trainer API.
It enables Distributed Training at Scale. It's a wrapper for sharding Module parameters across data parallel workers. This is inspired by Xu et al. as well as the ZeRO Stage 3 from DeepSpeed. PyTorch FSDP will focus more on production readiness and long-term support. This includes better integration with ecosystems and improvements on performance, usability, reliability, debuggability and composability.
New example scripts were added for image classification and semantic segmentation. Both now have versions that leverage the Trainer API and Accelerate.
To continue democratizing good machine learning, we're making the Transformers documentation more accessible to non-English speakers; starting with Spanish (572M speakers worldwide).
DataCollatorWithPadding by @secsilm in #16662SpmConverter for sentencepiece's model using the byte fallback feature by @SaulLu in #16629no_trainer scripts by @muellerzr in #16703Tapex in table question answering pipeline. by @Narsil in #16663.from_pretrained] Raise a warning if model weights are not in float32 by @sanchit-gandhi in #16762from_pretrained(..., low_cpu_mem_usage=True) + tests by @stas00 in #16657LayoutLMv2 tokenization docstrings by @qqaatw in #16187rum_clm.py seeking text column name twice by @dandelin in #16624attention_mask on gpt2 by @wiio12 in #16829Speech2TextTokenizer by Speech2TextFeatureExtractor in some docstrings by @SaulLu in #16835num_return_sequences>1. by @Narsil in #16828array key in raw dictionnaries in ASR pipeline. by @Narsil in #16827convert_file_size_to_int by @mariosasko in #16891distributed_concat with scalar tensor by @Yard1 in #16963decoder_module by @sanchit-gandhi in #17036token_type_ids by @deutschmn in #17082mobilebert onnx configs by @manandey in #17029question_answering pipeline. by @Narsil in #17143Trainer by @Yard1 in #17166The following contributors have made significant changes to the library over the last release:
Replace all deprecated jax.ops operations with jnp's at by @sanchit-gandhi in https://github.com/huggingface/transformers/pull/16078
You'll notice that we are starting to add several older models in vision. This is because those models are used as backbones in recent architectures. While we could rely on existing libraries for such pretrained models, we will ultimately need some support for those backbones in PyTorch/TensorFlow and Jax, and there is currently no library that supports those three frameworks. This is why we are starting to add those models to Transformers directly (here ResNet and VAN)
The GLPN model was proposed in Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth by Doyeon Kim, Woonghyun Ga, Pyungwhan Ahn, Donggyu Joo, Sehwan Chun, Junmo Kim. GLPN combines SegFormer’s hierarchical mix-Transformer with a lightweight decoder for monocular depth estimation. The proposed decoder shows better performance than the previously proposed decoders, with considerably less computational complexity.
The ResNet model was proposed in Deep Residual Learning for Image Recognition by Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun. Our implementation follows the small changes made by Nvidia, we apply the stride=2 for downsampling in bottleneck’s 3x3 conv and not in the first 1x1. This is generally known as “ResNet v1.5”.
ResNet introduced residual connections, they allow to train networks with an unseen number of layers (up to 1000). ResNet won the 2015 ILSVRC & COCO competition, one important milestone in deep computer vision.
The VAN model was proposed in Visual Attention Network by Meng-Hao Guo, Cheng-Ze Lu, Zheng-Ning Liu, Ming-Ming Cheng, Shi-Min Hu.
This paper introduces a new attention layer based on convolution operations able to capture both local and distant relationships. This is done by combining normal and large kernel convolution layers. The latter uses a dilated convolution to capture distant correlations.
The VisionTextDualEncoderModel can be used to initialize a vision-text dual encoder model with any pretrained vision autoencoding model as the vision encoder (e.g. ViT, BEiT, DeiT) and any pretrained text autoencoding model as the text encoder (e.g. RoBERTa, BERT). Two projection layers are added on top of both the vision and text encoder to project the output embeddings to a shared latent space. The projection layers are randomly initialized so the model should be fine-tuned on a downstream task. This model can be used to align the vision-text embeddings using CLIP like contrastive image-text training and then can be used for zero-shot vision tasks such image-classification or retrieval.
In LiT: Zero-Shot Transfer with Locked-image Text Tuning it is shown how leveraging pre-trained (locked/frozen) image and text model for contrastive learning yields significant improvment on new zero-shot vision tasks such as image classification or retrieval.
DiT was proposed in DiT: Self-supervised Pre-training for Document Image Transformer by Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, Furu Wei. DiT applies the self-supervised objective of BEiT (BERT pre-training of Image Transformers) to 42 million document images, allowing for state-of-the-art results on tasks including:
The DPT model was proposed in Vision Transformers for Dense Prediction by René Ranftl, Alexey Bochkovskiy, Vladlen Koltun. DPT is a model that leverages the Vision Transformer (ViT) as backbone for dense prediction tasks like semantic segmentation and depth estimation.
Large models are becoming more and more the norm and having a checkpoint in a single file is challenging for several reasons:
That's why the save_pretrained method will know automatically shard a checkpoint in several files when you go above a 10GB threshold for PyTorch models. from_pretrained will handle such sharded checkpoints as if there was only one file.
GPT-J and ViTMAE are now available in TensorFlow.
The IA migration is wrapped up with a new conceptual guide available.
torch.diag by @Narsil in https://github.com/huggingface/transformers/pull/15890MODEL_FOR_INSTANCE_SEGMENTATION_MAPPING by @Narsil in https://github.com/huggingface/transformers/pull/15934max_length in BeamScorer.finalize()) by @cwkeam in https://github.com/huggingface/transformers/pull/15555ForInstanceSegmentation models to image-segmentation pipelines by @Narsil in https://github.com/huggingface/transformers/pull/15937pos optional in PerceiverAudioPreprocessor to avoid crashing PerceiverModel operation by @basilevh in https://github.com/huggingface/transformers/pull/15972'torch.dtype' has str-type value in config and all nested dicts for JSON serializability by @feifang24 in https://github.com/huggingface/transformers/pull/16065HF_ENDPOINT for custom endpoints by @sgugger in https://github.com/huggingface/transformers/pull/16139hidden_states by @ydshieh in https://github.com/huggingface/transformers/pull/16167jax.ops operations with jnp's at by @sanchit-gandhi in https://github.com/huggingface/transformers/pull/16078has_attentions as done in PyTorch side by @ydshieh in https://github.com/huggingface/transformers/pull/16259add-new-model-like work in an env without all frameworks by @sgugger in https://github.com/huggingface/transformers/pull/16239cached_download ∘ hf_hub_url is hf_hub_download by @julien-c in https://github.com/huggingface/transformers/pull/16375torch.distributed process group if one is already initailized by @Yard1 in https://github.com/huggingface/transformers/pull/16487segmentation_maps by @FrancescoSaverioZuppichini in https://github.com/huggingface/transformers/pull/15964run_qa_no_trainer.py by @bhadreshpsavani in https://github.com/huggingface/transformers/pull/16508__init__.py: modeling_xglm -> modeling_flax_xglm by @stancld in https://github.com/huggingface/transformers/pull/16556convert_tokens_to_string's output by @SaulLu in https://github.com/huggingface/transformers/pull/16540unpack_inputs-related changes by @gante in https://github.com/huggingface/transformers/pull/16499_load_pretrained_model_low_mem static + bug fix by @FrancescoSaverioZuppichini in https://github.com/huggingface/transformers/pull/16548The community contributors below have significantly contributed to the v4.18.0 release. Thank you!
@sayakpaul, for contributing the TensorFlow version of ViTMAE @stancld, for contributing the TensorFlow version of of GPT-J
Full Changelog: https://github.com/huggingface/transformers/compare/v4.17.0...v4.18.0
Prepare deprecated ONNX exporter for torch v1.11 by @lewtun in https://github.com/huggingface/transformers/pull/15388
The XGLM model was proposed in Few-shot Learning with Multilingual Language Models by Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li.
XGLM is a GPT3-like multilingual model trained on a balanced corpus covering a diverse set of languages.
The ConvNeXT model was proposed in A ConvNet for the 2020s by Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie.
ConvNeXT is a pure convolutional model (ConvNet), inspired by the design of Vision Transformers, that claims to outperform them.
The PoolFormer model was proposed in MetaFormer is Actually What You Need for Vision by Sea AI Labs.
The PLBART model was proposed in Unified Pre-training for Program Understanding and Generation by Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, Kai-Wei Chang.
This is a BART-like model which can be used to perform code-summarization, code-generation, and code-translation tasks. The pre-trained model plbart-base has been trained using multilingual denoising task on Java, Python and English.
The Data2Vec model was proposed in data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language by Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu and Michael Auli.
Data2Vec proposes a unified framework for self-supervised learning across different data modalities - text, audio and images. Importantly, predicted targets for pre-training are contextualized latent representations of the inputs, rather than modality-specific, context-independent targets.
The MaskFormer model was proposed in Per-Pixel Classification is Not All You Need for Semantic Segmentation by Bowen Cheng, Alexander G. Schwing, Alexander Kirillov.
MaskFormer addresses semantic segmentation with a mask classification paradigm instead of performing classic pixel-level classification.
This is a new experimental feature added to the library. It allows you to share a custom model (with configuration, tokenizer, feature extractor, processor) with anyone through the Model Hub while still using the Auto-classes API of the Transformers library.
See the documentation for more information!
We are working on updating the existing guides in the documentation, and writing more!
Speech models that have been trained with the CTC loss (Wav2Vec2, XLS-R, HuBERT, WavLM, ...) can now output the time
stamp in addition to the transcription of the input audio. E.g. one can retrieve the start and end time for every transcribed word
via the Wav2Vec2CTCTokenizer.decode method or the Wav2Vec2ProcessorWithLM.decoder method. See the documentation here and here respectively.
This feature can also be directly used via the ASR pipeline - see here and this example.
Unfortunately, some bugs had crept into CLIPTokenizerFast : the tokenization produced by CLIPTokenizer and CLIPTokenizerFast were not equal. CLIPTokenizerFast has been corrected to encode the text with the same strategy as CLIPTokenizer.
What does this mean for you ? You need to use the tokenizer that was used to train the CLIP template you are using. For example:
openai/clip-vit-base-patch32, openai/clip-vit-base-patch16 or openai/clip-vit-large-patch14 , before v4.17.0 the good version of the tokenizer was CLIPTokenizer. From v4.17.0, you can use both CLIPTokenizer and CLIPTokenizerFast.CLIPTokenizerFast. Your tokenizer is no longer a CLIPTokenizerFast and we recommend you to load your tokenizer.json in a PreTrainedTokenizerFast directly or to continue to use a version prior to v4.17.0.CLIPTokenizer. Now, you can produce a fast equivalent of your tokenizer by doing CLIPTokenizerFast.from_pretrained("Path to local folder or Hub repo with slow tokenizer files", from_slow=True).To make CLIPTokenizerFast identical to CLIPTokenizer, the template of the tokenization of a sentence pair (A,B) has been modified. The previous template was <|startoftext|> A B <|endoftext|> and the new one is <|startoftext|> A <|endoftext|> <|endoftext|> B <|endoftext|>.
batch_size and num_return_Sequences in text-generation pipeline by @Narsil in https://github.com/huggingface/transformers/pull/15318bad_words_ids not working with sentencepiece-based tokenizers by @ngoquanghuy99 in https://github.com/huggingface/transformers/pull/15343pr_check by @ngoquanghuy99 in https://github.com/huggingface/transformers/pull/15380PreTrainedTokenizerBase __init__ by @SaulLu in https://github.com/huggingface/transformers/pull/15454tokenizer_config.json file for the slow tokenizer when a fast version is available by @SaulLu in https://github.com/huggingface/transformers/pull/15319Trainer.push_to_hub always tries to push to the Hub by @sgugger in https://github.com/huggingface/transformers/pull/15463microphone streaming within pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/15046__init__ of PreTrainedTokenizerBase by @SaulLu in https://github.com/huggingface/transformers/pull/15456requires_backends by @tkukurin in https://github.com/huggingface/transformers/pull/15636tokenizers>=0.11.1 by @aphedges in https://github.com/huggingface/transformers/pull/15266KeyDataset. by @Narsil in https://github.com/huggingface/transformers/pull/15645decoder_kwargs to send to LM on asr pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/15646HfDeepSpeedConfig argument in Trainer by @jaketae in https://github.com/huggingface/transformers/pull/15711HfArgumentParser when passing a generator by @bryant1410 in https://github.com/huggingface/transformers/pull/15758hf-internal-testing/tiny-clip for instance) by @Narsil in https://github.com/huggingface/transformers/pull/15782image-segmentation on AutoModelForSemanticSegmentation by @Narsil in https://github.com/huggingface/transformers/pull/15647The community contributors below have significantly contributed to the v4.17.0 release. Thank you!
@sayakpaul, for contributing the TensorFlow version of ConvNext @gchhablani, for contributing PLBart @edugp, for contributing Data2Vec
Full Changelog: https://github.com/huggingface/transformers/compare/v4.16.0...v4.17.0
[Hotfix] Fix Swin model outputs (huggingface#15414)
Full Changelog: https://github.com/huggingface/transformers/compare/v4.16.1...v4.16.2
Add init to BORT (#15378) by @LysandreJik
Add init to BORT (#15378) by @LysandreJik
Deprecates AdamW and adds --optim by @manuelciosici in https://github.com/huggingface/transformers/pull/14744
The Nyströmformer model was proposed in Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention by Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh.
The Nyströmformer model overcomes the quadratic complexity of self-attention on the input sequence length by adapting the Nyström method to approximate standard self-attention, enabling longer sequences with thousands of tokens as input.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=nystromformer
The REALM model was proposed in REALM: Retrieval-Augmented Language Model Pre-Training by Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat and Ming-Wei Chang.
It’s a retrieval-augmented language model that firstly retrieves documents from a textual knowledge corpus and then utilizes retrieved documents to process question answering tasks.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=realm
The ViTMAE model was proposed in Masked Autoencoders Are Scalable Vision Learners by Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick.
The paper shows that, by pre-training a Vision Transformer (ViT) to reconstruct pixel values for masked patches, one can get results after fine-tuning that outperform supervised pre-training.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=vit_mae
The ViLT model was proposed in ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision by Wonjae Kim, Bokyung Son, Ildoo Kim.
ViLT incorporates text embeddings into a Vision Transformer (ViT), allowing it to have a minimal design for Vision-and-Language Pre-training (VLP).
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=vilt
The Swin Transformer was proposed in Swin Transformer: Hierarchical Vision Transformer using Shifted Windows by Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, Baining Guo.
The Swin Transformer serves as a general-purpose backbone for computer vision. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=swin
The YOSO model was proposed in You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli Sampling by Zhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya, Glenn Fung, Vikas Singh.
YOSO approximates standard softmax self-attention via a Bernoulli sampling scheme based on Locality Sensitive Hashing (LSH). In principle, all the Bernoulli random variables can be sampled with a single hash.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=yoso
To help contributors add new models more easily to Transformers, there is a new command that will clone an existing model and set the various hooks in the library, so that you only have to write the tweaks needed to the modeling file. Just run transformers-cli add-new-model-like and fill the questionnaire!
New training scripts were introduced, for speech seq2seq models and an image pre-training script leveraging the ViTMAE models. Finally, an image captioning example in Flax gets added to the library.
Adding support for long files on automatic-speech-recognition (ASR) as well as supporting audio models with LM which increases the WER on many tasks See the blogpost.
Also continuously increasing homogeneity in arguments, framework support on all pipelines.
TF on image-classification pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/15030batch_size enabled on zero-cls and qa pipelines. by @Narsil in https://github.com/huggingface/transformers/pull/14225The ELECTRA model can now be used as a decoder, enabling an ELECTRA encoder-decoder model.
ElectraForCausalLM -> Enable Electra encoder-decoder model by @stancld in https://github.com/huggingface/transformers/pull/14729<FILL ME>
The vision encoder decoder model can now be used in TensorFlow.
CLIP gets ported to TensorFlow.
RoFormer gets ported to Flax.
--optim by @manuelciosici in https://github.com/huggingface/transformers/pull/14744The documentation has been fully migrated to MarkDown, if you are making contribution, make sure to read the upgraded guide on how to write good docstrings.
AttributeError from PreTrainedTokenizerFast.decoder by @aphedges in https://github.com/huggingface/transformers/pull/14691run_name in MLflowCallback by @YangDong2002 in https://github.com/huggingface/transformers/pull/14894num_return_sequences support for text2text generation. by @Narsil in https://github.com/huggingface/transformers/pull/14988tokenizers upgrade. by @Narsil in https://github.com/huggingface/transformers/pull/14941chunk_length_s instead of _ms. by @Narsil in https://github.com/huggingface/transformers/pull/15029batch_size arg (like others enabled everywhere). by @Narsil in https://github.com/huggingface/transformers/pull/15027with torch.no_grad() to DistilBERT integration test forward pass by @jaketae in https://github.com/huggingface/transformers/pull/14979tokenize_chinese_chars arg by @SaulLu in https://github.com/huggingface/transformers/pull/15158np.ndarray optional arguments by @gante in https://github.com/huggingface/transformers/pull/15074is_ctc needs to be updated to `self.type == "ctc". by @Narsil in https://github.com/huggingface/transformers/pull/15194from_encoder_decoder_pretrained in encoder-decoder models by @jsnfly in https://github.com/huggingface/transformers/pull/15056The community contributors below have significantly contributed to the v4.16.0 release. Thank you!
Full Changelog: https://github.com/huggingface/transformers/compare/v4.15.0...v4.16.0
[ImageGPT] Deprecate pixel_values input name to input_ids by @patrickvonplaten in https://github.com/huggingface/transformers/pull/14801
WavLM was proposed in WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing by Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Furu Wei.
WavLM sets a new SOTA on the SUPERB benchmark.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=wavlm
Wav2Vec2Phoneme was proposed in Simple and Effective Zero-shot Cross-lingual Phoneme Recognition by Qiantong Xu, Alexei Baevski, Michael Auli. Wav2Vec2Phoneme allows to do phoneme classification as part of automatic speech recognition
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=phoneme-recognition
Unispeech-SAT was proposed in UNISPEECH-SAT: UNIVERSAL SPEECH REPRESENTATION LEARNING WITH SPEAKER AWARE PRE-TRAINING by Sanyuan Chen, Yu Wu, Chengyi Wang, Zhengyang Chen, Zhuo Chen, Shujie Liu, Jian Wu, Yao Qian, Furu Wei, Jinyu Li, Xiangzhan Yu.
UniSpeech-SAT is especially good at speaker related tasks.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=unispeech-sat
Unispeech was proposed in UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data by Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, Xuedong Huang.
Three new models are released as part of the ImageGPT integration: ImageGPTModel, ImageGPTForCausalImageModeling, ImageGPTForImageClassification, in PyTorch.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=unispeech
Wav2Vec2-like architecture now have a speaker diarization and speaker verification head added to their architectures. You can try out the new task here: https://huggingface.co/spaces/microsoft/wavlm-speaker-verification
require_datasets testing utility by @LysandreJik in https://github.com/huggingface/transformers/pull/14795stopping_criteria and logits_processor to generate by @lvwerra in https://github.com/huggingface/transformers/pull/14779FlaxMarianMTModel return block. by @sgugger in https://github.com/huggingface/transformers/pull/14873add_prefix_space and trim_offsets in backend_tokenizer.post_processor of RobertaTokenizerFast by @SaulLu in https://github.com/huggingface/transformers/pull/14752Full Changelog: https://github.com/huggingface/transformers/compare/v4.14.0...v4.15.0
Fixes a circular import when TensorFlow and Onnx are both installed
Fixes a circular import when TensorFlow and Onnx are both installed (#14787)
The Perceiver model was released in the previous version:
The Perceiver model was released in the previous version:
Perceiver
Eight new models are released as part of the Perceiver implementation:
PerceiverModel,PerceiverForMaskedLM,PerceiverForSequenceClassification,PerceiverForImageClassificationLearned,PerceiverForImageClassificationFourier,PerceiverForImageClassificationConvProcessing,PerceiverForOpticalFlow,PerceiverForMultimodalAutoencoding, in PyTorch.The Perceiver IO model was proposed in Perceiver IO: A General Architecture for Structured Inputs & Outputs by Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, João Carreira.
- Add Perceiver IO by @NielsRogge in https://github.com/huggingface/transformers/pull/14487
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=perceiver
Version v4.14.0 adds support for Perceiver in multiple pipelines, including the fill mask and sequence classification pipelines.
The Keras push to hub callback now generates model cards when pushing to the model hub. Additionally to the callback, model cards will be generated by default by the model.push_to_hub() method.
Fix : wrong link in the documentation (ConvBERT vs DistilBERT) by @Tikquuss in https://github.com/huggingface/transformers/pull/14705
Put back open in colab markers by @sgugger in https://github.com/huggingface/transformers/pull/14684
Fix doc examples: KeyError by @ydshieh in https://github.com/huggingface/transformers/pull/14699
Fix doc examples: 'CausalLMOutput...' object has no attribute 'last_hidden_state' by @ydshieh in https://github.com/huggingface/transformers/pull/14678
Adding Perceiver to AutoTokenizer. by @Narsil in https://github.com/huggingface/transformers/pull/14711
Fix doc examples: unexpected keyword argument by @ydshieh in https://github.com/huggingface/transformers/pull/14689
Automatically build doc notebooks by @sgugger in https://github.com/huggingface/transformers/pull/14718
Fix special character in MDX by @sgugger in https://github.com/huggingface/transformers/pull/14721
Fixing tests for perceiver (texts) by @Narsil in https://github.com/huggingface/transformers/pull/14719
[doc] document MoE model approach and current solutions by @stas00 in https://github.com/huggingface/transformers/pull/14725
[Flax examples] remove dependancy on pytorch training args by @patil-suraj in https://github.com/huggingface/transformers/pull/14636
Update bug-report.md by @patrickvonplaten in https://github.com/huggingface/transformers/pull/14715
[Adafactor] Fix adafactor by @patrickvonplaten in https://github.com/huggingface/transformers/pull/14713
Code parrot minor fixes/niceties by @ncoop57 in https://github.com/huggingface/transformers/pull/14666
Fix doc examples: modify config before super().init by @ydshieh in https://github.com/huggingface/transformers/pull/14697
Improve documentation of some models by @NielsRogge in https://github.com/huggingface/transformers/pull/14695
Skip Perceiver tests by @LysandreJik in https://github.com/huggingface/transformers/pull/14745
Add ability to get a list of supported pipeline tasks by @codesue in https://github.com/huggingface/transformers/pull/14732
Fix the perceiver docs by @LysandreJik in https://github.com/huggingface/transformers/pull/14748
[CI/pt-nightly] switch to cuda-11.3 by @stas00 in https://github.com/huggingface/transformers/pull/14726
Swap TF and PT code inside two blocks by @LucienShui in https://github.com/huggingface/transformers/pull/14742
Fix doc examples: cannot import name by @ydshieh in https://github.com/huggingface/transformers/pull/14698
Fix: change tooslow to slow by @ydshieh in https://github.com/huggingface/transformers/pull/14734
Small fixes for the doc by @sgugger in https://github.com/huggingface/transformers/pull/14751
Update transformers metadata by @sgugger in https://github.com/huggingface/transformers/pull/14724
Mention no images added to repository by @LysandreJik in https://github.com/huggingface/transformers/pull/14738
Avoid using tf.tile in embeddings for TF models by @ydshieh in https://github.com/huggingface/transformers/pull/14735
Change how to load config of XLNetLMHeadModel by @josutk in https://github.com/huggingface/transformers/pull/14746
Improve perceiver by @NielsRogge in https://github.com/huggingface/transformers/pull/14750
Convert Trainer doc page to MarkDown by @sgugger in https://github.com/huggingface/transformers/pull/14753
Update Table of Contents by @sgugger in https://github.com/huggingface/transformers/pull/14755
Fixing tests for Perceiver by @Narsil in https://github.com/huggingface/transformers/pull/14739
Make data shuffling in run_clm_flax.py respect global seed by @bminixhofer in https://github.com/huggingface/transformers/pull/13410
Adding support for multiple mask tokens. by @Narsil in https://github.com/huggingface/transformers/pull/14716
Fix broken links to distillation on index page of documentation by @amitness in https://github.com/huggingface/transformers/pull/14722
[doc] performance: groups of operations by compute-intensity by @stas00 in https://github.com/huggingface/transformers/pull/14757
Fix the doc_build_test job by @sgugger in https://github.com/huggingface/transformers/pull/14774
Fix preprocess_function in run_summarization_flax.py by @ydshieh in https://github.com/huggingface/transformers/pull/14769
Simplify T5 docs by @xhlulu in https://github.com/huggingface/transformers/pull/14776
Update Perceiver code examples by @NielsRogge in https://github.com/huggingface/transformers/pull/14783
Full Changelog: https://github.com/huggingface/transformers/compare/v4.13.0...v4.14.0
fix deprecated tf method by @ZOHETH in https://github.com/huggingface/transformers/pull/14671
Eight new models are released as part of the Perceiver implementation: PerceiverModel, PerceiverForMaskedLM, PerceiverForSequenceClassification, PerceiverForImageClassificationLearned, PerceiverForImageClassificationFourier, PerceiverForImageClassificationConvProcessing, PerceiverForOpticalFlow, PerceiverForMultimodalAutoencoding, in PyTorch.
The Perceiver IO model was proposed in Perceiver IO: A General Architecture for Structured Inputs & Outputs by Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, João Carreira.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=perceiver
The mLUKE tokenizer is added. The tokenizer can be used for the multilingual variant of LUKE.
The mLUKE model was proposed in mLUKE: The Power of Entity Representations in Multilingual Pretrained Language Models by Ryokan Ri, Ikuya Yamada, and Yoshimasa Tsuruoka. It's a multilingual extension of the LUKE model trained on the basis of XLM-RoBERTa.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=luke
Three new models are released as part of the ImageGPT integration: ImageGPTModel, ImageGPTForCausalImageModeling, ImageGPTForImageClassification, in PyTorch.
The ImageGPT model was proposed in Generative Pretraining from Pixels by Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, Ilya Sutskever. ImageGPT (iGPT) is a GPT-2-like model trained to predict the next pixel value, allowing for both unconditional and conditional image generation.
Compatible checkpoints can be found on the hub: https://huggingface.co/models?other=imagegpt
Eight new models are released as part of the QDQBert implementation: QDQBertModel, QDQBertLMHeadModel, QDQBertForMaskedLM, QDQBertForSequenceClassification, QDQBertForNextSentencePrediction, QDQBertForMultipleChoice, QDQBertForTokenClassification, QDQBertForQuestionAnswering, in PyTorch.
The QDQBERT model can be referenced in Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation by Hao Wu, Patrick Judd, Xiaojie Zhang, Mikhail Isaev and Paulius Micikevicius.
The semantic Segmentation models' API is unstable and bound to change between this version and the next.
The first semantic segmentation models are added. In semantic segmentation, the goal is to predict a class label for every pixel of an image. The models that are added are SegFormer (by NVIDIA) and BEiT (by Microsoft Research). BEiT was already available in the library, but this release includes the model with a semantic segmentation head.
The SegFormer model was proposed in SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers by Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, Ping Luo. The model consists of a hierarchical Transformer encoder and a lightweight all-MLP decode head to achieve great results on image segmentation benchmarks such as ADE20K and Cityscapes.
The BEiT model was proposed in BEiT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong, Furu Wei. Rather than pre-training the model to predict the class of an image (as done in the original ViT paper), BEiT models are pre-trained to predict visual tokens from the codebook of OpenAI’s DALL-E model given masked patches.
Adds VisionTextDualEncoder model in PyTorch and Flax to be able to load any pre-trained vision (ViT, DeiT, BeiT, CLIP's vision model) and text (BERT, ROBERTA) model in the library for vision-text tasks like CLIP.
This model pairs a vision and text encoder and adds projection layers to project the embeddings to another embeddings space with similar dimensions. which can then be used to align the two modalities.
CodeParrot, a model trained to generate code, has been open-sourced in the research projects by @lvwerra.
See https://huggingface.co/patrickvonplaten/wav2vec2-xlsr-53-es-kenlm for more information.
Adds Flax version of the vision encoder-decoder model, and adds a Flax version of GPT-J.
Vision transformers are here! Convnets are so 2012, now that ML is converging on self-attention as a universal model.
Want to handle real-world tables, where text and data are positioned in a 2D grid? TAPAS is now here for both TensorFlow and PyTorch.
Automatic checkpointing and cloud saves to the HuggingFace Hub during training are now live, allowing you to resume training when it's interrupted, even if your initial instance is terminated. This is an area of very active development - watch this space for future developments, including automatic model card creation and more.
A new class to automatically select processors is added: AutoProcessor. It can be used for all models that require a processor, in both computer vision and audio.
A new documentation frontend is out for the transformers library! The goal with this documentation is to be better aligned with the rest of our website, and contains tools to improve readability. The documentation can now be written in markdown rather than RST.
The LayoutLMv2 feature extractor now supports non-English languages, and LayoutXLM gets its own processor.
You can now take advantage of the Ampere hardware with the Trainer:
--bf16 - do training or eval in mixed precision of bfloat16--bf16_full_eval - do eval in full bfloat16--tf32 control having TF32 mode on/offbatch_size support for (almost) all pipelines by @Narsil in https://github.com/huggingface/transformers/pull/13724BlenderbotTokenizerFast by @stancld in https://github.com/huggingface/transformers/pull/13720handle_long_generation paramters for text-generation pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/14118image-segmentation tests. by @Narsil in https://github.com/huggingface/transformers/pull/14223load_image function in image_utils.py & fix image rotation issue by @mishig25 in https://github.com/huggingface/transformers/pull/14062truncation parameter on feature-extraction pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/14193ignore_labels. by @Narsil in https://github.com/huggingface/transformers/pull/14274DPRPretrainedModel from docs by @xhlulu in https://github.com/huggingface/transformers/pull/14300pipeline. by @Narsil in https://github.com/huggingface/transformers/pull/14316BatchFeature: Convert List[np.ndarray] to np.ndarray before converting to pytorch tensors by @eladsegal in https://github.com/huggingface/transformers/pull/14306pipeline function. by @Narsil in https://github.com/huggingface/transformers/pull/14322generator in addition to Dataset for pipelines by @Narsil in https://github.com/huggingface/transformers/pull/14352AlbertConverter for FNet instead of using FNet's own converter by @qqaatw in https://github.com/huggingface/transformers/pull/14365inputs_embeds as an input by @patrickvonplaten in https://github.com/huggingface/transformers/pull/14443hidden_states and attentions in unbatching support. by @Narsil in https://github.com/huggingface/transformers/pull/14420Narsil to hf-internal-testing. by @Narsil in https://github.com/huggingface/transformers/pull/14463add-new-pipeline docs a bit by @stancld in https://github.com/huggingface/transformers/pull/14485~/.cache/torch_extensions between builds by @stas00 in https://github.com/huggingface/transformers/pull/14520__call__ method by @xhlulu in https://github.com/huggingface/transformers/pull/14379Full Changelog: https://github.com/huggingface/transformers/compare/v4.12.0...v4.13.0
Reverts a commit that introduced other issues:
Reverts a commit that introduced other issues:
Fix gradient_checkpointing backward compatibility
Add PushToHubCallback in main init
Fixes an issue with the image segmentation pipeline and PyTorch's inference mode.
Fixes an issue with the image segmentation pipeline and PyTorch's inference mode.
Nothing published for this version
One new model is released as part of the TrOCR implementation: TrOCRForCausalLM, in PyTorch. It comes along a new VisionEncoderDecoderModel class, whi
One new model is released as part of the TrOCR implementation: TrOCRForCausalLM, in PyTorch. It comes along a new VisionEncoderDecoderModel class, which allows to mix-and-match any vision Transformer encoder with any text Transformer as decoder, similar to the existing SpeechEncoderDecoderModel class.
The TrOCR model was proposed in TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models, by Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, Furu Wei.
The TrOCR model consists of an image transformer encoder and an autoregressive text transformer to perform optical character recognition in an end-to-end manner.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?other=trocr
SEW and SEW-D (Squeezed and Efficient Wav2Vec) were proposed in Performance-Efficiency Trade-offs in Unsupervised Pre-training for Speech Recognition by Felix Wu, Kwangyoun Kim, Jing Pan, Kyu Han, Kilian Q. Weinberger, Yoav Artzi.
SEW and SEW-D models use a Wav2Vec-style feature encoder and introduce temporal downsampling to reduce the length of the transformer encoder. SEW-D additionally replaces the transformer encoder with a DeBERTa one. Both models achieve significant inference speedups without sacrificing the speech recognition quality.
Compatible checkpoints are available on the Hub: https://huggingface.co/models?other=sew and https://huggingface.co/models?other=sew-d
DistilHuBERT was proposed in DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT, by Heng-Jui Chang, Shu-wen Yang, Hung-yi Lee.
DistilHuBERT is a distilled version of the HuBERT model. Using only two transformer layers, the model scores competitively on the SUPERB benchmark tasks.
Compatible checkpoint is available on the Hub: https://huggingface.co/ntu-spml/distilhubert
Several bug fixes and UX improvements for TensorFlow
Introduction of a Keras callback to push to the hub each epoch, or after a given number of steps:
The encoder-decoder framework is now available in TensorFlow, allowing mixing and matching different encoders and decoders together into a single encoder-decoder architecture!
Besides this, the EncoderDecoderModel classes have been updated to work similar to models like BART and T5. From now on, users don't need to pass decoder_input_ids themselves anymore to the model. Instead, they will be created automatically based on the labels (namely by shifting them one position to the right, replacing -100 by the pad_token_id and prepending the decoder_start_token_id). Note that this may result in training discrepancies if fine-tuning a model trained with versions anterior to 4.12.0 that set the decoder_input_ids = labels.
To make it easier to extend the Transformers library, every Auto class a new register method, that allows you to register your own custom models, configurations or tokenizers. See more in the documentation
run_glue.py] missing requirements scipy, sklearn by @stas00 in https://github.com/huggingface/transformers/pull/13768PreTrainedModel.framework attribute by @StellaAthena in https://github.com/huggingface/transformers/pull/13817find_unused_parameters in Trainer when gradient checkpointing is enabled by @patrickvonplaten in https://github.com/huggingface/transformers/pull/13961pad_to_multiple_of by @affjljoo3581 in https://github.com/huggingface/transformers/pull/13949hf-internal testing ... by @patrickvonplaten in https://github.com/huggingface/transformers/pull/14008modeling_speech_to_text by @mishig25 in https://github.com/huggingface/transformers/pull/14044to_tensor() in TF inline example by @Rocketknight1 in https://github.com/huggingface/transformers/pull/14140Full Changelog: https://github.com/huggingface/transformers/compare/v4.11.0...v4.12.0
This patch release fixes a few issues encountered since the release of v4.11.2:
This patch release fixes a few issues encountered since the release of v4.11.2:
# v4.11.2: Patch release Fix the Trainer API on TPU: - Fix gather for TPU #13813
Fix the Trainer API on TPU:
Patch release with a few bug fixes:
Patch release with a few bug fixes:
*We strive for no breaking changes between releases - however, some bugs are not discovered for long periods of time, and users may eventually rely on…
Three new models are released as part of the GPT-J implementation: GPTJModel, GPTJForCausalLM, GPTJForSequenceClassification, in PyTorch.
The GPT-J model was released in the kingoflolz/mesh-transformer-jax repository by Ben Wang and Aran Komatsuzaki. It is a GPT-2-like causal language model trained on the Pile dataset.
It was contributed by @StellaAthena, @kurumuz, @EricHallahan, and @leogao2.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=gptj
One new model is released as part of the Speech2Text2 implementation: Speech2Text2ForCausalLM, in PyTorch.
The Speech2Text2 model is used together with Wav2Vec2 for Speech Translation models proposed in Large-Scale Self- and Semi-Supervised Learning for Speech Translation by Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski, Michael Auli, Alexis Conneau.
Speech2Text2 is a decoder-only transformer model that can be used with any speech encoder-only, such as Wav2Vec2 or HuBERT for Speech-to-Text tasks. Please refer to the SpeechEncoderDecoder class on how to combine Speech2Text2 with any speech encoder-only model.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?other=speech2text2
Eight new models are released as part of the FNet implementation: FNetModel, FNetForPreTraining, FNetForMaskedLM, FNetForNextSentencePrediction, FNetForSequenceClassification, FNetForMultipleChoice, FNetForTokenClassification, FNetForQuestionAnswering, in PyTorch.
The FNet model was proposed in FNet: Mixing Tokens with Fourier Transforms by James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, Santiago Ontanon. The model replaces the self-attention layer in a BERT model with a fourier transform which returns only the real parts of the transform. The model is significantly faster than the BERT model because it has fewer parameters and is more memory efficient. The model achieves about 92-97% accuracy of BERT counterparts on GLUE benchmark, and trains much faster than the BERT model.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?other=fnet
Several bug fixes and UX improvements for Tensorflow:
Changes to compile() and train_step()
Associated PRs:
The pipelines underwent a large refactor that should make contributing pipelines much simpler, and much less error-prone. As part of this refactor, PyTorch-based pipelines are now optimized for GPU performance based on PyTorch's Datasets and DataLoaders.
See below for an example leveraging the superb dataset.
pipe = pipeline("automatic-speech-recognition", model="facebook/wav2vec2-base-960h", device=0)
dataset = datasets.load_dataset("superb", name="asr", split="test")
# KeyDataset (only `pt`) will simply return the item in the dict returned by the dataset item
# as we're not interested in the `target` part of the dataset.
for out in tqdm.tqdm(pipe(KeyDataset(dataset, "file"))):
print(out)
# {"text": "NUMBER TEN FRESH NELLY IS WAITING ON YOU GOOD NIGHT HUSBAND"}
# {"text": ....}
# ....
Additionally, an additional pipeline is available, for audio classification.
AudioClassificationPipeline #13342 (@anton-l)pipeline for audio-classification. #13376 (@Narsil)Version v4.11.0 introduces setters for common configuration properties. Different configurations have different properties as coming from different implementations.
One such example is the BertConfig having the hidden_size attribute, while the GPT2Config has the n_embed attribute, which are essentially the same.
The newly introduced setters allow setting such properties through a standardized naming scheme, even on configuration objects that do not have them by default.
See the following code sample for an example:
from transformers import GPT2Config
config = GPT2Config()
config.hidden_size = 4 # Failed previously
config = GPT2Config(hidden_size =4) # Failed previously
config.n_embed # returns 4
config.hidden_size # returns 4
An experimental feature adding support for model files hosted on the hub is added as part of this release. A walkthrough is available in the PR description.
:warning: This means that code files will be fetched from the hub to be executed locally. An additional argument, trust_remote_code is required when instantiating the model from the hub. We heavily encourage you to also specify a revision if using code from another user's or organization's repository.
The Trainer has received several new features, the main one being that models are uploaded to the Hub each time you save them locally (you can specify another strategy). This push is asynchronous, so training continues normally without interruption.
Also:
Trainer API as an opt-in component.Trainer API now supports fine-tuning on distributed CPUs.Associated PRs:
The memory required to load a model in memory using PyTorch's torch.load requires twice the amount of memory necessary. An experimental feature allowing model loading while requiring only the model size in terms of memory usage is out in version v4.11.0.
It can be used by using the low_cpu_mem_usage=True argument with PyTorch pretrained models.
from_pretrained #13466 (@stas00)The GPT-Neo local attention was greatly simplified with no loss of performance.
We strive for no breaking changes between releases - however, some bugs are not discovered for long periods of time, and users may eventually rely on such bugs. We document here such changes that may affect users when updating to a recent version.
The overflowing tokens returned by the slow tokenizers were returned in the wrong order. This is changed in the PR below.
Updates the behavior of aggregation_strategy to more closely mimic the deprecated grouped_entities pipeline argument.
The changes in v4.10 (#12804) introduced a bug in inputs normalization for non-padded tensors that affected Wav2Vec2 fine-tuning. This is fixed in the PR below.
Trainer #12806 (@SaulLu)Hubert to the AutoFeatureExtractor #13366 (@anton-l)SpeechEncoderDecoderModel #13422 (@anton-l)packaging package #13454 (@shivdhar)LabelSmoother #13464 (@sgugger)TFWav2Vec2 and TFHubert #13480 (@anton-l)past_key_values during GPT-Neo inference #13521 (@aphedges)perplexity.rstto make the notebook executable on google collaboratory #13541 (@SaulLu)fill-mask (wasn't tested in theses test suites apparently). #13540 (@Narsil)config_dict_or_path for deepspeed.zero.Init #13614 (@aphedges)forward #13665 (@stas00)extractor.pad() dtype backwards compatibility #13693 (@anton-l)float16 checkpoints in integration tests #13676 (@anton-l)UnicodeDecodeError when loading config file #13717 (@qqaatw)examples/pytorch/speech-recognition #13620 (@patrickvonplaten)distributed_concat() #13746 (@Renovamen)Patches an issue with the serialization of the TrainingArguments
Patches an issue with the serialization of the TrainingArguments
[Wav2Vec2] Fix dtype 64 bug #13517 (@patrickvonplaten)
[Wav2Vec2] Fix normalization for non-padded tensors #13512 (@patrickvonplaten)
Four new models are released as part of the LatourLM-v2 implementation: LayoutLMv2ForSequenceClassification, LayoutLMv2Model, LayoutLMv2ForTokenClassi
Four new models are released as part of the LatourLM-v2 implementation: LayoutLMv2ForSequenceClassification, LayoutLMv2Model, LayoutLMv2ForTokenClassification and LayoutLMv2ForQuestionAnswering, in PyTorch.
The LayoutLMV2 model was proposed in LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding by Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. LayoutLMV2 improves LayoutLM to obtain state-of-the-art results across several document image understanding benchmarks:
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=layoutlmv2
Three new models are released as part of the BEiT implementation: BeitModel, BeitForMaskedImageModeling, and BeitForImageClassification, in PyTorch.
The BEiT model was proposed in BEiT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong and Furu Wei. Inspired by BERT, BEiT is the first paper that makes self-supervised pre-training of Vision Transformers (ViTs) outperform supervised pre-training. Rather than pre-training the model to predict the class of an image (as done in the original ViT paper), BEiT models are pre-trained to predict visual tokens from the codebook of OpenAI’s DALL-E model given masked patches.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=beit
The Wav2Vec2 and HuBERT models now have a sequence classification head available.
The DeBERTa and DeBERTa-v2 models have been converted from PyTorch to TensorFlow.
EncoderDecoder, DistilBERT, and ALBERT, now have support in Flax!
A new example has been added in TensorFlow: multiple choice! Data collators have become framework agnostic and can now work for both TensorFlow and NumPy on top of PyTorch.
The Auto APIs have been disentangled from all the other mode modules of the Transformers library, so you can now safely import the Auto classes without importing all the models (and maybe getting errors if your setup is not compatible with one specific model). The actual model classes are only imported when needed.
When loading some kinds of corrupted state dictionaries of models, the PreTrainedModel.from_pretrained method was sometimes silently ignoring weights. This has now become a real error.
Improving pipeline tests #12784 (@Narsil)
Pin git python to <3.1.19 #12858 (@patrickvonplaten)
[tests] fix logging_steps requirements #12860 (@stas00)
[Sequence Feature Extraction] Add truncation #12804 (@patrickvonplaten)
add classifier_dropout to classification heads #12794 (@PhilipMay)
Fix barrier for SM distributed #12853 (@sgugger)
Add possibility to ignore imports in test_fecther #12801 (@sgugger)
Add accelerate to examples requirements #12888 (@sgugger)
Fix documentation of BigBird tokenizer #12889 (@sgugger)
Better heuristic for token-classification pipeline. #12611 (@Narsil)
Fix push_to_hub for TPUs #12895 (@sgugger)
Seq2SeqTrainer set max_length and num_beams only when non None #12899 (@cchen-dialpad)
[FLAX] Minor fixes in CLM example #12914 (@stefan-it)
Correct validation_split_percentage argument from int (ex:5) to float (0.05) #12897 (@Elysium1436)
Fix typo in the example of MobileBertForPreTraining #12919 (@buddhics)
Add option to set max_len in run_ner #12929 (@sgugger)
Fix QA examples for roberta tokenizer #12928 (@sgugger)
Print defaults when using --help for scripts #12930 (@sgugger)
Fix StoppingCriteria ABC signature #12918 (@willfrey)
Add missing @classmethod decorators #12927 (@willfrey)
fix distiller.py #12910 (@chutaklee)
Update generation_logits_process.py #12901 (@willfrey)
Update generation_logits_process.py #12900 (@willfrey)
Update tokenization_auto.py #12896 (@willfrey)
Fix docstring typo in tokenization_auto.py #12891 (@willfrey)
[Flax] Correctly Add MT5 #12988 (@patrickvonplaten)
ONNX v2 raises an Exception when using PyTorch < 1.8.0 #12933 (@mfuntowicz)
Moving feature-extraction pipeline to new testing scheme #12843 (@Narsil)
Add CpmTokenizerFast #12938 (@JetRunner)
fix typo in gradient_checkpointing arg #12855 (@21jun)
Log Azure ML metrics only for rank 0 #12766 (@harshithapv)
Add substep end callback method #12951 (@wulu473)
Add multilingual documentation support #12952 (@JetRunner)
Fix division by zero in NotebookProgressPar #12953 (@sgugger)
[FLAX] Minor fixes in LM example #12947 (@stefan-it)
Prevent Trainer.evaluate() crash when using only tensorboardX #12963 (@aphedges)
Fix typo in example of DPRReader #12954 (@tadejsv)
Place BigBirdTokenizer in sentencepiece-only objects #12975 (@sgugger)
fix typo in example/text-classification README #12974 (@fullyz)
Fix template for inputs docstrings #12976 (@sgugger)
fix Trainer.train(resume_from_checkpoint=False) is causing an exception #12981 (@PhilipMay)
Cast logits from bf16 to fp32 at the end of TF_T5 #12332 (@szutenberg)
Update CANINE test #12453 (@NielsRogge)
pad_to_multiple_of added to DataCollatorForWholeWordMask #12999 (@Aktsvigun)
[Flax] Align jax flax device name #12987 (@patrickvonplaten)
[Flax] Correct flax docs #12782 (@patrickvonplaten)
T5: Create position related tensors directly on device instead of CPU #12846 (@armancohan)
Skip ProphetNet test #12462 (@LysandreJik)
Create perplexity.rst #13004 (@sashavor)
GPT-Neo ONNX export #12911 (@michaelbenayoun)
Update generate method - Fix floor_divide warning #13013 (@nreimers)
[Flax] Correct pt to flax conversion if from base to head #13006 (@patrickvonplaten)
[Flax T5] Speed up t5 training #13012 (@patrickvonplaten)
FX submodule naming fix #13016 (@michaelbenayoun)
T5 with past ONNX export #13014 (@michaelbenayoun)
Fix ONNX test: Put smaller ALBERT model #13028 (@LysandreJik)
Tpu tie weights #13030 (@sgugger)
Use min version for huggingface-hub dependency #12961 (@lewtun)
tfhub.de -> tfhub.dev #12565 (@abhishekkrthakur)
[Flax] Refactor gpt2 & bert example docs #13024 (@patrickvonplaten)
Add MBART to models exportable with ONNX #13049 (@LysandreJik)
Add to ONNX docs #13048 (@LysandreJik)
Fix small typo in M2M100 doc #13061 (@SaulLu)
Add try-except for torch_scatter #13040 (@JetRunner)
docs: add HuggingArtists to community notebooks #13050 (@AlekseyKorshuk)
Fix ModelOutput instantiation form dictionaries #13067 (@sgugger)
Roll out the test fetcher on push tests #13055 (@sgugger)
Fix fallback of test_fetcher #13071 (@sgugger)
Revert to all tests whil we debug what's wrong #13072 (@sgugger)
Use original key for label in DataCollatorForTokenClassification #13057 (@ibraheem-moosa)
[Doctest] Setup, quicktour and task_summary #13078 (@sgugger)
Add VisualBERT demo notebook #12263 (@gchhablani)
Install git #13091 (@LysandreJik)
Fix classifier dropout in AlbertForMultipleChoice #13087 (@ibraheem-moosa)
Doctests job #13088 (@LysandreJik)
Fix VisualBert Embeddings #13017 (@gchhablani)
Proper import for unittest.mock.patch #13085 (@sgugger)
Reactive test fecthers on scheduled test with proper git install #13097 (@sgugger)
Change a parameter name in FlaxBartForConditionalGeneration.decode() #13074 (@ydshieh)
[Flax/JAX] Run jitted tests at every commit #13090 (@patrickvonplaten)
Rely on huggingface_hub for common tools #13100 (@sgugger)
[FlaxCLIP] allow passing params to image and text feature methods #13099 (@patil-suraj)
Ci last fix #13103 (@sgugger)
Improve type checker performance #13094 (@bschnurr)
Fix VisualBERT docs #13106 (@gchhablani)
Fix CircleCI nightly tests #13113 (@sgugger)
Create py.typed #12893 (@willfrey)
Fix flax gpt2 hidden states #13109 (@ydshieh)
Moving fill-mask pipeline to new testing scheme #12943 (@Narsil)
Fix omitted lazy import for xlm-prophetnet #13052 (@minwhoo)
Fix classifier dropout in bertForMultipleChoice #13129 (@mandelbrot-walker)
Fix frameworks table so it's alphabetical #13118 (@osanseviero)
[Feature Processing Sequence] Remove duplicated code #13051 (@patrickvonplaten)
Ci continue through smi failure #13140 (@LysandreJik)
Fix missing seq_len in electra model when inputs_embeds is used. #13128 (@sararb)
Optimizes ByT5 tokenizer #13119 (@Narsil)
Add splinter #12955 (@oriram)
[AutoFeatureExtractor] Fix loading of local folders if config.json exists #13166 (@patrickvonplaten)
Fix generation docstrings regarding input_ids=None #12823 (@jvamvas)
Update namespaces inside torch.utils.data to the latest. #13167 (@qqaatw)
Fix the loss calculation of ProphetNet #13132 (@StevenTang1998)
Fix LUKE tests #13183 (@NielsRogge)
Add min and max question length options to TapasTokenizer #12803 (@NielsRogge)
SageMaker: Fix sagemaker DDP & metric logs #13181 (@philschmid)
correcting group beam search function output score bug #13211 (@sourabh112)
Change how "additional_special_tokens" argument in the ".from_pretrained" method of the tokenizer is taken into account #13056 (@SaulLu)
remove unwanted control-flow code from DeBERTa-V2 #13145 (@kamalkraj)
Fix load_tf_weights alias. #13159 (@qqaatw)
Add RemBert to AutoTokenizer #13224 (@LysandreJik)
Allow local_files_only for fast pretrained tokenizers #13225 (@BramVanroy)
fix AutoModel.from_pretrained(..., torch_dtype=...) #13209 (@stas00)
Fix broken links in Splinter documentation #13237 (@oriram)
Custom errors and BatchSizeError #13184 (@AmbiTyga)
Bump notebook from 6.1.5 to 6.4.1 in /examples/research_projects/lxmert #13226 (@dependabot[bot])
Update generation_logits_process.py #12671 (@willfrey)
Remove side effects of disabling gradient computaiton #13257 (@LysandreJik)
Replace assert statement with if condition and raise ValueError #13263 (@nishprabhu)
Better notification service #13267 (@LysandreJik)
Fix failing Hubert test #13261 (@LysandreJik)
Add CLIP tokenizer to AutoTokenizer #13258 (@LysandreJik)
Some model_types cannot be in the mapping #13259 (@LysandreJik)
Add require flax to MT5 Flax test #13260 (@LysandreJik)
Migrating conversational pipeline tests to new testing format #13114 (@Narsil)
fix tokenizer_class_from_name for models with - in the name #13251 (@stas00)
Add error message concerning revision #13266 (@BramVanroy)
Move image-classification pipeline to new testing #13272 (@Narsil)
[Hotfix] Fixing the test (warnings was incorrect.) #13278 (@Narsil)
Moving question_answering tests to the new testing scheme. Had to tweak a little some ModelTesterConfig for pipelines. #13277 (@Narsil)
Moving summarization pipeline to new testing format. #13279 (@Narsil)
Moving table-question-answering pipeline to new testing. #13280 (@Narsil)
Moving table-question-answering pipeline to new testing #13281 (@Narsil)
Hotfixing master tests. #13282 (@Narsil)
Moving text2text-generation to new pipeline testing mecanism #13283 (@Narsil)
Add DINO conversion script #13265 (@NielsRogge)
Moving text-generation pipeline to new testing framework. #13285 (@Narsil)
Moving token-classification pipeline to new testing. #13286 (@Narsil)
examples: add keep_linebreaks option to CLM examples #13150 (@stefan-it)
Moving translation pipeline to new testing scheme. #13297 (@Narsil)
Fix BeitForMaskedImageModeling #13275 (@NielsRogge)
Moving zero-shot-classification pipeline to new testing. #13299 (@Narsil)
Fixing mbart50 with return_tensors argument too. #13301 (@Narsil)
[Flax] Correct all return tensors to numpy #13307 (@patrickvonplaten)
examples: only use keep_linebreaks when reading TXT files #13320 (@stefan-it)
Slow tests - run rag token in half precision #13304 (@patrickvonplaten)
[Slow tests] Disable Wav2Vec2 pretraining test for now #13303 (@patrickvonplaten)
Announcing the default model used by the pipeline (with a link). #13276 (@Narsil)
use float 16 in causal mask and masked bias #13194 (@hwijeen)
✨ add citation file #13214 (@flaxel)
Improve documentation of pooler_output in ModelOutput #13228 (@navjotts)
fix: typo spelling grammar #13212 (@slowy07)
Check None before going through iteration #13250 (@qqaatw)
Use existing functionality for #13251 #13333 (@sgugger)
neptune.ai logger: add ability to connect to a neptune.ai run #13319 (@fcakyon)
Update label2id in the model config for run_glue #13334 (@sgugger)
:bug: fix small model card bugs #13310 (@nateraw)
Fall back to observed_batch_size when the dataloader does not know the batch_size. #13188 (@mbforbes)
Fixes #12941 where use_auth_token not been set up early enough #13205 (@bennimmo)
Correct wrong function signatures on the docs website #13198 (@qqaatw)
Fix release utils #13337 (@sgugger)
Add missing module spec #13321 (@laurahanu)
Use DS callable API to allow hf_scheduler + ds_optimizer #13216 (@tjruwase)
Tests fetcher tests #13340 (@sgugger)
[Testing] Add Flax Tests on GPU, Add Speech and Vision to Flax & TF tests #13313 (@patrickvonplaten)
Fixing a typo in the data_collator documentation #13309 (@Serhiy-Shekhovtsov)
Add GPT2ForTokenClassification #13290 (@tucan9389)
Doc mismatch fixed #13345 (@Apoorvgarg-creator)
Handle nested dict/lists of tensors as inputs in the Trainer #13338 (@sgugger)
[doc] correct TP implementation resources #13248 (@stas00)
Fix minor typo in parallelism doc #13289 (@jaketae)
Set missing seq_length variable when using inputs_embeds with ALBERT & Remove code duplication #13152 (@olenmg)
TF CLM example fix typo #13002 (@Rocketknight1)
Add generate kwargs to Seq2SeqTrainingArguments #13339 (@sgugger)
Tpu tie weights #13030 (@sgugger)
Fix barrier for SM distributed #12853 (@sgugger)
Fix barrier for SM distributed #12853 (@sgugger)
The TFTrainer is now entering deprecation - and it is replaced by Keras. With version v4.9.0 comes the end of a long rework of the TensorFlow examples…
This version introduces a new package, transformers.onnx, which can be used to export models to ONNX. Contrary to the previous implementation, this approach is meant as an easily extendable package where users may define their own ONNX configurations and export the models they wish to export.
python -m transformers.onnx --model=bert-base-cased onnx/bert-base-cased/
Validating ONNX model...
-[✓] ONNX model outputs' name match reference model ({'pooler_output', 'last_hidden_state'}
- Validating ONNX Model output "last_hidden_state":
-[✓] (2, 8, 768) matchs (2, 8, 768)
-[✓] all values close (atol: 0.0001)
- Validating ONNX Model output "pooler_output":
-[✓] (2, 768) matchs (2, 768)
-[✓] all values close (atol: 0.0001)
All good, model saved at: onnx/bert-base-cased/model.onnx
Four new models are released as part of the CANINE implementation: CanineForSequenceClassification, CanineForMultipleChoice, CanineForTokenClassification and CanineForQuestionAnswering, in PyTorch.
The CANINE model was proposed in CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation by Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting. It’s among the first papers that train a Transformer without using an explicit tokenization step (such as Byte Pair Encoding (BPE), WordPiece, or SentencePiece). Instead, the model is trained directly at a Unicode character level. Training at a character level inevitably comes with a longer sequence length, which CANINE solves with an efficient downsampling strategy, before applying a deep Transformer encoder.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=canine
This version introduces a new method to train a tokenizer from scratch based off of an existing tokenizer configuration.
from datasets import load_dataset
from transformers import AutoTokenizer
dataset = load_dataset("wikitext", name="wikitext-2-raw-v1", split="train")
# We train on batch of texts, 1000 at a time here.
batch_size = 1000
corpus = (dataset[i : i + batch_size]["text"] for i in range(0, len(dataset), batch_size))
tokenizer = AutoTokenizer.from_pretrained("gpt2")
new_tokenizer = tokenizer.train_new_from_iterator(corpus, vocab_size=20000)
The TFTrainer is now entering deprecation - and it is replaced by Keras. With version v4.9.0 comes the end of a long rework of the TensorFlow examples, for them to be more Keras-idiomatic, clearer, and more robust.
HuBERT is now implemented in TensorFlow:
When load_best_model_at_end was set to True in the TrainingArguments, having a different save_strategy and eval_strategy was accepted but the save_strategy was overwritten by the eval_strategy (the option to keep track of the best model needs to make sure there is an evaluation each time there is a save). This led to a lot of confusion with users not understanding why the script was not doing what it was told, so this situation will now raise an error indicating to set save_strategy and eval_strategy to the same values, and in the case that value is "steps", save_steps must be a round multiple of eval_steps.
--log_level feature #12365 (@bhadreshpsavani)print statement with logger.info in QA example utils #12368 (@bhadreshpsavani)einsum in Albert's attention computation #12394 (@mfuntowicz)push_to_hub #12391 (@patrickvonplaten)Repository import to the FLAX example script #12501 (@LysandreJik)model_kwargs when loading a model in pipeline() #12449 (@aphedges)_mask_hidden_states to avoid double masking #12692 (@mfuntowicz)config.mask_feature_prob > 0 #12705 (@mfuntowicz)list type of additional_special_tokens in special_token_map #12759 (@SaulLu)cls and checkpoint #12619 (@europeanplaice)datasets_modules ImportError with Ray Tune #12749 (@Yard1)save_steps=0|None and logging_steps=0 #12796 (@stas00)Rename detr targets to labels #12280 (@NielsRogge)
Fix default for TensorBoard folder
[tests] reset report_to to none, avoid deprecation warning #12293 (@stas00)
Our example scripts and Trainer are now optimized for publishing your model on the Hugging Face Hub, with Tensorboard training metrics, and an automatically authored model card which contains all the relevant metadata, including evaluation results.
Use --push_to_hub to create a model repo for your training and it will be saved with all relevant metadata at the end of the training.
Other flags are:
push_to_hub_model_id to control the repo namepush_to_hub_organization to specify an organizationBy default if you have tensorboard installed the training scripts will use it to log, and the logging traces folder is conveniently located inside your model output directory, so you can push them to your model repo by default.
Any model repo that contains Tensorboard traces will spawn a Tensorboard server:
which makes it very convenient to see how the training went! This Hub feature is in Beta so let us know if anything looks weird :)
See this model repo
The model card contains info about the datasets used, the eval results, ...
Many users were already adding their eval results to their model cards in markdown format, but this is a more structured way of adding them which will make it easier to parse and e.g. represent in leaderboards such as the ones on Papers With Code!
We use a format specified in collaboration with [PaperswithCode] (https://github.com/huggingface/huggingface_hub/blame/main/modelcard.md), see also this repo.
All models, tokenizers and configurations having a revamp push_to_hub() method as well as a push_to_hub argument in their save_pretrained() method. The workflow of this method is changed a bit to be more like git, with a local clone of the repo in a folder of the working directory, to make it easier to apply patches (use use_temp_dir=True to clone in temporary folders for the same behavior as the experimental API).
Flax/JAX is becoming a fully supported backend of the Transformers library with more models having an implementation in it. BART, CLIP and T5 join the already existing models, find the whole list here.
fill-mask pipeline. #12113 (@Narsil)generate method #12139 (@stancld)fix pt-1.9.0 add_ deprecation #12217 (@stas00)
Three new models are released as part of the DETR implementation: DetrModel, DetrForObjectDetection and DetrForSegmentation, in PyTorch.
DETR consists of a convolutional backbone followed by an encoder-decoder Transformer which can be trained end-to-end for object detection. It greatly simplifies a lot of the complexity of models like Faster-R-CNN and Mask-R-CNN, which use things like region proposals, non-maximum suppression procedure, and anchor generation. Moreover, DETR can also be naturally extended to perform panoptic segmentation, by simply adding a mask head on top of the decoder outputs.
DETR can support any timm backbone.
The DETR model was proposed in End-to-End Object Detection with Transformers by Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=detr
A new tokenizer is released as part of the ByT5 implementation: ByT5Tokenizer. It can be used with the T5 family of models.
The ByT5 model was presented in ByT5: Towards a token-free future with pre-trained byte-to-byte models by Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?search=byt5
14 new models are released as part of the RoFormer implementation: RoFormerModel, RoFormerForCausalLM, RoFormerForMaskedLM, RoFormerForSequenceClassification, RoFormerForTokenClassification, RoFormerForQuestionAnswering and RoFormerForMultipleChoice, TFRoFormerModel, TFRoFormerForCausalLM, TFRoFormerForMaskedLM, TFRoFormerForSequenceClassification, TFRoFormerForTokenClassification, TFRoFormerForQuestionAnswering and TFRoFormerForMultipleChoice, in PyTorch and TensorFlow.
RoFormer is a BERT-like autoencoding model with rotary position embeddings. Rotary position embeddings have shown improved performance on classification tasks with long texts. The RoFormer model was proposed in RoFormer: Enhanced Transformer with Rotary Position Embedding by Jianlin Su and Yu Lu and Shengfeng Pan and Bo Wen and Yunfeng Liu.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=roformer
HuBERT is a speech model that accepts a float array corresponding to the raw waveform of the speech signal.
HuBERT was proposed in HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units by Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, Abdelrahman Mohamed.
Two new models are released as part of the HuBERT implementation: HubertModel and HubertForCTC, in PyTorch.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=hubert
On Monday, June 14th, 2021, we released the first part of the Hugging Face Course. The course is focused on the Hugging Face ecosystem, including transformers. Most of the material in the course is now linked from the transformers documentation which now includes videos to explain singular concepts.
The Wav2Vec2 model can now be used in TensorFlow:
dataset_name to data_args and added accuracy metric #11760 (@philschmid)generate() function #11621 (@stancld)run_flax_glue.py #11820 (@patrickvonplaten)zero.Init in from_config #11805 (@stas00)generate() function #11775 (@stancld)max_new_tokens for generate. #11476 (@Narsil)self.assertEqual instead of assert in deberta v2 test. #11935 (@PhilipMay)nn.log_softmax in run_flax_glue.py #11920 (@n2cholas)DeepSpeedConfigHF from Trainer #11966 (@stas00)PreTrainedModel.generate #11049 (@stas00)forward method #11976 (@kouyk)run_flax_glue.py #11964 (@n2cholas)tests #12155 (@stas00)examples #12156 (@stas00)from_pretrained method #12145 (@LysandreJik)Fix regression in models for sequence classification used for regression tasks #11785 Fix checkpoint deletion when load_bert_model_at_end = True #1174
Fix regression in models for sequence classification used for regression tasks #11785 Fix checkpoint deletion when load_bert_model_at_end = True #11748 Fix evaluation in question answering examples #11746 Fix release utils #11784
fix: The 'warn' method is deprecated #11105 (@stas00)
Transformers aren't just for text - they can handle a huge range of input types, and there's been a flurry of papers and new models in the last few months applying them to vision tasks that had traditionally been dominated by convolutional networks. With this release, we're delighted to announce that several state-of-the-art pretrained vision and multimodal text+vision transformer models are now accessible in the huggingface/transformers repo. Give them a try!
Two new models are released as part of the ViT implementation: ViTModel and ViTForImageClassification, in PyTorch.
ViT is an image transformer-based model obtaining state-of-the-art results on image classification tasks. It was the first paper that successfully trained a Transformer encoder on ImageNet, attaining very good results compared to familiar convolutional architectures.
The Vision Transformer (ViT) model was proposed in An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale by Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=vit
Three new models are released as part of the DeiT implementation: DeiTModel, DeiTForImageClassification and DeiTForImageClassificationWithTeacher, in PyTorch.
DeiT is an image transformer model similar to the ViT model. DeiT (data-efficient image transformers) models are more efficiently trained transformers for image classification, requiring far less data and far less computing resources compared to the original ViT models.
The DeiT model was proposed in Training data-efficient image transformers & distillation through attention by Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, Hervé Jégou.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=deit
Three new models are released as part of the CLIP implementation: CLIPModel, CLIPVisionModel and CLIPTextModel, in PyTorch.
CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3.
The CLIP model was proposed in Learning Transferable Visual Models From Natural Language Supervision by Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=clip
BigBird is a sparse-attention-based transformer that extends Transformer based models, such as BERT to much longer sequences. In addition to sparse attention, BigBird also applies global attention as well as random attention to the input sequence. Theoretically, it has been shown that applying sparse, global, and random attention approximates full attention while being computationally much more efficient for longer sequences. As a consequence of the capability to handle longer context, BigBird has shown improved performance on various long document NLP tasks, such as question answering and summarization, compared to BERT or RoBERTa.
The BigBird model was proposed in Big Bird: Transformers for Longer Sequences by Zaheer, Manzil and Guruganesh, Guru and Dubey, Kumar Avinava and Ainslie, Joshua and Alberti, Chris and Ontanon, Santiago and Pham, Philip and Ravula, Anirudh and Wang, Qifan and Yang, Li and others.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=bigbird_pegasus
LUKE is based on RoBERTa and adds entity embeddings as well as an entity-aware self-attention mechanism, which helps improve performance on various downstream tasks involving reasoning about entities such as named entity recognition, extractive and cloze-style question answering, entity typing, and relation classification.
The LUKE model was proposed in LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention by Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=luke
The MegatronBERT model is added to the library, giving access to the 345m variants.
It is implemented comes with nine different models: MegatronBertModel, MegatronBertForMaskedLM, MegatronBertForCausalLM, MegatronBertForNextSentencePrediction, MegatronBertForPreTraining, MegatronBertForSequenceClassification, MegatronBertForMultipleChoice, MegatronBertForTokenClassification, MegatronBertForQuestionAnswering, in PyTorch.
The MegatronBERT model was proposed in Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism by Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper and Bryan Catanzaro.
The Hugging Face Hub integrates better within transformers, through two new added features:
Models, configurations and tokenizers now have a push_to_hub method to automatically push their state to the hub.
The Trainer can now automatically push its underlying model, configuration and tokenizer in a similar fashion. Additionally, it is able to create a draft of model card on the fly with the training hyperparameters and evaluation results.
Auto modelcard #11599 (@sgugger)
Trainer push to hub #11328 (@sgugger)
The Trainer now integrates two additional stages of ZeRO: ZeRO stage 3 for parameter partitioning, and ZeRO Infinity which extends CPU Offload with NVMe Offload.
Flax support is getting more robust, with model code stabilizing and new models being added to the library.
We welcome @Rocketknight1 as a TensorFlow contributor. This version includes a brand new TensorFlow example based on Keras, which will be followed by examples covering most tasks. Additionally, more TensorFlow setups are covered by adding support for AMD-based GPUs and M1 Macs.
Two new pipelines are added:
AutomaticSpeechRecognitionPipeline. #11337 (@Narsil)get_special_tokens_mask consider all tokens #11163 (@sgugger)which with who #11183 (@cronoik)language_modeling.py #11275 (@taepd)max_length from being mandatory within generate. #11314 (@Narsil)XLMRobertaTokenizer #11149 (@PhilipMay)cross_attentions to model output and fix cross-attention head masking #10699 (@stancld)PreTrainedTokenizerBase to check/handle batch length for text_pair parameter #11486 (@hamelsmu)sp_model_kwargs param missing at unpickle in XLMRobertaTokenizer #11430 (@PhilipMay)datasets submodule. #11563 (@LysandreJik)run_glue_no_trainer #11569 (@sgugger)generate. #11566 (@Narsil)AlbertTransformer) #11596 (@baeseongsu)Your coding agent can read these notes before it upgrades. Set up the MCP server →