NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #225 most downloaded on PyPI
Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Last release 13 days ago
09 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
6 versions withdrawn
withdrawn after publishing
10 years old
242 releases · first in 2016
Fix pipeline when used with private models
pipeline when used with private models (#11123)addressing vulnerability report in research project deps #10802 (@stas00)
Seven new models are released as part of the BigBird implementation: BigBirdModel, BigBirdForPreTraining, BigBirdForMaskedLM, BigBirdForCausalLM, BigBirdForSequenceClassification, BigBirdForMultipleChoice, BigBirdForQuestionAnswering in PyTorch.
BigBird is a sparse-attention based transformer which extends Transformer based models, such as BERT to much longer sequences. In addition to sparse attention, BigBird also applies global attention as well as random attention to the input sequence.
The BigBird model was proposed in Big Bird: Transformers for Longer Sequences by Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, Amr Ahmed.
It is released with an accompanying blog post: Understanding BigBird's Block Sparse Attention
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=big_bird
Two new models are released as part of the GPT Neo implementation: GPTNeoModel, GPTNeoForCausalLM in PyTorch.
GPT-Neo is the code name for a family of transformer-based language models loosely styled around the GPT architecture. EleutherAI's primary goal is to replicate a GPT-3 DaVinci-sized model and open-source it to the public.
The implementation within Transformers is a GPT2-like causal language model trained on the Pile dataset.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=gpt_neo
Features have been added to some examples, and additional examples have been added.
Based on the accelerate library, examples completely exposing the training loop are now part of the library. For easy customization if you want to try a new research idea!
examples/multiple-choice/run_swag_no_trainer.py #10934 (@stancld)examples/run_ner_no_trainer.py #10902 (@stancld)examples/language_modeling/run_mlm_no_trainer.py #11001 (@hemildesai)examples/language_modeling/run_clm_no_trainer.py #11026 (@hemildesai)Thanks to the amazing contributions of @bhadreshpsavani, all examples with Trainer are now standardized and all support the predict stage and will return/save metrics in the same fashion.
The Trainer now supports SageMaker model parallelism out of the box, the old SageMakerTrainer is deprecated as a consequence and will be removed in version 5.
FLAX support has been widened to support all model heads of the BERT architecture, alongside a general conversion script for checkpoints in PyTorch to be used in FLAX.
Auto models now have a FLAX implementation.
pipeline.framework would actually contain a fully qualified model. #10970 (@Narsil)One column per quarter.
Add support for detecting intel-tensorflow version
Nothing published for this version
Deprecate Wav2Vec2ForMaskedLM and add Wav2Vec2ForCTC #10089 (@patrickvonplaten)
Two new models are released as part of the S2T implementation: Speech2TextModel and Speech2TextForConditionalGeneration, in PyTorch.
Speech2Text is a speech model that accepts a float tensor of log-mel filter-bank features extracted from the speech signal. It’s a transformer-based seq2seq model, so the transcripts/translations are generated autoregressively.
The Speech2Text model was proposed in fairseq S2T: Fast Speech-to-Text Modeling with fairseq by Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, Juan Pino.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=speech_to_text
Two new models are released as part of the M2M100 implementation: M2M100Model and M2M100ForConditionalGeneration, in PyTorch.
M2M100 is a multilingual encoder-decoder (seq-to-seq) model primarily intended for translation tasks.
The M2M100 model was proposed in Beyond English-Centric Multilingual Machine Translation by Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, Armand Joulin.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=m2m_100
Six new models are released as part of the I-BERT implementation: IBertModel, IBertForMaskedLM, IBertForSequenceClassification, IBertForMultipleChoice, IBertForTokenClassification and IBertForQuestionAnswering, in PyTorch.
I-BERT is a quantized version of RoBERTa running inference up to four times faster.
The I-BERT framework in PyTorch allows to identify the best parameters for quantization. Once the model is exported in a framework that supports int8 execution (such as TensorRT), a speedup of up to 4x is visible, with no loss in performance thanks to the parameter search.
The I-BERT model was proposed in I-BERT: Integer-only BERT Quantization by Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney and Kurt Keutzer.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=ibert
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=speech_to_text
MBart-50 is created using the original mbart-large-cc25 checkpoint by extending its embedding layers with randomly initialized vectors for an extra set of 25 language tokens and then pretrained on 50 languages.
The MBart model was presented in Multilingual Translation with Extensible Multilingual Pretraining and Finetuning by Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, Angela Fan.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=mbart-50
Fixe new models are released as part of the DeBERTa-v2 implementation: DebertaV2Model, DebertaV2ForMaskedLM, DebertaV2ForSequenceClassification, DeberaV2ForTokenClassification and DebertaV2ForQuestionAnswering, in PyTorch.
The DeBERTa model was proposed in DeBERTa: Decoding-enhanced BERT with Disentangled Attention by Pengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu Chen. It is based on Google’s BERT model released in 2018 and Facebook’s RoBERTa model released in 2019.
It builds on RoBERTa with disentangled attention and enhanced mask decoder training with half of the data used in RoBERTa.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=deberta-v2
The XLSR-Wav2Vec2 model was proposed in Unsupervised Cross-Lingual Representation Learning For Speech Recognition by Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, Michael Auli.
The checkpoint corresponding to that model is added to the model hub: facebook/ wav2vec2-large-xlsr-53
A fine-tuning script showcasing how the Wav2Vec2 model can be trained has been added.
The Wav2Vec2 architecture becomes more stable as several changes are done to its architecture. This introduces feature extractors and feature processors as the pre-processing aspect of multi-modal speech models.
Most of the TensorFlow models are now compatible with automatic mixed precision and have XLA support.
We are rolling out experimental support for model parallelism on SageMaker with a new SageMakerTrainer that can be used in place of the regular Trainer. This is a temporary class that will be removed in a future version, the end goal is to have Trainer support this feature out of the box.
[trainer] deepspeed bug fixes and tests #10039 (@stas00)
Removing run_pl_glue.py from text classification docs, include run_xnli.py & run_tf_text_classification.py #10066 (@cbjuan)
remove token_type_ids from TokenizerBertGeneration output #10070 (@sadakmed)
[deepspeed tests] transition to new tests dir #10080 (@stas00)
Added integration tests for Pytorch implementation of the ELECTRA model #10073 (@spatil6)
Fix naming in TF MobileBERT #10095 (@jplu)
[examples/s2s] add test set predictions #10085 (@patil-suraj)
Logging propagation #10092 (@LysandreJik)
Fix some edge cases in report_to and add deprecation warnings #10100 (@sgugger)
Add head_mask and decoder_head_mask to TF LED #9988 (@stancld)
Replace strided slice with tf.expand_dims #10078 (@jplu)
Fix Faiss Import #10103 (@patrickvonplaten)
[RAG] fix generate #10094 (@patil-suraj)
Fix TFConvBertModelIntegrationTest::test_inference_masked_lm Test #10104 (@abhishekkrthakur)
doc: update W&B related doc #10086 (@borisdayma)
Remove speed metrics from default compute objective [WIP] #10107 (@shiva-z)
Fix tokenizers training in notebooks #10110 (@n1t0)
[DeepSpeed docs] new information #9610 (@stas00)
[CI] build docs faster #10115 (@stas00)
[scheduled github CI] add deepspeed fairscale deps #10116 (@stas00)
Line endings should be LF across repo and not CRLF #10119 (@LysandreJik)
Fix TF LED/Longformer attentions computation #10007 (@jplu)
remove adjust_logits_during_generation method #10087 (@patil-suraj)
[DeepSpeed] restore memory for evaluation #10114 (@stas00)
Update run_xnli.py to use Datasets library #9829 (@Qbiwan)
Add new community notebook - Blenderbot #10126 (@lordtt13)
[DeepSpeed in notebooks] Jupyter + Colab #10130 (@stas00)
[examples/run_s2s] remove task_specific_params and update rouge computation #10133 (@patil-suraj)
Fix typo in GPT2DoubleHeadsModel docs #10148 (@M-Salti)
[hf_api] delete deprecated methods and tests #10159 (@julien-c)
Revert propagation #10171 (@LysandreJik)
Conversion from slow to fast for BPE spm vocabs contained an error. #10120 (@Narsil)
Fix typo in comments #10157 (@mrm8488)
Fix typo in comment #10156 (@mrm8488)
[Doc] Fix version control in internal pages #10124 (@sgugger)
[t5 tokenizer] add info logs #9897 (@stas00)
Fix v2 model loading issue #10129 (@BigBird01)
Fix datasets set_format #10178 (@sgugger)
Fixing NER pipeline for list inputs. #10184 (@Narsil)
Add new model to labels that should not stale #10187 (@LysandreJik)
Check TF ops for ONNX compliance #10025 (@jplu)
[RAG] fix tokenizer #10167 (@patil-suraj)
Fix TF template #10189 (@jplu)
fix run_seq2seq.py; porting trainer tests to it #10162 (@stas00)
Specify dataset dtype #10195 (@LysandreJik)
[CI] make the examples sub-group of tests run always #10196 (@stas00)
[WIP][examples/seq2seq] move old s2s scripts to legacy #10136 (@patil-suraj)
set tgt_lang of MBart Tokenizer for summarization #10205 (@HeroadZ)
Store FLOS as floats to avoid overflow. #10213 (@sgugger)
Fix add_token_positions in custom datasets tutorial #10217 (@joeddav)
[trainer] fix ignored columns logger #10219 (@stas00)
Factor out methods #10215 (@LysandreJik)
Fix head masking for TFT5 models #9877 (@stancld)
[CI] 2 fixes #10248 (@stas00)
[trainer] refactor place_model_on_device logic, add deepspeed #10243 (@stas00)
[Trainer] doc update #10241 (@stas00)
Reduce the time spent for the TF slow tests #10152 (@jplu)
Introduce warmup_ratio training argument #10229 (@tanmay17061)
[Trainer] memory tracker metrics #10225 (@stas00)
Script for distilling zero-shot classifier to more efficient student #10244 (@joeddav)
[test] fix func signature #10271 (@stas00)
[trainer] implement support for full fp16 in evaluation/predict #10268 (@stas00)
[ISSUES.md] propose using google colab to reproduce problems #10270 (@stas00)
Introduce logging_strategy training argument #10267 (@tanmay17061)
[CI] Kill any run-away pytest processes #10281 (@stas00)
Patch zero shot distillation script cuda issue #10284 (@joeddav)
Move the TF NER example #10276 (@jplu)
Fix example links in the task summary #10291 (@sgugger)
fixes #10303 #10304 (@cronoik)
[ci] don't fail when there are no zombies #10308 (@stas00)
fix typo in conversion script #10316 (@tagucci)
Add note to resize token embeddings matrix when adding new tokens to voc #10331 (@LysandreJik)
Deprecate prepare_seq2seq_batch #10287 (@sgugger)
[examples/seq2seq] defensive programming + expand/correct README #10295 (@stas00)
[Trainer] implement gradient_accumulation_steps support in DeepSpeed integration #10310 (@stas00)
Loading from last checkpoint functionality in Trainer.train #10334 (@tanmay17061)
[trainer] add Trainer methods for metrics logging and saving #10266 (@stas00)
Fix evaluation with label smoothing in Trainer #10338 (@sgugger)
Fix broken examples/seq2seq/README.md markdown #10344 (@Wikidepia)
[bert-base-german-cased] use model repo, not external bucket #10353 (@julien-c)
[Trainer/Deepspeed] handle get_last_lr() before first step() #10362 (@stas00)
ConvBERT fix torch <> tf weights conversion #10314 (@abhishekkrthakur)
fix deprecated reference tokenizer.max_len in glue.py #10220 (@poedator)
[trainer] move secondary methods into a separate file #10363 (@stas00)
Run GA on every push even on forks #10383 (@LysandreJik)
GA: only run model templates once #10388 (@LysandreJik)
Bugfix: Removal of padding_idx in BartLearnedPositionalEmbedding #10200 (@mingruimingrui)
Remove unused variable in example for Q&A #10392 (@abhishekkrthakur)
Ignore unexpected weights from PT conversion #10397 (@LysandreJik)
Add support for ZeRO-2/3 and ZeRO-offload in fairscale #10354 (@sgugger)
Fix None in add_token_positions - issue #10210 #10374 (@andreabac3)
Make Barthez tokenizer tests a bit faster #10399 (@sgugger)
Fix run_glue evaluation when model has a label correspondence #10401 (@sgugger)
[ci, flax] non-existing models are unlikely to pass tests #10409 (@julien-c)
[LED] Correct Docs #10419 (@patrickvonplaten)
Add Ray Tune hyperparameter search integration test #10414 (@krfricke)
Ray Tune Integration Bug Fixes #10406 (@amogkam)
[examples] better model example #10427 (@stas00)
Fix conda-build #10431 (@LysandreJik)
[run_seq2seq.py] restore functionality: saving to test_generations.txt #10428 (@stas00)
updated logging and saving metrics #10436 (@bhadreshpsavani)
Introduce save_strategy training argument #10286 (@tanmay17061)
Adds terms to Glossary #10443 (@darigovresearch)
Fixes compatibility bug when using grouped beam search and constrained decoding together #10475 (@mnschmit)
Generate can return cross-attention weights too #10493 (@Mehrad0711)
Fix typos #10489 (@WybeKoper)
[T5] Fix speed degradation bug t5 #10496 (@patrickvonplaten)
feat(docs): navigate with left/right arrow keys #10481 (@ydcjeff)
Refactor checkpoint name in BERT and MobileBERT #10424 (@sgugger)
remap MODEL_FOR_QUESTION_ANSWERING_MAPPING classes to names auto-generated file #10487 (@stas00)
Fix the bug in constructing the all_hidden_states of DeBERTa v2 #10466 (@felixgwu)
Smp grad accum #10488 (@sgugger)
Remove unsupported methods from ModelOutput doc #10505 (@sgugger)
Not always consider a local model a checkpoint in run_glue #10517 (@sgugger)
Removes overwrites for output_dir #10521 (@philschmid)
Rework TPU checkpointing in Trainer #10504 (@sgugger)
[ProphetNet] Bart-like Refactor #10501 (@patrickvonplaten)
Fix example of custom Trainer to reflect signature of compute_loss #10537 (@lewtun)
Fixing conversation test for torch 1.8 #10545 (@Narsil)
Fix torch 1.8.0 segmentation fault #10546 (@LysandreJik)
Fixed dead link in Trainer documentation #10554 (@jwa018)
Typo correction. #10531 (@cliang1453)
Fix embeddings for PyTorch 1.8 #10549 (@sgugger)
Stale Bot #10509 (@LysandreJik)
Refactoring checkpoint names for multiple models #10527 (@danielpatrickhug)
offline mode for firewalled envs #10407 (@stas00)
fix tf doc bug #10570 (@Sniper970119)
[run_seq2seq] fix nltk lookup #10585 (@stas00)
Fix typo in docstring for pipeline #10591 (@silvershine157)
wrong model used for BART Summarization example #10582 (@orena1)
[M2M100] fix positional embeddings #10590 (@patil-suraj)
Enable torch 1.8.0 on GPU CI #10593 (@LysandreJik)
tokenization_marian.py: use current_spm for decoding #10357 (@Mehrad0711)
[trainer] fix double wrapping + test #10583 (@stas00)
Fix version control with anchors #10595 (@sgugger)
offline mode for firewalled envs (part 2) #10569 (@stas00)
[examples tests] various fixes #10584 (@stas00)
Added max_sample_ arguments #10551 (@bhadreshpsavani)
[examples tests on multigpu] resolving require_torch_non_multi_gpu_but_fix_me #10561 (@stas00)
Check layer types for Optimizer construction #10598 (@sgugger)
Speedup tf tests #10601 (@LysandreJik)
[docs] How to solve "Title level inconsistent" sphinx error #10600 (@stas00)
[FeatureExtractorSavingUtils] Refactor PretrainedFeatureExtractor #10594 (@patrickvonplaten)
fix flaky m2m100 test #10604 (@patil-suraj)
[examples template] added max_sample args and metrics changes #10602 (@bhadreshpsavani)
Fairscale FSDP fix model save #10596 (@sgugger)
Fix tests of TrainerCallback #10615 (@sgugger)
Fixes an issue in text-classification where MNLI eval/test datasets are not being preprocessed. #10621 (@allenwang28)
[M2M100] remove final_logits_bias #10606 (@patil-suraj)
Add new GLUE example with no Trainer. #10555 (@sgugger)
Copy tokenizer files in each of their repo #10624 (@sgugger)
Document Trainer limitation on custom models #10635 (@sgugger)
Fix Longformer tokenizer filename #10653 (@LysandreJik)
Update README.md #10647 (@Arvid-pku)
Ensure metric results are JSON-serializable #10632 (@sgugger)
S2S + M2M100 should be available in tokenization_auto #10657 (@LysandreJik)
Remove special treatment for custom vocab files #10637 (@sgugger)
[S2T] fix example in docs #10667 (@patil-suraj)
W2v2 test require torch #10665 (@LysandreJik)
Fix Marian/TFMarian tokenization tests #10661 (@LysandreJik)
Fixes Pegasus tokenization tests #10671 (@LysandreJik)
Onnx fix test #10663 (@mfuntowicz)
Fix integration slow tests #10670 (@sgugger)
Specify minimum version for sacrebleu #10662 (@LysandreJik)
Add DeBERTa to MODEL_FOR_PRETRAINING_MAPPING #10668 (@jeswan)
Fix broken link #10656 (@WybeKoper)
fix typing error for HfArgumentParser for Optional[bool] #10672 (@bfineran)
MT5 integration test: adjust loss difference #10669 (@LysandreJik)
Adding new parameter to generate: max_time. #9846 (@Narsil)
TensorFlow tests: having from_pt set to True requires torch to be installed. #10664 (@LysandreJik)
Add auto_wrap option in fairscale integration #10673 (@sgugger)
fix: #10628 expanduser path in TrainingArguments #10660 (@PaulLerner)
Pass encoder outputs into GenerationMixin #10599 (@ymfa)
[wip] [deepspeed] AdamW is now supported by default #9624 (@stas00)
[Tests] RAG #10679 (@patrickvonplaten)
enable loading Mbart50Tokenizer with AutoTokenizer #10690 (@patil-suraj)
Wrong link to super class #10709 (@cronoik)
Distributed barrier before loading model #10685 (@sgugger)
GPT2DoubleHeadsModel made parallelizable #10658 (@ishalyminov)
split seq2seq script into summarization & translation #10611 (@theo-m)
Adding required flags to non-default arguments in hf_argparser #10688 (@Craigacp)
Fix backward compatibility with EvaluationStrategy #10718 (@sgugger)
Tests run on Docker #10681 (@LysandreJik)
Rename zero-shot pipeline multi_class argument #10727 (@joeddav)
Add minimum version check in examples #10724 (@sgugger)
independent training / eval with local files #10710 (@riklopfer)
Flax testing should not run the full torch test suite #10725 (@patrickvonplaten)
This patch fixes an issue with the conversion for ConvBERT models: https://github.com/huggingface/transformers/pull/10314.
This patch fixes an issue with the conversion for ConvBERT models: https://github.com/huggingface/transformers/pull/10314.
This patch release fixes the RAG model (#10094) and the detection of whether faiss is available
This patch release fixes the RAG model (#10094) and the detection of whether faiss is available (#10103)
Wav2Vec2ForMaskedLM is kept for backwards compatibility but is deprecated.
This patch release modifies the API of the Wav2Vec2 model: the Wav2Vec2ForCTC was added as a replacement of Wav2Vec2ForMaskedLM. Wav2Vec2ForMaskedLM is kept for backwards compatibility but is deprecated.
Deprecate model_path in Trainer.train #9854 (@sgugger)
Two new models are released as part of the Wav2Vec2 implementation: Wav2Vec2Model and Wav2Vec2ForMaskedLM, in PyTorch.
Wav2Vec2 is a multi-modal model, combining speech and text. It's the first multi-modal model of its kind we welcome in Transformers.
The Wav2Vec2 model was proposed in wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations by Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=wav2vec2
Available notebooks:
Contributions:
Future Additions
The ConvBERT model was proposed in ConvBERT: Improving BERT with Span-based Dynamic Convolution by Zihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen, Jiashi Feng, Shuicheng Yan.
Six new models are released as part of the ConvBERT implementation: ConvBertModel, ConvBertForMaskedLM, ConvBertForSequenceClassification, ConvBertForTokenClassification, ConvBertForQuestionAnswering and ConvBertForMultipleChoice. These models are available both in PyTorch and TensorFlow.
Contributions:
The BORT model was proposed in Optimal Subarchitecture Extraction for BERT by Amazon's Adrian de Wynter and Daniel J. Perry. It is an optimal subset of architectural parameters for the BERT, which the authors refer to as “Bort”.
The BORT model can be loaded directly in the BERT architecture, therefore all BERT model heads are available for BORT.
Contributions:
When executing a script with Trainer using Amazon SageMaker and enabling SageMaker's data parallelism library, Trainer will automatically use the smdistributed library. All maintained examples have been tested with this functionality. Here is an overview of SageMaker data parallelism library.
A new Community Page has been added to the docs. These contain all the notebooks contributed by the community, as well as some community projects built around Transformers. Feel free to open a PR if you want your project to be showcased!
DeBERTa now has more model heads available.
BART, mBART, Marian, Pegasus and Blenderbot now have decoder-only model architectures. They can therefore be used in decoder-only settings.
ProphetNetForCausalLM #9128 (@sadakmed)None.
past_key_values in GPT-2 #9596 (@forest1988)report_to training arguments to control the integrations used #9735 (@sgugger)Trainer.hyperparameter_search docstring #9762 (@sorami)skip_special_tokens=True to FillMaskPipeline #9783 (@Narsil)test_head_masking = True flags in test files #9858 (@stancld)return_full_text parameter to TextGenerationPipeline. #9852 (@Narsil)from_slow in fast tokenizers build and fixes some bugs #9987 (@sgugger)encoder_no_repeat_ngram_size to generate. #9984 (@Narsil)Deprecate model_path in Trainer.train #9854 (@sgugger)
Two new models are released as part of the Wav2Vec2 implementation: Wav2Vec2Model and Wav2Vec2ForMaskedLM, in PyTorch.
Wav2Vec2 is a multi-modal model, combining speech and text. It's the first multi-modal model of its kind we welcome in Transformers.
The Wav2Vec2 model was proposed in wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations by Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=wav2vec2
Available notebooks:
Contributions:
Future Additions
The ConvBERT model was proposed in ConvBERT: Improving BERT with Span-based Dynamic Convolution by Zihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen, Jiashi Feng, Shuicheng Yan.
Six new models are released as part of the ConvBERT implementation: ConvBertModel, ConvBertForMaskedLM, ConvBertForSequenceClassification, ConvBertForTokenClassification, ConvBertForQuestionAnswering and ConvBertForMultipleChoice. These models are available both in PyTorch and TensorFlow.
Contributions:
The BORT model was proposed in Optimal Subarchitecture Extraction for BERT by Amazon's Adrian de Wynter and Daniel J. Perry. It is an optimal subset of architectural parameters for the BERT, which the authors refer to as “Bort”.
The BORT model can be loaded directly in the BERT architecture, therefore all BERT model heads are available for BORT.
Contributions:
When executing a script with Trainer using Amazon SageMaker and enabling SageMaker's data parallelism library, Trainer will automatically use the smdistributed library. All maintained examples have been tested with this functionality. Here is an overview of SageMaker data parallelism library.
A new Community Page has been added to the docs. These contain all the notebooks contributed by the community, as well as some community projects built around Transformers. Feel free to open a PR if you want your project to be showcased!
DeBERTa now has more model heads available.
BART, mBART, Marian, Pegasus and Blenderbot now have decoder-only model architectures. They can therefore be used in decoder-only settings.
ProphetNetForCausalLM #9128 (@sadakmed)None.
past_key_values in GPT-2 #9596 (@forest1988)report_to training arguments to control the integrations used #9735 (@sgugger)Trainer.hyperparameter_search docstring #9762 (@sorami)skip_special_tokens=True to FillMaskPipeline #9783 (@Narsil)test_head_masking = True flags in test files #9858 (@stancld)return_full_text parameter to TextGenerationPipeline. #9852 (@Narsil)from_slow in fast tokenizers build and fixes some bugs #9987 (@sgugger)encoder_no_repeat_ngram_size to generate. #9984 (@Narsil)[TF Led] Fix wrong decoder attention mask behavior #9601 (@patrickvonplaten )
This patch contains two fixes:
This patch contains three fixes:
This patch contains three fixes:
There are no breaking changes between the previous version and this one. This will be the first version to require TensorFlow >= 2.3.
Four new models are released as part of the LED implementation: LEDModel, LEDForConditionalGeneration, LEDForSequenceClassification, LEDForQuestionAnswering, in PyTorch. The first two models have a TensorFlow version.
LED is the encoder-decoder variant of the Longformer model by allenai.
The LED model was proposed in Longformer: The Long-Document Transformer by Iz Beltagy, Matthew E. Peters, Arman Cohan.
Compatible checkpoints can be found on the Hub: https://huggingface.co/models?filter=led
Available notebooks:
Contributions:
The PyTorch generation function now allows to return:
scores - the logits generated at each stepattentions - all attention weights at each generation stephidden_states - all hidden states at each generation stepby simply adding return_dict_in_generate to the config or as an input to .generate()
Tweet:
Notebooks for a better explanation:
PR:
The TensorFlow version of the BERT-like models have been updated and are now twice as fast as the previous versions.
This version introduces a new API for TensorFlow saved models, which can now be exported with model.save_pretrained("path", saved_model=True) and easily loaded into a TensorFlow Serving environment.
Initial support for DeepSpeed to accelerate distributed training on several GPUs. This is an experimental feature that hasn't been fully tested yet, but early results are very encouraging (see this comment). Stay tuned for more details in the coming weeks!
The encoder-decoder version of the templates is now part of Transformers! Adding an encoder-decoder model is made very easy with this addition. More information can be found in the README.
The initialization process has been changed to only import what is required. Therefore, when using only PyTorch models, TensorFlow will not be imported and vice-versa. In the best situations the import of a transformers model now takes only a few hundreds of milliseconds (~200ms) compared to more than a few seconds (~3s) in previous versions.
Some models now have improved documentation. The LayoutLM model has seen a general overhaul in its documentation thanks to @NielsRogge.
The tokenizer-only models Bertweet, Herbert and Phobert now have their own documentation pages thanks to @Qbiwan.
There are no breaking changes between the previous version and this one. This will be the first version to require TensorFlow >= 2.3.
label_smoothing_factor training arg #9282 (@sgugger)past_key_values return a tuple of tuple as a default #9381 (@patrickvonplaten)--model_parallel #9451 (@stas00)prepare_seq2seq_batch #9524 (@sgugger)[examples/seq2seq] fix PL deprecation warning #8577 (@stas00)
Four new models are released as part of the TAPAS implementation: TapasModel, TapasForQuestionAnswering, TapasForMaskedLM and TapasForSequenceClassification, in PyTorch.
TAPAS is a question answering model, used to answer queries given a table. It is a multi-modal model, joining text for the query and tabular data.
The TAPAS model was proposed in TAPAS: Weakly Supervised Table Parsing via Pre-training by Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno and Julian Martin Eisenschlos.
Six new models are released as part of the MPNet implementation: MPNetModel, MPNetForMaskedLM, MPNetForSequenceClassification, MPNetForMultipleChoice, MPNetForTokenClassification, MPNetForQuestionAnswering, in both PyTorch and TensorFlow.
MPNet introduces a novel self-supervised objective named masked and permuted language modeling for language understanding. It inherits the advantages of both the masked language modeling (MLM) and the permuted language modeling (PLM) to addresses the limitations of MLM/PLM, and further reduce the inconsistency between the pre-training and fine-tuning paradigms.
The MPNet model was proposed in MPNet: Masked and Permuted Pre-training for Language Understanding by Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, Tie-Yan Liu.
Model parallelism is introduced, allowing users to load very large models on two or more GPUs by spreading the model layers over them. This can allow GPU training even for very large models.
Transformers welcome their first conda releases, with v4.0.0, v4.0.1 and v4.1.0. The conda packages are now officially maintained on the huggingface channel.
For the first time, very large models can be uploaded to the model hub, by using multi-part uploads.
We introduced a refactored SQuAD example & notebook, which is faster and simpler than the previous scripts.
The example directory has been re-ordered as we introduce the separation between "examples", which are maintained examples showcasing how to do one specific task, and "research projects", which are bigger projects and maintained by the community.
We introduce support for fariscale's ShardedDDP in the Trainer, allowing reduced memory usage when training models in a distributed fashion.
The BARThez model is a French variant of the BART model. We welcome its specific tokenizer to the library and multiple checkpoints to the modelhub.
disable_ngram_loss fix for prophetnet #8554 (@Zhylkaaa)PretrainedConfig and PreTrainedModel #8770 (@gcompagnoni)evalutate_during_training #8852 (@sgugger)parallel_mode property to TrainingArguments #8877 (@sgugger)modeling_bert import resulting in ImportError #8931 (@machelreid)ModelOutput pickle-able #8989 (@sgugger)encoder_hidden_states and encoder_attention_mask #8972 (@guillaume-be)Nothing published for this version
This patch releases introduces:
This patch releases introduces:
Version v4.0.0 introduces several breaking changes that were necessary.
Version v4.0.0 introduces several breaking changes that were necessary.
The python and rust tokenizers have roughly the same API, but the rust tokenizers have a more complete feature set. The main breaking change is the handling of overflowing tokens between the python and rust tokenizers.
grouped_entities flag.use_fast flag by setting it to False:In version v3.x:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("xxx")
to obtain the same in version v4.x:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("xxx", use_fast=False)
The requirement on the SentencePiece dependency has been lifted from the setup.py. This is done so that we may have a channel on anaconda cloud without relying on conda-forge. This means that the tokenizers that depend on the SentencePiece library will not be available with a standard transformers installation.
This includes the slow versions of:
XLNetTokenizerAlbertTokenizerCamembertTokenizerMBartTokenizerPegasusTokenizerT5TokenizerReformerTokenizerXLMRobertaTokenizerIn order to obtain the same behavior as version v3.x, you should install sentencepiece additionally:
In version v3.x:
pip install transformers
to obtain the same in version v4.x:
pip install transformers[sentencepiece]
or
pip install transformers sentencepiece
The past and foreseeable addition of new models means that the number of files in the directory src/transformers keeps growing and becomes harder to navigate and understand. We made the choice to put each model and the files accompanying it in their own sub-directories.
This is a breaking change as importing intermediary layers using a model's module directly needs to be done via a different path.
In order to obtain the same behavior as version v3.x, you should update the path used to access the layers.
In version v3.x:
from transformers.modeling_bert import BertLayer
to obtain the same in version v4.x:
from transformers.models.bert.modeling_bert import BertLayer
return_dict argument to True by defaultThe return_dict argument enables the return of named-tuples-like python objects containing the model outputs, instead of the standard tuples. This object is self-documented as keys can be used to retrieve values, while also behaving as a tuple as users may retrieve objects by index or by slice.
This is a breaking change as the limitation of that tuple is that it cannot be unpacked: value0, value1 = outputs will not work.
In order to obtain the same behavior as version v3.x, you should specify the return_dict argument to False, either in the model configuration or during the forward pass.
In version v3.x:
outputs = model(**inputs)
to obtain the same in version v4.x:
outputs = model(**inputs, return_dict=False)
Attributes that were deprecated have been removed if they had been deprecated for at least a month. The full list of deprecated attributes can be found in #8604.
Here is a list of these attributes/methods/arguments and what their replacements should be:
In several models, the labels become consistent with the other models:
masked_lm_labels becomes labels in AlbertForMaskedLM and AlbertForPreTraining.masked_lm_labels becomes labels in BertForMaskedLM and BertForPreTraining.masked_lm_labels becomes labels in DistilBertForMaskedLM.masked_lm_labels becomes labels in ElectraForMaskedLM.masked_lm_labels becomes labels in LongformerForMaskedLM.masked_lm_labels becomes labels in MobileBertForMaskedLM.masked_lm_labels becomes labels in RobertaForMaskedLM.lm_labels becomes labels in BartForConditionalGeneration.lm_labels becomes labels in GPT2DoubleHeadsModel.lm_labels becomes labels in OpenAIGPTDoubleHeadsModel.lm_labels becomes labels in T5ForConditionalGeneration.In several models, the caching mechanism becomes consistent with the other models:
decoder_cached_states becomes past_key_values in all BART-like, FSMT and T5 models.decoder_past_key_values becomes past_key_values in all BART-like, FSMT and T5 models.past becomes past_key_values in all CTRL models.past becomes past_key_values in all GPT-2 models.Regarding the tokenizer classes:
max_len becomes model_max_length.return_lengths becomes return_length.is_pretokenized becomes is_split_into_words.Regarding the Trainer class:
Trainer argument tb_writer is removed in favor of the callback TensorBoardCallback(tb_writer=...).Trainer argument prediction_loss_only is removed in favor of the class argument args.prediction_loss_only.Trainer attribute data_collator should be a callable.Trainer method _log is deprecated in favor of log.Trainer method _training_step is deprecated in favor of training_step.Trainer method _prediction_loop is deprecated in favor of prediction_loop.Trainer method is_local_master is deprecated in favor of is_local_process_zero.Trainer method is_world_master is deprecated in favor of is_world_process_zero.Regarding the TFTrainer class:
TFTrainer argument prediction_loss_only is removed in favor of the class argument args.prediction_loss_only.Trainer method _log is deprecated in favor of log.TFTrainer method _prediction_loop is deprecated in favor of prediction_loop.TFTrainer method _setup_wandb is deprecated in favor of setup_wandb.TFTrainer method _run_model is deprecated in favor of run_model.Regarding the TrainerArgument and TFTrainerArgument classes:
TrainerArgument argument evaluate_during_training is deprecated in favor of evaluation_strategy.TFTrainerArgument argument evaluate_during_training is deprecated in favor of evaluation_strategy.Regarding the Transfo-XL model:
tie_weight becomes tie_words_embeddings.reset_length becomes reset_memory_length.Regarding pipelines:
FillMaskPipeline argument topk becomes top_k.Version 4.0.0 will be the first to include the experimental feature of model templates. These model templates aim to facilitate the addition of new models to the library by doing most of the work: generating the model/configuration/tokenization/test files that fit the API, with respect to the choice the user has made in terms of naming and functionality.
This release includes a model template for the encoder model (similar to the BERT architecture). Generating a model using the template will generate the files, put them at the appropriate location, reference them throughout the code-base, and generate a working test suite. The user should then only modify the files to their liking, rather than creating the model from scratch.
Feedback welcome, get started from the README here.
The T5v1.1 is an improved version of the original T5 model, see here: https://github.com/google-research/text-to-text-transfer-transformer/blob/master/released_checkpoints.md
The multilingual T5 model (mT5) was presented in https://arxiv.org/abs/2010.11934 and is based on the T5v1.1 architecture.
Multiple pre-trained checkpoints have been added to the library:
Relevant pull requests:
The DPR model has been added in TensorFlow to match its PyTorch counterpart by @ratthachat
Additional heads have been added to the TensorFlow Longformer implementation: SequenceClassification, MultipleChoice and TokenClassification
pipeline docstring #8428 (@bryant1410)return_dict to True by default. #8530 (@sgugger)evaluate_during_training #8852 (@sgugger)Nothing published for this version
Fix a typo that raised an error instead of a deprecation warning.
Fix a typo that raised an error instead of a deprecation warning.
Remove deprecated arguments from new run_clm #8197 (@sgugger)
generate methodWe host more and more of the community's models which is awesome ❤️. To scale this sharing, we needed to change the infra to both support more models, and unlock new powerful features.
To that effect, we have rebuilt the storage backend that we use for models (currently S3), to our own git repos (using S3 as a git-lfs endpoint for large files), with one model = one repo.
The benefits of this switch are:
Let's dive in to the actual changes:
You'll now see a "Browse files and versions" tab or button on each model page. (design is not final, we'll make it more prominent/streamlined in the near future)
This is what this page looks like:
The UX should look familiar and self-explanatory, but we'll add more ML-specific features in the future.
You can:
The PR to enable this new storage mode in the transformers library is available here: https://github.com/huggingface/transformers/pull/8324
This PR has two parts:
1. changes to the file downloading code used in from_pretrained() methods to use the new file URLs.
Large files are stored in an S3 bucket and served by Cloudfront so downloads should be as fast as they are right now.
In addition, you now have a way to pin a specific version of a model, to a commit hash, tag or branch.
For instance:
tokenizer = AutoTokenizer.from_pretrained(
"julien-c/EsperBERTo-small",
revision="v2.0.1" # tag name, or branch name, or commit hash
)
Finally, the networking code is more robust and doesn't gobble up errors anymore, so in case you have trouble downloading a specific file you'll know exactly why.
2. changes to the model upload CLI to create a model repo then be able to git clone and git push to it.
We are intentionally not wrapping git too much because we expect most model authors to be familiar with git (and possibly git-lfs), let us know if not the case.
To create a repo:
transformers-cli repo create your-model-name
Then you'll get a repo url that you'll be able to clone:
git clone https://huggingface.co/username/your-model-name
# Then commit as usual
cd your-model-name
echo "hello" >> README.md
git add . && git commit -m "Update from $USER"
A nice side effect of the new system on the upload side is that file uploading should be more robust for very large files (hello T5!) as git-lfs handles the networking code.
By the way, again, every model is its own repo. So you can git clone any public model if you'd like:
git clone https://huggingface.co/gpt2
But you won't be able to push unless it's one of your models (or one of your orgs').
v3.5.0, or to build from master.
Alternatively, in the next week or so we'll add the ability to create a repo from the website directly so you'll be able to push even without the transformers library.We'working on giving examples on how to leverage the 🤗 Datasets library and the Trainer API. Those scripts are meant as examples easy to customize, with lots of comments explaining the various steps. The following tasks are now covered:
A child of Trainer specialized for training seq2seq models, from @patil-suraj, @stas00 and @sshleifer. Accessible through examples/seq2seq/finetune_trainer.py. API is similar to examples/seq2seq/finetune.py, but API support is better. Example scripts are in examples/seq2seq/builtin_trainer.
Re-run experiments from the paper here
generate() functionThe generate() method now has a new design so that the user can directly call upon the methods
sample(), greedy_search(), beam_search() and beam_sample(). The code was made more readable, and beam search was sped-up by ca. 5-10%.
Refactoring the generate() function #6949 (@patrickvonplaten)
do_sample=True is not deterministic. #7947 (@patrickvonplaten)CallbackHandler.callback_list #8052 (@harupy)Update Code example according to deprecation of AutoModeWithLMHead #7555 (@jshamg)
Two new models are released as part of the ProphetNet implementation: ProphetNet and XLM-ProphetNet.
ProphetNet is an encoder-decoder model and can predict n-future tokens for “ngram” language modeling instead of just the next token.
XLM-ProphetNet is an encoder-decoder model with an identical architecture to ProhpetNet, but the model was trained on the multi-lingual “wiki100” Wikipedia dump.
The ProphetNet model was proposed in ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training, by Yu Yan, Weizhen Qi, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, Ming Zhou on 13 Jan, 2020.
It was added to the library in PyTorch with the following checkpoints:
microsoft/xprophetnet-large-wiki100-cased-xglue-ntgmicrosoft/prophetnet-large-uncasedmicrosoft/prophetnet-large-uncased-cnndmmicrosoft/xprophetnet-large-wiki100-casedmicrosoft/xprophetnet-large-wiki100-cased-xglue-qgContributions:
Blenderbot is an encoder-decoder model for open-domain chat. It uses a standard seq2seq model transformer-based architecture.
The Blender chatbot model was proposed in Recipes for building an open-domain chatbot Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M. Smith, Y-Lan Boureau, Jason Weston on 30 Apr 2020.
It was added to the library in PyTorch with the following checkpoints:
facebook/blenderbot-90Mfacebook/blenderbot-3BContributions:
The SqueezeBERT model was proposed in SqueezeBERT: What can computer vision teach NLP about efficient neural networks? by Forrest N. Iandola, Albert E. Shaw, Ravi Krishna, Kurt W. Keutzer. It’s a bidirectional transformer similar to the BERT model. The key difference between the BERT architecture and the SqueezeBERT architecture is that SqueezeBERT uses grouped convolutions instead of fully-connected layers for the Q, K, V and FFN layers.
It was added to the library in PyTorch with the following checkpoints:
squeezebert/squeezebert-mnlisqueezebert/squeezebert-uncasedsqueezebert/squeezebert-mnli-headlessContributions:
The DeBERTa model was proposed in DeBERTa: Decoding-enhanced BERT with Disentangled Attention by Pengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu Chen It is based on Google’s BERT model released in 2018 and Facebook’s RoBERTa model released in 2019.
It was added to the library in PyTorch with the following checkpoints:
microsoft/deberta-basemicrosoft/deberta-largeContributions:
Support for SentencePiece is now part of the tokenizers library! Thanks to this we now have near-full support of fast tokenizers in the library.
With this new feature, we slightly change the paradigm regarding installation:
SentencePiece is now an optional dependency, paving the way to a fully-featured conda install in the near future
Tokenizers is now also an optional dependency, making it possible to install and use the library even when rust cannot be compiled on the machine.
[Dependencies|tokenizers] Make both SentencePiece and Tokenizers optional dependencies #7659 (@thomwolf)
The main __init__ has been improved to always import the same functions and classes. If someone then tries to use a class that requires an optional dependency, an ImportError will be raised at init (with instructions on how to install the missing dependency) #7537 (@sgugger)
TrainerThe Trainer API has been improved to work with models requiring several labels or returning several outputs, and to have clearer progress tracking. A new TrainerCallback class has been added to allow the user to easily customize the default training loop.
store_xxx on optional bools #7786 (@sgugger)A child of Trainer specialized for training seq2seq models, from @patil-suraj and @sshleifer. Accessible through examples/seq2seq/finetune_trainer.py.
examples/seq2seq/builtin_trainer/examples/seq2seq/finetune.py, but better TPU support.model.generate in pytorch on a large dataset and split the work across multiple GPUs, using examples/seq2seq/run_distributed_eval.pySquadProcessor #7616 (@phiyodr)cross_entropy #7841 (@katarinaslama)decoder_config used before intialisation #7903 (@ayubSubhaniya)Fixes errors due to the name conflicts between the datasets library and local folder or modules named datasets.
Fixes errors due to the name conflicts between the datasets library and local folder or modules named datasets.
The RAG model is a retrieval-augmented generation model that can be leveraged for question-answering tasks using RagTokenForGeneration or RagSequenceF
The RAG model is a retrieval-augmented generation model that can be leveraged for question-answering tasks using RagTokenForGeneration or RagSequenceForGeneration as proposed in Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks by Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela.
It was added to the library in PyTorch with the following checkpoints:
facebook/rag-token-nqfacebook/rag-sequence-nqfacebook/rag-token-basefacebook/rag-sequence-baseattention_mask to RAG generate #7373 (@patrickvonplaten)num_beams and bos_token_id in Rag Sequence generation #7386 (@patrickvonplaten)examples/seq2seq from rag #7395 (@ola13)no_... to their positive form #7075 (@fmcurti)Trainer hung while saving model in distributed training #7365 (@TevenLeScao)type instead of isinstance #7363 (@LysandreJik)fix deprecation warnings #7033 (@stas00)
The BertGeneration model is a BERT model that can be leveraged for sequence-to-sequence tasks using EncoderDecoderModel as proposed in Leveraging Pre-trained Checkpoints for Sequence Generation Tasks by Sascha Rothe, Shashi Narayan, Aliaksei Severyn.
It was added to the library in PyTorch with the following checkpoints:
google/roberta2roberta_L-24_bbcgoogle/roberta2roberta_L-24_gigawordgoogle/roberta2roberta_L-24_cnn_daily_mailgoogle/roberta2roberta_L-24_discofusegoogle/roberta2roberta_L-24_wikisplitgoogle/bert2bert_L-24_wmt_de_engoogle/bert2bert_L-24_wmt_en_deFSMT (FairSeq MachineTranslation) models were introduced in Facebook FAIR’s WMT19 News Translation Task Submission by Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, Sergey Edunov.
It was added to the library in PyTorch, with the following checkpoints:
facebook/wmt19-en-rufacebook/wmt19-en-defacebook/wmt19-ru-enfacebook/wmt19-de-enThe LayoutLM model was proposed in LayoutLM: Pre-training of Text and Layout for Document Image Understandin by Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. It’s a simple but effective pre-training method of text and layout for document image understanding and information extraction tasks, such as form understanding and receipt understanding.
It was added to the library in PyTorch with the following checkpoints:
layoutlm-base-uncasedlayoutlm-large-uncasedThe Funnel Transformer model was proposed in the paper Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing. It is a bidirectional transformer model, like BERT, but with a pooling operation after each block of layers, a bit like in traditional convolutional neural networks (CNN) in computer vision.
It was added to the library in both PyTorch and TensorFlow, with the following checkpoints:
funnel-transformer/smallfunnel-transformer/small-basefunnel-transformer/mediumfunnel-transformer/medium-basefunnel-transformer/intermediatefunnel-transformer/intermediate-basefunnel-transformer/largefunnel-transformer/large-basefunnel-transformer/xlargefunnel-transformer/xlarge-baseThe LXMERT model was proposed in LXMERT: Learning Cross-Modality Encoder Representations from Transformers by Hao Tan & Mohit Bansal. It is a series of bidirectional transformer encoders (one for the vision modality, one for the language modality, and then one to fuse both modalities) pre-trained using a combination of masked language modeling, visual-language text alignment, ROI-feature regression, masked visual-attribute modeling, masked visual-object modeling, and visual-question answering objectives. The pretraining consists of multiple multi-modal datasets: MSCOCO, Visual-Genome + Visual-Genome Question Answering, VQA 2.0, and GQA.
It was added to the library in TensorFlow with the following checkpoints:
unc-nlp/lxmert-base-uncasedunc-nlp/lxmert-vqa-uncasedunc-nlp/lxmert-gqa-uncasedThe following pipeline was added to the library:
The following community notebooks were contributed to the library:
An additional encoder-decoder architecture was added:
ModelOutputs #6735 (@patrickvonplaten)None, s/True/ :obj:`True/, etc. #6956 (@stas00)None #6984 (@LysandreJik)[mbart] prepare_translation_batch passes kwargs to allow DeprecationWarning #5581 (@sshleifer)
The Pegasus model from PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization by Jingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. Liu, was added to the library in PyTorch.
Model implemented as a collaboration between Jingqing Zhang and @sshleifer in #6340
The DPR model from Dense Passage Retrieval for Open-Domain Question Answering by Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih was added to the library in PyTorch.
The DeeBERT model from DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference by Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, Jimmy Lin has been added to the examples/ folder alongside its training script, in PyTorch.
As well as returning tuples, PyTorch and TensorFlow models now return a subclass of ModelOutput that is appropriate. A ModelOutput is a dataclass containing all model returns. This allows for easier inspection, and for self-documenting model outputs.
Models return tuples by default, and return self-documented outputs if the return_dict configuration flag is set to True or if the return_dict=True keyword argument is passed to the forward/call method.
Summary of the behavior:
# The new outputs are opt-in, you have to activate them explicitly with `return_dict=True`
# Either at instantiation
model = BertForSequenceClassification.from_pretrained('bert-base-cased', return_dict=True)
# Or when calling the model
output = model(**inputs, return_dict=True)
# You can access the elements of the outputs with
# (1) named attributes
loss = outputs.loss
logits = outputs.logits
# (2) their names as strings like a dict
loss = outputs["loss"]
logits = outputs["logits"]
# (3) their index as integers or slices in the pre-3.1.0 outputs tuples
loss = outputs[0]
logits = outputs[1]
loss, logits = outputs[:2]
# One **breaking behavior** of these new outputs (which is the reason you have to opt-in to use these new outputs:
# Iterating on the outputs now return the names (keys) instead of the values:
print([element for element in outputs])
>>> ['loss', 'logits']
# Thus you cannot unpack the output like pre-3.1.0 (you get the string names instead of the values):
# (But you can query a slice like indicated in (3) above)
loss_keys, logits_key = outputs
The encoder-decoder framework has been enhanced to allow more encoder decoder model combinations, e.g.: Bert2Bert, Bert2GPT2, Roberta2Roberta, Longformer2Roberta, ....
As we continue working towards having TensorFlow be a first-class citizen, we continually improve on our TensorFlow API and models.
The mBART model from Multilingual Denoising Pre-training for Neural Machine Translation was can now be accessed through MBartForConditionalGeneration.
prepare_seq2seq_batch method that makes batches for sequence to sequence trianing.PRs:
Several new documentation pages have been added and older documentation has been tweaked to be more accurate and understandable. An open in colab button has been added on the tutorial pages.
New additions to the Trainer
The following model architectures have been added to the library
Thanks to @zcain117 we now have access to TPU CI for the PyTorch/xla framework. This enables regression testing on the TPU aspects of the Trainer, and offers very simple regression testing on model training performance.
New pipelines have been added:
Logging is now centralized. The library offers methods to handle the verbosity level of all loggers contained in the library. [Link to logging doc here]:
[Reformer] Adapt Reformer MaskedLM Attn mask #5560 (@patrickvonplaten)
Make T5 compatible with ONNX #5518 (@abelriboulot)
[Bart] enable test_torchscript, update test_tie_weights #5457 (@sshleifer)
[docs] fix model_doc links in model summary #5566 (@patil-suraj)
[Benchmark] Readme for benchmark #5363 (@patrickvonplaten)
Fix Inconsistent NER Grouping (Pipeline) #4987 (@enzoampil)
QA pipeline BART compatible #5496 (@mfuntowicz)
More explicit error when failing to tensorize overflowing tokens #5633 (@LysandreJik)
Should check that torch TPU is available #5636 (@LysandreJik)
Add forum link in the docs #5637 (@sgugger)
Fixed TextGenerationPipeline on torch + GPU #5629 (@TevenLeScao)
Fixed use of memories in XLNet (caching for language generation + warning when loading improper memoryless model) #5632 (@TevenLeScao)
[squad] add version tag to squad cache #5669 (@lazovich)
Deprecate old past arguments #5671 (@sgugger)
Pipeline model type check #5679 (@JetRunner)
rename the functions to match the rest of the test convention #5692 (@stas00)
doc improvements #5688 (@stas00)
Fix Trainer in DataParallel setting #5685 (@sgugger)
[Longformer] fix longformer global attention output #5659 (@patrickvonplaten)
[Fix] github actions CI by reverting #5138 #5686 (@sshleifer)
[Reformer classification head] Implement the reformer model classification head for text classification #5198 (@as-stevens)
Cleanup bart caching logic #5640 (@sshleifer)
[AutoModels] Fix config params handling of all PT and TF AutoModels #5665 (@patrickvonplaten)
[cleanup] T5 test, warnings #5761 (@sshleifer)
[fix] T5 ONNX test: model.to(torch_device) #5769 (@mfuntowicz)
[Benchmark] fix benchmark non standard model #5801 (@patrickvonplaten)
[Benchmark] Fix models without architectures param in config #5808 (@patrickvonplaten)
[Longformer] fix longformer slow-down #5811 (@patrickvonplaten)
[seq2seq] pack_dataset.py rewrites dataset in max_tokens format #5819 (@sshleifer)
[seq2seq] Don't copy self.source in sortishsampler #5818 (@sshleifer)
[cleanups] make Marian save as Marian #5830 (@sshleifer)
[Reformer] - Cache hidden states and buckets to speed up inference #5578 (@patrickvonplaten)
Lightning Updates for v0.8.5 #5798 (@nateraw)
Update tokenizers to 0.8.1.rc to fix Mac OS X issues #5867 (@sepal)
Xlnet outputs #5883 (@TevenLeScao)
DataParallel fixes #5733 (@stas00)
[cleanup] squad processor #5868 (@sshleifer)
Improve doc of use_cache #5912 (@sgugger)
[Fix] seq2seq pack_dataset.py actually packs #5913 (@sshleifer)
Add AlbertForPretraining to doc #5914 (@sgugger)
DataParallel fix: multi gpu evaluation #5926 (@csarron)
Clarify arg class #5916 (@sgugger)
[CI] self-scheduled runner tests examples/ #5927 (@sshleifer)
Update doc to new model outputs #5946 (@sgugger)
[CI] Install examples/requirements.txt #5956 (@sshleifer)
Expose padding_strategy on squad processor to fix QA pipeline performance regression #5932 (@mfuntowicz)
[docs] Add integration test example to copy pasta template #5961 (@sshleifer)
Cleanup Trainer and expose customization points #5982 (@sgugger)
Avoid unnecessary warnings when loading pretrained model #5922 (@sgugger)
Ensure OpenAI GPT position_ids is correctly initialized and registered at init. #5773 (@mfuntowicz)
[CI] Don't test apex #6021 (@sshleifer)
add a summary report flag for run_examples on CI #6035 (@stas00)
don't complain about missing W&B when WANDB_DISABLED=true #6036 (@stas00)
Allow to set Adam beta1, beta2 in TrainingArgs #5592 (@gonglinyuan)
Fix the return documentation rendering for all model outputs #6022 (@sgugger)
Fix typo (model saving TF) #5734 (@Colanim)
Add new AutoModel classes in pipeline #6062 (@patil-suraj)
[pack_dataset] don't sort before packing, only pack train #5954 (@sshleifer)
CL util to convert models to fp16 before upload #5953 (@sshleifer)
Add fire to setup.cfg to make isort happy #6066 (@sgugger)
[fix] no warning for position_ids buffer #6063 (@sshleifer)
Pipelines should use tuples instead of namedtuples #6061 (@LysandreJik)
Moving transformers package import statements to relative imports in some files #5796 (@afcruzs)
github issue template suggests who to tag #5790 (@sshleifer)
Make all data collators accept dict #6065 (@sgugger)
Add inference widget examples #5825 (@clmnt)
[s2s] Delete useless method, log tokens_per_batch #6081 (@sshleifer)
Logs should not be hidden behind a logger.info #6097 (@LysandreJik)
Fix zero-shot pipeline single seq output shape #6104 (@joeddav)
[fix] add bart to LM_MAPPING #6099 (@sshleifer)
[Fix] position_ids tests again #6100 (@sshleifer)
Fix deebert tests #6102 (@sshleifer)
Use FutureWarning to deprecate #6111 (@sgugger)
Added capability to quantize a model while exporting through ONNX. #6089 (@mfuntowicz)
XLNet PLM Readme #6121 (@LysandreJik)
Fix TF CTRL model naming #6134 (@jplu)
Use google style to document properties #6130 (@sgugger)
Test TF Flaubert + Add {XLM, Flaubert}{TokenClassification, MultipleChoice} #5614
Rework TF trainer #6038 (@jplu)
Actually the extra_id are from 0-99 and not from 1-100 #5967 (@orena1)
add another e.g. to avoid confusion #6055 (@orena1)
Tf trainer cleanup #6143 (@sgugger)
Switch from return_tuple to return_dict #6138 (@sgugger)
Fix FlauBERT GPU test #6142 (@LysandreJik)
Enable ONNX/ONNXRuntime optimizations through converter script #6131 (@mfuntowicz)
Add Pytorch Native AMP support in Trainer #6151 (@prajjwal1)
enable easy checkout switch #5645 (@stas00)
Replace mecab-python3 with fugashi for Japanese tokenization #6086 (@polm)
parse arguments from dict #4869 (@patil-suraj)
Harmonize both Trainers API #6157 (@sgugger)
Model output test #6155 (@sgugger)
[s2s] clean up + doc #6184 (@stas00)
Add script to convert BERT tf2.x checkpoint to PyTorch #5791 (@mar-muel)
Empty assert hunt #6056 (@TevenLeScao)
Fix saved model creation #5468 (@jplu)
Adds train_batch_size, eval_batch_size, and n_gpu to to_sanitized_dict output for logging. #5331 (@jaymody)
[DataCollatorForLanguageModeling] fix labels #6213 (@patil-suraj)
Fix _shift_right function in TFT5PreTrainedModel #6214 (@maurice-g)
Remove outdated BERT tips #6217 (@JetRunner)
run_hans label fix #6221 (@VictorSanh)
Make the order of additional special tokens deterministic #5704 (@gonglinyuan)
cleanup torch unittests #6196 (@stas00)
test_tokenization_common.py: Remove redundant coverage #6224 (@sshleifer)
[Reformer] fix reformer fp16 test #6237 (@patrickvonplaten)
[Reformer] Make random seed generator available on random seed and not on model device #6244 (@patrickvonplaten)
Update to match renamed attributes in fairseq master #5972 (@LilianBordeau)
[WIP] lightning_base: support --lr_scheduler with multiple possibilities #6232 (@stas00)
Trainer + wandb quality of life logging tweaks #6241 (@TevenLeScao)
Add strip_accents to basic BertTokenizer. #6280 (@PhilipMay)
Argument to set GPT2 inner dimension #6296 (@TevenLeScao)
[Reformer] fix default generators for pytorch < 1.6 #6300 (@patrickvonplaten)
Remove redundant line in run_pl_glue.py #6305 (@xujiaze13)
[Fix] text-classification PL example #6027 (@bhashithe)
fix the shuffle agrument usage and the default #6307 (@stas00)
CI dependency wheel caching #6287 (@LysandreJik)
Patch GPU failures #6281 (@LysandreJik)
fix consistency CrossEntropyLoss in modeling_bart #6265 (@idoh)
Add a script to check all models are tested and documented #6298 (@sgugger)
Fix the tests for Electra #6284 (@jplu)
[examples] consistently use --gpus, instead of --n_gpu #6315 (@stas00)
refactor almost identical tests #6339 (@stas00)
Small docfile fixes #6328 (@sgugger)
Patch models #6326 (@LysandreJik)
Ci GitHub caching #6382 (@LysandreJik)
Fix links for open in colab #6391 (@sgugger)
[EncoderDecoderModel] add a add_cross_attention boolean to config #6377 (@patrickvonplaten)
Feed forward chunking #6024 (@Pradhy729)
add pl_glue example test #6034 (@stas00)
testing utils: capturing std streams context manager #6231 (@stas00)
Fix tokenizer saving and loading error #6026 (@yobekiko)
Warn if debug requested without TPU #6390 (@dmlap)
[Performance improvement] "Bad tokens ids" optimization #6064 (@guillaume-be)
pl version: examples/requirements.txt is single source of truth #6309 (@stas00)
[s2s] wmt download script use less ram #6405 (@stas00)
[pl] restore lr logging behavior for glue, ner examples #6314 (@stas00)
lr_schedulers: add get_polynomial_decay_schedule_with_warmup #6361 (@stas00)
[examples] add pytest dependency #6425 (@sshleifer)
[test] replace capsys with the more refined CaptureStderr/CaptureStdout #6422 (@stas00)
Fixes to make life easier with the nlp library #6423 (@sgugger)
Move prediction_loss_only to TrainingArguments #6426 (@sgugger)
Activate check on the CI #6427 (@sgugger)
cleanup tf unittests: part 2 #6260 (@stas00)
Fix docs and bad word tokens generation_utils.py #6387 (@ZhuBaohe)
Test model outputs equivalence #6445 (@LysandreJik)
add LongformerTokenizerFast in AutoTokenizer #6463 (@patil-suraj)
add BartTokenizerFast in AutoTokenizer #6464 (@patil-suraj)
Add POS tagging and Phrase chunking token classification examples #6457 (@vblagoje)
Clean directory after script testing #6453 (@JetRunner)
Use hash to clean the test dirs #6475 (@JetRunner)
Sort unique_no_split_tokens to make it deterministic #6461 (@lhoestq)
Fix TPU Convergence bug #6488 (@jysohn23)
Support additional dictionaries for BERT Japanese tokenizers #6515 (@singletongue)
[doc] Summary of the models fixes #6511 (@stas00)
Remove deprecated assertEquals #6532 (@JetRunner)
[testing] a new TestCasePlus subclass + get_auto_remove_tmp_dir() #6494 (@stas00)
[sched] polynomial_decay_schedule use default power=1.0 #6473 (@stas00)
Fix flaky ONNX tests #6531 (@mfuntowicz)
[doc] make the text more readable, fix some typos, add some disambiguation #6508 (@stas00)
[doc] multiple corrections to "Summary of the tasks" #6509 (@stas00)
replace _ with __ rst links #6541 (@stas00)
Fixed label datatype for STS-B #6492 (@amodaresi)
fix incorrect codecov reports #6553 (@stas00)
[docs] Fix wrong newline in the middle of a paragraph #6573 (@romainr)
[docs] Fix number of 'ug' occurrences in tokenizer_summary #6574 (@romainr)
add BartConfig.force_bos_token_to_be_generated #6526 (@sshleifer)
Fix bart base test #6587 (@sshleifer)
Feed forward chunking others #6365 (@Pradhy729)
tf generation utils: remove unused kwargs #6591 (@sshleifer)
[BartTokenizerFast] add prepare_seq2seq_batch #6543 (@patil-suraj)
[doc] lighter 'make test' #6512 (@stas00)
[docs] Copy code button misses '...' prefixed code #6518 (@romainr)
removed redundant arg in prepare_inputs #6614 (@prajjwal1)
add intro to nlp lib & dataset links to custom datasets tutorial #6583 (@joeddav)
Add tests to Trainer #6605 (@sgugger)
TFTrainer dataset doc & fix evaluation bug #6618 (@joeddav)
Add tests/test_tokenization_reformer.py #6485 (@D-Roberts)
[Tests] fix attention masks in Tests #6621 (@patrickvonplaten)
XLNet Bug when training with apex 16-bit precision #6567 (@johndolgov)
Move threshold up for flaky test with Electra #6622 (@sgugger)
Regression test for pegasus bugfix #6606 (@sshleifer)
Trainer automatically drops unused columns in nlp datasets #6449 (@sgugger)
[Docs model summaries] Add pegasus to docs #6640 (@patrickvonplaten)
[Doc model summary] add MBart model summary #6649 (@patil-suraj)
Specify config filename in HfArgumentParser #6626 (@jarednielsen)
Don't reset the dataset type + plug for rm unused columns #6683 (@sgugger)
Fixed DataCollatorForLanguageModeling not accepting lists of lists #6685 (@TevenLeScao)
Update repo to isort v5 #6686 (@sgugger)
Fix PL token classification examples #6682 (@vblagoje)
Lat fix for Ray HP search #6691 (@sgugger)
Create PULL_REQUEST_TEMPLATE.md #6660 (@stas00)
[doc] remove BartForConditionalGeneration.generate #6659 (@stas00)
Move unused args to kwargs #6694 (@sgugger)
[fixdoc] Add import to pegasus usage doc #6698 (@sshleifer)
Fix hyperparameter_search doc #6695 (@sgugger)
Remove hard-coded uses of float32 to fix mixed precision use #6648 (@schmidek)
Add DPR to models summary #6690 (@lhoestq)
Add typing.overload for convert_ids_tokens #6637 (@tamuhey)
Allow tests in examples to use cuda or fp16,if they are available #5512 (@Joel-hanson)
ci/gh/self-scheduled: add newline to make examples tests run even if src/ tests fail #6706 (@sshleifer)
Use separate tqdm progressbars #6696 (@sgugger)
More tests to Trainer #6699 (@sgugger)
Add tokenizer to Trainer #6689 (@sgugger)
tensor.nonzero() is deprecated in PyTorch 1.6 #6715 (@mfuntowicz)
[Albert] Add position ids to allowed uninitialized weights #6719 (@patrickvonplaten)
Fix ONNX test_quantize unittest #6716 (@mfuntowicz)
[squad] make examples and dataset accessible from SquadDataset object #6710 (@lazovich)
Fix pegasus-xsum integration test #6726 (@sshleifer)
T5Tokenizer adds EOS token if not already added #5866 (@sshleifer)
Install nlp for github actions test #6728 (@sgugger)
[Torchscript] Fix docs #6740 (@patrickvonplaten)
Add "tie_word_embeddings" config param #6692 (@patrickvonplaten)
Fix tf boolean mask in graph mode #6741 (@JayYip)
Fix TF optimizer #6717 (@jplu)
[TF Longformer] Improve Speed for TF Longformer #6447 (@patrickvonplaten)
add init.py to utils #6754 (@joeddav)
[s2s] run_eval.py QOL improvements and cleanup #6746 (@sshleifer)
s2s distillation uses AutoModelForSeqToSeqLM #6761 (@sshleifer)
Add AdaFactor optimizer from fairseq #6722 (@moscow25)
Adds Adafactor to the docs and slightly fixes the formatting #6765 (@LysandreJik)
Fix the TF Trainer gradient accumulation and the TF NER example #6713 (@jplu)
Fix run_squad.py to work with BART #6756 (@tomgrek)
Add NLP install to self-scheduled CI #6767 (@sshleifer)
[testing] replace hardcoded paths to allow running tests from anywhere #6523 (@stas00)
[test schedulers] adjust to test the first step's reading #6429 (@stas00)
new Makefile target: docs #6510 (@stas00)
[transformers-cli] fix logger getter #6777 (@stas00)
PL: --adafactor option #6776 (@sshleifer)
[style] set the minimal required version for black #6784 (@stas00)
Transformer-XL: Improved tokenization with sacremoses #6322 (@RafaelWO)
prepare_seq2seq_batch makes labels/ decoder_input_ids made later. #6654 (@sshleifer)
t5 model should make decoder_attention_mask #6800 (@sshleifer)
[s2s] Test hub configs in self-scheduled CI #6809 (@sshleifer)
[bart] rename self-attention -> attention #6708 (@sshleifer)
[tests] fix typos in inputs #6818 (@stas00)
Fixed open in colab link #6825 (@PandaWhoCodes)
clarify shuffle #6312 (@xujiaze13)
TF Flaubert w/ pre-norm #6841 (@LysandreJik)
Fix resuming training for Windows #6847 (@sgugger)
Only access loss tensor every logging_steps #6802 (@jysohn23)
Add checkpointing to Ray Tune HPO #6747 (@krfricke)
Split hp search methods #6857 (@sgugger)
Fix marian slow test #6854 (@sshleifer)
Bart can make decoder_input_ids from labels #6758 (@sshleifer)
add a final report to all pytest jobs #6861 (@stas00)
Restore PaddingStrategy.MAX_LENGTH on QAPipeline while no v2. #6875 (@mfuntowicz)
[Generate] Facilitate PyTorch generate using ModelOutputs #6735 (@patrickvonplaten)
Fixes bugs introduced by v3.0.0 and v3.0.1 in tokenizers.
Fixes bugs introduced by v3.0.0 and v3.0.1 in tokenizers.
…truncation and padding API but still led to two breaking changes that could have been avoided.
Version v3.0.0, included a refactoring of the tokenizers' backend to allow a simpler and more flexible user-facing API.
This refactoring was conducted with a particular focus on keeping backward compatibility for the v2.X encoding, truncation and padding API but still led to two breaking changes that could have been avoided.
This patch aims to bring back better backward compatibility, by implementing the following updates:
prepare_for_model method is now publicly exposed again for both slow and fast tokenizers with an API compatible with both the v2.X truncation/padding API and the v3.0 recommended API.longest_first instead of first_only.Deprecate any argument that's not labels (like masked_lm_labels, lm_labels, etc.) to labels.
v2BertForMaskedLM and BertLMHeadModel. BertForMaskedLM therefore cannot do causal language modeling anymore, and cannot accept the lm_labels argument.Trainer data collator is now a method instead of a classtokenizer.mask_token = '<mask>' now only associate the token to the attribute of the tokenizer but doesn't add the token to the vocabulary if it is not in the vocabulary. Tokens are only added by using the tokenizer.add_special_tokens() and tokenizer.add_tokens() methodsprepare_for_model method was removed as part of the new tokenizer API.only_first by default.The tokenizers has evolved quickly in version 2, with the addition of rust tokenizers. It now has a simpler and more flexible API aligned between Python (slow) and Rust (fast) tokenizers. This new API let you control truncation and padding deeper allowing things like dynamic padding or padding to a multiple of 8.
The redesigned API is explained in detail here #4510 and here: https://huggingface.co/transformers/master/preprocessing.html
Notable changes:
tokenizer.__call__ can be used for all case (single sequence, pair of sequences to groups, batches, etc...)AddedToken can be used to have a more fine-grained control on how added tokens behave during tokenization. In particular the user can control (1) whether left and right spaces are removed around the token during tokenization (2) whether the token will be identified inside another word and (3) whether the token will be recognized in normalized forms (e.g. in lower case if the tokenizer uses lower-casing)return_tensors parameter on tokenizers.TensorType to map all the possible tensor backends we support: TensorType.TENSORFLOW, TensorType.PYTORCH, TensorType.NUMPYTensorType enum on encode(...), encode_plus(...), batch_encode_plus(...) tokenizer method for return_tensors parameters.BatchEncoding new property is_fast indicates if the BatchEncoding comes from a Python (slow) tokenizer or a Rust (fast) tokenizer.BatchEncoding.Several PRs to make the API more stable have been made:
pad_to_multiple_of on tokenizers (reimport) #5054 (@mfuntowicz)Very big release for TensorFlow!
TFPretrainedModel.compute_loss method. #4530We welcome @sgugger as a team member in New York. He already introduced a lot of very cool documentation changes:
The MobileBERT from MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices by Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, Denny Zhou, was added to the library for both PyTorch and TensorFlow.
A single checkpoint is added: mobilebert-uncased which is the uncased_L-24_H-128_B-512_A-4_F-4_OPT checkpoint converted to our API.
This model was first implemented in PyTorch by @lonePatient, ported to the library by @vshampor, then finalized and implemented in Tensorflow by @LysandreJik.
The examples/eli5 folder contains training code for the dense retriever and to fine-tune a BART model, the jupyter notebook for the blog post, and the code for the live demo.
The RetriBert model implements the dense passage retriever. It's basically a wrapper for two Bert models and projection matrices, but it does gradient checkpointing in a way that is very different from a concurrent PR and Yacine thought it would be easier to write its own class for now and see if we can merge into the BART code later.
examples/seq2seq folder is a combination of the old examples/summarization and examples/translation folders.--freeze_encoder and --freeze_embeds options. These options make finetuning BART 5x faster on the cnn/dailymail dataset.bart-large-cnn and bart-large-xsum. They can be loaded using BartForConditionalGeneration.from_pretrained('sshleifer/distilbart-xsum-12-6'), for example See this tweet for more info on available models and their speed/performance.examples/seq2seq folderAdd BERT Loses Patience (Patience-based Early Exit) based on the paper https://arxiv.org/abs/2006.04152 and the official implementation https://github.com/JetRunner/PABEE
label arguments (@sgugger) #4722labels (like masked_lm_labels, lm_labels, etc.) to labels.Introduce a new tensor type for return_tensors on tokenizer for NumPy.
As we're introducing more than two tensor backend alternatives I created an enum TensorType listing all the possible tensor we can create TensorType.TENSORFLOW, TensorType.PYTORCH, TensorType.NUMPY. This might help newcomers who don't know about "tf", "pt". Note: TensorType are compatible with previous "tf", "pt" and now "np" str to allow backward compatibility (+unittest)
Numpy is now a possible target when creating tensors. This is usefull for JAX.
The benchmark script was consolidated and some features were added:
Adds the functionality to measure the following functionalities for TF and PT (#4912):
Tensorflow:
PyTorch:
[Benchmark] Add encoder decoder to benchmark and clean labels #4810
[Benchmark] add tpu and torchscipt for benchmark #4850
[Benchmark] Extend Benchmark to all model type extensions #5241
[Benchmarks] improve Example Plotter #5245
Before v3.0.0, the way to handle attentions, model hidden states, and whether to use the cache in models that have it for sequential decoding was to specify an argument in the configuration. In version v3.0.0, while we do maintain that argument for backwards compatibility, we introduce a new way of handling these through the forward and call methods.
AutoModels (@patrickvonplaten)The AutoModelWithLMHead encompasses all models with a language modeling head, not making the distinction between causal, masked and seq2seq models. Three new auto models are added:
AutoModelForCausalLM for Autoregressive modelsAutoModelForMaskedLM for Autoencoding modelsAutoModelForSeq2SeqCausalLM for Sequence-to-sequence models with causal LM for the decoderd_head is already in the configuration #4747 (@LysandreJik)sphinx-rtd-theme #5128 (@LysandreJik)The max_len attribute is now more robust, and warns the user about deprecation (@mfuntowicz, #4528)
Archive maps were dictionaries linking pre-trained models to their S3 URLs. Since the arrival of the model hub, these have become obsolete.
⚠️ This PR is breaking for the following models: BART, Flaubert, bert-japanese, bert-base-finnish, bert-base-dutch. ⚠️ Those models now have to be instantiated with their full model id:
"cl-tohoku/bert-base-japanese" "cl-tohoku/bert-base-japanese-whole-word-masking" "cl-tohoku/bert-base-japanese-char" "cl-tohoku/bert-base-japanese-char-whole-word-masking" "TurkuNLP/bert-base-finnish-cased-v1" "TurkuNLP/bert-base-finnish-uncased-v1" "wietsedv/bert-base-dutch-cased" "flaubert/flaubert_small_cased" "flaubert/flaubert_base_uncased" "flaubert/flaubert_base_cased" "flaubert/flaubert_large_cased"
all variants of "facebook/bart"
Update: ⚠️ This PR is also breaking for ALBERT from Tensorflow. See issue #4806 for discussion and resolution ⚠️
max_len attribute is now more robust, and warns the user about deprecation (@mfuntowicz, #4528)modeling_utils.py (@bglearning, #3911)nn.Module as a superclass (@shoarora, #4533)transformers-cli is now cross-platform (@BramVanroy, #4131) + (@patrickvonplaten, #4614)input_ids and past of variable length (@patrickvonplaten, #4581)--do_lower_case to SQuAD examples.add_special_tokens on fast tokenizers (@n1t0, #4531)num_labels (@julien-c, direct commit to master)Removed warning of deprecation (@Colanim)
Added a new model "Reformer": https://arxiv.org/abs/2001.04451 to the library. Original trax code: https://github.com/google/trax/tree/master/trax/models/reformer was translated to PyTorch.
Reformer uses chunked attention and reversible layers to model sequences as long as 500,000 tokens.
Reformer is currently available as a casual language model and will soon also be available as encoder only ("Bert"-like) model.
Two pretrained weights are uploaded: https://huggingface.co/models?search=google%2Freformer
https://huggingface.co/google/reformer-enwik8 is the first char lm in the library
ElectraForSequenceClassification was added by @liuzzinn.DataParallel support compatibility for PyTorch v1.5.0tokenizer.decode_batch, to decode an entire batch (@sshleifer)We've started adding community notebooks to the repository. Three notebooks have made their way into our codebase:
-Adds predict stage for glue tasks, and generate result files which can be submitted to gluebenchmark.com (@stdcoutzyx)
p_mask in SQuAD pre-processing (@LysandreJik)torch==1.4.0 (@mfuntowicz)None values in GradientAccumulator (@jarednielsen, improved by @jplu)run_language_modeling fix: actually use the overwrite_cache argument (@borisdayma)A new model architecture, MarianMTModel with 1,008+ pretrained weights is available for machine translation in PyTorch.
MarianMTModel with 1,008+ pretrained weights is available for machine translation in PyTorch.MarianTokenizer uses a prepare_translation_batch method to prepare model inputs.Helsinki-NLP/opus-mt-{src}-{tgt}A new model architecture has been added: AlbertForPreTraining in both PyTorch and TensorFlow
Changes have been made to both the TensorFlow scripts and our internals so that we are compatible with TensorFlow 2.2
This introduces a breaking change, in that it increases the default output length of T5Model and T5ForConditionalGeneration from 4 to 5 (including the…
Version 2.9 introduces a new Trainer class for PyTorch, and its equivalent TFTrainer for TF 2.
This let us reorganize the example scripts completely for a cleaner codebase.
The main features of the Trainer are:
The TFTrainer was largely contributed by awesome community member @jplu! 🔥 🔥
A few additional features of the example scripts are:
Documentation for the Trainer is still work-in-progress, please consider contributing improvements.
torch.distributed.New BART checkpoint converted: this adds mbart-en-ro model, a BART variant finetuned on english-romanian translation.
huggingface/tokenizershuggingface/tokenizers tokenizers. (@mfuntowicz, @thomwolf)Auto-regressive decoding for T5 has been greatly sped up by storing past key/value states. Work done on both PyTorch and TensorFlow.
This introduces a breaking change, in that it increases the default output length of T5Model and T5ForConditionalGeneration from 4 to 5 (including the past_key_value_states).
Question Answering support for Albert and Roberta in TF with (@Pierrci):
Implements a text generation pipeline, GenerationPipeline, which works on any ModelWithLMHead head.
output_past everywhere and replace by use_cache argument (@patrickvonplaten)PreTrainedModel (@sshleifer)PretrainedTokenizer (@sshleifer)qas_id to SquadResult and SquadExample (@jarednielsen)num_labels in configuration objectsELECTRA is a new method for self-supervised language representation learning. It can be used to pre-train transformer networks using relatively little
ELECTRA is a new method for self-supervised language representation learning. It can be used to pre-train transformer networks using relatively little compute. ELECTRA models are trained to distinguish "real" input tokens vs "fake" input tokens generated by another neural network, similar to the discriminator of a GAN. At small scale, ELECTRA achieves strong results even when trained on a single GPU. At large scale, ELECTRA achieves state-of-the-art results on the SQuAD 2.0 dataset.
This release comes with 6 ELECTRA checkpoints:
google/electra-small-discriminatorgoogle/electra-small-generatorgoogle/electra-base-discriminatorgoogle/electra-base-generatorgoogle/electra-large-discriminatorgoogle/electra-large-generatorRelated:
Thanks to the author @clarkkev for his help during the implementation.
Thanks to community members @hfl-rc @stefan-it @shoarora for already sharing more fine-tuned Electra variants!
generate (@patrickvonplaten)The generate method now has a bad word filter.
T5 is a powerful encoder-decoder model that formats every NLP problem into a text-to-text format. It achieves state of the art results on a variety of
T5 is a powerful encoder-decoder model that formats every NLP problem into a text-to-text format. It achieves state of the art results on a variety of NLP tasks (Summarization, Question-Answering, ...).
Five sets of pre-trained weights (pre-trained on a multi-task mixture of unsupervised and supervised tasks) are released. In ascending order from 60 million parameters to 11 billion parameters:
t5-small, t5-base, t5-large, t5-3b, t5-11b
T5 can now be used with the translation and summarization pipeline.
Related:
Big thanks to the original authors, especially @craffel who helped answer our questions, reviewed PRs and tested T5 extensively.
bart-large-xsum (@sshleifer)These weights are from BART finetuned on the XSum abstractive summarization challenge, which encourages shorter (more abstractive) summaries. It achieves state of the art.
New example: BART for summarization, using Pytorch-lightning. Trains on CNN/DM and evaluates.
A new pipeline is available, leveraging the T5 model. The T5 model was added to the summarization pipeline as well.
In an effort to have the same memory footprint and same computing power necessary to run inference on BART, several improvements have been made on the model:
evaluate_cnn example.Supports JSON serialization of Keras layers by overriding get_config, so that they can be sent to Tensorboard to display a conceptual graph of the model. TensorFlow models may now be saved using model.save, as other Keras models.
A new head was added to XLM: XLMForTokenClassification.
Bart is one of the first Seq2Seq models in the library, and achieves state of the art results on text generation tasks, like abstractive summarization
Bart is one of the first Seq2Seq models in the library, and achieves state of the art results on text generation tasks, like abstractive summarization. Three sets of pretrained weights are released:
bart-large: the pretrained base modelbart-large-cnn: the base model finetuned on the CNN/Daily Mail Abstractive Summarization Taskbart-large-mnli: the base model finetuned on the MNLI classification task.Related:
Big thanks to the original authors, especially Mike Lewis, Yinhan Liu, Naman Goyal who helped answer our questions.
The huggingface API for model upload now supports organisations.
A few beginner-oriented notebooks were added to the library, aiming at demystifying the two libraries huggingface/transformers and huggingface/tokenizers. Contributors are welcome to contribute links to their notebooks as well.
Examples leveraging pytorch-lightning were added, led by @srush. The first example that was added is the NER example. The second example is a lightning GLUE example, added by @nateraw.
CamembertForQuestionAnswering was added to the library and to the SQuAD script @maximeilluinAlbertForTokenClassification was added to the library and to the NER example @marmaMost of these fixes were done in the patch 2.5.1. Fast tokenizers should now have the exact same API as the python ones, with some additional functionalities.
Docker images for transformers were added.
no_repeat_ngram_size kwarg to avoid redundant generations (@sshleifer)Models such as DistilBERT and RoBERTa do not make use of token type IDs. These inputs are not returned by the encoding methods anymore, except if explicitly mentioned during the tokenizer initialization.
bart-large-cnn, with the generation parameters published in the paper.Previously all attempts to load a model from a pre-trained checkpoint would check that the S3 etag corresponds to the one hosted locally. This has been updated so that an argument local_files_only prevents this, which can be useful when a firewall is involved.
In a continuing effort to onboard new users (new to the lib or new to NLP in general), some usage examples were added to the documentation. These usage examples showcase how to do inference on several tasks:
CI now runs on GPU. PyTorch and TensorFlow.
Older tokenizers could pad even when no padding token was defined, which has been updated in this version to match the expected behavior, which is the FastTokenizers' behavior: add a pad token or raise an error when trying to batch without one.
We're now dropping Python 3.5 support.
add_special_tokens with the fast tokenizer methods of encoding (@LysandreJik)encode_plus was modified and tested to have the exact same behaviour as encode, but batches inputF.gelu for torch >= 1.4.0 (@sshleifer)get_vocab method to tokenizers, which can be used to retrieve all the vocabulary from the tokenizers. (@joeddav)special_tokens_mask when add_special_tokens=False (@LysandreJik)Model2LSTM and Model2Model which was not workingnum_labels initialization (@LysandreJik)run_ner.py script @eripAutoTokenizer has been put back to False by default so as to not have a breaking change between 2.4.x and 2.5.x
AutoTokenizer has been put back to False by default so as to not have a breaking change between 2.4.x and 2.5.x
Bug fixes
Bug fixes related to batch_encode_plus
Tokenizers for Bert, Roberta, OpenAI GPT, OpenAI GPT2, TransformerXL are now leveraging tokenizers library for fast tokenization :rocket:
Known Issues:
The distilled version of the bert-base-cased BERT checkpoint has been released.
Model cards are now stored directly in the repository
We now host a CLI script that gathers all the environment information when reporting an issue. The issue templates have been updated accordingly.
The main contributors as identified by Sourcerer are now visible directly on the repository.
The language fine-tuning script has been renamed from run_lm_finetuning to run_language_modeling as it is now also able to train language models from scratch.
cached_path (@thomwolf )Slight modification to cached_path so that zip and tar archives can be automatically extracted.
extract_compressed_file=True when calling cached_file.force_extract=True, in which case the cached extraction directory is removed and the archive is extracted again.Several activation functions (relu, swish, gelu, tanh and gelu_new) can now be accessed from the activations.py file and be used in the different PyTorch models.
test_attention_weights (@sshleifer)run_language_modeling (@LysandreJik )TokenClassificationPipeline, which is an alias over NerPipeline (@julien-c )Patched an issue where FlauBERT couldn't be loaded with AutoModel and AutoTokenizer classes.
Patched an issue where FlauBERT couldn't be loaded with AutoModel and AutoTokenizer classes.
MMBT was added to the list of available models, as the first multi-modal model to make it in the library. It can accept a transformer model as well as
wietsedv/bert-base-dutch-cased identifier. Added by @wietsedv. Model pageblack, isort and flake8. A test was added, check_code_quality, which checks that the contributions respect the contribution guidelines related to those tools.setup.py for the necessary dev dependencies.make style
make quality
The documentation was uniformized and some better guidelines have been defined. This work is part of an ongoing effort of making transformers accessible to a larger audience. A glossary has been added, adding definitions for most frequently used inputs.
Furthermore, some tips are given concerning each model in their documentation pages.
The code samples are now tested on a weekly basis alongside other slow tests.
The source code was moved from ./transformers to ./src/transformers. Since it changes the location of the source code, contributors must update their local development environment by uninstalling and re-installing the library.
Version 2.3.0 was the last version to support Python 2. As we begin the year 2020, official Python 2 support has been dropped.
Tests can now be run in parallel
An abstract method was added to PreTrainedModel, which is implemented in all models trained with CLM. This abstract method is generate, which offers an API for text generation:
Previously, when stopping a training the only saved values would be the model weights/configuration. Now the different scripts save several other values: the global step, current epoch, and the steps trained in the current epoch. When resuming a training, all those values will be leveraged to correctly resume the training.
This applies to the following scripts: run_glue, run_squad, run_ner, run_xnli.
The run_lm_finetuning.py script now handles training from scratch.
The configuration files now contain the architecture they're referring to. There is no need to have the architecture in the file name as it was necessary before. This should ease the naming of community models.
A new type of AutoModel was added: AutoModelForPreTraining. This model returns the base model that was used during the pre-training. For most models it is the base model alongside a language modeling head, whereas for others it is a different model, e.g. BertForPreTraining for BERT.
The HANS dataset was added to the examples. It allows for testing a model with adversarial evaluation of natural language.
When using PyTorch, certain values can be ignored when computing the loss. In order for the loss function to understand which indices must be ignored, those have to be set to a certain value. Most of our models required those indices to be set to -1. We decided to set this value to -100 instead as it is PyTorch's default value. This removes the discrepancy between user-implemented losses and the losses integrated in the models.
Further help from @r0mainK.
generate method which did not correctly handle the repetition penalty (@patrickvonplaten)repeating_words_penalty_for_language_generation (@patrickvonplaten)run_generation now leverages cached past input for models that have access to it (@patrickvonplaten)DistilHeisenBug or HeisenDistilBug (@LysandreJik, with the help of @julien-c and @thomwolf).run_tf_ner (@karajan1001).run_generation (@alberduris)run_gluerp_bucket (@mschrimpf)prepare_for_model when tensorizing and returning token type IDs (@LysandreJik).never_split parameters (@DeNeutoy)We have added a new class called Pipeline to simply run and use models for several down-stream NLP tasks.
Pipeline (beta): easily run and use models on down-stream NLP tasksWe have added a new class called Pipeline to simply run and use models for several down-stream NLP tasks.
A Pipeline is just a tokenizer + model wrapped so they can take human-readable inputs and output human-readable results.
The Pipeline will take care of :
tokenizing inputs strings => convert in tensors => run in the model => post-process output
Currently, we have added the following pipelines with a default model for each:
There are three ways to use pipelines:
from transformers import pipeline
# Test the default model for QA (Bert large finetuned on SQuAD 1.0)
nlp = pipeline('question-answering')
nlp(question= "Where does Amy live ?", context="Amy lives in Amsterdam.")
>>> {'answer': 'Amsterdam', 'score': 0.9657156007786263, 'start': 13, 'end': 21}
# Test a specific model for NER (XLM-R finetuned by @stefan-it on CoNLL03 English)
nlp = pipeline('ner', model='xlm-roberta-large-finetuned-conll03-english')
nlp("My name is Amy. I live in Paris.")
>>> [{'word': 'Amy', 'score': 0.9999586939811707, 'entity': 'I-PER'},
{'word': 'Paris', 'score': 0.9999983310699463, 'entity': 'I-LOC'}]
bash $ echo -e "Where does Amy live?\tAmy lives in Amsterdam" | transformers-cli run --task question-answering
{'score': 0.9657156007786263, 'start': 13, 'end': 22, 'answer': 'Amsterdam'}
transformers-cli serve --task question-answering
This new feature is currently in beta and will evolve in the coming weeks.
Users can now create accounts on the huggingface.co website and then login using the transformers CLI. Doing so allows users to upload their models to our S3 in their respective directories, so that other users may download said models and use them in their tasks.
Users may upload files or directories.
It's been tested by @stefan-it for a German BERT and by @singletongue for a Japanese BERT.
The run_squad script has been massively refactored. The reasons are the following:
It now leverages the full capacity of encode_plus, easing the addition of new models to the script. A new method squad_convert_examples_to_features encapsulates all of the tokenization.
This method can handle tensorflow_datasets as well as squad v1 json files and squad v2 json files.
A contribution by @rlouf building on the encoder-decoder mechanism to do abstractive summarization.
from_pretrained())@alexzubiaga added XLNetForTokenClassification and TFXLNetForTokenClassification
run_bertology.py was updated with correct imports and the ability to overwrite the cacherun_lm_finetuning.py @bkkagglefrom_pretrained() can now load from urls directly.run_squad.py @ethanjperezPatched error where the tokenizers would split the special tokens.
Patched error where the tokenizers would split the special tokens.
This patch fixes a bug related to the input shape in several models in TensorFlow.
This patch fixes a bug related to the input shape in several models in TensorFlow.
A tokenization message was too present and overloaded the output, hiding the relevant information. It was removed.
Four new models have been added in v2.2.0
Four new models have been added in v2.2.0
We welcome the possibility to create fully seq2seq models by incorporating Encoder-Decoder architectures using a PreTrainedEncoderDecoder class that can be initialized from pre-trained models. The base BERT class has be modified so that it may behave as a decoder.
Furthermore, a Model2Model class that simplifies the definition of an encoder-decoder when both encoder and decoder are based on the same model has been added. @rlouf
Works by @tlkh and @LysandreJik aiming to benchmark the library models with different technologies: with TensorFlow and Pytorch, with mixed precision (AMP and FP-16) and with model tracing (Torchscript and XLA). A new section was created in the documentation: benchmarks pointing to Google sheets with the results.
Tokenizers now add special tokens by default. @LysandreJik
Model templates to ease the addition of new models to the library have been added. @thomwolf
A new input has been added to all models' forward (for Pytorch) and call (for TensorFlow) methods. These inputs_embeds are a direct embedded representation. This is useful as it gives more control over how to convert input_ids indices into associated vectors than the model's internal embedding lookup matrix. @julien-c
A new API for the input and output embeddings are available. These methods are model-independent and allow easy acquisition/modification of the models' embeddings. @thomwolf
New model architectures are available, namely: DistilBertForTokenClassification, CamembertForTokenClassification @stefan-it
Two new models have been added since release 2.0.
Two new models have been added since release 2.0.
Several updates have been made to the distillation script, including the possibility to distill GPT-2 and to distill on the SQuAD task. By @VictorSanh.
The run_glue.py example script can now run on a Pytorch TPU.
Several example scripts have been improved and refactored to use the full potential of the new tokenizer functions:
run_multiple_choice.py has been refactored to include encode_plus by @julien-c and @erenuprun_lm_finetuning.py has been improved with the help of @dennymarcels, @jinoobaek-qz and @LysandreJikrun_glue.py has been improved with the help of @brian41005Enhancements have been made on the tokenizers. Two new methods have been added: get_special_tokens_mask and truncate_sequences .
The former returns a mask indicating which tokens are special tokens in a token list, and which are tokens from the initial sequences. The latter truncate sequences according to a strategy.
Both of those methods are called by the encode_plus method, which itself is called by the encode method. The encode_plus now returns a larger dictionary which holds information about the special tokens, as well as the overflowing tokens.
Thanks to @julien-c, @thomwolf, and @LysandreJik for these additions.
The two methods add_special_tokens_single_sequence and add_special_tokens_sequence_pair have been removed. They have been replaced by the single method build_inputs_with_special_tokens which has a more comprehensible name and manages both sequence singletons and pairs.
The boolean parameter truncate_first_sequence has been removed in tokenizers' encode and encode_plus methods, being replaced by a strategy in the form of a string: 'longest_first', 'only_second', 'only_first' or 'do_not_truncate' are accepted strategies.
When the encode or encode_plus methods are called with a specified max_length, the sequences will now always be truncated or throw an error if overflowing.
New contributing guidelines have been added, alongside library development requirements by @rlouf, the newest member of the HuggingFace team.
tensorflow_datasets. This work has been done by @agrinh and @philipp-eisen.Nothing published for this version
model(inputs_ids, attention_mask=attention_mask, token_type_ids=token_type_ids), this should not cause any breaking change.
Following the extension to TensorFlow 2.0, pytorch-transformers => transformers
Install with pip install transformers
Also, note that PyTorch is no longer in the requirements so don't forget to install TensorFlow 2.0 and/or PyTorch to be able to use (and load) the models.
All the PyTorch nn.Module classes now have their counterpart in TensorFlow 2.0 as tf.keras.Model classes. TensorFlow 2.0 classes have the same name as their PyTorch counterparts prefixed with TF.
The interoperability between TensorFlow and PyTorch is actually a lot deeper than what is usually meant when talking about libraries with multiple backends:
from_pretrained can load weights saved from models saved in one or the other framework),Training on TPU using free TPUs provided in the TensorFlow Research Cloud (TFRC) program is possible but requires to implement a custom training loop (not possible with keras.fit at the moment). We will add an example of such a custom training loop soon.
Tokenizers have been improved to provide extended encoding methods encoding_plus and additional arguments to encoding. Please refer to the doc for detailed usage of the new options.
To be able to better use Torchscript both on CPU and GPUs (see #1010, #1204 and #1195) the specific order of some models keywords inputs (attention_mask, token_type_ids...) has been changed.
If you used to call the models with keyword names for keyword arguments, e.g. model(inputs_ids, attention_mask=attention_mask, token_type_ids=token_type_ids), this should not cause any breaking change.
If you used to call the models with positional inputs for keyword arguments, e.g. model(inputs_ids, attention_mask, token_type_ids), you should double-check the exact order of input arguments.
PyTorch is no longer in the requirements so don't forget to install TensorFlow 2.0 and/or PyTorch to be able to use (and load) the models.
The method add_special_tokens_sentence_pair has been renamed to the more appropriate name add_special_tokens_sequence_pair.
The same holds true for the method add_special_tokens_single_sentence which has been changed to add_special_tokens_single_sequence.
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →