NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3437 most downloaded on PyPI
Optimum Library is an extension of the Hugging Face Transformers library, providing a framework to integrate third-party libraries from Hardware Partners and interface with their specific functionality.
Last release 2 months ago
04 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
84 releases · first in 2021
Fix GPT-QModel compat and deprecate AutoGPTQ by @Qubitium in #2426
Full Changelog: v2.2.0...v2.3.0
One column per quarter.
Remove deprecated INC and IPEX from documentation by @echarlaix in https://github.com/huggingface/optimum/pull/2396
hf_device_map after Transformers optimization by @ZX-ModelCloud in https://github.com/huggingface/optimum/pull/2401Full Changelog: https://github.com/huggingface/optimum/compare/v2.1.0...v2.2.0
hf_device_map after Transformers optimization by @ZX-ModelCloud in #2401Full Changelog: v2.1.0...v2.2.0
Fully deprecate AutoGPTQ for GPT-QModel by @ZX-ModelCloud in https://github.com/huggingface/optimum/pull/2385
Full Changelog: https://github.com/huggingface/optimum/compare/v2.0.0...v2.1.0
Full Changelog: v2.0.0...v2.1.0
v2.0.0 introduces several breaking changes
v2.0.0 introduces several breaking changes
ONNX export and ONNX Runtime inference related integrations were moved to Optimum ONNX https://github.com/huggingface/optimum/pull/2298
How to obtain the same behavior as v1.x
pip install "optimum-onnx[onnxruntime]"
or equivalently
pip install "optimum[onnxruntime]"
🚨 You shouldn't install optimum without an extra if you want to be able to export your model to ONNX, please follow the installation instructions from the documentation 🚨
ONNX Runtime Training officially deprecated, more information on this in optimum-onnx v0.0.1 release notes
TF Lite export officially deprecated https://github.com/huggingface/optimum/pull/2340
BetterTransformer officially deprecated https://github.com/huggingface/optimum/pull/2305
Full Changelog: https://github.com/huggingface/optimum/compare/v1.27.0...v2.0.0
v2.0.0 introduces several breaking changes
ONNX export and ONNX Runtime inference related integrations were moved to Optimum ONNX #2298
How to obtain the same behavior as v1.x
pip install "optimum-onnx[onnxruntime]"or equivalently
pip install "optimum[onnxruntime]"🚨 You shouldn't install optimum without an extra if you want to be able to export your model to ONNX, please follow the installation instructions from the documentation 🚨
ONNX Runtime Training officially deprecated, more information on this in optimum-onnx v0.0.1 release notes
TF Lite export officially deprecated #2340
BetterTransformer officially deprecated #2305
Full Changelog: v1.27.0...v2.0.0
Deprecated support for TFLite, BetterTransformer, and ONNXRuntime‑Training, these integrations will be fully removed in v2.
Full Changelog: https://github.com/huggingface/optimum/compare/v1.26.1...v1.27.0
Full Changelog: v1.26.1...v1.27.0
Add back from_transformers for base model by @echarlaix in https://github.com/huggingface/optimum/pull/2288
Add back from_transformers for base model by @echarlaix in https://github.com/huggingface/optimum/pull/2288
D-FINE support by @xenova in https://github.com/huggingface/optimum/pull/2249
Fix ORT pipelines by @echarlaix in https://github.com/huggingface/optimum/pull/2274
Full Changelog**: https://github.com/huggingface/optimum/compare/v1.25.2...v1.25.3
Full Changelog**: v1.25.2...v1.25.3
Upgrade optimum-intel in setup extras by @echarlaix in https://github.com/huggingface/optimum/pull/2271
Full Changelog: https://github.com/huggingface/optimum/compare/v1.25.1...v1.25.2
Full Changelog: v1.25.1...v1.25.2
Updated readme/pypi page by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/2268
Full Changelog: https://github.com/huggingface/optimum/compare/v1.25.0...v1.25.1
Full Changelog: v1.25.0...v1.25.1
Remove deprecated ORTModel class by @echarlaix in https://github.com/huggingface/optimum/pull/2187
from optimum.onnxruntime import ORTModelForCausalLM
model_id = "meta-llama/Llama-3.2-1B"
- model = ORTModelForCausalLM.from_pretrained(model_id, export=True)
+ model = ORTModelForCausalLM.from_pretrained(model_id)
A huge thank you to our first-time contributors:
<details>
Tensor.repeat_interleave) by @xenova in https://github.com/huggingface/optimum/pull/2162CLIPSdpaAttention had dropped since v4.48 by @hans00 in https://github.com/huggingface/optimum/pull/2245<details>
We’re excited to announce the release of Optimum v1.24.0. This update expands ONNX-based model capabilities and includes several improvements, bug fix
We’re excited to announce the release of Optimum v1.24.0. This update expands ONNX-based model capabilities and includes several improvements, bug fixes, and new contributions from the community.
ORTQuantizer now supports models with ONNX subfolders.ORTDiffusionPipeline enabling latest diffusion-based models.ModelPatcher, SDXL refiner export, and device checks for improved reliability.A huge thank you to our first-time contributors:
Your contributions make Optimum better! :tada:
For a detailed list of all changes, please check out the full changelog.
:rocket: Happy optimizing!
<details>
fix] Allow ORTQuantizer over models with subfolder ONNX files by @tomaarsen in https://github.com/huggingface/optimum/pull/2094DummyTimestepInputGenerator by @JingyaHuang in https://github.com/huggingface/optimum/pull/2107ModelPatcher returns empty outputs by @LoSealL in https://github.com/huggingface/optimum/pull/2109model_kwargs when exporting a SentenceTransformers model by @sjrl in https://github.com/huggingface/optimum/pull/2126PatchTST by @xenova in https://github.com/huggingface/optimum/pull/2101</details>
Add sentence-transformers and timm documentation example by @echarlaix in https://github.com/huggingface/optimum/pull/2072
Fix compatibility with diffusers < 0.25.0 #2063 @echarlaix
Full Changelog: https://github.com/huggingface/optimum/compare/v1.23.1...v1.23.2
Fix doc build by @regisss in https://github.com/huggingface/optimum/pull/2050
Adding ORTDiffusionPipeline to simplify diffusers model loading by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1960 and https://g
Adding ORTDiffusionPipeline to simplify diffusers model loading by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1960 and https://github.com/huggingface/optimum/pull/2021
model_id = "runwayml/stable-diffusion-v1-5"
- pipeline = ORTStableDiffusionPipeline.from_pretrained(model_id, revision="onnx")
+ pipeline = ORTDiffusionPipeline.from_pretrained(model_id, revision="onnx")
image = pipeline("sailing ship in storm by Leonardo da Vinci").images[0]
Transformers v4.45 support by @echarlaix in https://github.com/huggingface/optimum/pull/2023 and https://github.com/huggingface/optimum/pull/2045
Remove the restriction for the model's config to be in the model's subfolder by @echarlaix in https://github.com/huggingface/optimum/pull/2044
Full Changelog: https://github.com/huggingface/optimum/compare/v1.22.0...v1.23.0
Deprecate ORTModel class by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1939
inputs_names to input_names by @J4BEZ in https://github.com/huggingface/optimum/pull/2010evaluation_strategy by @muellerzr in https://github.com/huggingface/optimum/pull/1819deepspeed and is_torch_xla_available by @Rohan138 in https://github.com/huggingface/optimum/pull/2012Full Changelog: https://github.com/huggingface/optimum/compare/v1.21.4...v1.22.0
Update Habana extra in setup.py by @regisss in #1991
Full Changelog: https://github.com/huggingface/optimum/compare/v1.21.3...v1.21.4
Deprecate ORTModel class by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1939
Full Changelog: https://github.com/huggingface/optimum/compare/v1.21.2...v1.21.3
Remove inplace op in mistral patcher by @IlyasMoutawwakil in #1938
Full Changelog: https://github.com/huggingface/optimum/compare/v1.21.1...v1.21.2
Fix sentence transformers model patching by @echarlaix in https://github.com/huggingface/optimum/pull/1936
Full Changelog: https://github.com/huggingface/optimum/compare/v1.21.0...v1.21.1
Deprecated use_auth_token by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1837
use_auth_token by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1837Full Changelog: https://github.com/huggingface/optimum/compare/v1.20.0...v1.21.0
VITS ONNX export by @echarlaix in https://github.com/huggingface/optimum/pull/1607
Bump transformers version by @echarlaix in https://github.com/huggingface/optimum/pull/1824
Remove call to apt update before apt purge in the main doc build workflow by @regisss in https://github.com/huggingface/optimum/pull/1830
Update github workflows by @echarlaix in https://github.com/huggingface/optimum/pull/1829
Remove bad PPA in main doc build workflow by @regisss in https://github.com/huggingface/optimum/pull/1831
Fix TPU doc build by @regisss in https://github.com/huggingface/optimum/pull/1834
Fix sentence transformers models infer library by @echarlaix in https://github.com/huggingface/optimum/pull/1832
Fix random initialization of bias when using GPTQ quantization with models without bias by @B-201 in https://github.com/huggingface/optimum/pull/1827
Update the Transformers dependency in the Habana extra by @regisss in https://github.com/huggingface/optimum/pull/1851
Make stable diffusion unet and vae number of channels static by @eaidova in https://github.com/huggingface/optimum/pull/1840
Fix compatibility with transformers v4.41.0 for ONNX by @echarlaix in https://github.com/huggingface/optimum/pull/1860
Fix FX CI by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1866
Fix Utils CI by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1867
Fix BT CI by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1872
Fix ORTConfig loading by @mr-sarthakgupta in https://github.com/huggingface/optimum/pull/1879
Update ORT doc for ROCM 6.0 by @mht-sharma in https://github.com/huggingface/optimum/pull/1862
Fix ort config instantiation (from_pretrained) and saving (save_pretrained) by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1865
Fix ORT CI by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1875
Update optimum intel extra by @echarlaix in https://github.com/huggingface/optimum/pull/1882
Bump transformers version for neuron extras by @JingyaHuang in https://github.com/huggingface/optimum/pull/1881
Full Changelog: https://github.com/huggingface/optimum/compare/v1.19.0...v1.20.0
Update the Transformers dependency in the Habana extra #1851 @regisss
Full Changelog: https://github.com/huggingface/optimum/compare/v1.19.1...v1.19.2
Bump transformers version by @echarlaix in https://github.com/huggingface/optimum/pull/1824
apt update before apt purge in the main doc build workflow by @regisss in https://github.com/huggingface/optimum/pull/1830Full Changelog: https://github.com/huggingface/optimum/compare/v1.19.0...v1.19.1
Musicgen and MarkupLM models from Transformers can now be exported to ONNX through optimum-cli export onnx. Musicgen ONNX export is used to run the mo
Musicgen and MarkupLM models from Transformers can now be exported to ONNX through optimum-cli export onnx. Musicgen ONNX export is used to run the model locally in a browser through transformers.js.
Full Changelog: https://github.com/huggingface/optimum/compare/v1.18.0...v1.19.0
Improve the installation of optimum-neuron through optimum extras #1778
Full Changelog: https://github.com/huggingface/optimum/compare/v1.18.0...v1.18.1
Fix starcoder ORT integration by @fxmarty in #1722
v4.39.0 by @echarlaix in #1764Update Transformers dependency in Habana extra #1700
Full Changelog: https://github.com/huggingface/optimum/compare/v1.17.0...v1.17.1
A function is exposed to programmatically export any nn.Module (e.g. models coming from Transformers, but modified). This is useful in case you need t
nn.ModuleA function is exposed to programmatically export any nn.Module (e.g. models coming from Transformers, but modified). This is useful in case you need to do some modifications on models loaded from the Hub before exporting. Example:
from transformers import AutoModelForImageClassification
from optimum.exporters.onnx import onnx_export_from_model
model = AutoModelForImageClassification.from_pretrained("google/vit-base-patch16-224")
# Here one could do any modification on the model before the export.
onnx_export_from_model(model, output="vit_onnx")
The Optimum ONNX export CLI allows to disable dynamic shape for inputs/outputs:
optimum-cli export onnx --model timm/ese_vovnet39b.ra_in1k out_vov --no-dynamic-axes
This is useful if the exported model is to be consumed by a runtime that does not support dynamic shapes. The static shape can be specified e.g. with --batch_size 1 . See all the shape options in optimum-cli export onnx --help.
The Optimum ONNX export now supports BF16 export on CPU and GPU. Beware though that ONNX Runtime is most often not able to consume the models as some operation are not implemented in this data type, although the exported models comply with ONNX standard. This is useful if you are developing a runtime that consomes BF16 ONNX models.
Example:
optimum-cli export onnx --model bert-base-uncased --dtype bf16 bert_onnx
You can now export to ONNX table-transformer, bart for text-classification.
trust_remote_code to sentence transformers export by @xenova in https://github.com/huggingface/optimum/pull/1677Timm models can now be run through ONNX Runtime with the class ORTModelForImageClassification:
from urllib.request import urlopen
import timm
import torch
from PIL import Image
from optimum.onnxruntime import ORTModelForImageClassification
# Export the model to ONNX under the hood with export=True.
model = ORTModelForImageClassification.from_pretrained("timm/resnext101_64x4d.c1_in1k", export=True)
# Get model specific transforms (normalization, resize).
data_config = timm.data.resolve_data_config(pretrained_cfg=model.config.pretrained_cfg)
transforms = timm.data.create_transform(**data_config, is_training=False)
img = Image.open(
urlopen("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png")
)
output = model(transforms(img).unsqueeze(0)).logits
top5_probabilities, top5_class_indices = torch.topk(torch.softmax(output, dim=1) * 100, k=5)
input_labels input to SAM model export by @xenova in https://github.com/huggingface/optimum/pull/1638_export by @JingyaHuang in https://github.com/huggingface/optimum/pull/1652onnx_export by @fxmarty in https://github.com/huggingface/optimum/pull/1685Full Changelog: https://github.com/huggingface/optimum/compare/v1.16.0...v1.17.0
Fix ORT training compatibility for transformers v4.36.0 by @AdamLouly https://github.com/huggingface/optimum/pull/1586
Fix ORT training compatibility for transformers v4.36.0 by @AdamLouly https://github.com/huggingface/optimum/pull/1586
Fix ONNX export compatibility for transformers v4.37.0 by @echarlaix https://github.com/huggingface/optimum/pull/1641
The features from BetterTransformer for Llama, Falcon, Whisper and Bart have been upstreamed in Transformers. Please use transformers>=4.36 and torch>
The features from BetterTransformer for Llama, Falcon, Whisper and Bart have been upstreamed in Transformers. Please use transformers>=4.36 and torch>=2.1.1 to use by default PyTorch's scaled_dot_product_attention.
More details: https://github.com/huggingface/transformers/releases/tag/v4.36.0
Full Changelog: https://github.com/huggingface/optimum/compare/v1.16.0...v1.16.1
Notably, the ONNX exports aten::scaled_dot_product_attention in a standardized way for the compatible models.
Notably, the ONNX exports aten::scaled_dot_product_attention in a standardized way for the compatible models.
Work in progress.
modules_in_block_to_quantize arg for gptq by @SunMarc in https://github.com/huggingface/optimum/pull/1585Full Changelog: https://github.com/huggingface/optimum/compare/v1.15.0...v1.16.0
The Optimum ONNX Runtime integration is extended to officially support `ROCMExecutionProvider`. See more details in the documentation.
The Optimum ONNX Runtime integration is extended to officially support ROCMExecutionProvider. See more details in the documentation.
The Swin2sr, DPT, GLPN, ConvNextv2 are now supported in the ONNX export.
convnextv2 onnx export by @xenova in https://github.com/huggingface/optimum/pull/1560delete_doc_comment workflows by @regisss in https://github.com/huggingface/optimum/pull/1565Full Changelog: https://github.com/huggingface/optimum/compare/v1.14.0...v1.15.0
Update optimum-intel required version by @echarlaix in https://github.com/huggingface/optimum/pull/1521
Remove SharedDDP as it was deprecated from Transformers by @AdamLouly in https://github.com/huggingface/optimum/pull/1443
Enable LCMs (available in in diffusers since v0.22.0) ONNX export and ORT inference by @echarlaix in https://github.com/huggingface/optimum/pull/1469
from optimum.onnxruntime import ORTLatentConsistencyModelPipeline
pipe = ORTLatentConsistencyModelPipeline.from_pretrained("SimianLuo/LCM_Dreamshaper_v7", export=True)
prompt = "sailing ship in storm by Leonardo da Vinci"
images = pipe(prompt=prompt, num_inference_steps=4, guidance_scale=8.0).images
Also enable ONNX export using the CLI :
optimum-cli export onnx --model SimianLuo/LCM_Dreamshaper_v7 lcm_onnx/
with-past in the ONNX export by @fxmarty in https://github.com/huggingface/optimum/pull/1358Patch release for transformers==4.34.1 compatibility. We will do a release next week for transformers==4.35 compatibility and new features. Please bea
Patch release for transformers==4.34.1 compatibility. We will do a release next week for transformers==4.35 compatibility and new features. Please bear with us!
Fix provider availability check on ORT 1.16.0 release by @fxmarty in https://github.com/huggingface/optimum/pull/1403
Fix ONNX fp16 export that broke in 1.13.0.
Fix ONNX fp16 export that broke in 1.13.0.
Workaround for a bug in the PyTorch ONNX export that does not deduplicate the Embedding and LM head shared weight: https://github.com/pytorch/pytorch/
Workaround for a bug in the PyTorch ONNX export that does not deduplicate the Embedding and LM head shared weight: https://github.com/pytorch/pytorch/issues/108342. For small enough models, this results in up to 50% ONNX serialized model size decrease.
ONNX Runtime integration now supports Pix2Struct and MPT architectures. Donut now supports IO Binding. Encoder-Decoder models are now supported as well.
Additionally, the model SAM is now be default exported as a vision_encoder.onnx, and prompt_encoder_mask_decoder.onnx.
BetterTransformer] Add falcon to BetterTransformer by @younesbelkada in https://github.com/huggingface/optimum/pull/1343The function exllama_set_max_input_length from auto-gptq can now be used with Transformers GPTQ models.
Update version to 1.12.1.dev0 following release by @fxmarty in https://github.com/huggingface/optimum/pull/1312
Add GPTQ prefill benchmark by @fxmarty in https://github.com/huggingface/optimum/pull/1313
Precise ORTModel documentation by @fxmarty in https://github.com/huggingface/optimum/pull/1268
Improve BetterTransformer backward compatibility by @fxmarty in https://github.com/huggingface/optimum/pull/1314
Improve ORTModel documentation by @fxmarty in https://github.com/huggingface/optimum/pull/1245
Add bitsandbytes benchmark by @fxmarty in https://github.com/huggingface/optimum/pull/1320
fix typo in log message by @AAnirudh07 in https://github.com/huggingface/optimum/pull/1322
Support customize dtype for dummy generators by @JingyaHuang in https://github.com/huggingface/optimum/pull/1307
Fix opset custom onnx export by @mht-sharma in https://github.com/huggingface/optimum/pull/1331
Replace mpt to ernie custom export by @mht-sharma in https://github.com/huggingface/optimum/pull/1332
Fix BT benchmark script by @fxmarty in https://github.com/huggingface/optimum/pull/1344
Add name_or_path for donut generation by @fxmarty in https://github.com/huggingface/optimum/pull/1345
send both negative prompt embeds to ORT SDXL by @ssube in https://github.com/huggingface/optimum/pull/1339
add vae image processor by @echarlaix in https://github.com/huggingface/optimum/pull/1219
add negative prompt test by @echarlaix in https://github.com/huggingface/optimum/pull/1347
Add GPT BigCode to the BT documentation by @fxmarty in https://github.com/huggingface/optimum/pull/1356
Add BT dummy objects by @fxmarty in https://github.com/huggingface/optimum/pull/1355
Add text2text-generation-with-past test for encoder-decoder model by @mht-sharma in https://github.com/huggingface/optimum/pull/1338
Fix sentence transformer export by @mht-sharma in https://github.com/huggingface/optimum/pull/1366
Full Changelog: https://github.com/huggingface/optimum/compare/v1.12.0...v1.13.0
Part of AutoGPTQ library has been integrated in Optimum, with utilities to ease the integration in other Hugging Face libraries. Reference: https://hu
Part of AutoGPTQ library has been integrated in Optimum, with utilities to ease the integration in other Hugging Face libraries. Reference: https://huggingface.co/docs/optimum/llm_quantization/usage_guides/quantization
BetterTransformer now supports BLOOM and GPT-BigCode architectures.
Full Changelog: https://github.com/huggingface/optimum/compare/v1.11.2...v1.12.0
Remove the Transformers version constraint on optimum[habana].
Remove the Transformers version constraint on optimum[habana].
Full Changelog: https://github.com/huggingface/optimum/compare/v1.11.1...v1.11.2
Minor fix: documentation building for 1.11.
Minor fix: documentation building for 1.11.
Full Changelog: https://github.com/huggingface/optimum/compare/v1.11.0...v1.11.1
Add ONNX export and ONNX Runtime inference support for gpt bigcode.
Add ONNX export and ONNX Runtime inference support for gpt bigcode.
BetterTransformer now supports Llama 2 and bark.
Training and autocast are now supported for most architectures, please refer to the documentation for more details: https://huggingface.co/docs/optimum/main/en/bettertransformer/overview
Full Changelog: https://github.com/huggingface/optimum/compare/v1.10.0...v1.11.0
Fix OwlViT exporter by @regisss in https://github.com/huggingface/optimum/pull/1188
Fix OwlViT exporter by @regisss in https://github.com/huggingface/optimum/pull/1188
Fix SD loading when safetensors weights only by @echarlaix in https://github.com/huggingface/optimum/pull/1232
Fix optimum-intel version requirements by @echarlaix in https://github.com/huggingface/optimum/pull/1234
Full Changelog: https://github.com/huggingface/optimum/compare/v1.10.0...v1.10.1
Enable SD XL ONNX export and ONNX Runtime inference by @echarlaix in https://github.com/huggingface/optimum/pull/1168
Enable SD XL ONNX export and ONNX Runtime inference by @echarlaix in https://github.com/huggingface/optimum/pull/1168
optimum-cli export onnx --model stabilityai/stable-diffusion-xl-base-0.9 --task stable-diffusion-xl ./sd_xl_onnx
from optimum.onnxruntime import ORTStableDiffusionXLPipeline
model_id = "stabilityai/stable-diffusion-xl-base-0.9"
pipeline = ORTStableDiffusionXLPipeline.from_pretrained(model_id, export=True)
prompt = "sailing ship in storm by Leonardo da Vinci"
image = pipeline(prompt).images[0]
pipeline.save_pretrained("onnx-sd-xl-base-0.9")
Enable image-to-image and inpainting pipelines for ONNX Runtime inference by @echarlaix in https://github.com/huggingface/optimum/pull/1121
More examples in documentation
attention_mask=None for BetterTransformer in the inference batched case for gpt2 & gpt-neo by @fxmarty in https://github.com/huggingface/optimum/pull/1180Full Changelog: https://github.com/huggingface/optimum/compare/v1.9.0...v1.10.0
Fix stable diffusion ONNX export for diffusers>=v0.18.0 by @echarlaix in https://github.com/huggingface/optimum/pull/1173
diffusers>=v0.18.0 by @echarlaix in https://github.com/huggingface/optimum/pull/1173Full Changelog: https://github.com/huggingface/optimum/compare/v1.9.0...v1.9.1
Remove deprecated argument from tests and examples by @echarlaix in https://github.com/huggingface/optimum/pull/1072
Lower memory usage during the ONNX export. This is especially useful to export large models, or on cuda device. Until PyTorch 2.1 release, we recommend to use PyTorch nightly in case memory issues are encountered, as two major bugs were fixed on PyTorch side: https://github.com/pytorch/pytorch/pull/101134 https://github.com/pytorch/pytorch/pull/101148
The ONNX export now supports the sam, lilt, pix2struct, cvt and owlvit architectures.
The method main_export now supports two arguments model_kwargs and custom_onnx_configs that allow for a more custom export for advanced users. Reference.
IO Binding is useful not only to avoid RAM/device memory copies, but also simply between numpy tensors and OrtValue. Thus, for autoregressive tasks we enable IO Binding as a default on CPUExecutionProvider as well, which may bring >10% speedup for large context lengths.
OptimizationConfig by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1036attention_mask in ORTModelForxxx by @IlyasMoutawwakil in https://github.com/huggingface/optimum/pull/1045input_points data type by @michaelbenayoun in https://github.com/huggingface/optimum/pull/1048masked-im output name fix for transformers >= 4.29.0 by @michaelbenayoun in https://github.com/huggingface/optimum/pull/1049ORTQuantizer.quantize call for static quantization when no calibration range is provided by @fxmarty in https://github.com/huggingface/optimum/pull/1094Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.0...v1.9.0
Fix stable diffusion ONNX export following diffusers breaking change by @fxmarty in https://github.com/huggingface/optimum/pull/1116
transformers>=v4.30.0 by @echarlaix in https://github.com/huggingface/optimum/pull/1102Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.7...v1.8.8
Restrict transformers version by @echarlaix in https://github.com/huggingface/optimum/pull/1097
Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.6...v1.8.7
Fix CLI for exporting models to TFLite by @regisss #1059
Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.5...v1.8.6
Add transformers<4.29.0 in Habana extra by @regisss in #1047
transformers<4.29.0 in Habana extra by @regisss in #1047Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.4...v1.8.5
Set onnx requirement by @echarlaix @regisss in https://github.com/huggingface/optimum/pull/1037
Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.3...v1.8.4
Fix Stable Diffusion model ONNX export by @echarlaix in https://github.com/huggingface/optimum/pull/1020
optimum-neuron extra by @michaelbenayoun in https://github.com/huggingface/optimum/pull/1021Full Changelog: https://github.com/huggingface/optimum/compare/v1.8.2...v1.8.3
Various improvements in the PyTorch BetterTransformer integration.
Various improvements in the PyTorch BetterTransformer integration.
BetterTransformer support for ProphetNet by @hirotasoshu in https://github.com/huggingface/optimum/pull/923BT] Improve docs by @younesbelkada in https://github.com/huggingface/optimum/pull/944Instead of using two separate decoder_model.onnx and decoder_with_past_model.onnx models, a single decoder can be used for encoder-decoder models: decoder_model_merged.onnx. This allows to avoid duplicated weights in the two without/with past ONNX models.
By default, if available, the decoder_model_merged.onnx will be used in the ORTModel integration. This can be disabled with the option --no-post-process in the ONNX export CLI, and with use_merged=False in the ORTModel.from_pretrained method.
Example:
optimum-cli export onnx --model t5-small t5_onnx
will give:
└── t5_onnx
├── config.json
├── decoder_model_merged.onnx
├── decoder_model.onnx
├── decoder_with_past_model.onnx
├── encoder_model.onnx
├── generation_config.json
├── special_tokens_map.json
├── spiece.model
├── tokenizer_config.json
└── tokenizer.json
And decoder_model_merged.onnx is enough to be used for inference. We strongly recommend to inspect the subgraphs with netron to understand what are the inputs/outputs, in case the exported model is to be used with an other engine than ONNX Runtime in the Optimum integration.
The TasksManager replaces legacy tasks names by the canonical ones used on the Hub and in transformers metadata:
sequence-classification becomes text-classification,causal-lm becomes text-generation,seq2seq-lm becomes text2text-generation,speech2seq-lm and audio-ctc becomes automatic-speech-recognition,default becomes feature-extraction,masked-lm becomes fill-mask,vision2seq-lm becomes image-to-textThis should not break anything except if you rely on private methods and attributes from TasksManager.
optimun-cli onnxruntime quantize / optimize output argument is now required by @michaelbenayoun in https://github.com/huggingface/optimum/pull/927optimum-cli print the help of subcommands by @michaelbenayoun in https://github.com/huggingface/optimum/pull/940Full Changelog: https://github.com/huggingface/optimum/compare/v1.7.3...v1.8.2
Nothing published for this version
Nothing published for this version
This patch releases fixes a few bugs with PyTorch 2.0 release, and include a few new features as well.
This patch releases fixes a few bugs with PyTorch 2.0 release, and include a few new features as well.
We removed some constant past key values outputs from encoder-decoder models in the ONNX export. Beware that this could potentially break your existing code, but we recommend to use the new exported models as this removes unnecessary Identity nodes in the models.
torch.nn.functional.scaled_dot_product_attention support for decoders in BetterTransformerPytorch 2.0 introduces in beta torch.nn.functional.scaled_dot_product_attention, a fastpath for attention extending their accelerated transformer features. This is included in optimum.bettertransformer to be used with the following architectures: Bart, Blenderbot, GPT2, GTP-J, M2M100, Marian, Mbart, OPT, Pegasus, T5.
Beware that this is still experimental and speedups have yet to be validated on all architectures.
PyTorch's scaled_dot_product_attention allows to use flash attention and memory efficient attention natively in PyTorch.
Usage is as follow:
from transformers import AutoTokenizer, AutoModelForCausalLM
from optimum.bettertransformer import BetterTransformer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")
model = BetterTransformer.transform(model) # modify transformers modeling to use native scaled_dot_product_attention
# do you inference or training here
model = BetterTransformer.reverse(model) # go back to using canonical transformers modeling
model.save_pretrained("gpt2_model")
Inference benchmark (on fp16):
| Model | batch size | Input sequence length | Generated tokens | Latency eager (s) | Latency BT (s) | Speedup | Peak memory eager (MB) | Peak memory BT (MB) | Memory savings |
|---|---|---|---|---|---|---|---|---|---|
| gpt2 | 1 | 64 | 256 | 1.800 | 1.607 | 12.0% | 569.90 | 569.89 | 0% |
| gpt2 | 64 | 64 | 256 | 2.159 | 1.617 | 33.5% | 2067.45 | 2093.80 | 0% |
| opt-1.3b | 1 | 64 | 256 | 3.010 | 2.667 | 12.9% | 5408.238 | 5408.238 | 0% |
| gpt-neox-20b | 1 | 64 | 256 | 10.869 | 9.937 | 9.4% | 83670.67 | 83673.53 | 0% |
Training benchmark (on fp16):
| Model | batch size | Sequence length | time/epoch (eager, s) | time/epoch (BT, s) | Speedup | Peak memory eager (MB) | Peak memory BT (MB) | Memory savings |
|---|---|---|---|---|---|---|---|---|
| gpt2 | 8 | 1024 | 17.732 | 14.037 | 26.3% | 13291.16 | 10191.52 | 30.4% |
| gpt2 | 32 | 1024 | 17.336 | 13.309 | 30.3% | 52834.83 | 38858.56 | 36.0% |
| gpt2 | 64 | 1024 | OOM | 14.067 | / | OOM | 75600.08 | / |
Benchmarks can be reproduced using the inference script and training script:
python benchmark_bettertransformer.py --model-name gpt2 --use-half --use-cuda --is_decoder --num-batches 5 --max_token 256
python benchmark_bettertransformer.py --model-name gpt2 --use-half --use-cuda --is_decoder --num-batches 5 --max_token 256 --seqlen-stdev 0
BT] add decoder benchmark script by @younesbelkada in https://github.com/huggingface/optimum/pull/857BT] Fix bt benchmark by @younesbelkada in https://github.com/huggingface/optimum/pull/858BT] Add fp16 support by @younesbelkada in https://github.com/huggingface/optimum/pull/859BT] Add decoder training support by @younesbelkada in https://github.com/huggingface/optimum/pull/860BT] add accelerate_test markers by @younesbelkada in https://github.com/huggingface/optimum/pull/864Three additional architectures are supported in the ONNX export: ImageGPT, RegNet, OPT.
Continued progress in the TFLite export with quantization support. This is work in progress and not documented yet.
TasksManager by @michaelbenayoun in https://github.com/huggingface/optimum/pull/898Full Changelog: https://github.com/huggingface/optimum/compare/v1.2.0...v1.7.2
Nothing published for this version
Temporarily fix a critical bug in BetterTransformer https://github.com/huggingface/optimum/pull/849
Temporarily fix a critical bug in BetterTransformer https://github.com/huggingface/optimum/pull/849
Full Changelog: https://github.com/huggingface/optimum/compare/v1.7.0...v1.7.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →